Non-intrusive Load Monitoring Method Based on Hybrid Federated Learning and Cross Transformer

By introducing hybrid federated learning and cross-transformer models into NILM technology, dynamically switch training modes, and combining graph attention networks and voting mechanisms, the problems of insufficient accuracy, insufficient data privacy protection and scalability in the existing technology are solved, and more efficient, secure and flexible load monitoring is achieved.

CN119962582BActive Publication Date: 2025-06-10JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510436461.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-10
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing NILM technology is insufficient in handling diversified load behaviors, and centralized data processing methods have data privacy protection and security risks, and federated learning is insufficient in scalability and privacy protection in NILM.

Method used

A non-invasive load monitoring method based on hybrid federated learning and cross-transformer is adopted to improve the adaptability and privacy protection capabilities of the model by preprocessing power signal data, dynamically switching centralized and decentralized federated learning modes, and combining graph attention networks and voting mechanisms.

Benefits of technology

It improves the accuracy and adaptability of load monitoring, ensures user data privacy, enhances the scalability and attack resistance of the system, and reduces communication overhead and resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962582B_ABST
    Figure CN119962582B_ABST
Patent Text Reader

Abstract

A non-intrusive load monitoring method based on hybrid federated learning and cross Transformer belongs to the technical field of power system management and energy efficiency, and solves the problems of insufficient accuracy, scalability, and privacy protection in the prior art. The method of the present invention includes: a hybrid federated learning framework that combines centralized federated learning and decentralized federated learning; constructing a communication topology for decentralized federated learning through a graph neural network, and using the graph neural network to encode deep graph structure information into the embedding representation of the client; enhancing the privacy protection of federated learning through layer sensitivity pruning and a robust aggregation method; processing the input bus power sequence through a one-dimensional convolutional layer and a shared Transformer encoder, and adopting multiple load branch structures to achieve the prediction of multiple load working modes; in the decentralized federated learning training, a voting mechanism is adopted to judge whether the model converges. The present invention is applicable to monitoring the energy consumption of various load devices in the monitoring environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of power system management and energy efficiency, and particularly to non-intrusive load monitoring in home and industrial environments. Background Art

[0002] With the continuous growth of global energy demand, especially in the field of building energy management, how to effectively reduce energy consumption and optimize energy use has become a hot issue in academia and industry. Non-Intrusive Load Monitoring (NILM) technology, by analyzing bus data to estimate the energy consumption of each load, provides a low-cost and convenient solution. NILM technology enables users and energy managers to achieve more refined energy consumption monitoring and management, reduce energy waste, improve the level of energy automation management, and enhance energy use efficiency. Although NILM technology has broad application prospects, in practical applications, it still faces many challenges in terms of accuracy, scalability, and privacy protection. Early NILM methods mainly relied on hardware devices to collect high-frequency voltage and current, and the high cost of data acquisition limited their large-scale application. With the rapid development of machine learning and deep learning, NILM methods based on technologies such as Recurrent Neural Networks (RNNs), Convolutional Neural Networks (CNNs), Generative Adversarial Networks (GANs), and attention mechanisms have achieved certain success. However, when dealing with diverse load behaviors in home and commercial environments, these methods often struggle to maintain high accuracy. There are still deficiencies in the accuracy and generalization ability of the models, and a more effective load prediction and monitoring method is needed.

[0003] In addition, existing NILM models usually rely on a central server to centrally process a large amount of raw power data, which brings huge data privacy and security risks. The traditional centralized NILM data processing method can no longer meet the strict privacy protection requirements. Therefore, a method that can effectively monitor loads while ensuring user data privacy must be found.

[0004] As an emerging technical paradigm, federated learning allows models to be locally trained on multiple clients, thereby avoiding data centralization and safeguarding user privacy. However, in the practical deployment of NILM, federated learning still faces problems such as insufficient personalization performance, large communication overhead, and complex privacy protection mechanisms. Existing privacy protection technologies such as homomorphic encryption and secure multi-party computation can theoretically solve privacy problems, but they have high requirements for system resources and are not conducive to practical deployment; while the noise addition of differential privacy will affect the accuracy and convergence speed of the model.

[0005] In summary, the existing technologies have the following problems:

[0006] Insufficient accuracy of the NILM model: Existing NILM technologies are difficult to handle diverse load working characteristics and require a method that can improve accuracy and adaptability.

[0007] Data privacy protection and security risks: The centralized data processing method is prone to exposing user data privacy and has high security risks. There is an urgent need for a NILM technology solution to protect user data privacy.

[0008] Insufficient scalability and privacy protection of federated learning: The practicality of traditional federated learning methods in NILM tasks is limited, facing problems such as large communication overhead and difficulty in balancing privacy protection. It is necessary to improve the scalability, communication efficiency, and anti-attack ability of federated learning in NILM on the basis of ensuring privacy protection. Summary of the Invention

[0009] The purpose of the present invention is to solve the above technical problems and provide a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer.

[0010] The present invention is realized through the following technical solutions. On the one hand, the present invention provides a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer, and the method includes:

[0011] Step 1: Preprocess the power signal data;

[0012] Step 2: Determine the federated learning training mode according to the monitoring network bandwidth, latency, and data privacy requirements. The federated learning training mode includes centralized federated learning and decentralized federated learning; if centralized federated learning is determined, execute Step 3, otherwise execute Step 4;

[0013] Step 3: Centralized federated learning training. Deploy the CrossTransNILM model on the client side and the central server, and aggregate all the updated client models participating in the training on the central server; the CrossTransNILM model includes multiple layers of Transformer encoders and a load branch module, and predicts the power consumption of multiple loads through the bus power.

[0014] Step 4: Decentralized federated learning training. Each client independently trains the CrossTransNILM model and exchanges model parameters through peer-to-peer communication, optimizes the communication topology through a graph attention network, and judges the convergence of the global model through a voting mechanism.

[0015] Step 5: Non-intrusive load monitoring and real-time feedback. Deploy the trained global CrossTransNILM model to the monitoring terminal, collect the power signals of the main line in real time, and input them into the model for end-to-end monitoring after preprocessing; the model outputs the power decomposition results of each electrical device through the load branch module, and maps the load power to the start-stop states of specific electrical devices; the monitoring terminal regularly triggers the model parameter update process according to the federated learning mode selection mechanism, and maintains the model's adaptability to new electrical appliances and electricity consumption behaviors through federated training.

[0016] Further, in Step 3, the CrossTransNILM model specifically includes: a one-dimensional convolutional layer, positional encoding, a shared Transformer encoder, and multiple load branches, and each load branch processes one load type; each load branch contains a branch Transformer encoder layer, a cross Transformer layer, and two fully connected layers.

[0017] The input bus sequence of the model is where is the sequence length and is an odd number, used to locate the center point; the learning objective of the model is to predict the power of each load at the midpoint moment of the sequence, and the learning objective is expressed as:

[0018] ,

[0019] where is the hypothesis space, represents the total number of sampling points, is the mean square error loss function, is the total number of load types;

[0020] The one-dimensional convolutional layer is used to extract the local features of the sequence, and the convolution operation formula is expressed as:

[0021] ,

[0022] Among them, represents a one-dimensional convolution operation, is the input bus power sequence, is the output after convolution, representing the extracted local features;

[0023] Position encoding is used for each position The position encoding in the sequence is generated by sine and cosine functions, and the calculation formula is:

[0024]

[0025] Among them, is the model dimension, represents the sequence position index;

[0026] The embedding layer is used to add the output of the convolutional layer and the position encoding to obtain the embedded representation of the sequence :

[0027] ,

[0028] Among them, is the output after the convolutional layer, is the position information of the sine and cosine encoding;

[0029] The shared Transformer encoder is used to process and analyze the input bus power sequence. Through its multi-layer structure and multi-head self-attention mechanism, it captures the long-range dependencies in the sequence data.

[0030] Furthermore, the shared Transformer encoder consists of multiple encoding layers, and each layer includes a multi-head self-attention mechanism and a feed-forward neural network.

[0031] Furthermore, step 3 includes:

[0032] Step 3.1: The central server first initializes the global model parameters , and then distributes them to all participating clients. Mathematically, it is expressed as:

[0033] , among which, represents the total number of clients;

[0034] Step 3.2: The client uses local data to train and update the model parameters , and the update rule is:

[0035] ,

[0036] Among them, is the model parameter of the client in the round of local training, is the learning rate of model training, is the gradient of the mean square error loss function with respect to the parameter;

[0037] Define the model of each client in the round of local training as , where represents the client in the round of training, the parameter of the layer, is the total number of layers of the CrossTransNILM model. At this time, the mean change amount of the parameters of the layer is , called layer sensitivity, and is defined as:

[0038] ;

[0039] Calculate the sensitivity of each layer and sort it, and select the first layers with the largest sensitivity, where is the transmission ratio coefficient, and the value range is ;

[0040] Step 3.3: The server collects the model parameters uploaded by all clients , where is the round index of global training; for each client , calculate the Euclidean distance sum between its update and the updates of all other clients:

[0041] ,

[0042] where represents the index of other clients participating in training except client , represents the model parameter of client in the round of global training; The smaller the score, the closer is to the updates of other clients, and the more likely it is to be regarded as a normal update; then, sort and remove the highest layer updates, where is the aggregation ratio coefficient, and the value range is ;

[0043] Aggregate the remaining normal updates. For each layer of the CrossTransNILM model , initialize the weight accumulator and the counter . For the weights uploaded by each client , update the accumulator and the counter , and calculate the layer-average update:

[0044] ;

[0045] Add the average update to the corresponding layer of the global model :

[0046] ;

[0047] Step 3.4: The central server distributes the updated global model parameters to all clients to continue the next round of training:

[0048] , .

[0049] Furthermore, Step 4 includes:

[0050] Step 4.1: Use the graph attention network to learn the similarity between clients, and filter the most relevant edges according to the attention weights to obtain the communication topology;

[0051] Step 4.2: The central server generates the initial model parameters , which include all the network layers and parameters to be trained for the NILM task; then, the central server distributes the initialized global model parameters to all clients participating in the training, so that each client has the same starting point before the training begins. Mathematically, it is expressed as follows:

[0052] ,

[0053] where represents the model parameters of client at initialization, and is the total number of clients;

[0054] Step 4.3: Conduct local model training. Each client shares the model parameters with its neighbor clients according to the constructed communication topology ; Each client After receiving the updated parameters from its neighbor clients, aggregate the client model parameters through weighted averaging;

[0055] Step 4.4:

[0056] Evaluate the convergence and record the judgment result as a binary signal:

[0057] ,

[0058] wherein, indicates that the client believes the current model has converged, indicates that the model has not converged; indicates the client on the local validation set loss value; is a preset threshold, is the target accuracy, is the model on the local validation set accuracy;

[0059] Step 4.5: After each round of training is completed, all clients send their respective convergence signals to the central server; the central server determines whether to stop training based on the statistical voting results, and the total number of votes ; set the voting threshold as a static fixed value. When , the system stops training; otherwise, it enters the next round of training.

[0060] Furthermore, Step 4.1 includes:

[0061] Step 4.1.1: Each client first extracts key statistical features according to the local data , and the key statistical features include the mean, standard deviation, minimum value, and maximum value, which are mathematically expressed as:

[0062] ,

[0063] wherein, the features of all clients are expressed as , ;

[0064] Let represent the communication topology graph, the client set , and the edge set represent the potential communication connection relationship between clients. Initialize the edge set as a fully connected structure without self-loops , that is, initially set as a complete directed graph ;

[0065] Step 4.1.2: Use a three-layer graph attention network to construct communication connections between clients. For a client and its neighbor clients , the graph attention network first performs a learnable linear mapping on the client features ; then, calculate the attention coefficient for each edge , and its form is:

[0066] ,

[0067] where is the activation function, is the learnable attention vector, represents vector concatenation, and represent the feature matrices of client and its neighbor client respectively;

[0068] Then, through softmax normalization, the final attention coefficient is obtained:

[0069] ;

[0070] In the forward propagation process, the updated node representation and the set of attention coefficients of all edges are further obtained; represents the connection weight between client and ;

[0071] Call , where is the updated client embedding matrix, is the set of attention coefficients;

[0072] Step 4.1.3: Set the attention weight threshold to be the median of , and retain the topological edges of : , and remove the self-loops ;

[0073] For each directed edge in , only retain the undirected form , and obtain the undirected edge set ;

[0074] For any client , Count its degree , where represents the neighbor clients connected to ; If appears, it means that the client has no connection edges in , then randomly select a neighbor client for and incorporate the new edge into ;

[0075] Let , and finally the undirected graph is the decentralized communication topology learned based on GAT.

[0076] Furthermore, step 2 specifically includes:

[0077] When the network bandwidth , the latency , and the data privacy requirement is low, the system selects centralized federated learning; otherwise, it switches to decentralized federated learning;

[0078] If the conditions for centralized federated learning training are met, the central server uniformly aggregates and optimizes the model updates of all clients;

[0079] If the conditions for decentralized federated learning training are triggered, the clients independently train in a distributed environment and exchange model updates with neighbor clients.

[0080] In a second aspect, the present invention provides a non-intrusive load monitoring system based on hybrid federated learning and cross Transformer, and the system includes:

[0081] A preprocessing module for preprocessing power signal data;

[0082] A federated learning training mode selection module for determining the federated learning training mode according to the monitoring network bandwidth, latency, and data privacy requirements, and the federated learning training mode includes centralized federated learning and decentralized federated learning; if centralized federated learning is determined, execute the centralized federated learning module, otherwise execute the decentralized federated learning module;

[0083] The centralized federated learning module for centralized federated learning training, deploying the CrossTransNILM model on the clients and the central server, aggregating all the model updates of the participating clients to the central server; the CrossTransNILM model includes multiple layers of Transformer encoders and a load branch module, and predicts the power consumption of multiple loads through the bus power;

[0084] A decentralized federated learning module for decentralized federated learning training. Each client independently trains the CrossTransNILM model and exchanges model parameters through peer-to-peer communication, optimizes the communication topology through a graph attention network, and determines the convergence of the global model through a voting mechanism;

[0085] A monitoring module for non-intrusive load monitoring and real-time feedback. The trained CrossTransNILM model is deployed to the monitoring terminal, the power signal of the main line is collected in real time, and after preprocessing, it is input into the model for end-to-end monitoring; the power decomposition results of each electrical device are output through the load branch module, and the load power is mapped to the start-stop state of the electrical device; the monitoring terminal updates the model parameters regularly according to the federated learning mode selection mechanism.

[0086] In a third aspect, the present invention provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor runs the computer program stored in the memory, it executes the steps of a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer as described above.

[0087] In a fourth aspect, the present invention provides a computer-readable storage medium. Multiple computer instructions are stored in the computer-readable storage medium, and the multiple computer instructions are used to cause a computer to execute a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer as described above.

[0088] The beneficial effects of the present invention:

[0089] The present invention provides a non-intrusive load monitoring method based on a hybrid federated learning framework and a cross Transformer structure, aiming to solve the deficiencies in accuracy, scalability, and privacy protection in the prior art. The present invention provides a flexible hybrid federated learning framework that combines the advantages of centralized and decentralized federated learning to adapt to different network conditions and privacy protection requirements, and designs a cross Transformer model to simultaneously predict the power consumption of multiple loads through the bus power. The method of the present invention specifically includes:

[0090] (1) A hybrid federated learning framework that combines centralized federated learning and decentralized federated learning: By dynamically switching between centralized and decentralized federated learning modes to adapt to different network conditions and privacy requirements, the adaptability and robustness of the model are enhanced.

[0091] (2) Constructing the communication topology of decentralized federated learning through graph neural networks: Using graph neural networks to encode deep graph structure information into the embedding representations of clients, capturing complex communication topologies and long-range dependencies between clients, improving communication efficiency and stability, while enhancing the system's resistance to Byzantine attacks.

[0092] (3) NILM privacy protection and attack resistance strategies: Enhancing the privacy protection of federated learning through layer sensitivity pruning and robust aggregation methods. Among them, layer sensitivity pruning reduces communication requirements by pruning key model parameters, enhancing data security; using the Krum algorithm's robust aggregation method to resist Byzantine attacks and ensure the reliability of model updates.

[0093] (4) Design of the CrossTransformer structure (CrossTransNILM): Processing the input bus power sequence through a one-dimensional convolutional layer and a shared Transformer encoder, and adopting multiple load branch structures to achieve the prediction of multiple load working modes. Among them, the input is the bus power sequence, and the output is the power values of multiple loads at the midpoint position of the sequence.

[0094] (5) Voting mechanism consensus for judging the convergence of the NILM model: In decentralized federated learning training, a voting mechanism is adopted to judge whether the model converges, increasing the adaptability of model training.

[0095] The present invention provides a brand-new solution for the NILM field by introducing a hybrid federated learning framework, a CrossTransformer structure, a layer sensitivity pruning strategy, Krum robust aggregation, graph neural networks (GNNs), and a model convergence voting mechanism, with significant technical advantages, especially having important application prospects in smart grids and home energy consumption management. The present invention aims to monitor the energy consumption of various load devices in the monitored environment by using a hybrid federated learning framework and a CrossTransformer model, without the need to connect sensors to each load separately.

[0096] The present invention is applicable to environments that require privacy protection and data security, such as data centers and residential houses, where the confidentiality of user data and the security of the system are crucial. In addition, the application of the present invention can also be extended to the construction of smart cities and smart grids to provide support for urban energy management. Brief Description of the Drawings

[0097] To more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0098] Figure 1 Schematic diagram of the method of the present invention;

[0099] Figure 2 Schematic diagram of the NILM model CrossTransNILM model of the present invention;

[0100] Figure 3 Schematic diagram of the shared Transformer encoder of the present invention;

[0101] Figure 4 Schematic diagram of the cross Transformer encoder layer;

[0102] Figure 5 Loss curves of the CrossTransNILM model at different pruning ratios. Detailed implementation manners

[0103] The following details the implementation manners of the present invention. Examples of the implementation manners are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The implementation manners described below by referring to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0104] Detailed implementation manners: As Figure 1 shown, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is introduced in detail below for the working principle, working process and each step of the non-intrusive load monitoring method.

[0105] Step 1: Data preprocessing

[0106] Data preprocessing is a key step to ensure the effectiveness of subsequent model training, validation and testing. The power signals collected by sensors or electricity meters need to be subjected to a series of processes to improve the quality and consistency of the data. The following are the detailed preprocessing steps:

[0107] Step 1.1, Data cleaning and alignment

[0108] First, visualize the power of the main bus and each load to check the integrity and consistency of the data. By observing the power waveforms, identify and delete abnormal parts, such as missing data with a duration exceeding half an hour. For short-term missing data (from a few seconds to a few minutes), since it has little impact on the overall data distribution, retain it and use it as a noise source for model regularization to enhance the robustness of the model.

[0109] Then, unify the sampling rates of the main bus power and each load power to achieve consistency in time steps. The constructed dataset contains the power of each load and the main bus power. The main bus power Calculated from the sum of the powers of each load plus a random noise term:

[0110] ,

[0111] where, is the total number of load types, represents the load at time power, is the noise term.

[0112] Step 1.2, Data Standardization

[0113] To make the loads and bus power comparable and eliminate the influence of dimensions on model training, data standardization is required. Taking the bus power as an example, the standardized power is calculated by the following formula:

[0114] ,

[0115] where, is the mean value of the power, is the standard deviation of the power.

[0116] Step 2: Federal Learning Training Mode Selection

[0117] Hybrid federated learning combines centralized and decentralized training modes, dynamically selects the optimal mode to adapt to different network conditions and privacy requirements, thereby improving the robustness of the system. The system determines whether to use centralized federated learning (Step 3) or decentralized federated learning (Step 4) for the current CrossTransModel model training by monitoring the network bandwidth , latency and data privacy requirements.

[0118] Step 2.1: Federal Learning Mode Switching Conditions

[0119] Set as the network bandwidth threshold, as the latency threshold. By monitoring , and privacy requirements, the system dynamically selects the training mode:

[0120] .

[0121] Specifically, when the network bandwidth , latency , and when the data privacy requirement is low, the system selects centralized federated learning to ensure high communication efficiency and global synchronization, and accelerate the aggregation and optimization of the global model. Otherwise, the system switches to decentralized federated learning to reduce the network communication burden and latency and minimize the transmission of sensitive data.

[0122] Step 2.2: Processing flow after mode selection

[0123] If the conditions for centralized federated learning training are met, go to Step 3, centralized federated learning training, where the central server uniformly aggregates and optimizes the model updates of all clients to obtain higher model consistency and overall performance.

[0124] If the conditions for decentralized federated learning training are triggered, go to Step 4, decentralized federated learning training, where clients independently train in a distributed environment and exchange model updates with neighboring clients to improve the fault tolerance and privacy protection level of the system.

[0125] Step 3: Centralized federated learning training

[0126] Centralized federated learning aggregates the model updates of all participating clients at the central server to ensure model consistency and efficiency. This process includes four steps: client model initialization, local model training and layer sensitivity pruning and uploading, server model parameter aggregation, and global model distribution:

[0127] Step 3.1: Client model initialization

[0128] Step 3.1.1: CrossTransModel model structure design

[0129] The NILM model is deployed on the client and the central server. The NILM model in this embodiment is a prediction model based on the cross-Transformer structure, denoted as CrossTransNILM. This model combines multiple layers of Transformer encoders and load branch modules to predict the power consumption of multiple loads through the bus power.

[0130] The structure of the CrossTransNILM model is as Figure 2 shown, including: a one-dimensional convolutional layer, position encoding, a shared Transformer encoder, and multiple load branches. Among them, each load branch contains a branch Transformer encoder layer, a cross-Transformer layer, and two fully connected layers. The overall structure is designed to effectively capture the load operation rules and perform sequence-to-point multi-load power prediction for the NILM task.

[0131] Step 3.1.1.1: Model input

[0132] In this embodiment, a sequence-to-point learning strategy is adopted to improve the prediction accuracy and reduce the computational complexity. First, the input bus sequence is defined as , where is the sequence length. To facilitate the positioning of the center point, is selected as an odd number. The learning objective of the model is to predict the power of each load at the midpoint of the sequence. Therefore, the learning objective can be expressed as:

[0133] ,

[0134] where is the hypothesis space, represents the total number of sampling points, is the mean squared error (MSE) loss function, is the total number of load types.

[0135] Step 3.1.1.2, One-dimensional convolutional layer

[0136] The one-dimensional convolutional layer is used to extract the local features of the sequence to reduce the data dimension and retain the significant feature patterns in the sequence. The convolution operation formula is expressed as follows:

[0137] ,

[0138] where represents the one-dimensional convolution operation, is the input bus power sequence, is the output after convolution, representing the extracted local features.

[0139] Step 3.1.1.3, Positional encoding

[0140] To enable the NILM model to recognize the positional relationship in the sequence, the positional encoding of each position in the sequence is generated by sine and cosine functions, and the calculation formula is as follows:

[0141] ,

[0142] ,

[0143] where is the model dimension, represents the sequence position index, and the model's perception ability of time series is enhanced through the sine and cosine positional encoding.

[0144] Step 3.1.1.4, Embedding Layer

[0145] Add the output of the convolutional layer to the positional encoding to obtain the embedded representation of the sequence :

[0146] ,

[0147] wherein, is the output after passing through the convolutional layer, is the positional information of the sine-cosine encoding. This representation method combines positional encoding and local feature information, enhancing the representation ability of the model.

[0148] Step 3.1.1.5, Shared Transformer Encoder

[0149] The shared Transformer encoder adopted in this embodiment is designed to process and analyze the input bus power sequence. Through its multi-layer structure and multi-head self-attention mechanism, it can capture the long-range dependencies in the sequence data and achieve a global understanding of the entire sequence. The schematic diagram of the shared Transformer encoder in this embodiment is as shown in Figure 3 shown. This encoder consists of multiple encoding layers, and each layer includes a multi-head self-attention mechanism and a feed-forward neural network.

[0150] (1) Multi-Head Attention Mechanism

[0151] The multi-head attention mechanism allows the model to learn the internal structure of the input data in parallel in different representation spaces. Each head performs different linear transformations on the data and independently applies the attention mechanism. In each encoder layer, the sequence after positional encoding is mapped into a query matrix ( ), a key matrix ( ) and a value matrix ( ):

[0152] ,

[0153] ,

[0154] ,

[0155] wherein, , , are the parameter matrices learned during the training process.

[0156] The attention mechanism can help the model extract important information from the input. The calculation method is as follows:

[0157] ,

[0158] Among them, is the dimension of the key matrix which adjusts the computational stability through the scaling factor By executing in parallel with multiple heads, the attention output corresponding to each head is calculated as follows:

[0159] ,

[0160] Among them, , , are independent transformation matrices corresponding to each head, is the number of heads in each encoder layer. Then, by concatenating the results of all heads and performing a linear transformation through the output weight matrix :

[0161] ,

[0162] This design enables the model to integrate multiple feature expressions and improve the information integration ability.

[0163] By introducing the multi-head self-attention mechanism, the shared Transformer encoder of this embodiment can not only effectively process and analyze data, but also enhance the expressiveness and prediction accuracy of the model by processing various features in parallel.

[0164] (2) Residual connection and layer normalization

[0165] To improve the efficiency and effect of training deep networks, a residual connection is added after each attention output:

[0166] ,

[0167] The residual connection helps the gradient flow directly to deeper network layers, preventing the problem of gradient disappearance during training.

[0168] Then, layer normalization is applied after each residual connection, and the normalization process improves the stability of model training:

[0169] ,

[0170] Among them, is the mean obtained with the columns of the matrix as the unit, is the dimension of the rows of the matrix, represents the -th element in the -th column of the matrix. is the variance obtained with the columns of the matrix is a very small number to prevent the denominator from being zero during the calculation process.

[0171] (3) Feedforward neural network, residual connection, and layer normalization

[0172] Construct a basic neural network structure that includes two linear layers and a ReLU (Rectified Linear Unit) activation function. Through the combination of linear layers and the ReLU activation function, the input is subjected to feature transformation and non-linear mapping to obtain the output of the feedforward neural network :

[0173] ,

[0174] Among them, the ReLU activation function , helps to introduce non-linear characteristics into the network, alleviate the vanishing gradient problem, and accelerate the convergence speed of the network.

[0175] Then, perform residual connection and layer normalization in the same way as in (2):

[0176] ,

[0177] to obtain the output of the shared Transformer .

[0178] Step 3.1.1.6, Multi-load branch

[0179] To predict the power of different loads simultaneously, the CrossTransNILM model of this embodiment designs multiple load branches, and each branch processes one load type. This design allows the model to handle the unique characteristics of each load more carefully and improve the prediction efficiency. Each load branch internally contains an independent Transformer encoder layer, a cross-attention layer, and two fully connected layers.

[0180] (I) Transformer encoder layer

[0181] Each load branch internally contains an independent Transformer encoder layer, which has a similar structure to the shared Transformer encoder, including multi-head self-attention and a feedforward network. Feature extraction and learning are performed for each load, enabling a more in-depth analysis and understanding of the characteristics of each load.

[0182] (II) Cross Transformer encoder layer

[0183] The cross Transformer encoder layer utilizes the cross-attention mechanism, enabling each load branch to refer to the global information from the shared Transformer encoder during power prediction. The design of the cross-attention mechanism in this embodiment is based on the view that although each type of load has its unique load characteristics, there are often some connections among them in practical applications. Through cross-attention, the model can take into account the interactions and influences with other loads when predicting the power of one load. The schematic diagram of the cross Transformer encoder layer is as shown in Figure 4 shown below.

[0184] Specifically, the query matrix of the cross-attention is generated by the output of the Transformer encoder layer of the current load branch, while the key matrix and the value matrix come from the output of the shared Transformer encoder, which enables each load branch to consider the interactions and influences with other loads during power prediction.

[0185] (III) Fully Connected Layer and Output

[0186] The output of each load branch undergoes final feature transformation and non-linear mapping through the combination of two linear layers and a (Gaussian Error Liner Unit) activation function. Among them, the activation function , where \(x\) is the input, and is the standard normal cumulative distribution function of \(x\), and \(erf(x)\) is a non-linear error function used to smoothly transition the input value. The Gaussian Error Liner Unit provides a smoother non-linear transformation, so it can more finely adjust the transmission of the input signal in the output of the activation function, especially in the weak change region of the signal. This helps to capture more subtle pattern changes and improve the sensitivity and accuracy of the model in dealing with data details.

[0187] Finally, the predicted power value of load \(l\) at the midpoint position in the input sequence is obtained: :

[0188] ,

[0189] where is the output of the cross Transformer encoder layer.

[0190] Step 3.1.2, Model Initialization and Synchronization

[0191] Before the start of training, the central server first initializes the global model parameters , and then distributes them to all participating clients, so as to ensure that each client starts training from the same initial state of the NILM model. Mathematically, it is expressed as:

[0192] ,

[0193] where represents the total number of clients. This ensures that in the 0th round of training, the initial states of all client models are the same.

[0194] Step 3.2, Local model training, layer sensitivity pruning and uploading

[0195] In each round of centralized federated learning global training, each client uses local data for independent model training. The specific steps are as follows:

[0196] First, the client uses local data for training and updates the model parameters . The update rule is:

[0197] ,

[0198] where is the model parameter of client in the th round of local training, is the learning rate of model training, is the mean square error loss function for the gradient of the parameter.

[0199] Then, due to the problem of redundant weight parameters in the deep neural network model, this embodiment proposes to perform layer sensitivity pruning when uploading parameters, and only upload the parameters of some sensitive layers to reduce the amount of transmitted data, improve communication efficiency and enhance privacy protection. Define the model of each client in the th round of local training as , where represents the parameter of client in the th round of training at the th layer, is the total number of layers of the CrossTransNILM model. At this time, the mean change of the th layer parameter , called layer sensitivity, is defined as:

[0200] ,

[0201] The greater the layer sensitivity, the more important the change in this layer may be for improving the model performance.

[0202] Finally, calculate the sensitivity of each layer and sort them, and select the top layers with the greatest sensitivity, where is the transmission ratio coefficient, and the value range is . Through this pruning strategy, the client only needs to upload the parameters of the partially sensitive layers, reducing the communication overhead, protecting sensitive information, and enhancing the model security.

[0203] Step 3.3. Server Model Parameter Aggregation

[0204] To address malicious attacks and abnormal parameter updates in centralized federated learning, the Krum aggregation algorithm is adopted on the server side in this embodiment.

[0205] First, the server collects the model parameters uploaded by all clients , where is the index of the global training round. For each client , calculate the sum of the Euclidean distances between its update and the updates of all other clients :

[0206] ,

[0207] where, represents the index of other clients participating in the training except client , represents the model parameters of client in the round of global training. The smaller the score, the closer is to the updates of other clients and the more likely it is to be regarded as a normal update. Then, sort and remove the top layer updates, where is the aggregation ratio coefficient, and the value range is .

[0208] Finally, aggregate the remaining normal updates. Since each client has performed layer sensitivity pruning, the central server only receives the weights of partial layers of each client. For each layer of the CrossTransNILM model, initialize the weight accumulator and the counter . For the weights uploaded by each client, update the accumulator and the counter ​。Calculate the average update of the layer:

[0209] 。

[0210] Add the average update to the corresponding layer of the global model :

[0211] 。

[0212] Step 3.4, Global model distribution

[0213] After aggregating the model parameters, the central server distributes the updated global model parameters to all clients to continue the next round of training:

[0214] , ,

[0215] Ensure that the client model parameters in centralized federated learning are effectively optimized and updated.

[0216] Step 4: Decentralized federated learning training

[0217] When the model training switches to the decentralized federated learning mode, each client independently trains the CrossTransNILM model and exchanges model parameters through peer-to-peer communication to enhance privacy protection. The communication topology is optimized through the Graph Attention Network (GAT), and the convergence of the global model is judged through a voting mechanism. The decentralized structure improves the system's ability to resist Byzantine attacks and the scalability of communication, and reduces the dependence on the single point of failure of the central server. The specific steps are as follows:

[0218] Step 4.1, Communication topology construction

[0219] In the decentralized federated learning setting, the communication topology determines the direction of model parameter flow during the model training process. To optimize the communication efficiency in decentralized federated learning and enhance the robustness of the network, this embodiment uses GAT to learn the similarity between clients and filters the most relevant edges according to the attention weights to obtain the communication topology structure. The system dynamically adjusts the communication structure after a certain number of training rounds to cope with changes in network conditions and communication link failures. The following are the detailed steps for constructing the communication topology:

[0220] Step 4.1.1, Initialize node features and communication topology graph

[0221] Each client firstly, according to the local data , extract key statistical features. These features include the mean, standard deviation, minimum, and maximum, thus forming a feature vector to describe the data characteristics of each client. The mathematical representation is as follows:

[0222] ,

[0223] Among them, the features of all clients are represented as , .

[0224] Let represent the communication topology graph, the client set , and the edge set represent the potential communication connection relationship between clients. To form an effective learning and information exchange topology network and avoid isolated clients, in this embodiment, the edge set is initialized as a fully connected structure , without self-loops , that is, initially set as a complete directed graph .

[0225] Step 4.1.2, Construction of Communication Topology Based on GAT Model

[0226] Use a three-layer GAT to construct the communication connections between clients. For client and its neighbor client , GAT first performs a learnable linear mapping on the client features . Then, for each edge , calculate the attention coefficient , and its general form is:

[0227] ,

[0228] Among them, is the activation function, is the learnable attention vector, represents vector concatenation, and respectively represent the feature matrices of client and its neighbor client . Then, after softmax normalization, the final attention coefficient is obtained:

[0229] .

[0230] In the forward propagation process, further obtain the updated node representation and the set of attention coefficients of all edges . In the decentralized federated learning in this paper is represented as the client The connection weight with .

[0231] Call , where is the updated client embedding matrix, is the set of attention coefficients.

[0232] Step 4.1.3, Communication topology edge screening

[0233] First, set the attention weight threshold to be the median of , and retain the topological edges of , and remove self-loops .

[0234] Then, for each directed edge in , only retain the undirected form , and obtain the undirected edge set .

[0235] Next, for any client , count its degree , where represents the neighbor clients connected to . If appears, it means that this client has no connection edges in , then randomly select a neighbor client for and incorporate the new edge into to avoid its isolation.

[0236] Finally, let . The final undirected graph is the decentralized communication topology learned based on GAT.

[0237] To adapt to changing network conditions and client privacy protection requirements, the system will periodically re-evaluate and optimize the communication topology. After every rounds of training, the system will re-use GAT to calculate client embeddings and re-construct the communication topology based on these new embeddings to ensure efficient information exchange paths and improve training efficiency and overall performance.

[0238] Step 4.2, Client initialization

[0239] Before the start of training, it is first necessary to ensure that all clients start training from the same model state to maintain consistency during the training process. Specifically:[[]]

[0240] First, the central server generates the initial model parameters , the model parameters include all the network layers and parameters to be trained and are used for the NILM task. Then, the central server distributes the initialized global model parameters to all the participating clients, so that each client has the same starting point before the training begins. The mathematical representation is as follows:

[0241] ,

[0242] where, represents the model parameters of client at initialization, and is the total number of clients.

[0243] Step 4.3, Client Local Model Training and Update

[0244] In the decentralized federated learning mode, each client follows a process similar to the centralized mode and first conducts local model training. After each round of local training is completed, each client shares its model parameters with its neighbor clients according to the constructed communication topology to reduce the amount of data transmission and enhance data privacy protection. Then, each client aggregates the model parameters of each client through weighted averaging after receiving the updated parameters from its neighbor clients.

[0245] Step 4.4, Convergence Evaluation and Voting Mechanism

[0246] In decentralized federated learning, the design of the convergence evaluation and voting mechanism is to ensure that the training stops when sufficient model performance is achieved, so as to save computing resources and prevent overfitting.

[0247] Step 4.4.1, Convergence Evaluation

[0248] The convergence evaluation in decentralized federated learning includes the following two aspects: on the one hand, after each client completes local training, the model is obtained. When the loss change on the local validation set is less than the preset threshold (such as ), it indicates that the model has tended to be stable in multiple rounds of training. On the other hand, when the accuracy of the model on the local validation set reaches or exceeds the target accuracy , the model is considered to have achieved the expected performance and can be regarded as converged.

[0249] Each client evaluates based on the above criteria and records the judgment result as a binary signal:

[0250] ,

[0251] Among them, indicates that the client believes that the current model has converged, indicating that the model has not converged yet. Indicates the client on the local validation set of the loss value.

[0252] To evaluate , this embodiment selects the Mean Absolute Error (MAE), Normalized Signal Aggregate Error (SAE), and Energy per Day (EpD) to evaluate the algorithm:

[0253] 1) MAE: Evaluate the load at each time point of the power prediction value and the actual value the mean of the absolute error between them. The smaller it is, the more accurate the prediction result of the model. Among them, represents the total number of test samples.

[0254] .

[0255] 2) SAE: Represents the relative error of the total energy. The smaller the SAE, the smaller the prediction error of the model and the better the model performance.

[0256] ,

[0257] Among them, represents the total energy consumption of the load , represents the predicted total energy consumption of the load , is the sampling interval.

[0258] 3) EpD: Represents the absolute error of the predicted energy within a day. The smaller the EpD, the smaller the absolute error of the predicted energy of the model within a day and the better the model performance.

[0259] ,

[0260] Among them, , represents the sample the total number of days included. , represents the number of sampling points included in a day. , represents the load The sum of the actual energy consumed within a one-day period, , representing the load The total predicted energy within a one-day period.

[0261] Step 4.4.2, Voting mechanism

[0262] After each round of training is completed, all clients send their respective convergence signals to the central server. The central server determines whether to stop training based on the statistical voting results. The total number of votes . The system sets a voting threshold as a static fixed value. When , the system stops training; otherwise, it enters the next round of training. In summary, the training process will terminate when any of the following conditions is met:

[0263] 1) Reaching the predetermined maximum number of rounds .

[0264] 2) Exceeding the voting threshold (for example ).

[0265] 3) The global model performance reaches the target (such as when the loss change on the global validation set is less than the preset threshold ).

[0266] Through the convergence evaluation and voting mechanism, this embodiment effectively avoids overfitting and ensures that the training process stops after reaching the established goal. At the same time, the voting mechanism also improves the robustness of the system, avoiding model deviations caused by improper training of a few clients or communication problems, thereby improving the overall performance and stability of the system.

[0267] Conventionally, model convergence is based on the change in loss or reaching the specified number of rounds. However, in decentralized federated learning, due to aggregating models from multiple clients, the loss may have relatively large fluctuations. Using only the loss function or the specified number of training rounds may make it difficult to judge convergence or increase the number of training rounds for convergence. Therefore, a convergence judgment and voting mechanism are added here.

[0268] To verify the effectiveness of the proposed hybrid federated learning framework and CrossTransNILM model of the present invention, the present invention designs and conducts a number of experiments to evaluate the performance of the present invention under different conditions. The present invention selects two publicly available household electricity datasets, namely REFIT and UK-DALE:

[0269] 1) REFIT: The REFIT data was collected from 20 buildings in a certain country, with a time span from 2013 to 2015. The sampling period is once every 8 seconds, and it contains power data of the main bus and each load.

[0270] 2) UK - DALE: The UK - DALE dataset contains the power data of 5 households in a certain country from 2013 to 2015. The sampling period of the total power supply is once every 1 second, and the sampling period of each load is once every 6 seconds.

[0271] In the simulation design, five typical household loads (microwave oven, refrigerator, dishwasher, washing machine, and electric kettle) were selected to verify the performance of the NILM model of the present invention. In the present invention, the data format is defined as: (washing machine power, dishwasher power, microwave oven power, refrigerator power, electric kettle power, bus power), where the bus power is the sum of the powers of the above five loads.

[0272] Before applying the NILM algorithm, the dataset needs to be uniformly sampled to achieve the alignment of time steps. In the present invention, 1 / 8 Hz is selected as the standard sampling rate. In addition, the two datasets are also divided into federated learning clients, and the specific division method is shown in Table 1.

[0273] Table 1 Client data division

[0274]

[0275] Finally, REFIT and UK - DALE are standardized. In the simulation of the present invention, the mean and standard deviation values for standardization are shown in Table 2.

[0276] Table 2 Data standardization parameters

[0277]

[0278] It should be noted that the parameters in Table 2 are only used for standardizing the model training, validation, and test data, and not for determining the actual mean and variance of these loads. The following are the specific experimental results and the description of the invention effects.

[0279] I. Performance comparison between the CrossTransNILM model of the present invention and other NILM models

[0280] To evaluate the effectiveness of the proposed CrossTransNILM model of the present invention, two currently state - of - the - art NILM models are selected for comparative experiments using local training: the sequence - to - point model based on convolutional neural network (sequence - to - point S2P) and the BERT4NILM model that first applies Transformer. The specific model descriptions are as follows:

[0281] Sequence-to-point model (sequence-to-point S2P) based on convolutional neural network: This model contains multiple one-dimensional convolutional layers and ReLU activation functions, as well as two fully connected layers, which are used to process the input active power sequence. The features of the input bus power sequence are extracted through the convolutional layer, and the power estimates of the five loads corresponding to the center position of the sequence are output through the fully connected layer. Transformer-based BERT4NILM model: BERT4NILM combines the encoder-decoder architecture and the attention mechanism, and uses the advantages of Transformer in processing time series data to achieve the key capture of sequence information through the attention mechanism, especially when processing power data with complex patterns. Compared with the model without attention mechanism, it has more advantages. In order to ensure fairness, the hyperparameter settings of the above two comparison models are consistent with the CrossTransNILM model of the present invention, and all three models are trained until convergence. The general hyperparameter settings of the algorithm are shown in Table 3, and the data partitioning used is shown in Table 4.

[0282] Table 3 General hyperparameter settings of the algorithm

[0283]

[0284] Table 4. Local model training data division

[0285]

[0286] Table 5 is the comparison results of the three groups of experiments. It can be seen from Table 5 that the CrossTransNILM model shows significant advantages in most load types and three indicators. Especially in the loads of washing machines, microwave ovens and electric kettles, the performance of CrossTransNILM is significantly better than the other two models. In addition, CrossTransNILM also showed the best results in MAE of dishwashers and SAE and EpD of refrigerators. This shows that CrossTransNILM can more accurately capture the energy consumption characteristics of different loads and effectively balance the prediction accuracy and energy error. It is suitable for the sequence-to-point multi-load prediction scenario of NILM, which reflects the advanced nature of the CrossTransNILM model of the present invention.

[0287] Table 5 Performance comparison of CrossTransNILM local training and other models

[0288]

[0289] 2. Pruning ratio experiment

[0290] To test the impact of different pruning ratios on the model performance, we set four pruning ratios: 0%, 5%, 10%, and 20%. The experimental results are shown in Table 6. Compare the results of each evaluation metric under different pruning ratios for different load types.

[0291] Table 6 Performance of CrossTransNILM Model under Different Pruning Ratios

[0292]

[0293] From the results in Table 6, it can be seen that under the pruning ratio of 5% - 10%, the performance of the CrossTransNILM model on most load types is comparable to that of the unpruned model, and it can reduce the communication volume while maintaining a high prediction accuracy. Especially under the pruning ratio of 5%, in the prediction of the electric kettle, its MAE, SAE and other metrics are even better than the unpruned results. This may be because pruning introduces a certain regularization effect, which helps to prevent overfitting. However, when the pruning ratio reaches 20%, the model performance drops significantly, indicating that the effect of pruning does not increase infinitely, but will have a negative impact on the model's prediction ability to a certain extent. Therefore, in the selection of pruning, it is necessary to balance the model performance and communication efficiency and reasonably set the pruning ratio.

[0294] To evaluate the impact of layer sensitivity pruning on model training, the Loss curves of the CrossTransNILM model in centralized federated learning training under different layer sensitivity pruning ratios are compared, as shown in Figure 5.

[0295] From Figure 5 the experimental results, it can be seen that the performance of the unpruned model (0% pruning ratio) is the best, and the loss value drops rapidly and remains at the lowest level. In contrast, the model with a 5% pruning ratio can also converge well, but there are some fluctuations during the training process, and the final loss is slightly higher than that of the unpruned model. As the pruning ratio increases, the model with a 10% pruning ratio can converge, but the loss value is slightly higher and the fitting ability is weakened. The model with a 20% pruning ratio has large fluctuations during the convergence process, and the final loss is the highest, significantly affecting the learning ability of the model.

[0296] Figure 5 The results further show that the choice of pruning ratio has a direct impact on model performance. Moderate pruning (such as a 5% pruning ratio) can reduce the computational complexity while retaining the model performance, while excessive pruning (such as a 20% pruning ratio) may weaken the fitting ability of the model. Therefore, through moderate pruning, the present invention can maintain a high model accuracy and a good convergence speed while reducing the communication overhead, which shows obvious advantages under limited bandwidth conditions.

[0297] III. Federal Learning Performance Comparison Experiments

[0298] To evaluate the effectiveness of the federated learning of various models of the present invention, the performance of local training, centralized federated learning, and decentralized federated learning was compared. All models were uniformly tested using the data of the client numbered 1. The experimental results are shown in Table 7.

[0299] Table 7 Performance Comparison of CrossTransNILM Local Training and Other Models

[0300]

[0301] By comparing the performance of local training, 5% pruning centralized federated learning, and decentralized federated learning in Table 7 under five load types: washing machine, dishwasher, microwave oven, refrigerator, and electric kettle, the results show that local training outperforms the federated learning methods in all performance metrics. The centralized federated learning performs better than the decentralized federated learning in all load types, which may be because the centralized method can better coordinate the information of each client and reduce model inconsistency. However, the performance degradation of these federated learning methods is not significant, indicating that after introducing federated learning, the model can still maintain a high prediction accuracy. This verifies the effectiveness of the various federated learning models proposed in the present invention in dealing with different load types. Although there is a certain gap compared with local training, its performance is sufficient to support its feasibility and advantages in practical applications.

[0302] IV. Summary

[0303] Through the above experiments, the present invention effectively demonstrates the performance of the hybrid federated learning framework and the CrossTransNILM model in various environments. The results of each experiment verify the significant advantages of the present invention in communication efficiency, model accuracy, and anti-attack ability. The invention can not only achieve efficient load monitoring in scenarios with limited bandwidth and privacy protection, but also show strong adaptability and stability in the face of limited communication conditions and security threats. These characteristics make the present invention applicable to a wide range of application scenarios, including fields such as smart grids and home energy consumption management.

[0304] It should be noted that the above design is only an example of the implementation of the present invention, and can be adjusted and optimized according to specific requirements in actual applications. For example:

[0305] (1) Communication Topology and Federated Learning Mode Selection: When designing a decentralized federated learning system, the present invention adopts a communication topology structure and a hybrid federated learning mode inferred by a graph neural network. However, according to network conditions and data privacy requirements, different communication topologies (such as fully connected or rule-based connections) and federated learning modes (such as centralized, decentralized, or hybrid modes) can be flexibly selected.

[0306] (2) Pruning strategy and pruning ratio: When transmitting model parameters, the present invention adopts layer sensitivity pruning technology to reduce the amount of transmitted data and improve data security. However, the pruning ratio and strategy can be flexibly adjusted according to the performance requirements of the actual application scenario. For applications with high requirements for communication efficiency, a higher pruning ratio can be adopted to significantly reduce communication overhead; while for scenarios pursuing higher model accuracy, the pruning ratio can be appropriately reduced to maintain model performance.

[0307] (3) Flexibility of the cross-Transformer structure: The cross-Transformer structure (CrossTransNILM) of the present invention extracts features from the bus power sequence through a one-dimensional convolutional layer and a shared Transformer encoder, and adopts multiple load branch structures to achieve accurate prediction of complex load working modes. In practical applications, the number of Transformer layers, the size of the convolutional kernel, and the branch structure can be adjusted according to the characteristics of different loads to adapt to different load characteristics and ensure the prediction accuracy and computational efficiency of the model in complex scenarios.

[0308] It should be noted that although federated learning and graph neural network technologies have been successfully applied in many fields, such as privacy computing, recommendation systems, and network analysis, their applications in the NILM field are still in the exploratory stage. Therefore, the proposed hybrid federated learning framework combining centralized and decentralized approaches, which can dynamically switch to cope with different network conditions and privacy requirements, can provide a new flexible solution for the NILM field.

[0309] In addition, the cross-Transformer structure (CrossTransNILM) in the present invention utilizes the powerful feature extraction ability of the Transformer model to efficiently predict the complex patterns of electrical loads. This innovative architecture design, by combining client embedding optimization of graph neural networks and cross-attention mechanisms, can not only effectively improve the load prediction accuracy, but also better capture the potential connections between loads, providing a new approach for deep learning modeling of NILM tasks.

[0310] Generally speaking, the present invention, based on the innovative combination of a hybrid federated learning framework, a cross-Transformer structure, graph neural networks, and privacy protection technologies, is expected to bring new technological breakthroughs and application prospects to non-intrusive load monitoring in the fields of smart grids and home energy consumption management.

Claims

1. A non-intrusive load monitoring method based on hybrid federated learning and cross-Transformer, characterized in that: The method comprises: Step 1: Preprocess the power signal data; Step 2: Determine the federated learning training mode based on the monitoring network bandwidth, latency, and data privacy requirements. The federated learning training mode includes centralized federated learning and decentralized federated learning. If it is determined to be centralized federated learning, execute step 3, otherwise execute step 4. Step 3: Centralized federated learning training, deploy the CrossTransNILM model on the client and the central server, and centralize all client model updates participating in the training to the central server for aggregation; the CrossTransNILM model includes a multi-layer Transformer encoder and a load branch module, and predicts the power consumption of multiple loads through bus power; the CrossTransNILM model specifically includes: a one-dimensional convolutional layer, a positional encoding, a shared Transformer encoder, and multiple load branches, each load branch processes a load type; each load branch contains a branch Transformer encoder layer, a cross Transformer layer, and two fully connected layers; The shared Transformer encoder consists of multiple encoding layers, each of which includes a multi-head self-attention mechanism and a feed-forward neural network; Step 4: Decentralized federated learning training, each client independently trains the CrossTransNILM model and exchanges model parameters through point-to-point communication, optimizes the communication topology through the graph attention network, and judges the convergence of the global model through the voting mechanism; Step 5: Non-intrusive load monitoring and real-time feedback. The trained CrossTransNILM model is deployed to the monitoring terminal to collect the power signal of the main line in real time. After preprocessing, it is input into the model for end-to-end monitoring. The power decomposition results of each electrical equipment are output through the load branch module, and the load power is mapped to the start and stop status of the electrical equipment. The monitoring terminal regularly updates the model parameters according to the federated learning mode selection mechanism.

2. According to claim 1, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: In step 3, the input bus sequence of the CrossTransNILM model is ,in is the length of the sequence, and is an odd number, used to locate the center point; the learning goal of the model is to predict the midpoint of the sequence Each load Power , the learning objective is expressed as: , in, is the hypothesis space, represents the total number of sampling points, is the mean square error loss function, is the total number of load types; The one-dimensional convolution layer is used to extract the local features of the sequence. The convolution operation formula is expressed as: , in, represents a one-dimensional convolution operation, is the input bus power sequence, is the output after convolution, representing the extracted local features; Positional encoding is used for each position Encoding the position in the sequence Generated by sine and cosine functions, the calculation formula is: in, is the model dimension, Represents the sequence position index; The embedding layer is used to add the convolutional layer output to the positional encoding to obtain the embedded representation of the sequence. : , in, is the output after the convolutional layer. is the position information encoded by sine and cosine; The shared Transformer encoder is used to process and analyze the input bus power sequence. Through its multi-layer structure and multi-head self-attention mechanism, it captures long-distance dependencies in the sequence data.

3. According to claim 2, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: Step 3 includes: Step 3.1: The central server first initializes the global model parameters , and then send it to all clients participating in the training. The mathematical expression is: ,in, Indicates the total number of clients; Step 3.2: Client Using local data Perform training and update model parameters , the update rule is: , in, Is the client In the Model parameters for round 1 local training, is the learning rate for model training, is the mean square error loss function Gradients with respect to parameters; Define each client In the The model trained locally is ,in, Represents the client In the In the training round The parameters of the layer, is the total number of layers of the CrossTransNILM model. Layer parameter mean change , called layer sensitivity, is defined as: ; Calculate the sensitivity of each layer and sort it, and select the one with the highest sensitivity. Layer, where is the transmission ratio coefficient, the value range is ; Step 3.3: The server collects all model parameters uploaded by the client ,in is the index of the round number of global training; for each client , calculate the Euclidean distance between it and all other client updates and : , in, Indicates that except client The index of other clients participating in the training. Indicates Client in round global training Model parameters of The smaller the score, The closer the update is to other clients; then, Sort by removing the highest Layer update, where is the aggregation ratio coefficient, and its value range is ; Aggregate the remaining normal updates for each layer of the CrossTransNILM model , initialize the weight accumulator and counter , for each client upload weight , update the accumulator and counter , the average update of the calculation layer: ; Add the average update to the corresponding layer of the global model superior: ; Step 3.4: The central server updates the global model parameters Distribute to all clients to continue the next round of training: , 。 4. According to claim 1, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: Step 4 includes: Step 4.1: Use the graph attention network to learn the similarity between clients and select the most relevant edges based on the attention weights to obtain the communication topology. Step 4.2: The central server generates initial model parameters , the model parameters contain all the network layers and parameters that need to be trained for the NILM task; then, the central server initializes the global model parameters Distribute to all clients participating in the training so that each client has the same starting point before the training begins. The mathematical expression is as follows: ; in, Represents the client The model parameters at initialization, is the total number of clients; Step 4.3: Perform local model training. Each client builds a communication topology based on the Share model parameters with neighboring clients; each client After receiving the updated parameters from its neighboring clients, the model parameters of each client are aggregated by weighted averaging; Step 4.4: Evaluate convergence and record the judgment result as a binary signal: , in, Indicates that the client believes that the current model has converged. Indicates that the model has not yet converged; Represents the client Validate the set locally The loss value on ; is the preset threshold, is the target accuracy, For Model Validate the set locally The accuracy of Step 4.5: After each round of training is completed, all clients send their own convergence signals Sent to the central server; the central server determines whether to stop training based on the statistical voting results. ; Set voting threshold is a static fixed value. , the system stops training, otherwise it enters the next round of training.

5. According to claim 4, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: Step 4.1 includes: Step 4.1.1: Each client First, based on local data , extract key statistical features, including mean, standard deviation, minimum and maximum values, mathematically expressed as: , Among them, the characteristics of all clients are expressed as , ; make Represents the communication topology, client set , edge set Represents the potential communication connection relationship between clients, and initializes the edge set to a fully connected structure , does not include self-loops , that is, initially set to a completely directed graph ; Step 4.1.2: Use a three-layer graph attention network to build communication connections between clients. With its neighbor clients , the graph attention network first performs a learnable linear mapping on the client features ; Then, for each edge Calculate the attention coefficient , which is of the form: , in, is the activation function, is the learnable attention vector, represents vector concatenation, and Respectively represent the client and its neighbor clients The feature matrix of Then, after softmax normalization, the final attention coefficient is obtained : ; In the forward propagation process, the updated node representation is further obtained. And the set of attention coefficients of all edges ; Represented as client and The connection weight of Call , in, Embedding matrix for the updated client, is the set of attention coefficients; Step 4.1.3: Set the attention weight threshold for The median of The topological edges of: , and remove self-loops ; for Each directed edge in , only the undirected form is retained , the obtained undirected edge set ; For any client , Count its degree ,in, Representation and Connected neighbor client; if there is , indicating that the client is If there is no connecting edge in Pick a neighbor client and merge the new edge into it ; make , the final undirected graph This is the decentralized communication topology obtained based on GAT learning.

6. According to claim 1, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: Step 2 specifically includes: When network bandwidth ,Delay , and the data privacy requirements are low, the system chooses centralized federated learning, otherwise, it switches to decentralized federated learning; If the conditions for centralized federated learning training are met, the central server will aggregate and optimize the model updates of all clients; If the conditions for decentralized federated learning training are triggered, the client trains independently in a distributed environment and exchanges model updates with neighboring clients.

7. A non-intrusive load monitoring system based on hybrid federated learning and cross-Transformer, characterized in that: The system comprises: A preprocessing module, used for preprocessing power signal data; A federated learning training mode selection module is used to determine the federated learning training mode according to the monitoring network bandwidth, latency and data privacy requirements, wherein the federated learning training mode includes centralized federated learning and decentralized federated learning; if centralized federated learning is determined, the centralized federated learning module is executed, otherwise the decentralized federated learning module is executed; The centralized federated learning module is used for centralized federated learning training. The CrossTransNILM model is deployed on the client and the central server, and all client model updates participating in the training are centralized on the central server for aggregation. The CrossTransNILM model includes a multi-layer Transformer encoder and a load branch module, which predicts the power consumption of multiple loads through bus power. The CrossTransNILM model specifically includes: a one-dimensional convolutional layer, a positional encoding, a shared Transformer encoder, and multiple load branches, each of which processes a load type. Each load branch contains a branch Transformer encoder layer, a cross Transformer layer, and two fully connected layers. The shared Transformer encoder consists of multiple encoding layers, each of which includes a multi-head self-attention mechanism and a feed-forward neural network; Decentralized federated learning module, used for decentralized federated learning training. Each client independently trains the CrossTransNILM model and exchanges model parameters through point-to-point communication. The communication topology is optimized through the graph attention network, and the convergence of the global model is judged through the voting mechanism. The monitoring module is used for non-intrusive load monitoring and real-time feedback. It deploys the trained CrossTransNILM model to the monitoring terminal, collects the power signal of the main line in real time, and inputs it into the model for end-to-end monitoring after preprocessing. The power decomposition results of each electrical equipment are output through the load branch module, and the load power is mapped to the start and stop status of the electrical equipment. The monitoring terminal regularly updates the model parameters according to the federated learning mode selection mechanism.

8. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor runs the computer program stored in the memory, the steps of the method according to any one of claims 1 to 6 are performed.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of computer instructions, and the plurality of computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • A federated learning-based non-intrusive enterprise load decomposition method

    CN113962314A

  • Load monitoring model construction method and device, computer equipment, medium and product

    CN118839988A