A non-intrusive load monitoring method and system based on end-cloud collaboration splitting
Patent Information
- Application Number
- CN202611271703.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-22
AI Technical Summary
[0008]本发明的目的是提出一种基于端云协同拆分的非侵入式负荷监测方法及系统,通过构建物理先验门控多尺度拆分卷积网络、云端类别原型库以及跨设备梯度回传机制,解决非侵入式负荷监测(NILM)在边缘应用中面临的数据隐私泄露、通信带宽占用过高、通用CNN特征表达不足以及静态模型难以应对负荷特性演变(概念漂移)等核心技术难题
第一,通过多通道V-I物理轨迹、物理先验门控多尺度特征提取、类别原型漂移检测和联合损失约束,提高负荷特征的可分性以及模型在数据分布漂移条件下的持续适应能力,并抑制历史知识遗忘;
Smart Images

Figure CN122802557A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart grid and artificial intelligence cross-application technology, and in particular to a non-intrusive load monitoring (NILM) method and system based on edge-cloud collaborative decomposition and online continuous learning. This method is mainly used for online adaptive load identification, state feature extraction, and dynamic model fine-tuning of resource-constrained edge monitoring nodes (such as smart meters, home energy gateways, distributed electrical equipment, etc.), and is applicable to ubiquitous power Internet of Things, smart home energy management, and abnormal state monitoring of electrical equipment. Background Technology
[0002] With the rapid development of smart grids and home energy management systems, non-intrusive load monitoring (NILM) technology has been widely used in power efficiency monitoring and electricity consumption behavior analysis because it eliminates the need for individual sensor installations at each electrical device; it only requires collecting electrical signals at the main incoming line to achieve load decomposition and status identification. However, in real-world power grid applications, the core challenge for the successful implementation of NILM technology lies in how to efficiently, accurately, and stably identify various types of electrical loads over the long term.
[0003] Existing NILM system physical architectures are mainly divided into two modes: "pure cloud-based centralized identification" and "pure edge-based local identification." However, both have revealed insurmountable technical bottlenecks in actual physical deployment. For the "pure cloud" architecture, the system needs to directly upload the original discrete waveforms of voltage and current sampled at high frequencies to the cloud server for simulation via long connections. This not only consumes a great deal of communication bandwidth resources, but more seriously, the original high-frequency electrical waveforms contain a large number of extremely sensitive features such as family routines and device usage preferences, posing a significant risk of user privacy leaks. For the "pure edge" architecture, limited by the extremely low computing power of edge microcontrollers and the extremely limited SRAM static random access memory, edge devices can often only deploy lightweight static pre-trained models that have been extremely pruned. Their feature extraction capabilities are weak, and the accuracy of blind testing has a clear upper limit.
[0004] Furthermore, power load data exhibits high time-varying and non-stationarity in actual deployment environments. On the one hand, with the daily wear and tear of household appliances and capacitor aging, their physical characteristics, such as voltage-current (VI) waveform trajectories, will slowly evolve; on the other hand, users will continuously connect new models or unknown brands of appliances in their daily lives. This phenomenon, where the distribution of on-site power consumption data deviates from the initial training data distribution over time, is known as "concept drift" in the field of machine learning. Because existing lightweight edge monitoring models are "fixed from the factory," once they encounter concept drift in the physical world, their recognition accuracy will plummet over time, exhibiting extremely weak system generalization ability and environmental adaptability, resulting in a large number of false alarms and missed alarms.
[0005] To address concept drift, some existing cutting-edge technologies attempt to introduce online learning or edge-local continuous learning mechanisms. However, traditional online fine-tuning of neural networks requires devices to possess full automatic differentiation computing power in order to build a large and complex backpropagation graph locally to update parameters at all levels. This fully parametric backpropagation process incurs enormous overhead in terms of memory and computing power, completely exceeding the physical hardware limits of most low-cost IoT monitoring nodes, making it impossible for traditional continuous learning algorithms to be implemented on edge MCU devices.
[0006] Existing non-intrusive load monitoring technologies struggle to simultaneously address privacy protection, communication bandwidth limitations, low-computing-power hardware bottlenecks, and the adaptive capability to resist "concept drift." Therefore, for highly nonlinear and time-varying load data, there is an urgent need to research a novel non-intrusive load monitoring method and system that can break down the cloud-edge computing power barrier, continuously absorb new field data for model fine-tuning after device deployment with extremely low computational cost, and achieve "increasing accuracy with use and personalized results for each user." This represents an effective way to address the challenges currently facing smart electricity technologies.
[0007] However, existing technologies have not yet solved the problem of how to achieve continuous online evolution of edge-side models on low-cost MCUs with SRAM significantly less than 512KB. Standard deep learning training mechanisms require edge nodes to store forward activations and build backpropagation computation graphs, whose memory overhead far exceeds the physical limits of low-cost MCUs. This means that those skilled in the art can usually only use static pre-trained models or centralized inference solutions in the cloud, which makes it difficult to cope with concept drift after deployment. Summary of the Invention
[0008] The purpose of this invention is to propose a non-intrusive load monitoring method and system based on edge-cloud collaborative splitting. By constructing a physically prior gated multi-scale splitting convolutional network, a cloud-based category prototype library, and a cross-device gradient backhaul mechanism, it solves the core technical challenges faced by non-intrusive load monitoring (NILM) in edge applications, such as data privacy leakage, excessive communication bandwidth consumption, insufficient general CNN feature representation, and the inability of static models to cope with load characteristic evolution (concept drift).
[0009] A non-intrusive load monitoring method based on edge-cloud collaborative splitting is proposed. The method uses a collaborative splitting convolutional network for load monitoring. The collaborative splitting convolutional network is a physically prior-gated multi-scale splitting convolutional network. The collaborative splitting convolutional network is divided into an edge-side sub-network deployed at the edge monitoring node and a cloud-side sub-network deployed on the cloud server according to the network splitting point. Includes the following steps: Step 1: Data acquisition and preprocessing; Voltage and current discrete sequences are acquired at the user-side incoming line end, and the voltage and current discrete sequences are denoised, periodically sliced, and time-domain aligned to obtain time-domain aligned voltage and current discrete sequences. Step 2: Obtain hidden layer feature tensors based on collaborative splitting convolutional networks; Construct a multi-channel VI physical trajectory tensor based on time-domain aligned discrete sequences of voltage and current; The multi-channel VI physical trajectory tensor is input into the edge sub-network of the collaborative split convolutional network deployed on the edge monitoring node; the collaborative split convolutional network is a physical prior gated multi-scale split convolutional network (i.e., PGMS-SplitCNN). Multi-scale trajectory features are extracted by multi-scale convolutional branches in the collaborative split convolutional network, and the outputs of the multi-scale convolutional branches are weighted and fused by the physical prior gating fusion module in the collaborative split convolutional network to obtain the hidden layer feature tensor. Step 3, calculate the prototype offset; The edge monitoring node only sends the hidden layer feature tensor as the upload data object of the current sample to the cloud server; the physical prior vector corresponding to the current sample is mapped and encoded to the dedicated channel of the hidden layer feature tensor through a predetermined mapping, and is not uploaded as an independent data object; the cloud server uses the cloud-side sub-network to output the load class probability, and calculates the prototype offset of the hidden layer feature tensor relative to the predicted class prototype center based on the class prototype library; The prototype offset , is the normalized hidden layer feature tensor With the Prototype Center of Predictive Categories The Euclidean distance between them; prototype drift loss is used to constrain the stability of class prototypes; Step 4: Calculate the joint loss and the gradient at the split point; The "prototype offset" represents the distance between the normalized hidden layer feature tensor and the center of the predicted class prototype; the "prototype drift loss" is used as a loss term in the joint loss to constrain the stability of the class prototype. The edge-cloud network is a collaborative neural network deployed on edge monitoring nodes and cloud servers. The convolutional network is the specific network structure of the edge-cloud network, and it is divided into edge-side sub-networks and cloud-side sub-networks according to the network split point. The edge-cloud network split point is the "network split point between the edge-side sub-network and the cloud-side sub-network", wherein the cloud-side sub-network is deployed on the cloud server. When the maximum value of the load category probability is lower than the dynamic safety threshold, or the prototype offset is higher than the drift threshold, the cloud server calculates the joint loss based on the cross-entropy loss, prototype drift loss, and knowledge retention regularization term. The cloud server first calculates the feature gradient of the hidden layer feature tensor at the network split point based on the joint loss, and then generates the parameter gradient corresponding to the updatable parameters of the edge sub-network according to the parameter mapping relationship of the edge sub-network. The load category probabilities are output by the cloud-side sub-network. Specifically, the cloud server inputs the hidden layer feature tensor into the cloud-side sub-network, which then passes through a deep discriminant network and a Softmax classification layer (both mature existing technologies) to obtain the probabilities of each load category. The maximum value among these probabilities is taken as the maximum load category probability. The dynamic safety threshold ranges from 0.80 to 0.99, preferably 0.90, and can be adjusted based on the moving average of the maximum load category probability within the current update window. The drift threshold ranges from 1.0 to 1.5, preferably 1.20. The dynamic security threshold is denoted as The value ranges from 0.80 to 0.99, preferably 0.90, and is adjusted based on the moving average of the probability of the maximum load category within the current update window to meet the following requirements. ;in, This is the initial safety threshold, typically set to 0.90. This means that a maximum load category probability of 90% is required before the classification result is considered to have high reliability. This is the moving average of the probability of the highest load category within the current update window. To ensure that the reference probability remains consistent with the initial safety threshold, The adjustment coefficient is typically set to 0.1 to avoid excessively large threshold adjustments. , ; For the j-th sample within the current update window, first calculate the probability of the maximum load category: .in, This represents the probability that the j-th sample belongs to the k-th load category. A moving average is calculated using a sliding window of length K, typically K is 16. ; The joint loss is used to jointly constrain the classification accuracy, class prototype stability, and historical knowledge preservation of the load. It obtains the feature gradient of the hidden layer feature tensor at the network split point through backpropagation, and then generates the parameter gradient for updating the parameters of the edge sub-network based on the parameter mapping relationship of the edge sub-network. In this specification, the above feature gradient and the parameter gradient generated by its mapping are collectively referred to as the split point gradient. In the context of gradient distribution and parameter update in step 5, the split point gradient specifically refers to the mapped parameter gradient. The edge-cloud network is a collaborative split convolutional network deployed on edge monitoring nodes and cloud servers. The edge-side sub-network belongs to the edge side part of the collaborative split convolutional network, and the cloud-side sub-network is deployed on the cloud server. The two networks are split by the hidden layer feature tensor transmission position. The collaborative split convolutional network is divided into edge-side sub-networks and cloud-side sub-networks according to the network split point; the cloud server is located at the network split point between the edge-side sub-networks and the cloud-side sub-networks. The gradient at the split point originates from the backpropagation gradient of the joint loss L onto the hidden layer feature tensor H at the split point, i.e. The cloud server calculates the joint loss based on the hidden layer feature tensor and the corresponding payload label, using cross-entropy loss, prototype drift loss, and knowledge retention regularization term, and obtains the feature gradient at the split point through backpropagation; for the first... Let there be a sampling window, and let the hidden layer feature tensor at the cut point be... The corresponding joint loss is The gradient of the split point is then expressed as: The gradient at the split point is converted into the gradients of the physical prior gating parameters, low-rank drift adaptation parameters, and multi-scale convolution parameters selected based on the parameter mapping relationship of the edge-side sub-network, and then sent to the edge monitoring node for updating the edge-side sub-network parameters in step 5. The cloud server maps this gradient and the topology relationship of the edge-side sub-network into the gradients of the physical prior gating parameters, low-rank drift adaptation parameters, and optional multi-scale convolution parameters, and sends these gradients to the edge monitoring node. Therefore, the split point gradient received in step 5 is generated from the joint loss calculated in step 4 and is used to update the edge-side sub-network parameters. Step 5: Update the parameters of the edge-side sub-network based on the gradient of the split point, and implement non-intrusive load monitoring; The edge monitoring node receives the gradient of the split point and writes it into the grouped asynchronous gradient buffer. When the grouped asynchronous gradient buffer reaches the batch threshold, the batch threshold is 8 to 64, preferably 32. The physical prior gating parameters and low-rank drift adaptation parameters in the edge sub-network are updated online. When the preset drift conditions are met, the multi-scale convolution parameters selected by the cloud are updated to achieve continuous learning under the condition that the forward activation values are not saved on the edge side, the backpropagation computation graph is not constructed, and the local loss calculation is not performed. After the online update is completed, the updated edge sub-network continues to extract features from the subsequently acquired multi-channel VI physical trajectory tensors and uploads the hidden layer feature tensors to the cloud sub-network. The cloud sub-network then outputs the load category, operating status, and confidence level, thereby achieving non-intrusive load monitoring.
[0010] In step 5, the edge monitoring node receives the gradient of the split point generated by the cloud based on the joint loss and after parameter mapping, which is the parameter gradient corresponding to the updatable parameters of the edge sub-network. After the gradient buffer reaches the batch threshold, the edge sub-network parameters are updated according to the parameter gradient.
[0011] The edge-side sub-network is a shallow feature extraction network deployed on the edge monitoring node, including a multi-scale convolutional branch, a physical prior gating fusion module, and a low-rank drift adaptation layer. It is used to perform low-computational feature extraction on the multi-channel VI physical trajectory tensor locally and output the hidden layer feature tensor. The cloud-side sub-network is a deep discriminative network deployed on the cloud server, including a deep residual convolutional layer, a global pooling layer, a class prototype constraint layer, and a Softmax classification layer. It is used to output the load class probability, calculate the prototype offset, determine the drift state, and generate the gradient of the split point based on the hidden layer feature tensor.
[0012] The final load monitoring is completed collaboratively by the edge sub-network and the cloud sub-network, where the edge sub-network is responsible for front-end feature extraction, and the cloud sub-network is responsible for classification and outputting the final load monitoring results.
[0013] The multi-channel VI physical trajectory tensor includes at least trajectory density channel data, orientation gradient channel data, transient transition channel data, and harmonic phase weight channel data; Among them, the trajectory density channel data is obtained by accumulating the normalized VI trajectory coordinate density, the orientation gradient channel data is obtained by encoding the motion direction of adjacent sampling points in the VI plane, the transient transition channel data is obtained by marking sampling points whose current mutation exceeds the adaptive threshold, and the harmonic phase weight channel data is obtained by mapping the fundamental phase shift and higher harmonic energy.
[0014] The multi-scale convolution branches include local trajectory convolution branches, closed-shape convolution branches, and cross-region convolution branches; wherein, the local trajectory convolution branches use 3×3 convolution kernels, the closed-shape convolution branches use 5×5 convolution kernels, and the cross-region convolution branches use 3×3 dilated convolution kernels with a dilation rate of 2.
[0015] The physical prior gating fusion includes: constructing a physical prior vector from the VI trajectory closure area, fundamental phase offset, current transient slope peak, harmonic distortion index, and trajectory sparsity; generating a gating vector based on the physical prior vector; and performing weighted fusion on the outputs of each multi-scale convolution branch.
[0016] The gate vector (the gate vector is formed by the physical prior vector through the gate weight matrix) and gated bias vector (After mapping and sigmoid activation, the resulting vector is used as a weighted fusion vector for the multi-scale convolution branch.) satisfy ; in: The Sigmoid activation function satisfies This is used to restrict each element of the gate vector to between 0 and 1; Represents the gate weight matrix. Let represent the gating bias vector; both are trainable parameters of the gating mapping layer and are not calculated temporarily from the input samples. Let the physical prior vector be... The dimension is The number of branches in multi-scale convolution is ,but: , ; First, calculate the gated pre-activation vector: ; Then, the gating vector is obtained through the activation function: ; in, Indicates the number of branches in multi-scale convolution, avoiding compatibility with low-rank fitting matrices. Confusing. and During the offline training phase, the joint loss is updated via backpropagation. During the online continuous learning phase, the gradient of the split point generated in the cloud is written into the gating parameter buffer and updated. Represents the area of the VI locus closure, and according to Calculation, where Let be the coordinates of the nth normalized VI trajectory, where and Let N represent the nth normalized current coordinate and the nth normalized voltage coordinate, respectively; the VI trajectory within one power frequency cycle satisfies the start-end closure condition, i.e., the end sampling point is connected to the start sampling point; N is the number of sampling points in one power frequency cycle, and... ;and This means connecting the end sampling point of the VI trajectory within one power frequency cycle to the start sampling point to ensure the physical prior quantity of the closed shape. The calculation is based on the VI trajectory with closed beginning and end. This represents the physical prior quantity of the closed-form calculated from the normalized VI trajectory, used to characterize the degree of closure and hysteresis characteristics of the VI trajectory within one power frequency cycle; this physical prior quantity is not used alone for load category determination, but is used in conjunction with the fundamental phase offset. Peak value of transient slope of current Harmonic distortion index and trajectory sparsity Together they form a physical prior vector, which is used to generate a gating vector and to perform weighted fusion of the VI trajectory graphic features extracted by the multi-scale convolution branches; This indicates the fundamental phase shift. Indicates the peak value of the transient slope of the current. Indicates harmonic distortion index; Represents the sparsity of the trajectory, and according to Calculation, where The number of non-zero accessed grid cells in the trajectory density matrix is denoted as M×M, and the total number of grid cells in the trajectory density matrix is M×M. M represents the number of grid cells divided along the normalized voltage axis and the normalized current axis, respectively. Therefore, the trajectory density matrix contains M×M grid cells.
[0017] The low-rank drift adaptation parameters include low-rank adaptation matrices A and B, such that A·B constitutes a low-rank increment matrix; in this specification, low-rank adaptation layer is an abbreviation for low-rank drift adaptation layer, and the two refer to the same network layer. in, , r is less than the edge-side stable convolution parameter The smaller of the corresponding matrix dimensions m and n; r is a positive integer from 1 to 16, preferably 4 or 8, and r is less than the smaller of the corresponding matrix dimensions m and n for the edge-side stable convolution parameter; The low-rank adaptation layer stabilizes the convolution parameters on the edge sides based on low-rank adaptation matrices A and B. Perform low-rank increment adjustment to obtain the equivalent edge-side convolution parameters. The low-rank adaptation layer satisfies: ; The low-rank adaptation layer is a trainable parameter layer set in the edge-side sub-network. It consists of low-rank adaptation matrices A and B. The low-rank increment matrix A·B is used to correct the parameters of the stable convolutional parameters on the edge side to adapt to concept drift. in, Let A and B be the edge-side stable convolution parameters, and let A and B be the low-rank adaptation matrices. To and The dimension-matched low-rank increment matrix is used to characterize the amount of edge-side convolution parameter correction caused by data drift; The equivalent edge-side convolution parameters after low-rank drift adaptation are used for subsequent forward feature extraction of the edge-side subnetwork; The initial values of A and B are obtained from offline training samples through backpropagation. During the online continuous learning phase, the cloud server calculates the gradients of A and B based on the joint loss and sends them to the edge monitoring nodes. A and B are iteratively updated through gradient descent, so that... Forming a pair Low-rank increments; priority updates during the online continuous learning phase. , And physical prior gating parameters, including the gating weight matrix. and gated bias vector Freeze or weak update The weak update refers to... learning rate Set the learning rates for A and B 0.01 to 0.1 times, preferably 0.05 times; freezing hour, It is 0.
[0018] The joint loss L satisfies: ; in: Cross-entropy loss; ;in, For the unique thermal encoding of the actual load label, The load category probability output by the cloud-side sub-network; k is the load category index, k=1,2,…,K, where K is the total number of known load categories; the one-hot encoding of the actual load label. Satisfy: If the first If each category is a true category, then ,otherwise ; The first output of the cloud-based side network The probability of each load category; Hidden layer feature tensor With Category Prototype Center The prototype drift loss between; The multi-channel VI physical trajectory tensor is sequentially computed through multi-scale convolutional branches, physical prior-gated weighted fusion, and low-rank drift adaptation, and then output by the edge-side sub-network; the category prototype center It is a center vector used to characterize the typical feature distribution position of category y load in the hidden layer feature space. It is represented by the mean of the hidden layer feature tensors of the historical correctly classified samples of category y. The initial value is determined by averaging the hidden layer feature tensors of category y in the offline training set, and is updated by moving average based on the hidden layer feature tensors of newly added correctly classified samples during the online continuous learning phase. To preserve regular terms for knowledge, where The parameter set consists of the current low-rank drift adaptation parameters and the physical prior gating parameters. Parameters from low-rank adaptation layer and physical prior gating parameters Specifically, it is expressed as: ; This is the set of corresponding parameters saved before this online update; and The weight coefficients for prototype drift loss and knowledge retention regularization are determined through a grid search on the validation set. , .
[0019] When the cloud server generates the gradient of the split point, it first calculates the feature gradient of the hidden layer feature tensor at the network split point based on the hidden layer feature tensor corresponding to the current sample of the edge sub-network, the physical prior vector extracted from the dedicated channel of the hidden layer feature tensor, the topology of the edge sub-network, and the prototype offset. Then, it calculates the parameter gradient corresponding to the physical prior gating parameter, the low-rank drift adaptation parameter, and the multi-scale convolution parameter selected by the cloud according to the drift degree according to the parameter mapping relationship. The above feature gradient and parameter gradient are collectively referred to as the split point gradient. The parameter gradient is grouped and serialized according to the parameter type and then sent to the edge monitoring node.
[0020] The grouped asynchronous gradient buffer includes a gating parameter buffer, a low-rank drift adaptation parameter buffer, and a convolution parameter buffer; the gating parameter buffer and the low-rank drift adaptation parameter buffer are updated preferentially, while the convolution parameter buffer is only activated when the prototype offset exceeds the drift threshold for multiple consecutive rounds. The prototype offset is the Euclidean distance between the normalized hidden layer feature tensor and the prototype center of the corresponding predicted class; the drift threshold is 1.0 to 1.5, preferably 1.20.
[0021] A non-intrusive load monitoring system based on edge-cloud collaborative splitting is used to implement the aforementioned method; The system includes edge monitoring nodes and a cloud server. The edge monitoring nodes include a high-frequency sampling module, a physical trajectory construction module, an edge-side sub-network module, a grouped asynchronous gradient buffer, and an edge communication module. The cloud server includes a cloud-side sub-network module, a category prototype library, a split-point gradient generation module, and a cloud communication module. Furthermore, the system also includes a deadlock prevention handshake protocol based on a single-byte hardware synchronization permission order, used to confirm the completion of receiving gradient data frames at the split point and trigger the edge monitoring node to unblock and listen. In step 1, multi-core modular smart meters or electrical signal acquisition modules with high-frequency sampling capabilities are used to read instantaneous voltage and current data from the user-side power bus in real time. The sampling frequency should satisfy the sampling theorem, preferably 128 sampling points per power frequency cycle or higher. For the acquired raw discrete signal sequence, time-domain alignment and denoising are performed, and a sliding window filtering algorithm is used to eliminate high-frequency noise interference from the power grid, constructing a standardized user-side power consumption sample database. Step 2 involves meshing and edge-side physical prior gated convolution feature extraction based on multi-channel VI physical trajectories. The preprocessed voltage and current sequences are normalized and mapped to a multi-channel VI trajectory tensor reflecting the physical operating characteristics of the electrical equipment. This tensor includes not only traditional trajectory density maps but also directional gradient maps, transient transition maps, and harmonic phase weight maps. The edge monitoring node locally invokes the Collaborative Split Convolutional Network (PGMS-SplitCNN) to perform forward computation, extracting hidden layer feature tensors containing key electrical features of the load and possessing privacy-preserving properties. (1) in, It is a non-linear activation function. After calculation, the edge monitoring nodes only report the hidden layer feature tensor. The original high-frequency waveform sequence does not undergo network transmission; Step 3: Cloud-based collaborative deep load extrapolation and identification result judgment. The cloud server receives the hidden layer feature tensor. The data is then fed into a deep classification network model. This model consists of multiple deep convolutional layers and fully connected layers, used to perform high-dimensional nonlinear mapping on abstract features. Finally, the Softmax function outputs the posterior probability distribution of the current load belonging to each known electrical equipment category. (2) The cloud server identifies the load category based on the highest probability value and outputs the identification results and corresponding operating status parameters. Step 4: Continuous Learning Instruction Triggering and Split Point Gradient Backpropagation. The system dynamically determines whether to enable online continuous learning mode based on the maximum load class probability and prototype offset inferred from the cloud. Online continuous learning is triggered when the maximum load class probability is lower than the dynamic safety threshold or the prototype offset is higher than the drift threshold. After obtaining the true load label of the current sample, the cloud server calculates the joint loss based on the prediction result and the true load label, i.e.: (3) The triggering conditions for continuous learning mode are: the maximum load category probability is lower than the dynamic safety threshold, or the prototype offset is higher than the drift threshold, and the on-site real-time feedback label is used to provide the true label y; The gradient differentiation of deep networks is performed in the cloud using the backpropagation algorithm. First, the feature gradient of the hidden layer feature tensor at the network split point is calculated. Then, the parameter gradient corresponding to the updatable parameters is generated based on the parameter mapping relationship of the edge sub-networks. The feature gradient and the parameter gradient generated by their mapping are collectively referred to as the split point gradient; the loss is the joint loss. ,in: (4) For prototype drift loss, Preserve regularization terms for knowledge; backpropagation first obtains the gradient of the split point. Then, according to the chain rule, it is mapped to the gradient of the updateable parameters on the edge side; Step 5: Asynchronous Update of Edge Parameters and Local Model Evolution. After receiving the error gradient components, the edge monitoring node stores them in a pre-defined asynchronous gradient buffer within its local static random access memory (SRAM). A batch accumulation mechanism is used to vector-sum the gradients of multiple samples, generating the accumulated gradient. When the accumulated number of samples reaches the preset batch threshold. At that time, the weight parameters of the front-end network are updated online using a local parameter fine-tuning algorithm: (5) in, This involves controlled online fine-tuning of the learning rate. After the update, the edge monitoring node completes local model evolution and is able to perform the next load feature extraction with the updated parameters.
[0022] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: First, by using multi-channel VI physical trajectories, physical prior gated multi-scale feature extraction, category prototype drift detection, and joint loss constraints, the separability of load features and the model's continuous adaptability under data distribution drift conditions are improved, and historical knowledge forgetting is suppressed. Second, the edge monitoring node only uploads the hidden layer feature tensor containing the physical prior vector encoding channel, while the original voltage and current waveforms are retained on the edge side, thereby reducing communication bandwidth usage and reducing the risk of leakage of original power consumption data. Third, the cloud performs joint loss calculation, feature gradient differentiation at network split points, and parameter gradient mapping, while the edge side only performs lightweight forward computation and parameter writing, reducing the computational burden on edge monitoring nodes. Fourth, by prioritizing the updating of low-rank drift adaptation parameters and physical prior gating parameters, and using grouped asynchronous gradient buffers for batch accumulation, the SRAM usage of online continuous learning is reduced while keeping the convolution parameters frozen or weakly updated. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the implementation of the present invention;
[0024] Figure 2 This is a diagram of the edge-cloud collaborative system architecture.
[0025] Figure 3 A flowchart for data processing and collaborative reasoning;
[0026] Figure 4 This is a graph evaluating the multi-class classification performance of a load monitoring model based on a confusion matrix. Detailed Implementation
[0027] The following will combine Figures 1-4 The technical solution of the present invention will be further described in detail below.
[0028] This embodiment is based on a physically-gated multi-scale splitting convolutional network, and describes the specific implementation of edge-cloud collaborative inference, prototype drift triggering, split-point gradient generation, and asynchronous update of edge-side groups. Unless otherwise specified, the technical terms in this embodiment have the same meaning as those in the foregoing invention.
[0029] A non-intrusive load monitoring method based on edge-cloud collaborative splitting and online continuous learning, such as Figure 1 As shown, firstly, high-frequency voltage and current signals are collected in real time using edge monitoring nodes deployed on the load bus. A multi-channel VI physical trajectory tensor is constructed using time-domain alignment and normalization algorithms. Subsequently, forward convolution calculations are performed locally on the edge monitoring nodes using edge-side sub-networks to generate a dimensionality-reduced hidden layer feature tensor. The physical prior vector is mapped to a dedicated channel of the hidden layer feature tensor via a predetermined mapping. The edge monitoring nodes only upload the hidden layer feature tensor to the cloud server. The cloud server outputs the load category probability and calculates the prototype offset. After triggering online continuous learning and obtaining the real load labels, the cloud server calculates the joint loss, first obtaining the feature gradient of the hidden layer feature tensor at the network split point, and then mapping it to the parameter gradient of the edge-side updatable parameters. The above gradients are collectively referred to as split point gradients, and the mapped parameter gradients are sent back. Finally, the edge monitoring nodes use a grouped asynchronous gradient buffer for batch accumulation and parameter writing, thereby achieving adaptive adjustment of the load feature offset.
[0030] The specific implementation steps of the non-intrusive load monitoring method based on edge-cloud collaborative splitting and online continuous learning are as follows: Step 1: Collect high-frequency voltage and current data of user electricity consumption using edge devices; Step 2: Construct a multi-channel VI physical trajectory tensor at the edge and use the edge sub-network of the physical prior gated multi-scale split convolutional network to extract features from the electricity consumption data and generate a hidden layer feature tensor; the physical prior vector corresponding to the current sample is mapped and encoded to a dedicated channel of the hidden layer feature tensor through a predetermined mapping, and the edge monitoring node only uploads the hidden layer feature tensor to the cloud; Step 3: The cloud server receives the hidden layer feature tensor, performs forward inference using a deep classification network, and outputs the load identification result. Step 4: In continuous learning mode, the cloud combines real feedback, recognition confidence and category prototype offset to calculate joint loss and perform backpropagation. First, the feature gradient of the hidden layer feature tensor at the network split point is calculated, and then the parameter gradient corresponding to the updatable parameter is generated according to the parameter mapping relationship of the edge sub-network. The above gradients are collectively referred to as split point gradients, and the mapped parameter gradients are sent to the edge device. Step 5: The edge device receives the gradient of the split point as the parameter gradient after mapping, and performs batch accumulation through the grouped asynchronous gradient buffer. It prioritizes updating the physical prior gating parameters and low-rank drift adaptation parameters, and updates the multi-scale convolution parameters selected by the cloud when the preset drift conditions are met.
[0031] Step 1: A multi-dimensional electricity consumption data collection and basic database was constructed; Edge monitoring nodes (such as intelligent acquisition terminals based on embedded microcontrollers or modular smart meters) deployed on the user's incoming line side are used to read the bus voltage and current signals on the user side in real time. The edge monitoring nodes perform analog-to-digital conversion and data parsing on the collected analog signals to ensure that the voltage and current signals are strictly aligned in the time domain and processed in a dimension-unified manner, thereby constructing a basic database of user-side electricity consumption samples for edge computing. In step 1, the core processing chip of the edge monitoring node used in this embodiment is a microcontroller based on the ARM Cortex-M core (such as the STM32 series chip). The data acquired through the high-performance analog front end is the original discrete sequence of voltage and current. The sampling frequency is set to sample 128 points per power frequency cycle (i.e., for a 50Hz power grid, the sampling frequency is 6.4kHz) to ensure that the high-frequency transient characteristics and high-order harmonic information at the moment of load connection can be captured. This embodiment covers 11 typical household electrical devices in the experimental verification phase. Through massive data collection under different operating conditions, a sample set of electrical devices was constructed to verify the model performance. The device types are shown in Table 1: .
[0032] Step 2: Obtain the hidden layer feature tensor H based on the collaborative split convolutional network; Edge monitoring nodes construct multi-channel VI physical trajectory tensors based on time-domain aligned discrete voltage and current sequences. These multi-channel VI physical trajectory tensors include at least a trajectory density channel, a gradient orientation channel, a transient transition channel, and a harmonic phase weight channel. The trajectory density channel is obtained by accumulating the density of normalized VI trajectory coordinates; the gradient orientation channel is obtained by encoding the motion direction of adjacent sampling points in the VI plane; the transient transition channel is obtained by marking sampling points whose current abrupt changes exceed an adaptive threshold; and the harmonic phase weight channel is obtained by mapping the fundamental phase shift and higher harmonic energy. The edge subnetwork includes local trajectory convolution branches, closed-shape convolution branches, and cross-region convolution branches. The local trajectory convolution branch employs... Convolution kernel, closed-shape convolution branches adopt Convolution kernel, cross-region convolution branches adopt Hollow convolution kernel with a void ratio Each branch is used to extract local trajectories, closed patterns, and cross-regional correlation features, respectively. The physical prior gating fusion module consists of the VI trajectory closure area. Fundamental phase offset Peak value of transient slope of current Harmonic distortion index (THD) and trajectory sparsity Construct a physical prior vector and generate a gating vector based on it. : (6) in, It is the Sigmoid activation function. For the gated weight matrix, Let be the gated bias vector, and: (7) Let the dimension of the physical prior vector be . The number of branches in multi-scale convolution is ,but , The physical prior vector in this embodiment consists of the five physical quantities mentioned above, therefore Gating vector The elements are used to perform weighted fusion of the outputs of the corresponding multi-scale convolution branches; Area of VI trajectory closure Calculate according to the following formula: (8) The VI trajectory within one power frequency cycle satisfies the head-tail closure condition. .in, For the first A normalized VI trajectory coordinate, Represents normalized voltage coordinates. This represents the normalized current coordinates. Let the trajectory density matrix be divided along the two coordinate axes as follows: The number of grid cells with non-zero access counts in the trajectory density matrix is . Then the sparsity of the trajectory Calculate according to the following formula: (9) The outputs of each multi-scale convolution branch are gated by a vector. The weighted fusion is then input into a low-rank drift adaptation layer. The edge-side sub-network is equipped with a low-rank drift adaptation layer, which includes low-rank drift adaptation parameters, including a low-rank adaptation matrix. and The dimension of the low-rank adaptation matrix. Among them, the low-rank value The integer is a positive integer from 1 to 16, and is chosen as 8. Let the edge-side stable convolution parameters be... The equivalent edge-side convolution parameters after low-rank drift adaptation are: ,but: (10) Offline training phase determined , , , and Initial values; prioritized updates during the online continuous learning phase. , , and and freeze or weakly update. The hidden feature tensor of the edge subnetwork after low-rank drift adaptation is output. During the normal inference phase, edge monitoring nodes only upload data. The original discrete voltage and current sequences are not uploaded to the cloud.
[0033] Step 3: Calculate the prototype offset and output the load category probability; The cloud server receives the hidden layer feature tensor. ,Will The input to the cloud-side sub-network is processed through a deep discriminant network and a Softmax classification layer to obtain the probability of each load category. The output of the cloud-side sub-network can be represented as follows: , its first Each component is denoted as The probability of the maximum load category is ; The category prototype library stores the prototype centers for each load category. The normalized hidden layer feature tensor is still denoted as... The corresponding category prototype center is denoted as Then the prototype offset Calculated using Euclidean distance: (11) The dynamic security threshold is denoted as Its value ranges from 0.80 to 0.99, with an initial value of 0.90. The adjustment formula for the dynamic safety threshold is: (12) in, , , , The threshold adjustment factor is set to 0.1. Maintain consistency with the initial security threshold. For the current update window... For each sample, calculate first: (13) Then use a length of The sliding window is used to calculate the moving average of the probability of the maximum load category: (14) In this embodiment, the sliding window length is selected as [value]. The prototype drift threshold is set between 1.0 and 1.5, and 1.20 is chosen. This is when the probability of the maximum load category is not lower than... And prototype offset When the load type is not higher than the prototype drift threshold, the cloud directly outputs the load category, operating status, and confidence level; when the maximum load category probability is lower than... or prototype offset When the value exceeds the prototype drift threshold, an online continuous learning process is triggered. After triggering online continuous learning, the cloud server acquires or receives the true load label of the current sample through on-site calibration, user confirmation, or association with equipment control records. Supervisory joint loss is not calculated if a true load label is not obtained; proceed to step 4 after obtaining the true load label.
[0034] Step 4: Calculate the joint loss L and the gradient at the split point; The cloud server is based on real load labels, load category probabilities, and hidden layer feature tensors. and category prototype center Calculate joint loss The notation and expression for joint loss are as follows: (3) Among them, cross-entropy loss for: (4) In the formula, For the unique thermal encoding of the actual load label, The first output of the cloud-based side network Each load category probability. Prototype drift loss. for: (15) Knowledge retention regularization for: (16) Among them, the set of currently updatable parameters satisfy . This is the set of corresponding parameters saved before this online update. and These are the weighting coefficients for prototype drift loss and knowledge retention regularization, respectively. In this embodiment, we select... , ; The main backpropagation process is completed on a cloud server. (Regarding the first...) A sampling window, the hidden layer feature tensor at the split point is denoted as . The corresponding joint loss is denoted as : (17) The cloud server uses the feature gradient of the hidden layer feature tensor at the network segmentation point. The physical prior vector, edge subnetwork topology, and prototype offset extracted from the dedicated channel of the hidden layer feature tensor Physical prior gating parameters are generated by mapping according to known computational relationships and the chain rule. , The parameter gradient, the low-rank fitness matrix in the low-rank drift fitting parameters. , The parameter gradients, as well as a small number of multi-scale convolution parameter gradients selected by the cloud under continuous drift conditions, are mapped. The above-described parameter gradients are used as the segmentation point gradients sent to the edge monitoring nodes, and are grouped and serialized according to the gating parameters, low-rank drift adaptation parameters, and convolution parameters. To avoid gradient mismatch caused by asynchronous updates of the edge-side model, the edge monitoring node and the cloud server pre-establish a correspondence between the model version and the edge-side sub-network structure during model deployment or parameter synchronization. During the online inference phase, the model version number, physical prior vector, or edge-side sub-network structure identifier is not uploaded separately with the current sample. The physical prior vector corresponding to the current sample is encoded into a dedicated channel of the hidden layer feature tensor through a predetermined mapping. The cloud server generates gradient data frames according to the existing session configuration. The edge monitoring node only receives and writes gradients when the local model version is consistent with the session configuration; if the versions are inconsistent, parameter synchronization is performed first or the current gradient is discarded.
[0035] Step 5: Update the parameters of the edge-side sub-network based on the gradient of the split point and continue to implement load monitoring; The edge monitoring node is equipped with a grouped asynchronous gradient buffer, which includes a gated parameter buffer, a low-rank drift adaptation parameter buffer, and a convolution parameter buffer. The gated parameter buffer is used for accumulation. and The gradient; the low-rank drift adaptation parameter buffer is used for accumulation. and The gradient; the convolution parameter buffer is only used to accumulate the gradient of a small number of multi-scale convolution parameters selected by the cloud according to the degree of drift; Batch threshold for grouped asynchronous gradient buffers Choose from 8 to 64. Each time an edge monitoring node receives the grouped gradient of a sample, it accumulates the gradient into the corresponding buffer, and the accumulated result is recorded as... When the accumulated number of samples reaches the batch threshold At that time, write the updated execution parameters according to the aforementioned instructions: (5) in, This indicates the parameters before the update. Indicates the updated parameters. For controlled online fine-tuning of the learning rate, gating parameters and low-rank drift adaptation parameters are updated first; stable convolution parameters are kept frozen or weakly updated. During weak updates, the learning rate of stable convolution parameters is... Learning rate for adapting parameters to low-rank drift The value is 0.01 to 0.1 times the value of the selected value. When frozen The convolution parameter buffer is only applied at the prototype offset. Enabled when the prototype drift threshold is exceeded for multiple consecutive rounds; The edge monitoring node only performs gradient accumulation, learning rate scaling, and parameter writing locally; it does not calculate local loss, save the complete forward activation, or construct the complete backpropagation computation graph. After the update, the corresponding gradient buffer and counter are zeroed, and the updated edge sub-network is used to continue extracting the hidden layer feature tensors of subsequent samples. This creates a continuous learning loop; In a preferred communication implementation, the cloud server appends a single-byte synchronization permission command (0xAA) to the end of the gradient data frame. After the edge monitoring node fully receives the gradient data frame, passes the verification, and detects 0xAA, it writes the gradient into the corresponding buffer and enters the next processing cycle. This synchronization permission command is a preferred deadlock-preventing communication implementation. With 4 input channels, 16 output channels, and a low rank value and batch threshold For example, the edge side mainly stores gating parameters, low-rank drift adaptation parameters, and a small number of optional convolution parameters. By using 16-bit fixed-point storage and releasing the forward intermediate tensor layer by layer, the peak value of parameters and gradient buffers can be kept below 64 KB, and the peak value of a single forward intermediate activation can be kept below 128 KB, thus adapting to edge monitoring nodes with SRAM less than 512 KB; The validation was performed using 1100 edge-condition blind test samples from the 11 types of equipment listed in Table 1, with 100 samples from each type of equipment. The concept drift sample was constructed by applying a combination of amplitude perturbation, fundamental phase perturbation, and higher harmonic energy perturbation to the original steady-state VI trajectory, and was used to represent the load characteristic shift caused by equipment aging, motor wear, and changes in grid impedance. The online continuous learning function in steps 4 and 5 was turned off, and only the static model was used for inference, resulting in the results shown in Table 2. ; As shown in Table 2, due to the extremely similar current waveform distortion and start-up transient characteristics of categories 0 (electric fans), 2 (air conditioners), and 3 (electric water heaters), the static model, being "fixed at the factory," is unable to capture new physical boundaries after a data shift occurs. This results in a severe collapse in recall (only 0.5667 for category 0 and only 0.6333 for category 2), limiting the system's macroscopic test accuracy to 86.06%.
[0036] Experimental group: Evolutionary performance of the model of this invention (online continuous learning mode); Activating the online continuous learning closed-loop mechanism of this invention, the edge monitoring node begins to receive gradients and perform dynamic fine-tuning (learning rate). (The cloud server will handle the joint loss.) First, calculate the feature gradient of the hidden layer feature tensor at the network split point. Then, based on the parameter mapping relationship, parameter gradients for distribution are generated; edge monitoring nodes update using grouped asynchronous gradient buffers. , , and After the online absorption of this batch of data stream, the classification results of the system are shown in Table 3. .
[0037] The following is a comparison of physics experiment data before and after continuous online learning: 1) Effectively improves concept drift and similar waveform confusion: For nonlinear loads with similar electrical characteristics that are prone to confusion (such as single-phase asynchronous motors of category 0 and variable frequency compressors of category 2), this system enhances the expression of directional, transient, and harmonic differences through multi-channel VI physical trajectory input, amplifies key load features through a physical prior-gated multi-scale convolution module, and fine-tunes the segmentation point parameters by continuously absorbing error gradients. This improves the recall rate from 0.5667 and 0.6333 in the static model to 0.9000 and 0.9300, respectively. This indicates that the model's underlying parameters can effectively fit the evolution characteristics of data distribution in the physical environment. 2) The effectiveness of the knowledge retention mechanism and the low-rank drift adaptation mechanism was verified: while improving the recognition rate of newly added drift features, the system's recognition index for steady-state resistive loads (such as category 9) remained stable and reached 1.0000. This result verifies that the retention strategy proposed in this invention, based on knowledge retention coefficient, category prototype constraint, and low-rank drift adaptation parameter update, can effectively suppress the catastrophic forgetting phenomenon in the incremental learning process of neural networks; 3) Improve the overall recognition accuracy of the system: Without increasing the cost of additional sensing hardware and relying solely on the existing microcontroller, this method improved the overall macroscopic accuracy of the system from 86.06% of the static baseline to 96.82% in a blind test of 1100 samples; 4) It should be noted that the precision of category 8 in the static model was 1.0000, while it was 0.9899 in the online continuous learning mode. This change is a minor fluctuation caused by boundary redistribution; the decrease is approximately 1.01 percentage points, which is a small fluctuation on a statistical scale of 100 test samples per class. Meanwhile, the recall of category 8 remained at 0.9800, and the F1-Score was 0.9849. In contrast, the recall of categories 0 and 2, which showed significant drift, increased from 0.5667 and 0.6333 to 0.9000 and 0.9300, respectively, and the overall system precision increased from 0.8606 to 0.9682. Therefore, this minor fluctuation does not affect the overall technical performance.
[0038] In summary, the specific embodiments and comparative experimental data show that the present invention significantly improves the long-term recognition accuracy and robustness of the non-intrusive load monitoring system in concept drift scenarios while maintaining extremely low edge hardware overhead and communication bandwidth usage.
[0039] The foregoing has shown and described the basic principles, main features, and advantages of this embodiment. Those skilled in the art should understand that this embodiment is not limited to the specific embodiments described above. The specific embodiments and descriptions in the specification are merely for further illustrating the principles of this embodiment. Various changes and modifications can be made to this embodiment without departing from the spirit and scope of this embodiment, and all such changes and modifications fall within the scope of this embodiment as claimed. The scope of protection of this embodiment is defined by the claims and their equivalents.
Claims
1. A non-intrusive load monitoring method based on edge-cloud collaborative splitting, characterized in that, Load monitoring is performed using a collaborative split convolutional network, which is a physically prior-gated multi-scale split convolutional network. The collaborative split convolutional network is divided into an edge-side sub-network deployed at the edge monitoring node and a cloud-side sub-network deployed on the cloud server according to the network split point. Includes the following steps: Step 1: Data acquisition and preprocessing; Voltage and current discrete sequences are acquired at the user-side incoming line end, and the voltage and current discrete sequences are denoised, periodically sliced, and time-domain aligned to obtain time-domain aligned voltage and current discrete sequences. Step 2: Obtain hidden layer feature tensors based on collaborative splitting convolutional networks; Construct a multi-channel VI physical trajectory tensor based on time-domain aligned discrete sequences of voltage and current; The multi-channel VI physical trajectory tensor is input into the edge sub-network of the collaborative split convolutional network deployed on the edge monitoring node; the collaborative split convolutional network is a physical prior gated multi-scale split convolutional network. Multi-scale trajectory features are extracted by multi-scale convolutional branches in the collaborative split convolutional network, and the outputs of the multi-scale convolutional branches are weighted and fused by the physical prior gating fusion module in the collaborative split convolutional network to obtain the hidden layer feature tensor. Step 3, calculate the prototype offset; The edge monitoring node only uploads the hidden layer feature tensor to the cloud server. The cloud server uses the cloud-side sub-network to output the load class probability and calculates the prototype offset of the hidden layer feature tensor relative to the predicted class prototype center based on the class prototype library. The prototype offset , is the normalized hidden layer feature tensor With the Prototype Center of Predictive Categories The Euclidean distance between them; Step 4: Calculate the joint loss and the gradient at the split point; When the maximum value of the load category probability is lower than the dynamic safety threshold, or the prototype offset is higher than the drift threshold, the cloud server calculates the joint loss based on the cross-entropy loss, prototype drift loss, and knowledge retention regularization term. The collaborative split convolutional network is divided into edge-side sub-networks and cloud-side sub-networks according to the network split points; the cloud server generates the split point gradient corresponding to the updatable parameters of the edge-side sub-network based on the joint loss at the network split points between the edge-side sub-networks and the cloud-side sub-networks. Step 5: Update the parameters of the edge sub-network based on the gradient of the split point, and implement non-intrusive load monitoring.
2. The non-invasive load monitoring method according to claim 1, characterized in that, The multi-channel VI physical trajectory tensor includes at least trajectory density channel data, orientation gradient channel data, transient transition channel data, and harmonic phase weight channel data; Among them, the trajectory density channel data is obtained by accumulating the normalized VI trajectory coordinate density, the orientation gradient channel data is obtained by encoding the motion direction of adjacent sampling points in the VI plane, the transient transition channel data is obtained by marking sampling points whose current mutation exceeds the adaptive threshold, and the harmonic phase weight channel data is obtained by mapping the fundamental phase shift and higher harmonic energy.
3. The non-invasive load monitoring method according to claim 1, characterized in that, The multi-scale convolution branches include local trajectory convolution branches, closed-shape convolution branches, and cross-region convolution branches; wherein, the local trajectory convolution branches use 3×3 convolution kernels, the closed-shape convolution branches use 5×5 convolution kernels, and the cross-region convolution branches use 3×3 dilated convolution kernels with a dilation rate of 2.
4. The non-invasive load monitoring method according to claim 1, characterized in that, The physical prior gating fusion includes: constructing a physical prior vector from the VI trajectory closure area, fundamental phase offset, current transient slope peak, harmonic distortion index, and trajectory sparsity; generating a gating vector based on the physical prior vector; and performing weighted fusion on the outputs of each multi-scale convolution branch.
5. The non-invasive load monitoring method according to claim 4, characterized in that, The gate vector satisfy ; in: The Sigmoid activation function satisfies This is used to restrict each element of the gate vector to between 0 and 1; Represents the gate weight matrix. Represents the gate bias vector; Represents the area of the VI locus closure, and according to Calculation, where Let be the coordinates of the nth normalized VI trajectory, where and Let N represent the nth normalized current coordinate and the nth normalized voltage coordinate, respectively; the VI trajectory within one power frequency cycle satisfies the start-end closure condition, i.e., the end sampling point is connected to the start sampling point; N is the number of sampling points in one power frequency cycle, and... ; This indicates the fundamental phase shift. Indicates the peak value of the transient slope of the current. Indicates harmonic distortion index; Represents the sparsity of the trajectory, and according to Calculation, where The number of non-zero accessed grid cells in the trajectory density matrix is M×M, where M represents the total number of grid cells in the trajectory density matrix. M represents the number of grid cells divided along the normalized voltage axis and the normalized current axis, so the trajectory density matrix contains M×M grid cells.
6. The non-invasive load monitoring method according to claim 1, characterized in that, The edge-side sub-network is provided with a low-rank drift adaptation layer, which includes low-rank drift adaptation parameters, including low-rank adaptation matrices A and B, such that... Construct a low-rank increment matrix; in, , r is less than the edge-side stable convolution parameter The smaller of the corresponding matrix dimensions m and n; The low-rank adaptation layer stabilizes the convolution parameters on the edge sides based on low-rank adaptation matrices A and B. Perform low-rank increment adjustment to obtain the equivalent edge-side convolution parameters. The low-rank adaptation layer satisfies: ; The initial values of A and B are obtained from offline training samples through backpropagation. During the online continuous learning phase, the cloud server calculates the gradients of A and B based on the joint loss and sends them to the edge monitoring nodes. A and B are iteratively updated through gradient descent, so that... Forming a pair Low-rank increments; priority updates during the online continuous learning phase. , And physical prior gating parameters, including the gating weight matrix. and gated bias vector Freeze or weak update .
7. The non-invasive load monitoring method according to claim 1, characterized in that, The joint loss L satisfies: ; in: Cross-entropy loss; ;in, For the unique thermal encoding of the actual load label, The load category probability output by the cloud-side sub-network; k is the load category index, k=1,2,…,K, where K is the total number of known load categories; the one-hot encoding of the actual load label. Satisfy: If the first If each category is a true category, then ,otherwise ; The first output of the cloud-based side network The probability of each load category; Hidden layer feature tensor With Category Prototype Center Prototype drift loss between; To preserve regular terms for knowledge, where The parameter set consists of the current low-rank drift adaptation parameters and the physical prior gating parameters; This is the set of corresponding parameters saved before this online update; and The weight coefficients for prototype drift loss and knowledge retention regularization are determined through a grid search on the validation set. , .
8. The non-invasive load monitoring method according to claim 1, characterized in that, When the cloud server generates the gradient of the segmentation point, it calculates the gradient corresponding to the physical prior gating parameter, low-rank drift adaptation parameter and multi-scale convolution parameter selected by the cloud according to the drift degree based on the hidden layer feature tensor, physical prior vector, edge sub-network topology and prototype offset of the edge sub-network in the current sample. After grouping and serializing according to parameter type, the gradient is sent to the edge monitoring node.
9. The non-invasive load monitoring method according to claim 8, characterized in that, The edge monitoring node is equipped with a grouped asynchronous gradient buffer, which includes a gating parameter buffer, a low-rank drift adaptation parameter buffer, and a convolution parameter buffer. The gating parameter buffer and the low-rank drift adaptation parameter buffer are updated first, while the convolution parameter buffer is only enabled when the prototype offset exceeds the drift threshold for multiple consecutive rounds.
10. A non-intrusive load monitoring system based on edge-cloud collaborative splitting, characterized in that, For implementing the method according to any one of claims 1 to 9; The system includes edge monitoring nodes and a cloud server. The edge monitoring nodes include a high-frequency sampling module, a physical trajectory construction module, an edge-side sub-network module, a grouped asynchronous gradient buffer, and an edge communication module. The cloud server includes a cloud-side sub-network module, a category prototype library, a split-point gradient generation module, and a cloud communication module.