Non-intrusive load dynamic identification system and method based on multi-modal transfer learning

The non-intrusive load dynamic identification system, which utilizes multimodal perception and transfer learning, combined with multimodal data capture and edge computing, solves the bottlenecks of cross-domain migration and dynamic load identification. It achieves rapid adaptation, low power consumption, and high accuracy in load identification, overcoming the problems of difficult cross-domain migration, poor dynamic load identification, and energy efficiency costs associated with traditional NILM.

CN121167386APending Publication Date: 2025-12-19WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511196515.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing non-intrusive load monitoring technologies have bottlenecks in cross-domain migration, dynamic load resolution, and real-time energy consumption, making it difficult to dynamically identify complex operating loads under conditions of 'zero labeling, real-time operation, and low power consumption'.

Method used

A non-intrusive load dynamic identification system employing multimodal sensing, transfer learning, and edge computing captures multimodal data through coils, microphone arrays, accelerometers, and antenna arrays. It utilizes multimodal transfer learning networks for feature encoding and graph attention networks for load identification. Combined with IEEE 1588 synchronization technology and quantized sensing training, it achieves low-power edge deployment and incremental model optimization.

Benefits of technology

It enables rapid adaptation to new environments without manual annotation, improves cross-scenario robustness and recognition accuracy, lowers the threshold for engineering deployment, reduces the amount of training samples and energy consumption, and meets the needs of long-term online operation of large buildings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167386A_ABST
    Figure CN121167386A_ABST
Patent Text Reader

Abstract

The invention discloses a non-intrusive load dynamic identification system and method based on multi-mode transfer learning. The system comprises a coil, a diverter, a microphone array, an accelerometer and an antenna array. The coil and the diverter are wound outside a bus of a power distribution cabinet in a non-intrusive manner and are used for capturing waveforms of three-phase current I (t) and voltage V (t); the microphone array is mounted on the inner side of a cabinet door of the power distribution cabinet and is used for recording a Mel spectrum of start-stop noise S (t, f); the accelerometer is arranged on the power distribution cabinet body and is used for capturing a micro-vibration vector a (t) = [ax, ay, azz]; the antenna array is used for capturing a radiation pulse E (t) at the moment of on-off of the equipment. According to the method, at a cloud end, active learning-driven increment fine tuning and knowledge graph maintenance ensure that the recognition accuracy can reach 95% or above 30 seconds after system'zero labeling 'is online, and meanwhile, INT8 quantitative reasoning enables edge power consumption to be reduced by 40% or above.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence and smart grid, and relates to a non-intrusive load dynamic identification system and method, in particular to a non-intrusive load dynamic identification system and method based on multi-modal perception, transfer learning and edge computing. BACKGROUND

[0002] Non-intrusive load monitoring (NILM) splits the energy consumption of each electrical device by analyzing the current or power sequence of the main line of power distribution, but the traditional single-modal NILM is highly sensitive to the domain offset of building structure, access to power grid topology and electrical appliance generation characteristics, and often needs to be re-labeled and trained for cross-scene deployment, which takes several weeks.

[0003] Existing research attempts to combine transfer learning with sparse coding, convolutional network and other models to shorten the parameter adjustment time, but is still limited to a single electrical signal and cannot capture the sound-vibration-electromagnetic characteristics at the start-stop moment, resulting in insufficient identification accuracy of dynamic loads such as variable frequency air conditioners and elevators. At the same time, the Transformer architecture is introduced into NILM due to its long-range dependence modeling capability, but it still lacks explicit constraints on cross-domain distribution differences; although MMD and other metrics can measure the distribution difference between domains, they ignore the category prior bias. In addition, multi-modal sensing nodes are difficult to strictly align due to clock drift, and PTP-IEEE 1588 synchronization provides a reference with <1 µs precision for industrial Internet of Things, but has not been combined with NILM.

[0004] In summary, the existing technology cannot complete the dynamic identification of complex working condition loads under the conditions of "zero labeling, real-time and low power consumption", and breakthroughs are needed. SUMMARY

[0005] The purpose of the present application is to solve the core bottlenecks of existing NILM in cross-domain transfer, dynamic load resolution and real-time energy consumption, and to provide a non-intrusive load dynamic identification system and method based on multi-modal perception, transfer learning and edge computing.

[0006] The technical scheme of the system of the present application is: a non-intrusive load dynamic identification system based on multi-modal transfer learning, characterized by: comprising a coil, a shunt, a microphone array, an accelerometer and an antenna array. The coil and the shunt are non-intrusively wrapped around the outside of the bus of the power distribution cabinet for capturing three-phase current I(t) and voltage V(t) waveforms; the microphone array is installed on the inside of the cabinet door of the power distribution cabinet for recording the mel spectrum of the start-stop noise S(t,f); the accelerometer is arranged on the cabinet body of the power distribution cabinet for capturing the micro-vibration vector a(t)=[a x ,a y ,a zThe antenna array is used to capture a radiation pulse E(t) at a device on-off moment.

[0007] The method of the application adopts the technical scheme of: a non-invasive load dynamic identification method based on multi-modal transfer learning, comprising the following steps: Step 1: Collect multi-modal data at a preset frequency under the same time axis, including three-phase current I(t) and voltage V(t) waveforms, mel spectrum of start-stop noise S(t,f), micro-vibration vector a(t) = [a x ,a y ,a z ] and a radiation pulse E(t) at a device on-off moment; Step 2: Each modal data completes denoising and standardization processing in an independent thread, and outputs an equal-length TxC tensor; wherein T represents the sampling frame number in the data time window corresponding to each mode, reflecting the time sequence length of the signal; C represents the feature channel number in each frame, and different modes correspond to different dimensions: Step 3: Feature encoding is performed on the tensor based on a multi-modal transfer learning network; The multi-modal transfer learning network comprises: a 1-D convolution scattering module, which is used to divide an input TxC tensor into a Patch sequence with a length of S through one-dimensional convolution; a position encoding module, which is used to add sine and cosine position encoding pos to the divided small Patch to retain time sequence information; a core encoder module, which adopts a four-layer stacked time-frequency Vision-Transformer as a main Transformer-Encoder architecture to extract global context information; wherein each layer inserts an Adapter module A(•) with a parameter proportion of only X%, which is responsible for task-specific fine-tuning, and only the weights of 0.0XParam base need to be updated to complete the transfer, Param base represents the total amount of all trainable parameters of the frozen main Transformer-Encoder before transfer learning, and X is a preset value; and finally, the feature encoding is output through a multi-head attention layer; a multi-head attention aggregation module is used to perform attention weighted fusion at the end of the feature encoding, and an ultimate modal embedding vector representation is output; Step 4: The time sequence embedding output by the encoding is aggregated into nodes through a sliding window, and adjacent nodes are associated into three types of edges of "steady state, transient state and abnormality" according to the power slope, harmonic mutation and acoustic vibration pulse, so as to form a heterogeneous time sequence graph G(V,E); a graph attention network dynamically learns the weight a ijAnd the neighborhood information is weighted and aggregated, and finally the "device-state-power" triplets and real-time power sequence are output .

[0008] As preferred, in step 1, the system adopts IEEE 1588 Precision Time Protocol, estimates and compensates the propagation delay in the two-way measurement of the Sync message issued by the master clock and the Delay-Req / Resp returned by the slave node, ensures that the clock deviation εsync of each node is less than the preset value, and writes the hardware timestamp into the ring buffer; the edge gateway packages the data frame with w points into a frame and uploads it by encapsulating it with UDP-PTP, the frame sequence number n and the nanosecond-level timestamp τ n . n ) are constituted.

[0009] As preferred, in step 2, for the electrical signal, the cut-off frequency of the Butterworth low-pass filter F(s) is fc, and the filter order N, and then the Hamming window with a window length L is slid by half the window length to obtain the amplitude and phase matrix E STFT ; the acoustic spectrum is subjected to spectrum subtraction to eliminate the background machine room noise, and then the Mel filter bank energy is calculated; the vibration vector is subjected to Kalman filtering to extract the Hilbert envelope; the UWB pulse signal is subjected to single pulse energy P UWB .

[0010] As preferred, in step 3, the multi-head attention layer defines the input key, query, and value vectors as (K, Q, V), and the multi-head attention is: ; wherein d is the hidden dimension; the Adapter module A(•) is first projected downward to dr, activated by ReLU, and then projected upward back to d, and the output is obtained by element-wise addition with the residual of the sublayer.

[0011] As preferred, in step 4, ; wherein is the MMD Gaussian kernel width, and the adaptation is half of the median of the source-target batch inter-Euclidean distance; a is a trainable attention weight vector; Wh i is the linearly transformed feature of node i; is the adjacency set of node i.

[0012] As preferred, in step 3, the multi-modal transfer learning network is a trained network; the loss function L used in the training is: L= Ly + L mmd +Ladv ; wherein Ly is the main task cross-entropy; the kernel width adaptive maximum mean difference MMD(Xs,Xt) is taken as the measure loss L mmd ;L adv is the domain adversarial loss; β, λ balance the measure and the adversarial strength respectively; A domain discriminator is introduced in the training The gradient inversion layer Rλ is used in the training, and the inversion layer is multiplied by-λ in the back propagation to reverse the domain label gradient, so that the feature extractor learns the domain-inseparable features.

[0013] As a preferred, in the training, the edge side enables the quantization perception training, and the weights and activations are inserted into the pseudo quantization operation and mapped to the INT8 representation; in the inference, the TensorRT-QDQ pipeline is used to run the convolution-matrix multiplication unit in the INT8; the cloud GPU cluster retains the FP16 backbone for batch fine-tuning and knowledge base updating, and differentials weights Δθ are issued through the gRPC-TLS channel.

[0014] As a preferred, in the training, the system continuously monitors the confidence γ of the inference output, when γ is less than the threshold, the edge gateway automatically uploads the corresponding window segment, and the cloud selects the most representative and information amount data points from the uploaded data in the core-boundary sampling strategy to request artificial one-time labeling; then the system uses the 8-batch Mini-Replay incremental fine-tuning mechanism to quickly update the Adapter layer, wherein the 8-batch Mini-Replay incremental fine-tuning mechanism refers to that the newly added labeled samples and the historical memory samples are combined according to a fixed proportion to form 8 small batches, each batch contains the current newly labeled data and the labeled samples sampled from the long-term memory buffer; the system continuously performs forward propagation, back propagation and parameter update on the 8 batches to realize small-scale but efficient parameter fine-tuning.

[0015] As a preferred, in the training, the cloud stores the four-layer knowledge graph of "building-distribution area-equipment-working condition" in the graph database, and periodically calculates the KL-divergence to detect concept drift; when the drift degree D KL is greater than the threshold or the recognition index decreases by more than the threshold for three consecutive days, the "degradation" label is automatically triggered and a retraining task is generated; the differential update replaces the edge model within Y ms through the general flash mechanism, ensuring that there is no need for maintenance downtime, wherein Y is a preset value.

[0016] Compared with the existing non-intrusive load monitoring (NILM) technology, the present application has achieved a breakthrough in cross-scene robustness, recognition accuracy, system deployment efficiency and energy consumption control, and has the following beneficial effects: 1. "Zero-labeled" online, quickly adapt to new environment: The multi-modal transfer learning network proposed in the application combines the adversarial-metric joint training strategy, so that the model can complete initialization without manual labeling in the new building environment, and reach ≥95% accuracy within 30 seconds of online, significantly reducing the engineering deployment threshold.

[0017] 2. Multi-modal fusion enhances load discrimination ability: Compared with traditional electrical waveform driven NILM method, the application introduces auxiliary modalities such as sound, vibration and electromagnetic pulse, effectively distinguishes complex loads such as variable frequency air conditioners and elevators, improves dynamic load identification accuracy, and F1-score is improved by 18% compared with single modal method.

[0018] 3. Low-power edge deployment, significant energy saving: Through quantization perception training (QAT) and INT8 inference deployment of TensorRT-QDQ pipeline, the application reduces the inference power consumption of the edge gateway by 42%, meeting the energy consumption requirements of long-term online operation of large buildings.

[0019] 4. Incremental fine-tuning mechanism driven by active learning: Using 8-batch Mini-Replay strategy, combined with core-boundary sampling to select the most informative samples, only a small amount of labeling is required to maintain high accuracy, reducing the training sample size by 70%, achieving low-cost and high-efficiency model self-adaptive update.

[0020] 5. Continuously evolving knowledge graph driven governance: The system maintains the cloud knowledge graph and concept drift detection mechanism, realizes long-term learning and automatic updating of the dynamic relationship between "device-state-power", and significantly improves the model stability across cycles and across buildings.

[0021] In summary, the application has significant technical advantages in the five key dimensions of "low-labeled deployment", "multi-modal fusion identification", "edge energy-saving inference", "incremental model optimization" and "knowledge-driven autonomy", breaking through the core bottlenecks of traditional NILM such as cross-domain transfer difficulty, poor dynamic load identification and high energy cost, with outstanding technical advancement and industrial application value. BRIEF DESCRIPTION OF DRAWINGS

[0022] The technical solutions of the application are further illustrated below using examples and specific embodiments. In addition, some drawings are also used in the process of explaining the technical solutions. For those skilled in the art, other drawings and the intent of the application can also be obtained from these drawings without creative labor.

[0023] Figure 1 The system working principle diagram of the present embodiment is shown in the figure; Figure 2 The method flowchart of the present embodiment is shown in the figure; Figure 3 A multi-modal transfer learning network training schematic for the present embodiment. DETAILED DESCRIPTION

[0024] For the person skilled in the art to understand and implement the present application, the present application will be further described in detail below in conjunction with the drawings and examples. It should be understood that the embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0025] See Figure 1 The present embodiment provides a multi-modal transfer learning based non-intrusive load dynamic identification system, which comprises a coil, a shunt, a microphone array, a six-degree-of-freedom accelerometer, and an antenna array. The coil and the shunt are non-intrusively wrapped outside the busbar of the power distribution cabinet for capturing three-phase current I(t) and voltage V(t) waveforms; the microphone array is installed inside the cabinet door for recording the mel spectrum of start-stop noise S(t,f); the six-degree-of-freedom accelerometer is arranged on the cabinet body for capturing micro-vibration vector a(t)=[a x ,a y ,a z ]; and the antenna array is used to capture the radiation pulse E(t) at the moment of device on-off.

[0026] In one embodiment, first, a Rogowski coil and a 0.1Ω shunt are non-intrusively wrapped outside the busbar of the power distribution cabinet to capture three-phase current I(t) and voltage V(t) waveforms at a sampling frequency of 10 kHz; second, a 48 kHz MEMS microphone array is installed inside the cabinet door to record the mel spectrum of start-stop noise S(t,f); third, a six-degree-of-freedom accelerometer is pasted on the cabinet body steel plate to obtain micro-vibration vector a(t)=[a x ,a y ,a z ] at a sampling frequency of 20 kHz; and finally, a 3.1-10.6 GHz UWB antenna array is arranged on the top of the machine room to capture the radiation pulse E(t) at the moment of device on-off. This does not only inherit the simplicity of deployment of single-point NILM, but also compensates for similar loads that are difficult to distinguish by electrical waveforms through acoustic-vibration-electromagnetic information.

[0027] See Figure 2 The present embodiment provides a multi-modal transfer learning based non-intrusive load dynamic identification method, which comprises the following steps: Step 1: Collect multi-modal data at preset frequency on the same time axis, including three-phase current I(t) and voltage V(t) waveforms, mel-spectrogram of start-stop noise S(t,f), micro-vibration vector a(t) = [a x y z ] and radiation pulse E(t) of device on-off transient; In an embodiment, multi-modal data are aligned on the same time axis, the system adopts IEEE 1588 Precision Time Protocol, estimates and compensates propagation delay in two-way measurement of Delay-Req / Resp returned by slave node from Sync message issued by master clock, ensuring that clock deviation εsync≤1μs of each node. The node writes data into a ring buffer with hardware timestamp, and the edge gateway packages data frame with length w=1024 points into UDP-PTP encapsulation for uploading, frame sequence number and 64-bit nanosecond-level timestamp τ n constitute tuple (n,τ n ).

[0028] Step 2: Each modal data completes denoising and standardization processing in an independent thread, and outputs equal-length TxC tensor; wherein T represents the number of sampling frames in the data time window corresponding to each modal, reflecting the time sequence length of the signal; C represents the number of feature channels in each frame, different modalities correspond to different dimensions: For electrical waveforms, C represents the frequency domain amplitude and phase vector dimension after short-time Fourier transform (STFT); For sound spectrum, C represents the number of mel filter banks, typically 128; For vibration signals, C represents the three-axis vector dimension of Hilbert envelope after Kalman filtering; For pulse signals, C represents the number of energy features after Gabor wavelet packet decomposition.

[0029] The unified aligned TxC tensor ensures that each modal can be processed in parallel in the subsequent multi-modal encoding stage, and keeps the time dimension consistent, facilitating fusion modeling.

[0030] In an embodiment, after reaching the gateway, each modal first completes denoising and standardization in an independent thread. For electrical signals, the cutoff frequency of the Butterworth low-pass filter F(s) is fc=5kHz, and the filter order N=4, then the amplitude and phase matrix E STFT ​​The acoustic spectrum is subtracted from the background machine room noise by spectrum subtraction, and the 128-dimensional Mel filter bank energy is calculated again; the vibration vector enters the Kalman filter to extract the Hilbert envelope; the UWB pulse signal is calculated by the Gabor wavelet basis function to obtain the single pulse energy P UWB These processes uniformly compress each modality into an equal-length TxC tensor, facilitating subsequent parallel encoding.

[0031] Step 3: Feature encoding of the tensor based on the multi-modal transfer learning network; The multi-modal transfer learning network includes: a 1-D convolutional scattering module for dividing the input TxC tensor into small patch sequences with a length of S through one-dimensional convolution; a position encoding module for adding sine and cosine position encoding pos to the divided small patch to preserve the time sequence information; a core encoder module adopting a four-layer stacked time-frequency Vision-Transformer as the main Transformer-Encoder architecture for extracting global context information; wherein each layer inserts an Adapter module A(•) with a parameter ratio of only 3% to perform task-specific fine-tuning, which only needs to update 0.03xParam base weights to complete the transfer, Param base represents the total number of trainable parameters of the frozen main Transformer-Encoder before transfer learning, and X is a preset value; finally, the feature encoding is output through a multi-head attention layer; a multi-head attention aggregation module is used to perform attention weighted fusion at the end of the feature encoding to output the final modality embedding vector representation; In one embodiment, the multi-head attention layer defines the input key, query, and value vectors as (K, Q, V), and the multi-head attention is: ; Where d=512 is the hidden dimension, and the number of heads h=8; the Adapter module A(•) is first projected down to dr=64, activated by ReLU, and then projected back up to d, and added element-wise to the residual of the sublayer (i.e. the input vector of the aforementioned multi-head attention sublayer) to output.

[0032] Step 4: The encoded time sequence embedding z1:S is aggregated into nodes with a 50 ms sliding window, and adjacent nodes are associated as "steady, transient, and abnormal" edges according to the power slope, harmonic mutation, and acoustic vibration pulse, forming a heterogeneous time sequence graph G(V, E); the graph attention network calculates the weight α ijand weightedly aggregate neighborhood information, finally output "device-state-power" triplets with real-time power sequence .

[0033] In an embodiment, ; wherein, is the Gaussian kernel width in MMD, adaptive is half of the source-target batch inter-Euclidean distance median; a is a trainable attention weight vector with dimension 2d', which is specifically used to measure the relevance of node i and neighbor j feature combination, through inner product compress the high-dimensional spliced vector into a scalar score e ij This vector is shared within the entire layer, and all edges in the same layer reuse it. Wh i is the linearly transformed feature of node i, h i ∈R d is the original input feature; W∈R d′×d is the weight matrix shared within the layer, used to project the original dimension d to the new hidden dimension d'=d / h=64 (the present invention sets d=512, h=8). is the adjacency set of node i, in the "dynamic load behavior graph" of the present invention, it is automatically generated by the power slope, harmonic mutation, and sound-vibration impulse among 50 ms sliding window nodes; thus N(i) changes with time rolling, used for softmax normalization, making .

[0034] In an embodiment, please see Figure 3 , the multi-modal transfer learning network is a well-trained network; to maintain the discriminative ability in the "zero-labeled" environment of new buildings, the model introduces a domain discriminator and a gradient reversal layer Rλ. The reversal layer multiplies -λ in backpropagation to reverse the domain label gradient, forcing the feature extractor to learn domain-inseparable features; at the same time, the maximum mean difference MMD(Xs,Xt) with kernel width adaptation is used as the metric loss L mmd . The loss function L used in training is: L= Ly + L mmd +L adv ; wherein, Ly is the main task cross-entropy; the maximum mean difference MMD(Xs,Xt) with kernel width adaptation is used as the metric loss L mmd ; L advFor domain adversarial loss, the feature extractor is constrained to learn a common representation that is "domain-inseparable yet discriminative" between source buildings (labeled) and target buildings (zero-labeled). This loss, in conjunction with the gradient reversal layer, enables a game-theoretic training between the feature extractor and the domain discriminator; β and λ balance the metric and adversarial strength, respectively, and β = 0.3 and λ = 0.1 are taken in the experiment, which can effectively suppress feature collapse.

[0035] In an implementation, in the multi-modal transfer learning network training, the edge side enables Quantization-Aware Training, inserts the weight and activation into a Fake-Quantization operation (in forward propagation, the combination of "quantization-dequantization" is used to approximate the INT8 effect with floating point, and the differentiability of gradient calculation is retained.) and then mapped to the INT8 representation (8-bit integer data format. It maps the original 32-bit floating point (FP32) storage and calculation of weights and activation values to a discrete integer interval that can be represented by only 1 byte, thereby significantly reducing the model size, memory bandwidth and multiply-accumulate power consumption, while maintaining the accuracy close to FP32 with the help of QAT. Signed INT8 (default for deep learning) uses binary two's complement coding, which can express [-128, 127] [-128, 127] [-128, 127] a total of 256 discrete values); during inference, the TensorRT-QDQ pipeline is used to run the convolution-matrix multiplication unit entirely in INT8, keeping the accuracy degradation <0.3% while controlling the single-frame delay to 11 ms, and reducing power consumption by 42%; the cloud GPU cluster retains FP16 (the core network is still trained, fine-tuned and inferred in the cloud in 16-bit half-precision floating point (FP16) format, rather than being further compressed to INT8 like the edge side. This does both speed and accuracy, and provides a unified high-precision master version for subsequent INT8 differential weight pushing to the edge side) backbone is used for batch fine-tuning and knowledge base updating, and differential weight Δθ is delivered through a gRPC-TLS channel ("gRPC-TLS channel" refers to a two-way flow channel that uses gRPC protocol, is carried on HTTP / 2 long connection and uses TLS encryption authentication).

[0036] The quantization-aware training adopted in this embodiment is a training technique that compresses model weights and activations from FP32 to INT8 and other low bit-width formats while maintaining accuracy. It "simulates" quantization errors during inference by inserting fake-quant nodes, performing truncation, rounding, and counting quantization intervals during the training phase, so that the gradient updates can "see" these errors in real time, and thus learn a weight distribution and scale parameter that is robust to quantization. Compared with post-training quantization (PTQ), which only performs offline rounding after training, QAT can reduce model size, memory usage, and multiply-add power consumption, and significantly reduce edge inference latency, without sacrificing accuracy.

[0037] In an embodiment, in the multi-modal transfer learning network training, the system continuously monitors the confidence γ of the inference output. When γ < τ = 0.6, the edge gateway automatically uploads the corresponding window segment, and the cloud selects the most representative and information-rich data points from the uploaded data using the core-boundary sampling strategy and requests artificial one-time labeling. Subsequently, the system uses an 8-batch Mini-Replay incremental fine-tuning mechanism to quickly update the Adapter layer, wherein the 8-batch Mini-Replay incremental fine-tuning mechanism refers to grouping the newly added labeled samples and historical memory samples in a fixed proportion into 8 small batches, each batch containing current newly labeled data and labeled samples sampled from the long-term memory buffer. The system performs forward propagation, back propagation, and parameter update on the 8 batches in succession to achieve small-scale but efficient parameter fine-tuning.

[0038] This mechanism significantly reduces the cost of artificial labeling and training resource consumption while maintaining the discriminative ability of the model, achieving a 70% reduction in sample size while maintaining an accuracy of more than 95%.

[0039] In an embodiment, the 8-batch Mini-Replay is a minimal replay iteration strategy used in the active learning fine-tuning stage: after receiving one-time labeled core-boundary samples, the cloud does not perform complete retraining, but only uses 8 consecutive mini-batches for rapid replay training. Each batch mixes (1) newly labeled samples and (2) "old" samples uniformly sampled from the long-term memory buffer in a fixed proportion to perform 8x 1 forward-backward-parameter updates. In this way, the advantages of replay in preventing forgetting in incremental learning are fully utilized, and the cloud GPU computation is controlled at the level of tens of milliseconds, meeting the real-time requirements of "edge-cloud hot updates".

[0040] In an embodiment, in the multi-modal transfer learning network training, the cloud stores a four-layer knowledge graph of "building-distribution area-equipment-working condition" in a graph database Neo4j, periodically calculates KL-divergence to detect concept drift, and automatically triggers a "degradation" label and generates a retraining task when the drift degree D KL >0.1 or when the recognition index F1-score (weighted average of accuracy and recall) decreases by more than 3% for three consecutive days, an automatic "degradation" label is triggered and a retraining task is generated; differential update replaces the edge model within 50 ms through a general flash mechanism, ensuring no downtime maintenance.

[0041] The application is further described below through specific experiments.

[0042] In the experiment, the data sources include: (1) Source domain (pre-training): public NILM datasets REFIT ElectricLoad & UK DALE, a total of 5 residential buildings, about 2 years of multi-modal (current, voltage, sound spectrum, vibration) records.

[0043] (2) Target domain (test): a 9-story office building in Wuhan City is equipped with the sensor system of the application, and 2h of continuous data is collected, of which the first 30s is used as a "zero-label" starting window, and the rest of the period is used for performance evaluation.

[0044] Baseline comparison: single-modal current / voltage waveform data is collected synchronously at the same location, and a mainstream CNNLSTM NILM model is used as the baseline method.

[0045] The experimental process includes: (1) Pre-training of multi-modal transfer learning network with REFIT+UK DALE; (2) Target building online: multi-modal data enters the denoising, standardization, and tensorization process in real time; (3) Zero-label stage (first 30s): freeze the backbone, only inference; record the initial accuracy; (4) Active learning stage: trigger upload when the confidence γ<0.6; the cloud uses core boundary sampling, and manually annotates less than 0.3% of the segments at one time; (5) 8 batch Mini Replay: only fine-tune the Adapter layer; (6) Continuous 2h running, statistics of recognition index, inference delay, and power consumption.

[0046] The evaluation indexes include: (1) F1 score, Accuracy (ACC) (2) Mean Power Error (MPE) (3) Average Latency of Edge Inference (ms) (4) Average Power Consumption of Edge Inference (W) In the experiment, the electrical transient characteristic length: ; where f s is the sampling frequency (Hz), i.e. the number of sample points recorded per second by the ADC; T cycle is the time length of a single power frequency cycle of the power grid, T cycle = 20 ms when fgrid= 50 Hz; f grid is the power frequency of the power grid (Hz, the method takes 50 Hz); gcd(f s , f grid ) represents the greatest common divisor (Hz) of f s and f grid , which is used to ensure that the denominator is divisible, and let N be an integer.

[0047] Example: when f s = 10 kHz, fgrid= 50: ; ; If the fs= 20 kHz of the present scheme is used (i.e. "10 kHz x 2 times oversampling" is performed on the current / voltage branch), then ; in this way, 400-point transient characteristics aligned with the whole cycle can be obtained, providing consistent input length for subsequent STFT and patch scattering.

[0048] In the experiment, the filter group delay is 0.18 ms, which does not affect the synchronization accuracy. N tap is the filter tap number, i.e. the FIR coefficient length. Take 8 (corresponding to the equivalent 8-tap symmetric FIR after window transformation of the 4th Butterworth low-pass). The Gaussian kernel width σ in MMD is adaptively set to half of the median of the Euclidean distance between the source and target batches: , which can be dynamically updated during training to prevent under-fitting or over-fitting.

[0049] The experimental results show that: (1) The present invention can reach 93.1% F1 score within 30s of zero labeling; after 8 batch fine-tuning, it rises to 95.3%, which is 18.2pp higher than the single-modal baseline.

[0050] (2) The tracking error of dynamic load (elevator, variable frequency air conditioner) is reduced from 9.4% to 3.2%.

[0051] (3) Active learning reduces the target domain annotation amount by about 70% (only 0.3%) while maintaining high accuracy.

[0052] (4) INT8 inference power consumption is reduced from 4.0W to 2.3W (42%), meeting the building 24x7 operation requirements.

[0053] Experiments verify the significant advantages of the present application in cross-domain migration, dynamic load identification and low-power deployment.

[0054] From the experimental results, it can be seen that the present application reaches ≥95% recognition accuracy within 30 seconds of starting in a new building with "zero annotation", reduces the training sample requirement by 70% compared to traditional methods, and reduces the edge inference power consumption by ≥40%.

[0055] From the experimental results, it can be seen that compared with single-modal NILM, the present application can reach 95.3% F1-score in a new building without annotation within 30 seconds, with an accuracy improvement of 18%, a training sample requirement reduction of 70%, and a variable frequency device energy curve tracking error reduction from 9.4% to 3.2%. Edge quantization inference reduces power consumption from 4W to 2.3W, saving 42% energy, meeting the 24x7 operation requirements of large buildings. Multi-modal fusion and adversarial-metric migration significantly improve cross-domain robustness.

[0056] It should be understood that the above-described embodiments are part of the embodiments of the present application, not all of the embodiments. In addition, the technical features of each embodiment or single embodiment provided by the present application can be combined with each other to form a feasible technical solution, and such combination is not restricted by the order of steps and / or structure composition mode, but must be based on the realization by ordinary skilled in the art, when the combination of technical solutions appears contradictory or cannot be realized, it should be considered that the combination of technical solutions does not exist, and is not within the protection scope of the present application.

[0057] It should be understood that the above description of the preferred embodiments is more detailed, and therefore should not be considered as a limitation on the scope of patent protection of the present application. Ordinary skilled in the art can make substitutions or modifications without departing from the scope of the claims of the present application, which fall within the scope of protection of the present application. The scope of protection of the present application should be subject to the appended claims.

Claims

1. A non-intrusive load dynamic identification system based on multimodal transfer learning, characterized in that: Includes coils, shunts, microphone arrays, accelerometers, and antenna arrays; The coil and shunt are non-intrusively wrapped around the outside of the distribution cabinet busbar to capture the waveforms of the three-phase current I(t) and voltage V(t); the microphone array is installed inside the distribution cabinet door to record the Mel spectrum of the start-stop noise S(t,f); the accelerometer is mounted on the distribution cabinet body to capture the micro-vibration vector a(t) = [a x ,a y ,a z The antenna array is used to capture the radiation pulse E(t) at the moment the device is switched on or off.

2. A non-intrusive load dynamic identification method based on multimodal transfer learning, applied to the system described in claim 1; characterized in that, Includes the following steps: Step 1: Acquire multi-mode data at a preset frequency along the same time axis, including the waveforms of three-phase current I(t) and voltage V(t), the Mel spectrum of start-stop noise S(t,f), and the micro-vibration vector a(t) = [a x ,a y ,a z The radiation pulse E(t) at the moment the equipment is switched on and off; Step 2: Denoising and standardization of each modality's data are performed in an independent thread, outputting a constant-length T × C tensor; where T represents the number of sampling frames within the data time window corresponding to each modality, reflecting the temporal length of the signal; C represents the number of feature channels in each frame, with different dimensions corresponding to different modalities. Step 3: Encode features of tensors based on a multimodal transfer learning network; The multimodal transfer learning network includes: a 1-D convolutional scattering module, used to divide the input T × C tensor into patches through one-dimensional convolution, resulting in a patch sequence of length S; a position encoding module, used to add sine and cosine position codes pos to the divided patches to preserve temporal order information; a core encoder module, using a four-layer stacked temporal-frequency domain Vision-Transformer as the backbone Transformer-Encoder architecture, used to extract global context information; and an Adapter module A(•), whose parameter insertion ratio is only X% in each layer, responsible for task-specific fine-tuning, requiring only an update of 0.0 X×Param. base The migration can be completed by adjusting the weights of Param. base X represents the total number of trainable parameters of the backbone Transformer-Encoder that are kept frozen before transfer learning, where X is a preset value; finally, the feature encoding is output after passing through a multi-head attention layer; the multi-head attention aggregation module is used to perform attention-weighted fusion at the end of the feature encoding and output the final modality embedding vector representation; Step 4: The temporal embeddings of the encoded output are aggregated into nodes through a sliding window, and adjacent nodes are associated with three types of edges: "steady-state, transient, and anomalous" based on power slope, harmonic abrupt changes, and acoustic vibration pulses, thus forming a heterogeneous temporal graph G(V,E); the graph attention network calculates the dynamically learned weight α of node i when receiving information from neighbor j at each node. ij It then weights and aggregates neighborhood information, ultimately outputting a "device-status-power" triplet and a real-time power sequence. .

3. The non-intrusive load dynamic identification method based on multimodal transfer learning according to claim 2, characterized in that: In step 1, the system uses the IEEE 1588 Precision Time Protocol to estimate and compensate for propagation delay in the bidirectional measurement process of the master clock sending Sync messages and the slave nodes returning Delay-Req / Resp messages, ensuring that the clock deviation εsync of each node is less than a preset value. Nodes write hardware timestamps into a circular buffer, and the edge gateway packages and uploads data frames of length w using UDP-PTP encapsulation. The frame sequence number n and the nanosecond-level timestamp τ are used for data transmission. n Constructing a tuple (n, τ) n ).

4. The non-intrusive load dynamic identification method based on multimodal transfer learning according to claim 2, characterized in that: In step 2, for the electrical signal, the cutoff frequency of the Butterworth low-pass filter F(s) is set to fc, and the filter order is N. Then, the Hamming window of length L is halved, and a short-time Fourier transform is performed to obtain the amplitude-phase matrix E. STFT ; Background noise from the computer room is eliminated by spectral subtraction, followed by calculation of the energy of the Mel filter group; the vibration vector is then subjected to Kalman filtering to extract the Hilbert envelope; the single-pulse energy P of the UWB pulse signal is calculated using Gabor wavelet basis functions. UWB .

5. The non-intrusive load dynamic identification method based on multimodal transfer learning according to claim 2, characterized in that: In step 3, the multi-head attention layer defines the input key, query, and value vector as (K, Q, V), and the multi-head attention is as follows: ; Where d is the hidden dimension; the Adapter module A(•) first projects down to dr, then up to d after ReLU activation, and finally adds it element by element to the residual of the sub-layer before outputting it.

6. The non-intrusive load dynamic identification method based on multimodal transfer learning according to claim 2, characterized in that: In step 4, ; in, is the Gaussian kernel width in MMD, adaptively set to half the median of the source-target inter-batch Euclidean distance; 'a' is the trainable attention weight vector; Wh i The feature of node i after linear transformation; Let i be the set of adjacencies of node i.

7. The non-intrusive load dynamic identification method based on multimodal transfer learning according to any one of claims 2-6, characterized in that: The multimodal transfer learning network mentioned in step 3 is a pre-trained network; the loss function L used during training is: L= Ly + L mmd +L adv ; Among them, Ly is the cross-entropy of the main task; the maximum mean difference MMD(Xs,Xt) with adaptive kernel width is used as the loss metric L. mmd L adv For domain adversarial loss; β and λ are the balance metrics and adversarial strength, respectively; Introducing a domain discriminator during training With the gradient inversion layer Rλ, the inversion layer multiplies by -λ during backpropagation to invert the gradient of the domain label, forcing the feature extractor to learn the domain-inseparable features.

8. The non-invasive load dynamic identification method based on multimodal transfer learning according to claim 7, characterized in that: During training, quantization-aware training is enabled on the edge side, and the weights and activations are mapped to INT8 representation after sham quantization. During inference, the TensorRT-QDQ pipeline is used to run all convolution-matrix multiplication units on INT8. The cloud GPU cluster retains the FP16 backbone for batch fine-tuning and knowledge base updates, and distributes differential weights Δθ through the gRPC-TLS channel.

9. The non-invasive load dynamic identification method based on multimodal transfer learning according to claim 7, characterized in that: During training, the system continuously monitors the confidence level γ of the inference output. When γ < the threshold, the edge gateway automatically uploads the corresponding window fragment, and the cloud selects the most representative and informative data points from the uploaded data using a core-boundary sampling strategy to request manual annotation once. The system then employs an 8-batch Mini-Replay incremental fine-tuning mechanism to rapidly update the Adapter layer. This mechanism involves grouping newly labeled samples and historical memory samples into eight small batches at a fixed ratio. Each batch contains newly labeled data and labeled samples sampled from the long-term memory buffer. The system continuously performs forward propagation, backward propagation, and parameter updates on these eight batches, achieving small-scale but efficient parameter fine-tuning.

10. The non-intrusive load dynamic identification method based on multimodal transfer learning according to claim 7, characterized in that: During training, a four-layer knowledge graph of "building-power distribution area-equipment-operating condition" is stored in the cloud using a graph database, and KL divergence is periodically calculated to detect concept drift; when the drift degree D KL When the index exceeds the threshold or the identification index drops below the threshold for three consecutive days, a "degradation" flag is automatically triggered and a retraining task is generated; differential updates replace the edge model within Y ms through a general flashover mechanism, ensuring that no downtime maintenance is required, where Y is a preset value.