Non-intrusive load decomposition method and device for energy scene

Through the ARDF-EN algorithm, the ARDF-EN algorithm is integrated with Transformer and RDF, combined with the dynamic feedback mechanism, the limitations and robustness of traditional models in high-dimensional timing data processing are solved, and high-precision load decomposition is achieved.

CN120337161AActive Publication Date: 2025-07-18SHANGHAI LINGANG HONGBO NEW ENERGY DEV CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510827908.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

When processing high-dimensional timing data in the prior art, traditional random decision forest models are difficult to capture global dependencies, while pure Transformer models lack local noise robustness, resulting in high load superposition error rate and fixed parameter models cannot adapt to complex working conditions.

Method used

The ARDF-EN algorithm is adopted, and the Transformer feature extraction and RDF integrated learning is integrated, and the dynamic feedback mechanism is introduced to improve the load decomposition accuracy by adaptively adjusting parameters and building equipment correlation diagrams.

Benefits of technology

It significantly improves the load decomposition accuracy in large-scale heterogeneous energy scenarios, dynamically adapts to complex working conditions, reduces calculation complexity, and effectively solves the load superposition problem caused by coordinated start and stop of equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337161A_ABST
    Figure CN120337161A_ABST
Patent Text Reader

Abstract

The invention provides a non-intrusive load decomposition method and device for an energy scene. The method relates to the technical field of intelligent power grid and energy management, and comprises the following steps: S1, obtaining aggregation load time sequence data and carrying out standardization processing; s2, inputting the standardized aggregation load time sequence data into a feature extraction module, and sequentially performing high-dimensional vector embedding, position code adding, multi-head self-attention mechanism processing and nonlinear transformation; s3, inputting the high-dimensional feature vector after nonlinear transformation into a load decomposition prediction module, sequentially performing maximum pooling dimension reduction processing, training a random decision forest model by using the features after dimension reduction, performing prediction, and outputting an operation load prediction value of the individual equipment; and S4, calculating the error of the operation load predicted value of the single equipment, and adjusting the model parameters according to the error adaptive feedback. According to the method, the load decomposition precision in a large-scale heterogeneous energy scene is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart grid and energy management, and particularly to a non-intrusive load decomposition method and device for energy scenarios. Background Art

[0002] In the construction of the energy Internet and new power systems, non-intrusive load monitoring (NILM) technology realizes device-level energy consumption decomposition by analyzing total load data, which plays a key role in demand response, anomaly detection, and energy efficiency optimization. Traditional methods such as Random Decision Forest (RDF) achieve load feature classification through ensemble learning, but there are problems with insufficient feature expression ability when dealing with high-dimensional time-series data. In recent years, attention mechanism models based on Transformer have shown advantages in sequence modeling. However, traditional Transformer models are less sensitive to local feature fluctuations, and the fixed model parameters result in weak scenario adaptability. For example, using fixed attention weights cannot adapt to the dynamic changes in device correlations (such as the coordinated start and stop of charging piles), leading to a high misjudgment rate of load superposition.

[0003] The existing technologies have the following defects: limitations of single models: pure RDF models are difficult to capture the global dependencies of time-series data, while pure Transformer models lack robustness to local noise; insufficient dynamic adaptation ability: fixed-parameter models cannot adapt to complex working conditions such as sudden changes in charging pile power and fluctuations in photovoltaic power output; feature fusion defects: the time-series position encoding and device operation mode correlation features are not effectively combined.

[0004] The present invention proposes the ARDF-EN algorithm, which significantly improves the load decomposition accuracy in large-scale heterogeneous energy scenarios by integrating the advantages of Transformer feature extraction and RDF ensemble learning and introducing a dynamic feedback mechanism. Summary of the Invention

[0005] In view of the above problems, the present invention proposes an Adaptive Random Decision Forest Encoder Network (ARDF-EN) algorithm, which can be applied to a non-intrusive load decomposition system in high-proportion energy scenarios. This algorithm can decompose individual operating devices from the input aggregated load data and introduce a dynamic feedback mechanism to achieve adaptive parameter adjustment, significantly improving the load decomposition accuracy in large-scale heterogeneous energy scenarios.

[0006] To achieve the above object, the following technical solutions are adopted: According to a first aspect of the present invention, there is provided a non-intrusive load decomposition method for energy scenarios, including: Step S1: Obtain the aggregated load time series data and perform normalization processing; Step S2: Input the normalized aggregated load time series data into the feature extraction module, and successively perform high-dimensional vector embedding, adding positional encoding, multi-head self-attention mechanism processing, and non-linear transformation to obtain the high-dimensional feature vector after non-linear transformation; Step S3: Input the high-dimensional feature vector after the non-linear transformation into the load decomposition prediction module, perform max-pooling dimensionality reduction processing successively to obtain the dimensionality-reduced features, and use the dimensionality-reduced features to train a random decision forest model, and then use the trained random decision forest model for prediction to output the operating load prediction value of a single device; Step S4: Calculate the error between the operating load prediction value of a single device and the true value, and adaptively feedback and adjust the hyperparameters of the feature extraction module and the number of trees and depth of the random decision forest model according to the error.

[0007] Further, the Step S1: Obtain the aggregated load time series data and perform normalization processing includes, Step S101: The input layer collects the aggregated load data in the format of a time series , where is the aggregated load data, is the load value at time , and M is the length of the time series; ;

[0008] In the formula, and are the mean and standard deviation of the aggregated load data respectively.

[0009] Further, among them, the Step S2 includes: Step S201: The embedding layer converts the normalized aggregated load time series data into a high-dimensional vector embedding representation , where is the real number field, and

[0010] is the embedding dimension; Step S202: The positional encoding layer adds positional encoding to the high-dimensional vector embedding representation E to obtain the embedding representation after adding positional encoding:

[0011] Among them, the positional encoding uses a sine-cosine function, is the time step at the position the sine positional encoding, is the time step at the position the cosine positional encoding, is the index of the time step, is the dimension index of the positional encoding; Step S203: The multi-head self-attention layer performs a multi-head self-attention mechanism process on the embedded representation after adding the positional encoding, calculates the self-attention, and obtains the feature representation after the multi-head self-attention mechanism process ; Step S204: The feed-forward neural network layer uses the ReLU activation function to perform a non-linear transformation on the feature representation A after the multi-head self-attention mechanism process, and obtains a high-dimensional feature vector after the non-linear transformation ;

[0012] In the formula, is the feature vector after the non-linear transformation, ReLU is the output of the activation function applied to the first layer, is the output of the self-attention layer, is the weight matrix of the first layer, is the bias vector of the first layer, is the weight matrix of the second layer, is the bias vector of the second layer.

[0013] Furthermore, the step S3 includes: Step S301: The max pooling layer of the load decomposition prediction module performs max pooling dimensionality reduction processing on the high-dimensional feature vector after the non-linear transformation, including normalization and dimensionality reduction processing, to obtain a dimensionality-reduced feature vector; Step S302: Use the dimensionality-reduced feature vector to train a random decision forest model to obtain a trained random decision forest model; Step S303: Use the trained random decision forest model to output the running load prediction value of a single device by integrating the prediction results of multiple decision trees.

[0014] Furthermore, in the step S302, each decision tree is trained using the dimensionality-reduced feature vector, and each decision tree makes a prediction independently;

[0015] In the formula, is the number of decision trees, is the prediction function of the j-th decision tree, is the feature vector after dimensionality reduction; In step S303, the prediction results of each decision tree, that is, the load prediction values, are integrated by majority voting or weighted average:

[0016] In the formula, is the predicted operating load value of a single device output by prediction, is the number of decision trees, is the prediction function of the j-th decision tree, is the processed feature vector.

[0017] Furthermore, step S4 includes: Calculating the root mean square error between the predicted operating load value of a single device and the true value; According to the root mean square error, adaptively and feedback-adjusting at least one of the number of Transformer layers, the number of attention heads, and the learning rate of the feature extraction module; synchronously optimizing the number of trees and the maximum depth of the random decision forest model.

[0018] Furthermore, when the feedback mechanism detects that the error exceeds the threshold: Start the graph neural network (GNN) to analyze the spatial correlation of devices and construct a device association graph, generating a node relationship matrix ; According to the node relationship matrix Reconstruct the weight assignment of the multi-head self-attention layer:

[0019] Among them, Q is the query matrix, K is the key matrix, and V is the value matrix, is the key vector dimension, and ⊙ represents the Hadamard product.

[0020] Furthermore, the step of starting the graph neural network (GNN) to analyze the device correlation and construct a device association graph, generating a node relationship matrix , includes: Collect the historical operating true load time series of all devices ; According to the historical operating true load time series of all devices , calculate the Pearson correlation coefficient between device i' and device j' ; Construct an initial undirected graph structure G=(V,E), where each device i' is used as a graph node , and the edge weight ; Optimize the edge weight through the graph neural network: , , where is the original edge weight between devices; is a differentiable edge weight learning structure; is the device node feature vector; For the original edge weight between devices perform non-linear normalization on the non-diagonal elements: ( ), , thus mapping to the interval [0, 1]; where α is a trainable scaling factor, is the trainable edge weight parameter; And, when and only when the load fluctuation of device i' or device j' exceeds the preset variance , recalculate the Pearson correlation coefficient ; The graph neural network only optimizes the edges that satisfy , is the sparsification threshold.

[0021] Furthermore, in the random decision forest model, the maximum depth of a single decision tree does not exceed 10, the minimum number of samples in a leaf node is 2, and the ensemble strategy uses a weighted average method to fuse the prediction results of each tree.

[0022] According to the second aspect of the present invention, there is also provided a non-intrusive load decomposition device for an energy scenario, which is used to implement the non-intrusive load decomposition method for an energy scenario as described in the first aspect, including: a data input module, a feature extraction module, a load decomposition prediction module, and a dynamic feedback module connected in sequence; The data input module includes an input layer and a normalization layer, and is used to collect and standardize the aggregated load time series data; The feature extraction module includes an embedding layer, a positional encoding layer, and multiple Transformer blocks, and is used to process the standardized aggregated load time series data to generate a high-dimensional feature vector after non-linear transformation; wherein, the Transformer block includes a multi-head self-attention layer and a feed-forward neural network layer; The load decomposition prediction module includes a max pooling layer and a random decision forest model, and is used to perform max pooling dimensionality reduction processing on the high-dimensional feature vector after non-linear transformation, and use the dimensionality-reduced features to train and predict the random decision forest model, and output the running load prediction value of a single device; The dynamic feedback module includes an error calculation unit and a parameter adjustment unit. The error calculation unit calculates the error between the predicted value and the true value; the parameter adjustment unit: dynamically optimizes the parameters of the feature extraction module and the load decomposition prediction module based on the error, and realizes the adaptive feedback adjustment of the model; Among them, the feedforward neural network layer includes two fully connected layers and a ReLU activation function; the random decision forest includes multiple decision trees and is trained through a random subset of features and a random subset of data.

[0023] Compared with the prior art, the present invention has the following beneficial effects: 1. By integrating the global feature extraction ability of Transformer and the local pattern recognition advantage of the random decision forest (RDF), the present invention enables Transformer to capture the long-range dependence relationship of time-series data, while RDF enhances the robustness to local noise through ensemble learning, breaking through the limitations of a single model.

[0024] 2. The present invention introduces a dynamic feedback mechanism to adaptively adjust the number of Transformer layers, the number of attention heads, and the depth of the RDF tree according to the real-time error. For example, when a load mutation is detected, the model can dynamically increase the number of attention heads to strengthen feature extraction, and at the same time optimize the RDF tree structure to suppress overfitting. While ensuring the accuracy, the computational complexity is reduced. For example, the number of Transformer layers and the depth of the RDF tree are dynamically adjusted according to the scene complexity to avoid redundant calculations.

[0025] 3. The present invention constructs a device association graph through a graph neural network (GNN) to generate a dynamic adjacency matrix to correct the attention weight distribution. This mechanism effectively solves the problem of load superposition caused by the coordinated start and stop of devices. For example, when a group of charging piles start and stop simultaneously, traditional models are prone to misjudging the superposed load as the behavior of a single device.

[0026] 4. By combining the time-series position encoding with the device operation mode correlation features, the present invention enables the model to distinguish real load fluctuations from noise interference. For example, the characteristics of photovoltaic power output fluctuations and charging loads are effectively decoupled through position encoding and GNN correlation analysis.

[0027] In summary, through technological innovations such as algorithm fusion, dynamic feedback, and device association modeling, the present invention has made breakthroughs in key indicators such as load decomposition accuracy, dynamic adaptability, anti-interference ability, and computational efficiency, providing a solution with both high accuracy and strong robustness for the fields of smart grid and energy management. The present invention has significant technical advantages and application values in complex heterogeneous energy scenarios.

[0028] It should be understood that the content described in the invention content section is not intended to limit the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The accompanying drawings are used to better understand the solution and do not limit the present invention. In the drawings, the same or similar reference numerals denote the same or similar elements, where: Figure 1 is a schematic flowchart of a non-intrusive load decomposition method for an energy scenario according to an embodiment of the present invention; Figure 2 is a schematic diagram of the data flow and network architecture of a non-intrusive load decomposition method and device for an energy scenario according to an embodiment of the present invention; Figure 3 is a schematic diagram of the modules of a non-intrusive load decomposition device for an energy scenario according to an embodiment of the present invention; Figure 4 is a schematic diagram of the structure of a Transformer encoding block according to an embodiment of the present invention; Figure 5 is a schematic flowchart of step S5 according to an embodiment of the present invention. Detailed implementation manners

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] In addition, the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0032] The purpose of this application is to provide an Adaptive Random Decision Forest Encoder Network (ARDF-EN), which is applied to the non-intrusive load decomposition and prediction task in a high-proportion energy scenario. It can be used for the non-intrusive load decomposition system in a high-proportion energy scenario. For example, in scenarios where the new energy penetration rate > 20% and the equipment composition includes a photovoltaic inverter group, a wind power converter, and a charging pile cluster, etc., it can solve the problem that the fluctuations of wind and light cover the charging load characteristics. Or, for example, in the case of a large-scale electric vehicle charging facility where the number of charging piles ≥ 100 and the power mutation rate > 30% / min, and the equipment composition includes DC fast charging piles, AC slow charging piles, and V2G piles, etc., it can solve the problem of load superposition caused by the coordinated start and stop of the pile group. Through the self-developed adaptive random decision forest large model algorithm, the load conditions of individual devices or energy consumption components are decomposed from the input aggregated load data, and parameter adaptive adjustment is realized by adding a feedback mechanism.

[0033] Load decomposition prediction refers to accurately predicting the load conditions of individual devices or energy consumption components from the overall energy load data, which is of great significance for energy management and optimization.

[0034] Embodiment 1: Figure 1 It shows a schematic flowchart of a non-intrusive load decomposition method for an energy scenario according to an embodiment of the present invention; Figure 2 It is a schematic diagram of the data flow and network architecture of a non-intrusive load decomposition method and device for an energy scenario according to an embodiment of the present invention. As Figure 1 and Figure 2 shown, a non-intrusive load decomposition method 100 for an energy scenario is applied to a non-intrusive load decomposition device 200 as shown in Figure 3 shown, and includes the following steps: Step S1: Obtain the aggregated load time series data and perform standardization processing; Step S2: Input the standardized aggregated load time series data into the feature extraction module, and successively perform high-dimensional vector embedding, add position encoding, multi-head self-attention mechanism processing, and non-linear transformation to obtain a non-linearly transformed high-dimensional feature vector; Step S3: Input the non-linearly transformed high-dimensional feature vector into the load decomposition prediction module, perform max-pooling dimensionality reduction processing successively to obtain the dimensionality-reduced feature, and use the dimensionality-reduced feature to train a random decision forest model, and then use the trained random decision forest model for prediction to output the running load prediction value of an individual device; Step S4: Calculate the error between the predicted value and the true value of the operating load of a single device, and adaptively adjust the hyperparameters of the feature extraction module, the number of trees and the depth of the random decision forest model according to the error.

[0035] Figure 3 It is a schematic diagram of the modules of a non-intrusive load decomposition device for an energy scenario according to an embodiment of the present invention. As Figure 3 shown, a non-intrusive load decomposition device 200 for an energy scenario according to an embodiment of the present invention is used to implement the above non-intrusive load decomposition method 100 for an energy scenario, and includes: a data input module 210, a feature extraction module 220, a load decomposition prediction module 230, and a dynamic feedback module 240 that are connected in sequence.

[0036] The data input module 210 includes an input layer and a normalization layer, and is used to collect and standardize the aggregated load time series data; The feature extraction module 220 includes an embedding layer (fully connected layer), a positional encoding layer, and multiple Transformer blocks, and is used to process the standardized aggregated load time series data to generate a high-dimensional feature vector after non-linear transformation; wherein, the Transformer block includes a multi-head self-attention layer and a feed-forward neural network layer; Among them, each Transformer block contains: a multi-head self-attention layer (Multi-Head Self-Attention); a feed-forward neural network layer (Feed-Forward Network, FFN). The FFN completes feature transformation inside the Transformer block, and its output is directly used as the input (or final output) of the next Transformer block. Among them, the feed-forward neural network layer includes two fully connected layers and a ReLU activation function; The load decomposition prediction module 230 includes a max pooling layer and a random decision forest model, and is used to perform max pooling dimensionality reduction processing on the high-dimensional feature vector after non-linear transformation, and use the dimensionality-reduced features to train and predict the random decision forest model, and output the predicted value of the operating load of a single device; Among them, the random decision forest includes multiple decision trees and is trained through a random subset of features and a random subset of data.

[0037] The dynamic feedback module 240 includes an error calculation unit and a parameter adjustment unit. The error calculation unit calculates the error between the predicted value and the true value; the parameter adjustment unit: dynamically optimizes the parameters of the feature extraction module 220 and the load decomposition prediction module 230 based on the error, such as adjusting at least one of the number of Transformer layers, the number of attention heads, and the learning rate, and synchronously optimizing the number of trees and the maximum depth of the random decision forest model to achieve adaptive feedback adjustment of the model; Further, step S1: Obtain the aggregated load time series data and perform normalization processing, specifically including: Step S101: The input layer collects the aggregated load data in the format of a time series , which is the aggregated load data, and is the load value at time , where M is the length of the time series; specifically, the aggregated load data received by the input layer from multiple devices (such as charging piles and battery swapping stations in charging facilities; photovoltaic storage devices and wind turbines in energy storage systems, etc.) is in the form of a time series, containing load data points (such as power values and energy consumption) for multiple time steps. For example, M = 1440, representing 24-hour data at 1-minute intervals.

[0038] Step S102: The normalization layer performs normalization processing on the aggregated load data to obtain the normalized time series data ;

[0039] In the formula, and are the mean and standard deviation of the aggregated load data respectively.

[0040] The time series data normalized through step S102 is used for subsequent feature extraction.

[0041] Further, step S2 specifically includes: Step S201: The embedding layer converts the normalized aggregated load time series data into a high-dimensional vector embedding representation , where is the real number field, and is the embedding dimension, for example, taking 256;

[0042] Through this step S201, the input time series data is converted into a high-dimensional vector representation, making it suitable for subsequent neural network processing.

[0043] Step S202: The position encoding layer adds position encoding to the high-dimensional vector embedding representation E , obtaining the embedding representation after adding position encoding:

[0044] where the position encoding uses a sine-cosine function, is the sine position encoding of time step at position , is the sine position encoding of time step at position Cosine position encoding is the index of the time step is the dimension index of the position encoding

[0045] In step S202, position encoding P is added to the high-dimensional vector embedding representation E to retain the order information of the time series, ensure that the model can capture the position information between time steps, and avoid affecting the prediction effect due to the loss of time step order information

[0046] Step S203: The multi-head self-attention layer performs a multi-head self-attention mechanism on the embedding representation after adding the position encoding, calculates the self-attention, and obtains the feature representation after the multi-head self-attention mechanism processing :

[0047] In the formula, Q is the query matrix, E is the high-dimensional vector embedding representation, W is the linear transformation matrix, K is the key matrix, V is the value matrix is the feature representation A after the multi-head self-attention mechanism processing is the dot product of the query matrix and the key matrix is the dimension of the key matrix, and softmax is the normalization function that converts the dot product result into a weight distribution

[0048] This step S203 is used to capture the global dependencies in the time series data. By calculating the attention weights between each time step and all other time steps, important features and relationships in the input sequence are extracted

[0049] Step S204: The feed-forward neural network layer uses the ReLU activation function to perform a non-linear transformation on the feature representation A after the multi-head self-attention mechanism processing, and obtains the high-dimensional feature vector after the non-linear transformation ;

[0050] In the formula is the feature vector after the non-linear transformation, ReLU is the activation function applied to the output of the first layer is the output of the self-attention layer is the weight matrix of the first layer is the bias vector of the first layer is the weight matrix of the second layer is the bias vector of the second layer. In the internal structure of the FFN, Linear(256→512) + ReLU + Linear(512→256)

[0051] Further, in some embodiments, the feature extraction module includes N (e.g., 4) cascaded Transformer encoding blocks, as Figure 4 shown. Each encoding block sequentially includes: a layer normalization unit, a multi-head self-attention calculation unit, a first residual connection unit, a second layer normalization unit, a feed-forward neural network unit (including two fully-connected layers and ReLU activation), and a second residual connection unit; Figure 4 where y in

[0052] is the output of the first residual connection unit. Among them, the parameters of the feed-forward neural network unit satisfy: , , the hidden layer dimension of the feed-forward layer , c is the input of the FFN, that is, the output A of the multi-head attention layer.

[0053] This step S204 performs a non-linear transformation on the output of the self-attention layer through the feed-forward neural network layer to enhance the expression ability of the model and further extract high-order features.

[0054] Further, step S3 specifically includes: Step S301: The max-pooling layer of the load decomposition prediction module performs max-pooling dimensionality reduction processing on the non-linearly transformed high-dimensional feature vector, including normalization and dimensionality reduction processing, to obtain a dimensionality-reduced feature vector; specifically including: Step S3011: Normalize the non-linearly transformed feature, that is, the output high-dimensional feature vector, to obtain the normalized feature:

[0055] In the formula, is the normalized feature, is the non-linearly transformed high-dimensional feature vector, is the mean of the feature vector, is the standard deviation of the feature vector.

[0056] Step S3012: Perform dimensionality reduction processing on the normalized feature to obtain the dimensionality-reduced feature:

[0057] In the formula, is the dimensionality-reduced feature; is the dimensionality reduction algorithm, which is used to reduce the dimension of the feature vector.

[0058] The max-pooling layer performs dimensionality reduction processing on the high-dimensional feature vector output from the Transformer model, and the processed feature vector. Through the above feature dimensionality reduction and other processing steps, it is convenient for the subsequent training of the random decision forest model.

[0059] Step S302: Train a random decision forest model using the dimensionality-reduced feature vectors to obtain a trained random decision forest model; The random decision forest model includes N decision trees and is trained using a random subset of features and a random subset of data. The specific steps are as follows: Decision tree training: Train N decision trees using the feature vectors, and each tree makes predictions independently. Ensemble prediction: Integrate the prediction results of each decision tree through an ensemble strategy such as majority voting or weighted averaging. Output: The predicted value of the operating load of a single device.

[0060] In this random decision forest model, the maximum depth of a single decision tree does not exceed 10, and the minimum number of samples in a leaf node is 2.

[0061] Train each decision tree using the dimensionality-reduced feature vectors, and each decision tree makes predictions independently;

[0062] where is the number of decision trees, is the prediction function of the j-th decision tree, is the dimensionality-reduced feature vector; In step S303, use the trained random decision forest model to output the predicted value of the operating load of a single device by integrating the prediction results of multiple decision trees; This step S303 integrates the prediction results of each decision tree, i.e., the load prediction value, by means of majority voting or weighted averaging:

[0063] where is the predicted value of the operating load of a single device output by the prediction, is the number of decision trees, is the prediction function of the j-th decision tree, is the processed feature vector.

[0064] Furthermore, step S4 specifically includes: Calculate the error between the predicted value of the operating load of a single device and the true value, such as the root mean square error:

[0065] where is the mean square error, is the predicted load value of the t-th sample, is the true load value of the t-th sample.

[0066] According to the root mean square error Adaptively adjust at least one of the number of Transformer layers, the number of attention heads, and the learning rate of the feature extraction module; synchronously optimize the number of trees and the maximum depth of the random decision forest model.

[0067] For example, when the root mean square error Error exceeds the threshold adopt the gradient descent method to update the number of Transformer layers L and the number of attention heads H, and the update formulas are: L = L - α * ∂Error / ∂L H = H - β * ∂Error / ∂H where α and β are the learning rates, which are dynamically adjusted according to the historical error (α ∈ [0.01, 0.1], β ∈ [0.005, 0.05]).

[0068] Through the updated model parameters, this step S4 can significantly improve the prediction accuracy.

[0069] Preferably, in some embodiments of the present invention, the method further includes: S5: Based on the dynamic topology reconstruction mechanism, when the feedback mechanism detects that the error exceeds the threshold correct the attention weight distribution through the Hadamard product. Among them, the threshold is dynamically set according to the 95th percentile of the historical error and updated every 24 hours. Specifically, it includes the following steps: S501: Start the graph neural network (GNN) to analyze the device spatial correlation and construct a device association graph, generating a node relationship matrix (i.e., a dynamic adjacency matrix) ; Z is the total number of devices to be decomposed (for example, the charging pile numbers from 1 to Z), with a value range of [0, 1], representing the association strength (topological structure) between device i' and device j'. The node relationship matrix is the prior knowledge of the device dimension, such as describing the collaborative start-stop rules between charging piles.

[0070] Preferably, when it is detected that the error > 0.25, trigger the GNN to generate the node relationship matrix .

[0071] S502: Reconstruct the weight distribution of the multi-head self-attention layer according to the node relationship matrix :

[0072] where Q is the query matrix, K is the key matrix, V is the value matrix, is the key vector dimension, and ⊙ represents the Hadamard product.

[0073] The dynamic topology mechanism in step S5 can solve technical problems such as the dynamic change of equipment correlation caused by the start and stop of a charging pile group, which is difficult to adapt to with a fixed model, or the load superposition caused by the simultaneous start and stop of the pile group, which cannot be distinguished by traditional attention. By separating the superimposed load, the recognition rate of the model for sudden load events can be effectively improved, and the prediction value error can be reduced.

[0074] Further, step S501: Start the Graph Neural Network (GNN) to analyze the equipment correlation and construct an equipment correlation graph, generating a node relationship matrix , specifically including: S5011: Collect the historical running true load time series Y = , that is, equipment metadata (type / position / rated power), .

[0075] S5012: According to the historical running true load time series of all equipment , calculate the Pearson correlation coefficient between equipment i' and equipment j' ; Calculate the covariance:

[0076] Pearson correlation coefficient:

[0077] S5013: Construct an initial undirected graph structure G=(V,E), where each equipment i' is used as a graph node , and the edge weight ; S5014: Optimize the edge weight through the graph neural network: , , where, is the original edge weight between equipment; is the differentiable edge weight learning structure; is the equipment node feature vector ( is the hidden layer dimension), which fuses the equipment dynamic power characteristics ( ) and static attributes (type / position), and the pre-trained can quickly adapt to new scenarios (such as newly added charging piles) and has transferability.

[0078] Moreover, the Graph Neural Network (GNN) model structure: adopts 2 layers of GCN (Graph Convolutional Network), with an output dimension of 128 for each layer, and the activation function is ReLU; Design the node feature ; for example, if equipment i' and j' are charging piles, then is the average power of the equipment, is the power standard deviation; device type embedding (e.g., fast charging pile / slow charging pile); geographical location embedding (e.g., charging station area A); Message passing formula: , where represents the feature vector of node i' at l the +1 layer; represents the feature vector of j' at l the layer; represents the neighbor set of node i'. If node i' has 3 neighbors, i.e., associated devices, then ; ReLU is a non-linear activation function, ReLU(⋅)=max(0,⋅); is the trainable weight matrix of the l layer; is a linear feature transformation; represents neighbor feature aggregation; represents the normalization coefficient to avoid information bias caused by differences in the number of neighbors. If device j' has 10 neighbors and device i' has 2 neighbors, then the weight decay ; The combined pattern of neighbor features is extracted through the learnable matrix For example, capture the light-load coupling characteristics of photovoltaic energy storage devices or learn the periodic start-stop pattern of charging pile groups. When the device association changes (such as adding a charging pile), only the neighbor set needs to be updated, and there is no need to retrain the entire model.

[0079] S5015: Perform non-linear normalization on the non-diagonal elements of the original edge weight between devices: ( ), , thereby mapping to the [0,1] interval; where α is a trainable scaling factor, for example, α∈[2,5], is a trainable edge weight parameter; And, when and only when the load fluctuation of device i' or device j' exceeds the preset variance , recalculate the Pearson correlation coefficient ; The graph neural network only optimizes the edges that satisfy , is the sparsification threshold. For example, neighbor aggregation: only aggregate devices with an association degree (such as β = 0.3).

[0080] Preferably, when recalculating the Pearson correlation coefficient , the covariance update formula is:

[0081] Among them, γ∈[0,1] is the forgetting factor.

[0082] According to the non-intrusive load decomposition method and device for energy scenarios provided by the above embodiment of the present invention, by designing an adaptive random decision forest large model (Adaptive Random Decision Forest Encoder Network, ARDF-EN) algorithm, it is possible to decompose individual operating devices from the input aggregated load data, and introduce a dynamic feedback mechanism to achieve parameter adaptive adjustment, significantly improving the load decomposition accuracy in large-scale heterogeneous energy scenarios; and the feature extraction module (Transformer) and the load decomposition module (random forest) are trained and optimized through an end-to-end gradient backpropagation joint training optimization mechanism (not staged training), the Transformer feature layer is targetedly optimized, the random forest decision boundary is matched, and the error feedback adjusts the Transformer hyperparameters (such as the number of attention heads) and the forest structure (tree depth) in real time, thereby improving the decomposition accuracy and the recognition rate of sudden load scenarios. The present invention provides a load decomposition solution that takes into account high precision, strong anti-interference, low latency, and reliability in the field of smart grids.

[0083] Embodiment 2: The present invention discloses a non-invasive load decomposition device for energy scenarios, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. The memory and the processor form an electronic terminal. When the processor executes the computer program, the steps of a non-invasive load decomposition method for energy scenarios of an embodiment are implemented.

[0084] Embodiment three: The present invention discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a non-invasive load decomposition method for energy scenarios in embodiment 1 are implemented.

[0085] Comparative experiment: Experiments were conducted on a system equipped with an NVIDIA Tesla V100 GPU using Python 3.8, TensorFlow 2.5, and Scikit-learn 0.24. The Transformer model was configured with 4 layers, 8 attention heads, 256-dimensional embedding, a learning rate of 0.001, a batch size of 32, and 50 epochs. The random decision forest model included 100 decision trees, a maximum depth of 10, and a minimum sample split of 2. The root mean square error (RMSE) was used as the main evaluation metric to evaluate the performance of the two models in load decomposition forecasting. The results are shown in Table 1.

[0086] Table 1: Experimental results

[0087] According to the results of this comparative experiment, it can be seen that the present invention captures temporal dependencies through a feature extraction module (Transformer), enhances robustness with a random decision forest, and dynamically optimizes through a feedback mechanism. Compared with a single model, it effectively reduces the root mean square error (RMSE), improves the decomposition accuracy, especially has significant advantages in high-interference scenarios, and dynamic topology reconstruction suppresses the load superposition of associated devices, with strong anti-interference ability.

[0088] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the described steps can refer to the corresponding processes in the foregoing system embodiments and will not be elaborated herein.

[0089] It should be noted that the embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0090] It should also be noted that in the embodiments of the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0091] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined in the embodiments of the present application can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown in the embodiments of the present application, but will conform to the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.

Claims

1. A non-intrusive load decomposition method for an energy scenario, characterized in that Including: Step S1: Obtain the aggregated load time series data and perform normalization processing; Step S2: Input the normalized aggregated load time series data into the feature extraction module, and successively perform high-dimensional vector embedding, add positional encoding, multi-head self-attention mechanism processing, and non-linear transformation to obtain the non-linearly transformed high-dimensional feature vector; Step S3: Input the non-linearly transformed high-dimensional feature vector into the load decomposition prediction module, successively perform max pooling dimensionality reduction processing to obtain the dimensionality-reduced features, use the dimensionality-reduced features to train a random decision forest model, and then use the trained random decision forest model for prediction to output the operating load prediction value of a single device; Step S4: Calculate the error between the operating load prediction value of a single device and the true value, and adaptively feedback and adjust the hyperparameters of the feature extraction module, the number of trees and depth of the random decision forest model according to the error.

2. The non-intrusive load decomposition method for an energy scenario according to claim 1, wherein The said Step S1: Obtaining the aggregated load time series data and performing normalization processing includes, Step S101: The input layer collects aggregated load data in the format of time series , is the aggregated load data, is the load value at time , and M is the length of the time series; Step S102: The normalization layer performs standardization processing on the aggregated load data to obtain the standardized time series data ; In the formula, and are the mean and standard deviation of the polymerization load data, respectively.

3. The non-intrusive load decomposition method for an energy scenario according to claim 2, wherein Wherein, The said Step S2 includes: Step S201: The embedding layer converts the standardized aggregated load time series data into a high-dimensional vector embedding representation , where is the real number field, and is the embedding dimension; Step S202: The position encoding layer adds position encoding to the high-dimensional vector embedding representation E , obtaining the embedding representation E+P after adding the position encoding: Among them, the positional encoding uses a sine-cosine function, is the time step at the position the sine positional encoding, is the time step at the position the cosine positional encoding, is the index of the time step, is the dimension index of the positional encoding; Step S203: The multi-head self-attention layer processes the embedded representation after adding the positional encoding through the multi-head self-attention mechanism, calculates the self-attention, and obtains the feature representation after the multi-head self-attention mechanism processing ; Step S204: The feedforward neural network layer uses the ReLU activation function to perform a non-linear transformation on the feature representation A processed by the multi-head self-attention mechanism, obtaining a high-dimensional feature vector after the non-linear transformation ; In the formula, is the feature vector after non - linear transformation, ReLU is the activation function applied to the output of the first layer, is the output of the self - attention layer, is the weight matrix of the first layer, is the bias vector of the first layer, is the weight matrix of the second layer, is the bias vector of the second layer.

4. The non-intrusive load decomposition method for an energy scenario according to claim 3, wherein The said Step S3 includes: Step S301: The max pooling layer of the load decomposition prediction module performs max pooling dimensionality reduction processing on the non-linearly transformed high-dimensional feature vector, including normalization and dimensionality reduction processing, to obtain the dimensionality-reduced feature vector; Step S302: Use the dimensionality-reduced feature vector to train a random decision forest model to obtain the trained random decision forest model; Step S303: Use the trained random decision forest model to output the operating load prediction value of a single device by integrating the prediction results of multiple decision trees.

5. The non-intrusive load decomposition method for an energy scenario according to claim 4, characterized in that In the said Step S302, use the dimensionality-reduced feature vector to train each decision tree, and each decision tree makes an independent prediction; Wherein, is the number of decision trees, is the prediction function of the j-th decision tree, is the feature vector after dimensionality reduction; In the said Step S303, integrate the prediction results of each decision tree, that is, the load prediction value, by means of majority voting or weighted average: Wherein, is the predicted operating load value of a single device for the predicted output, is the number of decision trees, is the prediction function of the j-th decision tree, is the processed feature vector.

6. The non-intrusive load decomposition method for an energy scenario according to claim 1, wherein The said Step S4 includes: Calculate the root mean square error between the operating load prediction value of a single device and the true value; According to the root mean square error, adaptively feedback and adjust at least one of the number of Transformer layers, the number of attention heads, and the learning rate of the feature extraction module; synchronously optimize the number of trees and the maximum depth of the random decision forest model.

7. The non-intrusive load decomposition method for an energy scenario according to claim 6, wherein When the feedback mechanism detects that the error exceeds the threshold: Start the graph neural network (GNN) to analyze the spatial correlation of devices and construct a device association graph, and generate a node relationship matrix ; According to the node relationship matrix Reconstruct the weight assignment of the multi-head self-attention layer: where Q is the query matrix, K is the key matrix, and V is the value matrix. is the key vector dimension, and ⊙ represents the Hadamard product.

8. The non-intrusive load decomposition method for an energy scenario according to claim 7, characterized in that The startup graph neural network (GNN) analyzes the relevance of devices and constructs a device association graph to generate a node relationship matrix , including: Collect the real load time series of all equipment's historical operations ; According to the true load time series of all the device historical operations , calculate the Pearson correlation coefficient between device i' and device j' ; Construct an initial undirected graph structure G=(V, E), where each device i’ serves as a graph node , edge weight ; Optimizing Edge Weights through Graph Neural Networks: , , where is the original edge weight between devices; is a differentiable edge weight learning structure; is the device node feature vector; Non-linearly normalize the non-diagonal elements of the original edge weights in the equipment room : ( ), , thus mapping to the interval [0, 1]; where α is a trainable scaling factor, and is a trainable edge weight parameter; Moreover, when and only when the load fluctuation of device i' or device j' exceeds a preset variance the Pearson correlation coefficient is recalculated ; the graph neural network only optimizes the edges that satisfy where is the sparsification threshold.

9. The non-intrusive load decomposition method for an energy scenario according to claim 6, wherein In the said random decision forest model, the maximum depth of a single decision tree does not exceed 10, the minimum number of samples in a leaf node is 2, and the integration strategy uses the weighted average method to fuse the prediction results of each tree.

10. A non-intrusive load decomposition device for an energy scenario, which is used to implement the non-intrusive load decomposition method for an energy scenario as described in any one of claims 1 to 9, characterized in that, Including: A data input module, a feature extraction module, a load decomposition prediction module, and a dynamic feedback module connected in sequence; The said data input module includes an input layer and a normalization layer, and is used to collect and normalize the aggregated load time series data; The said feature extraction module includes an embedding layer, a positional encoding layer, and multiple Transformer blocks, and is used to process the normalized aggregated load time series data to generate a non-linearly transformed high-dimensional feature vector; wherein, the said Transformer block includes a multi-head self-attention layer and a feed-forward neural network layer; The load decomposition prediction module includes a max pooling layer and a random decision forest model, which are used to perform max pooling dimensionality reduction processing on the high-dimensional feature vectors after non-linear transformation, and use the dimensionality-reduced features to train and predict the random decision forest model, and output the operating load prediction value of a single device; The dynamic feedback module includes an error calculation unit and a parameter adjustment unit. The error calculation unit calculates the error between the predicted value and the true value; the parameter adjustment unit: dynamically optimizes the parameters of the feature extraction module and the load decomposition prediction module based on the error to achieve adaptive feedback adjustment of the model; Among them, the feedforward neural network layer includes two fully connected layers and a ReLU activation function; the random decision forest includes multiple decision trees and is trained through random subsets of features and random subsets of data.

Citation Information

Patent Citations

  • Graph neutral networks with attention

    CN112119412A

  • Transform and RF fused power load prediction method

    CN119448265A

  • Non-intrusive load monitoring method based on hybrid federated learning and cross Transform

    CN119962582A

  • Graph neutral networks with attention

    US20210081717A1

  • Non-Intrusive Load Decomposition Method Based on Informer Model Coding Structure

    US20220397874A1