Building load prediction method based on deep neural network and non-invasive monitoring

By combining deep neural networks with non-invasive monitoring using wavelet transform and Hilbert-Huang transform, multi-scale and transient features are extracted, solving the problem that existing building load forecasting methods cannot accurately decompose to the equipment level. This achieves high-precision total load and equipment-level load forecasting, supporting refined energy management.

CN121301901APending Publication Date: 2026-01-09SHANGHAI HOSPITAL OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511441535.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing building load forecasting methods cannot accurately decompose to the equipment level. Traditional methods rely on invasive monitoring, which is costly and complex to implement. Non-invasive methods have shortcomings in feature extraction and prediction models, making it difficult to achieve high-precision simultaneous forecasting of total load and equipment-level load.

Method used

By combining deep neural networks with non-invasive monitoring, multi-scale and transient features are extracted through wavelet transform and Hilbert-Huang transform to construct a multi-dimensional feature matrix. The matrix is ​​then trained using a bidirectional long short-term memory network and an attention mechanism layer to achieve high-precision prediction of total load and equipment-level load.

Benefits of technology

It enables high-precision prediction of total building load and equipment load, provides a macro-to-micro view of energy consumption, reduces system deployment complexity and cost, and supports refined energy management and demand response strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301901A_ABST
    Figure CN121301901A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of building energy management, and discloses a building load prediction method based on a deep neural network and non-invasive monitoring. The method comprises the following steps: firstly, acquiring total load data and non-invasive monitoring data of a building, including voltage, current, power signals and equipment switching events; performing multi-scale decomposition on the power signal by using wavelet transform to obtain an approximation coefficient and detail coefficient sub-sequence; and extracting instantaneous frequency and instantaneous amplitude characteristics of the power signal based on Hilbert-Huang transform. And fusing the approximation coefficient, the detail coefficient, the instantaneous frequency and the instantaneous amplitude features to construct a multi-dimensional feature matrix. And inputting the multi-dimensional feature matrix into a deep neural network model containing a bidirectional long-short-term memory network layer and an attention mechanism layer for training. And decomposing the predicted total load into specific equipment-level load prediction results by using a non-invasive load monitoring technology. According to the method, the building load prediction precision and the equipment-level load identification capability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of building energy management technology, specifically a building load prediction method based on deep neural networks and non-invasive monitoring. Background Technology

[0002] Building energy consumption accounts for a significant proportion of global energy consumption, and accurate load forecasting is essential for optimizing energy dispatch, improving energy efficiency, reducing operating costs, and enhancing grid stability. Traditional building load forecasting methods primarily rely on historical total load data, combined with meteorological and temporal information, and utilize statistical models or early machine learning models for prediction. While these methods have achieved some success, their predictions only remain at the overall building level, failing to reveal the operating modes and energy consumption details of various internal electrical devices, thus hindering the development of sophisticated energy management and demand response strategies.

[0003] To obtain device-level energy consumption information, early methods primarily employed invasive load monitoring. This method requires installing dedicated sensors in the circuits of each monitored electrical device to directly measure its energy consumption. While invasive methods can theoretically obtain accurate device-level data, they have numerous limitations in practical applications. Large-scale sensor deployment incurs high equipment purchase and installation costs, retrofitting existing buildings is complex and may disrupt normal power supply for users, and subsequent maintenance is extensive. These factors severely restrict its application in a wide range of scenarios.

[0004] The emergence of non-intrusive load monitoring (NILM) technology offers a new approach to solving equipment-level monitoring problems. This technology requires only a small number of sensors (typically monitoring total voltage and current) installed at the building's main entrance, identifying the switching status and energy consumption of internal equipment by analyzing changes in the total power waveform. NILM avoids the drawbacks of invasive methods, lowering the implementation threshold. However, existing NILM-based building load forecasting methods still face challenges. Many methods focus on steady-state characteristic analysis, underutilizing transient characteristics during equipment start-up and shutdown, which often contain important equipment identification and status information. At the feature extraction level, a single feature extraction method may struggle to fully capture the multi-scale, non-stationary characteristics inherent in power signals. Regarding prediction models, simple model structures may fail to effectively learn the complex temporal dependencies and nonlinear patterns in load changes, resulting in limited prediction accuracy. Furthermore, effectively and accurately decomposing the total load forecast results to the equipment level is a key area for improvement in existing methods. Therefore, developing a high-precision prediction method that can fully utilize non-intrusive data information, integrate multi-scale and transient features, possess strong time series modeling capabilities, and simultaneously output total load and equipment-level load is an urgent problem to be solved in the field of building energy management. Summary of the Invention

[0005] The purpose of this invention is to provide a building load prediction method based on deep neural networks and non-invasive monitoring to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a building load prediction method based on deep neural networks and non-invasive monitoring, the method comprising: Acquire total building load data and non-invasive monitoring data, including voltage signals, current signals, power signals, and equipment switching event data; The power signal is decomposed into approximate coefficient subsequences and detail coefficient subsequences by wavelet transform. Instantaneous frequency and instantaneous amplitude characteristics of power signals are extracted based on Hilbert-Huang transform; By integrating the approximate coefficient subsequence, detail coefficient subsequence, instantaneous frequency features, and instantaneous amplitude features, a multidimensional feature matrix is ​​constructed. The multidimensional feature matrix is ​​input into a deep neural network model for training. The deep neural network model includes a bidirectional long short-term memory network layer and an attention mechanism layer. The total load forecast results are decomposed into equipment-level load forecast results using non-invasive load monitoring technology, and the forecast results of total load and equipment-level load are output.

[0007] Preferably, the multi-scale decomposition of the power signal using wavelet transform specifically includes: A hybrid Morette wavelet basis function is used to perform continuous wavelet transform on the power signal. By adjusting the scaling and translation parameters, approximate coefficient subsequences and detail coefficient subsequences with different decomposition levels are extracted.

[0008] Preferably, the extraction of instantaneous frequency features based on Hilbert-Huang transform includes: Empirical mode decomposition is performed on the power signal to obtain multiple intrinsic mode functions and residuals; Perform a Hilbert transform on each intrinsic mode function to generate an analytic signal and calculate its instantaneous frequency and instantaneous amplitude.

[0009] Preferably, the construction of the multidimensional feature matrix includes: Calculate the maximum value, root mean square value, and variance of the approximate coefficient subsequence and the detailed coefficient subsequence within each time window; The statistical features, instantaneous frequency features, and instantaneous amplitude features are combined to generate a multidimensional feature matrix.

[0010] Preferably, the structure of the deep neural network model includes: The input layer receives a multidimensional feature matrix; Bidirectional long short-term memory network layers are used to capture temporal dependencies and output hidden state sequences; The attention mechanism layer performs weighted aggregation on the hidden state sequence and then outputs the total load prediction and device-level load probability distribution through the fully connected layer.

[0011] Preferably, the deep neural network model uses a joint loss function to optimize the model parameters. The joint loss function is a weighted combination of the mean square error of total load prediction and the cross-entropy loss of equipment-level decomposition, and the contributions of the two types of losses are balanced by hyperparameters.

[0012] Preferably, the dynamically updated feature matrix includes: The wavelet coefficients and instantaneous frequency characteristics of the current time window are calculated based on the real-time power signal. The feature matrix is ​​then updated and the input data is renormalized using a sliding window mechanism.

[0013] Preferably, the equipment-level load decomposition method includes: Detect sudden changes in power signals and determine changes in device status based on threshold values; A pre-trained database of device features is used to match power mutation patterns to generate device-level load prediction results.

[0014] Preferably, the method for adjusting the model weight parameters includes: The learning rate and the weight allocation of the attention mechanism layer are dynamically adjusted based on the exponentially weighted moving average of historical prediction errors.

[0015] Preferably, the present invention also includes a building load prediction system based on deep neural networks and non-invasive monitoring, the system comprising: The data acquisition module is used to acquire total building load data and non-intrusive monitoring data, including voltage signals, current signals, power signals and equipment switching event data. The signal processing module is used to perform multi-scale decomposition of the power signal through wavelet transform to obtain approximate coefficient subsequences and detail coefficient subsequences; The feature extraction module is used to extract the instantaneous frequency and instantaneous amplitude features of the power signal based on the Hilbert-Huang transform. The feature fusion module is used to fuse the approximate coefficient subsequence, the detail coefficient subsequence, the instantaneous frequency feature, and the instantaneous amplitude feature to construct a multidimensional feature matrix; The model training module is used to input the multidimensional feature matrix into a deep neural network model for training. The deep neural network model includes a bidirectional long short-term memory network layer and an attention mechanism layer. The load decomposition and output module is used to decompose the total load forecast results into equipment-level load forecast results using non-intrusive load monitoring technology, and output the forecast results of total load and equipment-level load.

[0016] Compared with the prior art, the beneficial effects of the present invention are: The proposed building load forecasting method significantly enriches the information dimensions extracted from the original power signal by integrating the multi-scale decomposition capability of wavelet transform and the ability of Hilbert-Huang transform to capture the instantaneous features of non-stationary signals. Wavelet transform decomposes the power signal into subsequences of approximate and detail coefficients at different scales, respectively characterizing the long-term trend and short-term fluctuation details of the load. Hilbert-Huang transform, on the other hand, deeply analyzes the instantaneous frequency and amplitude changes of the power signal at key event points such as equipment switching, revealing the unique characteristics of equipment operation. Fusing the multi-scale decomposition results with the instantaneous features to construct a multi-dimensional feature matrix provides more comprehensive and refined input information for subsequent models, enhancing the model's ability to understand the inherent laws governing load changes.

[0017] In terms of model architecture, a deep neural network is employed, integrating a bidirectional long short-term memory (LSTM) network layer and an attention mechanism layer, effectively enhancing the model's ability to handle complex time-series data. The LSTM network can simultaneously learn the long-term dependencies of load data along the time axis, fully capturing the historical inertia of load changes and the potential connection to future trends. The introduction of the attention mechanism allows the model to dynamically focus on the most critical features in the input sequence for the current prediction time, automatically distinguishing the importance of different feature points, especially improving the model's sensitivity and modeling accuracy to abrupt changes such as equipment switching events. This model combination effectively overcomes the shortcomings of traditional models in handling long-term dependencies and identifying key transient events.

[0018] The core advantage of this method lies in its prediction results simultaneously covering both total building load and equipment-level load. A trained deep neural network model is first used to obtain a high-precision total load forecast. Then, combined with non-invasive load monitoring technology, the predicted total load waveform is decomposed and restored to its constituent equipment-level load forecasts. This integrated prediction framework eliminates the disconnect between total load forecasting and equipment-level decomposition, allowing equipment-level forecasts to directly reflect the trends and detailed changes in total load forecasts. The final output of total load and equipment-level load forecast information provides building managers with a complete energy consumption view from macro to micro levels.

[0019] From an application perspective, this method relies entirely on non-invasive monitoring data, eliminating the need to install dedicated sensors on various circuits within a building. This significantly reduces the complexity and cost of system deployment, facilitating its implementation in various new and existing buildings. Its output device-level forecasts enable the identification of high-energy-consuming equipment, analysis of equipment usage patterns, and evaluation of the effectiveness of energy-saving measures. This information is crucial for developing personalized energy-saving strategies, implementing precise demand response, and optimizing equipment operation plans. Simultaneously, accurate total load forecasting helps buildings better participate in grid interaction, promoting energy supply and demand balance. While improving the accuracy and practicality of building load forecasting, this method's non-invasive nature and device-level forecasting capabilities create favorable conditions for refined building energy management. Attached Figure Description

[0020] Figure 1 This is a schematic diagram illustrating the working principle of the building load prediction method based on deep neural networks and non-invasive monitoring as described in this invention. Figure 2 To extract feature maps of instantaneous frequencies based on Hilbert-Huang transform; Figure 3 A flowchart illustrating the structure and operation of a deep neural network model; Figure 4 A flowchart for optimizing the joint loss function of a deep neural network model. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1 This invention provides a building load prediction method based on deep neural networks and non-invasive monitoring, and its specific implementation will be described in detail below.

[0023] The building load forecasting system collects total load data and non-invasive monitoring data, including power signals generated from voltage and current signals and equipment switching events. The power signals are processed as follows: multi-scale decomposition is performed using wavelet transform to generate approximate coefficient subsequences and detail coefficient subsequences; instantaneous frequency and amplitude features are extracted using Hilbert-Huang transform. All the above subsequences and features are fused to construct a multi-dimensional feature matrix. This matrix is ​​input into a deep neural network model, whose structure includes a bidirectional long short-term memory network layer to capture temporal dependencies and an attention mechanism layer to weight the hidden states. After the model outputs the total load forecast, the total load is decomposed into equipment-level load probability distributions based on non-invasive load monitoring technology, finally outputting the total load and equipment-level load forecast results.

[0024] Example 1: Wavelet Transform Multi-Scale Decomposition of Power Signals. This method employs continuous wavelet transform, specifically using hybrid Morlet wavelet basis functions. The input power signal is time-series data, such as building power consumption monitoring values, calculated from voltage and current signals. The wavelet transform method is chosen based on its ability to separate signal features at different frequency scales, allowing the capture of both long-term trends and short-term fluctuations in the power signal within the building load. Hybrid Morlet wavelet basis functions are particularly suitable for this scenario because their design provides a good balance of resolution in time-frequency analysis, making them suitable for complex transient events contained in the power signal, such as equipment startup or shutdown. The basis function has a complex form, providing phase information, which facilitates the capture of the signal's transient behavior. Power signals are typically represented as discrete-time series, and the sampling frequency must meet the signal characteristics requirements, commonly ranging from 1Hz to 1kHz to cover the range of building load variations.

[0025] In the implementation of continuous wavelet transform, the scaling and translation parameters control the multi-scale decomposition process by adjusting the number of decomposition levels. The scaling parameter determines the frequency resolution, with smaller scales corresponding to high-frequency components and larger scales corresponding to low-frequency components; the translation parameter determines the analysis window position in time. The signal is decomposed into a fixed number of levels to achieve coverage from low-order trends to high-order details. A typical setting is 6 to 8 decomposition levels, covering a frequency band from 0.1 Hz to 100 Hz. This ensures effective separation of slow low-frequency changes and rapid high-frequency disturbances in building load. The number of levels is determined based on the signal length and noise level; for example, in a typical 10-minute sampling window, 7 decomposition levels are set. The scaling parameter values ​​start from the lowest scale corresponding to the highest frequency, increasing logarithmically to ensure uniform coverage in the frequency space. The minimum scale value is set to 0.1 seconds, corresponding to the highest frequency of 100 Hz; the maximum scale value is set to 10 seconds, corresponding to the lowest frequency of 0.1 Hz; the scale interval grows logarithmically, with each increment based on a proportional factor such as 1.5 times. The translation parameters cover the entire signal time range, and the translation interval is dynamically set according to the signal sampling rate; for example, the translation step size is set to 1 second when the sampling rate is 1Hz, and the step size is set to 0.01 seconds when the sampling rate is 100Hz.

[0026] The approximate coefficient subsequence and detail coefficient subsequence for each level are calculated directly using wavelet transform coefficient values. The core mathematical expression of wavelet transform is defined as: In this formula, the specific meanings of the symbols are as follows: : Indicates the scale parameter Translation parameters The wavelet coefficient values ​​calculated below; Scale parameter, a real-valued variable used to control the width of the frequency scale; smaller values ​​are preferred. Corresponding to high-frequency components, larger Corresponding to low-frequency components; Translation parameter, a real number variable, used to specify the starting position of the wavelet analysis window on the time axis; : Original power signal function, representing time The input signal value is represented in real number form; : Represents the complex conjugate form of the wavelet basis functions used (mixed Morette wavelet basis functions); : Time variable, real number variable, covering the entire time domain of the signal; Integration operation: executed from negative infinity to positive infinity, indicating continuous analysis over the entire time range of the signal.

[0027] In this calculation process, the approximate coefficient subsequence extracts low-frequency trend components and is formed by serializing coefficient values ​​at larger scales (such as higher levels); the detail coefficient subsequence extracts high-frequency transient components and is formed by serializing coefficient values ​​at smaller scales (such as lower levels). The coefficient subsequences for each decomposition level are calculated and stored independently. When constructing multi-level sequences, the approximate coefficient subsequence for a specific level is generated by applying coefficient values ​​at the corresponding level's scale; the detail coefficient subsequence is processed similarly. The coefficient value sequence is stored as a time series of equal length to the original signal, with each time point corresponding to one coefficient value. The sequence length is consistent with the number of sampling points of the input power signal. A zero-padding strategy is used for signal boundary processing, automatically filling in zero values ​​when the integration range exceeds the signal boundary to avoid boundary distortion. The sequence storage format is set to array form, with two arrays output for each level: one for the approximate coefficient sequence and the other for the detail coefficient sequence; data storage uses floating-point format to preserve precision. The output of the multi-level decomposition is organized as a hierarchical index structure, storing the coefficient sequences sequentially from level one to level six or level eight.

[0028] The number of decomposition levels is set by adjusting the number of scale parameters. After determining the total number of levels, scale values ​​are allocated in a geometric sequence: the minimum scale value is the system-defined lower limit, the maximum scale value is the system-defined upper limit, and intermediate values ​​are generated exponentially. For example, if the minimum scale... Maximum scale If the total number of levels is 6, then the scale value of each level is calculated logarithmically as follows: in For hierarchical indexes, This represents the total number of layers. The translation parameter interval dynamically depends on the sampling frequency; if the sampling frequency is... Hz, the translation step size is usually set to The calculation of coefficient values ​​at each time point involves a discrete integration process, which is actually implemented using numerical integration methods to process the sampled data. The input power signal is preprocessed, including detrending and noise reduction. Detrending removes linear or slowly changing baselines, and noise reduction uses a smoothing filter to reduce the impact of high-frequency noise. During the calculation, wavelet coefficient values ​​are optimized for performance using a fast convolution algorithm. The convolution kernel is the complex conjugate form of the wavelet basis function, and discrete convolution is performed with the signal sampled point sequence. The resulting values ​​are stored in the corresponding positions of the output sequence. Approximate coefficients reflect the overall power trend characteristics at higher levels, while detail coefficients reflect transient event characteristics such as device switching at lower levels. The output coefficient subsequences are stored in a two-dimensional array structure, with the dimension being the time point multiplied by the number of levels, ensuring that each subsequence is independently accessible.

[0029] When calculating detail coefficients, high-frequency fluctuation information corresponding to high-frequency bands at low scales is preserved; when calculating approximation coefficients, low-frequency trend information corresponding to low-frequency bands at high scales is captured. The output sequence is stored as a floating-point array of time-series data. At each time point, the coefficient value sequence at each level is stored, aligned with the original signal time. After the sequence is generated, it is directly used for subsequent feature fusion without further processing. The level allocation logic includes frequency range mapping, converting scale values ​​into frequency values. Small-scale high-frequency wavelet coefficients mainly contribute to the detail subsequence, while large-scale low-frequency coefficients contribute to the approximation subsequence, ensuring that the decomposition result reflects the multi-resolution characteristics of the signal. The scale-frequency correspondence is determined based on the center frequency of the wavelet basis function used. The center frequency of the mixed Morette wavelet is a specific constant by default, usually set to a typical value to maintain versatility. The frequency band coverage must include the main characteristics of the building load, achieved by adjusting the minimum and maximum scale values. Time shift control ensures that the coefficient values ​​cover the entire time domain, with the shift step equal to the sampling interval to avoid information loss. The calculation process is performed on a digital signal processor. After the sampled signal is discretized, numerical integration is performed, using the trapezoidal rule to approximate the continuous integral formula. The integration variable is discretized at each time point. In real-time processing mode where the sequence generation period is synchronized with the signal input period, all coefficients are recalculated when a new signal segment is added. The output subsequence data structure includes an index array to identify the hierarchical type and a timestamp to ensure collaborative management of multi-level sequences. The signal processing includes error control measures, maintaining numerical calculation accuracy above single precision and avoiding boundary effects during integration. The final generated approximate coefficient sequence and detailed coefficient sequence have the same length and correspondence in time points, and are directly input into subsequent steps.

[0030] Example 2: See Figure 2 This embodiment focuses on feature extraction and multidimensional feature matrix construction using the Hilbert-Huang transform. After the building power signal is input, empirical mode decomposition (EMD) is first performed. This process adaptively generates several intrinsic mode functions (EMFs) and the final residual components based on the signal's local extremum characteristics. All local maxima and minima of the signal are identified, and upper and lower envelopes are constructed using cubic spline interpolation. The arithmetic mean of the upper and lower envelopes is calculated to obtain the envelope mean signal. Subtracting this envelope mean from the original signal yields the first candidate component. The candidate component is verified to meet two basic conditions for EMFs: the difference between the number of extrema and the number of zero-crossings does not exceed one, and the symmetry of the upper and lower envelopes at any given time point satisfies a set threshold. If the conditions are met, it is recorded as a valid EMF; otherwise, the candidate component is used as a new input signal, and the envelope construction and filtering process is repeated. Subtracting this EMF from the original signal forms a residual signal. The same decomposition operation is iteratively performed based on this residual signal until the residual component becomes a monotonic function or a constant term. The decomposition termination conditions include the residual amplitude being less than a preset threshold or the component energy change rate being less than 0.1% after three consecutive iterations. The final output includes multiple sets of intrinsic mode functions and one residual component.

[0031] The intrinsic mode functions (IMFs) are processed using a Hilbert transform: For each IMF time series, its Hilbert transform result is calculated, implemented using the Cauchy principal value integral principle. In the discrete form processing, a finite impulse response (FIR) filter is used to process the IMF sampling point sequence. The filter coefficients are designed based on the Hilbert kernel function, and their length is set to one-tenth of the number of sampling points. The original IMFs and their Hilbert transform results are combined to form a complex analytic signal. The real part of the analytic signal equals the original function value, and the imaginary part corresponds to its Hilbert transform result. The magnitude of this complex analytic signal is calculated to obtain the instantaneous amplitude characteristics at each time point. The magnitude is calculated using the Euclidean norm form: taking the square root of the sum of the squares of the real and imaginary parts. Simultaneously, the instantaneous phase angle of the complex analytic signal is calculated using the four-quadrant arctangent function to calculate the phase angle in radians corresponding to the ratio of the real to the imaginary part. The first-order difference of the phase angle difference between consecutive time points is calculated and divided by twice the parameter of pi to obtain the instantaneous frequency characteristic sequence. Each intrinsic mode function generates an instantaneous amplitude time series and an instantaneous frequency time series. Typically, the results of the three intrinsic mode functions with the highest energy proportions are selected as the final output features, while other components are discarded as appropriate.

[0032] In terms of time dimension processing, a fixed-duration window is set to divide the signal, with a uniform window width of ten minutes and 50% overlap between adjacent windows. Statistical feature calculations are performed on all wavelet decomposition products within the window (including the approximate coefficient subsequence and detail coefficient subsequence output in Example 1). Each coefficient subsequence is processed independently: the maximum absolute value within the window is extracted as the amplitude feature; the square value sequence of all coefficients is calculated, and the square root of the arithmetic mean of this sequence is taken as the root mean square feature; the dispersion of the coefficient value sequence from its arithmetic mean is calculated, specifically by dividing the sum of the squares of the differences between each coefficient value and the mean by the number of coefficients to obtain the variance feature. These three statistical features are generated separately for each subsequence of each wavelet decomposition level. If the total number of wavelet decomposition levels is six, twelve sets of coefficient subsequences are generated (six levels of approximate coefficients plus six levels of detail coefficients), and a total of thirty-six features are calculated for each subsequence using the three statistical measures.

[0033] The instantaneous features generated by the Hilbert-Huang transform are simultaneously windowed: within the same time window, the mean and standard deviation of each instantaneous amplitude sequence are calculated; the median and interquartile range of each instantaneous frequency sequence are also calculated. Each intrinsic mode function corresponds to two sets of features (one set for instantaneous amplitude features and one set for instantaneous frequency features). If three intrinsic mode functions are selected, six sets of window statistics are generated. Outliers in instantaneous frequencies need to be filtered before window statistics are performed. The physical frequency constraint range is set to 0.1-100 Hz, and instantaneous frequency values ​​outside this range are replaced by linear interpolation of adjacent effective values.

[0034] The feature normalization process is implemented according to global statistical characteristics: the mean and standard deviation of each feature dimension are calculated using the complete training dataset. Each dimension of all features is processed independently; each feature value is subtracted from the mean of the training set for that feature dimension and then divided by its training set standard deviation. If the feature distribution has significant skewness, a logarithmic transformation is applied for preprocessing before standardization. The normalization parameters are calculated and stored during the training phase.

[0035] The multidimensional feature matrix is ​​constructed according to the time alignment principle: using a ten-minute window as the basic unit, each time window corresponds to one row of data in the feature matrix. The feature dimensions include thirty-six wavelet statistical features (three statistics for each of the six-level decomposed dual sequences), six Hilbert-Yellow window statistical features (two features for each of the three intrinsic mode functions), and device on / off event counts. Device on / off event data is obtained from independent monitoring units, and the cumulative number of device state transitions (on→off or off→on) within the ten-minute window is used as an additional dimension. The final feature matrix has forty-four columns per row (36+6+1=43), and the time dimension is the total monitoring duration divided by the window step size (five-minute step). The matrix data is stored as a two-dimensional floating-point array indexed by timestamps, and missing values ​​are handled using a forward-filling strategy.

[0036] The feature storage structure and transmission protocol follow specific specifications: matrix data is stored in binary format, and the file header contains feature dimension definitions and normalization parameters. Real-time transmission uses fixed-length data packets, each containing the feature vector for a complete time window. Timestamp accuracy is down to the millisecond level, synchronized with the original power signal acquisition clock. Matrix data is directly input into subsequent deep neural network models without additional preprocessing steps. The data buffer uses a circular queue management system, where new window data overwrites the oldest data, maintaining a fixed-length time series. A data range monitoring mechanism for each feature dimension detects numerical overflow in real time, automatically switching to a backup normalization parameter set when an anomaly is triggered.

[0037] Example 3: See Figure 3 This embodiment relates to the structural implementation of a deep neural network model. The model input is a multi-dimensional feature matrix, the dimensions of which are determined by the number of samples in the batch, the time step, and the feature dimensions. The data organization is set as a three-dimensional tensor: the first dimension represents the number of samples in the batch, the second dimension corresponds to the total number of time steps, and the third dimension equals the total number of features. After receiving this tensor, the input layer directly passes it to the subsequent processing layer.

[0038] The core of the network comprises a bidirectional Long Short-Term Memory (LSTM) layer. This layer consists of two independent LSM sequence processing modules: a forward processing module and a backward processing module. The forward processing module processes the input sequence in forward time order, receiving the current feature vector and the hidden state from the previous time step at each time step. After updating the gating mechanism to control the information flow, it outputs a new hidden state. The backward processing module processes the sequence in reverse time order, receiving the current feature vector and the hidden state from the next time step at each time step. It also generates the hidden state after passing through forget gates, input gates, and output gates. The hidden state dimension is set the same for both processing methods, typically a 128-dimensional or 256-dimensional floating-point vector. The output hidden state sequences of the forward and backward modules are concatenated in the time dimension: the forward and backward hidden state vectors at each time step are joined end-to-end to form a complete bidirectional hidden state representation. The concatenated result constitutes a new tensor with a dimension equal to the time step size multiplied by twice the unit dimension.

[0039] The attention mechanism layer processes the aforementioned bidirectional hidden state sequence. This layer introduces a trainable weight parameter matrix, the dimension of which is equal to the length of the hidden state vector. An attention score is calculated for each time step by multiplying the corresponding bidirectional hidden state vector by the weight matrix and then transforming it using a non-linear activation function. The attention weight allocation is represented by the following calculation: Meaning of each character in the formula: : No. Normalized attention weights at each time step, scalar output; Time step index, an integer value ranging from 1 to the total time step. ; : Natural exponential function; : No. The unnormalized attention score at each time step, a scalar value; Hyperbolic tangent activation function; : No. The bidirectional hidden state vector at each time step, with dimension ; : The trainable weight matrix of the attention layer, dimension ; : Bias vector of the attention layer, dimension ; : Hidden state dimension of unidirectional long short-term memory unit; : The total time step of the input sequence.

[0040] The attention weight sequence is weighted and summed with the original hidden state sequence to generate a context vector. This vector integrates global information from the time series, while maintaining the same dimensionality as the hidden state vector.

[0041] The fully connected layer receives a context vector and performs multi-task output processing. The first output branch predicts the total load: the context vector is input to an intermediate fully connected layer with 64 neurons, linearly activated, and then connected to a single-neuron output layer. The output value is a scalar floating-point number representing the predicted total building load power. The second output branch processes device-level load probabilities: the context vector is input to another fully connected layer with the number of neurons matching the total number of device types. The output layer uses the Softmax function to calculate the probability distribution for each class. Each neuron's output value represents the probability that the corresponding device is active at the current time, and the sum of all device probabilities is strictly 1.0.

[0042] The network parameters are initialized as the starting point for training. The weights are initialized using a Xavier normal distribution strategy: sampled from a normal distribution with a mean of zero and a standard deviation equal to the reciprocal of the square root of the number of input and output neurons. The bias term is initialized to a constant value of zero. This initialization method maintains the stability of the output variance of the activation function.

[0043] The tensor operations performed by the model follow specific dimensionality transformation rules. The input feature tensor dimension is batch size multiplied by time step size multiplied by the number of features. After passing through a bidirectional long short-term memory (LSTM) layer, the output dimension becomes batch size multiplied by time step size multiplied by twice the number of hidden units. The attention layer compresses it into a two-dimensional tensor of batch size multiplied by twice the number of hidden units. The fully connected layer operates independently on each sample in the batch: the total load output branch produces a tensor of batch size multiplied by 1 dimension; the device-level output branch produces a tensor of batch size multiplied by the number of device types. Gradient calculation is automatically performed based on the backpropagation principle, and all tensor transformations are differentiable operations. The device-level probability distribution output uses one-hot encoding to label the target value, and the relative entropy loss is calculated during training. The entire network structure is implemented in the form of a computational graph; intermediate variable values ​​are not retained during inference, only the final output prediction result is stored.

[0044] Example 4: See Figure 4 This embodiment focuses on joint loss function optimization and real-time feature updating. During the operation of the building load forecasting system, assume that the system is deployed in an office building, and the current time is 10:00 AM on July 28, 2025. The system processing flow is as follows: Newly acquired power signal sequences (sampling frequency 1Hz) are continuously input into the buffer, triggering the dynamic feature update mechanism. The system sets the time window length to 10 minutes. When the latest 6 sampling point data arrives (corresponding to the time period from 10:00:00 to 10:00:05), the following operations are performed: Feature recalculation: Extract the power signal within the current window (09:50:00-10:00:05): Recalculate the wavelet multi-scale decomposition: Use the hybrid Morette wavelet basis function to generate the approximate coefficients A7(t) and detail coefficients D7(t) at the latest scale. Perform Hilbert-Huang transform update: extract the instantaneous amplitudes a1(t)-a3(t) and instantaneous frequencies f1(t)-f3(t) of the first three eigenmode functions in the current window. Sliding window update: The feature matrix removes the oldest 5-minute data rows (corresponding to 60 rows from 09:50:00 to 09:54:55) and adds 60 newly generated feature data rows (09:55:00 to 10:00:05). After the update, the feature matrix maintains 240 rows (covering 40 minutes of data).

[0045] Normalization: Standardize the newly added features by calling pre-stored normalization parameters (mean and standard deviation are from the training set): Subtract the mean of the training set from each feature dimension The result divided by the standard deviation of the training set For example, the normalization of the wavelet detail coefficient D5 is: (D5-0.12) / 1.8 The following is an example of the feature matrix after real-time updates (only some time-point data for key columns are shown): Model training optimization details: When the system is in training mode, joint loss function optimization is performed. The training dataset for the office building contains 30 days of load records. Each iteration follows this process: The mean square error of the total load forecast branch is calculated by comparing the predicted total power output by the model with the actual meter reading. For example, if the actual value is 40.2kW and the model output is 39.8kW, then the model contributes a loss component of (39.8-40.2)²=0.16.

[0046] Device-level decomposition branch calculation of cross-entropy: When the actual state of the air conditioning device is "on" (marked [1,0]), and the model output probability distribution is [0.85,0.15], the cross-entropy is calculated as: -(1×log(0.85)+0×log(0.15))≈0.163 Joint loss weighting: Set the total load loss weight λ = 0.7, and the equipment-level loss weight 1 - λ = 0.3. Assuming the average total load MSE of the current batch of samples is 0.25 and the average equipment cross-entropy is 0.42, then the joint loss value = 0.7 × 0.25 + 0.3 × 0.42 = 0.301 The optimizer uses the Adam algorithm with a base learning rate of 0.001. The dynamic learning rate adjustment mechanism is based on historical error: maintaining the exponentially weighted average (EMA) of the validation set mean squared error. When the EMA cumulatively increases by more than 5% over three consecutive iterations, for example: After 100 iterations: EMA = 0.215 After 101 iterations: EMA = 0.219 (increase of 1.86%) After 102 iterations: EMA=0.227 (an increase of 3.65%, accumulating to over 5%), the learning rate decays to 0.0005 with a decay factor of 0.5.

[0047] Real-time prediction process: The updated feature matrix is ​​immediately fed into the pre-trained model for inference. The model output includes: Total load forecast: 42.3kW at the current time. Device state probability distribution: Air conditioning unit: [Probability of operation: 0.92] Lighting system: [Probability of being off: 0.85] Elevator equipment: [Operation probability 0.15] When a sudden change in total load is detected (the difference between adjacent time points exceeds the dynamic threshold δ=1.5×σ_24h), such as a sudden increase in power of 3.8kW (δ=2.6kW) at 10:00:10: Extracting abrupt waveform features: duration 4.2 seconds, slope 0.91 kW / s Matching device library: The typical waveform similarity of the air conditioner start-up event reaches 87% (matching threshold 70%). Update device status: Change air conditioner status from off to on. The system outputs the final prediction results to the monitoring platform, including the time-series total load curve and equipment status switching event records. All parameter update operations are completed within 200 milliseconds, meeting the real-time prediction time constraint. Historical prediction errors are recorded in a recurring database for hyperparameter tuning during monthly model retraining.

[0048] Example 5: This example focuses on the equipment-level load decomposition and dynamic adjustment mechanism of model parameters. Equipment status identification is initiated based on the detection of power signal mutations. This process sets a dynamic threshold mechanism: the system automatically calculates the standard deviation of the total building load data in the most recent 24 hours, and multiplies this standard deviation by a coefficient of 1.5 as the mutation judgment threshold. Taking a commercial building monitoring scenario as an example, the calculated standard deviation of the total load from 11:00 on a certain day to 11:00 on the next day is 2.3 kW, so the current threshold δ = 3.45 kW. When the difference between two adjacent 30-second intervals of total load exceeds this threshold, it is marked as a mutation event and a timestamp is recorded. The time point at which the power mutation is detected is the center position of the mutation event, and an analysis window is formed by extending a fixed duration (default 0.5 seconds) before and after it.

[0049] Power waveform feature extraction is performed within this analysis window: a power sampling sequence is extracted from 0.5 seconds before to 0.5 seconds after the abrupt change point, and key feature parameters are calculated. The power change is the algebraic difference between the power value at the end of the window and the power value at the beginning; the waveform duration is defined as the time required for the power change to reach 90% of its final value; the waveform rise / fall slope is calculated by dividing the power change by the duration. Simultaneously, waveform curvature parameters are calculated: the power change curve is divided into five equal segments based on time, and the variance of the slope in each segment is compared. For example, if a power increase of 4.8 kW is detected in a sudden event, with a duration of 3.2 seconds, then the slope = 1.5 kW / s; the curvature variance is calculated to be 0.18.

[0050] The device feature library contains the electrical fingerprint parameters of pre-registered devices. Each device stores three sets of core features: the steady-state power average value records the power level during normal operation; the power fluctuation standard deviation reflects the intensity of fluctuations during operation; and the typical switching event waveform template stores time-power data sequences of more than 10 normal switching cycles. These features are generated through individual tests during device registration: air conditioning equipment tests show an average power of 3.5 kW and a standard deviation of 0.4 kW; lighting systems show an average power of 0.8 kW and a standard deviation of 0.05 kW. The waveform template collects complete waveforms from at least 10 normal start-stop events.

[0051] The feature matching process employs a multi-level screening strategy: The first level filters based on power change, selecting candidate devices from the device library whose average power differs from the sudden change by less than 20%. For example, if a power increase of 4.8 kW is detected, devices with an average power between 3.84 and 5.76 kW are selected. The second level calculates waveform similarity using a dynamic time warping algorithm to measure the similarity between the current event waveform and the candidate device template waveform. Similarity scores range from 0 to 100%, and the device with the highest similarity score is considered the match. For example, if the air conditioner waveform has a similarity of 87% and the elevator waveform has a similarity of 43%, it is determined to be an air conditioner startup event.

[0052] Matching results trigger device state transitions: When the matching similarity exceeds a threshold of 85%, the target device state is updated. State transition records contain a four-tuple of data: timestamp, device type, operation type (on / off), and power change. The system maintains a state cache table for all devices, updated once per second. Device state changes trigger load allocation logic: new power changes are attributed to the device undergoing the state transition; other devices maintain their load forecasts from the previous cycle. Device-level prediction results are output in a state list format: Air conditioning = running, elevator = out of service, lighting = on.

[0053] The dynamic adjustment of model parameters is implemented in two parts: The learning rate adjustment is based on the exponentially weighted moving average of the validation set error. The mean of the latest window prediction error is calculated hourly and compared with the exponentially averaged mean of the errors over the previous 24 hours. A decay factor of 0.9 is used in the current error exponential average calculation formula; when the current value exceeds the historical average by 5%, the system automatically reduces the learning rate to 50% of its original value. The attention mechanism layer weights are automatically optimized through gradient backpropagation: after each new sample input, the error between the predicted output layer and the true label is calculated, and the error signal is backpropagated to the attention weight matrix to trigger parameter updates. The attention weight adjustment cycle is synchronized with model training.

[0054] Model retraining cycles are triggered by the device status misclassification rate: when the number of device status misclassifications exceeds 10% of the total number of events within 24 consecutive hours, full model retraining is initiated. Retraining uses data from the most recent 14 days, retaining 20% ​​as a validation set. Training terminates when the validation set loss shows no improvement for three consecutive rounds. After training, the model is automatically updated and deployed. The new model must undergo a 72-hour observation period to verify its stability before it can be officially replaced.

[0055] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0056] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A building load prediction method based on deep neural networks and non-invasive monitoring, characterized in that, include: Acquire total building load data and non-invasive monitoring data, including voltage signals, current signals, power signals, and equipment switching event data; The power signal is decomposed into approximate coefficient subsequences and detail coefficient subsequences by wavelet transform. Instantaneous frequency and instantaneous amplitude characteristics of power signals are extracted based on Hilbert-Huang transform; By integrating the approximate coefficient subsequence, detail coefficient subsequence, instantaneous frequency features, and instantaneous amplitude features, a multidimensional feature matrix is ​​constructed. The multidimensional feature matrix is ​​input into a deep neural network model for training. The deep neural network model includes a bidirectional long short-term memory network layer and an attention mechanism layer. The total load forecast results are decomposed into equipment-level load forecast results using non-invasive load monitoring technology, and the forecast results of total load and equipment-level load are output.

2. The building load prediction method based on deep neural networks and non-invasive monitoring according to claim 1, characterized in that, The multi-scale decomposition of the power signal using wavelet transform specifically includes: A hybrid Morette wavelet basis function is used to perform continuous wavelet transform on the power signal. By adjusting the scaling and translation parameters, approximate coefficient subsequences and detail coefficient subsequences with different decomposition levels are extracted.

3. The building load prediction method based on deep neural networks and non-invasive monitoring according to claim 2, characterized in that, The extraction of instantaneous frequency features based on Hilbert-Huang transform includes: Empirical mode decomposition is performed on the power signal to obtain multiple intrinsic mode functions and residuals; Perform a Hilbert transform on each intrinsic mode function to generate an analytic signal and calculate its instantaneous frequency and instantaneous amplitude.

4. The building load prediction method based on deep neural networks and non-invasive monitoring according to claim 3, characterized in that, The construction of the multidimensional feature matrix includes: Calculate the maximum value, root mean square value, and variance of the approximate coefficient subsequence and the detailed coefficient subsequence within each time window; The statistical features, instantaneous frequency features, and instantaneous amplitude features are combined to generate a multidimensional feature matrix.

5. The building load prediction method based on deep neural networks and non-invasive monitoring according to claim 4, characterized in that, The structure of the deep neural network model includes: The input layer receives a multidimensional feature matrix; Bidirectional long short-term memory network layers are used to capture temporal dependencies and output hidden state sequences; The attention mechanism layer performs weighted aggregation on the hidden state sequence and then outputs the total load prediction and device-level load probability distribution through the fully connected layer.

6. The building load prediction method based on deep neural networks and non-invasive monitoring according to claim 5, characterized in that, The deep neural network model uses a joint loss function to optimize the model parameters. The joint loss function is a weighted combination of the mean square error of total load prediction and the cross-entropy loss of equipment-level decomposition. The contributions of the two types of losses are balanced through hyperparameters.

7. The building load prediction method based on deep neural networks and non-invasive monitoring according to claim 6, characterized in that, The dynamically updated feature matrix includes: The wavelet coefficients and instantaneous frequency characteristics of the current time window are calculated based on the real-time power signal. The feature matrix is ​​then updated and the input data is renormalized using a sliding window mechanism.

8. The building load prediction method based on deep neural networks and non-invasive monitoring according to claim 7, characterized in that, The equipment-level load decomposition method includes: Detect sudden changes in power signals and determine changes in device status based on threshold values; A pre-trained database of device features is used to match power mutation patterns to generate device-level load prediction results.

9. The building load prediction method based on deep neural networks and non-invasive monitoring according to claim 8, characterized in that, The method for adjusting the model weight parameters includes: The learning rate and the weight allocation of the attention mechanism layer are dynamically adjusted based on the exponentially weighted moving average of historical prediction errors.

10. A building load prediction system based on deep neural networks and non-invasive monitoring, characterized in that, include: The data acquisition module is used to acquire total building load data and non-intrusive monitoring data, including voltage signals, current signals, power signals and equipment switching event data. The signal processing module is used to perform multi-scale decomposition of the power signal through wavelet transform to obtain approximate coefficient subsequences and detail coefficient subsequences; The feature extraction module is used to extract the instantaneous frequency and instantaneous amplitude features of the power signal based on the Hilbert-Huang transform. The feature fusion module is used to fuse the approximate coefficient subsequence, the detail coefficient subsequence, the instantaneous frequency feature, and the instantaneous amplitude feature to construct a multidimensional feature matrix; The model training module is used to input the multidimensional feature matrix into a deep neural network model for training. The deep neural network model includes a bidirectional long short-term memory network layer and an attention mechanism layer. The load decomposition and output module is used to decompose the total load forecast results into equipment-level load forecast results using non-intrusive load monitoring technology, and output the forecast results of total load and equipment-level load.