Power consumption dynamic prediction method based on physical-semantic dual-channel modulation
The dynamic power consumption prediction method based on physical-semantic dual-path modulation, which combines physical and semantic sensing paths, solves the problems of insufficient accuracy and adaptability of existing power load prediction methods when dealing with high-frequency, non-stationary, and multi-user heterogeneous data, and achieves high-precision and robust power load prediction.
Patent Information
- Application Number
- CN202511851328.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2045-12-10
AI Technical Summary
Existing power load forecasting methods struggle to simultaneously capture local abrupt change patterns and long-term semantic dependencies when processing high-frequency, non-stationary, and multi-user heterogeneous power consumption data, resulting in insufficient forecast accuracy and adaptability.
A dynamic prediction method for power consumption based on physical-semantic dual-path modulation is adopted. The method extracts time-frequency physical features and contextual semantic features in parallel through physical sensing and semantic sensing paths, and performs dynamic adaptive prediction by combining the liquid core computing architecture.
It achieves high-precision and robust prediction of non-stationary power load data, improving the model's prediction performance in complex power load scenarios, especially in challenging scenarios such as sudden changes in power consumption peaks and holiday patterns.
Smart Images

Figure CN121303468B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence and power grid management, and particularly relates to a power consumption dynamic prediction method based on physical-semantic dual-channel modulation. BACKGROUND
[0002] Power load prediction is the core link of power system dispatching, energy management and market transaction. With the popularization of smart meters, high-frequency and multi-user power consumption data show strong non-stationarity, multi-time scale characteristics and user behavior heterogeneity, and traditional time series models (such as LSTM, RNN) are difficult to capture both local mutation patterns and long-term semantic dependencies.
[0003] The existing power load prediction method has four significant limitations when dealing with high-frequency, non-stationary and multi-user heterogeneous power consumption data. First, the serial architecture of traditional wavelet transform and neural network leads to the separation of feature extraction and prediction modeling: the time-frequency features obtained by wavelet decomposition are input into the network in a static way, which cannot be integrated into the network dynamic evolution process, and this "decomposition first and then prediction" pipeline blocks the end-to-end gradient propagation, so that the physical priori cannot be optimized adaptively according to the prediction target. Second, although the model based on attention mechanism can capture the global dependence of the sequence, it lacks understanding of the physical nature of power signals, and its pure data-driven characteristics lead to a lag response to power load mutation, periodic oscillation and other physical phenomena, making it difficult to achieve fast and instinctive transient response. Third, the fixed network structure of traditional LSTM and RNN determines the time constant or receptive field after training, and lacks a mechanism for dynamic modulation according to the characteristics of the input signal, so it cannot simultaneously achieve "fast reflection" for sudden changes and "deep thinking" for long-term patterns in a single model. In addition, although the existing liquid neural network has continuous time dynamics, its parameters are usually constant throughout the time series, lacking the ability to adapt to the multi-scale spatio-temporal characteristics of power load, making it difficult to adaptively switch between rapidly changing load scenarios and stable normal modes. These systematic deficiencies collectively limit the accuracy and adaptability of existing methods in complex non-stationary power load prediction. SUMMARY
[0004] In order to overcome the defects of insufficient dynamic adaptability of power load prediction in the prior art, the present application proposes a power consumption dynamic prediction method based on physical-semantic dual-channel modulation, which can integrate physical priori and semantic understanding, has dynamic adaptive ability, and realizes high-precision and high-robustness prediction of non-stationary power load data.
[0005] The power consumption dynamic prediction method based on physical-semantic dual-channel modulation proposed by the present application first trains a power consumption prediction model based on power load time series data predicts power consumption ;
[0006] The electricity consumption prediction model includes: a physical sensing path, a semantic sensing path, a fusion module, and a liquid core computing architecture; the physical sensing path extracts the input sequence based on wavelet transform. The time-frequency characteristics are used to generate physical modulation signals. The semantic perception pathway is based on the attention mechanism for signals. Processing is performed to generate semantically modulated signals. The physical modulation signal and the semantic modulation signal are fused by the fusion module, and then processed by the liquid core computing architecture to obtain the power consumption prediction result. ;
[0007] Then, let the electricity consumption prediction model iterate through t={t0,t0+1,……,t0+C}, based on {x t-w+1 ,x t-w+2 ,…,x t Predict x t+1 The electricity consumption prediction sequence {x} is obtained. t0+1 , x t0+2 , …, x t0+C}; where t0 is the current time step, C is the set total number of prediction steps; w is the window width; x t-w+1 x t-w+2 x t x t+1 x t0+1 x t0+2 and x t0+C These represent the power consumption at time steps t-w+1, t-w+2, t, t+1, t0+1, t0+2, and t0+C, respectively.
[0008] Preferably, the physical sensing pathway includes a wavelet transform module, an instantaneous feature extraction module, and a physical dynamic parameter generator connected in sequence;
[0009] The wavelet transform module performs a discrete wavelet transform on the input sequence to decompose the wavelet coefficients, including approximation coefficients. and multi-scale detail coefficients { ,1≤j≤J}, where J is the total number of wavelet decomposition layers;
[0010] The instantaneous feature extraction module calculates the frequency band energy based on wavelet coefficients. ,1≤j≤J}、Center of Spectrum And the transient impact index TI(t), and construct the physical feature vector. ; , and These represent the energy of the first, j, and J-th frequency bands, respectively. the energy of the approximation coefficients of the j-th layer;
[0011] The physical dynamic parameter generator transforms the physical feature vector into a physical modulation signal .
[0012] Preferably,
[0013] ;
[0014] ;
[0015] ;
[0016] ;
[0017] wherein, is the length of the j-th layer detail coefficient ; is the k-th value in ; is the k-th value in the j-th layer approximation coefficient ; is the center frequency weight of the j-th frequency band, .
[0018] Preferably, the semantic perception pathway comprises a context encoder, a multi-head attention module and a semantic dynamic parameter generator connected in sequence;
[0019] The context encoder encodes the signal x t to obtain its hidden state representation h t ;
[0020] The multi-head attention module processes the hidden state representation h t based on a scaled dot-product attention mechanism to obtain an attention head output fusion that evaluates the importance of the context, and transforms into a context perception representation ;
[0021] ;
[0022] wherein, and represent the outputs of the 1st and h-th attention heads respectively, h being the number of attention heads, is an output projection matrix; represents concatenation;
[0023] The semantic dynamic parameter generator maps into a semantic modulation signal .
[0024] Preferably, the fusion module first processes the concatenation vector of the physical modulation signal and the semantic modulation signal using a gating network, then activates the gating signal, and finally performs weighted fusion on the physical modulation signal and the semantic modulation signal using the gating signal as the weight.
[0025] Preferably, the processing process of the liquid core computing unit is represented by the following formula:
[0026] ;
[0027] ;
[0028] ;
[0029] wherein, is the power consumption at time step t, is the power consumption prediction value at time step t+1, is the step size; is the weighted fusion result output by the fusion module, and tanh is the hyperbolic tangent function; is a nonlinear activation function; and are the system time constant and the bias term, respectively; and are learnable static parameters; and are modulation gain hyperparameters; and x(t) is a hidden state vector calculated using an ordinary differential equation.
[0030] Preferably, the power consumption prediction model is trained on a data set {X}, X = {x1, x2, …, xT}, X is a historical sequence, x1, x2, and xT are power consumption sample values at time steps 1, 2, and T, respectively. T T
[0031] The power consumption model is based on the sequence The power consumption at time step t+1 is denoted as ; the power consumption model iterates through t = w, w+1, …, T-1 to generate a prediction sequence = {x' w+1 , x' w+2 , …, x' T}, x' w+1 , x' w+2 , and x' T These are the predicted power consumption values at time steps w+1, w+2, and T, respectively.
[0032] In the historical sequence y={x w+1 ,x w+2 ,…,x T} and predicted sequence ={x' w+1 ,x' w+2 ,…,x' T The loss function is calculated on the x-axis, and the model parameters are updated through backpropagation using the loss function until the model converges; w+1 and x w+2 These are the actual power consumption values at time steps w+1 and w+2, respectively.
[0033] Preferably, the loss function for:
[0034] ;
[0035] ;
[0036] ;
[0037] in, This represents the average loss across variables. Represents the regularization term;
[0038] C represents the total number of predictor variables, and N represents the total number of users; and Let these represent the actual value and the predicted value of the c-th variable for the i-th user, respectively. for and Huber's losses; This represents all learning parameters; Represents the learning parameters, It is the regularization coefficient.
[0039] The present invention proposes a dynamic power consumption prediction system based on physical-semantic dual-path modulation, comprising a memory and a processor. The memory stores a computer program, and the processor is connected to the memory. The processor is used to execute the computer program to realize the dynamic power consumption prediction method based on physical-semantic dual-path modulation.
[0040] The present invention proposes a storage medium storing a computer program, which, when executed, is used to implement the power consumption dynamic prediction method based on physical-semantic dual-path modulation.
[0041] The advantages of this invention are:
[0042] (1) The power consumption dynamic prediction method based on physical-semantic dual-channel modulation is proposed, a liquid-state neural network architecture of physical-semantic dual-channel modulation is constructed, and the time-frequency physical characteristics and context semantic characteristics of the non-stationary power load sequence are extracted and dynamically fused in parallel. The physical perception channel gives the model the ability to quickly respond to transient patterns such as load mutations, while the semantic perception channel ensures deep mining of long-term periodic patterns, and the two are adaptively coordinated through gate fusion, thereby comprehensively improving the modeling and prediction accuracy of complex non-stationary power load sequences.
[0043] (2) In the present application, the time-frequency physical characteristics of the signal are extracted in real time through wavelet transform, which overcomes the limitations of traditional time series models in responding to the essential physical structure of the signal, and realizes the deep fusion of physical priori and data-driven. The present application significantly improves the prediction performance of the model in complex power load scenarios, especially in handling challenging scenarios such as peak load mutations and holiday patterns.
[0044] (3) The present application designs a unique parallel feature extraction and gate fusion mechanism in the dual-channel modulation system. In the physical perception channel, the time-frequency energy distribution, spectral center of gravity and transient impact indicators of the signal are extracted through multi-scale wavelet transform, which have clear physical meaning; in the semantic perception channel, the context semantic importance of the sequence is extracted through the multi-head attention mechanism. The dual-channel output is dynamically weighted through a learnable gate network to form the final modulation signal. This design has both the ability to quickly respond to load mutations and the ability to deeply understand long-term periodic patterns, and the gate mechanism realizes the adaptive balance of channel contribution, greatly improving the modeling accuracy of the model for the dynamic characteristics of non-stationary power loads.
[0045] (4) The liquid-state neural network core used in the present application has dynamic adjustable system characteristics, which adjusts the system time constant and bias item in real time through the modulation signal output by the fusion module, so that the network can adaptively switch between "fast reflection" and "deep thinking" modes. The entire system uses an end-to-end training method and is optimized based on a multivariate Huber loss function, ensuring training stability and robustness to outliers. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 The power consumption prediction model structure diagram proposed by the present application;
[0047] Figure 2(a) is the RMSE statistics of different models in the embodiment;
[0048] Figure 2(b) is the MAE statistics of different models in the embodiment;
[0049] Figure 2(c) is the R 2statistics;
[0050] Figure 2(d) shows the MAPE statistics for different models in the examples;
[0051] Figure 3 This is a comparison chart of time series predictions using the model of the present invention in the embodiments;
[0052] Figure 4 This is a scatter plot showing the correlation between predicted and actual values of the model of the present invention in this embodiment.
[0053] Figure 5(a) shows the residual statistics of the predicted values of the model of the present invention in the embodiment;
[0054] Figure 5(b) shows the residual distribution of the model of the present invention in the embodiment;
[0055] Figure 5(c) is the residual QQ plot of the model of the present invention in the embodiment;
[0056] Figure 6 This invention provides a model training method;
[0057] Figure 7 This invention provides a method for dynamic prediction of power consumption based on physical-semantic dual-path modulation. Detailed Implementation
[0058] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0059] like Figure 1 As shown, this invention proposes a power consumption prediction model based on power load time-series data. Predicting electricity consumption . and They are time steps t- +1、t- Electricity consumption at +2, t, and t+1.
[0060] The electricity consumption prediction model includes: a physical sensing path, a semantic sensing path, a fusion module, and a liquid core computing architecture; the physical sensing path extracts the input sequence based on wavelet transform. The time-frequency characteristics are used to generate physical modulation signals. Semantic perception pathways extract signals based on attention mechanisms. Based on the contextual semantic importance, generate semantic modulation signals. The physical modulation signal and the semantic modulation signal are fused by a fusion module, and then are processed by a liquid core computing architecture to obtain a power consumption prediction result .
[0061] Let an input sequence at time t be denoted as I(t), then:
[0062] (1)
[0063] wherein, is an analysis window size, which can be adjusted according to signal characteristics, and in the embodiment, is set to .
[0064] The physical perception channel is a "reflex arc" system of the model, responsible for low-delay, instinctive physical feature extraction and modulation of the input signal. The channel converts the physical nature of the signal into modulatable parameters of the neural network dynamics through wavelet transform.
[0065] The physical perception channel includes sequentially connected wavelet transform modules, instantaneous feature extraction modules, and physical dynamic parameter generators.
[0066] The wavelet transform module performs discrete wavelet transform on the input sequence I(t) to decompose it into wavelet coefficients, including multi-scale detail coefficients and approximation coefficients, denoted as:
[0067] (2)
[0068] wherein, is the jth layer of detail coefficients (Detail Coefficients), which captures the high-frequency component of the signal at scale j, and the detail coefficients correspond to the frequency band one-to-one; 1≤j≤J; that is, 、 and represent the high-frequency components of the input sequence at scale 1, scale 2, and scale J, respectively; is the Jth layer of approximation coefficients (Approximation Coefficients), which captures the low-frequency trend component of the signal at scale J; J is the total number of wavelet decomposition layers, i.e., the number of frequency bands, and in the embodiment, J=5.
[0069] The mathematical expression of wavelet decomposition is:
[0070] (3)
[0071] wherein, represents the kth coefficient in the wavelet basis obtained by shifting and scaling the mother wavelet , represents the scale function The k-th coefficient in the wavelet basis function obtained after translation and scaling. In this embodiment, the DB4 (Daubechies-4) wavelet basis is selected because it has a good balance between time-frequency localization and computational efficiency. and They are respectively and The k-th value in the middle.
[0072] The instantaneous feature extraction module calculates the physical feature vector P(t) based on wavelet coefficients, quantizing the time-frequency structure of the input sequence into features with clear physical meaning. P(t) contains three core components: energy distribution in each frequency band { } 1≤j≤J Spectral centroid FC(t) and transient impact index TI(t).
[0073] Frequency band energy is denoted as { The energy of the j-th frequency band (detail coefficient) is denoted as ,1≤j≤J}. :
[0074] (4)
[0075] in, For the detail coefficients of the j-th layer Length, It is the k-th element of the level j detail coefficients. Let be the k-th element of the approximation coefficients at level J, where J is the total number of wavelet decomposition levels (J=5 in this embodiment). This formula calculates the normalized proportion of the energy of each frequency band (detail coefficient) in the total energy (the sum of the energies of all detail coefficients and approximation coefficients). This reflects the distribution of signal energy at different frequency scales and is a direct representation of the signal's spectral structure.
[0076] The centroid of the spectrum is denoted as:
[0077] (5)
[0078] in, The center frequency weight of the j-th frequency band (detail coefficients) is... This weighting strategy assigns higher frequency components (smaller j) to larger j. Values, low-frequency components (larger j) correspond to smaller values. Value. Spectral centroid It represents the main location where signal energy is concentrated in the frequency domain. The larger the value, the higher the dominant frequency component of the signal.
[0079] The transient impact index (high-frequency energy percentage) is denoted as:
[0080] (6)
[0081] The transient impact metric TI(t) calculates the proportion of energy in the two highest frequency bands (j=1,2) of the total energy. High-frequency energy is typically associated with signal abrupt changes, transients, or noise components. Therefore, TI(t) can effectively detect sudden events or impacts in a signal, providing the network with the ability to respond quickly to such patterns.
[0082] Finally, the physical feature vector P(t) is formed by concatenating the above features:
[0083] (7)
[0084] in, and These represent the energy of the first and J-th frequency bands, respectively. For approximate coefficient energy, .
[0085] It can be seen that the dimension of P(t) is In formula (7), the superscript T indicates transpose.
[0086] The physical dynamic parameter generator generates physical feature vectors. Through a learnable mapping function Perform nonlinear transformation to generate a physical modulation signal The mapping function is implemented using a two-layer multilayer perceptron (MLP):
[0087] (8)
[0088] in, and The weights and biases of the first layer MLP, and For the weights and biases of the second-layer MLP, Let P(t) be the dimension of the physical feature vector. ;d represents the hidden layer dimension of the liquid core computing architecture, i.e., the hidden layer dimension of the model backbone; This refers to the hidden layer dimension of the MLP inside the physical dynamic parameter generator, which is also the modulation signal. The dimensions; specific settings are available. The parameter is less than d to facilitate dimensionality reduction or feature compression and avoid over-parameterization; ReLU is the activation function used to introduce nonlinearity; the final output uses the tanh (hyperbolic tangent) activation function to... Each element of x is restricted in the range [-1, 1]. This design is crucial, which ensures the bounded amplitude of the modulated signal, thus maintaining the stability of the liquid-core computing unit dynamical system.
[0089] The semantic-aware pathway includes sequentially connected context encoder, multi-head attention module and semantic dynamic parameter generator.
[0090] The context encoder encodes the signal x using GRU (Gated Recurrent Unit) to obtain its hidden state representation, i.e. the final hidden state h for subsequent computation. t t The specific calculation process of GRU at time step t is as follows:
[0091] The reset gate is used to combine new input information with past memory:
[0092] (9)
[0093] where, represents an activation function, and represent the weight and bias of the reset gate at time step t, and represent the weight and bias of the reset gate at time step t-1, represents the hidden state of the data x t-1 processed by the context encoder.
[0094] The update gate is used to determine the degree of past information retained to the current state:
[0095] (10)
[0096] where, represents an activation function, and represent the weight and bias of the update gate at time step t, and represent the weight and bias of the update gate at time step t-1.
[0097] The candidate hidden state uses the reset gate to control the inflow of historical information:
[0098] (11)
[0099] where tanh is the hyperbolic tangent function; is the dot product; represents the reset gate; and denote the weights and biases at time step t, and denote the weights and biases at time step t-1;
[0100] the final hidden state is obtained by weighting sum of the previous time step hidden state and the candidate hidden state
[0101] (12)
[0102] In equations (9)-(12), the activation function is specifically a Sigmoid function, denotes element-wise multiplication, and W and b are the learnable weight matrix and bias term, respectively. where is the hidden layer dimension of the context encoder.
[0103] The multi-head attention module computes the encoded hidden representation based on the Scaled Dot-Product Attention mechanism, which evaluates the importance of each time step relative to the current context.
[0104] First, the query (Q), key (K), and value (V) vectors are obtained by linear projection:
[0105] (13)
[0106] where is the learnable projection matrix, , is the hidden layer dimension of the context encoder, is the dimension after projection.
[0107] The dot product of the query vector Q and the key vector K is computed, and a scaling factor is applied to stabilize the gradients, resulting in a score matrix S:
[0108] (14)
[0109] The score matrix S is normalized by row through the Softmax function to obtain the attention weight matrix A:
[0110] (15)
[0111] where, A ij and S ij are the elements of the i-th row and j-th column of the attention weight matrix A and score matrix S, respectively; j max is the column total, i.e., the length of the input sequence, and I(t) is the input sequence, max = .
[0112] Finally, the output of the attention mechanism, head, is obtained by weighted summing the value vectors V:
[0113] (16)
[0114] To enhance the expressive power of the model, a multi-head attention mechanism is adopted. h independent attention heads are used to compute in parallel, and their outputs are concatenated and fused by a linear transformation:
[0115] (17)
[0116] (18)
[0117] where, , and are the attention weight matrix and score matrix of the i-th attention head, , is the output of the i-th attention head, is the hidden representation of the context encoder output, and h is the number of attention heads; denotes the output fusion of the attention heads, and Concat denotes the concatenation operation; is the output projection matrix, is the hidden layer dimension of the context encoder, denotes the value vector output dimension of a single attention head, In this embodiment, the number of attention heads h = 4 is set.
[0118] In this embodiment, the context-aware representation z(t) output by the multi-head attention module adopts the global representation after pooling).
[0119] Semantic dynamic parameter generator: the context-aware representation z(t) output by the multi-head attention module is transformed by a learnable mapping function to generate a semantic modulation signal . The mapping function is usually implemented by a two-layer multi-layer perceptron (MLP):
[0120] (19)
[0121] where, This is the weight matrix. The bias term is denoted by , and ReLU is the activation function. The final output of the semantic dynamic parameter generator is constrained to the range [-1, 1] using the tanh activation function to ensure its stability as a modulated signal.
[0122] The fusion module adopts a gating fusion mechanism, the core of which lies in dynamically and adaptively balancing the modulation effects of physical perception pathways and semantic perception pathways on the liquid core computing architecture.
[0123] The fusion module receives modulated signals from the physical sensing path. and modulated signals from semantic perception pathways Where d is the hidden layer dimension of the liquid core computing architecture. First, the two modulation signals are concatenated along the feature dimension to form a joint feature vector that integrates physical and semantic information:
[0124] (20)
[0125] Subsequently, the gating signal is calculated through a gate network. This network is a small feedforward neural network whose purpose is to... (The sentence is incomplete and requires more context to translate accurately.) An optimal fusion weight is learned. The calculation process is as follows:
[0126] (twenty one)
[0127] (twenty two)
[0128] in, Here is the weight matrix of the gated network. For bias terms, This is a transitional term; The sigmoid activation function has an output range between [0,1], ensuring the gating signal... Each element represents a weight between 0 and 1.
[0129] Finally, the calculated gating signal is used. For physical modulation signals and semantic modulation signal Weighted fusion is performed to generate the final dynamic parameters used to modulate the liquid core. :
[0130] (twenty three)
[0131] Where σ is the Sigmoid function and W is the learnable weight matrix. denotes element-wise multiplication.
[0132] The gating mechanism realizes adaptive trade-off between physical and semantic awareness: when the system is biased towards the physical awareness channel, suitable for handling sudden changes; when the system is biased towards the semantic awareness channel, suitable for handling periodic patterns. The formula (23) ensures that for each dimension of the hidden state, the modulation parameter is a convex combination of the contributions from the physical channel and the semantic channel.
[0133] The liquid core computing architecture is the core to extract dynamic characteristics, whose behavior is described by a modulatable ordinary differential equation (ODE). The unique feature of this ODE is that its key parameters can be modulated by the final dynamic parameters in real-time and online, so that the dynamic characteristics of the network can adapt to the changes of the signal x t .
[0134] The continuous-time dynamics of this liquid core computing unit is described by the following ordinary differential (ODE) equation:
[0135] (24)
[0136] where, denotes the computed hidden state vector of the liquid core computing unit at time t; is the dynamically modulated system time constant, used to control the evolution speed of the state . The smaller the value, the faster the system response; The larger the value, the smoother the system dynamics; is the dynamically modulated bias term, used to provide a baseline or drive for the system dynamics. is a nonlinear activation function (such as tanh or ReLU), used to introduce a nonlinear transformation.
[0137] The dynamic modulation mechanism is the key to realize the adaptive capability of the present application. The basic parameters and are learnable static parameters, which, combined with the dynamic modulation signal , generate time-varying system parameters, i.e., the system time constant and the bias term .
[0138] The modulation formula for the system time constant is:
[0139] (25)
[0140] where, For the modulation gain hyperparameter, an embodiment sets it to 0.5, which controls the range of modulation strength; For the settable learnable base time constant, a good rule of thumb is to set it to Initialize it to an order of magnitude close to half of the dominant periodic component (the daily cycle), which is 24 hours, and can be set to = 12.
[0141] The tanh function ensures that the modulation quantity is limited in , so that and the system dynamics are stable. This design makes it so that when indicates that a fast response is needed (for example, the physical channel is dominant), decreases, and the system state changes faster; when indicates that a smooth processing is needed (for example, the semantic channel is dominant), increases, and the system dynamics are more gentle, which helps to capture long-term dependencies.
[0142] The modulation formula for the bias term is:
[0143] (26)
[0144] where is the modulation gain of the bias term, and can be set to 0.2. This allows the equilibrium point or driving signal of the network to be dynamically adjusted according to the context (physical or semantic). is the base bias term, which can be randomly set or not set, or can be initialized to 0.
[0145] In order to implement and perform end-to-end gradient backpropagation on a digital computer, the above continuous-time ODE needs to be discretized. The present application uses the forward Euler method, which is simple in form and computationally efficient. Let the discretization step size be (in this embodiment, for simplicity and to align with the data sampling interval, it is often taken to be ), then the state update formula after discretization is:
[0146] (27)
[0147] Substituting , we get the iterative formula for actual calculation:
[0148] (28)
[0149] This formula is executed at each time step t, so as to complete the transition from to State evolution.
[0150] With reference Figure 6 to the foregoing, the training method of the power consumption prediction model comprises the following steps:
[0151] S1, historical data is obtained and preprocessed and standardized to form a historical sequence X={x1, x2, …, x T} and input into the power consumption prediction model; x1, x2 and x T are actual power consumption values at time steps 1, 2 and T respectively.
[0152] S2, the power consumption prediction model traverses t=w, w+1, …, T-1 to generate a prediction sequence ={x' w+1 , x' w+2 ,…,x' T}, x' w+1 , x' w+2 and x' T are power consumption prediction values at time steps w+1, w+2 and T respectively.
[0153] It is worth noting that the historical sequence X={x1, x2, …, x t ,…,x T} is input at one time during the training process, batch processing can be performed in the semantic perception channel, thereby obtaining hidden representation H={h1, h2, …, h t ,…,h T}, h1, h2, h t and h T are hidden state representations at time steps 1, 2, t and T respectively; similarly, h t in formulas 13, 17 and 18 can be replaced by H to complete batch processing, thereby obtaining:
[0154] ={ };
[0155] At this time, the column total j max in formula 15=T.
[0156] S3, a loss function is calculated on the historical sequence y={x w+1 , x w+2 ,…,x T} and the prediction sequence ={x' w+1 , x' w+2 ,…,x' T}; x w+1 and x w+2are the actual power consumption values at time steps w+1 and w+2, respectively, y e X.
[0157] S4, updating the model parameters by backpropagation through the loss function;
[0158] S5, repeating the above steps S1-S4 until the model converges.
[0159] During training, each batch of samples can contain historical data of multiple users, and the loss function is the sum of the losses of multiple users. In subsequent embodiments, the model input is the hourly electricity consumption data of 321 users, and the output is the load prediction value in a 24-hour sequence.
[0160] To effectively train the dual-channel modulation liquid state neural network system and evaluate its prediction performance, the present application adopts a loss function that combines multivariate regression and robust optimization, namely the multivariate Huber loss:
[0161] The design of this loss function aims to address the following challenges simultaneously:
[0162] Multivariate output: C load values (i.e., the number of prediction time steps is C) of N=321 users need to be predicted and error evaluated simultaneously.
[0163] Data heterogeneity: There are differences in power consumption patterns between different users and different times, and there may be noise and outliers.
[0164] Training stability: a balance needs to be struck between the smooth gradient of mean squared error (MSE) and the outlier robustness of mean absolute error (MAE).
[0165] The point-by-point definition of the Huber loss function is the basis for constructing the total loss. For the true value y and the predicted value of a single data point (a single time step of a single user and a single variable), the Huber loss is defined as:
[0166] (29)
[0167] where is a hyperparameter, called the threshold (in this embodiment, it is taken as ), which determines the boundary between the quadratic loss (MSE) and the linear loss (MAE). When the prediction error is small (≤ ), the loss function behaves as MSE, which has a continuous and non-zero gradient near zero, which is conducive to fine updating of model parameters and rapid convergence; when the prediction error is large (> 1.5 ), the loss function behaves as MAE, whose gradient amplitude is constant (± ), which avoids the gradient explosion problem of MSE when facing outliers, and enhances the robustness of the model.
[0168] Extending to the multivariate multi-user scenario, the total loss function needs to aggregate the loss over all users (N), all predicted variables (prediction results over C prediction time steps). Its calculation follows the following steps:
[0169] 1. Calculate the average loss of each variable-user pair: for the cth predicted variable (e.g., the load at the cth hour), calculate its average Huber loss over all N users:
[0170] (30)
[0171] where and represent the true value and predicted value of the cth variable on the ith user, respectively.
[0172] 2. Cross-variable average: average the average loss of all C predicted variables again to ensure that each prediction target has equal importance in the total loss:
[0173] (31)
[0174] represents the average data fitting error of the model on all prediction tasks, is the prediction loss on the cth variable.
[0175] 3. Introduce regularization term: To prevent overfitting and improve the generalization ability of the model, an L2 regularization term (also known as weight decay) is added to the loss function. The penalty term is the square sum (L2 norm) of all learnable parameters of the model:
[0176] (32)
[0177] where represents the learning parameter; is the regularization coefficient, used to control the regularization strength.
[0178] 4. Total loss function: The final total loss is the weighted sum of the data fitting loss and the regularization loss :
[0179] (33)
[0180] Refer to Figure 6The application provides a power consumption dynamic prediction method based on physical-semantic dual-channel modulation.
[0181] St1, acquiring a current sequence y={x t-w+1 ,x t-w+2 ,…,x t} and inputting the power consumption prediction model to obtain a predicted value x t+1 ; x t is an initial value of the latest collected value at a current time step t0; when the initial time t=t0, h t-1 is a 0 vector;
[0182] St2, judging whether a difference between the time step t+1 and the time t0 is greater than or equal to a set value C;
[0183] No, updating t to t+1, and then returning to step St1;
[0184] Yes, outputting a predicted sequence { x t0+1 , x t0+2 , …, x t0+C}.
[0185] The above power consumption dynamic prediction method based on physical-semantic dual-channel modulation (hereinafter referred to as the application method) is verified in combination with specific embodiments.
[0186] This embodiment takes the prediction of 321 user future 24-hour power load as an example. The overall network structure is shown in Figure 1 .
[0187] In this embodiment, the hourly electricity consumption data of 321 users from July 1, 2016 to 2019 is used to construct a multivariate time series, each user is regarded as a node, and the node feature is the normalized electricity consumption; the model input is the historical 120-hour electricity consumption data, and the output is the future 24-hour load prediction value. In terms of network structure, the DB4 wavelet basis is selected for 5-layer wavelet decomposition in the physical perception channel, the energy feature is calculated, and the multilayer perception machine is used to generate the physical modulation signal Mp(t), the semantic perception channel adopts the 4-head self-attention mechanism to encode the context information and generate the semantic modulation signal Ms(t), the dual-channel output is dynamically balanced by the gate fusion mechanism, and finally the liquid core calculation unit is modulated. The liquid core calculation unit adopts a 32-dimensional hidden state, and the dynamics of the ordinary differential equation is solved by the Euler method. In the training setting, the AdamW optimizer (learning rate 1e-3) and the Huber loss function (δ=1.0) are used, and the early stopping strategy (the training is terminated when the validation set loss does not decrease for 10 consecutive rounds) is adopted to ensure the stability and generalization performance of the model training.
[0188] Figure 1 The middle physical perception pathway is a "reflex arc" system of the method of the present application, responsible for low-delay, instinctive physical feature extraction and modulation of input signals. The pathway converts the physical essence of the signal into modulatable parameters of neural network dynamics through wavelet transform.
[0189] To verify the effectiveness of the power consumption prediction model (hereinafter referred to as the model of the present application) proposed by the present application, the model of the present application (hereinafter referred to as LiqDuet-Net) is compared with the following representative baseline models on the multi-user power load data set provided in the present embodiment:
[0190] Long Short-Term Memory Network (LSTM): Captures long-term dependencies through gating mechanisms, alleviating the gradient vanishing problem;
[0191] Attention Recurrent Neural Network (Attention-RNN): RNN combined with attention mechanism, adaptively focusing on important historical information;
[0192] Gated Recurrent Unit (GRU): A lightweight variant of LSTM, simplifying the structure while maintaining the ability to model long-term dependencies;
[0193] Neural Ordinary Differential Equation (Neural ODE): Based on ordinary differential equations to construct continuous-time dynamics, providing smoother sequence evolution modeling;
[0194] Temporal Convolutional Network (TCN): A modern sequence model based on dilated causal convolution, with stable gradient flow and parallel computing advantages;
[0195] Liquid Neural Network (LNN): A continuous-time dynamics model based on neural ordinary differential equations, realizing adaptive evolution of time series information through liquid state;
[0196] Attention Enhanced Liquid Neural Network (Attention-LNN): Introduces attention mechanism in the basic liquid neural network, enhancing the perception ability of key time steps;
[0197] Wavelet Enhanced Liquid Neural Network (Wavelet-LNN): Integrates wavelet transform module in the front end of liquid neural network, extracts time-frequency physical features of signal;
[0198] Multi-Scale Wavelet Liquid Neural Network (Multi-Wavelet-LNN): Uses multi-scale wavelet decomposition strategy to more comprehensively capture the physical properties of signals in different frequency bands.
[0199] All comparative models were trained with the same training set, validation set and test set, and were hyper-parameter tuned to ensure their performance was optimal. The evaluation metrics of model prediction performance were root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE) and coefficient of determination (R²), and the specific performance of each model on the test set is shown in Table 1.
[0200] Table 1: Comparison of prediction performance of different models on the test set
[0201] ;
[0202] As can be seen from Table 1, the LiqDuet-Net model of the present application is significantly better than all baseline models in all four quantitative evaluation indicators of RMSE, MAE, MAPE and R². This fully verifies the effectiveness of the method of the present application.
[0203] To comprehensively evaluate the prediction performance of the LiqDuet-Net model proposed in the present application, the present embodiment performs systematic display and quantitative analysis from multiple dimensions. Figures 2(a), 2(b), 2(c) and 2(d) clearly reveal that the LiqDuet-Net model of the present application is significantly better than a series of baseline models in RMSE, MAE, MAPE and R² through multi-index comparison, verifying the advancement of its physical-semantic dual-channel architecture. Figure 3 Further, the high consistency between the predicted values and the true values of the model is intuitively presented on typical time series segments, especially in the accurate tracking of load mutations and periodic patterns. To quantify this consistency, Figure 4 The predicted value-true value scatter plot shows that the sample points are closely distributed on both sides of the diagonal line, proving that the prediction results are overall unbiased and accurate. Finally, the prediction residual distribution plots of Figures 5(a), 5(b) and 5(c) show that the model errors are concentrated around zero and approximately normally distributed, statistically confirming the robustness and stability of the prediction performance of LiqDuet-Net. In summary, the figures from macroscopic comparison to microscopic analysis collectively constitute strong evidence for the excellent prediction ability and reliability of the LiqDuet-Net model.
[0204] Of course, for those skilled in the art, the present application is not limited to the details of the above exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be regarded as exemplary and non-limiting from any point of view, and the scope of the present application is defined by the appended claims rather than the above description, and therefore all changes falling within the meaning and scope of the essential elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0205] In addition, it should be understood that, although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the description of the specification is only for the sake of clarity, and those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that those skilled in the art can understand.
[0206] The technologies, shapes, and structural parts not described in detail in the present application are well-known technologies.
Claims
1. A method for dynamic prediction of power consumption based on physical-semantic dual-path modulation, characterized in that, First, a power consumption prediction model is trained based on time-series power load data. Predicting electricity consumption ; The electricity consumption prediction model includes: a physical sensing path, a semantic sensing path, a fusion module, and a liquid core computing architecture; the physical sensing path extracts the input sequence based on wavelet transform. The time-frequency characteristics are used to generate physical modulation signals. The semantic perception pathway is based on the attention mechanism for signals. Processing is performed to generate semantically modulated signals. The physical modulation signal and the semantic modulation signal are fused by the fusion module, and then processed by the liquid core computing architecture to obtain the power consumption prediction result. ; Then, let the electricity consumption prediction model iterate through t={t0,t0+1,……,t0+C}, based on {x t-w+1 ,x t-w+2 ,…,x t Predict x t+1 The electricity consumption prediction sequence {x} is obtained. t0+1 , x t0+2 , …, x t0+C }; where t0 is the current time step, C is the set total number of prediction steps; w is the window width; x t-+1 x t-w+2 x t x t+1 x t0+1 x t0+2 and x t0+C These represent the power consumption at time steps t-w+1, t-w+2, t, t+1, t0+1, t0+2, and t0+C, respectively. The semantic awareness pathway includes a sequentially connected context encoder, a multi-head attention module, and a semantic dynamic parameter generator; Context encoder for signal x t Encode the hidden state representation h. t ; Multi-head attention module: Based on the scaled dot product attention mechanism for the hidden state representation h t The output of the attention head is processed to obtain the fusion of the attention head that evaluates the importance of the context. and will Transform into context-aware representation ; in, and These represent the outputs of the 1st and hth attention heads, respectively, where h is the number of attention heads. To output the projection matrix; Indicates splicing; The semantic dynamic parameter generator will The signal is mapped to a semantic modulation signal through a nonlinear transformation. ; The processing formula for the liquid core computing unit is expressed as follows: in, The power consumption at time step t. This is the predicted power consumption value at time step t+1. Step size; The weighted fusion result is the output of the fusion module, where tanh is the hyperbolic tangent function; It is a non-linear activation function; and These are the system time constant and the bias term, respectively. and These are learnable static parameters; and is the modulation gain hyperparameter; x(t) is the hidden state vector calculated using ordinary differential equations.
2. The power consumption dynamic prediction method based on physical-semantic dual-path modulation as described in claim 1, characterized in that, The physical sensing pathway includes a wavelet transform module, an instantaneous feature extraction module, and a physical dynamic parameter generator connected in sequence; The wavelet transform module performs a discrete wavelet transform on the input sequence to decompose the wavelet coefficients, including approximation coefficients. and multi-scale detail coefficients { ,1≤j≤J}, where J is the total number of wavelet decomposition layers; The instantaneous feature extraction module calculates the frequency band energy based on wavelet coefficients. ,1≤j≤J}、Center of Spectrum And the transient impact index TI(t), and construct the physical feature vector. ; , and These represent the energy of the first, j, and J-th frequency bands, respectively. The approximate coefficient energy of the Jth layer; The physical dynamic parameter generator transforms physical feature vectors through nonlinear transformation. Converted into a physical modulation signal .
3. The power consumption dynamic prediction method based on physical-semantic dual-path modulation as described in claim 2, characterized in that: in, For the detail coefficients of the j-th layer Length, for The k-th value in the middle; The approximation coefficients of the Jth layer The k-th value in the middle; The center frequency weight of the j-th frequency band is... .
4. The method for dynamic prediction of power consumption based on physical-semantic dual-path modulation as described in claim 1, characterized in that, The fusion module first uses a gating network to modulate the physical modulation signal. and semantic modulation signal The concatenated vector is processed, then activated to obtain a gating signal, which is then used as a weight to modulate the physical modulation signal. and semantic modulation signal Perform weighted fusion.
5. The method for dynamic prediction of power consumption based on physical-semantic dual-path modulation as described in claim 1, characterized in that, The electricity consumption prediction model is trained on the dataset {X}, where X = {x1, x2, ..., x}. T Let X be the historical sequence, x1, x2, and x... T These are the power consumption samples at time steps 1, 2, and T; Electricity consumption model based on sequence The predicted power consumption at time step t+1 is denoted as The electricity consumption model iterates through t=w, w+1...T-1 to generate a prediction sequence. ={x' w+1 ,x' w+2 ,…,x' T }, x' w+1 、x' w+2 and x' T These are the predicted power consumption values at time steps w+1, w+2, and T, respectively. In the historical sequence y={x w+1 ,x w+2 ,…,x T } and predicted sequence ={x' w+1 ,x' w+2 ,…,x' T The loss function is calculated on the x-axis, and the model parameters are updated through backpropagation using the loss function until the model converges; w+1 and x w+2 These are the actual power consumption values at time steps w+1 and w+2, respectively.
6. The method for dynamic prediction of power consumption based on physical-semantic dual-path modulation as described in claim 5, characterized in that, loss function for: in, This represents the average loss across variables. Represents the regularization term; C represents the total number of predictor variables, and N represents the total number of users; and Let these represent the actual value and the predicted value of the c-th variable for the i-th user, respectively. for and Huber's losses; This represents all learning parameters; Represents the learning parameters, It is the regularization coefficient.
7. A dynamic power consumption prediction system based on physical-semantic dual-path modulation, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, the processor is connected to the memory, and the processor is used to execute the computer program to implement the power consumption dynamic prediction method based on physical-semantic dual-path modulation as described in any one of claims 1-6.
8. A storage medium, characterized in that, The device contains a computer program that, when executed, is used to implement the power consumption dynamic prediction method based on physical-semantic dual-path modulation as described in any one of claims 1-6.
Citation Information
Patent Citations
Power consumption time sequence prediction method and system based on self-attention mechanism, medium and processor
CN120911734A
High-voltage SVG adaptive voltage control method based on deep learning
CN120934058A