Error prediction method of electric energy metering device

Through the method of adaptive decomposition and feature modeling, the problems of non-stationarity and insufficient environmental adaptability in the error prediction of electric energy metering devices are solved, and high-precision and explainable error prediction is achieved.

CN120654187APending Publication Date: 2025-09-16实德电气集团有限公司

Patent Information

Application Number
CN202510760735.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The error prediction method of electric energy metering devices in the existing technology has deficiencies in handling non-stationarity, feature fusion depth and adaptability to complex environments, resulting in insufficient prediction accuracy and interpretability.

Method used

Adaptive decomposition technology is used to decompose the error sequence into three components: trend, cycle and random. Feature modeling is performed through time convolutional network, frequency attention long short-term memory network and probabilistic variational autoencoder. Prediction is performed by combining physical constraints and elastic weight integration algorithm, and the model is dynamically updated to adapt to environmental changes.

Benefits of technology

It improves the accuracy and interpretability of error prediction, enhances the adaptability of the model in complex environments, and reduces prediction error fluctuations and key feature forgetting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654187A_ABST
    Figure CN120654187A_ABST
Patent Text Reader

Abstract

The invention discloses an error prediction method of an electric energy metering device, which solves the problems of insufficient non-stationarity processing and poor environmental adaptability in the traditional method by adaptively decomposing an error sequence and applying a physical constraint fusion prediction component and combining an elastic model updating mechanism. The method has the advantages of improving error prediction precision and enhancing model environment adaptability. Spectral residual transform and adaptive wavelet threshold decomposition are adopted for cooperative work, the defects of noise sensitivity and trend term extraction deviation of a traditional method are effectively overcome, the detection capacity of transient characteristics is enhanced through spectral residual transform, and the distortion problem of principal component analysis under strong interference is avoided; the adaptive wavelet threshold algorithm dynamically optimizes the decomposition granularity, and precisely separates the components of a trend term, a periodic term and a random term.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of metering device data processing, and in particular to an error prediction method for an electric energy metering device. Background Art

[0002] The accuracy of electric energy metering devices is the cornerstone of fair trading and settlement in the electricity market and the safe operation of the power grid. Traditional management relies on static verification and periodic rotation, treating errors as fixed values. However, in actual operation, errors are affected by factors such as ambient temperature and humidity, secondary loads, and equipment aging, exhibiting significant non-stationary dynamic characteristics. Static methods cannot capture real-time changes, and manual periodic rotations have lags, which can lead to error accumulation and even metering inaccuracies, resulting in huge electricity bill disputes or system malfunctions. With the development of smart grids, high-precision dynamic error prediction technology is urgently needed to enable online monitoring and early warning.

[0003] Although existing dynamic prediction methods have made some progress, they still have obvious limitations. Chinese invention patent CN116663728A proposes a prediction scheme that integrates differential processing, principal component analysis (PCA), and multi-model weighting: first, the non-stationary sequence is converted into a stationary sequence through differential processing, and then PCA is used to extract the residual sequence of historical errors to highlight nonlinear characteristics; finally, rigid prediction (fixed time window) and flexible prediction (adaptive time window) are used to process the stationary sequence respectively, and the nonlinear prediction results of the residual sequence are combined for weighted fusion. Although this method takes into account both linear and nonlinear characteristics, PCA is sensitive to noise, and the residual extraction is easily distorted under strong interference. In addition, the model robustness is not optimized for complex noise environments.

[0004] Another Chinese invention patent, CN109061544A, focuses on noise suppression. It decomposes the error sequence into low-frequency components (trend terms) and high-frequency components (noise and details) through wavelet transforms, and constructs a robust extreme learning machine (RELM) model. This model inputs the decomposed components along with real-time metering data, utilizing 1-norm constraints on training errors and an augmented Lagrangian algorithm to improve noise immunity. However, this method does not pre-process the non-stationarity of the original sequence, and direct decomposition may lead to deviations in the trend term. Furthermore, the independent prediction of each component ignores the correlation between multi-scale features, limiting the potential for accuracy improvement.

[0005] The common bottlenecks of the two are: insufficient handling of non-stationarity, insufficient feature fusion depth and weak adaptability to complex environments. Summary of the Invention

[0006] (1) Technical issues to be resolved

[0007] To solve the above problems, the present invention proposes an error prediction method for an electric energy metering device, aiming to solve the problems in the prior art of insufficient non-stationarity processing, insufficient feature fusion depth and weak adaptability to complex environments.

[0008] (2) Technical solution

[0009] An error prediction method for an electric energy metering device according to the present invention comprises:

[0010] Acquire an original error sequence and environmental factor data of an electric energy metering device, and adaptively decompose the original error sequence to obtain a trend term component, a period term component, and a random term component;

[0011] Based on the environmental factor data, the trend term component, the periodic term component, and the random term component are respectively input into an error prediction model to obtain corresponding prediction components, and the prediction components are subjected to physical constraint fusion to obtain an error state prediction value, wherein the physical constraints include a point-to-point constraint of equipment aging and an energy conservation constraint;

[0012] When data distribution drift is detected, the prediction model is updated based on an elastic weight integration algorithm.

[0013] In the present invention, the adaptive decomposition of the original error sequence includes:

[0014] extracting transient features of the original error sequence through spectral residual transformation;

[0015] Adaptive wavelet threshold algorithm is used to decompose the transient feature into the trend term component, the period term component and the random term component.

[0016] In the present invention, the error prediction model includes:

[0017] Temporal convolutional networks for processing trend term components;

[0018] Frequency-attention long short-term memory network for processing periodic term components;

[0019] Probabilistic variational autoencoder for handling random term components.

[0020] In the present invention, the physical constraint fusion process includes:

[0021] Dynamically assign fusion weights to each prediction component through a gating network;

[0022] Applying monotonicity constraints ensures that the trend item prediction component satisfies the time monotonically increasing characteristic;

[0023] Energy conservation constraints are applied to ensure that the total predicted value does not exceed the upper limit of the device's physical range.

[0024] In the present invention, the detection of data distribution drift includes:

[0025] Calculate the KL divergence of the real-time error feature and the historical feature distribution;

[0026] When the KL divergence exceeds a preset threshold, a model update mechanism is triggered.

[0027] In the present invention, the updating of the prediction model based on the elastic weight integration algorithm includes:

[0028] Construct the Fisher information matrix including the importance of historical parameters;

[0029] A regularization term based on Fisher information is added to the model update loss function to protect key knowledge from being covered.

[0030] The present invention also includes:

[0031] When the KL divergence exceeds the secondary threshold, the domain adaptation expert module is activated to reconstruct the feature space;

[0032] The new and old feature distributions are aligned via the maximum mean difference loss function.

[0033] In the present invention, the environmental factor data includes:

[0034] Three-dimensional time series data of temperature, humidity and load current;

[0035] The dynamic feature interaction between environmental factors and error components is achieved through the cross-modal attention mechanism.

[0036] Another computing device of the present invention comprises:

[0037] at least one processor; and

[0038] a memory storing executable instructions;

[0039] When the instruction is executed by at least one processor, the method described in any one of the above technical solutions is implemented.

[0040] Another non-transitory machine-readable storage medium of the present invention stores executable instructions, which, when executed, enable a machine to perform the method as described in any one of the above technical solutions.

[0041] (3) Beneficial effects

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] (1) The present invention adopts spectral residual transform and adaptive wavelet threshold decomposition to work together, effectively overcoming the defects of traditional methods such as sensitivity to noise and deviation in trend term extraction. Spectral residual transform enhances the detection capability of transient features and avoids the distortion problem of principal component analysis under strong interference. The adaptive wavelet threshold algorithm dynamically optimizes the decomposition granularity and accurately separates the trend term, periodic term and random term components.

[0044] (2) In the present invention, the temporal convolutional network processes the long-term dependency of the trend term, the frequency attention long short-term memory network captures the spectral characteristics of the periodic term, and the probabilistic variational autoencoder models the uncertainty of the random term; further, the weights of each prediction component are dynamically allocated through the gating network, and the monotonicity constraint of equipment aging and the energy conservation constraint are introduced to ensure that the predicted value of the forced trend term increases monotonically over time and the total predicted value does not exceed the upper limit of the physical range.

[0045] (3) The present invention constructs a continuous learning framework based on the elastic weight integration algorithm. Drift is detected by calculating the KL divergence of the real-time error characteristics and the historical distribution. When the threshold is exceeded, the Fisher information matrix is ​​used to protect the historical key parameters. A regularization term is added to the model update loss function, which greatly reduces the key knowledge forgetting rate. When the KL divergence exceeds the secondary threshold, the domain adaptive expert module is activated to reconstruct the feature space. The maximum mean difference loss is used to align the new and old data distributions, reducing the fluctuation range of the prediction error in strong drift scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1 Schematic diagram of the overall structure of the error prediction method for electric energy metering device;

[0048] Figure 2 Schematic diagram of the process structure of adaptive decomposition;

[0049] Figure 3 This is a schematic diagram of the comparison curve between the prediction results of this method and the traditional method;

[0050] Figure 4 Schematic diagram of the curve for data drift detection and model update response;

[0051] Figure 5 Schematic diagram of the comparison curve of long-term prediction performance;

[0052] Figure 6Schematic diagram of the framework structure of the execution equipment.

[0053] 1. Processor, 2. Memory, 3. Communication interface, 4. Communication bus. DETAILED DESCRIPTION

[0054] In existing technologies, dynamic error prediction for electric energy metering devices faces technical bottlenecks such as inadequate handling of non-stationary characteristics, insufficient depth of multi-component feature fusion, and weak adaptability to complex environments. Traditional methods extract features through differential processing or wavelet decomposition, but differential processing is susceptible to noise interference, resulting in residual distortion, and wavelet decomposition does not consider the non-stationarity of the original sequence, which can easily lead to trend term deviations. Existing prediction models often use a fixed-weight fusion approach and fail to incorporate constraints based on the physical characteristics of the equipment, resulting in prediction results that deviate from actual operating patterns. When environmental factors mutate, the model parameter update mechanism lacks knowledge protection capabilities, resulting in the forgetting of key features.

[0055] To address these issues, the inventors discovered that existing methods for processing nonstationary signals suffer from a disconnect between eigenvalue decomposition and physical laws, and also lack interpretability constraints during the model fusion phase. Analyzing operational data from electricity metering devices revealed that error evolution exhibits time dependence, environmental sensitivity, and the cumulative effects of equipment aging. Based on this, they proposed decomposing the error sequence into three independent components: trend, cycle, and randomness, modeling them independently. Physical constraints were then introduced to guide predictive fusion. To address data drift, a flexible parameter update mechanism was employed to balance new and existing knowledge.

[0056] Example 1

[0057] like Figure 1-Figure 5 The error prediction method of an electric energy metering device shown includes the following specific steps:

[0058] S100 , obtaining an original error sequence and environmental factor data of an electric energy metering device, and adaptively decomposing the original error sequence to obtain a trend term component, a period term component, and a random term component.

[0059] S200. Based on the environmental factor data, the trend item component, the period item component and the random item component are input into the error prediction model respectively to obtain the corresponding prediction components, and the prediction components are subjected to physical constraint fusion to obtain the error state prediction value.

[0060] S300: When data distribution drift is detected, the prediction model is updated based on the elastic weight integration algorithm.

[0061] Adaptive decomposition involves using signal processing techniques to decompose nonstationary error sequences into components with distinct physical meanings. This can be achieved using a spectral residual transform combined with an adaptive wavelet threshold algorithm to eliminate noise interference while preserving mutation characteristics. The error prediction model incorporates specialized network structures tailored to the characteristics of different components. For example, a time convolutional network is used to handle the temporal dependence of trend terms, a frequency attention network captures the fluctuation patterns of periodic terms, and a probabilistic variational autoencoder models the uncertainty of random terms. Physical constraint fusion involves modifying prediction results based on the physical characteristics of the device. This is achieved by dynamically adjusting component weights through a gating network, imposing monotonicity constraints to ensure the increasing nature of aging trends, and applying energy conservation constraints to limit the physical range of the predicted values. The elastic weight integration algorithm calculates a parameter importance matrix to protect key features from being overwritten during model updates.

[0062] Specifically, after the original error sequence is transformed through the spectral residual to extract transient features, it is decomposed into three components: trend, periodic, and random. Each component is input into a dedicated prediction model. The trend term model learns the aging patterns of equipment, the periodic term model correlates with changes in environmental factors, and the random term model captures transient disturbances. During the fusion phase, the gating network dynamically adjusts the weights of each component based on environmental factors, while also imposing the monotonically increasing error constraint caused by equipment aging and the principle of energy conservation to limit the maximum predicted value. When changes in the distribution of environmental factors such as temperature and humidity are detected, the elastic weight algorithm selectively updates the model parameters to preserve important features from the historical data.

[0063] Compared with existing technologies, existing solutions are sensitive to noise when extracting residual features using principal component analysis. This solution, however, enhances the ability to extract mutation features through spectral residual transformation. Traditional wavelet decomposition fails to account for the non-stationarity of the original sequence. This solution extracts transient features before decomposition, improving the physical interpretability of the decomposition results. Existing fusion methods use fixed weighting, while this solution uses physical constraints to ensure that the predicted values ​​conform to the actual operating principles of the device. When the environment undergoes sudden changes, traditional model updates overwrite historical knowledge. This solution, however, implements a flexible weighting mechanism to preserve this knowledge.

[0064] Through the above technical solution, this application achieves physical characteristic decomposition of non-stationary error sequences. A dedicated model captures the evolution of different components, and the aging characteristics of equipment and the principle of energy conservation guide prediction fusion, improving the interpretability and accuracy of prediction results. As environmental factors change, the elastic parameter update mechanism effectively balances new and old knowledge, enhancing the model's adaptability under complex working conditions.

[0065] The present application further proposes to adaptively decompose the original error sequence, including extracting the transient features of the original error sequence through spectral residual transformation, and decomposing the transient features into trend term components, periodic term components and random term components using an adaptive wavelet threshold algorithm.

[0066] Among them, spectral residual transform refers to the calculation of phase spectrum residual based on Fourier transform. Specifically, the fast Fourier transform algorithm can be used to convert the original error sequence into the frequency domain, calculate the residual component of the phase spectrum, and realize the enhanced extraction of transient features. This transform improves the signal-to-noise ratio of transient features by separating the mutation component and background noise in the signal. The adaptive wavelet threshold algorithm refers to a wavelet denoising method that dynamically adjusts the threshold parameters according to the number of decomposition layers. Specifically, a layered soft threshold function can be used to set differentiated threshold rules at different decomposition scales. This algorithm eliminates the excessive suppression of high-frequency detail information by the fixed threshold through adaptive adjustment, retaining the true fluctuation characteristics of the signal.

[0067] Specifically, the spectral residual transform first performs a fast Fourier transform on the original error sequence, extracts the phase spectrum component and calculates its residual to generate a transient component containing mutation characteristics. This process effectively eliminates the stationary background noise in the signal and retains the non-stationary fluctuations caused by equipment anomalies or environmental mutations. The transient component is then input into the adaptive wavelet threshold algorithm for multi-scale decomposition. The optimal wavelet basis function is dynamically selected according to the number of decomposition layers, and a soft threshold parameter that matches the noise intensity is used in each decomposition layer. For example, a larger threshold is used in the low-frequency trend term decomposition layer to suppress random noise, a medium threshold is used in the periodic term decomposition layer to retain periodic fluctuations, and a small threshold is used in the random term decomposition layer to capture detailed fluctuations. Through a hierarchical optimized threshold strategy, accurate separation of trend terms, periodic terms, and random terms is achieved.

[0068] This application further proposes an error prediction model including a temporal convolutional network, a frequency attention long short-term memory network and a probabilistic variational autoencoder, which respectively process trend term components, periodic term components and random term components.

[0069] A temporal convolutional network refers to a deep neural network with an expanded causal convolution structure, which can be implemented using multi-layer convolution kernels with exponentially expanded receptive fields to capture slowly changing features across long periods in trend term components; a frequency attention long short-term memory network refers to a frequency domain feature extraction mechanism embedded in a recurrent neural network, which can be implemented through a Fourier basis function projection layer and dynamic calculation of attention weights to identify key frequency patterns in periodic term components; a probabilistic variational autoencoder refers to a generative model that establishes a probability distribution in the latent space, which can be implemented through an encoder-decoder architecture combined with KL divergence constraints to model the statistical distribution characteristics of random term components.

[0070] Specifically, in the processing of trend components, dilated causal convolution gradually expands the time window perception range through hierarchical superposition, effectively capturing the long-term correlation of slowly changing processes such as equipment aging; in the processing of periodic components, Fourier basis functions map time series data to the frequency domain space, and combined with the attention mechanism, automatically screen the dominant frequency components, eliminating the dependence on manually set period lengths; in the processing of random components, variational inference models the noise distribution through latent variables and uses reparameterization technology to generate prediction results that conform to statistical laws, avoiding the oversensitivity of deterministic models to instantaneous noise. Through feature decoupling and collaborative prediction, the three achieve accurate feature extraction of error components in multi-dimensional space.

[0071] The expression of the error prediction model is:

[0072]

[0073] is the trend term forecast value at time t, which represents the long-term change dominated by the aging of the output device. Represents the input sequence at time step The value of , d is the expansion factor, which increases exponentially with the number of network layers, τ represents the index in the convolution kernel and τ=0,1,...T-1, W τ represents the convolution kernel weight matrix, K is the convolution kernel size, which is used to determine the length of the time window, b is the bias term, σ(·) is the activation function, which is the Sigmoid function in this embodiment. It represents the integer part of τ divided by K, which determines the exponent of expansion;

[0074] represents the predicted value of the periodic term at time t, W o is the output weight matrix, which is used to represent the final prediction, h t represents the hidden state at time t, and V i Represents a value matrix used to store historical cycle characteristics, α t,i is the frequency attention weight, used to identify the dominant periodic component, and F(·) represents the fast Fourier transform, which is used to convert the time domain signal into the frequency domain signal. Represents the basic hidden state of LSTM The frequency domain vector obtained by fast Fourier transform, F(H) represents the frequency domain matrix obtained by fast Fourier transform on each column of the historical hidden state matrix H, d k is the scaling factor, H is the historical hidden state matrix, which usually contains the hidden states of multiple time steps, for example, H = [h1,h2,...,h t-1], each column is a hidden state vector, Softmax(·) represents Softmax normalization along the row or column to obtain the attention weight,

[0075] is a one-step forward prediction, which indicates the output of the burst interference prediction value, z t is the latent variable sampling, representing the probability distribution of random items, e t is the environmental feature vector, which represents the injection of real-time data such as temperature, humidity, and load. represents the vector concatenation operation, fusing latent variables and environmental factors, and Decoder(·) represents the conditional decoder, which is used to characterize the deconvolution network.

[0076] Furthermore, the latent variable in

[0077] Encoder(·) is a variational encoder, T t ,H t ,L t They represent the real-time sensor input data of temperature, humidity, and load current respectively, MLP(·) represents the environmental factor encoder, and X residual (t-6:t) represents the sequence of random items in the last 7 steps, representing the sliding window, μ t represents the latent variable mean vector, represents the logarithmic variance vector, and N(·) represents the Gaussian distribution.

[0078] Through the above technical solutions, this application solves the problems of inaccurate error component feature extraction and lack of multi-scale correlation. The temporal convolutional network ensures accurate characterization of the equipment aging process, the frequency attention mechanism improves the dynamic recognition accuracy of periodic features, and the probabilistic variational model suppresses random noise interference, thereby comprehensively improving the robustness and accuracy of error prediction of electric energy metering devices.

[0079] This application further proposes physical constraint fusion processing, including dynamically allocating the fusion weights of each prediction component through a gating network, applying monotonicity constraints to ensure that the trend item prediction component satisfies the time monotonically increasing characteristics, and applying energy conservation constraints to ensure that the total prediction value does not exceed the upper limit of the physical range of the device.

[0080] The gated network refers to a dynamic weight allocation mechanism built based on the gated recurrent unit. Specifically, the input feature gating function can be combined with the historical weight state to achieve weight adjustment, which is used to dynamically adjust the contribution ratio of each prediction component according to changes in environmental factors.

[0081] The monotonicity constraint refers to adding a mathematical constraint term that the time series is monotonically increasing to the loss function of the trend item prediction component. Specifically, it can be implemented using a penalty function based on the non-negativity of the first-order difference, which is used to force the trend item component to meet the irreversible characteristic of error accumulation caused by equipment aging.

[0082] The energy conservation constraint refers to setting a range limit function in the predicted total value output layer. Specifically, it can be implemented using an activation function with saturation characteristics or a projection optimization algorithm to map the fusion result to the maximum allowable error range of the electric energy metering device.

[0083] During the error prediction component fusion phase, the prediction results for the trend, period, and random terms are first input into a gated network. The gating mechanism analyzes the correlation between the eigenvectors of each component and environmental factors, generating dynamic weight coefficients that vary with time. During model training, the trend prediction component is subjected to gradient direction constraints through a monotonicity constraint module to ensure that its predicted value remains non-decreasing in the time dimension, consistent with the physical law that equipment errors increase with age. During the final fusion prediction phase, the weighted summation results are nonlinearly transformed using energy conservation constraints, limiting the predicted value to the range defined by the equipment's technical specifications and eliminating anomalous predictions that exceed physical limits.

[0084] Traditional error prediction methods use a fixed-weight fusion strategy, which is unable to adapt to the varying contributions of each prediction component under varying environmental conditions and fails to account for equipment aging patterns and physical range limitations. The existing approach of predicting residual sequences and processing components independently can easily lead to inaccurate weight allocation and physical law conflicts. This solution, through a dynamic gating mechanism and dual physical constraints, ensures prediction accuracy while enforcing compliance with the actual operating characteristics of the equipment.

[0085] This application further proposes a computational method for detecting data distribution drift, including measuring the distribution difference by calculating the KL divergence of the real-time error feature and the historical feature distribution, and a response strategy for initiating a model update mechanism when the divergence exceeds a preset threshold.

[0086] Calculating the KL divergence of the real-time error characteristics and the historical characteristic distribution refers to using the relative entropy in information theory to measure the degree of difference between the two probability distributions. Specifically, the sliding window method can be used to collect real-time samples, construct the empirical distribution of the multi-dimensional feature vector, and compare the entropy value with the historical benchmark distribution. This calculation method can accurately quantify the potential distribution shift caused by sudden changes in environmental factors or equipment aging. The preset threshold refers to the dynamic trigger boundary set according to the actual operating scenario. It can be determined specifically by statistical simulation. The KL divergence distribution under historical normal operating conditions is simulated by the Monte Carlo method, and the 95% quantile is selected as the initial threshold. This threshold setting mechanism not only avoids false triggering caused by noise disturbances, but also can capture significant distribution changes in a timely manner.

[0087] Specifically, during the continuous operation of the electric energy metering device, the probability distribution of the error characteristics will undergo irreversible changes as the environmental parameters drift. The statistical characteristics of the current error sequence, including indicators such as mean, variance, skewness, and kurtosis, are extracted through a sliding time window to form a real-time feature vector. The kernel density estimation method is used to construct its real-time feature distribution, and the KL divergence is calculated with the pre-stored historical benchmark distribution. When the calculation result continuously exceeds the preset threshold, it indicates that the error data generation mechanism has undergone substantial changes, and the model update process is automatically activated. This detection mechanism quantifies the degree of data evolution through the rate of change of information entropy, overcoming the defect that the traditional mean-variance test is insensitive to high-order moment characteristics.

[0088] Compared to existing technologies, traditional methods typically rely on fixed-period manual verification or simple statistical tests for model maintenance, which can lead to delayed response and excessive maintenance. This solution uses KL divergence to implement non-parametric dynamic distribution monitoring, capable of capturing implicit correlation changes in multidimensional feature spaces. The threshold trigger mechanism introduces statistical confidence control, making update decisions probabilistically interpretable and avoiding the subjective bias inherent in empirical threshold setting in existing technologies.

[0089] This application further proposes a specific implementation method for updating the prediction model based on the elastic weight integration algorithm, including constructing a Fisher information matrix containing the importance of historical parameters, and adding a regularization term based on the matrix to the model update loss function.

[0090] The elastic weight integration algorithm refers to an incremental learning mechanism that achieves knowledge protection by quantifying the importance of model parameters to historical tasks. Specifically, it can be achieved by calculating the Fisher information of the parameters on historical training data. This information reflects the degree of influence of parameter changes on the model prediction results.

[0091] The Fisher information matrix refers to a diagonal matrix that stores the importance of parameters at each layer of the model. It can be constructed by calculating the expected value of the second-order derivative of the loss function with respect to the parameters. This matrix is ​​used to identify weight parameters that have a key impact on historical error characteristics.

[0092] Among them, the regularization term refers to the penalty term that constrains the update of model parameters. It can be implemented in the form of the product of the parameter change and the Fisher information matrix. This design suppresses the update amplitude of important parameters, while non-critical parameters can be freely adjusted to adapt to the new data distribution.

[0093] Specifically, after detecting data distribution drift, the Fisher information of each model parameter is first calculated based on the historical training data to form a matrix structure that reflects the importance of the parameters. The elements with larger values ​​in this matrix correspond to the key weight parameters that affect the identification of historical error patterns. When constructing the model update loss function, in addition to the conventional new data fitting loss, an additional regularization term based on Fisher information is added. This regularization term imposes a stronger update resistance on important parameters by multiplying the parameter update amount with the corresponding Fisher information value. As a result, when the model adapts to the new data distribution, the key weights remain relatively stable, and the secondary weights can be flexibly adjusted to achieve collaborative optimization of new and old knowledge. This dynamic constraint mechanism not only avoids the problem of historical knowledge coverage caused by traditional full parameter updates, but also overcomes the rigidity of the fixed parameter freezing strategy.

[0094] This application further proposes a specific implementation method for updating the prediction model based on the elastic weight integration algorithm, including constructing a Fisher information matrix containing the importance of historical parameters, and adding a regularization term based on the matrix to the model update loss function.

[0095] The elastic weight integration algorithm refers to an incremental learning mechanism that achieves knowledge protection by quantifying the importance of model parameters to historical tasks. Specifically, it can be achieved by calculating the Fisher information of the parameters on historical training data. This information reflects the degree of influence of parameter changes on the model prediction results.

[0096] The Fisher information matrix refers to a diagonal matrix that stores the importance of parameters at each layer of the model. It can be constructed by calculating the expected value of the second-order derivative of the loss function with respect to the parameters. This matrix is ​​used to identify weight parameters that have a key impact on historical error characteristics.

[0097] The regularization term refers to the penalty term that constrains the update of model parameters. It can be implemented in the form of the product of the parameter change and the Fisher information matrix. This design suppresses the update amplitude of important parameters, while non-critical parameters can be freely adjusted to adapt to the new data distribution.

[0098] Specifically, after detecting data distribution drift, the Fisher information of each model parameter is first calculated based on the historical training data to form a matrix structure that reflects the importance of the parameters. The elements with larger values ​​in this matrix correspond to the key weight parameters that affect the identification of historical error patterns. When constructing the model update loss function, in addition to the conventional new data fitting loss, an additional regularization term based on Fisher information is added. This regularization term imposes a stronger update resistance on important parameters by multiplying the parameter update amount with the corresponding Fisher information value. As a result, when the model adapts to the new data distribution, the key weights remain relatively stable, and the secondary weights can be flexibly adjusted to achieve collaborative optimization of new and old knowledge. This dynamic constraint mechanism not only avoids the problem of historical knowledge coverage caused by traditional full parameter updates, but also overcomes the rigidity of the fixed parameter freezing strategy.

[0099] This application further proposes environmental factor data including three-dimensional time series data of temperature, humidity and load current, and realizes the dynamic feature interaction between environmental factors and error components through a cross-modal attention mechanism.

[0100] Three-dimensional time series data refers to continuous monitoring data of three physical quantities, temperature, humidity, and load current, aligned in time. This data can be collected using a timestamped sensor array, covering the core environmental parameters that influence error variations during device operation. This data construction method can fully characterize the multi-dimensional coupling of environmental disturbance sources.

[0101] The cross-modal attention mechanism establishes a dynamic weighted association model between the modalities of environmental factors and error components. This is achieved using a multi-head attention layer combined with a gated network. By calculating the attention scores for each error component across different environmental dimensions, it adaptively captures nonlinear interactions. This mechanism enables the model to automatically focus on key influencing factors under different operating conditions.

[0102] Specifically, three-dimensional time series data of temperature, humidity, and load current are fed into an encoder to extract feature vectors, which are then fed into the cross-modal attention module along with the feature vectors of the error components. An attention weight matrix is ​​generated by calculating the similarity score between the environmental feature vectors and the error feature vectors. This weight matrix dynamically adjusts the influence of environmental information on the error components, for example, increasing the weight of the humidity dimension under high temperature and high humidity conditions or reinforcing the interaction of the current dimension when the load changes suddenly. After multiple rounds of attention iteration, a representation of the error components that incorporates the dynamic characteristics of the environment is ultimately formed.

[0103] Through the above technical solution, this application realizes the dynamic feature interaction between environmental factors and error components, enabling the model to adjust the contribution of different factors to error prediction according to real-time environmental changes, solving the prediction deviation problem caused by insufficient environmental correlation modeling in traditional methods, and improving the error prediction accuracy under complex working conditions.

[0104] Example 2

[0105] An embodiment of the present invention provides a computer-readable storage medium.

[0106] The computer-readable storage medium provided in the embodiment of the present invention stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned error prediction methods for electric energy metering devices can be implemented.

[0107] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes.

[0108] For an introduction to the computer-readable storage medium provided in an embodiment of the present invention, please refer to the above method embodiment, and the present invention will not elaborate on it here.

[0109] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0110] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0111] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0112] Example 3

[0113] An embodiment of the present invention provides an execution device.

[0114] Please refer to Figure 6 , Figure 6 This is a schematic diagram of the structure of an execution device provided by the present invention, which may include:

[0115] memory for storing computer programs;

[0116] The processor is configured to implement the steps of any of the above-mentioned error prediction methods for electric energy metering devices when executing a computer program.

[0117] like Figure 6 FIG2 is a schematic diagram of the structure of the execution device, which may include a processor 1, a memory 2, a communication interface 3, and a communication bus 4. The processor 1, the memory 2, and the communication interface 3 communicate with each other via the communication bus 4.

[0118] In the embodiment of the present invention, the processor 1 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices.

[0119] The processor 1 may call a program stored in the memory 2. Specifically, the processor 1 may execute the operations in the embodiment of the push button switch fault detection method.

[0120] The memory 2 is used to store one or more programs. The programs may include program codes, and the program codes include computer operating instructions. In the embodiment of the present invention, the memory 2 stores at least a program for implementing the following functions:

[0121] Acquire an original error sequence and environmental factor data of an electric energy metering device, and adaptively decompose the original error sequence to obtain a trend term component, a period term component, and a random term component;

[0122] Based on the environmental factor data, the trend term component, the periodic term component, and the random term component are respectively input into an error prediction model to obtain corresponding prediction components, and the prediction components are subjected to physical constraint fusion to obtain an error state prediction value, wherein the physical constraints include a point-to-point constraint of equipment aging and an energy conservation constraint;

[0123] When data distribution drift is detected, the prediction model is updated based on an elastic weight integration algorithm.

[0124] In one possible implementation, the memory 2 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function, etc.; the data storage area may store data created during use.

[0125] In addition, the memory 2 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device.

[0126] The communication interface 3 may be an interface of a communication module, used for connecting to other devices or systems.

[0127] Of course, it needs to be explained that Figure 6 The structure shown does not constitute a limitation on the execution device in the embodiment of the present invention. In actual applications, the execution device may include Figure 6 More or fewer components than shown, or combinations of certain components.

[0128] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the concept and scope of the present invention. Any modifications and improvements made to the technical solution of the present invention by a person of ordinary skill in the art without departing from the design concept of the present invention shall fall within the scope of protection of the present invention. The technical content for which protection is sought in the present invention is fully set forth in the claims.

Claims

1. A method for predicting an error of an electric energy metering device, characterized in that: include: Acquire an original error sequence and environmental factor data of an electric energy metering device, and adaptively decompose the original error sequence to obtain a trend term component, a period term component, and a random term component; Based on the environmental factor data, the trend term component, the periodic term component, and the random term component are respectively input into an error prediction model to obtain corresponding prediction components, and the prediction components are subjected to physical constraint fusion to obtain an error state prediction value, wherein the physical constraints include a point-to-point constraint of equipment aging and an energy conservation constraint; When data distribution drift is detected, the prediction model is updated based on an elastic weight integration algorithm.

2. The method according to claim 1, characterized in that The adaptive decomposition of the original error sequence comprises: extracting transient features of the original error sequence through spectral residual transformation; Adaptive wavelet threshold algorithm is used to decompose the transient feature into the trend term component, the period term component and the random term component.

3. The method according to claim 1 or 2, characterized in that The error prediction model includes: Temporal convolutional networks for processing trend term components; Frequency-attention long short-term memory network for processing periodic term components; Probabilistic variational autoencoder for handling random term components.

4. The method according to claim 3, characterized in that ,The physical constraint fusion processing includes: Dynamically assign fusion weights to each prediction component through a gating network; Applying monotonicity constraints ensures that the trend item prediction component satisfies the time monotonically increasing characteristic; Energy conservation constraints are applied to ensure that the total predicted value does not exceed the upper limit of the device's physical range.

5. The method according to claim 4, characterized in that The detection data distribution drift includes: Calculate the KL divergence of the real-time error feature and the historical feature distribution; When the KL divergence exceeds a preset threshold, a model update mechanism is triggered.

6. The method according to claim 5, characterized in that The updating of the prediction model based on the elastic weight integration algorithm includes: Construct the Fisher information matrix including the importance of historical parameters; A regularization term based on Fisher information is added to the model update loss function to protect key knowledge from being covered.

7. The method according to claim 6, characterized in that Also includes: When the KL divergence exceeds the secondary threshold, the domain adaptation expert module is activated to reconstruct the feature space; The new and old feature distributions are aligned via the maximum mean difference loss function.

8. The method according to claim 7, characterized in that The environmental factor data include: Three-dimensional time series data of temperature, humidity and load current; The dynamic feature interaction between environmental factors and error components is achieved through the cross-modal attention mechanism.

9. A computing device, characterized in that include: at least one processor; as well as a memory storing executable instructions; When the instructions are executed by at least one processor, the method according to any one of claims 1 to 8 is implemented.

10. A non-transitory machine-readable storage medium storing executable instructions, characterized in that: When executed, the instructions cause the machine to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Electric energy measuring error estimation method

    CN109061544A

  • Method and device for predicting error state of electric energy metering device

    CN116663728A

  • Short-term power load prediction method based on long-term and short-term memory neural network

    CN111815065A

  • Electric energy meter metering error prediction method and device and storage medium

    CN114626019A

  • Electronic voltage transformer error prediction method based on Prophet, self-attention mechanism and time sequence convolutional network

    CN115438576A

Cited By

  • Acrylic plate performance evaluation method based on double-spectrum in-situ detection

    CN121207957A