Electric energy metering data reliability verification system and method based on fusion of random forest and deep belief network
By integrating the model of random forests and deep belief networks in the power metering data verification system, the problem that traditional methods are difficult to identify and deal with outliers in the power grid is solved, and high-precision and high-reliability verification of the power metering data is achieved, ensuring the stable operation of the power system.
Patent Information
- Application Number
- CN202510208479.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional power metering data verification methods are difficult to effectively identify and process outliers in the power grid caused by factors such as aging electrical equipment, fluctuations in the power grid status, electromagnetic interference, etc., resulting in the impact of data accuracy and reliability, and cannot meet the strict requirements of modern power systems for high accuracy and high reliability.
The reliability verification system of the electrical energy metering data fusion based on the fusion of random forests and deep belief networks is adopted. By collecting and preprocessing data in real time, key features are extracted, and a fusion model is constructed to identify the errors of the metering device, so as to achieve accurate capture and real-time monitoring of the electrical energy metering data.
Real-time acquisition and monitoring of the errors of metering instruments under the conditions of no power outage and no physical standard instruments, significantly improving the reliability of power metering data, reducing power transaction risks and potential risks of power grid operation, and ensuring the economic security of the power system and fairness and justice of transactions.
Smart Images

Figure CN120123936A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electric energy metering, and particularly to a system and method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network. Background Art
[0002] Against the backdrop of the continuous development of the modern power system and the accelerating intelligentization process, the importance of electric energy metering data has become increasingly prominent, and its reliability has become a core factor for the steady progress of the power industry. With the rapid expansion of the scale of the power grid and the sharp increase in complexity, the electric energy metering data shows a trend of massive growth. Facing such a huge data scale, the traditional verification methods relying on manual inspection and simple statistical analysis are becoming increasingly inadequate.
[0003] Manual inspection is limited by human and time costs and is difficult to conduct a comprehensive and timely screening of large-scale data, resulting in many potential data anomalies not being discovered in time. Simple statistical analysis methods can only touch the surface features of the data and lack the ability to effectively mine and analyze the complex correlation structures, dynamic change laws, and potential risk hazards hidden deep in the data.
[0004] In actual power grid operation, factors such as aging of electrical equipment, fluctuations in grid conditions, and electromagnetic interference frequently cause outliers in electric energy metering data. These outliers seriously affect the accuracy and reliability of the data, may lead to deviations in power trading settlements, interfere with the load forecasting and dispatching decisions of the power grid, and pose a serious threat to the safe and stable operation of the power system. However, the traditional verification methods have obvious deficiencies in identifying and processing these outliers and cannot meet the strict requirements of modern power systems for high-precision and high-reliability electric energy metering data. There is an urgent need for new technical means to solve this problem.
[0005] Although certain progress has been made in automated data processing technologies in the power field, when dealing with electric energy metering data, most existing technical solutions fail to fully consider the spatio-temporal coupling relationships among electrical topologies, transducer characteristics, and the random processes of electrical parameters. This makes it difficult for them to accurately characterize the errors of metering devices under complex power conditions, and it is even impossible to obtain and effectively evaluate the errors in real time without power interruption and without physical standard devices, greatly limiting their role in ensuring the reliability of electric energy metering data.
[0006] Therefore, this application proposes a system and method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network. Summary of the Invention
[0007] The objective of the present invention is to fully exploit the advantages of advanced machine learning algorithms by deeply integrating the random forest and the deep belief network, and to deeply explore the inherent spatio-temporal correlations among the electrical topology, the characteristics of transformers, and the stochastic processes of electrical parameters. A highly robust intelligent model and an efficient computational intelligence solution mechanism are constructed to achieve the accurate capture and real-time monitoring of the errors of measuring instruments, successfully overcome the problem of real-time acquisition of errors under the conditions of power-on and without physical standard instruments, and propose a power metering data reliability verification system and method based on the fusion of the random forest and the deep belief network.
[0008] The technical solution of the present invention: A power metering data reliability verification system and method based on the fusion of the random forest and the deep belief network includes the following steps:
[0009] Omnidirectional power metering data and environmental data are collected in real time, and the 3σ rule screening, wavelet transform, and adaptive filtering techniques are used to remove abnormal data, and then time synchronization processing and normalization processing are performed to obtain preprocessed data;
[0010] Key features are extracted from the preprocessed data, and the relationships between the features are analyzed by the Pearson correlation coefficient, the mutual information method, and the grey relational analysis;
[0011] A fusion model of the random forest and the deep belief network is constructed and trained. Its input layer receives multi-source heterogeneous data that have been preprocessed and feature-extracted, and after calculation, the error value of the measuring instrument is output at the output layer, and monitoring and early warning are performed according to a preset threshold.
[0012] Optionally, omnidirectional power metering data and environmental data are collected in real time through a variety of sensor devices, and the sensor devices include smart meters, voltage transformers, current transformers, power factor sensors, and environmental sensor matrices.
[0013] Optionally, the sensor device collects a voltage data sequence V = [v 1 , v 2 , …, v n and a current data sequence I = [i 1 , i 2 , …, i n at a certain time period, and the method of calculating the voltage data and the current data using the 3σ rule is the same;
[0014] Calculate the mean value and the standard deviation of the voltage data. If there is v j that satisfies it is determined as a suspected outlier and removed.
[0015] Optionally, when processing the voltage data and the current data, wavelet basis functions are used for decomposition to obtain wavelet coefficients at different scales where \(j\) represents the scale and \(k\) represents the position;
[0016] By setting a threshold \(T\), the high-frequency coefficients are processed using the soft thresholding method where \(\sigma\) is the estimated value of the noise standard deviation and \(n\) is the data length;
[0017] That is, when then Then the denoised voltage data or current data is reconstructed.
[0018] Optionally, during the time synchronization process, let the timestamp of the data collected by the smart meter be \(t\) m , and the timestamp of the data collected by the environmental sensor be \(t\) e . Through the time calibration algorithm based on the Network Time Protocol, the time of all data is unified to the same reference time \(t\) 0 , and at the same time, the data is uniformly converted according to the standardized data format.
[0019] Optionally, the normalization process uses the Min-Max normalization method to process the voltage, current, power, and environmental factor-related data respectively;
[0020] Let the minimum value of the voltage data be \(v\) min , and the maximum value be \(v\) max , then the normalized voltage data is expressed as:
[0021]
[0022] Map it to the interval \([0, 1]\).
[0023] Optionally, the extraction of key features includes power metering data, transformer characteristic data, and electrical topology data, where
[0024] For the power metering data, let a voltage data \(V = [v(t\) 1 ), \(v(t\) 2 ),..., \(v(t\) m )] with a duration of \(T\), and \(m\) be the number of sampling points in this time period. Calculate its amplitude feature as: \(A\) v = \(t\) i represents the \(i\)-th sampling moment within the duration \(T\);
[0025] For the transformer characteristic data, it includes:
[0026] For the voltage transformer characteristic, let the nominal transformation ratio of the voltage transformer be \(K\) n , and the transformation ratio obtained through actual measurement be \(K\) m , then the transformation ratio accuracy feature is expressed as:
[0027] Phase angle error feature, the phase angle error value θ is measured by a phase angle measuring device e ;
[0028] Load characteristic feature, by measuring the output error of the mutual inductor under different load conditions, constructing a load characteristic curve E(I):
[0029] E(I) = k load I + b load
[0030] where I is the load current, which is one of the key factors affecting the output error of the mutual inductor. k load is the slope of the curve, and its physical meaning is to reflect the sensitivity of the output error of the mutual inductor to the change of the load current. b load is the intercept of the curve, which reflects the inherent error or initial deviation of the mutual inductor when the load current is zero;
[0031] For electrical topology data, the adjacency matrix A in graph theory is used to represent the electrical topology structure, A ij represents the connection relationship between node i and node j. If there is a connection between node i and node j, then A ij = 1, otherwise A ij = 0; The connection relationship characteristics of the nodes are described by calculating the degree of the nodes where N is the total number of nodes; The branch impedance characteristic is expressed according to Ohm's law as: where V ij is the voltage drop across both ends of branch ij, I ij is the branch current, calculate the impedance of each branch, and statistically calculate the mean value and standard deviation σ Z , for the case of n branches are respectively expressed as:
[0032]
[0033] Optionally, the calculation of the Pearson correlation coefficient to analyze the relationship between features includes:
[0034] Let the voltage amplitude feature vector be A v , and the power factor feature vector be PF, then the Pearson correlation coefficient between them is expressed as:
[0035]
[0036] where PF i is the power factor of the i-th sample, is the value obtained by averaging all PF i , which generally describes the average level of the power factor in the studied system or time period. is the voltage amplitude obtained from each sampling or measurement, which reflects the specific magnitude of the voltage at different times or positions. is the average value of the voltage amplitude.
[0037] Calculate the correlation between the fluctuations of electrical parameters and the characteristic data of the instrument transformer through the mutual information method. Let the characteristic vector of the electrical parameter fluctuations be X, and the characteristic vector of the voltage transformer be Y. Then the mutual information between them is expressed as:
[0038]
[0039] where P(x,y) is the joint probability distribution of X and Y, and P(x) and P(y) are the marginal probability distributions of X and Y respectively;
[0040] Let the change sequence of the electrical topology data be T, and the fluctuation sequence of the measurement accuracy of the instrument transformer be E. Initialize the data, and then calculate the correlation coefficient as:
[0041]
[0042] where ρ is the resolution coefficient, which is used to adjust the calculation of the correlation coefficient and controls the influence of the differences between different sample data on the calculation of the correlation degree to a certain extent. Generally, it is taken as 0.5; T i (k) represents the specific value of the i-th sample in the change sequence T of the electrical topology structure data at the k-th moment. In a power grid containing multiple substations, T i (k) represents the change value of a certain quantization index of the electrical connection mode of the i-th substation at the k-th time point. E j (k) represents the measurement accuracy fluctuation value of the j-th sample in the fluctuation sequence E of the measurement accuracy of the instrument transformer at the k-th moment. For instrument transformers of different types or positions, E j (k) records the deviation value relative to the standard measurement accuracy at a specific time or working condition;
[0043] The calculation of the grey correlation degree is expressed as:
[0044]
[0045] Optionally, in the fusion model, for the random forest part, let the training set be D = {(x 1 ,y 1 ),(x 2 ,y 2 ),…,(x n ,y n )}, where x i is the input sample, and y i is the corresponding error value of the measuring instrument. Generate m subsets D 1 ,D2 , …, D m , each subset has a size of n, and for each subset D j , construct a decision tree T j , at each split of the tree, randomly select k features from all features, where k < total number of features, and find the optimal split point among the k features based on the Gini index or mean squared error.
[0046] Optionally, in the fusion model, the deep belief network part includes:
[0047] Train the first layer of restricted Boltzmann machine. Let the input data x be an n - dimensional vector, and the hidden layer h 1 has m neurons. The energy function of the restricted Boltzmann machine is:
[0048]
[0049] where a i , b j are the biases of the visible layer and the hidden layer respectively, w ij is the connection weight, v i is the input node of the visible layer in the model structure of the deep belief network, representing the value of the i - th dimension of the input data x, and h j represents the output value of the j - th neuron of the hidden layer h 1 ; learn the parameters of the hidden layer h by maximizing the likelihood function 1 , where θ = {a, b, ω} are the parameters of the restricted Boltzmann machine;
[0050] Use the contrastive divergence algorithm for parameter update, which is used for the parameter update of the restricted Boltzmann machine. The weight update formula is:
[0051] Δw ij = ∈(<v i h j > data - <v i h j > recon )
[0052] where ∈ is the learning rate, <·> data represents the expectation over the data distribution, and <·> recon represents the expectation over the reconstruction distribution;
[0053] Take the hidden layer h 1 as the input of the next layer of restricted Boltzmann machine, and repeat the training process until all layers are trained;
[0054] Adjust the parameters of the entire network through the backpropagation algorithm. During the backpropagation process, calculate the gradients of the loss function with respect to the weights and biases of the fusion model. Let the cross-entropy loss function be:
[0055]
[0056] where y i is the true label, and
[0057] is the predicted output of the model; Let the mean squared error loss function be CE The total loss function L = αL MSE +(1 - α)L
[0058] where α is the weight coefficient; Update the parameters using the gradient descent method to globally adjust the network parameters to minimize the total loss function. The weight update formula is t where η is the learning rate, w
[0059]
[0060] Optionally, the multi-source heterogeneous data includes instrument transformer characteristic parameters, electrical topology feature vectors, electrical parameter time series data, and environmental factor vectors;During the training process of the fusion model, combine data augmentation techniques to randomly crop the power metering data, randomly intercept it according to a certain time window or the number of data points in the original data sequence. Let the length of the original voltage data sequence be N, and the random cropping length be M (M < N). Then randomly select the starting position s (0 ≤ s ≤ N - M) from the original sequence to intercept the new voltage data sequence:
[0061] V new = [v(s), v(s + 1), …, v(s + M - 1)]
[0062] Perform a rotation operation on the data to change the order of the electrical parameter time series or the arrangement of the feature dimensions. For a two-dimensional feature matrix X, rotate it 90° clockwise to obtain a new matrix X rot ;
[0063] By adding Gaussian noise, let the original data be x, and the added Gaussian noise follows the N(0, σ 2 ) distribution, and the noise data obtained is:
[0064] x noisy = x + σ·randn(size(x))
[0065] Among them, randn is a function for generating random numbers with a standard normal distribution, and σ is the standard deviation of Gaussian noise, which determines the degree of dispersion of the noise data relative to the original data.
[0066] Optionally, after the fusion model is trained, the power metering data that is collected, preprocessed, and feature-extracted in real time is input into the fusion model. Let the input data be x text , and the estimated error value of the metering device output by the model is
[0067] Set the error threshold ∈. If then the data is determined to be reliable data.
[0068] Optionally, when it is detected that there is a reliability risk in the data, the device immediately activates a multi-channel early warning mechanism and sends the information containing the anomaly to the management end through a reminder service.
[0069] Compared with the prior art, the present application includes at least one of the following beneficial technical effects:
[0070] 1. The present invention adopts a fusion architecture of random forest and deep belief network, and shows excellent performance in the discrimination of outliers and reliability verification of power metering data. Compared with traditional methods, it can lock potential abnormal points with higher accuracy and faster speed. In practical applications, the present invention can quickly respond and accurately identify anomalies that may require long-term manual analysis or complex statistical calculations by traditional methods, and immediately start the early warning processing process, significantly reducing the risks of power transactions and hidden dangers in power grid operation, and ensuring the economic safety of the power system and the fairness and justice of transactions.
[0071] 2. By deeply mining the spatio-temporal correlation of electrical topology, transformer characteristics, and the random process of electrical parameters, the potential information value of the data is fully released, greatly improving the model performance and environmental adaptability. In complex industrial power usage environments (such as strong electromagnetic interference and high-load impact conditions in the steel and chemical industries), commercial power usage scenarios (frequent load fluctuations and frequent device switching), and residential power usage situations (dispersed and random power usage demands), it can stably and accurately verify the reliability of power metering data, capture subtle anomalies and potential risks, and provide a solid technical shield for the stable operation of the power system.
[0072] 3. The present application successfully breaks through the technical bottleneck of real-time acquisition and monitoring of the error of metering devices under the conditions of power outage and without physical standard devices. Through real-time online data analysis and model calculation, potential problems of metering devices can be detected in a timely manner, significantly shortening the time interval for fault discovery and handling, reducing power outage time and power losses, improving the power supply reliability and power quality of the power system, enhancing user satisfaction and trust, and promoting the sustainable development of the power industry. Description of the Drawings
[0073] Figure 1It is the system architecture diagram of this embodiment;
[0074] Figure 2 It is the data processing flow chart of this embodiment;
[0075] Figure 3 It is the structural schematic diagram of the fusion model of this embodiment. Detailed implementation manners
[0076] The technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. The components of the embodiments of the present disclosure described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0077] As Figure 1 shown is the system architecture diagram of the power metering data reliability verification system based on the fusion of random forest and deep belief network of the present invention, clearly presenting the main components of the system of the present invention, as well as their data flow and functional relationships. The system mainly consists of a data acquisition module, a preprocessing and feature extraction module, a fusion model calculation module, a data reliability verification and warning module, and a data storage and management module. The data acquisition module is connected to the power system through a sensor network, acquires power metering and environmental data, and transmits it to the preprocessing and feature extraction module. After performing operations such as data cleaning, denoising, normalization, and feature extraction on the data, the processed data is input into the fusion model calculation module. The fusion model calculation module performs calculations based on the fusion model of random forest and deep belief network, and outputs the error estimation value of the metering device to the data reliability verification and warning module. This module makes judgments according to preset thresholds and rules. If abnormalities are found, alarms are sent to the operation and maintenance personnel by means of text messages, emails, pop-ups, and voices, and the relevant data is stored in the data storage and management module.
[0078] Please refer to Figure 2 and Figure 3 , the power metering data reliability verification method based on the fusion of random forest and deep belief network proposed by the present invention specifically includes:
[0079] Step 1: Real-time collect all-round power metering data and environmental data through a variety of sensor devices. The sensor devices include smart meters, voltage transformers, current transformers, power factor sensors, and environmental sensor matrices. Among them, the selected smart meter has an accuracy of up to 0.2S level, the accuracy of voltage transformers and current transformers is 0.2 level, the temperature measurement accuracy of environmental sensors can reach ±0.5°C, the humidity measurement accuracy is ±3%RH, and the electromagnetic interference intensity measurement range covers common interference frequency bands, ensuring the accuracy and comprehensiveness of the collected data.
[0080] Specifically, when the voltage data sequence V = [v 1 , v 2 , …, v n and the current data sequence I = [i 1 , i 2 , …, i n of a certain period are collected, first perform preliminary processing using the 3σ rule. Calculate the mean and the standard deviation of the voltage data. If there exists v j that satisfies , then it is determined as a suspected outlier and removed. The same processing method is adopted for the current data.
[0081] Adopt a combination of wavelet transform and adaptive filtering technology for denoising processing, effectively remove abnormal data, improve data quality, and provide a solid foundation for data analysis. For voltage data, select an appropriate wavelet basis function (such as db4 wavelet) for decomposition to obtain wavelet coefficients (j represents the scale, k represents the position) at different scales. By setting a threshold T (such as using the soft threshold method where σ is the estimated value of the noise standard deviation and n is the data length), process the high-frequency coefficients, that is, when , Then perform reconstruction to obtain the denoised voltage data. The denoising process for current data is similar.
[0082] Perform time synchronization processing on the collected data from different sensors. Let the timestamp of the data collected by the smart meter be t m , and the timestamp of the data collected by the environmental sensor be t e . Through the time calibration algorithm based on the Network Time Protocol (NTP), unify the time of all data to the same reference time t 0 . At the same time, convert the data into a unified format according to the standardized data format (such as CSV format) for subsequent data processing and analysis.
[0083] Use the Min-Max normalization method to normalize the data. For voltage data, let its minimum value be vmin , with the maximum value being v max , then the normalized voltage data
[0084]
[0085] is mapped to the interval [0, 1]. Similarly, similar normalization operations are performed on other data such as current and power.
[0086] Step 2: Extract key features from the preprocessed data, such as the amplitude, phase, harmonic spectrum, and power factor fluctuation of electrical energy measurement, and use data mining and machine learning algorithms to analyze the correlations between the features. By means of Pearson correlation coefficient, mutual information method, and grey relational analysis, etc., reveal the interactions between electrical topology, transducer characteristics, and electrical parameters, providing rich feature inputs for model training. Specifically:
[0087] For electrical energy measurement data, taking a voltage data V = [v(t 1 ), v(t 2 ), …, v(t m )] (m is the number of sampling points within this time period) with a time duration of T as an example, calculate its amplitude feature:
[0088]
[0089] t i represents the i-th sampling moment within the time duration T.
[0090] The phase feature obtains its phase information sequence by methods such as Hilbert transform on the voltage signal For current data, the amplitude and phase features can be calculated similarly. Calculate the harmonic spectrum features of voltage and current data through fast Fourier transform (FFT). Let V(k) and I(k) be the FFT transformation results of voltage and current respectively (k is the frequency component index), then the harmonic content feature is expressed as:
[0091]
[0092] where N is the number of FFT points. Here, calculate the ratio of the harmonic amplitude except the fundamental wave to the fundamental wave amplitude as the harmonic content feature. For the current harmonic content feature H i the calculation method is similar. The power factor fluctuation feature is measured by calculating the standard deviation σ PF of the power factor over a period of time. Let the power factor sequence be PF = [PF(t 1 ), PF(t 2 ), …, PF(t q )], then:
[0093]
[0094] wherein is the average value of the power factor.
[0095] For the characteristic data of the instrument transformer, assume that the nominal transformation ratio of a certain voltage transformer is K n , and the transformation ratio obtained through actual measurement is K m , then the transformation ratio accuracy characteristic The phase angle error characteristic measures the phase angle error value θ through a dedicated phase angle measuring device e . The load characteristic characteristic measures the output error of the instrument transformer under different load conditions (such as the load current changing from I min to I max in a certain step), and constructs the load characteristic curve E(I):
[0096] E(I) = k load I + b load
[0097] wherein, I is the load current, which is one of the key factors affecting the output error of the instrument transformer. k load is the slope of the curve, and its physical meaning is to reflect the sensitivity of the output error of the instrument transformer to the change of the load current. b load is the intercept of the curve, which reflects the inherent error or initial deviation of the instrument transformer when the load current is zero.
[0098] For the electrical topology data, the electrical topology structure is represented by the adjacency matrix A in graph theory. A ij represents the connection relationship between node i and node j. If there is a connection between node i and node j, then A ij = 1; otherwise, A ij = 0. The connection relationship characteristics of the nodes are described by calculating the degree of the nodes (N is the total number of nodes). For the branch impedance characteristics, according to Ohm's law (V ij is the voltage drop across both ends of branch ij, and I ij is the branch current), the impedance of each branch is calculated, and the mean value and standard deviation σ Z of the branch impedance are statistically analyzed. For the case of n branches, they are respectively expressed as:
[0099]
[0100] Calculate the Pearson correlation coefficient to analyze the relationship between characteristics. Assume that the voltage amplitude characteristic vector is A v , and the power factor characteristic vector is PF. Then the Pearson correlation coefficient between them is:
[0101]
[0102] wherein, PF i is the power factor of the i-th sample, is the value obtained by averaging all PF i It describes the average level of the power factor of the system or time period under study as a whole. is the voltage amplitude obtained by each sampling or measurement, which reflects the specific magnitude of the voltage at different times or positions, is the average value of the voltage amplitude.
[0103] The correlation between the electrical parameter fluctuations and the characteristics of the current transformer is calculated by the mutual information method. Let the feature vector of the electrical parameter fluctuations be X, and the feature vector of the current transformer characteristics be Y. Then the mutual information between them:
[0104]
[0105] where P(x,y) is the joint probability distribution of X and Y, and P(x) and P(y) are the marginal probability distributions of X and Y respectively. Grey relational analysis is used to study the relationship between the changes in the electrical topology structure and the fluctuations in the measurement accuracy of the current transformer. Let the change sequence of the electrical topology data be T, and the fluctuation sequence of the current transformer measurement accuracy be E. First, the data is initialized, and then the correlation coefficient is calculated:
[0106]
[0107] where ρ is the resolution coefficient, which is used to adjust the calculation of the correlation coefficient and controls the influence of the differences between different sample data on the calculation of the correlation degree to a certain extent. Generally, it takes 0.5; T i (k) represents the specific value of the i-th sample in the electrical topology structure data change sequence T at the k-th moment. In a power grid containing multiple substations, T i (k) represents the change value of a certain quantization index of the electrical connection mode of the i-th substation at the k-th time point. E j (k) represents the measurement accuracy fluctuation value of the j-th sample in the current transformer measurement accuracy fluctuation sequence E at the k-th moment. For current transformers of different types or positions, E j (k) records the deviation value relative to the standard measurement accuracy at a specific time or working condition.
[0108] Finally, the grey correlation degree is calculated:
[0109]
[0110] Step 3: Construct and train a fusion model of random forest and deep belief network to process multi-source heterogeneous data. The fusion model effectively identifies data anomaly patterns and non-linear relationships through ensemble learning and deep feature extraction. During the training process, data augmentation techniques such as random cropping, rotation, and adding noise are adopted to improve the generalization ability of the model. The input layer of the fusion model receives multi-source heterogeneous data that has undergone preprocessing and feature extraction, and after calculation, the error value of the measuring instrument is output at the output layer.
[0111] In the random forest part of the fusion model, let the training set be D = {(x 1 ,y 1 ),(x 2 ,y 2 ),…,(x n ,y n )}, where x i is the input sample (including electrical topology feature vector, transformer characteristic parameters, electrical parameter time series data, and environmental factor vector), and y i is the corresponding error value of the measuring instrument. m subsets D 1 ,D 2 ,…,D m are generated through bootstrap sampling, and the size of each subset is n (sampling with replacement). For each subset D j , a decision tree T j is constructed. At each split of the tree, k features (k < total number of features) are randomly selected from all features, and the optimal split point is found based on the Gini index (for classification problems) or mean squared error (for regression problems) among these k features. For example, for a binary classification problem (judging whether the data is abnormal), the Gini index is expressed as:
[0112]
[0113] where C is the number of classes, and p c is the proportion of samples belonging to class c), and the feature and split point that minimize the Gini index are selected for splitting, and continue splitting until the stopping condition is met (such as reaching the maximum depth or the number of samples in the node is insufficient).
[0114] In the deep belief network part of the fusion model, the layer-by-layer greedy training algorithm is adopted. First, train the first layer of restricted Boltzmann machine (RBM). Let the input data x be an n-dimensional vector, and the hidden layer h 1 has m neurons. The energy function of the RBM is:
[0115]
[0116] where a i ,b jThe biases of the visible layer and the hidden layer respectively, w ij are the connection weights, v i is the input node of the visible layer in the model structure of the deep belief network, representing the value of the i-th dimension of the input data x, h j represents the hidden layer h 1 the output value of the j-th neuron; by maximizing the likelihood function (where θ = {a, b, ω} are the RBM parameters) to learn the parameters of h 1 . The contrastive divergence (CD) algorithm is used for parameter update, which is used for the parameter update of the restricted Boltzmann machine. The weight update formula is:
[0117] Δw ij = ∈(<v i h j > data - <v i h j > recon )
[0118] where ∈ is the learning rate, <·> data represents the expectation over the data distribution, <·> recon represents the expectation over the reconstruction distribution. Then h 1 is used as the input of the next layer of RBM, and the training process is repeated until all layers are trained.
[0119] Finally, the parameters of the entire network are fine-tuned through the backpropagation algorithm to better fit the data. During the backpropagation process, the gradients of the loss function (such as the combination of the cross-entropy loss function and the mean squared error loss function) with respect to the model weights and biases are calculated. Let the cross-entropy loss function be:
[0120]
[0121] where, y i is the true label, is the model prediction output), and the mean squared error loss function is The total loss function L = αL CE + (1 - α)L MSE (α is the weight coefficient). The gradient descent method is used to update the parameters for overall adjustment of the network parameters to minimize the total loss function. The weight update formula is (η is the learning rate, w t is the weight at the t-th iteration).
[0122] Furthermore, during the training process of the fusion model, data augmentation techniques are combined. For example, random cropping is performed on the power metering data. Randomly intercept according to a certain time window or the number of data points in the original data sequence. Suppose the length of the original voltage data sequence is N, and the random cropping length is M (M < N), then randomly select the starting position s (0 ≤ s ≤ N - M) from the original sequence, and intercept to obtain a new voltage data sequence:
[0123] V new =[v(s), v(s + 1), …, v(s + M - 1)]
[0124] Perform a rotation operation on the data to change the order of the electrical parameter time series or the arrangement of the feature dimensions. For example, for a two-dimensional feature matrix X, rotate it 90° clockwise to obtain a new matrix X rot . By adding Gaussian noise, assume the original data is x, and the added Gaussian noise follows N(0, σ 2 ) distribution, and obtain the noise data:
[0125] x noisy =x + σ·randn(size(x))
[0126] where randn is a function that generates random numbers with a standard normal distribution, and σ is the standard deviation of the Gaussian noise, which determines the degree of dispersion of the noise data relative to the original data. Through these data augmentation operations, the sample set is expanded, and the generalization performance of the model is strengthened.
[0127] Step Four: Data Reliability Verification and Early Warning:
[0128] The fusion model has been trained. Input the power metering data collected, preprocessed, and feature-extracted in real time into the model. Assume the input data is x text , and the estimated value of the measuring instrument error output by the model is Set an error threshold ∈ (which can be adjusted according to the actual situation). If then it is determined that there is a reliability problem with the data. Or set the error change rate within a short time Δt = 10 minutes as:
[0129]
[0130] where is the set change rate threshold, which can be adjusted, and it is also determined that there is a reliability problem with the data.
[0131] When a data reliability risk is detected, the device immediately activates a multi-channel early warning mechanism. Through notification services such as text messages or phone calls, it timely sends abnormal information, such as the time and location of the abnormality, the type of error, and the approximate severity level, to the mobile phones or computers of the operation and maintenance personnel at the management end for early warning, ensuring the reliability of the power metering data and the stable operation of the system.
[0132] Furthermore, it should be noted that to verify the stability and reliability of the system of the present invention, the system of the present invention was successfully deployed in the complex power grid environment of a large industrial power park. This park covers various types of industrial enterprises, and the electrical equipment is complex and diverse, with extremely high requirements for the accuracy and reliability of power metering.
[0133] The device sensors are scientifically arranged according to the power grid topology structure of the park and the distribution of metering key nodes, continuously collecting power metering and environmental data. During the production process of a steelmaking workshop in a steel enterprise, a large arc furnace suddenly short-circuited, instantly causing a large drop in the grid voltage, a violent impact on the current, and strong electromagnetic interference. The device sensors quickly collected a series of abnormal change data.
[0134] These abnormal data first enter the data acquisition module. The 3σ rule quickly eliminates a large number of data points that are significantly deviated from the normal range, initially purifying the data. Wavelet transform and adaptive filtering techniques effectively remove high-frequency noise interference and improve data stability. After time synchronization and format unification processing, the data is transmitted to the preprocessing and feature extraction module. In this module, key features such as the voltage dip amplitude, current harmonic distortion, transformer saturation characteristics, and significant changes in the impedance of relevant branches in the electrical topology structure are accurately extracted, and through means such as Pearson correlation coefficient, mutual information method, and grey relational analysis, the close correlation relationships between these features are deeply revealed. For example, the voltage dip amplitude is positively correlated with the current harmonic distortion and has a causal relationship with the transformer saturation characteristics, providing a key basis for subsequent model judgment.
[0135] The processed characteristic data is input into the fusion model calculation module. Thanks to the accumulation of a large amount of training data and optimized parameter configuration in the early stage, the model quickly determines some data as outliers based on the learned complex data patterns and mapping relationships, and immediately triggers the early warning mechanism. Finally, in the data reliability verification and early warning module, the early warning information is sent to the management terminal, and the operation and maintenance personnel receive the emergency early warning information through text messages, pop-up windows, and voice alerts in a short time, and obtain a detailed abnormal data report in the email. According to the early warning information, the operation and maintenance personnel quickly rush to the scene with professional detection equipment and conduct a comprehensive and in-depth inspection of the electric energy metering device and related electrical equipment of the steel enterprise. After careful investigation, it is determined that a key voltage transformer is affected by strong electromagnetic interference and has a core saturation fault, resulting in serious deviation of the metering data. The operation and maintenance personnel promptly replace the faulty transformer and recalibrate and debug the equipment to restore the accuracy and reliability of the electric energy metering, successfully avoiding power trading disputes and power grid operation risks, and ensuring the normal production of the enterprise and the stable operation of the power system in the park.
[0136] The above specific embodiments are only several alternative embodiments of the present invention. Based on the technical solution of the present invention and the relevant revelations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.
Claims
1. A reliability verification method for electric energy metering data based on the fusion of random forest and deep belief network, characterized in that: It includes the following steps: Collect omnidirectional power metering data and environmental data in real time, and use the 3σ rule screening, wavelet transform and adaptive filtering technology to remove abnormal data, and then perform time synchronization processing and normalization processing to obtain preprocessed data; Extract key features from the preprocessed data, and analyze the relationship between features through Pearson correlation coefficient, mutual information method and grey relational analysis; Construct and train a fusion model of random forest and deep belief network. Its input layer receives multi-source heterogeneous data after preprocessing and feature extraction, calculates and outputs the error value of the metering device at the output layer, and performs monitoring and early warning according to the preset threshold.
2. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 1 is characterized in that: Collect omnidirectional power metering data and environmental data in real time through a variety of sensor devices, and the sensor devices include smart meters, voltage transformers, current transformers, power factor sensors and environmental sensors.
3. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 2 is characterized in that: The sensor device collects a voltage data sequence V=[v1, v2, ..., v n ] and the current data sequence i=[i1,i2,…,i n ], where the voltage data and current data are calculated in the same way using the 3σ rule; Calculate the mean of the voltage data and standard deviation If there is v j satisfy It is then judged as a suspected outlier and removed.
4. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 3 is characterized in that: When the voltage data and current data are processed, the wavelet basis function is used to decompose the data to obtain wavelet coefficients at different scales. Where j represents the scale and k represents the position; By setting the threshold T, the high-frequency coefficients are processed using the soft threshold method. Where σ is the estimated value of the noise standard deviation, and n is the data length; When hour, Then the denoised voltage data or current data is reconstructed.
5. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 2 is characterized in that: In the time synchronization process, the timestamp of the data collected by the smart meter is assumed to be t m , the timestamp of the environmental sensor collecting data is t e , through the time calibration algorithm based on the network time protocol, the time of all data is unified to the same reference time t0, and the data is converted into a unified format according to the standardized data format.
6. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 1 is characterized in that: The normalization processing uses the Min-Max normalization method to process the data related to voltage, current, power and environmental factors respectively; Assume the minimum value of voltage data is v min , the maximum value is v max , then the normalized voltage data is expressed as: Map it to the interval [0,1].
7. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 2 is characterized in that: The extraction of key features includes power metering data, transformer characteristic data, and electrical topology data, where For the electric energy metering data, suppose a voltage data V = [v(t1), v(t2), …, v(t m )], m is the number of sampling points in this time period, and its amplitude characteristics are calculated as follows: t i represents the i-th sampling time within the duration T; For the transformer characteristic data, it includes: Voltage transformer characteristics, assuming the nominal transformation ratio of the voltage transformer is K n , the transformation ratio obtained through actual measurement is K m , then the ratio accuracy characteristic is expressed as: Phase angle error characteristics, the phase angle error value θ is measured by the phase angle measurement equipment e ; Load characteristic features. By measuring the output error of the transformer under different load conditions, construct a load characteristic curve E(I): E(I)=k load I+b load Where I is the load current, k load is the slope of the curve, b load is the intercept of the curve; For electrical topology data, the adjacency matrix A in graph theory is used to represent the electrical topology structure. ij Represents the connection relationship between node i and node j. If there is a connection between node i and node j, then A ij =1, otherwise A ij =0; by calculating the degree of the node To describe the node connection relationship characteristics, where N is the total number of nodes; its branch impedance characteristics are expressed according to Ohm's law as: Where V ij is the voltage drop across the branch ij, I ij is the branch current, calculate the impedance of each branch, and calculate the average value of the branch impedance and standard deviation σ Z , for the case with n branches, they are expressed as:
8. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 7 is characterized in that: The calculation of the Pearson correlation coefficient to analyze the relationship between features includes: Assume that the voltage amplitude characteristic vector is A v , the power factor eigenvector is PF, then the Pearson correlation coefficient between them is expressed as: Among them, PF i is the power factor of the i-th sample, For all PF i The average calculated value is is the voltage amplitude obtained by each sampling or measurement, is the average value of the voltage amplitude; Calculate the correlation between the fluctuation of electrical parameters and the transformer characteristic data through the mutual information method. Let the electrical parameter fluctuation feature vector be X and the voltage transformer feature vector be Y, then their mutual information is expressed as: Where P(x,y) is the joint probability distribution of X and Y, and P(x) and P(y) are the marginal probability distributions of X and Y respectively; Let the electrical topology data change sequence be T and the transformer measurement accuracy fluctuation sequence be E. Initialize the data, and then calculate the correlation coefficient as: Where ρ is the resolution coefficient, which is used to adjust the calculation of the correlation coefficient; T i (k) represents the specific value of the i-th sample in the electrical topology data change sequence T at the k-th moment; in a power grid containing multiple substations, T i (k) represents the change value of a certain quantitative index of the electrical connection mode of the i-th substation at the k-th time point, E j (k) represents the measurement accuracy fluctuation value of the jth sample in the mutual inductor measurement accuracy fluctuation sequence E at the kth moment; for mutual inductors of different types or different locations, E j (k) record the deviation from the standard measurement accuracy at a specific time or under specific working conditions; The calculation of the grey correlation degree is expressed as:
9. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 1 is characterized in that: In the fusion model, for the random forest part, the training set is D = {(x1, y1), (x2, y2), …, (x n ,y n )}, where x i is the input sample, y i is the corresponding measuring instrument error value, and m subsets D1, D2, …, D are generated by self-service sampling. m , each subset size is n, for each subset D j , construct a decision tree T j At each split of the tree, k features are randomly selected from all features, where k < the total number of features, and the optimal split point is found among the k features based on the Gini index or mean square error.
10. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 1 is characterized in that: In the fusion model, the deep belief network part includes: Train the first-layer restricted Boltzmann machine. Let the input data x be an n-dimensional vector, and the hidden layer h1 has m neurons. The energy function of the restricted Boltzmann machine is: where a i ,b j are the biases of the visible layer and the hidden layer, respectively, w ij is the connection weight, v i is the input node of the visible layer in the model structure of the deep belief network, representing the value of the i-th dimension of the input data x, h j Represents the output value of the jth neuron in the hidden layer h1; by maximizing the likelihood function To learn the parameters of the hidden layer h1, where θ = {a, b, ω} is the restricted Boltzmann machine parameter; Use the contrastive divergence algorithm to update the parameters, and the weight update formula is: Δw ij =∈(<v i h j > data -<v i h j > recon ) where ∈ is the learning rate, <·> data represents the expectation on the data distribution, <·> recon represents the expectation on the reconstructed distribution; Take the hidden layer h1 as the input of the next layer of the restricted Boltzmann machine, and repeat the training process until all layers are trained; Fine-tune the parameters of the entire network through the backpropagation algorithm. During the backpropagation process, calculate the gradients of the loss function with respect to the weights and biases of the fusion model. Let the cross-entropy loss function be: Among them, y i is the true label, Predict output for the model; Assume the mean square error loss function is The total loss function L = αL CE +(1-α)L MSE , where α is the weight coefficient; Using the gradient descent method to update the parameters, the weight update formula is: Where η is the learning rate, w t is the weight of the tth iteration.
11. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 10 is characterized in that: The multi-source heterogeneous data includes transformer characteristic parameters, electrical topology feature vectors, electrical parameter time series data and environmental factor vectors; During the training process of the fusion model, combine the data augmentation technology to randomly crop the power metering data, randomly intercept according to a certain time window or the number of data points in the original data sequence. Let the length of the original voltage data sequence be N, and the random cropping length be M (M < N), then randomly select the starting position s (0 ≤ s ≤ N - M) from the original sequence, and intercept to obtain a new voltage data sequence: V new =[v(s),v(s+1),…,v(s+M-1)] Rotate the data to change the order of the electrical parameter time series or the arrangement of the feature dimensions. For a two-dimensional feature matrix X, rotate it 90° clockwise to obtain a new matrix X. rot ; By adding Gaussian noise, let the original data be x, and the added Gaussian noise obeys N(0,σ 2 ) distribution, the noise data is: x noisy =x+σ·randn(size(x)) Where randn is a function that generates standard normally distributed random numbers, and σ is the standard deviation of Gaussian noise.
12. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 1 is characterized in that: After the fusion model is trained, the electric energy metering data collected in real time, preprocessed and feature extracted is input into the fusion model. Suppose the input data is x text The output error value of the measuring instrument is Set the error threshold ∈, if The data are judged to be reliable data.
13. The method for verifying the reliability of electric energy metering data based on the fusion of random forest and deep belief network according to claim 12 is characterized in that: When reliability risks are detected in the data, the device immediately activates a multi-channel early warning mechanism and sends information containing abnormalities to the management end through a reminder service.
14. The reliability verification system of electric energy metering data based on the fusion of random forest and deep belief network is characterized by: include: Data acquisition module, collecting power metering and environmental data; The preprocessing and feature extraction module receives the electric energy metering and environmental data, and performs cleaning, denoising, normalization and feature extraction operations in sequence to obtain preprocessed data; A fusion model calculation module, which inputs the preprocessed data into the fusion model calculation module and performs calculation based on a random forest and deep belief network fusion model; The data reliability verification and early warning module receives the measuring instrument error value calculated by the fusion model calculation module, analyzes and judges according to preset thresholds and rules, and outputs abnormal information early warning prompts.
15. The electric energy metering data reliability verification system based on the fusion of random forest and deep belief network according to claim 14 is characterized in that: It also includes a data storage and management module, where the management end receives abnormal warning messages and stores warning data at the same time.
Citation Information
Cited By
Intelligent electric meter performance test method and system
CN120802163A