Electric energy metering error correction method and system based on data fusion
By combining multi-dimensional data collection and deep learning models, the problems of large measurement errors and poor versatility in traditional electricity metering methods have been solved, achieving high precision and intelligent adaptive correction of electricity metering.
Patent Information
- Application Number
- CN202510725851.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Traditional electricity metering methods suffer from large measurement errors and lack comprehensive consideration of multiple factors, resulting in decreased measurement accuracy. Furthermore, they have poor versatility and are difficult to adapt to the electricity metering needs under different environments and conditions.
A data fusion-based method for correcting electricity metering errors is adopted. This method involves multi-dimensional data acquisition, feature extraction using convolutional neural networks and recurrent neural networks, and error correction using a deep learning model. A closed-loop feedback mechanism is established to dynamically adjust model parameters to adapt to environmental changes.
It significantly improves the accuracy and robustness of electricity metering, reduces errors, enhances the system's adaptability and intelligence, and lowers maintenance costs.
Smart Images

Figure CN120234767B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application provides an electric energy metering error correction method and system based on data fusion, and belongs to the technical field of electric energy metering. BACKGROUND
[0002] In the power system, electric energy metering is an important basis for power transaction, cost accounting and power management. However, the traditional electric energy metering method has many problems, resulting in large metering error. On the one hand, the electric energy meter itself has mechanical wear, temperature influence, electromagnetic interference and other factors, which reduces the metering accuracy of the electric energy meter. For example, during the long-term operation of the electric energy meter, the friction torque between the rotating disc and the bearing will change, resulting in a non-proportional relationship between the rotating speed and the power, thereby generating an error. On the other hand, the performance of the mutual inductor also affects the accuracy of electric energy metering. The current transformer will cause certain data deviation due to the change of the current and voltage inside the line, and the early used mutual inductor equipment has low accuracy and cannot meet the advanced detection work standard. In addition, the load change of the secondary circuit system will also cause voltage drop, further increasing the metering error.
[0003] At present, although there are some correction methods for electric energy metering error, most of them are limited to the correction of a single factor and lack consideration of the comprehensive influence of multiple factors. Moreover, these methods often rely on specific equipment or parameters, have poor universality, and are difficult to adapt to the electric energy metering requirements under different environments and conditions. Therefore, a more comprehensive, accurate and universal electric energy metering error correction method is needed. SUMMARY
[0004] The application provides an electric energy metering error correction method and system based on data fusion, which solves the problems mentioned in the background.
[0005] The electric energy metering error correction method based on data fusion provided by the application comprises the following steps:
[0006] S1, collecting multi-dimensional data and preprocessing the collected multi-dimensional data;
[0007] S2, using a convolutional neural network to extract features from the operation data and mine spatial features in the data; using a recurrent neural network to extract time sequence features of the secondary circuit and environmental data and establishing time sequence association between the data;
[0008] S3, fusing the extracted spatial features and time sequence features; deeply mining and analyzing the multi-dimensional data;
[0009] S4, training the built-in error correction model; evaluating and verifying the trained error correction model; feeding back the evaluation results to form a closed-loop feedback mechanism.
[0010] The application provides a data fusion-based electric energy metering error correction system, including a memory, a processor and a computer program stored in the memory and capable of running on the memory, and the processor executes the program to realize the data fusion-based electric energy metering error correction method according to any one of the above.
[0011] The application has the following beneficial effects: through different sizes of convolution kernels and deep separable convolution technology, spatial feature multi-level extraction of equipment operation data and environmental data is realized while reducing the calculation complexity, and the comprehensiveness and efficiency of feature expression are improved; the time window is dynamically adjusted based on the temperature change rate, and a physical quantity-window size joint optimization model is combined, so that the method can flexibly cope with environmental mutations, and the robustness and stability of the model in a complex environment are significantly improved. The spatial features of the secondary loop data and the environmental data are fused through the cross-modal feature interaction layer, the time sequence dependence is captured through the combination of the cyclic unit, the joint modeling of the space-time features is realized, and the accuracy of the error correction is improved. The channel attention mechanism dynamically allocates weights, and the key channel features are strengthened; the deep separable convolution reduces the parameter quantity, and the gradient monitoring and truncation mechanism are combined to effectively alleviate the gradient vanishing / explosion problem, and the training efficiency and the model convergence speed are improved. The window size is dynamically adjusted through the sliding window and the autocorrelation coefficient decay rate, the key time points are focused through the combination of the self-attention mechanism, the short-term fluctuations and the long-term trends are considered, and the processing capacity for irregular time sequence data is significantly improved. The environmental data is mapped to the convolution kernel initialization parameter, and is input into the cyclic unit in combination with the secondary loop data, a physical quantity-time sequence coupling model is constructed, the correlation between the environmental factors and the electrical quantities is fully tapped, and the physical rationality of the error correction is improved. The dynamic sequence length optimization based on reinforcement learning automatically adjusts the input sequence length according to the model performance, balances the calculation overhead and the correction accuracy, and is suitable for diversified actual application scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 The method steps of the application are described. DETAILED DESCRIPTION
[0013] The preferred embodiments of the application are described below with reference to the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the application, and are not used to limit the application.
[0014] An embodiment of the application is shown in Figure 1 The data fusion-based electric energy metering error correction method includes the following steps.
[0015] S1, collecting multi-dimensional data through a multi-dimensional data acquisition system, and preprocessing the collected multi-dimensional data;
[0016] S2, using a convolutional neural network to extract features from the operation data of electric energy meters, transformers and other devices, and to mine spatial features in the data; using a recurrent neural network to extract time sequence features of secondary circuit and environmental data, and to establish time sequence correlation between data;
[0017] S3, the extracted spatial features and time sequence features are fused, a deep learning model is used to establish complex mapping relationships between data; through data fusion, multi-dimensional data is deeply mined and analyzed;
[0018] S4, according to the deep data fusion result, the built-in error correction model is trained using historical data, an optimization algorithm is used to determine the parameters of the model; a dynamic adaptive mechanism is introduced into the error correction model, and the model parameters are adjusted in real time according to the actual operation data; the trained error correction model is evaluated and verified by cross-validation, the performance index of the model is evaluated, and the real-time collected multi-dimensional data is input into the trained dynamic adaptive error correction model, the electric energy measurement data is corrected in real time; according to the evaluation of the correction result, the error correction model is further optimized and adjusted; and the evaluation result is fed back to form a closed-loop feedback mechanism.
[0019] The working principle of the above technical solution is: through a multi-dimensional data acquisition system, multiple data sources related to electric energy measurement are collected and preprocessed; the collected data includes:
[0020] Electric energy meter data: reflects the actual measurement of electric energy;
[0021] Transformer data: records the signals in the power system such as current and voltage;
[0022] Secondary circuit data: provides real-time state information of device control in the power system;
[0023] Environmental data: environmental parameters such as temperature and humidity, which will affect the operation of the device;
[0024] The spatial features are extracted from the operation data of devices such as electric energy meters and transformers by a convolutional neural network (CNN), so as to capture the spatial relationship between the device state and electric energy metering. The time series features in the secondary circuit and environmental data are extracted by a recurrent neural network (RNN), so as to identify the change trend of these data in the time dimension and the relevance between the change trend and the device operation state. After the feature extraction, the spatial features and the time series features are fused. At this time, the deep learning model will establish a complex mapping relationship between the multi-dimensional data. The process of data fusion can fully excavate the mutual influence between different data sources, reveal deep rules, and thus improve the accuracy of electric energy metering data. Based on the fused deep data, the error correction model is trained by using historical data. By using optimization algorithms such as back propagation algorithm and stochastic gradient descent method, the model parameters are continuously adjusted and optimized, so as to reduce the error in electric energy metering. A dynamic adaptive mechanism is introduced, so that the model can automatically adjust its parameters according to real-time operation data. After the model training is completed, the performance of the error correction model is evaluated by using cross-validation and other methods. Evaluation indexes such as mean square error and mean absolute error are used to measure the effect of the model. If the evaluation result does not meet the expectation, the model is optimized and adjusted until the performance meets the requirements. In actual operation, the multi-dimensional data collected in real time are input into the trained error correction model, and the online correction of electric energy metering data is performed, so as to ensure the real-time and accuracy of the correction result. At the same time, the correction effect is evaluated by analyzing the change of the error in electric energy metering data before and after the correction. In the process of electric energy metering correction, the effect of the model needs to be continuously monitored and fed back. Through the evaluation of the correction result, the error correction model is further optimized and adjusted, so as to form a closed-loop feedback mechanism.
[0025] The technical scheme has the effects that: through the collection and depth data fusion of multi-dimensional data, the error in electric energy metering can be effectively corrected, thereby significantly improving the accuracy of electric energy metering and reducing the error caused by traditional error correction methods; the dynamic adaptive mechanism combined with real-time data adjustment of model parameters can quickly respond during the operation of the power system, improve the real-time performance of error correction, reduce the correction delay, and ensure the timeliness of electric energy metering; the dynamic adaptive mechanism enables the error correction model to automatically adjust according to real-time operation data, enhances the adaptability of the system under different working conditions, and adapts to the changing and complex operating environment of the power system; through the combination of the deep learning model and historical data, rich information can be extracted from multi-dimensional data, the robustness of the error correction model to various interference factors and noise is improved, and the electric energy metering can still maintain high accuracy under different working scenarios; through the automation of model training and optimization process, the necessity of manual intervention is reduced, the influence of human error is reduced, the automation level of the electric energy metering system is improved, and the cost of manual maintenance and adjustment is reduced; through the deep learning method, multi-dimensional data can be fused and analyzed from multiple angles to comprehensively understand the operating state of the power system, improve the data analysis capability of the power system, and provide more accurate basis for subsequent decision support; through cross-validation and other evaluation methods, the efficiency and accuracy of the error correction model in the training and application process are ensured. The model is continuously optimized and adjusted to improve the efficiency of error correction and reduce the impact of unqualified models on the system; through comparative analysis of data before and after correction, the performance of the electric energy metering system can be continuously improved, and the quality of corrected data can be continuously improved, which is helpful for more accurate electric energy billing and load forecasting; the implementation of the technical scheme makes the electric energy metering correction process of the power system more intelligent, and can adaptively adjust according to different environments and system states, further promoting the intelligent development of power operation management; through the use of automatic error correction method, manual intervention and equipment adjustment are reduced, the maintenance cost of the electric energy metering system is reduced, and the potential loss caused by error is reduced.
[0026] In one embodiment of the present application, the S1 comprises:
[0027] S11, a timing collection task is set to collect multi-dimensional data at a preset time interval, and a real-time collection channel is started to obtain data in time when an abnormal situation occurs;
[0028] S12, during the collection of data, the collected data is checked for integrity to determine whether there is data loss or damage; if the data is found to be incomplete, the data is supplemented or repaired by interpolation;
[0029] S13, pre-process the collected data; and label the pre-processed data to mark normal data and abnormal data.
[0030] The working principle of the above technical solution is: through the setting of the collection interval (such as every minute, every hour, etc.) of the timing task, the multi-dimensional data is collected regularly to ensure the acquisition of continuous and periodic data flow. At the same time, a real-time data collection channel is established, and when an abnormal situation (such as current surge) occurs in the electric energy meter or system, the system can immediately trigger the real-time collection mechanism to quickly record the relevant data. This combination of timing and real-time collection can ensure the comprehensiveness and timeliness of data collection and capture abnormal situations; during data collection, the system will perform integrity checking on each batch of data. During data collection, the system will verify whether the data is lost or damaged. For example, if some data is not collected due to communication failure or hardware problems, the system will be able to discover this problem in time and take measures to perform data re-collection or repair. Re-collection can be performed by requesting data again or using interpolation methods for repair to ensure the integrity and continuity of the entire data set; after data collection and completion of integrity checking, the system will pre-process the data, and at the same time, the system will also label the data to clearly indicate which data belongs to the normal range and which data belongs to the abnormal data. These labeled data can provide key reference information for subsequent analysis and modeling, especially in identifying and processing abnormal data.
[0031] The effect of the above technical scheme is that: through the combination of the periodic collection task and the real-time collection channel, the data can be periodically collected under normal circumstances, and the related data can be quickly acquired in response to the occurrence of sudden events such as current mutation of the electric energy meter, so that the data loss or delay caused by the sudden events is avoided, and the efficient operation of the system and the timeliness of the data are ensured; the integrity checking mechanism is added in the data collection process, which can effectively find and repair the data loss or damage. Through the repair by supplementing or interpolation, the continuity and accuracy of the collected data are ensured, so that the subsequent analysis deviation caused by the data loss or error is reduced; the difference between the normal data and the abnormal data is marked, which helps to improve the accuracy of the abnormal detection model, so that the system can quickly and accurately identify and respond to potential problems; through the real-time data collection and abnormal processing mechanism, even if an abnormality occurs in the collection process, the system can still quickly remedy, so that the system error or performance degradation caused by the data loss or damage is avoided, the stability and robustness of the system are enhanced, and the electric energy metering system can stably operate in various environments; the data is preprocessed and marked in the data collection and processing stage, so that the accuracy and usability of the data are ensured, and reliable basis is provided for subsequent data mining, analysis and decision support. The marked data helps to perform retrospective analysis on the historical data and provide data support for future possible abnormal conditions, so that the intelligent level of the system is effectively improved.
[0032] In one embodiment of the present application, the S2 comprises:
[0033] S21, slicing the running data of the device according to a certain time window, performing convolution operation on the sliced data by using convolution kernels of different sizes through multiple convolution layers, and extracting spatial features in the data;
[0034] S22, adding a pooling layer after the convolution layer to perform dimension reduction processing on the convolution features while retaining important feature information;
[0035] S23, constructing sequence data in time sequence according to the secondary circuit and the environmental data, and processing the sequence data through a recurrent unit;
[0036] S24, encoding the features output by the recurrent unit to convert the time sequence features into a fixed-length vector representation.
[0037] The working principle of the above technical solution is to slice the operation data of devices such as electric energy meters and transformers according to a certain time window. For example, the data of each hour can be divided into multiple 10-minute time windows. This slicing helps to capture more fine-grained local features and provides more abundant data for subsequent processing. Next, the convolution layer in the convolutional neural network is used for processing, and different size convolution kernels (such as 3x3, 5x5, etc.) are used for convolution operation on the sliced data. These convolution kernels can extract spatial features of different scales from the data and identify change patterns of different granularities; after convolution operation, a pooling layer (such as a maximum pooling layer or an average pooling layer) is added, which is used for dimension reduction processing of the features after convolution. In addition to device operation data, other environment data related to device operation (such as temperature, humidity, etc.) also need to be processed. These data form sequence data in chronological order. For example, the temperature and humidity data of each hour can form a sequence, representing the environmental changes in a certain period of time. Through the introduction of recurrent units such as recurrent neural networks or long short-term memory networks, these time series data are processed. These recurrent units can capture the time sequence dependence in the data, learn the long-term and short-term dependence features in the data, so that the model can identify and process dynamic change patterns in the time series; after processing by the recurrent unit, the output time sequence features are encoded and converted into a fixed-length vector representation. Through this encoding process, complex time series data can be compressed into a more representative vector, which is convenient for subsequent analysis, prediction or other tasks.
[0038] The effect of the above technical solution is that: by slicing the equipment operation data according to the time window, and using different size convolution kernels for multi-layer convolution operation, the multi-level and multi-scale spatial features can be extracted from the data. This fine feature extraction method can effectively capture the details and patterns in the equipment operation process, providing rich feature information for subsequent analysis and prediction; adding a pooling layer after the convolution layer for dimension reduction processing can reduce the dimension of the data, and in turn reduce the calculation amount and storage demand. The pooling operation not only reduces the complexity of the model, but also retains important feature information in the data, avoids information loss, and helps the model to maintain good performance with less computing resources; by using the recurrent neural network or long short-term memory network to process the secondary loop and environmental data, the time sequence dependence relationship in the data can be effectively captured. This method can learn the long-term and short-term dependence features in the time series, which is very helpful for analyzing and predicting the long-term behavior and instantaneous change of the equipment, especially in a dynamic changing environment, which can timely adjust the prediction strategy; the features output by the recurrent unit are encoded and converted into a fixed-length vector representation, so that complex time series data can be converted into a more concise and structured feature representation. This improves the efficiency of data processing, and also makes subsequent classification, regression or other machine learning tasks simpler and easier to perform, and can ensure efficient feature utilization; the extraction of spatial features and time sequence features can more comprehensively capture the dynamic information of equipment operation and environmental changes. By combining the advantages of convolutional neural networks and recurrent neural networks, this technical solution can improve the prediction accuracy of the equipment state, and can more accurately respond to changes in operation under different scenarios, enhancing the robustness and adaptability of the system; through the dimension reduction processing of the pooling layer and the fixed-length vector representation after encoding, the model not only improves the calculation efficiency, but also maintains high-precision prediction ability while reducing the consumption of computing and storage resources.
[0039] In an embodiment of the present application, the S21 comprises:
[0040] Align the equipment operation data and the environmental data according to a unified timestamp format; and use the timestamp interpolation method to perform linear interpolation or physical model-based interpolation on the missing timestamps;
[0041] Calculate the autocorrelation coefficient of each equipment operation data; and divide a plurality of time windows to capture features of different time granularities;
[0042] Calculate the data information entropy in each time window, introduce the environmental data as an auxiliary variable, and dynamically reduce the time window when the temperature change rate exceeds a threshold value;
[0043] Different size of convolution kernel is used to extract local, mesoscopic and global spatial features; depth separable convolution is introduced in the convolution layer, which decomposes the standard convolution into depth convolution and point-wise convolution.
[0044] Channel attention mechanism is embedded in the convolution layer to give different weights to different channel features and strengthen the extraction of key features; the environmental data is mapped to the initialization parameters of the convolution kernel.
[0045] A cross-modal feature interaction layer is added after the convolution layer to fuse the convolution features of device running data and environmental data; the change rate of environmental data is introduced as a sensitive index for window selection; when the temperature gradient exceeds the threshold, the time window is automatically reduced.
[0046] A joint optimization model of physical quantity-window size is constructed, which takes environmental data and window size as joint variables to determine the optimal window size through optimization algorithm.
[0047] The working principle of the above technical solution is as follows: the device operation data (such as current, voltage, etc.) and environmental data (such as temperature, humidity, etc.) are first aligned according to a unified timestamp. Through the timestamp interpolation method, the missing data points are filled in to ensure data continuity. The interpolation method includes linear interpolation and interpolation based on physical models (such as temperature interpolation based on thermodynamic equations) to ensure that all data points have consistent time markers; by calculating the autocorrelation coefficients of each device operation data (such as current, voltage, etc.), the time-dependent range of the data is determined. For example, when the autocorrelation of the current data decays to a set threshold within 10 minutes, 10 minutes is selected as the initial window size; at the same time, multiple time windows of different sizes (such as 5 minutes, 10 minutes, 15 minutes, etc.) are divided to capture features at different time granularities; by calculating the information entropy of the data in each time window, the effectiveness of the window is evaluated. Select a window with moderate information entropy to avoid information redundancy caused by a large window or feature loss caused by a small window. At the same time, dynamically adjust the window size, introduce environmental data (such as temperature change rate) as an auxiliary variable, and automatically reduce the time window when the temperature change rate exceeds the threshold to improve the response speed to environmental changes; use different size convolution kernels (such as 3x3, 7x7) to extract features of different scales. Small convolution kernels (such as 3x3) capture local features, while large convolution kernels (such as 7x7) capture global trends; by decomposing the convolution operation into depth convolution and point-wise convolution through depth separable convolution, the computational complexity is reduced while maintaining the ability of feature extraction. For example, a 5x5 convolution kernel can be decomposed into a 5x1 depth convolution and a 1x5 point-wise convolution, reducing the computational complexity by about 75%; embed a channel attention mechanism in the convolution layer to assign different weights to different channels according to their importance. Through this mechanism, the extraction of key features can be strengthened, especially when the harmonic component increases, the weight of the corresponding channel is automatically increased to enhance the ability to capture harmonic features; after the convolution layer, a cross-modal feature interaction layer is introduced to fuse the convolution features of the device operation data and the environmental data. Methods such as Bilinear Pooling are used to generate joint feature representations through outer product operations. The rate of change of environmental data (such as temperature gradient) is introduced as a sensitive indicator for window size selection. For example, when the temperature gradient exceeds a set threshold (such as 5℃ / h), the time window is automatically reduced to improve the response speed to environmental changes. This method of dynamically adjusting the window size can optimize data processing in real time according to actual environmental changes; a joint optimization model of physical quantity-window size is constructed, combining environmental data (such as temperature, humidity) and window size, and using optimization algorithms (such as genetic algorithms) to find the optimal window size. For example, under the condition of a temperature of 30℃ and a humidity of 80%, the optimization algorithm will select 8 minutes as the optimal window size.
[0048] The effect of the above technical scheme is that: by aligning the equipment operation data and the environment data with a unified timestamp format, and applying linear interpolation or interpolation technology based on a physical model, the continuity of data from different sources in the time dimension can be ensured. This helps to solve the problem of data missing or inconsistency, and improves the accuracy of subsequent analysis; by calculating the autocorrelation coefficient of the equipment operation data, the time dependence range of the data can be effectively determined, so that a reasonable time window is determined, avoiding the subjectivity of manually selecting the window size, and by setting different time windows, the characteristics of different time granularities can be captured, ensuring multi-dimensional extraction of information; the concept of information entropy is introduced to select a suitable time window, avoiding information redundancy caused by a too large window or feature loss caused by a too small window. In addition, the size of the time window is dynamically adjusted in combination with the environment data (such as temperature), improving the response speed of the model to environmental changes and enhancing the adaptability of data processing; different size convolution kernels are used to extract local, mesoscopic and global features, and deep separable convolution is used to reduce the amount of calculation while maintaining the ability of feature extraction. This way optimizes the calculation efficiency and effectively captures different levels of features of equipment operation. Combined with the channel attention mechanism, the feature weights of each channel can be dynamically adjusted to improve the extraction ability of key features; the equipment operation data and the environment data are fused into convolution features, and a bilinear pooling method is introduced for feature interaction, effectively fusing the feature information of the two types of data and enhancing the perception ability of the model to complex systems. In addition, the rate of change of the environment data as a sensitive indicator for window selection can quickly respond to environmental changes, improving the flexibility and timeliness of the model; by jointly optimizing the physical quantity (such as temperature, humidity) and the window size, the time window can be adaptively adjusted, optimizing the accuracy and efficiency of data processing. The introduction of optimization methods such as genetic algorithm provides a more intelligent window selection strategy, which can dynamically adjust according to the actual environmental changes. By introducing environment data, not only the division and dynamic adjustment ability of the time window are optimized, but also the expressiveness and adaptability of the model are enhanced in many aspects, improving the performance of the model in complex and variable environments, further promoting the intelligent analysis and prediction ability based on time series data.
[0049] In an embodiment of the present application, the S23 comprises:
[0050] Align the secondary circuit data and the environment data according to the timestamp, and use linear interpolation to complete the missing data;
[0051] Normalize the data of different dimensions to map them to a unified interval;
[0052] Divide the time series into fixed-length windows for extracting local time series features; on the basis of the fixed window, introduce a sliding window mechanism to dynamically adjust the window step; capture the dependence relationship of different time granularities;
[0053] Calculate the autocorrelation coefficient decay rate of the time series, determine the effective length of the current sequence when the autocorrelation decays to a threshold, and dynamically adjust the window size;
[0054] Introduce information entropy as an auxiliary index for window length selection, select a window with moderate information entropy, and balance feature richness and computational complexity;
[0055] Use a hybrid structure of long short-term memory network (LSTM) and GRU to process short-term and long-term dependencies respectively;
[0056] Introduce self-attention mechanism in the loop unit, give higher weight to key time points in the time series, and strengthen the response ability to abnormal events;
[0057] Take environmental data as external input, and input it into the loop unit together with the secondary loop data; construct a physical quantity-time series coupling model to map temperature, humidity and other physical quantities to a unified feature space with current and voltage data;
[0058] During the training process, dynamically monitor the gradient change of the loop unit, and automatically truncate the long sequence when the gradient vanishes or explodes; based on reinforcement learning method, dynamically adjust the sequence length according to the model performance.
[0059] The working principle of the above technical solution is: aligning secondary circuit data (such as current, voltage) and environmental data through timestamps to ensure time consistency between different data sources; for missing data, linear interpolation is used for completion to ensure the integrity of time series data; normalize different dimensional data (for example, map current units A and temperature units ℃ to the [0, 1] interval), and apply Z-score standardization to eliminate skewness of data distribution and improve the efficiency of subsequent model training; divide the time series data into fixed length windows (for example, each hour of data is divided into 6 10-minute windows); introduce a sliding window mechanism to dynamically adjust the window step (for example, set to 5 minutes); calculate the autocorrelation coefficient decay rate of the time series and set a threshold (such as 0.1) to determine the effective length of the sequence. Introduce information entropy as an auxiliary index for window length selection to balance the richness of features and computational complexity. Windows with higher information entropy represent more information, while lower entropy represents less redundant information; use a hybrid structure combining LSTM and GRU to process short-term and long-term time series dependencies. The LSTM layer mainly captures hour-level time series dependencies, while the GRU layer is used to process minute-level time series dependencies; through a hierarchical structure, multi-scale time series modeling is performed to capture both short-term and long-term dependencies; introduce a self-attention mechanism in the recurrent unit, allowing the model to automatically focus on key time points in the time series and assign higher weights to these time points. For current mutations and other situations, the self-attention mechanism can help the model focus on the time series information before and after the mutation, improving the response capability of abnormal events; environmental data is input as external input together with secondary circuit data into the recurrent unit. The model can adjust the sensitivity of current data in combination with changes in environmental conditions (such as changes in current when temperature rises); construct a physical quantity-time series coupling model, combine the thermodynamic model to associate temperature and current data, predict the impact of temperature on current fluctuations, and use this information as auxiliary input to the recurrent unit; during training, dynamically monitor the gradient changes of the recurrent unit to prevent gradient vanishing or explosion. When gradient vanishing or explosion occurs, use the method of truncating long sequences to avoid the model from learning effective features; based on reinforcement learning method, dynamically adjust the sequence length according to the prediction error of the model. When the prediction error is large, increase the sequence length to capture more historical information; when the error is small, shorten the sequence length to reduce the amount of calculation and improve efficiency.
[0060] The effect of the above technical scheme is: by aligning data of different sources according to timestamps, data synchronization is ensured, and data deviation caused by time misalignment is avoided; linear interpolation is performed on missing data, which helps to retain the continuity and stability of the data and reduce the prediction error caused by data loss; normalization processing ensures that data of different dimensions can be compared and processed on the same scale, avoiding the influence of dimension difference on model training; Z-score standardization eliminates the skewness of data distribution, so that the model can better process abnormal values in the data and improve the convergence speed of training; by dividing the time series into fixed length windows and introducing a sliding window mechanism, local time series features can be effectively extracted, and the window size can be dynamically adjusted to adapt to different time scale dependencies; the introduction of the autocorrelation coefficient decay rate helps to dynamically adjust the window length, capture the effective part of the sequence, and avoid unnecessary computational complexity caused by too long sequences; the introduction of information entropy as an auxiliary indicator effectively balances the richness of the features and the complexity of the calculation, ensuring that the selected window can contain enough information; the use of a hybrid structure of LSTM and GRU can capture both short-term and long-term time series dependencies, improving the time series modeling capability of the model, especially for processing complex power system data; the introduction of the self-attention mechanism enables the model to automatically identify key time points, improving the sensitivity of anomaly detection; environmental data is used as external input and is jointly input with power data, which helps the model to dynamically adjust the sensitivity to different factors and improve the adaptability to complex environmental changes. For example, when the temperature rises, the model can automatically adjust the prediction of current fluctuations; by constructing a coupling model of physical quantities and time series, the working state of the equipment under different environments can be simulated, and the prediction accuracy and robustness are enhanced; dynamically monitoring the gradient change and automatically truncating the long sequence can avoid the problem of gradient vanishing or explosion, and improve the stability of the training process; based on reinforcement learning, the sequence length is dynamically adjusted, so that the model can adaptively select the most suitable sequence length, thereby improving the prediction accuracy of the model and reducing the computational overhead.
[0061] In one embodiment of the present application, the S3 comprises:
[0062] S31, the spatial features extracted by the CNN and the time series features extracted by the RNN are spliced to form a comprehensive feature vector;
[0063] S32, on the basis of feature splicing, an attention mechanism is introduced to weight process the importance of different features and highlight key features;
[0064] S33, a deep learning model is used as the model architecture of data fusion; according to the characteristics of the data and the task requirements, the number of network layers and neurons is designed;
[0065] S34, using the fused data to train the deep learning model, using the back propagation algorithm to continuously adjust the parameters of the model, so that the model can learn the complex mapping relationship between the data;
[0066] S35, during the training process, the model is evaluated regularly by using cross-validation, and the structure and parameters of the model are optimized according to the evaluation results.
[0067] The working principle of the above technical solution is: through the convolutional neural network; extract spatial features to capture spatial structure information in the input data (such as images, spatial data, etc.). At the same time, use the recurrent neural network; extract the time sequence feature to capture the time dynamic change of the data; then, splice the two features (spatial features and time sequence features) to form a comprehensive feature vector. This spliced feature vector can contain spatial information and time sequence information at the same time, providing more comprehensive input data for subsequent model learning; on the basis of feature splicing, the attention mechanism is introduced. The purpose of the attention mechanism is to weight the features according to their importance, so as to highlight the key features and suppress the unimportant features; by calculating the weight of each feature, the model can automatically identify which features are more important for the task prediction, and accordingly enhance its influence on the model output. This helps to improve the performance of the model, especially when facing complex data, it can avoid overfitting and improve the generalization ability; use the fused feature data as input, design a deep learning model architecture to process the data. The deep learning model can be a multi-layer neural network, and the number of network layers and neurons will be adjusted according to the characteristics of the data and the requirements of the task; at this time, the fused data (including spatial features, time sequence features and features weighted by the attention mechanism) will be used as the input of the model, and the model will process the data and learn the features through multiple layers of neurons; use the fused data to train the deep learning model. During the training process, the back propagation algorithm is used; continuously adjust the model parameters to minimize the loss function, so that the model can learn the complex mapping relationship between the data; through back propagation, the model continuously adjusts its parameters according to the error feedback, so that the model can gradually improve the prediction accuracy; during the training process, the model is evaluated regularly by using cross-validation. Cross-validation can effectively detect whether the model has overfitting, and provide evaluation indicators to help us judge the generalization ability of the model; according to the evaluation results of cross-validation, further optimize the structure and parameters of the model. This includes adjusting the number of network layers, the number of neurons in each layer, the learning rate and other hyperparameters, so as to ensure that the model can perform optimally in practical applications.
[0068] The effect of the above technical solution is that: by splicing the spatial features extracted by the CNN and the time sequence features extracted by the RNN, the spatial and time sequence information in the data can be captured at the same time, so that the model can comprehensively understand the data. This feature fusion method effectively improves the representation ability of the model for complex data structures. The introduction of the attention mechanism can weight different features, highlight key features, and automatically identify and strengthen features that are more important for task prediction. This helps the model to avoid paying attention to irrelevant features, improve its sensitivity to key information, and thus improve prediction accuracy. Through the attention mechanism, the model can dynamically adjust the importance of different features according to the characteristics of the input data, without manually specifying which features are more important. This adaptive feature selection mechanism can find effective features in complex data and improve the generalization ability of the model. According to the characteristics of the data and the task requirements, the number of network layers and neurons of the deep learning model is designed, which can make the model better adapt to the requirements of specific tasks. The flexibility of the deep learning architecture makes this solution have strong adaptability and can be applied to different types of tasks and data. Through the back propagation algorithm to train the model and combined with cross-validation for regular evaluation, overfitting can be effectively avoided, and the accuracy and generalization ability of the model in practical application can be ensured. Cross-validation provides accurate evaluation indicators to help quickly identify and adjust potential problems in the model, thereby continuously improving the performance of the model. Combined with the back propagation and cross-validation method, the parameters of the model can be continuously optimized during the training process, so that the model can more efficiently learn the complex relationships in the data, improving the effectiveness of model training.
[0069] In one embodiment of the present application, the S31 comprises:
[0070] S311, obtain spatial features by extracting geometric or positional attributes of the original data; obtain time sequence features by processing the time correlation of the time sequence data; preprocess the obtained spatial data and time sequence data to obtain a spatial feature set and a time sequence feature set;
[0071] S312, dimensionally expand the spatial feature set to obtain an enhanced vector of the spatial features, the enhanced vector of the spatial features including coordinate information, spatial distance, and direction features; further extract the time sequence feature set to obtain an enhanced vector of the time sequence features; the enhanced vector of the time sequence features including a timestamp, a time interval, and a change rate;
[0072] S313, based on different splicing strategies, splice the enhanced vector of the spatial features and the enhanced vector of the time sequence features respectively to form a preliminary comprehensive feature vector;
[0073] S314, normalizing the formed preliminary comprehensive feature vector to obtain a final comprehensive feature vector; performing further feature analysis and processing according to the final comprehensive feature vector to obtain a final multi-modal feature set;
[0074] S315, based on the final multi-modal feature set, modeling and inferring a space-time dependency relationship to obtain a space-time dependency relationship graph; based on the space-time dependency relationship graph, performing relationship inference and optimization to obtain a final relationship inference score.
[0075] The working principle of the above technical solution is: by processing the geometric or position attributes of the original data, the space-related features are extracted, including the coordinate position, shape, size and other spatial information of the object; the time series data is processed to analyze the time correlation therein and extract the time sequence features. The time sequence features are usually related to time variation, such as periodicity, trend, fluctuation, etc.; the extracted spatial data and time sequence data are preprocessed to finally form a spatial feature set and a time sequence feature set; the spatial features are dimensionally expanded to enhance the vector representation of the spatial features, including coordinate information, spatial distance, direction features; similarly, the time sequence feature set is further extracted to obtain an enhanced vector of the time sequence features. The time sequence feature enhancement vector includes timestamp, time interval and change rate; different splicing strategies are used to splice the enhanced spatial feature vector and the time sequence feature vector to form a preliminary comprehensive feature vector. The formed preliminary comprehensive feature vector is normalized; different features are compared on the same scale, and the normalized feature vector finally forms a multi-modal feature set containing spatial and time sequence two-dimensional information, based on the final multi-modal feature set, modeling and inferring a space-time dependency relationship; based on the space-time dependency relationship graph, relationship inference and optimization are performed to obtain a final relationship inference score.
[0076] The effect of the above technical scheme is that: by respectively extracting the spatial features and the time sequence features and performing enhancement processing thereon, the understanding and analysis capability of the model for complex space-time data is effectively improved, and the spatial and time dependence relationship in the data can be more accurately captured; by performing normalization processing on the spatial features and the time sequence features, the scale difference between different dimension features is effectively avoided, the feature processing process of the subsequent model is simplified, and the training efficiency of the model is improved; in the electric energy measurement of the power system, different models are used, and different models have different requirements for input features. By performing different processing on different data and different attributes, a feature set suitable for different models can be generated, and the performance and generalization capability of the model are improved. By dimension expansion and enhancement of the spatial and time sequence features, a spatial enhancement vector containing coordinate information, spatial distance, direction feature and the like, and a time sequence enhancement vector containing timestamp, time interval, change rate and the like are generated, and the feature expression of the data is further enriched; by inference and optimization based on the space-time dependence relationship graph, the identification and prediction capability of the model for the potential law in the space-time data is enhanced, the relationship inference score is improved, and the accuracy is improved; by comprehensive application of the multi-modal feature set, the model can better adapt to different types of data scenes, and the generalization capability of the model in different fields is enhanced; the normalization processing and the feature splicing strategy reduce the interference of data noise, and the stability and robustness of the model in processing unbalanced or missing data are improved; by effective enhancement and fusion of the space-time features, the model can better distinguish important spatial and time sequence features, reduce the influence of redundant information and noise, thereby reducing the prediction error and improving the precision of the model; in combination with the space-time dependence relationship graph and the optimized inference, the model can deeply mine the potential complex relationship in the space-time data, enhance the depth and breadth of the space-time dependence relationship modeling, and improve the accuracy and reliability of the inference result.
[0077] In an embodiment of the present application, the S313 comprises:
[0078] The time information is represented clockwise by time dependence, and the timestamp is represented counterclockwise; the clockwise represented coordinate information and the counterclockwise represented timestamp are first spliced based on the relationship between the position and the time, to obtain a first layer splicing;
[0079] The spatial distance is divided into X balanced small segments by a reference point, where X>2, the odd numbered segments are represented in a positive direction, the even numbered segments are represented in a reverse direction, the spatial distances represented in different directions are spliced to obtain a first spatial distance; the time interval is divided into N equal parts by a gradual transition function, where N>2, the odd numbered parts and the even numbered parts are spliced respectively, and then the spliced results are spliced again to obtain a first time interval; the first spatial distance and the first time interval are second spliced based on the relative relationship between the positions and the time change, to obtain a second layer splicing;
[0080] The direction feature and the change rate are thirdly spliced based on the direction and the change speed of the object motion, to obtain a third layer splicing;
[0081] The first 30% dimensions of the first layer splicing are extracted and spliced with the second layer splicing, and the last 70% dimensions of the first layer splicing are extracted and spliced with the third layer splicing; the two spliced parts are spliced again to form a preliminary comprehensive feature vector.
[0082] The working principle of the technical solution is as follows: firstly, the time information is represented in a clockwise direction, and the timestamp is represented in an anticlockwise direction; then the coordinate information represented in the clockwise direction and the timestamp represented in the anticlockwise direction are spliced to form a first layer of splicing; in the power system, the position (coordinate information) of the electric energy metering device is closely related to the time information. For example, the power consumption mode of the power consumption equipment at different geographical positions is different in different time periods. The clockwise and anticlockwise representation methods can fuse the position and time information in a unique way, so that the spliced features can better reflect the space-time correlation, which helps the model better understand the influence of the space-time factors on the electric energy metering in the power system; the clockwise and anticlockwise representation methods introduce directionality differences for the time information, so that the features of different position and time combinations have higher distinguishability in the spliced features. In power load prediction and other tasks, such distinguishability can help the model more accurately identify the power consumption features under different space-time conditions and improve the prediction accuracy; through the reference point, the spatial distance is divided into balanced small segments, and the odd segments are represented in the positive direction and the even segments are represented in the negative direction. For example, if the spatial distance is divided into three segments and represented as [1, 2, 3, 4, 5, 6], the representation after splicing becomes [1, 2, 4, 3, 5, 6], thereby changing the representation order and directionality of the spatial features. In the power system, the spatial distance between devices has relativity. The positive direction and the negative direction represent different aspects of the spatial relationship between devices. For example, there is mutual influence between some devices, and such influence is different with the positive and negative directions of the distance. By representing the spatial distance in this way, the spatial interaction between devices can be more accurately described. Then, the time interval is divided into equal-length segments by a gradual transition function, and the odd-numbered segments and the even-numbered segments are spliced respectively; for example, if the time interval is divided into three segments and represented as [1, 2, 3, 4, 5, 6], the splicing forms [1, 2, 5, 6, 3, 4], which represents and fuses the time interval in different orders; the gradual transition function can more carefully analyze the changes in the time interval. Splicing the odd-numbered and even-numbered segments respectively and then splicing them as a whole can highlight the change features of different stages in the time interval, so that the model can better capture the dynamic change law in the time interval. In the power system, for example, the power load changes unevenly in a day, and this method can more accurately describe the rhythm and mode of load change; compared with directly using the original time interval, the first time interval after division and splicing can provide more information about the time change. The spatial distance and the time interval that have been processed are spliced for the second time according to the relative change relationship between the position and the time; the second splicing splices the first spatial distance and the first time interval based on the relative relationship between the positions and the time change.In the power system, spatial distance and time change are interrelated. For example, in the process of power transmission, the farther the spatial distance, the longer the transmission time, and the greater the loss. Through the second splicing, the two kinds of information can be integrated together, so that the model can comprehensively consider the influence of space-time distance factors on electric energy measurement; based on the moving direction and change speed of the object, the dynamic characteristics are spliced with other characteristics for the third time; the first 30% of the dimensions are extracted from the features spliced in the first layer, and spliced with the features after the second layer splicing; the last 70% of the dimensions are extracted from the features spliced in the first layer, and spliced with the features after the third layer splicing; the first splicing contains position and time information, the second layer splicing integrates space-time distance information, and the third layer splicing combines direction features and change rate. According to the 30%:70% splicing ratio, the weight of different types of feature information in the preliminary comprehensive feature vector can be balanced. In the power system, different types of features have different influences on electric energy measurement, and this proportioning can make the model more reasonably use various feature information; through this proportioning splicing method, the preliminary comprehensive feature vector can contain both the basic correlation information of position and time (the first splicing part) and highlight important features such as space-time distance and direction change (the second layer and the third layer splicing part); in this way, the model can more accurately combine the features at each level to form a preliminary comprehensive feature vector.
[0083] The effects of the above technical solutions are as follows: By representing time information in different directions, different dimensions of time change can be captured more accurately, effectively identifying patterns of time change, improving prediction accuracy, and enhancing the ability to recognize time series patterns; by dividing spatial distance into multiple segments and combining forward and reverse representations, the accuracy of spatial feature expression is improved; by reasonably dividing and splicing spatial distance and time intervals, redundant features are reduced, while the complexity of the calculation process is optimized, avoiding excessive feature accumulation in traditional feature engineering and improving computational efficiency; by selecting the dimension of the first layer of splicing, the accuracy of features can be ensured while maintaining efficient computation, avoiding the risk of overfitting that is easily caused by simple splicing in existing technologies; by combining the splicing of the object's motion direction and change speed, the model's ability to recognize dynamic changes in object behavior is enhanced, improving accuracy; the equal-length division and splicing of time intervals helps the model understand the changing patterns in the time series, improving the model's performance in dynamically changing tasks. By fusing features from the first, second, and third layers of splicing, complex spatial and temporal relationships can be captured at different levels, thereby enhancing the model's understanding of spatiotemporal data. Spatial and temporal splicing fully considers the relative relationships between locations and changes over time, thus improving the model's adaptability to dynamic spatiotemporal changes in complex environments. Through multi-layered and multi-dimensional splicing, the model can analyze data from multiple perspectives, helping to reduce dependence on specific patterns and enhancing its generalization ability. Reasonable feature selection and splicing strategies effectively avoid the introduction of excessive redundant information, reducing the possibility of overfitting and improving the model's performance on new data. By merging the splicing results from different levels, a comprehensive feature vector is ultimately formed, enhancing the model's expressive power. Multi-level feature optimization reduces the complexity required for each calculation, improving the system's processing speed and response time when handling real-time data.
[0084] In one embodiment of the present invention, step S4 includes:
[0085] S41. Based on the deep data fusion results, train the built-in error correction model;
[0086] S42. Evaluate and validate the trained error correction model;
[0087] S43. Provide feedback on the evaluation results to form a closed-loop feedback mechanism.
[0088] In one embodiment of the present invention, S41 includes:
[0089] S411, according to the result of depth data fusion, an error correction model is constructed; the fused features are taken as input, the electric energy metering error is taken as output, and a mapping relationship between the input and the output is established;
[0090] S422, the parameters of the error correction model are initialized, and a random initialization or a pre-trained model parameter initialization method is adopted to provide an initial state for the training of the model;
[0091] S433, a large amount of historical electric energy metering data is collected, the historical electric energy metering data includes multi-dimensional data and corresponding actual error values, and the historical electric energy metering data is taken as a data set for model training;
[0092] S444, the error correction model is trained, and hyperparameters are set, the hyperparameters include a learning rate and an iteration number, and the training process is controlled;
[0093] S455, based on an online learning algorithm or an incremental learning algorithm, the model can adjust parameters in real time according to actual operation data; and based on a preset parameter updating strategy, parameters of the model are dynamically adjusted according to differences between actual operation data and model prediction results.
[0094] The working principle of the above technical solution is as follows: different sources of data (such as spatial features, time sequence features, etc.) are fused by a deep learning algorithm to obtain a comprehensive feature set. Then, these fused features are used to construct an error correction model. The input of the model is the fused features, and the output is the electric energy metering error. By establishing the mapping relationship between the input and the output, the error correction model can automatically predict and correct the electric energy metering error; the parameters of the error correction model are initialized. Common methods include random initialization (assigning initial values to model parameters through random values) or using pre-trained model parameters for initialization. The purpose of initialization is to provide a reasonable starting point for model training, reduce the convergence time during training, and improve training efficiency; a large amount of historical electric energy metering data is collected, including multiple dimensions of data (such as current, voltage, temperature, etc.), and the actual error values corresponding to these data. These data will be used as a training set for the model to learn the relationship between input features and electric energy metering error. The quality and diversity of historical data are crucial to the training effect and generalization ability of the model; batch gradient descent or small batch gradient descent is used to train the error correction model. Batch gradient descent adjusts model parameters by performing a complete update on all training samples, while small batch gradient descent uses a portion of data (small batch) for update at each iteration, which is more efficient. In this process, appropriate hyperparameters such as learning rate and iteration number are set to control the training process of the model and ensure that the model converges to the optimal solution; in order to enable the model to adapt to the actual data changes in real-time operation, online learning algorithm or incremental learning algorithm is introduced. Through these algorithms, the model can immediately adjust its parameters after receiving new data to ensure that it can continuously update and improve according to new input data. Through the pre-set parameter update strategy, the model can dynamically adjust its parameters according to the difference between real-time running data and prediction results, maintaining efficient electric energy metering correction capability, especially under different environmental or load conditions.
[0095] The effect of the above technical scheme is that: through deep data fusion and training of the error correction model, the error in electric energy measurement can be effectively corrected, and the accuracy of electric energy measurement is improved. The model can dynamically adjust the correction parameters of electric energy measurement according to the data characteristics of different dimensions, thereby reducing the influence of external environment and equipment factors on the measurement results; by using online learning and incremental learning algorithms, the error correction model can update the parameters in real time according to the actual operation data, so as to ensure that the model can adapt to the changing environment and operating conditions in the power system; by using batch gradient descent method or small batch gradient descent method for training, the model can efficiently process a large amount of historical electric energy measurement data and converge to the optimal solution in a relatively short time. The reasonable setting of the hyperparameters can further accelerate the training process of the model and reduce unnecessary computational burden; deep data fusion comprehensively processes data characteristics of multiple dimensions, which helps to improve the generalization ability of the model and avoid overfitting problem caused by relying on a single feature. The diversity and richness of historical electric energy measurement data provide sufficient samples for the model, enhancing its prediction ability for unknown data; the strategy of adjusting the model parameters based on real-time data can ensure that the error correction model is optimized at any time. Through this dynamic adjustment mechanism, the model can continuously learn from new data, optimize the correction strategy, and improve the stability and reliability of electric energy measurement in long-term operation; since the scheme can perform real-time correction and optimization on the basis of automation, the need for manual intervention is reduced, and the cost of manual adjustment and maintenance is reduced. The system can learn and correct independently, improving the long-term stability and reliability of electric energy measurement and reducing errors caused by human operation.
[0096] In an embodiment of the present application, the S42 comprises:
[0097] S421, evaluate and verify the trained error correction model by using cross-validation method, divide the data set into multiple subsets, and take each subset as a test set in turn, and the remaining subsets as training sets, and calculate the average performance index of the model;
[0098] S422, evaluate the performance index of the model to judge the accuracy and reliability of the model; input the multi-dimensional data collected in real time into the trained dynamic adaptive error correction model to correct the electric energy measurement data in real time;
[0099] S423, output the corrected electric energy measurement data to the monitoring system or the measurement equipment; and evaluate and analyze the corrected electric energy measurement data to calculate the error change before and after correction.
[0100] The working principle of the above technical solution is: the cross-validation method is used to evaluate and verify the error correction model. By dividing the data set into multiple subsets and taking each subset as the test set in turn, the remaining subsets as the training set, the performance of the model is evaluated. This method can avoid overfitting and ensure the generalization ability of the model, thereby improving the performance of the model on different data sets. Multiple performance indicators of the model are evaluated. These indicators are used to measure the accuracy of the model's predictions, the size of the errors, and the degree of fitting of the model to data changes; Once the model is evaluated and obtains good performance indicators, the system inputs the real-time collected multi-dimensional data (such as voltage, current, load, etc.) into the trained error correction model, and the model will perform real-time electric energy metering correction according to these input data. The corrected electric energy metering data will be output to the monitoring system or metering equipment for further analysis and monitoring. The corrected data can more accurately reflect the true electric energy consumption, helping operation and maintenance personnel to perform more accurate power management. In addition, the corrected electric energy metering data also needs to be further evaluated and analyzed. Specifically, by calculating the error change (such as absolute error change, relative error change, etc.) before and after correction, the effect of error correction can be intuitively evaluated. For example, if the error before correction is large and the error after correction is significantly reduced, the effect of the correction model is considered successful. This analysis process also helps to continuously optimize the error correction model.
[0101] The effect of the above technical solution is: by training and verifying the error correction model, the error in electric energy metering can be significantly reduced, providing more accurate data. This enables the power system to be more accurate in metering, monitoring and management, thereby improving energy utilization efficiency and reducing unnecessary losses; by using a dynamic self-adaptive error correction model, real-time correction can be performed according to real-time collected multi-dimensional data (such as voltage, current, load, etc.). By evaluating and verifying the performance of the model through cross-validation and other methods, the overfitting phenomenon of the model can be effectively avoided, and the generalization ability of the model can be improved. The model is tested on different subsets to ensure its stability and reliability in various power system situations; the corrected electric energy metering data is output to the monitoring system or metering equipment for real-time monitoring and analysis by operation and maintenance personnel. This not only helps to improve the efficiency of daily management, but also provides accurate data support for decision-making, ensuring that the operation of the power system is more efficient and safe; evaluating and analyzing the corrected electric energy metering data can provide valuable feedback for further optimization of the model. By calculating the error change before and after correction, the correction effect can be understood in real time, and based on this feedback, the model can be further improved to improve the accuracy of the overall electric energy metering; by reducing the error in electric energy metering, unnecessary energy waste and electricity bill calculation errors are avoided, which can effectively reduce the operating costs of enterprises. In addition, accurate electric energy metering helps to achieve better load management, further improving energy saving effects.
[0102] In one embodiment of the present application, the S421 comprises:
[0103] The environmental data is introduced as a hierarchical basis, and the data set is divided into multiple environment-related layers; the hierarchical threshold is dynamically adjusted according to the change range of the environmental physical quantity;
[0104] In the time series data, a sliding window method is used to divide the subsets, and each subset contains consecutive time points; the sliding window length is dynamically adjusted according to the time resolution of the electric energy metering data;
[0105] The geographical location information in the multi-dimensional data is subjected to cluster analysis, and the data points with similar geographical locations are divided into the same subset;
[0106] Each feature in the multi-dimensional data is assigned a weight, and in the model training process, the weight is dynamically adjusted according to the importance change of the feature; the environmental data is used as auxiliary information for anomaly value detection; and the corresponding electric energy metering data is marked as a potential anomaly value;
[0107] The similarity of the feature distribution between different subsets is tested; when the distribution similarity between the subsets is too high, the hierarchical or clustering parameters are readjusted; the time series features are extracted from the environmental data and used as auxiliary inputs for cross-validation; and the time series features are embedded into the subset division process to enhance the time dependence of the subset division;
[0108] The influence degree of the change of the environmental physical quantity on the subset division result is calculated, and according to the sensitivity analysis result, a sensitivity threshold is set, and when the change of the environmental physical quantity exceeds the threshold, the re-optimization of the subset division is triggered;
[0109] In the model training process, the change of the environmental physical quantity is monitored in real time, and the parameters of the cross-validation are dynamically adjusted; and according to the change mode of the environmental physical quantity, the optimal cross-validation strategy is automatically selected;
[0110] The consistency of the environmental physical quantity distribution between different subsets is verified, and the consistency of the environmental physical quantity distribution between different subsets is quantified.
[0111] The working principle of the above technical solution is: introducing environmental data (such as temperature fluctuation rate, humidity gradient, etc.) as the basis, and performing hierarchical processing on the original data set. By dynamically adjusting the hierarchical threshold (such as increasing the number of layers when the temperature fluctuation rate exceeds ±2℃), it is ensured that each subset has consistency in environmental characteristics; the specific operation is based on the changes of environmental data to determine the division method of different data, further ensuring that the characteristics of each data subset are more uniform and consistent; the sliding window method is used to divide the time series data, ensuring that each subset contains consecutive time points. This method adapts to the time resolution of electric energy metering data (such as recording a data point every 15 minutes), and dynamically adjusts the window size according to the actual situation, avoiding mixing cross-day or cross-season data, and ensuring the time consistency of the data subset; cluster analysis is performed on the geographical location information in the data, and points with similar geographical locations are divided into the same subset. According to the electricity consumption characteristics of different regions (such as the difference in power use between urban and rural areas), the radius of clustering is dynamically adjusted to ensure that the spatial characteristics within the same subset are more consistent; weights are assigned to each feature (such as voltage, current, power factor), and the determination of the weight is based on the correlation between the feature and the target variable (such as electric energy metering error). During the training process, the weight is dynamically adjusted as the importance of the feature changes to avoid the adverse effects of irrelevant features on data division; environmental data (such as lightning activity frequency) is used to assist in detecting outliers. If lightning activity is frequent, it will affect the accuracy of electric energy metering data, thereby marking the related data as potential outliers; these outliers are repaired by interpolation or adjacent value replacement to ensure the integrity and consistency of the data within the subset; electric energy data from different environments or geographical locations may follow different distributions, and direct mixing training may cause model bias or decreased generalization ability, and may also lead to unreasonable subset division (such as too fine layering or too large clustering radius); the similarity of feature distribution between different subsets is tested by using KL divergence or JS divergence; if the feature distribution between some subsets is too similar, the layering or clustering parameters are adjusted to increase the difference between different subsets, enhancing the generalization ability of the model; the time series features in the environmental data (such as daily temperature variation, seasonal humidity fluctuation) are extracted and used as auxiliary inputs for cross-validation. Using methods such as autoencoder or LSTM (Long Short-Term Memory) network, these time series features are embedded into the subset division process, making the data division more consistent with the time dependence; the sensitivity analysis is used to calculate the influence of environmental physical quantity changes on subset division (such as the influence of temperature fluctuation rate on division results). When the change of environmental physical quantity exceeds a certain threshold, re-optimization of subset division is triggered to ensure that the division result is more sensitive to environmental changes; during the model training process, the changes of environmental physical quantities are monitored in real time, and the parameters of cross-validation (such as the number of subsets, the hierarchical threshold) are dynamically adjusted. According to the change pattern of environmental physical quantities, the cross-validation strategy is selected to improve the accuracy of the model; the consistency of environmental physical quantity distribution between subsets is verified.The consistency of environmental characteristics between subsets is quantified using statistical methods such as chi-square test or t-test, ensuring that the integrity of environmental characteristics is not damaged during data partitioning.
[0112] The above technical solution has the following effects: By dynamically adjusting the data layering threshold based on changes in environmental data such as temperature and humidity, the consistency of each subset in environmental characteristics is ensured, reducing the heterogeneity of data under different environmental conditions and improving the stability of model training; By using the sliding window method and geographical location clustering, the correlation between time series and geographical location is considered, so that each subset is not only continuous in time but also has consistent characteristics in space, thereby improving the spatio-temporal consistency of data modeling; By combining environmental factors (such as lightning activity frequency) for anomaly detection, potential abnormal data can be more accurately identified and repaired through interpolation or neighboring value replacement, avoiding the negative impact of abnormal data on model training; By introducing KL divergence or JS divergence and other measurement methods, the distribution difference between subsets can be monitored, and the layering or clustering parameters can be dynamically adjusted according to environmental changes, ensuring the independence and difference between subsets and preventing overfitting or uneven data distribution; By extracting the time series features of environmental data (such as temperature change patterns and humidity fluctuations) and using them as auxiliary inputs for cross-validation, the accuracy and generalization ability of the model are further improved, especially in terms of long time scales and seasonal changes; By monitoring the changes in environmental physical quantities in real time and dynamically adjusting the cross-validation strategy, the most suitable validation method can be selected under different environmental conditions, further improving the reliability and generalization ability of the model; By calculating the influence of environmental physical quantities on subset division, the sensitivity of the model can be evaluated, and the data partitioning strategy can be flexibly adjusted based on the sensitivity analysis results to avoid unnecessary errors. Through the above technical solution, not only the randomness of subset division is solved, but also the robustness and physical meaning of cross-validation are enhanced through the introduction of environmental physical quantities, multi-dimensional feature weighting, and time series embedding; At the same time, by combining the sensitivity and consistency of physical quantity changes, the reliability and interpretability of the cross-validation results are ensured.
[0113] In one embodiment of the present application, the S43 comprises:
[0114] S431, according to the evaluation of the correction result, further optimizing and adjusting the error correction model; if the performance of the model does not meet the requirements, the structure, parameters or training algorithm of the model can be adjusted to improve the correction effect of the model.
[0115] S432, optimizing the feature extraction method in the data fusion process, converting different feature combinations and feature selection methods to improve the quality and representativeness of the features;
[0116] S433, feedback the evaluation of the correction result to the model training and data acquisition link, provide basis for further optimization of the model and adjustment of the data acquisition strategy; form a closed-loop feedback mechanism, and continuously optimize and improve the error correction model.
[0117] The working principle of the above technical solution is that in the data processing and analysis process, the error correction model evaluates the difference between the actual result and the expected result, and feeds back the correction effect of the model. If the performance after correction does not meet the expectation, the system will further optimize the model. The optimization content can include adjusting the structure, parameters or training algorithm of the model. The purpose of this process is to improve the accuracy and adaptability of the correction model, thereby improving the accuracy of the overall system. In the data fusion process, feature extraction is a key step, and the quality and representativeness of the features directly affect the performance of the model. By optimizing different feature combinations and feature selection methods, the expression ability and discrimination of the features can be effectively improved. The optimized features can better reflect the core information of the data, thereby providing more effective input for the error correction model. While the correction model and the feature extraction method are continuously optimized, the correction result is also fed back to the model training and data acquisition link. By evaluating the correction result, the subsequent data acquisition strategy and model training process can be guided. For example, if the correction effect is poor, the data acquisition method needs to be adjusted or the model structure needs to be further adjusted. Through this feedback mechanism, a closed loop is formed to ensure that the system can self-adjust and continuously optimize.
[0118] The effect of the above technical solution is that by optimizing and adjusting the error correction model, the performance of the correction model can be improved for different data environments and requirements, ensuring that the model can more accurately correct errors and thereby improving the overall correction effect of the system. By optimizing the feature extraction method, different feature combinations and selection methods can be converted to extract more representative and discriminative features from the data, thereby improving the quality of the model input and enhancing the performance of the model in data fusion. Feedback of the correction result to the model training and data acquisition link not only provides basis for further optimization of the model, but also allows real-time adjustment of the data acquisition strategy. This closed-loop feedback mechanism ensures that the system can dynamically adapt to different environmental changes, continuously optimize the error correction effect, and thereby improve the intelligence and adaptability of the entire system. Through continuous optimization and adjustment, the system can improve itself to adapt to changing data environments and requirements. This self-adaptability improves the long-term effectiveness of the system, reduces the need for manual intervention, and reduces system maintenance costs; by continuously optimizing and adjusting the parameters and features of the model, the correction model can better adapt to different types of data and complex application scenarios, enhancing the generalization ability of the model and improving its stability and reliability in practical applications.
[0119] In an embodiment of the present application, the S432 comprises:
[0120] Introduce environmental physical quantities as a bridge to map different data sources to a unified physical feature space; distribution differences in space;
[0121] Align time series data using the DTW algorithm to address inconsistent sampling frequencies across different data sources; dynamically adjust interpolation strategies based on changes in environmental physical quantities;
[0122] Perform Granger causality tests on different data sources to identify causal relationships between data sources and avoid interference from irrelevant data sources during feature extraction; based on causal relationships, construct Bayesian networks or structural equation models to quantify the strength of dependence between data sources and guide feature selection;
[0123] Generate new features using environmental physical quantities and use polynomial regression or factor decomposition machines to model the interaction effects between different features and uncover nonlinear relationships between data sources;
[0124] Design a multi-modal autoencoder to input different data sources into a shared encoding layer to learn joint feature representations across modalities; introduce an attention mechanism in the autoencoder to dynamically allocate weights to different data sources, highlighting key features and suppressing noise features;
[0125] Select features that are highly correlated with physical quantity changes based on their correlation with the target variable;
[0126] Use environmental physical quantities as fusion weights to weight and fuse features from different data sources to generate a comprehensive feature vector; use mutual information or conditional entropy to evaluate the amount of information between the fused features and the target variable;
[0127] Verify the distribution consistency of fused features under different environmental physical quantity conditions to ensure that feature extraction does not disrupt the physical connections between data sources; use KL divergence or JS divergence to quantify the distribution differences of fused features under different physical quantity conditions; use environmental physical quantities as auxiliary evaluation indicators for model performance.
[0128] The working principle of the above technical solution is: by introducing environmental physical quantities (such as temperature, humidity, etc.) as a bridge, different data sources (such as voltage, current, meteorological data) are unified and mapped to a shared physical feature space. This allows different types of data to be compared and analyzed in the same physical context, reducing differences between different data sources; using non-negative matrix factorization (NMF) or generative adversarial network (GAN) technology, learn the optimal mapping matrix between data sources, minimize the distribution difference of different data sources in the physical feature space. This can better unify different data sources into a common representation space; use dynamic time warping (DTW) algorithm to align time series data (such as energy metering data and meteorological data), solve the problem of inconsistent sampling frequency between different data sources. Through this alignment, the consistency and comparability of the data in the time dimension can be ensured; according to the mutation of environmental physical quantities (such as temperature changes, etc.), dynamically adjust the data interpolation strategy (for example, use linear interpolation or spline interpolation), ensure that the smoothness of the data is not affected after time alignment; through Granger causality test, identify the causal relationship between different data sources, avoid irrelevant data sources interfering with feature extraction. Use the causal relationship to further build a Bayesian network or structural equation model to quantify the dependency relationship between different data sources and guide the subsequent feature selection process; use environmental physical quantities (such as light intensity, etc.) to generate new features, such as combining light intensity and voltage data to generate coupled features. Use polynomial regression or factorization machine (FM) to model the interaction effects between features and mine the nonlinear relationships between data sources; design a multi-modal autoencoder to input different data sources into a shared encoding layer to learn a joint feature representation across modalities. Introduce an attention mechanism in the autoencoder to dynamically adjust the weights of different data sources, highlighting key information and suppressing noise features; based on the correlation between environmental physical quantities (such as wind speed) and target variables (such as energy metering error), select highly correlated features. Use Shapley value or LIME method to dynamically evaluate the importance of features, avoiding the limitations of static feature selection; use environmental physical quantities as fusion weights to weight and fuse features from different data sources to generate a comprehensive feature vector. Evaluate the amount of information between the fused features and the target variables through mutual information or conditional entropy to ensure that the fused features have high effectiveness; verify the distribution consistency of the fused features under different environmental physical quantity conditions through KL divergence or JS divergence, etc. to ensure that feature extraction does not destroy the physical association between data sources. At the same time, environmental physical quantities can be used as auxiliary indicators for model performance evaluation to ensure model adaptability under different environmental conditions.
[0129] The effect of the above technical solution is: by introducing environmental physical quantities as a bridge, mapping data of different sources to a unified physical feature space can effectively reduce the differences between data sources and improve the consistency and comparability between different data sources; by learning the optimal mapping matrix between data sources through non-negative matrix factorization or adversarial generative network, the distribution difference of data sources in the feature space is minimized, making the feature extraction process more accurate and meaningful; by using the dynamic time warping (DTW) algorithm to process the alignment of time series data, the problem of inconsistent sampling frequency of different data sources is solved; and the interpolation strategy is dynamically adjusted according to the change of the environmental physical quantity, ensuring that the data after time alignment has good smoothness and stability; through Granger causality test, the causal relationship between different data sources can be identified, avoiding irrelevant data sources from interfering with feature extraction. At the same time, the dependence relationship between data sources is quantified using Bayesian networks or structural equation models, guiding the feature selection process and enhancing the interpretability of the model; by combining environmental physical quantities with other data sources to generate new features, and using polynomial regression or factor decomposition machine, the nonlinear relationship between data sources can be captured, thereby enriching the feature space and improving the expressiveness of the model; a multi-modal autoencoder is designed to input different data sources into a shared encoding layer, effectively learning a cross-modal joint feature representation. In addition, the introduction of the attention mechanism can dynamically allocate the weights of different data sources, highlight key features, reduce noise, and further improve the quality of feature selection; through Shapley value or LIME method, the importance of features under different environmental conditions can be dynamically evaluated, avoiding the limitations of traditional static feature selection methods. By selecting features highly related to the target variable according to the change of the environmental physical quantity, the model is more adaptable to environmental changes; the environmental physical quantity is used as a fusion weight to weight and fuse the features of different data sources, generating a more comprehensive feature vector. In addition, mutual information or conditional entropy is used to evaluate the information amount between the fused features and the target variable, ensuring that the generated fused features have strong relevance and effectiveness; by verifying the distribution consistency of the fused features under different environmental physical quantity conditions, it is ensured that the feature extraction does not destroy the physical association between data sources. By quantifying the distribution difference of the fused features through KL divergence or JS divergence, the robustness of the model under various environmental conditions is further improved; the environmental physical quantity is used as an auxiliary evaluation index of model performance, and when the environmental conditions change dramatically, the model can have higher fault tolerance and adaptability, thereby better coping with challenges in different actual environments.
[0130] An embodiment of the present application is a data fusion-based electric energy metering error correction system, comprising a memory, a processor, and a computer program stored on the memory and executable on the memory, wherein the processor executes the program to implement the data fusion-based electric energy metering error correction method as described in any of the above embodiments.
[0131] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A method for correcting the error of electric energy metering based on data fusion, characterized in that, The method comprises: S1, collecting multi-dimensional data and preprocessing the collected multi-dimensional data; S2, using a convolutional neural network to extract features from the operation data and mine spatial features in the data; using a recurrent neural network to extract time sequence features of the secondary loop and environmental data and establishing time sequence correlation between the data; S3, fusing the extracted spatial features and time sequence features; deeply mining and analyzing the multi-dimensional data; S4, training the built-in error correction model; evaluating and verifying the trained error correction model; feeding back the evaluation results to form a closed-loop feedback mechanism; The S4 comprises: S41, training the built-in error correction model according to the deep data fusion result; S42, evaluating and verifying the trained error correction model; S43, feeding back the evaluation results to form a closed-loop feedback mechanism; The S42 comprises: S421, evaluating and verifying the trained error correction model, dividing the data set into multiple subsets, taking each subset as a test set in turn, taking the remaining subsets as training sets, and calculating the average performance index of the model; S422, evaluating the performance index of the model; inputting the real-time collected multi-dimensional data into the trained dynamic self-adaptive error correction model to correct the electric energy metering data in real time; S423, outputting the corrected electric energy metering data to a monitoring system or a metering device; and evaluating and analyzing the corrected electric energy metering data to calculate the error change before and after correction.
2. The method for correcting the error of electric energy metering based on data fusion according to claim 1, characterized in that, The S1 comprises: S11, collecting multi-dimensional data at a preset time interval; S12, checking the integrity of the collected data during data collection; S13, preprocessing the collected data.
3. The method for correcting the error of electric energy metering based on data fusion according to claim 1, characterized in that, The S2 comprises: S21, slicing the operation data of the device according to a certain time window, using different size convolution kernels to perform convolution operation on the sliced data through multiple convolution layers, and extracting spatial features in the data; S22, adding a pooling layer after the convolution layer to reduce the dimension of the convolution features while retaining the feature information; S23, constructing sequence data from the secondary loop and environmental data in time sequence, and processing the sequence data through a recurrent unit; S24, encoding the features output by the recurrent unit to convert the time sequence features into a fixed-length vector representation.
4. The method for correcting the error of electric energy metering based on data fusion according to claim 3, characterized in that, The S21 comprises: Aligning the operation data of the device and the environmental data according to a unified timestamp format; and using a timestamp interpolation method to perform linear interpolation on the missing timestamp; Calculating the autocorrelation coefficient of each device operation data to determine the time dependence range of the data; and dividing multiple time windows to capture features of different time granularities; Calculating the data information entropy in each time window, introducing environmental data as an auxiliary variable, and dynamically reducing the time window when the temperature change rate exceeds a threshold; Using different size convolution kernels to extract local, mesoscopic and global spatial features; introducing a depth separable convolution in the convolution layer to decompose the standard convolution into a depth convolution and a point-by-point convolution; Channel attention mechanism is embedded in the convolutional layer to give different weights to features of different channels and strengthen the extraction of key features. The environmental data is mapped to the initialization parameters of the convolution kernel. A cross-modal feature interaction layer is added after the convolutional layer to fuse the convolutional features of the device operation data and the environmental data. A physical quantity-window size joint optimization model is constructed to determine the optimal window size by optimizing the environmental data and the window size as joint variables.
5. The method for correcting the error of electric energy metering based on data fusion according to claim 3, characterized in that, The S23 comprises: The secondary circuit data and the environmental data are aligned according to the time stamp. The data of different dimensions are normalized to map them to a unified interval. The Z-score standardization method is used to eliminate the skewness of the data distribution. The time series is divided into fixed-length windows. On the basis of the fixed window, a sliding window mechanism is introduced to dynamically adjust the window step. The dependence relationship of different time granularities is captured. The autocorrelation coefficient decay rate of the time series is calculated. When the autocorrelation decays to a threshold value, the effective length of the current sequence is determined, and the window size is dynamically adjusted. A hybrid structure of long short-term memory network (LSTM) and GRU is used to process short-term and long-term dependence relationships respectively. The self-attention mechanism is introduced in the recurrent unit to give higher weights to key time points in the time series. The environmental data is input into the recurrent unit as external input together with the secondary circuit data. A physical quantity-time series coupling model is constructed to map the temperature-related physical quantities and current and voltage data to a unified feature space. During the training process, the gradient change of the recurrent unit is dynamically monitored. When the gradient vanishes or explodes, the excessively long sequence is automatically truncated. Based on the reinforcement learning method, the sequence length is dynamically adjusted according to the model performance.
6. The method for correcting the error of electric energy metering based on data fusion according to claim 1, characterized in that, The S3 comprises: S31, the spatial and temporal features are spliced to form a comprehensive feature vector; S32, based on the feature splicing, the attention mechanism is introduced to weight the importance of different features and highlight the key features; S33, according to the characteristics of the data and the task requirements, the number of network layers and neurons is designed; S34, the fused data is used to train the deep learning model; S35, during the training process, the model is evaluated regularly, and the structure and parameters of the model are optimized according to the evaluation results.
7. The method for correcting the error of electric energy metering based on data fusion according to claim 6, characterized in that, The S31 comprises: S311, the obtained spatial data and time series data are preprocessed to obtain a spatial feature set and a time series feature set; S312, the spatial feature set is dimensionally expanded to obtain an enhanced vector of the spatial feature, and the time series feature set is further extracted to obtain an enhanced vector of the time series feature; S313, based on different splicing strategies, the enhanced vector of the spatial feature and the enhanced vector of the time series feature are spliced respectively to form a preliminary comprehensive feature vector; S314, the preliminary comprehensive feature vector is normalized to obtain a final comprehensive feature vector. Further feature analysis and processing are performed on the final comprehensive feature vector to obtain a final multi-modal feature set. S315, based on the final multi-modal feature set, modeling and inferring the spatio-temporal dependency relationship to obtain a spatio-temporal dependency graph; based on the spatio-temporal dependency graph, relationship inference and optimization are performed to obtain a final relationship inference score.
8. The method for correcting the error of electric energy metering based on data fusion according to claim 7, characterized in that, The S313 comprises: The time information is represented clockwise by time dependency, and the timestamp is represented counterclockwise; the clockwise represented coordinate information and the counterclockwise represented timestamp are first spliced based on the relationship between the position and the time to obtain a first layer splicing; The spatial distance is divided into balanced X small segments by the reference point, where X>2, the odd-numbered segments are represented in the positive direction, and the even-numbered segments are represented in the reverse direction; the spatial distances represented in different directions are spliced to obtain a first spatial distance; the time interval is divided into N equal parts by the gradual transition function, where N>2; the odd-numbered parts and the even-numbered parts are spliced respectively, and then the spliced results are spliced again to obtain a first time interval; the first spatial distance and the first time interval are second spliced based on the relative relationship between the positions and the time change to obtain a second layer splicing; The direction feature and the change rate are third spliced based on the direction and the change speed of the object motion to obtain a third layer splicing; The first 30% dimensions of the first layer splicing are extracted and spliced with the second layer splicing; the last 70% dimensions of the first layer splicing are extracted and spliced with the third layer splicing; the two spliced parts are spliced again to form a preliminary comprehensive feature vector.
9. A system for correction of electric energy metering errors based on data fusion, characterized by The computer program stored on the memory and executable on the memory, and the processor executes the program to realize the data fusion based electric energy metering error correction method according to any one of claims 1-8.
Citation Information
Patent Citations
Intelligent temperature and humidity control system for power distribution cabinet
CN119336105A
Electrical equipment life prediction system and method based on deep learning
CN119830737A