A metering data processing method based on dynamic partition rebalancing and stream-batch collaboration

Through the metering data processing method that coordinates dynamic partition rebalancing and flow batches, the problems of unbalanced partitioning and low processing efficiency in metering data processing are solved, and efficient, real-time and accurate processing of power metering data is achieved, and the adaptability and stability of the system are improved.

CN119939362BActive Publication Date: 2025-07-04NANJING YISHUNHONG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510424426.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-04
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

When facing high-dimensional, high-time variability and multi-source heterogeneous data, existing metrological data processing systems have problems such as unbalanced partitioning, low processing efficiency, unstable service quality, and unscientific rebalancing strategies, which are difficult to meet the real-time and accuracy requirements of the power system.

Method used

The method of dynamic partition rebalancing and flow batch collaboration is adopted. By collecting multi-source heterogeneous data and adding quality marks, adaptive partition feature extraction and low-overhead partition rebalancing based on time-varying features are performed, combining flow batch collaborative processing and multi-dimensional analysis to achieve efficient processing of metrological data.

Benefits of technology

It improves the accuracy and adaptability of feature extraction, achieves a deep understanding and prediction of data behavior, reduces system interference, ensures service continuity, and improves system throughput and processing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939362B_ABST
    Figure CN119939362B_ABST
Patent Text Reader

Abstract

The present invention discloses a metering data processing method based on dynamic partition rebalancing and stream-batch collaboration, including: collecting multi-source heterogeneous metering data and adding quality marks; performing adaptive partition feature extraction and low-overhead partition rebalancing based on time-varying features to generate an optimized partition scheme; performing stream-batch collaborative processing on real-time data streams and historical batch data; and performing multi-dimensional analysis and application on the processing results. The present invention solves problems such as partition imbalance, low processing efficiency, and unstable service quality in metering data processing, and improves system throughput and processing accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent metering and data management, and in particular, to a metering data processing method based on dynamic partition rebalancing and stream-batch collaboration. Background Art

[0002] With the continuous deepening of the construction of the smart grid, metering data has become the core resource for the operation, management, and decision-making of the power system. Power metering data is characterized by high dimensionality, high time-variability, and strong correlation. The amount of data generated daily increases exponentially, posing a huge challenge to traditional data processing systems. Accurately and efficiently processing these massive metering data is of great significance for realizing the intelligent operation of the power grid, improving energy utilization efficiency, supporting refined management and service decision-making. The quality of metering data processing directly affects the operation efficiency and service quality of power enterprises, affects the fairness and transparency of the power market, and is also related to the safe and stable operation of the power grid and the realization of energy conservation and emission reduction goals. Therefore, constructing a processing method that can cope with the variability and complexity of metering data has far-reaching strategic significance for promoting the digital transformation and intelligent upgrading of the power industry.

[0003] Currently, the technical route in the field of metering data processing mainly adopts static partitioning and a single processing mode. In terms of data partitioning, mainstream technologies adopt static partitioning strategies based on fixed hash functions or predefined ranges, such as consistent hash partitioning, range partitioning, etc. These methods rarely adjust the partitioning boundaries after the initial design stage and lack the ability to adapt to dynamic data changes. In terms of data processing modes, existing systems mostly adopt a single mode of batch processing or stream processing. Batch processing systems such as Hadoop and Spark are good at processing large-scale historical data but have poor real-time performance. Stream processing systems such as Storm and Flink can process real-time data streams but have limited analysis depth. In addition, existing solutions are relatively weak in data quality awareness, often adopting a unified processing strategy and ignoring data quality differences, resulting in low-quality data consuming too many resources or high-quality data not being fully utilized. In terms of rebalancing strategies, traditional methods mostly adopt simple strategies of global data reallocation or triggered by preset thresholds. The rebalancing process causes great interference to the system and affects service continuity.

[0004] Despite certain progress in the existing technologies, there are still technical bottlenecks in several key aspects. First, in terms of time-varying feature extraction, existing methods are mostly based on static features or simple time-window statistical features, which cannot effectively capture the complex time-varying patterns and long-term evolution laws in metering data, resulting in a mismatch between the partition features and the actual characteristics of the data and poor partitioning effects. Second, in terms of phase space reconstruction and dynamic evolution pattern recognition, traditional analysis methods are limited to surface statistical feature analysis, making it difficult to reveal the internal dynamic structure and complex behavior mechanisms of the data and lacking the ability to predict future evolution trends. Third, in terms of partition imbalance evaluation and rebalancing decision-making, existing technologies mostly rely on single-index evaluation and empirical threshold decision-making, making it difficult to comprehensively and accurately evaluate the imbalance state of the system and resulting in insufficient scientificity and reliability of the decision-making. These technical bottlenecks severely restrict the adaptability of the metering data processing system to data characteristic changes and the resource utilization efficiency, leading to unstable system performance, uneven quality, and high operation and maintenance costs. Summary of the Invention

[0005] The object of the invention is to provide a metering data processing method based on dynamic partition rebalancing and stream-batch collaboration, in order to solve at least one technical problem existing in the existing technologies.

[0006] Technical solution: A metering data processing method based on dynamic partition rebalancing and stream-batch collaboration includes the following steps:

[0007] Collect metering data from multi-source heterogeneous data sources, preprocess the metering data to form a quality-labeled data set containing quality level labels; where the metering data includes structured metering records generated by smart electricity meters, water meters, and gas meters, as well as system logs and device status data;

[0008] Based on the quality-labeled data set, the dynamic partition management module performs adaptive partition feature extraction based on time-varying features and low-overhead partition rebalancing to generate an updated partition scheme;

[0009] Based on the updated partition scheme, perform stream-batch collaborative processing on the real-time data stream and pre-stored historical batch data in the quality-labeled data set, and generate a final processing result through multi-level data fusion;

[0010] Perform multi-dimensional analysis and application on the final processing result to achieve the storage, visual presentation, and anomaly warning of metering data.

[0011] Beneficial effects: The present invention improves the accuracy and adaptability of feature extraction, realizes in-depth understanding and prediction of data behavior; achieves a scientific and accurate decision-making mechanism through a multi-dimensional imbalance evaluation model combined with a rebalancing revenue-cost modeling; reduces system interference and ensures service continuity through a minimum interference rebalancing path planning and an incremental rebalancing execution mechanism; improves prediction accuracy and decision-making robustness by adopting a structural causal model and uncertainty-aware revenue prediction. Description of the Drawings

[0012] Figure 1 It is a flowchart of the steps of a metering data processing method based on dynamic partition rebalancing and stream-batch collaboration provided by an embodiment of the present application.

[0013] Figure 2 It is a flowchart of the steps of forming a quality marked data set including quality level marks provided by an embodiment of the present application.

[0014] Figure 3 It is a flowchart of the steps of performing adaptive partition feature extraction based on time-varying features and low-overhead partition rebalancing provided by an embodiment of the present application.

[0015] Figure 4 It is a flowchart of the steps of performing stream-batch collaborative processing provided by an embodiment of the present application.

[0016] Figure 5 It is a flowchart of the steps of performing multi-dimensional analysis and application provided by an embodiment of the present application. Detailed Embodiments

[0017] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] It should be specifically noted that, for clearly showing the step flow of the present application, serial numbers are marked for each step in the specification. These serial numbers are only for the convenience of description and do not limit the execution order of the steps. In actual operation, according to the technical requirements of the specific implementation scenario, the steps can be executed in an order different from that shown in the specification, and in some cases, parallel processing between steps can also be achieved.

[0019] As Figure 1 shown, a metering data processing method based on dynamic partition rebalancing and stream-batch collaboration includes the following steps:

[0020] S1. Collect metering data from multiple heterogeneous data sources, pre-process the metering data, and form a quality-labeled data set containing quality grade labels; the metering data includes structured metering records generated by smart electricity meters, water meters, and gas meters, as well as system logs and device status data;

[0021] S2. Based on the quality-labeled dataset, the dynamic partition management module performs adaptive partition feature extraction based on time-varying features and low-overhead partition rebalancing to generate an updated partition scheme;

[0022] S3, based on the updated partitioning scheme, performs stream-batch collaborative processing on the real-time data stream and pre-stored historical batch data in the quality labeling dataset, and generates the final processing result through multi-level data fusion;

[0023] S4. Conduct multi-dimensional analysis and application of the final processing results to achieve storage, visualization and abnormal warning of measurement data.

[0024] This embodiment realizes efficient processing of the entire life cycle of metering data by integrating dynamic partition rebalancing and stream-batch collaborative metering data processing methods. First, metering data is collected from multi-source heterogeneous data sources and preprocessed, and quality grade tags are added so that the system can distinguish data of different quality levels and process them in a targeted manner, thereby improving the reliability of subsequent analysis. Secondly, based on the quality tag data set, adaptive partitioning and low-overhead partition rebalancing based on time-varying features are performed. The system can dynamically adjust the partitioning scheme according to the time-varying features of the data, adapt to the dynamic change characteristics of the metering data, avoid the data tilt and resource utilization imbalance caused by traditional static partitioning, and solve the problem that the existing technology adopts a one-time large-scale migration strategy, the system overhead is large, the service quality is significantly reduced, and the minimum interference path planning and incremental execution mechanism are lacking. Third, stream-batch collaborative processing is performed on real-time data streams and historical batch data, and the final processing results are generated through multi-level data fusion, which solves the problem of separation between stream processing and batch processing in traditional methods and ensures the balance between real-time and accuracy of metering data processing. Finally, the processing results are analyzed and applied in multiple dimensions to realize the all-round value mining of metering data. Through closed-loop optimization design, this embodiment cooperates with each other to form a complete set of metering data processing solutions, which can effectively cope with the challenges of large-scale, variable and time-sensitive data processing in the field of power metering, improve system throughput, reduce processing delays, and ensure data processing quality.

[0025] like Figure 2 As shown, according to one aspect of the present application, step S1 is further:

[0026] S11. According to the preset data collection rules, access multiple heterogeneous data sources and collect metering data to generate original metering data;

[0027] S12. Perform protocol conversion and parsing on the original measurement data, as well as data cleaning and standardization processing to obtain normalized measurement data;

[0028] S13. Calculate the quality indicators of the normalized measurement data, and add quality level marks to each piece of normalized measurement data based on the quality indicators to form a quality-marked data set.

[0029] In an embodiment of the present application, the system collects original measurement data from multiple data sources, including structured measurement records generated by IoT devices such as smart meters, water meters, and gas meters, as well as semi-structured data such as system logs and device statuses. A multi-protocol adapter is used to process data streams of different communication protocols (such as MQTT, CoAP, HTTP), and uniformly convert them into standard data packets. The standard data packets are initially parsed to extract metadata information (such as device ID, timestamp, data type) and payload data to form an initial data set.

[0030] Perform null value processing on the initial data set, and generate a null-free data set using different strategies (such as time series interpolation, statistical value filling) according to the data type. Identify and remove outliers and duplicate values in the null-free data set, and output the cleaned data set. Convert the cleaned data set into a unified data model and measurement unit to generate a standardized data set.

[0031] Perform integrity check on the standardized data set, and calculate the data integrity rate indicator. Perform accuracy check on the standardized data set, verify through physical model constraints, and calculate the data accuracy rate indicator. Perform timeliness evaluation on the standardized data set, and calculate the data timeliness indicator. Combine the data integrity rate indicator, data accuracy rate indicator, and data timeliness indicator to generate a data quality score, and mark the quality level of the standardized data set according to the score to output a quality-marked data set.

[0032] This embodiment solves the problems of diverse measurement data sources, inconsistent standards, and unstable quality by constructing a multi-source heterogeneous data collection and quality marking system. It enables the system to distinguish high-quality data from low-quality data, and adopt differential processing strategies for different quality levels, which not only ensures the priority processing of high-quality data but also does not completely discard the useful information that may be contained in low-quality data, ultimately improving the overall data processing efficiency and result reliability. This embodiment reduces the error rate of subsequent analysis, improves the utilization value of measurement data, and lays a solid foundation for the accurate measurement and analysis of the power measurement system.

[0033] As Figure 3 shown, according to one aspect of the present application, step S2 is further:

[0034] S21. Based on the quality - marked data set, perform adaptive partition feature extraction of time - varying features to obtain partition features;

[0035] S22. Apply the partition features to perform an initial partition of the data, generate an initial partition scheme, and obtain the system partition state;

[0036] S23. Combine the pre - stored system operation status and load statistics data, perform low - overhead partition rebalancing, update the initial partition scheme, and generate an updated partition scheme.

[0037] In this embodiment, by implementing adaptive partition feature extraction and low - overhead partition rebalancing, the problems of data skew, low processing efficiency, and unbalanced resource utilization caused by static partitioning in traditional metering data processing are solved. The high system overhead of traditional rebalancing is minimized. Through a refined cost - benefit analysis, while ensuring an improvement in partition balance, the impact on system performance is minimized, achieving "low interference, high benefit" partition adjustment. It solves the problem that the prior art, based on simple correlation analysis rather than causal inference, inaccurately predicts the actual impact of rebalancing operations, resulting in a high rate of rebalancing decision - making errors. Overall, this embodiment achieves a balance between the dynamic adaptability of data partitioning and system stability. It can not only respond promptly to changes in data characteristics but also maintain the stable operation of the system, improving the throughput and response time of metering data processing, and is particularly suitable for scenarios such as electric power metering data with obvious time - varying characteristics.

[0038] According to one aspect of the present application, step S21 is further as follows:

[0039] S211. Input the quality - marked data set into the multi - scale time - varying feature decomposition unit to extract the dominant feature components;

[0040] S212. Based on the dominant feature components, perform feature sensitivity evaluation and screening to generate a refined feature set;

[0041] S213. Use the refined feature set to construct a dynamic evolution pattern recognizer and output a pattern prediction result;

[0042] S214. According to the pattern prediction result, perform adaptive feature fusion and weight assignment to generate weighted fusion features, that is, partition features;

[0043] S215. Dynamically optimize the feature extraction strategy based on the weighted fusion features and output a feature extraction parameter set.

[0044] In one embodiment of the present application, a received quality marked data set is used. Wavelet transform is applied to decompose the time series into components of different scales, obtaining multi-scale time series components. Trend extraction, periodic analysis, and noise assessment are respectively performed on the multi-scale time series components to form a time series feature component set. The energy proportion of each component in the time series feature component set is calculated to identify the dominant component, and the dominant feature component is output.

[0045] An association model between features and partitioning effects is constructed. The dominant feature component and historical partitioning effect data are input, and a feature-effect sensitivity matrix is output. Principal component analysis is applied to reduce the dimension of the feature-effect sensitivity matrix to identify key feature combinations and generate a key feature set. Redundancy analysis is performed on the key feature set to remove highly correlated features, and a refined feature set is output. A dynamic evolution pattern recognizer is constructed using the refined feature set, and a pattern prediction result is output.

[0046] The current data pattern recognized by the pattern predictor is received, and a suitable feature weight configuration is extracted from the pattern knowledge graph, and an initial weight configuration is output. Based on the current state of the system and historical feature effectiveness evaluation, the initial weight configuration is dynamically adjusted, and an optimized weight configuration is output. The optimized weight configuration is applied to weight-fuse the features in the refined feature set to generate weighted fusion features.

[0047] Partitioning effect feedback data is collected, and a feature extraction effect evaluation model is constructed in combination with the weighted fusion features, and a feature effect score is output. Based on the feature effect score, Bayesian optimization is applied to adjust the feature extraction parameters, and a feature extraction parameter set is output. The feature extraction parameter set is applied to the next round of feature extraction process to form a closed-loop optimization.

[0048] This embodiment solves the problem of low partitioning efficiency caused by the inability to cope with the time-varying characteristics of data in traditional metering data partitioning by implementing adaptive partitioning feature extraction of time-varying features. The feature extraction process can continuously improve itself, evolve with the change of data characteristics, and maintain long-term effectiveness. Overall, this embodiment realizes the intelligence and self-adaptability of partitioning feature extraction, improves the accuracy and dynamic adaptability of metering data partitioning, and provides an efficient and stable partitioning basis for complex and variable power metering data.

[0049] According to one aspect of the present application, step S213 is further as follows:

[0050] S2131. Receive the refined feature set, determine the optimal time delay parameter based on the principle of minimum mutual information, and determine the optimal embedding dimension in combination with the improved false nearest neighbor algorithm to reconstruct the one-dimensional time series into multi-dimensional phase space trajectory data; Kernel density estimation is used to generate a phase space density map;

[0051] S2132. Calculate the Lyapunov exponent, correlation dimension, and entropy rate of the phase space trajectory data based on the phase space density map, combine the sparse identification method and deep learning to identify the dynamic equation, and obtain the dynamic model parameters through L1 regularization optimization;

[0052] S2133. Construct a recurrence plot based on the phase space trajectory data and dynamic model parameters, analyze the network topology characteristics, combine the Wasserstein distance to calculate the distribution change rate, and identify the multi-scale transition points;

[0053] S2134. Segment the time series according to the multi-scale transition points, extract the pattern features to construct a knowledge graph, train a graph neural network to learn the pattern conversion rules, output a dynamic evolution pattern recognizer, and generate a pattern prediction result.

[0054] In an embodiment of the present application, the construction and application of the evolution pattern knowledge graph are as follows: Based on the multi-scale transition point set, the time series is segmented into multiple intervals, each interval corresponding to an evolution pattern, and a pattern division interval set is output. Extract the feature representation of each pattern from the pattern division interval set to form a pattern feature library. Construct an evolution pattern knowledge graph: Nodes: represent the identified typical evolution patterns, and the attributes include the features from the pattern feature library; Edges: represent the conversion relationships and probabilities between patterns, and output the pattern knowledge graph. Construct and train a graph neural network, input the pattern knowledge graph, learn the similarities and conversion rules between patterns, and output a pattern relationship model. Combine the pattern relationship model and the causal reasoning mechanism to achieve the inference ability for unseen patterns, and output a pattern predictor.

[0055] This embodiment solves the problem of insufficient grasp of the evolution law of complex time series data in traditional measurement data analysis by realizing dynamic evolution pattern recognition, and provides the ability to predict the future change trend of data. Connect discrete data patterns into an organic knowledge network, and realize the inference ability for unseen patterns. Overall, this embodiment provides unprecedented prediction depth for measurement data analysis, enabling the system to "predict the future", make resource preparations and strategy adjustments in advance, and improve the foresight and adaptability of the power measurement system.

[0056] According to one aspect of the present application, step S2131 is further as follows:

[0057] S21311. Receive the reduced feature set, and use the adaptive time delay parameter calculation method (based on the principle of minimum mutual information) to determine the optimal time delay parameter;

[0058] S21312. Based on the optimal time delay parameter and the improved false nearest neighbor algorithm, determine the optimal embedding dimension;

[0059] S21313. Reconstruct the one-dimensional time series into multi-dimensional phase space trajectory data by using the optimal time delay parameter and the optimal embedding dimension;

[0060] S21314. Calculate the probability density distribution of the phase space trajectory data by using the kernel density estimation method, and output the phase space density map.

[0061] Reconstruct the one-dimensional time series X(t) into an m-dimensional phase space by using the optimal time delay parameter and the optimal embedding dimension: X(t) → [X(t), X(t+τ), X(t+2τ),..., X(t+(m-1)τ)], and generate the phase space trajectory data, where τ is the optimal time delay parameter and m is the optimal embedding dimension.

[0062] According to one aspect of the present application, step S2132 is further:

[0063] Calculate the Lyapunov exponent of the phase space trajectory data (used to quantify the chaos degree of the system), and output the Lyapunov exponent set.

[0064] Calculate the correlation dimension of the phase space trajectory data (used to quantify the trajectory complexity), and output the correlation dimension value.

[0065] Calculate the entropy rate of the phase space trajectory data (used to quantify the information generation rate), and output the entropy rate value.

[0066] Combine the sparse identification method (SINDy) and deep learning to identify the implicit dynamic equation from the phase space trajectory data: dX / dt = F(X) = Ξ(X)·Θ, where X is the system state vector, F(X) is the dynamic function, Ξ(X) is the candidate function library, and Θ is the coefficient vector.

[0067] Solve Θ by L1 regularization optimization to obtain the sparse representation of the dynamic equation, and output the dynamic model parameters. Perform multi-modal transition point detection and classification based on the dynamic model parameters.

[0068] In this embodiment, through the phase space reconstruction and topological invariant extraction technology, the problem that the traditional time series analysis method is difficult to reveal the internal dynamic structure and complex behavior mechanism of the data is solved, and the in-depth mining of the essential characteristics of the measurement data is realized. The internal mechanism of the measurement data is revealed and mathematically expressed, the traditional black box analysis is transformed into a transparent model analysis, the understanding depth and prediction accuracy of the behavior of the power measurement system are improved, and a theoretical basis is provided for system optimization and anomaly diagnosis.

[0069] According to one aspect of the present application, step S2133 is further:

[0070] S21331. Construct a recurrence plot (RP) based on phase space trajectory data: RP(i, j) = Θ(ε - ||X(i) - X(j)||); where Θ is the Heaviside function, ε is the distance threshold, and X(i) is the state vector of the phase space trajectory data at time point i, and output the recurrence plot matrix.

[0071] S21332. Analyze the network topology characteristics of the recurrence plot matrix, calculate indicators such as node centrality and clustering coefficient, and form a topology characteristic vector.

[0072] S21333. Calculate the change rate of the probability distribution within continuous time windows based on the Wasserstein distance, identify distribution mutation points, and output a set of distribution mutation points.

[0073] S21334. Combine the change trend of the topology characteristic vector and the set of distribution mutation points, apply a hierarchical attention mechanism to identify transition points at different scales, and output a set of multi-scale transition points.

[0074] This embodiment solves the problems in traditional metrological data analysis that it is difficult to accurately identify system state change points and predict transition trends through multi-modal transition point detection and classification technology. This embodiment can identify transition points at different time scales, solving the problem that single-scale analysis cannot take into account both macroscopic trends and microscopic changes. Overall, this embodiment realizes the accurate identification and classification of state transitions in metrological data, improving the transition point detection accuracy rate from 70% of traditional methods to over 90%, providing key technical support for the predictive maintenance and anomaly warning of power metering systems.

[0075] According to one aspect of the present application, step S22 is further as follows:

[0076] S221. Design a data partitioning strategy based on weighted fusion features and the current system load condition, determine the number of partitions, partition keys, and partition boundaries, and output a partition strategy plan.

[0077] S222. Evaluate the difference between the partition strategy plan and the current partition state, calculate the migration cost, and output the partition adjustment cost.

[0078] S223. Decide whether to perform partition adjustment according to the partition adjustment cost and system tolerance, and output a partition execution decision.

[0079] S224. If the partition execution decision is to execute, apply the new partition strategy and update the system partition state.

[0080] In this embodiment, by designing a reasonable data partitioning strategy to determine the number of partitions, partition keys, and partition boundaries, data storage and querying are made more efficient, reducing waste of system resources. By evaluating the differences between the partitioning strategy plan and the current partition state and calculating the migration cost, when performing partition adjustment, the overhead and system downtime caused by data migration can be minimized. Based on the partition adjustment cost and system tolerance, it is decided whether to execute the partition adjustment to ensure that the system can flexibly adapt to changes and maintain efficient operation under different load conditions. By applying the new partitioning strategy and updating the system partition state, the performance and response speed of the system can be improved to meet the growth of business requirements. By dynamically adjusting the partitioning strategy, the system can be adjusted according to the actual situation, providing more accurate decision-making support and optimizing resource allocation. This embodiment can improve data processing efficiency, reduce migration costs, enhance system flexibility, optimize system performance, and thus improve the overall decision-making support ability through intelligent data partitioning strategy design and adjustment.

[0081] According to one aspect of the present application, step S23 is further as follows:

[0082] S231. Collect load statistical data of each partition, including data volume, access frequency, and computing resource occupancy; calculate the imbalance index based on the load statistical data, analyze the data access pattern and identify hot partitions, construct a multi-dimensional imbalance evaluation model, and output a comprehensive imbalance index;

[0083] S232. Construct a rebalancing benefit model based on the comprehensive imbalance index, construct an environment-aware weight adjustment mechanism in combination with the system partition state and system operation state, predict the rebalancing cost and perform uncertainty modeling, and optimize the rebalancing strategy;

[0084] S233. Construct a data dependency graph according to the optimized rebalancing strategy, calculate the shard association graph and construct a migration priority algorithm, and generate a migration path plan with minimal interference;

[0085] S234. Execute incremental partition data migration according to the migration path plan, monitor the execution status in real time and dynamically adjust the migration parameters;

[0086] S235. Based on the adjusted migration parameters, evaluate the improvement degree of the comprehensive imbalance index before and after rebalancing, calculate the actual resource consumption cost, update the parameters of the rebalancing decision model and generate an updated partition plan.

[0087] In one embodiment of the present application, load statistical data of each partition is collected, including indicators such as data volume, access frequency, and computing resource occupancy. Calculate the imbalance index of data distribution: Imbalance = (σ / μ) × 100%; where σ is the standard deviation of the load of each partition, and μ is the average value of the load of each partition, and output the load imbalance degree. Analyze the data access pattern, identify the hot partitions, calculate the access proportion of the hot partitions, and output the hot concentration degree. Integrate the load imbalance degree and the hot concentration degree to construct a multi-dimensional imbalance evaluation model and output a comprehensive imbalance index.

[0088] Perform rebalancing effect evaluation and feedback, specifically: compare the comprehensive imbalance indicators before and after rebalancing, calculate the improvement degree, and output the balance improvement degree. Monitor the resource consumption and system performance impact during the rebalancing process, calculate the actual cost, and output the actual rebalancing cost. Compare the actual rebalancing cost with the predicted cost estimate C, evaluate the accuracy of the cost model, and output the cost model error. Combine the balance improvement degree and the actual rebalancing cost to calculate the actual benefit of the rebalancing operation, and output the actual benefit value. Based on indicators such as the actual benefit value and the cost model error, update the parameters of the rebalancing decision model, and output the model update parameter set for guiding the next rebalancing decision.

[0089] This embodiment solves the problems of high system overhead and service interruption in the traditional rebalancing process by implementing low-overhead partition rebalancing, and realizes "low interference, high benefit" partition adjustment. It enables the system to continuously accumulate experience, optimize the rebalancing decision, and improve the long-term operation efficiency. Generally speaking, this embodiment realizes high efficiency and low impact of partition adjustment, and is particularly suitable for critical business scenarios with high availability requirements such as power metering.

[0090] According to one aspect of the present application, step S232 is further:

[0091] S2321. Construct a multi-objective rebalancing benefit function based on the comprehensive imbalance index, and output the equilibrium benefit value and the dynamic weight vector;

[0092] S2322. Combine the pre-stored current system load situation, and predict the rebalancing cost based on causal inference to obtain the cost estimate and the state change prediction;

[0093] S2323. Based on the equilibrium benefit value and the cost estimate, perform uncertainty-aware benefit prediction, and output the risk-adjusted benefit and the quantum optimization decision boundary;

[0094] S2324. Based on the risk-adjusted benefit, the quantum optimization decision boundary and the state change prediction, optimize the rebalancing strategy through reinforcement learning, and output the optimal rebalancing strategy.

[0095] In this embodiment, a multi-objective rebalancing benefit function is constructed through comprehensive imbalance indicators, and the system weights are dynamically adjusted so that the system can maintain a high degree of balance when the load and priority change. Combining the current system load situation, causal inference is used to predict the rebalancing cost, effectively evaluate and reduce the cost brought by the rebalancing operation, and improve the overall benefit of the system. Based on the equilibrium benefit value and the cost prediction value, an uncertainty-aware benefit prediction is performed, and the risk-adjusted benefit and the quantum optimization decision boundary are output, thereby improving the accuracy and reliability of the benefit prediction. Through the reinforcement learning algorithm, based on the risk-adjusted benefit, the quantum optimization decision boundary and the state change prediction, the rebalancing strategy is optimized to generate the optimal rebalancing plan, thereby improving the stability and performance of the system. This embodiment optimizes the rebalancing strategy, improves the system balance degree and benefit, reduces the rebalancing cost, improves the accuracy of the benefit prediction, and finally realizes the stable and efficient operation of the system.

[0096] According to one aspect of the present application, step S2321 is further:

[0097] S23211. Construct an equilibrium benefit model after rebalancing based on comprehensive imbalance indicators:

[0098] B = f(I_current, I_expected);

[0099] Where I_current is the current comprehensive imbalance indicator, I_expected is the expected imbalance indicator, f() is a function, and the equilibrium benefit value B is output;

[0100] S23212. Combine the information of the current state (peak / low period) of the system to construct an environment-aware adaptive weight adjustment mechanism:

[0101] W(t) = W_base + ΔW(load(t), priority(t));

[0102] Where W(t) is the weight vector at time t, W_base is the base weight, load(t) is the system load at time t, priority(t) is the service priority, ΔW is the weight change amount, and the dynamic weight vector W is output;

[0103] S23213. Construct a time series prediction model to predict the system load trend in the future period to obtain the load prediction value;

[0104] S23214. Adjust the dynamic weight vector W based on the load prediction value to obtain the pre-adjusted weight vector;

[0105] S23215. Combine the equilibrium benefit value B with the pre-adjusted weight vector to construct a complete multi-objective rebalancing benefit function:

[0106] E = α•B - β•C - γ•D;

[0107] Where α, β, and γ are weight coefficients from a pre-adjusted weight vector, B is the equilibrium revenue value, C is the rebalancing cost, and D is the business interference degree.

[0108] In this embodiment, by constructing an equilibrium degree revenue model based on a comprehensive imbalance index and dynamically adjusting the system weights, the system can maintain a high degree of balance under different loads and priorities. Through the comprehensive application of technical means such as imbalance indicators, system status, load prediction, and reinforcement learning, an intelligent and efficient rebalancing strategy optimization is achieved, which helps to improve the balance, flexibility, and overall efficiency of the system.

[0109] According to one aspect of the present application, step S2322 is further as follows:

[0110] S23221. Establish a fine-grained cost model for rebalancing operations, including: C_data: data migration bandwidth cost; C_index: index reconstruction calculation cost; C_query: query performance impact cost; C_io: I / O load increase cost; The total cost C = C_data + C_index + C_query + C_io is obtained comprehensively, and the cost estimate C is output.

[0111] S23222. Construct a structural causal model (SCM) of system state variables: S = f(X, do(R)); where S is the system state vector, X is the influence factor vector, and do(R) represents the intervention of performing the rebalancing operation R, and the causal structure diagram is output.

[0112] S23223. Use the causal structure diagram to calculate the system state difference between performing rebalancing and not performing rebalancing: ΔS = E[S|do(R = 1)] - E[S|do(R = 0)], and output the state change prediction ΔS, where E represents the expectation function.

[0113] S23224. Based on the state change prediction ΔS, accurately quantify the costs of each component of the system brought by rebalancing, and update the cost estimate C. Perform uncertainty-aware revenue prediction and rebalancing strategy optimization based on reinforcement learning.

[0114] In this embodiment, through the rebalancing cost prediction technology based on causal inference, the problem of prediction deviation caused by traditional cost models only considering surface correlation and ignoring deep causal relationships is solved, and the accuracy and interpretability of cost prediction are improved. The prediction model can learn and improve from experience and maintain prediction accuracy in the long term. Overall, this embodiment realizes the accurate prediction of the intervention effect of complex systems, improves the cost prediction accuracy rate from 75% of traditional methods to more than 90%, and provides precise cost control capabilities for the economic and efficient operation of the power metering system.

[0115] According to one aspect of the present application, step S2323 is further as follows:

[0116] S23231. Establish a Bayesian network to represent the probability dependence relationship between system states and rebalancing benefits, input the current system partition state and historical data, and output a Bayesian network model.

[0117] S23232. Use the Bayesian network model to design a Monte Carlo simulation method to generate multiple possible rebalancing result scenarios and output benefit distribution data.

[0118] S23233. Based on the benefit distribution data, calculate the risk-adjusted expected benefit: RAR = E[B] - λ × σ(B); where E[B] is the expected benefit, σ(B) is the standard deviation of the benefit, and λ is the risk aversion coefficient, and output the risk-adjusted benefit RAR.

[0119] S23234. Introduce the quantum probability theory framework to handle the complementary uncertainties that are difficult to express by traditional probability models and construct a quantum probability representation.

[0120] S23235. Based on the quantum probability representation, design a quantum information geometric optimization method to determine the optimal decision boundary and output the quantum optimization decision boundary.

[0121] In this embodiment, through the benefit prediction technology with uncertainty perception, the problem that the future benefit prediction in traditional rebalancing decisions is too deterministic and ignores risks and uncertainties is solved, and the robustness and reliability of the decision-making are improved. The comprehensive assessment and control of the rebalancing decision-making risk are realized, and the decision-making error rate is reduced from 15% of traditional methods to less than 5%, providing decision-making guarantees for the stable operation of the power metering system.

[0122] According to one aspect of the present application, step S2324 is further as follows:

[0123] S23241. Construct a reinforcement learning environment for rebalancing decisions: State space S: It includes indicators such as load imbalance degree, hotspot concentration, and system resource utilization rate; Action space A: {Execute rebalancing, Adjust rebalancing parameters, Delay rebalancing}; Reward function R: Calculate and output the RL environment model based on the multi-objective benefit function E = α·B - β·C - γ·D.

[0124] S23242. Design a multi-agent reinforcement learning architecture, configure independent decision-making agents for different data partitions, and form an agent network.

[0125] S23243. Use a graph attention network (GAT) to construct an inter-agent communication mechanism: h_i (l+1) = σ(∑_j∈N(i)α_ij W h_j l ); where h_i l is the state representation of the i-th agent at the l-th layer, α_ij is the attention weight, W is the parameter matrix, and the agent communication protocol is output.

[0126] S23244. Train the agent network, optimize the rebalancing decision-making strategy, and output the optimal rebalancing strategy.

[0127] In this embodiment, through the rebalancing strategy optimization technology based on reinforcement learning, the problems in traditional rebalancing decisions, such as relying on manual experience and being difficult to adapt to complex dynamic environments, are solved, and the automation and intelligence of decision-making are realized. The intelligence and self-adaptability of rebalancing decisions are achieved, the decision-making time is reduced from the minute level of traditional methods to the second level, and at the same time, the decision-making quality is improved, and the load balance degree is increased by more than 30%, providing intelligent decision-making support for the automated operation and maintenance of the power metering system.

[0128] According to one aspect of the present application, step S233 is further as follows:

[0129] S2331. Based on the optimal rebalancing strategy, determine the data shards that need to be migrated to form a migration task set.

[0130] S2332. Construct a data dependency graph, analyze the access relationships between shards, identify highly correlated shards, and output a shard association graph.

[0131] S2333. Based on the shard association graph, design a migration priority algorithm to ensure that shards with high correlation are migrated in the same batch as much as possible, and output a migration batch plan.

[0132] S2334. Evaluate the system impact of each migration path, select the path with the least interference, and generate a migration path plan.

[0133] Step S234 is further as follows:

[0134] S2341. Split the rebalancing task into multiple incremental steps according to the migration path plan and migration batch plan, and output the incremental execution plan.

[0135] S2342. Monitor the execution status of each incremental step in real time and collect execution status data.

[0136] S2343. Dynamically adjust the execution parameters of subsequent incremental steps, such as migration rate, parallelism, etc., based on the execution status data, and output the adjusted execution parameters.

[0137] S2344. Apply the adjusted execution parameters and continue to perform incremental rebalancing until all migration task sets are completed.

[0138] In this embodiment, through the minimum interference rebalancing path planning and incremental rebalancing execution technology, the problems of system performance degradation and service interruption caused by large-scale data migration in the traditional rebalancing process are solved, and a smooth and imperceptible rebalancing process is achieved. "Imperceptible" data rebalancing is realized, reducing the performance degradation in the traditional rebalancing process from more than 30% to less than 5%, while ensuring data consistency and service continuity, providing key technical support for the highly available operation and maintenance of the power metering system.

[0139] As Figure 4 shown, according to one aspect of the present application, step S3 is further:

[0140] S31. Construct a unified stream-batch processing model and apply the updated partition scheme to the real-time data stream in the quality-tagged dataset and the pre-stored historical batch data;

[0141] S32. Perform incremental feature update and real-time anomaly detection on the real-time data stream to generate real-time processing results;

[0142] S33. Perform time series analysis and quality verification on the historical batch data to obtain the verified batch processing results;

[0143] S34. Perform multi-level data fusion based on the real-time processing results and the verified batch processing results to generate the final processing results.

[0144] In an embodiment of the present application, receive the real-time data stream in the quality-tagged dataset, perform real-time distribution according to the system partition status, and output the partitioned data stream. Apply sliding window processing to the partitioned data stream, calculate the statistics within the window, and output the window statistics results. Process the key events in the partitioned data stream based on the event-driven model and output the event processing results. Fusion the window statistics results and the event processing results to generate the real-time processing results.

[0145] Load historical measurement data from the data warehouse, reorganize the data using the application system partition status, and output a batch processing dataset. Execute complex analysis algorithms on the batch processing dataset, such as time series analysis, pattern recognition, etc., and output the batch processing results. Perform quality verification on the batch processing results to ensure the accuracy of the results, and output the verified batch processing results.

[0146] Determine the overlapping time window of the real-time processing results and the verified batch processing results, and output the result overlapping interval. Compare the differences between the two results within the result overlapping interval, calculate the consistency index, and output the consistency score. Based on the consistency score and the predetermined rules, decide which strategy to use to fuse the results, and output the fusion strategy. Apply the fusion strategy to combine the real-time processing results and the verified batch processing results to generate the final processing results.

[0147] In this embodiment, by constructing a unified stream-batch processing model and implementing stream-batch collaborative processing, the problem of difficult to balance real-time performance and accuracy caused by the separation of stream processing and batch processing in traditional measurement data processing is solved. It can simultaneously meet the dual needs of real-time measurement monitoring and in-depth electricity consumption behavior analysis, providing more comprehensive and timely data support for power enterprises.

[0148] As Figure 5 shown, according to one aspect of the present application, step S4 is further as follows:

[0149] S41. Organize the final processing results according to a predetermined data model, write them into the persistent storage module, obtain the persistent result data, and establish a multi-dimensional index;

[0150] S42. Perform multi-dimensional statistical analysis based on the persistent result data to generate a visual resource set;

[0151] S43. Identify anomalies in the final processing results according to the preset anomaly rules and trigger corresponding warning messages.

[0152] In an embodiment of the present application, organize the final processing results according to a predetermined data model, write them into the persistent storage, and output the persistent result data. Establish a multi-dimensional index for the persistent result data to optimize the query performance, and output the index structure. Implement the data life cycle management strategy, perform hierarchical storage and archiving on the persistent result data, and output the hierarchical storage structure.

[0153] Perform multi-dimensional statistical analysis based on the persistent result data, such as time dimension, space dimension, user dimension, etc., and output the multi-dimensional analysis results. Generate various data visualization charts, such as trend charts, distribution charts, association charts, etc., and output the visual resource set. Construct an interactive data exploration interface, integrate the visual resource set, support in-depth analysis, and output the data exploration interface.

[0154] Based on the final processing results, apply anomaly detection algorithms to identify anomaly patterns and output an anomaly candidate set. Verify and classify the anomaly candidate set, filter false alarms, and output a confirmed anomaly set. Generate hierarchical warning information according to the severity of the confirmed anomaly set and output a warning message set. Push the warning message set to relevant personnel through pre-configured notification channels and record the processing status to form a closed-loop management.

[0155] This embodiment realizes multi-dimensional analysis and application of metering data, solves the problem of insufficient data value mining in traditional metering data processing, and maximizes the commercial value and decision-making support ability of metering data. It enables the system to detect and handle problems before they escalate, transforming passive response into proactive prevention, and improving the operation reliability and security of the power metering system. Overall, this embodiment transforms metering data from simple recorded numbers into actionable decision-making bases, providing strong data support for the refined management, energy conservation and emission reduction, and optimized operation of power enterprises.

[0156] In summary, the present invention collects multi-source heterogeneous metering data and adds quality marks; performs adaptive partition feature extraction based on time-varying features and low-overhead partition rebalancing to generate an optimized partition scheme; performs stream-batch collaborative processing on real-time data streams and historical batch data; and conducts multi-dimensional analysis and application of the processing results. For the problem of time-varying feature extraction, traditional methods usually adopt simple time window statistics or basic spectrum analysis, lacking multi-scale decomposition ability. The present invention adopts multi-scale time-varying feature decomposition, decomposes the time series into components of different scales through wavelet transform, extracts dominant features, and generates a refined feature set through feature sensitivity evaluation, improving the accuracy and adaptability of feature extraction. For the problem of dynamic evolution pattern recognition, existing technologies mostly rely on simple statistical models, while the present invention constructs a complete evolution pattern knowledge graph through phase space reconstruction, topological invariant extraction, and multi-modal transition point detection, achieving in-depth understanding and prediction of data behavior. For the problem of partition imbalance evaluation and rebalancing decision-making, traditional methods usually use a single index and a fixed threshold. The present invention constructs a multi-dimensional imbalance evaluation model, combines rebalancing revenue-cost modeling and reinforcement learning-based strategy optimization, realizes a scientific and accurate decision-making mechanism, and solves the problems that existing technologies are difficult to cope with multi-source uncertainties and risks in complex environments and lack a robust decision-making mechanism. For the problem of rebalancing execution, existing technologies usually adopt a one-time migration strategy. The present invention constructs a minimum interference rebalancing path planning and incremental rebalancing execution mechanism, reducing system interference and ensuring service continuity. For the problem of cost-benefit analysis, traditional models are based on simple correlation analysis, while the present invention adopts a structural causal model and uncertainty-aware benefit prediction, improving prediction accuracy and decision-making robustness.

[0157] Taking the processing of smart meter measurement data of a provincial power company as an example, the system collects multi-source heterogeneous measurement data including smart meter records, device status data, and system logs.

[0158] First, according to the preset data collection rules, the system processes data streams with different communication protocols through a multi-protocol adapter. Smart meter data is mainly transmitted using the DLT645-2007 protocol, device status data uses the MQTT protocol, and system logs use the HTTP protocol. The system uniformly converts all data into standard JSON format data packets:

[0159] { "device_id": "E10056782", "timestamp": "2024-02-15T08:30:00Z", "data_type": "power_reading", "value": 5.63, "unit": "kWh"};

[0160] Parse the standard data packet to extract metadata information and payload data, forming an initial data set. Then perform data cleaning processing, including: handling missing values: time series data (such as meter readings) is filled using linear interpolation; outlier detection: the Z-score method is used to identify outliers, and the rule is that |Z| > 3 is marked as an outlier; duplicate value removal: duplicate records are identified based on the combination of device ID and timestamp;

[0161] After cleaning, standardize the data by uniformly converting different measurement units (such as kWh, W, etc.) into a standard unit system.

[0162] Then calculate the data quality indicators: data integrity rate indicator CI = (1 - number of missing values / total number of records) × 100%; data accuracy rate indicator AI = (1 - number of outliers / total number of records) × 100%; data timeliness indicator TI = (1 - average delay time / maximum acceptable delay) × 100%; comprehensive quality score QS = 0.4×CI + 0.4×AI + 0.2×TI;

[0163] According to the comprehensive quality score QS, mark the quality level for each piece of data: QS ≥ 90%: High quality; 70% ≤ QS < 90%: Medium quality; QS < 70%: Low quality. Finally, form a quality-marked data set with quality level markings to provide data quality awareness for subsequent processing.

[0164] By performing adaptive extraction of time-varying features on the quality-labeled dataset to capture the characteristics of data changing over time. Specifically as follows: Input the quality-labeled dataset into the multi-scale time-varying feature decomposition unit. Taking the 24-hour power consumption time series X(t) of a single user as an example, the system uses discrete wavelet transform (DWT) for multi-scale decomposition: X(t) = A_J(t) + ∑ j=1 J D_j(t);

[0165] Where: X(t) is the original time series; A_J(t) is the J-level approximation component, representing the low-frequency trend of the sequence; D_j(t) is the j-level detail component, representing the fluctuations at different frequency scales; J is the decomposition level, and in this example, J = 5;

[0166] For power metering data, the Daubechies-4 wavelet basis function is selected for transformation because it performs well in capturing the characteristics of power loads. After transformation, an approximation component A_5 and five detail components D_1 to D_5 are obtained, corresponding to the characteristics of different time scales respectively: D_1: corresponding to the rapid fluctuations within 15 minutes; D_2: corresponding to the fluctuations from 15 to 30 minutes; D_3: corresponding to the fluctuations from 30 minutes to 1 hour; D_4: corresponding to the fluctuations from 1 to 2 hours; D_5: corresponding to the fluctuations from 2 to 4 hours; A_5: corresponding to the long-term trend over 4 hours;

[0167] Calculate the energy proportion of each component: E_i = (∑ t |C_i(t)|2) / (∑ i ∑ t |C_i(t)|2); where C_i(t) represents the coefficient value of the i-th component (A_J or D_j) at time t.

[0168] Identify the dominant components according to the energy proportion. The rule is: the components with an energy proportion exceeding 10% are regarded as dominant components. For the typical load curve of residential users, A_5, D_3, and D_4 are usually the dominant components, indicating that the long-term trend and the fluctuation characteristics from 30 minutes to 2 hours play a major role in the power load.

[0169] Perform feature sensitivity evaluation based on the dominant feature components. Build an association model between features and partitioning effects:

[0170] Extract statistical features from the dominant components, including mean, standard deviation, skewness, kurtosis, maximum value, minimum value, etc.

[0171] For each of the dominant components A_5, D_3, and D_4, calculate the following features: mean μ_i = (1 / N) * ∑ t=1 N\(C_i(t)\); the standard deviation \(\sigma_i=\sqrt{(1 / N)\sum}\) t=1 N \((C_i(t)-\mu_i)^2\); the skewness \(skew_i=(1 / N)\sum\) t=1 N \(((C_i(t)-\mu_i) / \sigma_i)\) 3 ; the kurtosis \(kurt_i=(1 / N)\sum\) t=1 N \(((C_i(t)-\mu_i) / \sigma_i)\) 4 ; the maximum value: \(max_i = max(C_i(t))\); the minimum value: \(min_i = min(C_i(t))\).

[0172] Calculate the feature - effect sensitivity matrix \(S_{i,j}=dP_j / dF_i\); where: \(P_j\) is the \(j\)-th partition effect index (such as load balancing degree, query response time); \(F_i\) is the \(i\)-th feature; \(S_{i,j}\) represents the sensitivity of the effect index \(P_j\) to the feature \(F_i\); Apply principal component analysis (PCA) to reduce the dimension of the feature - effect sensitivity matrix, and retain the principal components with the proportion of explained variance exceeding 85%; Calculate the redundancy for the feature combinations corresponding to the principal components, and remove the highly correlated features with Pearson correlation coefficient exceeding 0.85.

[0173] Through analysis, the identified reduced feature set includes: the mean and standard deviation of the long - term trend component (\(A_5\)); the standard deviation and kurtosis of the medium - term fluctuation component (\(D_3\)); the skewness and kurtosis of the short - medium - term fluctuation component (\(D_4\)); These features can effectively capture the time - varying characteristics of the power load and provide a basis for subsequent partitioning.

[0174] For each time - series feature in the reduced feature set, perform phase - space reconstruction. Taking the mean sequence of the long - term trend component \(A_5\) as an example:

[0175] Use the principle of minimum mutual information to calculate and determine the optimal time - delay parameter \(\tau\): \(I(X(t), X(t + \tau))=\sum_{x(t),x(t+\tau)}p(x(t),x(t+\tau))\times log(p(x(t),x(t+\tau)) / (p(x(t))\times p(x(t+\tau))))\); where: \(I(X(t), X(t + \tau))\) is the mutual information function; \(p(x(t),x(t+\tau))\) is the joint probability distribution; \(p(x(t))\) and \(p(x(t+\tau))\) are the marginal probability distributions.

[0176] Calculate the mutual information at different \(\tau\) values, and take the first local minimum point as the optimal time - delay parameter. For a typical daily load curve, \(\tau\) is usually 5 - 6 hours.

[0177] Determine the optimal embedding dimension \(m\): Calculate the embedding dimension using the improved false nearest neighbor algorithm (FNN). The core of the algorithm is to calculate the proportion of false nearest neighbors at different dimensions: \(FNN(m)=\sum_{i = 1}^{N - m\tau}\frac{\Theta(R_i(m + 1) / R_i(m)-R_{threshold})}{N - m\tau}\)

[0178] where: \(\Theta\) is the Heaviside step function; \(R_i(m)\) is the distance from point \(i\) to its nearest neighbor in the \(m\)-dimensional space; \(R_{threshold}\) is the threshold, usually set to 10;

[0179] Take the value of \(m\) when \(FNN(m)<0.01\) as the optimal embedding dimension. For power load data, \(m\) is usually 3 - 4.

[0180] Use the optimal time delay parameter \(\tau = 6\) hours and the optimal embedding dimension \(m = 4\) to reconstruct the time series \(X(t)\) into a 4-dimensional phase space: \(X(t)\to[X(t),X(t + 6),X(t + 12),X(t + 18)]\); Generate phase space trajectory data.

[0181] Use kernel density estimation (KDE) to calculate the phase space density \(f(x)=\frac{1}{nh}\sum_{i = 1}^nK(\frac{x - x_i}{h})\); where: \(K\) is the Gaussian kernel function; \(h\) is the bandwidth parameter, determined using the Silverman rule; \(x_i\) is the sample point

[0182] Calculate the dynamic characteristic indicators based on the phase space trajectory data:

[0183] Calculate the largest Lyapunov exponent \(\lambda=\lim_{t\to\infty}\frac{1}{t}\log(\frac{\|\delta Z(t)\|}{\|\delta Z(0)\|})\); where: \(\delta Z(t)\) is the trajectory separation vector at time \(t\); \(\delta Z(0)\) is the initial separation vector;

[0184] For residential users, \(\lambda\) is usually in the range of 0.02 - 0.05, indicating that the system has weak chaos.

[0185] Calculate the correlation dimension \(D_2=\lim_{r\to0}\frac{\log(C(r))}{\log(r)}\);

[0186] where: \(C(r)\) is the correlation integral, \(C(r)=\frac{2}{N(N - 1)}\sum\) i,j=1,i≠j N \(\Theta(r-\|x_i - x_j\|)\); \(\Theta\) is the Heaviside step function; For power load data, \(D_2\) is usually in the range of 2.3 - 2.8, reflecting the geometric complexity of the system.

[0187] Calculate the entropy rate \(h=\sum\) i=1 k \(\lambda_i (\lambda_i > 0)\); where \(\lambda_i\) is the Lyapunov exponent spectrum of the system.

[0188] For power loads, \(h\) is usually between 0.03 - 0.07, which reflects the information generation rate of the system.

[0189] Use the sparse identification method (SINDy) to identify the kinetic equation \(dX / dt = F(X)=\Xi(X)\cdot\Theta\) from the data

[0190] where: \(X\) is the system state vector; \(\Xi(X)\) is the candidate function library, including polynomials, trigonometric functions, etc.; \(\Theta\) is the coefficient vector;

[0191] Solve for \(\Theta\) through L1 regularization: \(\min||X\) * \(-\Xi(X)\Theta||_2^2+\alpha||\Theta||_1\); where \(\alpha\) is the regularization parameter, and in this example, it is taken as 0.1.

[0192] For typical residential power loads, the identified simplified kinetic equations may be in the form of: \(dx1 / dt = 0.05x1 - 0.02x1x2+0.01x3\); \(dx2 / dt = 0.03x2 + 0.04x1 - 0.01x2²\); \(dx3 / dt=-0.02x3 + 0.03x1x2+0.01\sin(x1)\);

[0193] Construct a recurrence plot and detect transition points based on the phase space trajectory data and kinetic model parameters:

[0194] Construct the recurrence plot \(RP(i,j)=\Theta(\varepsilon - ||X(i)-X(j)||)\); where: \(\Theta\) is the Heaviside function; \(\varepsilon\) is the distance threshold, taken as 10% of the phase space diameter; \(X(i)\) is the state vector of the phase space trajectory at time point \(i\). The generated recurrence plot matrix is a binary matrix, with 0 indicating that the distance between two state points is greater than the threshold, and 1 indicating that the distance is less than the threshold.

[0195] Analyze the topological properties of the recurrence plot: Calculate the node centrality: \(C_i=\sum_j RP(i,j) / N\); Calculate the clustering coefficient: \(CC_i=\sum_{j,k} RP(i,j)\cdot RP(j,k)\cdot RP(k,i) / \sum_{j,k} RP(i,j)\cdot RP(i,k)\);

[0196] Generate the topological property vector \([C_i, CC_i]\), which reflects the kinetic state of the system at different time points.

[0197] Calculating the Wasserstein distance to measure distribution changes: For consecutive time windows \(W_1\) and \(W_2\), calculate the Wasserstein distance between the phase space probability distributions \(P_1\) and \(P_2\): \(W(P_1, P_2)=\inf_{\gamma\in\Gamma(P_1,P_2)}\iint||x - y||_2d\gamma(x,y)\); where \(\Gamma(P_1,P_2)\) is the set of all joint distributions with marginal distributions \(P_1\) and \(P_2\).

[0198] Calculate the sequence of Wasserstein distances for the sliding window and identify the points with sudden increases in distance as distribution mutation points.

[0199] Combining the change trends of topological feature vectors and distribution mutation points, use a hierarchical attention mechanism to identify multi-scale transition points: Intra-day transition points: Usually corresponding to changes in the electricity consumption pattern within a season; Seasonal transition points: Corresponding to changes in the electricity consumption pattern caused by seasonal alternation; Annual transition points: Corresponding to the evolution of long-term electricity consumption behavior.

[0200] Based on the multi-scale transition points, segment the time series into different intervals, and each interval corresponds to an evolution pattern:

[0201] Extract the feature representations of each pattern to form a pattern feature library. It includes: Dynamic parameters: Lyapunov exponents, correlation dimensions, entropy rates, etc.; Statistical features: mean, variance, kurtosis, etc.; Spectral features: main frequency components and their amplitudes.

[0202] Construct an evolution pattern knowledge graph, including: Nodes representing the identified typical evolution patterns, with attributes containing features from the pattern feature library. Edges representing the transition relationships and probabilities between patterns.

[0203] For residential users, typical patterns include: "Normal weekday pattern": Obvious morning and evening peaks, and low valleys at noon and late at night. "Holiday pattern": The load is relatively stable throughout the day, without obvious peaks. "Seasonal transition pattern": The load gradually increases or decreases. "Extreme weather pattern": The load is abnormally high or low.

[0204] Use a graph attention network (GAT) to process the pattern knowledge graph. The node feature update formula is: \(h\) i l+1 \(=\sigma(\sum\) j∈N(i) \(\alpha_{ij}W h\) j l );where: \(h\) j lIt is the feature representation of the i-th node at the l-th layer; α_ij is the attention coefficient; W is the weight matrix; σ is the activation function. The model is trained by minimizing the prediction error. The finally output dynamic evolution pattern recognizer can: recognize the evolution pattern to which the current data belongs; predict the probability and time point of pattern conversion; infer the features of unseen patterns.

[0205] This recognizer provides forward-looking prediction of data behavior for partition rebalancing, enabling the system to adjust the partition strategy "clairvoyantly".

[0206] Based on the results of dynamic evolution pattern recognition, this embodiment performs low-overhead partition rebalancing to achieve dynamic optimization of data distribution.

[0207] First, collect the load statistical data of each partition and construct a multi-dimensional imbalance evaluation model:

[0208] Calculate the imbalance index of data distribution Imbalance = (σ / μ) × 100%; where: σ is the standard deviation of the load of each partition; μ is the average value of the load of each partition;

[0209] For example, at a certain moment, the data volumes of 5 partitions are [120, 85, 180, 60, 155] GB respectively, then: μ = 120 GB, σ = 46.04 GB, Imbalance = 38.37%.

[0210] Analyze the data access pattern and calculate the hotspot concentration Hotspot = (∑ i∈H v_i) / (∑ i=1 n v_i)

[0211] where: H is the set of hotspot partitions (the top 20% of partitions with the highest access volume); v_i is the access volume of the i-th partition.

[0212] If the number of queries per second for 5 partitions is [320, 150, 450, 130, 250] times, then the hotspot partition is the 3rd partition, Hotspot = 450 / (320 + 150 + 450 + 130 + 250) = 34.61%

[0213] Calculate the resource utilization imbalance ResImbalance = (σ_r / μ_r) × 100%;

[0214] where: σ_r is the standard deviation of the resource utilization of each partition; μ_r is the average value of the resource utilization of each partition;

[0215] If the CPU utilization of 5 partitions is [75%, 45%, 85%, 40%, 65%], then: μ_r = 62%, σ_r = 18.65%, ResImbalance = 30.08%.

[0216] Construct the comprehensive imbalance index I_comprehensive = w_1×Imbalance + w_2×Hotspot + w_3×ResImbalance. Where w_1, w_2, and w_3 are weight coefficients, determined according to the business importance, and take [0.3, 0.4, 0.3] in this example. Calculate to get I_comprehensive = 0.3×38.37% + 0.4×34.61% + 0.3×30.08% = 34.39%

[0217] Construct a rebalancing revenue model based on the comprehensive imbalance index:

[0218] Construct the balance degree revenue model B = f(I_current, I_expected)= k × (I_current - I_expected) / I_current; where: I_current is the current comprehensive imbalance index, which is 34.39% in this example; I_expected is the expected imbalance index, and the target value is set to 15%; k is the revenue coefficient, and the value is 100;

[0219] Calculate to get B = 100 × (34.39% - 15%) / 34.39% = 56.38;

[0220] Construct an environment-aware adaptive weight adjustment mechanism W(t) = W_base + ΔW(load(t), priority(t));

[0221] Where: W_base is the basic weight vector [0.6, 0.3, 0.1], corresponding to the balance revenue, rebalancing cost, and business interference degree; load(t) is the current system load, assumed to be 70%; priority(t) is the business priority, assumed to be "high";

[0222] When the load is high (>60%) and the priority is "high", the weight is adjusted to: ΔW = [-0.1, 0.0, 0.1], indicating reducing the balance revenue weight and increasing the business interference degree weight.

[0223] It is calculated that W(t) = [0.5, 0.3, 0.2]. Based on the time series prediction model, the future system load trend is predicted: The ARIMA(2,1,2) model is used to predict the load in the next 24 hours, and the results show that the load will drop from the current 70% to 50%.

[0224] Based on this prediction, the weights are adjusted: When the predicted load drops below <60%, the weights are adjusted to: ΔW_future = [0.1,0.0, -0.1]. It is calculated that the pre-adjusted weight vector W_adjusted = [0.6, 0.3, 0.1].

[0225] Construct a multi-objective rebalancing benefit function E = α·B - β·C - γ·D;

[0226] Where: α, β, γ are the components of the pre-adjusted weight vector, which are 0.6, 0.3, 0.1 respectively; B is the equilibrium benefit value, calculated as 56.38; C is the rebalancing cost, to be estimated; D is the business interference degree, to be estimated.

[0227] Construct a fine-grained cost model for rebalancing operations and use a Structural Causal Model (SCM) to predict system state changes:

[0228] Establish a fine-grained cost model:

[0229] C_data Data migration bandwidth cost = Amount of migrated data × Unit bandwidth cost; = 40GB × 0.05 = 2.0.

[0230] C_index Index reconstruction calculation cost = Index size × Reconstruction complexity coefficient = 5GB × 0.4 = 2.0.

[0231] C_query Query performance impact cost = Number of queries × Response time increase rate × Unit response time cost = 400qps × 0.15 × 0.05 = 3.0.

[0232] C_io: I / O load increase cost = I / O increment × Unit I / O cost = 200 IOPS × 0.01 = 2.0. The total cost C = C_data + C_index + C_query + C_io = 9.0.

[0233] Construct a Structural Causal Model (SCM) to represent the causal relationship between system state variables: S = f(X, do(R))

[0234] Where: S is the system state vector [response time, throughput, error rate]; X is the influencing factor vector [data volume, query complexity, hardware resources]; do(R) represents the intervention of executing the rebalancing operation R.

[0235] The causal diagram obtained through learning from historical data shows that the rebalancing operation directly affects the data distribution and index status, and thus affects the response time and throughput.

[0236] Calculate the difference in system state ΔS between executing rebalancing and not executing rebalancing: ΔS = E[S|do(R = 1)] - E[S|do(R = 0)]; where E represents the expectation function.

[0237] According to the prediction of the causal model, executing rebalancing will result in: the response time increases by 15% temporarily and decreases by 20% in the long term; the throughput decreases by 10% temporarily and increases by 25% in the long term; the change in the error rate is not significant;

[0238] Update the cost estimate based on the predicted state change ΔS: Considering the benefits of long-term performance improvement, the adjusted cost is: C_adjusted = C - long_term_benefit = 9.0 - 3.0 = 6.0.

[0239] Perform uncertainty-aware benefit prediction based on Bayesian networks and quantum probability theory:

[0240] Construct a Bayesian network to represent the probabilistic dependence relationship between the system state and the rebalancing benefit: The network nodes include: partition imbalance degree, system load, data growth rate, rebalancing degree, and final benefit

[0241] Example of conditional probability representation: P(benefit = high|imbalance degree = high, load = medium, data growth rate = low, rebalancing degree = medium) = 0.75.

[0242] Use Monte Carlo simulation to generate multiple possible results: Conduct 1000 simulations to obtain the benefit distribution: average benefit: E[B] = 48.5; benefit standard deviation: σ(B) = 12.3; 5% quantile: 26.8; 95% quantile: 68.7;

[0243] Calculate the risk-adjusted expected benefit: RAR = E[B] - λ × σ(B)= 48.5 - 1.2 × 12.3=33.7; where λ = 1.2 is the risk aversion coefficient, determined according to the business importance.

[0244] Introducing quantum probability representation to handle complementary uncertainties: The state of the system is represented by the density operator ρ, and the quantum observable A represents the return measurement. The expected return is calculated as: E[A] = Tr(ρA); this method can better express the interference effects between different decision options.

[0245] Determining the optimal decision boundary based on quantum information geometry: Calculate the quantum Fisher information matrix F_Q and the quantum Cramér-Rao bound to determine the optimal decision boundary. Decision rule: Rebalancing is performed when the risk-adjusted return RAR > 30 and the decision certainty > 0.8. In this example, RAR = 33.7 and the decision certainty = 0.85, meeting the execution conditions.

[0246] Using multi-agent reinforcement learning to optimize the rebalancing decision strategy:

[0247] Constructing a reinforcement learning environment: The state space S includes metrics such as load imbalance degree, hot spot concentration, and system resource utilization. The action space A {perform full rebalancing, perform partial rebalancing, adjust rebalancing parameters, delay rebalancing}. The reward function R is based on the multi-objective benefit function E = 0.6×B - 0.3×C - 0.1×D.

[0248] Designing a multi-agent architecture, configuring one decision agent for each of the 5 data partitions. Using a graph attention network (GAT) to construct the communication mechanism between agents: h i l+1 = σ(∑ j∈N(i) α_ij W h j l );

[0249] where: h i l is the state representation of the i-th agent at the l-th layer; α_ij is the attention weight, calculated by softmax(LeakyReLU(a T [Wh_i, Wh_j])); W is the parameter matrix; agents share local state information through this mechanism to coordinate decisions.

[0250] Training the agent network to obtain the optimal rebalancing strategy: For the current system state, the optimized decision is "perform partial rebalancing", and the specific operation is: Only migrate the data in the 3rd partition (hot spot) and the 4th partition (low load); the migration ratio is 30%; execute during the low load period; the maximum bandwidth limit is 50MB / s.

[0251] Perform minimum interference migration path planning based on an optimization strategy, specifically including: constructing a data dependency graph and analyzing the access relationships between shards. Represent the shards of data to be migrated as a graph G(V, E), where the nodes V represent shards and the edges E represent the access relationships between shards. The edge weight w_ij represents the access frequency between shards i and j. According to the analysis, there are high access relationships among the shards {S3.1, S3.2, S3.5} in the hot partition (the 3rd partition).

[0252] Calculate the shard association graph and design the migration priority: Detect the community structure based on the Louvain algorithm to ensure the co-migration of highly associated shards. The priority calculation formula: Priority(S_i) = w_1×Hot(S_i) + w_2×Size(S_i)+ w_3×Dependency(S_i). Where: Hot(S_i) is the heat of shard S_i; Size(S_i) is the size of shard S_i; Dependency(S_i) is the degree of dependence of shard S_i; w_1, w_2, w_3 are weight coefficients with values in [0.5, 0.3, 0.2].

[0253] The calculated migration priorities of the shards in the 3rd partition are: S3.1: 0.85; S3.2: 0.78; S3.5: 0.72.

[0254] Generate a migration path plan with minimum interference: Migration batch plan [{S3.1, S3.2}, {S3.5}, {S4.3, S4.7}]; Target partition assignment S3.1→P1, S3.2→P5, S3.5→P2, S4.3→P5, S4.7→P1. Execution time window 02:00 - 04:00 (low system load period).

[0255] Execute incremental partition data migration: Set the initial migration rate to 30MB / s; Degree of parallelism is 2. The waiting time between batches is 5 minutes.

[0256] Monitor the execution status in real time: System load monitoring: CPU, memory, I / O; Quality of service monitoring: Response time, throughput, error rate;

[0257] Adjust the execution parameters based on real-time monitoring data: After the first batch is completed, the CPU load is below the threshold, and the migration rate is adjusted to 40MB / s; During the execution of the second batch, it is found that the response time increases slightly, and the degree of parallelism is reduced to 1.

[0258] Evaluate the rebalancing effect after migration: Comprehensive imbalance index before rebalancing: 34.39%; Comprehensive imbalance index after rebalancing: 18.25%; Improvement degree: 47.0%; Actual resource consumption: CPU increased by 12% on average, bandwidth used 35 MB / s; Service impact: Response time increased temporarily by 8%, no increase in error rate.

[0259] Update the parameters of the rebalancing decision model: Based on the actual execution results, update the cost model parameters and the uncertainty model to optimize the next rebalancing decision.

[0260] Based on the updated partitioning scheme, this embodiment performs stream-batch collaborative processing on the real-time data stream and historical batch data in the quality-tagged dataset.

[0261] Build a unified stream-batch processing model, including:

[0262] Design a unified data model, including: Time field: providing the timestamp of the event occurrence; Entity ID: identifying the data source (such as the meter ID); Measurement value: recording the actual measurement data; Quality tag: indicating the data quality level; Processing status: marking the data processing stage.

[0263] Build a Lambda + Kappa architecture: Lambda layer: processes batch historical data and provides accurate but high-latency analysis results; Kappa layer: processes real-time data streams and provides near-real-time but slightly less accurate analysis results; Coordination layer: responsible for result fusion and consistency management.

[0264] Apply the updated partitioning scheme: Redistribute the data in 5 partitions according to the new scheme to ensure that the same partitioning strategy is used for stream processing and batch processing.

[0265] Perform incremental feature updates and real-time anomaly detection on the real-time data stream:

[0266] Real-time data distribution: Receive the real-time data stream in the quality-tagged dataset and perform real-time distribution according to the partitioning scheme

[0267] For meter data, determine the partition by taking the modulus of the hash of the meter ID with the number of partitions: Partition(meter_id) = Hash(meter_id) % num_partitions.

[0268] Sliding window processing: Window size is 10 minutes; Sliding step is 1 minute; Calculate window statistics as mean, standard deviation, rate of change, peak value;

[0269] Calculate for the 10-minute window data of meter E10056782: Mean is 4.8 kWh; Standard deviation is 0.7 kWh; Rate of change is +2.1%; Peak value is 5.9 kWh.

[0270] Incremental Feature Update: Update features using the Exponential Weighted Moving Average (EWMA) method: F_new = α × F_current + (1 - α) × F_old; where α is the smoothing coefficient, taking 0.3.

[0271] When a pattern change is detected, trigger a fast feature update: α_dynamic = min(0.8, base_α × change_magnitude).

[0272] Real-time Anomaly Detection: Detect point anomalies based on statistical control charts: If |x_t - μ| > k × σ, then mark as an anomaly where k is the sensitivity coefficient, taking 3.

[0273] Identify trend anomalies based on the change point detection algorithm: Use the CUSUM algorithm to monitor the cumulative deviation S_t = max(0, S_{t - 1} + (x_t - μ - δ)); If S_t > h, then trigger an alarm; where δ is the offset parameter and h is the threshold.

[0274] Detect pattern anomalies based on pattern matching: Use the Dynamic Time Warping (DTW) algorithm to calculate the distance between the current pattern and the normal pattern.

[0275] Generate real-time processing results: Processing timestamp is 2024-02-15T08:40:00Z; Entity ID is E10056782; Aggregate value is 4.8 kWh (10-minute average); Anomaly flag is normal; Updated features are {trend: increasing, pattern: morning_peak}.

[0276] Perform time series analysis and quality verification on historical batch data:

[0277] Load historical metering data for the past 30 days from the data warehouse and reorganize it according to the new partitioning scheme.

[0278] For the data in Partition 1: Data volume is 105GB; Number of records is approximately 840 million; Time range is from 2024-01-15 to 2024-02-14;

[0279] Seasonal Decomposition: Use the STL (Seasonal-Trend decomposition using Loess) method to decompose into trend, seasonal, and residual components. Use the Fast Fourier Transform (FFT) to identify hidden cycles. Use hierarchical clustering to identify typical load patterns.

[0280] Verify whether the metering data conforms to physical constraints and grid balance conditions. Identify missing intervals in the time series and evaluate the imputation effect. Conduct systematic anomaly detection on the batch processing results and filter out possible false positives.

[0281] Generate the verified batch processing results: Processing date is 2024-02-14; Data coverage is a 30-day complete record; Quality score is 92.5% (high quality); Anomaly summary is 21 anomaly points and 3 anomaly intervals detected; Typical patterns are working day pattern, weekend pattern, and holiday pattern identified.

[0282] Perform multi-level data fusion based on the real-time processing results and the verified batch processing results:

[0283] Determine the result overlap interval: There is an overlap between the most recent 24-hour data of the real-time processing and the last 24 hours of the batch processing results. Calculate the consistency metrics: Mean deviation |(μ_real - μ_batch) / μ_batch| × 100% = 2.3%. Proportion of anomaly points with consistent detection results = 92.7%. Consistency of the trend directions of the two results = 96.5%.

[0284] Determine the fusion strategy. At the data point level: When the quality flag of the real-time result is "high" and the deviation from the batch processing result is <5%, give priority to the real-time result. For trend analysis: When the confidence level of the batch processing result is >90%, give priority to the batch processing result. For anomaly detection: When the two results are inconsistent, apply a weighted voting mechanism to determine the final result.

[0285] Execute data fusion. Use the Bayesian fusion method to combine the distribution characteristics of the two results: p(z|x,y) ∝ p(x|z)p(y|z)p(z); where: z is the true value after fusion; x is the real-time processing result; y is the batch processing result; p(z) is the prior distribution;

[0286] Generate the final processing result: Processing time is 2024-02-15T08:40:00Z; Short-term trend is +2.3% (from real-time processing); Long-term pattern is the regular working day pattern (from batch processing); Anomaly status result is normal; Quality metric is 95.3% (improved after fusion).

[0287] In this embodiment, the final processing result is analyzed and applied in multiple dimensions to achieve the storage, visual presentation, and anomaly warning of metering data.

[0288] Data model organization. Design the data model according to the star schema, including: A fact table for storing metering readings and calculated metrics; Dimension tables including time dimension, space dimension, equipment dimension, etc.

[0289] Multi-level storage strategy. Hot data (last 7 days): stored in an in-memory database; warm data (last 90 days): stored on SSDs; cold data (over 90 days): stored on HDDs and compressed; archived data (over 1 year): stored in object storage;

[0290] Build multi-dimensional indexes. Time index: based on B+ tree to achieve efficient time range queries; spatial index: use R tree to support region queries; device index: adopt hash index to accelerate queries by device ID; composite index: support multi-condition combined queries.

[0291] Index update strategy. Hot data index: updated in real time; warm data index: batch updated hourly; cold data index: batch updated daily.

[0292] Execute multi-dimensional statistical analysis based on the persistent result data:

[0293] Time dimension analysis, daily load curve: 24-hour load variation; weekly load pattern: difference between weekday and weekend patterns; monthly trend: month-on-month load change; annual cycle: seasonal change pattern.

[0294] Spatial dimension analysis, regional load distribution: heat map shows electricity consumption density in different regions; propagation pattern: diffusion pattern of electricity consumption behavior in space; correlation analysis: load correlation between adjacent regions.

[0295] User dimension analysis, user clustering: grouping users based on electricity consumption behavior; abnormal user identification: detecting users with abnormal electricity consumption behavior; evolution of electricity consumption pattern: changes in users' electricity consumption behavior over time.

[0296] Generate a set of visualization resources, interactive load curve: supports zoom in, zoom out, and comparison functions; regional heat map: dynamically displays load distribution at different times; anomaly detection panel: highlights detected anomaly points; user profile dashboard: shows users' electricity consumption characteristics and behavior patterns.

[0297] Build an interactive data exploration interface, multi-level drill-down: from the global view to detailed data points; multi-dimensional filtering: filter data by time, region, user type, etc.; comparative analysis: supports data comparison between different time periods and different regions; prediction simulation: simulate future load scenarios based on historical data.

[0298] Perform anomaly identification and early warning on the final processing results:

[0299] Statistical anomalies: detect numerical anomalies based on the 3σ rule; pattern anomalies: detect pattern deviations based on the DTW algorithm; context anomalies: consider anomalies in environmental factors (such as temperature); collective anomalies: abnormal patterns shown by multi-device collaboration.

[0300] Using the decision tree algorithm to filter false alarms, the accuracy rate is increased to 92%; classifying anomalies into: sudden anomalies, gradual anomalies, periodic anomalies, and systematic anomalies;

[0301] Generating hierarchical warning information, level 1 (information): slight deviation, no immediate processing required; level 2 (warning): obvious deviation, attention needed; level 3 (severe): major deviation, timely processing required; level 4 (urgent): extreme deviation, immediate response required;

[0302] Warning push and response, information level: recorded in the system log; warning level: pushed to the monitoring interface; severe level: sending email and SMS notifications; urgent level: triggering automatic call and emergency response;

[0303] Closed-loop management, recording the warning handling status: unprocessed, processing, resolved, false alarm; statistical warning effectiveness: true positive rate, false positive rate; continuously optimizing warning rules: adjusting thresholds based on historical warning effectiveness.

[0304] The preferred embodiments of the present invention are described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A metering data processing method based on dynamic partition rebalancing and stream-batch collaboration, characterized in that It includes the following steps: Collect measurement data from multi-source heterogeneous data sources and perform preprocessing to form a quality-marked dataset containing quality level markings; Based on the quality-marked dataset, execute adaptive partition feature extraction based on time-varying features and low-overhead partition rebalancing through a dynamic partition management module to generate an updated partition scheme; specifically: Based on the quality-marked dataset, execute adaptive partition feature extraction of time-varying features to obtain partition features; Apply the partition features to perform an initial partition of the data, generate an initial partition scheme, and obtain the system partition state; Combine the pre-stored system operation state and load statistical data, execute low-overhead partition rebalancing, update the initial partition scheme, and generate an updated partition scheme; Based on the updated partition scheme, perform stream-batch collaborative processing on the real-time data stream in the quality-marked dataset and the pre-stored historical batch data, and generate a final processing result through multi-level data fusion; Perform multi-dimensional analysis and application on the final processing result to achieve storage, visual presentation, and anomaly warning of measurement data; Obtain partition features, including: Input the quality-marked dataset into a multi-scale time-varying feature decomposition unit to extract the dominant feature components; Perform feature sensitivity evaluation and screening based on the dominant feature components to generate a reduced feature set; Use the reduced feature set to construct a dynamic evolution pattern recognizer and output a pattern prediction result; Perform adaptive feature fusion and weight allocation according to the pattern prediction result to generate weighted fusion features, that is, partition features; Construct a dynamic evolution pattern recognizer, including: Receive the reduced feature set, determine the optimal time delay parameter based on the principle of minimum mutual information, combine the improved false nearest neighbor algorithm to determine the optimal embedding dimension, and reconstruct the one-dimensional time series into multi-dimensional phase space trajectory data; generate a phase space density map using kernel density estimation; Calculate the Lyapunov exponent, correlation dimension, and entropy rate of the phase space trajectory data based on the phase space density map, combine the sparse identification method and deep learning to identify the dynamic equation, and obtain the dynamic model parameters through L1 regularization optimization; Construct a recurrence plot based on the phase space trajectory data and dynamic model parameters, analyze the network topology characteristics, combine the Wasserstein distance to calculate the distribution change rate, and identify multi-scale transition points; Segment the time series according to the multi-scale transition points, extract pattern features to construct a knowledge graph, train a graph neural network to learn the pattern conversion rules, and output a dynamic evolution pattern recognizer.

2. The method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration according to claim 1, wherein The steps to form a quality-marked dataset containing quality level markings include: According to the preset data collection rules, access multi-source heterogeneous data sources and collect measurement data to generate raw measurement data; Perform protocol conversion and parsing on the raw measurement data, as well as data cleaning and standardization processing to obtain normalized measurement data; Calculate the quality indicators of the normalized measurement data, and add quality level markings to each piece of normalized measurement data based on the quality indicators to form a quality-marked dataset.

3. A metering data processing method based on dynamic partition rebalancing and stream-batch collaboration according to claim 1, characterized in that The steps to perform stream-batch collaborative processing to generate a final processing result include: Construct a unified stream-batch processing model and apply the updated partition scheme to the real-time data stream in the quality-marked dataset and the pre-stored historical batch data; Perform incremental feature updates and real-time anomaly detection on real-time data streams to generate real-time processing results; Perform time series analysis and quality verification on historical batch data to obtain verified batch processing results; Perform multi-level data fusion based on the real-time processing results and the verified batch processing results to generate final processing results.

4. A metering data processing method based on dynamic partition rebalancing and stream-batch collaboration according to claim 1, characterized in that The steps for performing multi-dimensional analysis and applications to achieve the storage, visual presentation, and anomaly warning of metering data include: Organize the final processing results according to a predetermined data model and write them into the persistent storage module to obtain persistent result data and establish a multi-dimensional index; Perform multi-dimensional statistical analysis based on the persistent result data to generate a set of visual resources; Identify anomalies in the final processing results according to preset anomaly rules and trigger corresponding warning messages.

5. The metering data processing method based on dynamic partition rebalancing and stream-batch collaboration according to claim 1, characterized in that The steps for performing low-overhead partition rebalancing and generating an updated partition scheme include: Collect load statistical data for each partition; calculate the imbalance index based on the load statistical data, analyze the data access pattern and identify hot partitions, construct a multi-dimensional imbalance evaluation model, and output a comprehensive imbalance index; Construct a rebalancing benefit model based on the comprehensive imbalance index, combine the system partition status and the system operation status to construct an environment-aware weight adjustment mechanism, predict the rebalancing cost and perform uncertainty modeling, and optimize the rebalancing strategy; Construct a data dependency graph according to the optimized rebalancing strategy, calculate the shard association graph and construct a migration priority algorithm to generate a migration path scheme with minimal interference; Execute incremental partition data migration according to the migration path scheme, monitor the execution status in real time and dynamically adjust the migration parameters; Based on the adjusted migration parameters, evaluate the improvement degree of the comprehensive imbalance index before and after rebalancing, calculate the actual resource consumption cost, update the parameters of the rebalancing decision model and generate an updated partition scheme.

6. A metering data processing method based on dynamic partition rebalancing and stream-batch collaboration according to claim 5, characterized in that The steps for predicting the rebalancing cost and performing uncertainty modeling to optimize the rebalancing strategy include: Construct a multi-objective rebalancing benefit function based on the comprehensive imbalance index and output an equilibrium benefit value and a dynamic weight vector; Combine the pre-stored current system load conditions and perform rebalancing cost prediction based on causal inference to obtain a cost prediction value and a state change prediction; Based on the equilibrium benefit value and the cost prediction value, perform uncertainty-aware benefit prediction and output a risk-adjusted benefit and a quantum optimization decision boundary; Based on the risk-adjusted benefit, the quantum optimization decision boundary, and the state change prediction, optimize the rebalancing strategy through reinforcement learning and output the optimal rebalancing strategy.

7. A metering data processing method based on dynamic partition rebalancing and stream-batch collaboration according to claim 6, characterized in that, The steps for constructing a multi-objective rebalancing benefit function include: Construct an equilibrium degree benefit model after rebalancing based on the comprehensive imbalance index: B = f(I_current, I_expected); where I_current is the current comprehensive imbalance index, I_expected is the expected imbalance index, f() is a function, and the equilibrium benefit value B is output; Combine the current state of the system and construct an environment-aware adaptive weight adjustment mechanism: W(t) = W_base + ΔW(load(t), priority(t)); Where \(W(t)\) is the weight vector at time \(t\), \(W_{base}\) is the basic weight, \(load(t)\) is the system load at time \(t\), \(priority(t)\) is the service priority, \(\Delta W\) is the weight change, and the dynamic weight vector \(W\) is output; Construct a time series prediction model to predict the system load trend in the future period to obtain the load prediction value; Adjust the dynamic weight vector \(W\) based on the load prediction value to obtain a pre-adjusted weight vector; Combine the equilibrium revenue value \(B\) with the pre-adjusted weight vector to construct a complete multi-objective rebalancing benefit function: \(E=\alpha\cdot B - \beta\cdot C - \gamma\cdot D\); Where \(\alpha\), \(\beta\), \(\gamma\) are the weight coefficients from the pre-adjusted weight vector, \(B\) is the equilibrium revenue value, \(C\) is the rebalancing cost, and \(D\) is the service interference degree.

Citation Information

Patent Citations

  • Data detection method and device for power grid regulation and control multi-source time sequence data batch stream fusion

    CN118069643A

  • Method for dynamically partitioning relational cluster database based on spatio-temporal data

    CN119226421A