Metering data processing method based on dynamic partition weight balance and flow batch collaboration
By adopting dynamic partition rebalancing and stream batch collaborative processing methods in power metering data processing, problems such as static partition in the prior art and stream batch splitting are solved, efficient and stable metering data processing and prediction capabilities are achieved, and operation and maintenance costs are reduced.
Patent Information
- Application Number
- CN202510424426.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
When processing power metering data with high dimensions, high time variability and strong correlation, the prior art has problems such as static partitioning inadequacy, stream batch processing splitting, insufficient data quality perception, and unscientific decision-making on balance, resulting in unstable system performance, uneven quality and high operation and maintenance costs.
The metered data processing method based on dynamic partition rebalancing and flow batch collaboration is adopted. The dynamic partition management module performs adaptive partition feature extraction and low-overhead partition rebalancing based on time-varying features, and combines multi-level data fusion to generate the final processing results, and the storage, visual presentation and abnormal warning of metered data are realized through multi-dimensional analysis and application.
It improves the accuracy and adaptability of feature extraction, achieves a deep understanding and prediction of data behavior, reduces system interference through a scientific and accurate decision-making mechanism, ensures service continuity, and improves prediction accuracy and decision-making robustness.
Smart Images

Figure CN119939362A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent metering and data management, and in particular to a metering data processing method based on dynamic partition rebalancing and stream-batch collaboration. Background Art
[0002] With the continuous deepening of smart grid construction, metering data has become the core resource for power system operation, management and decision-making. Power metering data has the characteristics of high dimensionality, high time variability and strong correlation. The amount of data generated daily is growing exponentially, which brings huge challenges to traditional data processing systems. Accurate and efficient processing of these massive metering data is of great significance for realizing the intelligent operation of power grids, improving energy utilization efficiency, and supporting refined management and service decision-making. The quality of metering data processing directly affects the operating efficiency and service quality of power companies, the fairness and transparency of the power market, and is also related to the safe and stable operation of the power grid and the realization of energy conservation and emission reduction goals. Therefore, building a set of processing methods that can cope with the variability and complexity of metering data has far-reaching strategic significance for promoting the digital transformation and intelligent upgrading of the power industry.
[0003] At present, the technical route of static partitioning and single processing mode is mainly adopted in the field of metering data processing. In terms of data partitioning, mainstream technologies adopt static partitioning strategies based on fixed hash functions or predefined ranges, such as consistent hash partitioning and range partitioning. These methods are rarely adjusted after the partition boundaries are determined in the initial design stage, and lack the ability to adapt to dynamic changes in data. In terms of data processing mode, existing systems mostly adopt a single mode of batch processing or stream processing. Batch processing systems such as Hadoop and Spark are good at processing large-scale historical data but have poor real-time performance. Stream processing systems such as Storm and Flink can process real-time data streams but have limited analysis depth. In addition, existing solutions are relatively weak in data quality perception. They often adopt a unified processing strategy and ignore data quality differences, resulting in low-quality data consuming too many resources or high-quality data not being fully utilized. In terms of rebalancing strategies, traditional methods mostly adopt simple strategies of global data redistribution or preset threshold triggering. The rebalancing process causes great interference to the system and affects service continuity.
[0004] Although the existing technology has made certain progress, there are still technical bottlenecks in many key links. First, in terms of time-varying feature extraction, the existing methods are mostly based on static features or simple time window statistical features, which cannot effectively capture the complex time-varying patterns and long-term evolution laws in the metering data, resulting in the mismatch between the partitioning features and the actual characteristics of the data, and the poor partitioning effect. Secondly, in terms of phase space reconstruction and dynamic evolution pattern recognition, traditional analysis methods are limited to surface statistical feature analysis, which makes it difficult to reveal the inherent dynamic structure and complex behavior mechanism of the data, and lacks the ability to predict future evolution trends. Third, in terms of partition imbalance assessment and rebalancing decision-making, the existing technologies are mostly based on single indicator evaluation and empirical threshold decision-making, which makes it difficult to comprehensively and accurately assess the imbalance state of the system, and the scientificity and reliability of the decision-making are insufficient. These technical bottlenecks seriously restrict the adaptability of the metering data processing system to changes in data characteristics and the efficiency of resource utilization, resulting in unstable system performance, uneven quality and high operation and maintenance costs. Summary of the invention
[0005] The purpose of the invention is to provide a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration, in order to solve at least one technical problem existing in the prior art.
[0006] Technical solution: A method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration, comprising the following steps:
[0007] Collect metering data from multiple heterogeneous data sources, pre-process the metering data, and form a quality-labeled data set containing quality grade labels; the metering data includes structured metering records generated by smart electricity meters, water meters, and gas meters, as well as system logs and device status data;
[0008] Based on the quality-labeled dataset, the dynamic partition management module performs adaptive partition feature extraction based on time-varying features and low-overhead partition rebalancing to generate an updated partition scheme.
[0009] Based on the updated partitioning scheme, the real-time data stream and pre-stored historical batch data in the quality labeling dataset are processed in a stream-batch collaborative manner, and the final processing result is generated through multi-level data fusion.
[0010] The final processing results are analyzed and applied in multiple dimensions to achieve storage, visualization and abnormal warning of measurement data.
[0011] Beneficial effects: The present invention improves the accuracy and adaptability of feature extraction, realizes a deep understanding and prediction of data behavior; realizes a scientific and accurate decision-making mechanism through a multi-dimensional imbalance assessment model combined with rebalancing benefit-cost modeling; reduces system interference and ensures service continuity through minimum interference rebalancing path planning and incremental rebalancing execution mechanism; improves prediction accuracy and decision robustness by adopting structural causal models and uncertainty-aware benefit prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A flowchart of the steps of a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration is provided in an embodiment of the present application.
[0013] Figure 2 A flow chart of the steps for forming a quality mark data set including quality grade marks provided in an embodiment of the present application.
[0014] Figure 3 A flowchart of the steps for performing adaptive partition feature extraction and low-overhead partition rebalancing based on time-varying features provided in an embodiment of the present application.
[0015] Figure 4 A flowchart of the steps for executing stream-batch collaborative processing provided in an embodiment of the present application.
[0016] Figure 5 A flowchart of the steps for performing multi-dimensional analysis and application provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0018] It should be noted that in order to clearly show the steps of this application, serial numbers are marked for each step in the specification. These serial numbers are only used for the convenience of explanation and do not limit the order of execution of the steps. In actual operation, according to the technical requirements of the specific implementation scenario, the steps can be executed in a different order than that shown in the specification, and in some cases, parallel processing between steps can be achieved.
[0019] like Figure 1 As shown, a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration includes the following steps:
[0020] S1. Collect metering data from multiple heterogeneous data sources, pre-process the metering data, and form a quality-labeled data set containing quality grade labels; the metering data includes structured metering records generated by smart electricity meters, water meters, and gas meters, as well as system logs and device status data;
[0021] S2. Based on the quality-labeled dataset, the dynamic partition management module performs adaptive partition feature extraction based on time-varying features and low-overhead partition rebalancing to generate an updated partition scheme;
[0022] S3, based on the updated partitioning scheme, performs stream-batch collaborative processing on the real-time data stream and pre-stored historical batch data in the quality labeling dataset, and generates the final processing result through multi-level data fusion;
[0023] S4. Conduct multi-dimensional analysis and application of the final processing results to achieve storage, visualization and abnormal warning of measurement data.
[0024] This embodiment realizes efficient processing of the entire life cycle of metering data by integrating dynamic partition rebalancing and stream-batch collaborative metering data processing methods. First, metering data is collected from multi-source heterogeneous data sources and preprocessed, and quality grade tags are added so that the system can distinguish data of different quality levels and process them in a targeted manner, thereby improving the reliability of subsequent analysis. Secondly, based on the quality tag data set, adaptive partitioning and low-overhead partition rebalancing based on time-varying features are performed. The system can dynamically adjust the partitioning scheme according to the time-varying features of the data, adapt to the dynamic change characteristics of the metering data, avoid the data tilt and resource utilization imbalance caused by traditional static partitioning, and solve the problem that the existing technology adopts a one-time large-scale migration strategy, the system overhead is large, the service quality is significantly reduced, and the minimum interference path planning and incremental execution mechanism are lacking. Third, stream-batch collaborative processing is performed on real-time data streams and historical batch data, and the final processing results are generated through multi-level data fusion, which solves the problem of separation between stream processing and batch processing in traditional methods and ensures the balance between real-time and accuracy of metering data processing. Finally, the processing results are analyzed and applied in multiple dimensions to realize the all-round value mining of metering data. Through closed-loop optimization design, this embodiment cooperates with each other to form a complete set of metering data processing solutions, which can effectively cope with the challenges of large-scale, variable and time-sensitive data processing in the field of power metering, improve system throughput, reduce processing delays, and ensure data processing quality.
[0025] like Figure 2 As shown, according to one aspect of the present application, step S1 is further:
[0026] S11. According to the preset data collection rules, access multiple heterogeneous data sources and collect metering data to generate original metering data;
[0027] S12, performing protocol conversion and parsing, as well as data cleaning and standardization processing on the original metering data to obtain standardized metering data;
[0028] S13. Calculate the quality index of the standardized measurement data, and add a quality grade mark to each piece of standardized measurement data based on the quality index to form a quality mark data set.
[0029] In one embodiment of the present application, the system collects raw metering data from multiple data sources, including structured metering records generated by IoT devices such as smart electricity meters, water meters, and gas meters, as well as semi-structured data such as system logs and device status. A multi-protocol adapter is used to process data streams of different communication protocols (such as MQTT, CoAP, and HTTP) and convert them into standard data packets. The standard data packets are preliminarily parsed to extract metadata information (such as device ID, timestamp, data type) and load data to form an initial data set.
[0030] Process the initial data set for null values, and use different strategies (such as time series interpolation and statistical value filling) according to the data type to generate a data set without null values. Identify and remove outliers and duplicate values in the data set without null values, and output the cleaned data set. Convert the cleaned data set into a unified data model and measurement unit to generate a standardized data set.
[0031] Perform integrity check on the standardized data set and calculate the data integrity rate index. Perform accuracy check on the standardized data set and calculate the data accuracy rate index through physical model constraint verification. Perform timeliness evaluation on the standardized data set and calculate the data timeliness index. Generate a data quality score by combining the data integrity rate index, data accuracy index and data timeliness index, and mark the quality level of the standardized data set according to the score, and output the quality marked data set.
[0032] This embodiment solves the problems of diversified metering data sources, inconsistent standards, and unstable quality by constructing a multi-source heterogeneous data collection and quality marking system. The system can distinguish between high-quality data and low-quality data, and adopt differentiated processing strategies for different quality levels, which not only ensures the priority processing of high-quality data, but also does not completely discard the useful information that may be contained in low-quality data, ultimately improving the overall data processing efficiency and result reliability. This embodiment reduces the error rate of subsequent analysis, improves the utilization value of metering data, and lays a solid foundation for the accurate measurement and analysis of the power metering system.
[0033] like Figure 3 As shown, according to one aspect of the present application, step S2 is further:
[0034] S21, based on the quality labeled data set, performing adaptive partition feature extraction of time-varying features to obtain partition features;
[0035] S22, applying the partition feature to initially partition the data, generating an initial partition scheme, and obtaining a system partition status;
[0036] S23. Based on the pre-stored system operation status and load statistics, perform low-overhead partition rebalancing, update the initial partition scheme, and generate an updated partition scheme.
[0037] This embodiment solves the problems of data skew, low processing efficiency and unbalanced resource utilization caused by static partitioning in traditional metering data processing by implementing adaptive partition feature extraction and low-overhead partition rebalancing. The high system overhead of traditional rebalancing is reduced to a minimum, and through a sophisticated cost-benefit analysis, the impact on system performance is minimized while ensuring the improvement of partition balance, thereby achieving "low interference, high efficiency" partition adjustment. This solves the problem that the existing technology is based on simple correlation analysis rather than causal inference, and the actual impact of the rebalancing operation is inaccurately predicted, resulting in a high error rate in rebalancing decisions. Overall, this embodiment achieves a balance between the dynamic adaptability of data partitioning and system stability, which can respond to changes in data features in a timely manner and maintain stable system operation, thereby improving the throughput and response time of metering data processing, and is particularly suitable for scenarios such as power metering data with obvious time-varying characteristics.
[0038] According to one aspect of the present application, step S21 is further:
[0039] S211, inputting the quality-marked data set into a multi-scale time-varying feature decomposition unit to extract dominant feature components;
[0040] S212, performing feature sensitivity evaluation and screening based on the dominant feature components to generate a simplified feature set;
[0041] S213, constructing a dynamic evolution pattern recognizer using a simplified feature set, and outputting a pattern prediction result;
[0042] S214, performing adaptive feature fusion and weight allocation according to the pattern prediction result to generate weighted fusion features, i.e., partition features;
[0043] S215. Dynamically optimize the feature extraction strategy based on the weighted fusion features and output a feature extraction parameter set.
[0044] In one embodiment of the present application, a quality label data set is received, and a wavelet transform is applied to decompose the time series into components of different scales to obtain multi-scale time series components. Trend extraction, period analysis and noise evaluation are performed on the multi-scale time series components to form a time series feature component set. The energy proportion of each component in the time series feature component set is calculated, the dominant component is identified, and the dominant feature component is output.
[0045] Construct a correlation model between features and partition effects, input the dominant feature components and historical partition effect data, and output the feature-effect sensitivity matrix. Apply principal component analysis to reduce the dimension of the feature-effect sensitivity matrix, identify key feature combinations, and generate key feature sets. Perform redundancy analysis on the key feature sets, remove highly correlated features, and output a simplified feature set. Use the simplified feature set to construct a dynamic evolution pattern recognizer and output the pattern prediction results.
[0046] Receive the current data pattern identified by the pattern predictor, extract the appropriate feature weight configuration from the pattern knowledge graph, and output the initial weight configuration. Based on the current state of the system and the historical feature effectiveness evaluation, dynamically adjust the initial weight configuration and output the optimized weight configuration. Apply the optimized weight configuration to perform weighted fusion on the features in the streamlined feature set to generate weighted fusion features.
[0047] Collect partition effect feedback data, build a feature extraction effect evaluation model based on weighted fusion features, and output feature effect scores. Based on the feature effect scores, apply Bayesian optimization to adjust feature extraction parameters and output feature extraction parameter sets. Apply the feature extraction parameter sets to the next round of feature extraction process to form a closed-loop optimization.
[0048] This embodiment solves the problem of low efficiency of partitioning caused by the inability to cope with the time-varying characteristics of data in traditional metering data partitioning by implementing adaptive partitioning feature extraction of time-varying characteristics. This enables the feature extraction process to continuously improve itself, evolve with the changes in data characteristics, and maintain long-term effectiveness. Overall, this embodiment realizes the intelligence and adaptability of partitioning feature extraction, improves the accuracy and dynamic adaptability of metering data partitioning, and provides an efficient and stable partitioning basis for complex and changeable power metering data.
[0049] According to one aspect of the present application, step S213 is further:
[0050] S2131, receiving a simplified feature set, determining an optimal time lag parameter based on the mutual information minimum principle, determining an optimal embedding dimension in combination with an improved pseudo-nearest neighbor algorithm, reconstructing a one-dimensional time series into multi-dimensional phase space trajectory data; and generating a phase space density map using kernel density estimation;
[0051] S2132. Calculate the Lyapunov exponent, correlation dimension and entropy rate of phase space trajectory data based on phase space density diagram, identify the dynamic equation by combining sparse recognition method and deep learning, and obtain the dynamic model parameters by L1 regularization optimization;
[0052] S2133, construct recursive graphs based on phase space trajectory data and dynamic model parameters, analyze network topology characteristics, calculate distribution change rate in combination with Wasserstein distance, and identify multi-scale transition points;
[0053] S2134. Segment the time series according to multi-scale transition points, extract pattern features to construct a knowledge graph, train the graph neural network to learn the pattern transition rules, output a dynamic evolution pattern recognizer and generate pattern prediction results.
[0054] In one embodiment of the present application, an evolutionary pattern knowledge graph is constructed and applied, specifically: based on a multi-scale transition point set, the time series is divided into multiple intervals, each interval corresponds to an evolutionary pattern, and a pattern partition interval set is output. The feature representation of each pattern is extracted from the pattern partition interval set to form a pattern feature library. Construct an evolutionary pattern knowledge graph: Node: represents the identified typical evolutionary pattern, and the attributes contain features from the pattern feature library; Edge: represents the conversion relationship and probability between patterns, and outputs a pattern knowledge graph. Construct and train a graph neural network, input the pattern knowledge graph, learn the similarities and conversion rules between patterns, and output a pattern relationship model. Combine the pattern relationship model and the causal reasoning mechanism to achieve the ability to infer unseen patterns and output a pattern predictor.
[0055] This embodiment solves the problem of insufficient grasp of the evolution law of complex time series data in traditional metering data analysis by realizing dynamic evolution pattern recognition, and provides the ability to predict future data change trends. Discrete data patterns are connected into an organic knowledge network to achieve the ability to reason about unseen patterns. Overall, this embodiment provides unprecedented prediction depth for metering data analysis, enabling the system to "foresee the future", prepare resources and adjust strategies in advance, and improve the foresight and adaptability of the power metering system.
[0056] According to one aspect of the present application, step S2131 is further:
[0057] S21311, receiving a simplified feature set, and determining an optimal time lag parameter using an adaptive time lag parameter calculation method (based on the mutual information minimum principle);
[0058] S21312. Determine the optimal embedding dimension based on the optimal time lag parameter and the improved pseudo nearest neighbor algorithm;
[0059] S21313, using the optimal time lag parameter and the optimal embedding dimension, reconstruct the one-dimensional time series into multi-dimensional phase space trajectory data;
[0060] S21314. Use the kernel density estimation method to calculate the probability density distribution of the phase space trajectory data and output the phase space density map.
[0061] Using the optimal time lag parameter and the optimal embedding dimension, the one-dimensional time series X(t) is reconstructed into an m-dimensional phase space: X(t) → [X(t), X(t+τ), X(t+2τ), ..., X(t+(m-1)τ)], generating phase space trajectory data, where τ is the optimal time lag parameter and m is the optimal embedding dimension.
[0062] According to one aspect of the present application, step S2132 is further:
[0063] Calculate the Lyapunov exponents (used to quantify the degree of chaos in the system) of the phase space trajectory data and output a set of Lyapunov exponents.
[0064] Computes the correlation dimension of phase space trajectory data (used to quantify trajectory complexity) and outputs the correlation dimension value.
[0065] Calculate the entropy rate of phase space trajectory data (used to quantify the rate of information generation) and output the entropy rate value.
[0066] Combining the sparse identification method (SINDy) and deep learning, the implicit dynamic equation is identified from the phase space trajectory data: dX / dt = F(X) = Ξ(X)·Θ where X is the system state vector, F(X) is the dynamic function, Ξ(X) is the candidate function library, and Θ is the coefficient vector.
[0067] The L1 regularized optimization is used to solve Θ, obtain the sparsely represented dynamic equation, and output the dynamic model parameters. Multimodal transition point detection and classification are performed based on the dynamic model parameters.
[0068] This embodiment solves the problem that traditional time series analysis methods are difficult to reveal the intrinsic dynamic structure and complex behavior mechanism of data through phase space reconstruction and topological invariant extraction technology, and realizes in-depth mining of the essential characteristics of metering data. It realizes the revelation and mathematical expression of the intrinsic mechanism of metering data, transforms traditional black box analysis into transparent model analysis, improves the depth of understanding and prediction accuracy of the behavior of power metering system, and provides a theoretical basis for system optimization and abnormal diagnosis.
[0069] According to one aspect of the present application, step S2133 is further:
[0070] S21331. Construct a recursive graph (RP) based on the phase space trajectory data: RP(i, j) = Θ(ε - ||X(i) - X(j)||); where Θ is the Heaviside function, ε is the distance threshold, X(i) is the state vector of the phase space trajectory data at time point i, and output the recursive graph matrix.
[0071] S21332. Analyze the network topological characteristics of the recursive graph matrix, calculate indicators such as node centrality and clustering coefficient, and form a topological characteristic vector.
[0072] S21333. Calculate the rate of change of the probability distribution in the continuous time window based on the Wasserstein distance, identify the distribution mutation points, and output the distribution mutation point set.
[0073] S21334. Combining the changing trend of the topological feature vector and the distribution mutation point set, a hierarchical attention mechanism is applied to identify transition points of different scales and output a multi-scale transition point set.
[0074] This embodiment solves the problem of the difficulty in accurately identifying system state change points and predicting conversion trends in traditional metering data analysis through multimodal conversion point detection and classification technology. This embodiment can identify conversion points at different time scales, solving the problem that single-scale analysis cannot take into account both macro trends and micro changes. Overall, this embodiment realizes the accurate identification and classification of state transitions in metering data, and improves the accuracy of conversion point detection from 70% of traditional methods to more than 90%, providing key technical support for predictive maintenance and abnormal warning of power metering systems.
[0075] According to one aspect of the present application, step S22 is further:
[0076] S221. Based on the weighted fusion features and the current system load, design a data partitioning strategy, determine the number of partitions, partition keys and partition boundaries, and output the partitioning strategy solution.
[0077] S222. Evaluate the difference between the partition strategy solution and the current partition status, calculate the migration cost, and output the partition adjustment cost.
[0078] S223. Determine whether to perform partition adjustment based on the partition adjustment cost and system tolerance, and output the partition execution decision.
[0079] S224: If the partition execution decision is to execute, then apply the new partition strategy and update the system partition status.
[0080] This embodiment determines the number of partitions, partition keys and partition boundaries by designing a reasonable data partitioning strategy, making data storage and query more efficient and reducing the waste of system resources. By evaluating the difference between the partitioning strategy scheme and the current partitioning state, the migration cost is calculated, so that the overhead and system downtime caused by data migration can be minimized when performing partition adjustments. According to the partition adjustment cost and system tolerance, it is decided whether to perform partition adjustment to ensure that the system can flexibly adapt to changes under different load conditions and maintain efficient operation. By applying a new partitioning strategy and updating the system partition state, the performance and response speed of the system can be improved to meet the growth of business needs. By dynamically adjusting the partitioning strategy, the system can be adjusted according to actual conditions, providing more accurate decision support and optimizing resource allocation. This embodiment can improve data processing efficiency, reduce migration costs, enhance system flexibility, and optimize system performance through intelligent data partitioning strategy design and adjustment, thereby improving the overall decision support capability.
[0081] According to one aspect of the present application, step S23 is further:
[0082] S231, collecting load statistics of each partition, including data volume, access frequency and computing resource usage; calculating the imbalance index based on the load statistics, analyzing the data access pattern and identifying hotspot partitions, building a multi-dimensional imbalance assessment model, and outputting a comprehensive imbalance index;
[0083] S232. Construct a rebalancing benefit model based on comprehensive imbalance indicators, build an environment-aware weight adjustment mechanism based on system partition status and system operation status, predict rebalancing costs and perform uncertainty modeling, and optimize rebalancing strategies;
[0084] S233. Build a data dependency graph according to the optimized rebalancing strategy, calculate the shard association graph and build a migration priority algorithm to generate a migration path plan with minimal interference;
[0085] S234, performing incremental partition data migration according to the migration path plan, monitoring the execution status in real time and dynamically adjusting the migration parameters;
[0086] S235. Based on the adjusted migration parameters, evaluate the degree of improvement of the comprehensive imbalance index before and after rebalancing, calculate the actual resource consumption cost, update the rebalancing decision model parameters and generate an updated partitioning scheme.
[0087] In one embodiment of the present application, the load statistics of each partition are collected, including indicators such as data volume, access frequency, and computing resource usage. The imbalance index of data distribution is calculated: Imbalance = (σ / μ) × 100%; where σ is the standard deviation of the load of each partition, μ is the average value of the load of each partition, and the load imbalance is output. The data access pattern is analyzed, the hotspot partitions are identified, the access ratio of the hotspot partitions is calculated, and the hotspot concentration is output. The load imbalance and hotspot concentration are combined to construct a multi-dimensional imbalance assessment model, and the comprehensive imbalance index is output.
[0088] Conduct rebalancing effect evaluation and feedback, specifically: compare the comprehensive imbalance indicators before and after rebalancing, calculate the degree of improvement, and output the balance improvement. Monitor the resource consumption and system performance impact of the rebalancing process, calculate the actual cost, and output the actual rebalancing cost. Compare the actual rebalancing cost with the predicted cost estimate C, evaluate the accuracy of the cost model, and output the cost model error. Combine the balance improvement and the actual rebalancing cost to calculate the actual benefit of the rebalancing operation and output the actual benefit value. Based on indicators such as the actual benefit value and the cost model error, update the rebalancing decision model parameters and output the model update parameter set to guide the next rebalancing decision.
[0089] This embodiment solves the problem of high system overhead and service interruption in the traditional rebalancing process by implementing low-overhead partition rebalancing, and realizes "low interference, high efficiency" partition adjustment. It enables the system to continuously accumulate experience, optimize rebalancing decisions, and improve long-term operating efficiency. Overall, this embodiment achieves high efficiency and low impact of partition adjustment, which is particularly suitable for key business scenarios that require high availability, such as power metering.
[0090] According to one aspect of the present application, step S232 is further:
[0091] S2321. Construct a multi-objective rebalancing benefit function based on comprehensive imbalance indicators and output equilibrium benefit value and dynamic weight vector;
[0092] S2322. Combine the pre-stored current system load situation, and perform rebalancing cost prediction based on causal inference to obtain a cost estimate and a state change prediction;
[0093] S2323. Based on the equilibrium benefit value and cost estimate, uncertainty-aware benefit prediction is performed, and risk-adjusted benefit and quantum-optimized decision boundary are output;
[0094] S2324. Based on risk-adjusted returns, quantum-optimized decision boundaries, and state change predictions, the rebalancing strategy is optimized through reinforcement learning, and the optimal rebalancing strategy is output.
[0095] This embodiment constructs a multi-objective rebalancing benefit function through comprehensive imbalance indicators, dynamically adjusts the system weights, and enables the system to maintain a high degree of balance when the load and priority change. Combined with the current system load situation, causal inference is used to predict the rebalancing cost, effectively evaluate and reduce the cost of rebalancing operations, and improve the overall system efficiency. Based on the equilibrium benefit value and cost estimate, uncertainty-aware benefit prediction is performed, and risk-adjusted benefits and quantum optimized decision boundaries are output, thereby improving the accuracy and reliability of benefit prediction. Through the reinforcement learning algorithm, based on risk-adjusted benefits, quantum optimized decision boundaries and state change predictions, the rebalancing strategy is optimized to generate the optimal rebalancing solution, thereby improving the stability and performance of the system. This embodiment optimizes the rebalancing strategy, improves the system balance and benefits, reduces the rebalancing cost, improves the accuracy of benefit prediction, and ultimately achieves stable and efficient operation of the system.
[0096] According to one aspect of the present application, step S2321 is further:
[0097] S23211. Construct a balanced return model after rebalancing based on comprehensive imbalance indicators:
[0098] B = f(I_current, I_expected);
[0099] Where I_current is the current comprehensive imbalance index, I_expected is the expected imbalance index, f() is a function that outputs the equilibrium benefit value B;
[0100] S23212. Combine the current system status (peak / trough period) information to build an adaptive weight adjustment mechanism for environmental perception:
[0101] W(t) = W_base + ΔW(load(t), priority(t));
[0102] Where W(t) is the weight vector at time t, W_base is the base weight, load(t) is the system load at time t, priority(t) is the service priority, ΔW is the weight change, and the dynamic weight vector W is output;
[0103] S23213. Build a time series forecasting model to predict the system load trend in the future period and obtain the load forecast value;
[0104] S23214. Adjust the dynamic weight vector W based on the load prediction value to obtain a pre-adjusted weight vector;
[0105] S23215. Combine the equilibrium benefit value B with the pre-adjusted weight vector to construct a complete multi-objective rebalancing benefit function:
[0106] E = α·B - β·C - γ·D;
[0107] Among them, α, β, γ are weight coefficients from the pre-adjusted weight vector, B is the equilibrium benefit value, C is the rebalancing cost, and D is the service interference degree.
[0108] This embodiment constructs a balance benefit model based on comprehensive imbalance indicators and dynamically adjusts the system weights, so that the system can maintain a high degree of balance under different loads and priorities. By comprehensively using imbalance indicators, system status, load prediction, and reinforcement learning, etc., intelligent and efficient rebalancing strategy optimization is achieved, which helps to improve the balance, flexibility and overall benefits of the system.
[0109] According to one aspect of the present application, step S2322 is further:
[0110] S23221. Establish a fine-grained cost model for the rebalancing operation, including: C_data: data migration bandwidth cost; C_index: index reconstruction computing cost; C_query: query performance impact cost; C_io: I / O load increase cost; the total cost C = C_data + C_index + C_query + C_io is obtained, and the cost estimate C is output.
[0111] S23222. Construct a structural causal model (SCM) of system state variables: S = f(X, do(R)); where S is the system state vector, X is the influencing factor vector, do(R) represents the intervention of performing the rebalancing operation R, and outputs a causal structure diagram.
[0112] S23223. Use the causal structure diagram to calculate the difference in system state between performing rebalancing and not performing rebalancing: ΔS = E[S|do(R=1)] - E[S|do(R=0)], and output the state change prediction ΔS, where E represents the expectation function.
[0113] S23224. Based on the state change prediction ΔS, accurately quantify the cost of each system component caused by rebalancing, and update the cost estimate C. Perform uncertainty-aware benefit prediction and optimization of rebalancing strategy based on reinforcement learning.
[0114] This embodiment solves the problem of prediction bias caused by traditional cost models that only consider surface correlations but ignore deep causal relationships through rebalancing cost prediction technology based on causal inference, and improves the accuracy and interpretability of cost prediction. It enables the prediction model to learn and improve from experience and maintain prediction accuracy for a long time. Overall, this embodiment achieves accurate prediction of the intervention effects of complex systems, increases the cost prediction accuracy from 75% of traditional methods to more than 90%, and provides precise cost control capabilities for the economical and efficient operation of the power metering system.
[0115] According to one aspect of the present application, step S2323 is further:
[0116] S23231. Establish a Bayesian network to represent the probabilistic dependency between system status and rebalancing benefits, input the current system partition status and historical data, and output the Bayesian network model.
[0117] S23232. Using the Bayesian network model, design a Monte Carlo simulation method to generate multiple possible rebalancing result scenarios and output the return distribution data.
[0118] S23233. Based on the return distribution data, calculate the risk-adjusted expected return: RAR = E[B] - λ × σ(B); where E[B] is the expected value of the return, σ(B) is the standard deviation of the return, and λ is the risk aversion coefficient. Output the risk-adjusted return RAR.
[0119] S23234. Introduce the quantum probability theory framework to deal with complementary uncertainties that are difficult to express in traditional probability models and construct quantum probability representation.
[0120] S23235. Based on quantum probability representation, design a quantum information geometry optimization method, determine the optimal decision boundary, and output the quantum optimized decision boundary.
[0121] This embodiment uses uncertainty-aware revenue prediction technology to solve the problem of over-certainty in future revenue prediction and neglect of risks and uncertainties in traditional rebalancing decisions, thereby improving the robustness and reliability of decisions. It achieves comprehensive assessment and control of rebalancing decision risks, reduces the decision error rate from 15% in traditional methods to less than 5%, and provides decision-making guarantee for the robust operation of the power metering system.
[0122] According to one aspect of the present application, step S2324 is further:
[0123] S23241. Construct a reinforcement learning environment for rebalancing decisions: state space S: includes indicators such as load imbalance, hotspot concentration, and system resource utilization; action space A: {execute rebalancing, adjust rebalancing parameters, delay rebalancing}; reward function R: calculates the output RL environment model based on the multi-objective benefit function E = α·B -β·C -γ·D.
[0124] S23242. Design a multi-agent reinforcement learning architecture, configure independent decision-making agents for different data partitions, and form an agent network.
[0125] S23243, Use Graph Attention Network (GAT) to build inter-agent communication mechanism: h_i (l+1) =σ(∑_j∈N(i)α_ij W h_j l ); where h_i l is the state representation of the ith agent in the lth layer, α_ij is the attention weight, W is the parameter matrix, and the output agent communication protocol.
[0126] S23244. Train the intelligent agent network, optimize the rebalancing decision strategy, and output the optimal rebalancing strategy.
[0127] This embodiment solves the problem of traditional rebalancing decision-making that relies on manual experience and is difficult to adapt to complex dynamic environments through the rebalancing strategy optimization technology based on reinforcement learning, and realizes the automation and intelligence of decision-making. It realizes the intelligence and adaptability of rebalancing decision-making, reduces the decision-making time from minutes to seconds in traditional methods, and improves the quality of decision-making, increases the load balancing degree by more than 30%, and provides intelligent decision-making support for the automated operation and maintenance of the power metering system.
[0128] According to one aspect of the present application, step S233 is further:
[0129] S2331. Based on the optimal rebalancing strategy, determine the data shards that need to be migrated and form a migration task set.
[0130] S2332. Build a data dependency graph, analyze the access relationship between shards, identify highly correlated shards, and output a shard association graph.
[0131] S2333. Based on the shard association graph, design a migration priority algorithm to ensure that shards with high correlation are migrated in the same batch as much as possible, and output the migration batch plan.
[0132] S2334. Evaluate the system impact of each migration path, select the path with minimum interference, and generate a migration path plan.
[0133] Step S234 is further as follows:
[0134] S2341. According to the migration path plan and the migration batch plan, split the rebalancing task into multiple incremental steps and output the incremental execution plan.
[0135] S2342. Monitor the execution status of each incremental step in real time and collect execution status data.
[0136] S2343. Based on the execution status data, dynamically adjust the execution parameters of subsequent incremental steps, such as migration rate, parallelism, etc., and output the adjusted execution parameters.
[0137] S2344. Apply the adjusted execution parameters and continue to perform incremental rebalancing until all migration task sets are completed.
[0138] This embodiment solves the problem of system performance degradation and service interruption caused by large-scale data migration in the traditional rebalancing process through minimum interference rebalancing path planning and incremental rebalancing execution technology, and realizes a smooth and imperceptible rebalancing process. It realizes "imperceptible" data rebalancing, reduces the performance degradation in the traditional rebalancing process from more than 30% to less than 5%, and ensures data consistency and service continuity, providing key technical support for high-availability operation and maintenance of power metering systems.
[0139] like Figure 4 As shown, according to one aspect of the present application, step S3 is further:
[0140] S31, building a unified stream-batch processing model, applying the updated partitioning scheme to the real-time data stream and pre-stored historical batch data in the quality-labeled dataset;
[0141] S32, performing incremental feature update and real-time anomaly detection on the real-time data stream to generate real-time processing results;
[0142] S33, performing time series analysis and quality verification on historical batch data to obtain verified batch processing results;
[0143] S34. Perform multi-level data fusion based on the real-time processing results and the verified batch processing results to generate the final processing results.
[0144] In one embodiment of the present application, a real-time data stream in a quality mark data set is received, and real-time distribution is performed according to the system partition state, and the partition data stream is output. Sliding window processing is applied to the partition data stream, statistics within the window are calculated, and window statistics results are output. Key events in the partition data stream are processed based on an event-driven model, and event processing results are output. The window statistics results and event processing results are integrated to generate real-time processing results.
[0145] Load historical measurement data from the data warehouse, reorganize the data using the system partition status, and output batch data sets. Execute complex analysis algorithms on batch data sets, such as time series analysis and pattern recognition, and output batch results. Perform quality verification on batch results to ensure the accuracy of the results, and output the verified batch results.
[0146] Determine the overlapping time window of the real-time processing results and the batch processing results after verification, and output the result overlapping interval. Compare the differences between the two results in the result overlapping interval, calculate the consistency index, and output the consistency score. Based on the consistency score and predetermined rules, decide which strategy to use to fuse the results, and output the fusion strategy. Apply the fusion strategy, combine the real-time processing results and the batch processing results after verification, and generate the final processing result.
[0147] This embodiment solves the problem of balancing real-time performance and accuracy caused by the separation of stream processing and batch processing in traditional metering data processing by building a unified stream-batch processing model and realizing stream-batch collaborative processing. It can simultaneously meet the dual needs of real-time metering monitoring and in-depth power consumption behavior analysis, and provide more comprehensive and timely data support for power companies.
[0148] like Figure 5 As shown, according to one aspect of the present application, step S4 is further:
[0149] S41, organizing the final processing results according to a predetermined data model, writing them into a persistent storage module, obtaining persistent result data and establishing a multidimensional index;
[0150] S42, performing multi-dimensional statistical analysis based on the persistent result data to generate a visualization resource set;
[0151] S43. Perform abnormal identification on the final processing result according to the preset abnormal rules, and trigger corresponding warning information.
[0152] In one embodiment of the present application, the final processing result is organized according to a predetermined data model, written to persistent storage, and persistent result data is output. A multidimensional index is established for the persistent result data, query performance is optimized, and the index structure is output. The data lifecycle management strategy is implemented, the persistent result data is hierarchically stored and archived, and the hierarchical storage structure is output.
[0153] Perform multi-dimensional statistical analysis based on persistent result data, such as time dimension, space dimension, user dimension, etc., and output multi-dimensional analysis results. Generate various data visualization charts, such as trend charts, distribution charts, association charts, etc., and output visualization resource sets. Build an interactive data exploration interface, integrate visualization resource sets, support in-depth analysis, and output a data exploration interface.
[0154] Based on the final processing results, anomaly detection algorithms are applied to identify abnormal patterns and output anomaly candidate sets. The anomaly candidate sets are verified and classified, false positives are filtered, and confirmed anomaly sets are output. According to the severity of the confirmed anomaly set, graded warning information is generated and a warning message set is output. The warning message set is pushed to relevant personnel through pre-configured notification channels, and the processing status is recorded to form a closed-loop management.
[0155] This embodiment solves the problem of insufficient data value mining in traditional metering data processing by realizing multi-dimensional analysis and application of metering data, and maximizes the commercial value and decision-making support capabilities of metering data. It enables the system to discover and deal with problems before they expand, transforming passive response into active prevention, and improving the operational reliability and safety of the power metering system. Overall, this embodiment transforms metering data from a simple record of numbers into an actionable basis for decision-making, providing strong data support for the refined management, energy conservation and emission reduction, and optimized operation of power companies.
[0156] In summary, the present invention collects multi-source heterogeneous metering data and adds quality tags; performs adaptive partition feature extraction and low-overhead partition rebalancing based on time-varying features to generate an optimized partitioning scheme; performs stream-batch collaborative processing on real-time data streams and historical batch data; and performs multi-dimensional analysis and application of the processing results. To address the problem of time-varying feature extraction, traditional methods usually use simple time window statistics or basic spectrum analysis, which lack multi-scale decomposition capabilities. The present invention uses multi-scale time-varying feature decomposition, decomposes the time series into components of different scales through wavelet transform, extracts dominant features, and generates a streamlined feature set through feature sensitivity evaluation, thereby improving the accuracy and adaptability of feature extraction. To address the problem of dynamic evolution pattern recognition, existing technologies mostly rely on simple statistical models, while the present invention constructs a complete evolution pattern knowledge graph through phase space reconstruction, topological invariant extraction, and multi-modal transition point detection, thereby achieving a deep understanding and prediction of data behavior. For the partition imbalance assessment and rebalancing decision-making problems, traditional methods usually use a single indicator and a fixed threshold. The present invention constructs a multi-dimensional imbalance assessment model, combines rebalancing benefit-cost modeling and reinforcement learning-based policy optimization, and realizes a scientific and accurate decision-making mechanism, solving the problem that the existing technology is difficult to deal with multi-source uncertainties and risks in complex environments and lacks a robust decision-making mechanism. For the rebalancing execution problem, the existing technology usually adopts a one-time migration strategy. The present invention constructs a minimum interference rebalancing path planning and an incremental rebalancing execution mechanism, which reduces system interference and ensures service continuity. For the cost-benefit analysis problem, the traditional model is based on simple correlation analysis, while the present invention adopts a structural causal model and uncertainty-aware benefit prediction to improve prediction accuracy and decision robustness.
[0157] Taking the smart meter data processing of a provincial power company as an example, the system collects multi-source heterogeneous metering data including smart meter records, equipment status data and system logs.
[0158] First, the system processes data streams of different communication protocols through a multi-protocol adapter according to the preset data collection rules. Smart meter data is mainly transmitted using the DLT645-2007 protocol, device status data uses the MQTT protocol, and system logs use the HTTP protocol. The system converts all data into standard JSON format data packets:
[0159] { "device_id": "E10056782", "timestamp": "2024-02-15T08:30:00Z", "data_type": "power_reading", "value": 5.63, "unit": "kWh"};
[0160] Parse the standard data packets, extract metadata information and load data, and form the initial data set. Then perform data cleaning, including: processing missing values: time series data (such as meter readings) are filled using linear interpolation; outlier detection: use the Z-score method to identify outliers, and the rule is |Z| > 3 is marked as anomaly; duplicate value removal: identify duplicate records based on the combination of device ID and timestamp;
[0161] After cleaning, the data is standardized and different measurement units (such as kWh, W, etc.) are converted into a standard unit system.
[0162] Then calculate the data quality indicators: data completeness index CI = (1 - number of missing values / total number of records) × 100%; data accuracy index AI = (1 - number of outliers / total number of records) × 100%; data timeliness index TI = (1 - average delay time / maximum acceptable delay) × 100%; comprehensive quality score QS = 0.4×CI + 0.4×AI+ 0.2×TI;
[0163] According to the comprehensive quality score QS, the quality level of each data is marked: QS ≥ 90%: high quality (High); 70% ≤ QS < 90%: good (Medium); QS < 70%: low quality (Low). Finally, a quality-marked dataset containing quality level marks is formed to provide data quality perception capabilities for subsequent processing.
[0164] By performing adaptive extraction of time-varying features on the quality tag dataset, the characteristics of the data changing over time are captured. Specifically, the quality tag dataset is input into the multi-scale time-varying feature decomposition unit. Taking the 24-hour power time series X(t) of a single user as an example, the system uses discrete wavelet transform (DWT) for multi-scale decomposition: X(t) = A_J(t) + ∑ j=1 J D_j(t);
[0165] Where: X(t) is the original time series; A_J(t) is the J-level approximate component, which represents the low-frequency trend of the series; D_j(t) is the j-level detail component, which represents the fluctuations of different frequency scales; J is the decomposition level, in this example, J=5;
[0166] For the power metering data, the Daubechies-4 wavelet basis function is selected for transformation because it performs well in capturing the characteristics of power load. After the transformation, an approximate component A_5 and five detail components D_1 to D_5 are obtained, corresponding to the characteristics of different time scales: D_1: corresponding to rapid fluctuations within 15 minutes; D_2: corresponding to fluctuations of 15-30 minutes; D_3: corresponding to fluctuations of 30 minutes to 1 hour; D_4: corresponding to fluctuations of 1-2 hours; D_5: corresponding to fluctuations of 2-4 hours; A_5: corresponding to long-term trends of more than 4 hours;
[0167] Calculate the energy proportion of each component: E_i = (∑ t |C_i(t)|2) / (∑ i ∑ t |C_i(t)|2); where C_i(t) represents the coefficient value of the i-th component (A_J or D_j) at time t.
[0168] The dominant component is identified based on the energy share, and the rule is: the component with an energy share of more than 10% is considered the dominant component. For the typical load curve of residential users, A_5, D_3 and D_4 are usually the dominant components, indicating that the long-term trend and the fluctuation characteristics of 30 minutes to 2 hours play a major role in the power load.
[0169] Perform feature sensitivity evaluation based on the dominant feature components. Build a model of the association between features and partition effects:
[0170] Extract statistical features from the dominant components, including mean, standard deviation, skewness, kurtosis, maximum, minimum, etc.
[0171] For each dominant component A_5, D_3 and D_4, the following features are calculated: Mean μ_i = (1 / N) * ∑ t=1 NC_i(t); standard deviation σ_i = sqrt((1 / N) * ∑ t=1 N (C_i(t) - μ_i)2); Skewness skew_i = (1 / N) * ∑ t=1 N ((C_i(t) - μ_i) / σ_i) 3 ; Kurtosis kurt_i = (1 / N) * ∑ t=1 N ((C_i(t) - μ_i) / σ_i) 4 ; Maximum value: max_i = max(C_i(t)); Minimum value: min_i = min(C_i(t)).
[0172] Calculate the feature-effect sensitivity matrix S_{i,j} = dP_j / dF_i; where: P_j is the j-th partition effect indicator (such as load balancing, query response time); F_i is the i-th feature; S_{i,j} represents the sensitivity of the effect indicator P_j to the feature F_i; apply principal component analysis (PCA) to reduce the dimension of the feature-effect sensitivity matrix, retain the principal components with an explained variance ratio of more than 85%; calculate the redundancy of the feature combination corresponding to the principal component, and remove highly correlated features with a Pearson correlation coefficient of more than 0.85.
[0173] After analysis, the identified streamlined feature set includes: the mean and standard deviation of the long-term trend component (A_5); the standard deviation and kurtosis of the medium-term fluctuation component (D_3); the skewness and kurtosis of the short-term and medium-term fluctuation component (D_4); these features can effectively capture the time-varying characteristics of the power load and provide a basis for subsequent partitioning.
[0174] For each time series feature in the streamlined feature set, the phase space is reconstructed. Take the mean sequence of the long-term trend component A_5 as an example:
[0175] The optimal time lag parameter τ is calculated and determined using the mutual information minimum principle: I(X(t), X(t+τ)) = ∑_{x(t),x(t+τ)} p(x(t),x(t+τ)) * log(p(x(t),x(t+τ)) / (p(x(t))*p(x(t+τ)))); where: I(X(t), X(t+τ)) is the mutual information function; p(x(t),x(t+τ)) is the joint probability distribution; p(x(t)) and p(x(t+τ)) are the marginal probability distributions.
[0176] Calculate the mutual information under different τ values and take the first local minimum point as the optimal time lag parameter. For a typical daily load curve, τ is usually 5-6 hours.
[0177] Determine the optimal embedding dimension m: Use the improved false nearest neighbor algorithm (FNN) to calculate the embedding dimension. The core of the algorithm is to calculate the false nearest neighbor ratio under different dimensions: FNN(m) = ∑_{i=1}^{N-mτ} Θ(R_i(m+1) / R_i(m) - R_threshold) / (N-mτ)
[0178] Where: Θ is the Heaviside step function; R_i(m) is the distance from point i to its nearest neighbor in the m-dimensional space; R_threshold is the threshold, usually set to 10;
[0179] The value of m when FNN(m) < 0.01 is taken as the optimal embedding dimension. For power load data, m is usually 3-4.
[0180] Using the optimal time lag parameter τ = 6 hours and the optimal embedding dimension m = 4, the time series X(t) is reconstructed into a 4-dimensional phase space: X(t) → [X(t), X(t+6), X(t+12), X(t+18)]; the phase space trajectory data is generated.
[0181] Use kernel density estimation (KDE) to calculate the phase space density f(x) = (1 / nh) * ∑_{i=1}^n K((x-xi) / h); where: K is the Gaussian kernel function; h is the bandwidth parameter, determined using the Silverman rule; x_i is the sample point
[0182] Calculate dynamic characteristic indicators based on phase space trajectory data:
[0183] Calculate the maximum Lyapunov exponent λ = lim_{t→∞} (1 / t) * log(||δZ(t)|| / ||δZ(0)||); where: δZ(t) is the trajectory separation vector at time t; δZ(0) is the initial separation vector;
[0184] For residential users, λ is usually in the range of 0.02-0.05, indicating that the system is weakly chaotic.
[0185] Calculate the correlation dimension D2 = lim_{r→0} (log(C(r)) / log(r));
[0186] Where: C(r) is the correlation integral, C(r) = (2 / N(N-1)) * ∑ i,j=1,i≠j N Θ(r-||x_i-x_j||); Θ is the Heaviside step function; for the power load data, D_2 is usually in the range of 2.3-2.8, reflecting the geometric complexity of the system.
[0187] Calculate the entropy rate h = ∑ i=1 k λ_i (λ_i > 0); where λ_i is the Lyapunov index spectrum of the system.
[0188] For electrical loads, h is typically between 0.03 and 0.07, reflecting the information generation rate of the system.
[0189] The sparse identification method (SINDy) was used to identify the kinetic equation dX / dt = F(X) = Ξ(X)·Θ from the data.
[0190] Where: X is the system state vector; Ξ(X) is the candidate function library, including polynomials, trigonometric functions, etc.; Θ is the coefficient vector;
[0191] Solve Θ by L1 regularization: min || X * - Ξ(X)Θ||_2^2 + α||Θ||_1; where α is the regularization parameter, which is 0.1 in this example.
[0192] For a typical residential electrical load, the simplified dynamic equations identified might look like this: dx1 / dt = 0.05x1- 0.02x1x2 + 0.01x3; dx2 / dt = 0.03x2 + 0.04x1 - 0.01x2²; dx3 / dt = -0.02x3 +0.03x1x2 + 0.01sin(x1);
[0193] Construct a recurrence graph based on phase space trajectory data and dynamic model parameters and detect transition points:
[0194] Construct a recursive graph RP(i,j) = Θ(ε - ||X(i) - X(j)||); where: Θ is the Heaviside function; ε is the distance threshold, which is 10% of the phase space diameter; X(i) is the state vector of the phase space trajectory at time point i. The generated recursive graph matrix is a binary matrix, where 0 indicates that the distance between two state points is greater than the threshold, and 1 indicates that the distance is less than the threshold.
[0195] Analyze the topological characteristics of the recursive graph: Calculate the node centrality: C_i = ∑_j RP(i,j) / N; Calculate the clustering coefficient: CC_i = ∑_{j,k} RP(i,j)·RP(j,k)·RP(k,i) / ∑_{j,k} RP(i,j)·RP(i,k);
[0196] Generate a topological characteristic vector [C_i, CC_i], which reflects the dynamic state of the system at different time points.
[0197] Calculate the Wasserstein distance to measure distribution changes: For consecutive time windows W_1 and W_2, calculate the Wasserstein distance between the phase space probability distributions P_1 and P_2: W(P_1, P_2) = inf_{γ∈Γ(P_1,P_2)} ∫∫ ||xy||_2 dγ(x,y); where Γ(P_1,P_2) is the set of joint distributions of all marginal distributions P_1 and P_2.
[0198] Calculate the Wasserstein distance sequence of the sliding window and identify the distance sudden increase points as distribution mutation points.
[0199] Combining the changing trend and distribution mutation points of the topological feature vector, a hierarchical attention mechanism is used to identify multi-scale transition points: intra-day transition points, which usually correspond to changes in electricity consumption patterns within a season; seasonal transition points, which correspond to changes in electricity consumption patterns caused by seasonal alternations; and annual transition points, which correspond to the evolution of long-term electricity consumption behavior.
[0200] The time series is divided into different intervals based on multi-scale transition points, each interval corresponds to an evolution mode:
[0201] Extract the characteristic representation of each mode to form a mode feature library, including: dynamic parameters: Lyapunov index, correlation dimension, entropy rate, etc.; statistical characteristics: mean, variance, kurtosis, etc.; spectrum characteristics: main frequency components and their amplitudes.
[0202] Construct an evolutionary pattern knowledge graph, including: nodes representing the typical evolutionary patterns identified, with attributes containing features from the pattern feature library, and edges representing the conversion relationships and probabilities between patterns.
[0203] For residential users, typical modes include: "Normal working day mode": obvious peaks in the morning and evening, and troughs at noon and late at night. "Holiday mode": relatively stable load throughout the day, with no obvious peaks. "Seasonal transition mode": load gradually increases or decreases. "Extreme weather mode": abnormally high or abnormally low load.
[0204] Using the graph attention network (GAT) to process the pattern knowledge graph, the node feature update formula is: i l+1 = σ(∑ j∈N(i) α_ij W h j l );where: h j lis the feature representation of the ith node in the lth layer; α_ij is the attention coefficient; W is the weight matrix; σ is the activation function. The model is trained by minimizing the prediction error. The final output dynamic evolution pattern recognizer can: identify the evolution mode to which the current data belongs; predict the probability and time point of mode transition; and infer the characteristics of unseen patterns.
[0205] The identifier provides forward-looking data behavior predictions for partition rebalancing, enabling the system to adjust partition strategies in advance.
[0206] This embodiment performs low-overhead partition rebalancing based on the results of dynamic evolution pattern recognition to achieve dynamic optimization of data distribution.
[0207] First, collect the load statistics of each partition and build a multi-dimensional imbalance assessment model:
[0208] Calculate the imbalance index of data distribution: Imbalance = (σ / μ) × 100%; where: σ is the standard deviation of the load of each partition; μ is the average value of the load of each partition;
[0209] For example, at a certain moment, the data volume of the five partitions is [120, 85, 180, 60, 155] GB respectively, then: μ = 120 GB, σ = 46.04 GB, Imbalance = 38.37%.
[0210] Analyze data access patterns and calculate hotspot concentration Hotspot = (∑ i∈H v_i) / (∑ i=1 n v_i)
[0211] Where: H is the set of hotspot partitions (partitions with the top 20% of visit volume); v_i is the number of visits to the i-th partition.
[0212] The query volume per second of the 5 partitions is [320, 150, 450, 130, 250] times, so the hotspot partition is the 3rd partition, Hotspot = 450 / (320+150+450+130+250) = 34.61%
[0213] Calculate resource utilization imbalance ResImbalance = (σ_r / μ_r) × 100%;
[0214] Among them: σ_r is the standard deviation of resource utilization of each partition; μ_r is the average value of resource utilization of each partition;
[0215] The CPU utilization of the five partitions is [75%, 45%, 85%, 40%, 65%], then: μ_r = 62%, σ_r = 18.65%, ResImbalance = 30.08%.
[0216] Construct the comprehensive imbalance index I_comprehensive = w_1×Imbalance + w_2×Hotspot +w_3×ResImbalance. Among them, w_1, w_2, and w_3 are weight coefficients, which are determined according to the importance of the business. In this example, they are [0.3, 0.4, 0.3]. Calculated I_comprehensive = 0.3×38.37% + 0.4×34.61% + 0.3×30.08% = 34.39%
[0217] Construct a rebalancing profit model based on comprehensive imbalance indicators:
[0218] Construct the equilibrium benefit model B = f(I_current, I_expected) = k × (I_current - I_expected) / I_current; where: I_current is the current comprehensive imbalance index, which is 34.39% in this example; I_expected is the expected imbalance index, and the target value is set to 15%; k is the benefit coefficient, which is 100;
[0219] Calculated B = 100 × (34.39% - 15%) / 34.39% = 56.38;
[0220] Construct an adaptive weight adjustment mechanism for environment perception W(t) = W_base + ΔW(load(t), priority(t));
[0221] Where: W_base is the basic weight vector [0.6, 0.3, 0.1], corresponding to the balance benefit, rebalancing cost and service interference; load(t) is the current system load, assumed to be 70%; priority(t) is the service priority, assumed to be "high";
[0222] When the load is high (>60%) and the priority is "high", the weight is adjusted to: ΔW = [-0.1, 0.0, 0.1], which means reducing the weight of balanced benefit and increasing the weight of business interference.
[0223] The calculated value W(t) = [0.5, 0.3, 0.2] is used to predict the future system load trend based on the time series forecasting model: the ARIMA(2,1,2) model is used to predict the load in the next 24 hours. The results show that the load will drop from the current 70% to 50%.
[0224] Based on this forecast, the weights are adjusted: When the forecast load drops to <60%, the weights are adjusted to: ΔW_future = [0.1, 0.0, -0.1]. The calculated pre-adjusted weight vector W_adjusted = [0.6, 0.3, 0.1].
[0225] Construct a multi-objective rebalancing benefit function E = α·B - β·C - γ·D;
[0226] Where: α, β, γ are the components of the pre-adjusted weight vector, which are 0.6, 0.3, and 0.1 respectively; B is the equilibrium benefit value, calculated as 56.38; C is the rebalancing cost, to be estimated; D is the service interference degree, to be estimated.
[0227] Build a fine-grained cost model of the rebalancing operation and use structural causal model (SCM) to predict system state changes:
[0228] Building a fine-grained cost model:
[0229] C_data data migration bandwidth cost = migration data volume × unit bandwidth cost; = 40GB × 0.05 = 2.0.
[0230] C_index index reconstruction calculation cost = index size × reconstruction complexity coefficient = 5GB × 0.4 = 2.0.
[0231] C_query query performance impact cost = query volume × response time increase rate × unit response time cost = 400qps × 0.15 × 0.05 = 3.0.
[0232] C_io: I / O load increase cost = I / O increment × unit I / O cost = 200 IOPS × 0.01 = 2.0. Total cost C = C_data + C_index + C_query + C_io = 9.0.
[0233] Construct a structural causal model (SCM) to represent the causal relationship between system state variables: S = f(X, do(R))
[0234] Where: S is the system state vector [response time, throughput, error rate]; X is the influencing factor vector [data volume, query complexity, hardware resources]; do(R) represents the intervention of performing the rebalancing operation R.
[0235] The causal graph learned from historical data shows that the rebalancing operation directly affects data distribution and index status, which in turn affects response time and throughput.
[0236] Calculate the system state difference between performing rebalancing and not performing rebalancing ΔS = E[S|do(R=1)] - E[S|do(R=0)]; where E represents the expected function.
[0237] According to the causal model prediction, rebalancing will result in: a temporary increase of 15% in response time, and a long-term decrease of 20%; a temporary decrease of 10% in throughput, and a long-term increase of 25%; and no significant change in error rate;
[0238] Update the cost estimate based on the predicted ΔS of the state change: Taking into account the benefit of long-term performance improvement, the adjusted cost is: C_adjusted = C - long_term_benefit = 9.0 - 3.0 = 6.0.
[0239] Uncertainty-aware profit forecasting based on Bayesian networks and quantum probability theory:
[0240] Construct a Bayesian network to represent the probabilistic dependency between system status and rebalancing benefits: Network nodes include: partition imbalance, system load, data growth rate, rebalancing degree and final benefits
[0241] Example of a conditional probability table: P(revenue=high|imbalance=high, load=medium, data growth rate=low, rebalancing=medium) = 0.75.
[0242] Monte Carlo simulation is used to generate multiple possible results: 1000 simulations are performed to obtain the following return distribution: average return: E[B] = 48.5; return standard deviation: σ(B) = 12.3; 5% quantile: 26.8; 95% quantile: 68.7;
[0243] Calculate the risk-adjusted expected return: RAR = E[B] - λ × σ(B) = 48.5 - 1.2 × 12.3 = 33.7; where λ = 1.2 is the risk aversion coefficient, which is determined based on the importance of the business.
[0244] Quantum probability representation is introduced to deal with complementary uncertainty: the density operator ρ is used to represent the system state, and the quantum observable A is used to represent the profit measurement. The expected profit is calculated as: E[A] = Tr(ρA); this method can better express the mutual interference effect between different decision options.
[0245] Determine the optimal decision boundary based on quantum information geometry: Calculate the quantum Fisher information matrix F_Q and the quantum Cramér-Rao bound to determine the optimal decision boundary. Decision rule: Rebalancing is performed when the risk-adjusted return RAR > 30 and the decision certainty > 0.8. In this example, RAR = 33.7 and decision certainty = 0.85, which meets the execution conditions.
[0246] Optimizing rebalancing decision strategies using multi-agent reinforcement learning:
[0247] Build a reinforcement learning environment: The state space S contains indicators such as load imbalance, hotspot concentration, and system resource utilization. The action space A is {perform full rebalancing, perform partial rebalancing, adjust rebalancing parameters, and delay rebalancing}. The reward function R is based on the multi-objective benefit function E = 0.6×B - 0.3×C - 0.1×D.
[0248] Design a multi-agent architecture and configure a decision agent for each of the five data partitions. Use the Graph Attention Network (GAT) to build an inter-agent communication mechanism: i l+1 = σ(∑ j∈N(i) α_ij W h j l );
[0249] Where: h i l is the state representation of the ith agent in the lth layer; α_ij is the attention weight, which is obtained by softmax(LeakyReLU(a T [Wh_i, Wh_j])) calculation; W is the parameter matrix; the agents share local state information and coordinate decisions through this mechanism.
[0250] Train the agent network to obtain the optimal rebalancing strategy: For the current system state, the optimized decision is "perform partial rebalancing", with the following specific operations: migrate only the data of the 3rd partition (hotspot) and the 4th partition (low load); the migration ratio is 30%; execute during the low load period; the maximum bandwidth limit is 50MB / s.
[0251] Based on the optimization strategy, the migration path with minimum interference is planned, which specifically includes: building a data dependency graph and analyzing the access relationship between shards: the data shards to be migrated are represented as a graph G(V,E), where the node V represents the shard, the edge E represents the access relationship between the shards, and the edge weight w_ij represents the access frequency of shards i and j. According to the analysis, there is a high access relationship between the shards {S3.1, S3.2, S3.5} in the hotspot partition (partition 3).
[0252] Calculate the shard association graph and design the migration priority: Detect the community structure based on the Louvain algorithm to ensure the coordinated migration of highly associated shards. Priority calculation formula: Priority(S_i) = w_1×Hot(S_i) + w_2×Size(S_i)+ w_3×Dependency(S_i). Where: Hot(S_i) is the heat of shard S_i; Size(S_i) is the size of shard S_i; Dependency(S_i) is the dependency of shard S_i; w_1, w_2, w_3 are weight coefficients, and the values are [0.5, 0.3, 0.2].
[0253] The migration priorities of the shards in partition 3 are calculated as follows: S3.1: 0.85; S3.2: 0.78; S3.5: 0.72.
[0254] Generate a migration path plan with minimal interference: migration batch plan [{S3.1, S3.2}, {S3.5}, {S4.3,S4.7}]; target partition allocation S3.1→P1, S3.2→P5, S3.5→P2, S4.3→P5, S4.7→P1. Execution time window 02:00-04:00 (system load trough period).
[0255] Perform incremental partition data migration: Set the initial migration rate to 30MB / s and the parallelism to 2. The waiting time between batches is 5 minutes.
[0256] Real-time monitoring of execution status: system load monitoring: CPU, memory, I / O; service quality monitoring: response time, throughput, error rate;
[0257] Adjust execution parameters based on real-time monitoring data: After the first batch was completed, the CPU load was lower than the threshold, and the migration rate was adjusted to 40MB / s. During the second batch, it was found that the response time increased slightly, and the degree of parallelism was reduced to 1.
[0258] After completing the migration, the rebalancing effect was evaluated: comprehensive imbalance index before rebalancing: 34.39%; comprehensive imbalance index after rebalancing: 18.25%; improvement degree: 47.0%; actual resource consumption: CPU increased by an average of 12%, bandwidth usage 35MB / s; service impact: response time temporarily increased by 8%, no error rate increased.
[0259] Update rebalancing decision model parameters: Based on the actual execution results, update the cost model parameters and uncertainty model to optimize the next rebalancing decision.
[0260] This embodiment performs stream-batch collaborative processing on the real-time data stream and historical batch data in the quality labeling data set based on the updated partitioning scheme.
[0261] Build a unified stream and batch processing model, including:
[0262] Design a unified data model, including: time field: that is, providing the timestamp of the event; entity ID: identifying the data source (such as meter ID); measurement value: recording the actual metering data; quality mark: indicating the data quality level; processing status: marking the data processing stage.
[0263] Build a Lambda+Kappa architecture: Lambda layer: processes batch historical data and provides accurate but high-latency analysis results; Kappa layer: processes real-time data streams and provides near-real-time but slightly lower-precision analysis results; Coordination layer: is responsible for result fusion and consistency management.
[0264] Apply the updated partitioning scheme: Redistribute the data of the five partitions according to the new scheme to ensure that the stream processing and batch processing use the same partitioning strategy.
[0265] Perform incremental feature updates and real-time anomaly detection on live data streams:
[0266] Real-time data distribution: Receive real-time data streams from quality labeled datasets and distribute them in real time according to the partitioning scheme
[0267] For meter data, the partition is determined by hashing the meter ID and taking the modulus of the number of partitions: Partition(meter_id) = Hash(meter_id) % num_partitions.
[0268] Sliding window processing: the window size is 10 minutes; the sliding step is 1 minute; the calculated window statistics are mean, standard deviation, rate of change, and peak value;
[0269] The 10-minute window data of electricity meter E10056782 is calculated as follows: the mean is 4.8 kWh; the standard deviation is 0.7 kWh; the rate of change is +2.1%; and the peak is 5.9 kWh.
[0270] Incremental feature update: Use the exponentially weighted moving average (EWMA) method to update features: F_new = α × F_current + (1-α) × F_old; where α is the smoothing coefficient, which is 0.3.
[0271] When a mode change is detected, a fast feature update is triggered: α_dynamic = min(0.8, base_α ×change_magnitude).
[0272] Real-time anomaly detection: Detect point anomalies based on statistical control charts: If |x_t - μ| > k × σ, mark it as an anomaly where k is the sensitivity coefficient, which is 3.
[0273] Identify trend anomalies based on change point detection algorithm: Use CUSUM algorithm to monitor cumulative deviation S_t = max(0,S_{t-1} + (x_t - μ - δ)); if S_t > h, trigger an alarm; where δ is the offset parameter and h is the threshold.
[0274] Detect pattern anomalies based on pattern matching: Use the dynamic time warping (DTW) algorithm to calculate the distance between the current pattern and the normal pattern.
[0275] Generate real-time processing results: the processing timestamp is 2024-02-15T08:40:00Z; the entity ID is E10056782; the aggregate value is 4.8 kWh (10-minute average); the anomaly is marked as normal; the updated feature is {trend: increasing, pattern: morning_peak}.
[0276] Perform time series analysis and quality verification on historical batch data:
[0277] Load the historical measurement data for the past 30 days from the data warehouse and reorganize it according to the new partitioning scheme.
[0278] For the data in partition 1: the data volume is 105 GB; the number of records is approximately 840 million; the time range is from 2024-01-15 to 2024-02-14;
[0279] Seasonal decomposition: Decompose into trend, seasonal and residual components using STL (Seasonal-Trend decomposition using Loess) method. Use Fast Fourier Transform (FFT) to identify hidden cycles. Use hierarchical clustering to identify typical load patterns.
[0280] Verify that metering data complies with physical constraints and grid balance conditions. Identify missing intervals in time series and evaluate interpolation performance. Perform systematic anomaly detection on batch results to filter possible false positives.
[0281] The batch processing results after verification are generated: the processing date is 2024-02-14; the data coverage is 30 days of complete records; the quality score is 92.5% (high quality); the anomaly summary is that 21 abnormal points and 3 abnormal intervals are detected; the typical pattern is the identification of weekday pattern, weekend pattern, and holiday pattern.
[0282] Perform multi-level data fusion based on real-time processing results and verified batch processing results:
[0283] Determine the overlap interval of results: The last 24 hours of real-time processing data overlap with the last 24 hours of batch processing results. Calculate the consistency index: mean deviation |(μ_real - μ_batch) / μ_batch| × 100% = 2.3%. The proportion of anomalies with consistent detection results = 92.7%. The consistency of the trend direction of the two results = 96.5%.
[0284] Determine the fusion strategy. For data point level: When the quality of real-time results is marked as "high" and the deviation from batch results is <5%, the real-time results are preferred. For trend analysis: When the confidence of batch results is >90%, the batch results are preferred. For anomaly detection: When the two results are inconsistent, a weighted voting mechanism is applied to determine the final result.
[0285] Perform data fusion. Use the Bayesian fusion method to combine the distribution characteristics of the two results: p(z|x,y) ∝ p(x|z)p(y|z)p(z); where: z is the true value after fusion; x is the real-time processing result; y is the batch processing result; p(z) is the prior distribution;
[0286] The final processing result is generated: the processing time is 2024-02-15T08:40:00Z; the short-term trend is +2.3% (from real-time processing); the long-term mode is the regular mode of weekdays (from batch processing); the abnormal state result is normal; the quality index is 95.3% (improved after fusion).
[0287] This embodiment conducts multi-dimensional analysis and application of the final processing results to achieve storage, visualization and abnormal warning of measurement data.
[0288] Data model organization. Design the data model based on the star schema, including: fact tables for storing meter readings and calculated indicators; dimension tables including time dimension, space dimension, equipment dimension, etc.
[0289] Multi-level storage strategy. Hot data (last 7 days): stored in the memory database; warm data (last 90 days): stored on SSD; cold data (more than 90 days): stored on HDD and compressed; archived data (more than 1 year): stored in object storage;
[0290] Establish multi-dimensional indexes. Time index: implement efficient time range query based on B+ tree; spatial index: use R tree to support regional query; device index: use hash index to accelerate query by device ID; composite index: support multi-condition combination query.
[0291] Index update strategy: Hot data index: real-time update; warm data index: hourly batch update; cold data index: daily batch update.
[0292] Perform multi-dimensional statistical analysis based on persistent result data:
[0293] Time dimension analysis, daily load curve: 24-hour load changes; weekly load pattern: difference between weekdays and weekend patterns; monthly trend: month-on-month load changes; annual cycle: seasonal change pattern.
[0294] Spatial dimension analysis, regional load distribution: heat map shows the electricity consumption density in different regions; propagation mode: the diffusion pattern of electricity consumption behavior in space; correlation analysis: correlation between loads in adjacent regions.
[0295] User dimension analysis, user clustering: user grouping based on electricity usage behavior; abnormal user identification: detection of users with abnormal electricity usage behavior; electricity usage pattern evolution: changes in user electricity usage behavior over time.
[0296] Generate a visual resource set, interactive load curve: support zoom in, zoom out, and compare functions; regional heat map: dynamically display the load distribution at different times; anomaly detection panel: highlight the detected anomalies; user portrait dashboard: display user power consumption characteristics and behavior patterns.
[0297] Build an interactive data exploration interface with multi-level drill-down: from a global view to detailed data points; multi-dimensional filtering: filter data by time, region, user type, etc.; comparative analysis: support data comparison in different time periods and regions; forecast simulation: simulate future load scenarios based on historical data.
[0298] Identify and warn of abnormalities in the final processing results:
[0299] Statistical anomalies: Detect numerical anomalies based on the 3σ rule; Pattern anomalies: Detect pattern deviations based on the DTW algorithm; Contextual anomalies: Anomalies that consider environmental factors (such as temperature); Collective anomalies: Abnormal patterns of collaborative performance of multiple devices.
[0300] The decision tree algorithm is used to filter false positives, with an accuracy rate of 92%. The anomalies are classified into: sudden anomalies, gradual anomalies, periodic anomalies, and systematic anomalies.
[0301] Generate graded warning information: Level 1 (information): minor deviation, no immediate processing required; Level 2 (warning): significant deviation, need attention; Level 3 (serious): major deviation, need timely processing; Level 4 (urgent): extreme deviation, need immediate response;
[0302] Early warning push and response, information level: recorded in system log; warning level: pushed to monitoring interface; severe level: sent email and SMS notification; emergency level: trigger automatic call and emergency response;
[0303] Closed-loop management, recording the warning processing status: unprocessed, in progress, resolved, false alarm; statistical warning effectiveness: true positive rate, false positive rate; continuous optimization of warning rules: adjusting thresholds based on historical warning effectiveness.
[0304] The preferred embodiments of the present invention are described in detail above; however, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration, characterized in that: The steps include: Collect metrological data from multiple heterogeneous data sources and pre-process them to form a quality-labeled dataset containing quality grade labels; Based on the quality-labeled dataset, the dynamic partition management module performs adaptive partition feature extraction based on time-varying features and low-overhead partition rebalancing to generate an updated partition scheme. Based on the updated partitioning scheme, the real-time data stream and pre-stored historical batch data in the quality labeling dataset are processed in a stream-batch collaborative manner, and the final processing result is generated through multi-level data fusion. The final processing results are analyzed and applied in multiple dimensions to achieve storage, visualization and abnormal warning of measurement data.
2. According to claim 1, a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration is characterized in that: The steps of forming a quality label dataset containing quality grade labels include: According to the preset data collection rules, access multiple heterogeneous data sources and collect metering data to generate original metering data; Perform protocol conversion and analysis, as well as data cleaning and standardization on the original metering data to obtain standardized metering data; Calculate the quality index of the normalized measurement data, add a quality grade mark to each normalized measurement data based on the quality index, and form a quality mark data set.
3. According to claim 1, a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration is characterized in that: The steps of performing adaptive partition feature extraction and low-overhead partition rebalancing based on time-varying features to generate an updated partition scheme include: Based on the quality labeled dataset, adaptive partition feature extraction of time-varying features is performed to obtain partition features; Apply the partition feature to initially partition the data, generate an initial partition scheme, and obtain the system partition status; Combined with the pre-stored system operation status and load statistics, low-overhead partition rebalancing is performed to update the initial partition scheme and generate an updated partition scheme.
4. According to claim 1, a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration is characterized in that: The steps to perform stream-batch collaborative processing and generate the final processing results include: Build a unified stream-batch processing model and apply the updated partitioning scheme to the real-time data stream and pre-stored historical batch data in the quality labeled dataset; Perform incremental feature updates and real-time anomaly detection on real-time data streams to generate real-time processing results; Perform time series analysis and quality verification on historical batch data to obtain verified batch results; Multi-level data fusion is performed based on real-time processing results and verified batch processing results to generate the final processing results.
5. According to claim 1, a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration is characterized in that: The steps for multi-dimensional analysis and application to achieve storage, visualization and abnormal warning of measurement data include: The final processing results are organized according to the predetermined data model, written into the persistent storage module, the persistent result data is obtained and a multi-dimensional index is established; Perform multi-dimensional statistical analysis based on persistent result data to generate a visualization resource set; According to the preset exception rules, the final processing results are identified as abnormal and the corresponding warning information is triggered.
6. According to claim 3, a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration is characterized in that: The steps of performing adaptive partition feature extraction of time-varying features to obtain partition features include: The quality-labeled data set is input into a multi-scale time-varying feature decomposition unit to extract the dominant feature components; Perform feature sensitivity evaluation and screening based on the dominant feature components to generate a streamlined feature set; Use the reduced feature set to build a dynamically evolving pattern recognizer and output pattern prediction results; Adaptive feature fusion and weight assignment are performed according to the pattern prediction results to generate weighted fusion features, namely partition features.
7. According to claim 3, a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration is characterized in that: The steps to perform a low-overhead partition rebalance and generate an updated partition scheme include: Collect load statistics of each partition; calculate the imbalance index based on the load statistics, analyze the data access pattern and identify the hotspot partitions, build a multi-dimensional imbalance assessment model, and output a comprehensive imbalance index; A rebalancing benefit model is built based on comprehensive imbalance indicators, and an environment-aware weight adjustment mechanism is built based on system partition status and system operation status. The rebalancing cost is predicted and uncertainty modeling is performed to optimize the rebalancing strategy. Build a data dependency graph based on the optimized rebalancing strategy, calculate the shard association graph and build a migration priority algorithm to generate a migration path plan with minimal interference; Perform incremental partition data migration according to the migration path plan, monitor the execution status in real time and dynamically adjust the migration parameters; Based on the adjusted migration parameters, the improvement of the comprehensive imbalance index before and after rebalancing is evaluated, the actual resource consumption cost is calculated, the rebalancing decision model parameters are updated, and the updated partitioning scheme is generated.
8. According to claim 6, a method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration is characterized in that: The steps to build a dynamically evolving pattern recognizer include: Receive a simplified feature set, determine the optimal time lag parameter based on the mutual information minimum principle, determine the optimal embedding dimension by combining the improved pseudo-nearest neighbor algorithm, and reconstruct the one-dimensional time series into multi-dimensional phase space trajectory data; use kernel density estimation to generate a phase space density map; Based on the phase space density map, the Lyapunov exponent, correlation dimension and entropy rate of the phase space trajectory data are calculated. The sparse recognition method and deep learning are combined to identify the dynamic equation, and the dynamic model parameters are obtained through L1 regularization optimization. Based on the phase space trajectory data and dynamic model parameters, a recursive graph is constructed to analyze the network topology characteristics, and the distribution change rate is calculated in combination with the Wasserstein distance to identify multi-scale transition points; The time series is segmented according to the multi-scale transition points, the pattern features are extracted to construct the knowledge graph, the graph neural network is trained to learn the pattern transition rules, and the dynamic evolution pattern recognizer is output.
9. The method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration according to claim 7 is characterized in that: The steps to predict rebalancing costs and model uncertainty and optimize the rebalancing strategy include: Construct a multi-objective rebalancing benefit function based on comprehensive imbalance indicators, and output the equilibrium benefit value and dynamic weight vector; Combined with the pre-stored current system load, rebalancing cost prediction is performed based on causal inference to obtain cost estimates and state change predictions; Based on the equilibrium benefit value and cost estimate, uncertainty-aware benefit prediction is performed to output risk-adjusted benefits and quantum-optimized decision boundaries; Based on risk-adjusted returns, quantum-optimized decision boundaries, and state change predictions, the rebalancing strategy is optimized through reinforcement learning to output the optimal rebalancing strategy.
10. A method for processing metering data based on dynamic partition rebalancing and stream-batch collaboration according to claim 9, characterized in that: The steps to construct a multi-objective rebalancing benefit function include: Based on the comprehensive imbalance index, the balanced return model after rebalancing is constructed: B = f(I_current, I_expected); Where I_current is the current comprehensive imbalance index, I_expected is the expected imbalance index, f() is a function that outputs the equilibrium benefit value B; Combined with the current state of the system, an environment-aware adaptive weight adjustment mechanism is constructed: W(t) = W_base + ΔW(load(t), priority(t)); Where W(t) is the weight vector at time t, W_base is the base weight, load(t) is the system load at time t, priority(t) is the service priority, ΔW is the weight change, and the dynamic weight vector W is output; Build a time series forecasting model to predict the system load trend in the future period and obtain the load forecast value; Adjust the dynamic weight vector W based on the load prediction value to obtain a pre-adjusted weight vector; Combine the equilibrium benefit value B with the pre-adjusted weight vector to construct a complete multi-objective rebalancing benefit function: E = α•B - β•C - γ•D; Among them, α, β, γ are weight coefficients from the pre-adjusted weight vector, B is the equilibrium benefit value, C is the rebalancing cost, and D is the service interference degree.
Citation Information
Patent Citations
Data detection method and device for power grid regulation and control multi-source time sequence data batch stream fusion
CN118069643A
Method for dynamically partitioning relational cluster database based on spatio-temporal data
CN119226421A
Distributed resource cooperative control method based on layering and partitioning autonomy of power distribution network
CN119448281A
Rural digital governance control method and system based on cloud platform sharing
CN119579029A
Node load-based dynamic data partitioning system
WO2021073083A1
Cited By
Road surface collapse risk monitoring method and system
CN120339962A
Data sharing method and system for multi-agent platform
CN120832328A