A Distributed Storage Method for Electric Energy Meter Data Based on Improved PoS Consensus
By introducing the weight computer system of grid load perception and the sharding strategy of space-time optimization in the power blockchain, the problem of insufficient correlation between power load fluctuations and consensus efficiency and low matching between space-time characteristics of power meter data with traditional sharding strategies is solved, and the system's resource allocation efficiency and data processing capabilities are improved.
Patent Information
- Application Number
- CN202510388027.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing technology has insufficient correlation with power load fluctuations and consensus efficiency, and the traditional PoS mechanism fails to dynamically adjust node weights, resulting in the disconnection of resource allocation of blockchain system and the demand for power grids. The spatial and temporal characteristics of the power meter data match the traditional sharding strategy, which increases the overhead of cross-shashing query.
By introducing a weight computer system for grid load-awareness, the node weights in the PoS consensus are dynamically adjusted, and the electricity meter clustering is constructed based on the spatiotemporal and spatiotemporal and spatial-temporal characteristics analysis of the electricity meter data, the sharding scheme for spatiotemporal optimization is designed, and the node grouping and data sharding are optimized.
It effectively solves the problem of insufficient performance and adaptability of power blockchain, improves the system's resource allocation efficiency and data processing capabilities, and reduces cross-shash query overhead.
Smart Images

Figure CN119902721B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to power electronics technology, and in particular to a distributed storage method for electric energy meter data based on improved PoS consensus. Background Art
[0002] With the in-depth advancement of smart grid construction, the amount of data generated by electricity meters has exploded, posing severe challenges to data storage and processing systems. Traditional centralized storage faces capacity bottlenecks, single point failure risks and data credibility issues, while blockchain distributed storage provides a new solution for electricity meter data management with its tamper-proof, traceable and decentralized characteristics. The construction of power blockchain is of great significance to promoting the transparency of power market transactions, improving the intelligence of power grid dispatching and ensuring data security and reliability. Especially in the context of smart microgrids and energy Internet, the blockchain-based electricity meter data storage method can realize decentralized collaboration involving multiple parties, provide infrastructure support for energy data value mining and intelligent decision-making, and has important research value for promoting the construction of energy Internet ecology and the digital transformation of power systems.
[0003] At present, the research on blockchain storage of electric energy meter data mainly focuses on basic consensus mechanism and simple sharding strategy. In terms of consensus mechanism, researchers mostly use standard PoW (proof of work) or PoS (proof of stake) algorithms, and some work attempts to introduce reputation mechanism or identity authentication to enhance security; in terms of data sharding, random sharding based on hash function or simple geographical partitioning strategy is mainly adopted. In data storage design, unified storage strategy is mostly adopted; verification mechanism mainly relies on general cryptographic verification. In terms of system adaptability, static configuration parameters are mostly used, and the response ability to changes in grid load is limited.
[0004] In summary, the existing technology faces two key technical problems: on the one hand, there is insufficient correlation between power load fluctuations and consensus efficiency. The traditional PoS mechanism does not take into account the operating status of the underlying physical system (such as the power grid), resulting in the resource allocation of the blockchain system being out of touch with the actual needs of the power grid, and unable to prioritize key data processing when the power grid load is peak or fluctuating violently; on the other hand, the temporal and spatial characteristics of meter data have a low match with the sharding strategy. The meter data has a unique temporal and spatial correlation (the electricity consumption patterns of adjacent meters and similar users are highly similar), while the traditional hash-based random sharding strategy disperses related data to different shards, increasing the cross-shard query overhead and reducing system efficiency. Summary of the invention
[0005] Purpose of the invention: To provide a distributed storage method for electric energy meter data based on improved PoS consensus, in order to solve the above technical problems.
[0006] Technical solution: A distributed storage method for electric energy meter data based on improved PoS consensus, comprising the following steps:
[0007] Obtain electric energy meter data, conduct spatio-temporal characteristic analysis, form electric meter clusters, and construct a spatio-temporal characteristic model of electric energy meter data;
[0008] Obtain blockchain network node information, and calculate node performance indicators and network latency between nodes;
[0009] Based on the spatio-temporal characteristic model of electric energy meter data, node performance indicators, and network latency, group the nodes to form node clusters, and calculate the affinity between electric meter clusters and node clusters;
[0010] According to the affinity, allocate the electric meter clusters to the node clusters and assign responsibilities to the nodes;
[0011] During operation, obtain and design a data heat evaluation function based on the data access characteristics of each electric meter cluster, formulate a hot data replication strategy and a data life cycle management strategy, and output a spatio-temporal optimized sharding scheme.
[0012] Beneficial effects: By introducing a weight calculation mechanism for power grid load perception, the present invention enables PoS consensus to dynamically adjust node weights according to the power grid load status. At the same time, based on the spatio-temporal characteristic analysis of electric energy meter data, electric meter clusters are constructed and spatio-temporal optimized sharding allocation is performed, effectively solving the above technical problems and significantly improving the performance and adaptability of the power blockchain. Description of the drawings
[0013] Figure 1 is the flowchart of the present invention.
[0014] Figure 2 is the flowchart of step S1 of the present invention.
[0015] Figure 3 is the flowchart of step S2 of the present invention.
[0016] Figure 4 is the flowchart of step S3 of the present invention.
[0017] Figure 5 is the flowchart of step S4 of the present invention.
[0018] Figure 6 is the flowchart of step S5 of the present invention. Detailed implementation manners
[0019] To make the technical problems of the present invention clearer, embodiments of the present invention will be provided below and a more comprehensive description will be given. However, the specific embodiments given here are only used to explain the present invention and are not used to limit the usage environment and scope of the invention.
[0020] When processing electricity meter data: First, electricity meter data has significant spatio-temporal correlation. The electricity consumption patterns of adjacent meters or similar users are often highly similar, but the current hash-based sharding strategy randomly scatters related data across different shards, increasing the cross-shard query overhead. Second, the power system load has obvious time-varying characteristics, but the existing PoS consensus lacks the ability to perceive this, and cannot dynamically adjust node weights according to the load status, resulting in a disconnect between system resource allocation and the actual needs of the power grid. Third, the access to electricity meter data shows unbalanced characteristics. Some real-time monitoring data is frequently accessed while historical data is less used. The unified storage strategy leads to congestion in accessing hot data and waste in storing cold data. Fourth, power data must conform to the physical constraints of the power system, but the existing verification mechanism only focuses on data integrity and source authentication, and cannot detect abnormal data that meets cryptographic requirements but violates physical laws. Fifth, the static sharding configuration cannot adapt to changes in the power grid topology and the evolution of load distribution. As time goes by, the system performance gradually degrades, but there is a lack of an effective adaptive adjustment mechanism. In addition, there is no clear mapping relationship between blockchain nodes and power grid physical facilities, making it difficult to optimize the resource allocation of related nodes according to the load status of substations, which affects the system's response ability in the emergency state of the power grid.
[0021] Therefore, a distributed storage method for electricity meter data based on improved PoS consensus is proposed, as Figure 1 shown, including:
[0022] Step S1: Improvement of the PoS consensus mechanism with power grid load perception, including:
[0023] Obtain the regional real-time load data and substation load data from the SCADA system, apply a sliding window average to the regional real-time load data, and calculate the load fluctuation index; based on the smoothed load data and the load fluctuation index, combined with a preset threshold, determine the current load status; obtain the original staking amount, regional information, and historical reliability index of the nodes from the blockchain network, and calculate the adjusted staking amount; calculate the node weights after load adjustment according to the load status, substation load data, adjusted staking amount, and node regional information for subsequent consensus processes; the specific process is as follows:
[0024] Step S11: Obtain the regional real-time load data and substation load data from the SCADA system, apply a 15-minute sliding window average to the regional real-time load data to obtain the smoothed load data, and obtain the load fluctuation index by calculating the square root formula;
[0025] Step S12: Receive the smoothed load data and the load fluctuation index, classify the power grid load status as high load and high fluctuation, high load and low fluctuation, low load and high fluctuation, or low load and low fluctuation according to the preset load threshold and fluctuation threshold, and output the current load status and substation load data;
[0026] Step S13: Obtain the original pledge amount, regional information, and historical reliability index of each node from the blockchain network, and calculate the adjusted pledge amount according to the historical reliability adjustment formula;
[0027] Step S14, receiving the load status, substation load data, adjusted pledge amount and node area information, associating the node with the corresponding substation, setting the weight adjustment factor according to different load status, calculating the load status adjustment item, and finally obtaining the node weight after load adjustment.
[0028] Step S2: Analysis of spatiotemporal characteristics of electric energy meter data and model construction, including:
[0029] The original data of historical electric energy meters are obtained from the electric energy meter database, and data cleaning and time regularization are performed to obtain regularized electric energy meter data; the daily load curve and weekly load curve are extracted from the regularized electric energy meter data, and the time similarity matrix is calculated; the geographical location information of the electric energy meters and the power grid topology are obtained from the power grid topology database, the electrical distance and geographical distance between the electric energy meters are calculated, and the spatial similarity matrix is generated; the temporal similarity matrix and the spatial similarity matrix are combined to calculate the spatiotemporal similarity, and the hierarchical clustering algorithm is applied to form electric energy meter clusters, extract clustering feature vectors, and construct the spatiotemporal characteristic model of electric energy meter data; the specific process is as follows:
[0030] Step S21, obtaining the historical electric energy meter original data of the past 30 days from the electric energy meter database, performing data cleaning, removing abnormal values, obtaining cleaned electric energy meter data, performing time regularization on the cleaned electric energy meter data, unifying the sampling interval to 15 minutes, and outputting the regularized electric energy meter data;
[0031] Step S22, receiving the regularized meter data, extracting the daily load curve and weekly load curve of each meter, extracting the spectrum characteristics of the meter data using Fourier transform, and calculating the time similarity matrix between the meters based on the spectrum characteristics and the load curve;
[0032] Step S23, obtaining the geographical location information of the electric meters and the topological structure of the electric grid from the electric grid topology database, calculating the electrical distance and geographical distance between the electric meters, and combining the two distance indicators to calculate the spatial similarity matrix between the electric meters;
[0033] Step S24, receiving the time similarity matrix and the space similarity matrix, applying weights to calculate the spatiotemporal similarity, using a hierarchical clustering algorithm to divide the electric meters into several clusters, extracting the eigenvector of each cluster, establishing a mapping relationship between the electric meters and the clusters, and constructing a spatiotemporal characteristic model of the electric energy meter data including the similarity matrix, clusters and eigenvectors.
[0034] Step S3: Design of spatiotemporal-aware sharding strategy, including:
[0035] Obtain the network node list and the computing power, storage capacity, and network bandwidth of each node from the blockchain network monitoring module, calculate the comprehensive performance index; obtain the node geographical location information and the network latency matrix; based on the spatio-temporal characteristic model of the electricity meter data, the node performance index, the node geographical location, and the network latency matrix, cluster the nodes to form node clusters, assign electricity meter clustering to each node cluster, and design the responsibilities of nodes within the shard; design the intra-shard communication protocol and the inter-shard communication protocol, construct the shard routing table and the inter-shard data consistency protocol; based on the electricity meter data access pattern, design a data heat evaluation function, formulate a hot data replication strategy and a data life cycle management strategy, and output a spatio-temporally optimized sharding scheme; the specific process is as follows:
[0036] Step S31: Obtain the list of currently active network nodes and the computing power, storage capacity, and network bandwidth of each node from the blockchain network monitoring module, calculate the comprehensive performance index of each node by applying weights, and obtain the geographical location information of each node and the network latency matrix between nodes;
[0037] Step S32: Receive the spatio-temporal characteristic model of the electricity meter data, the node performance index, the node geographical location, and the network latency matrix, cluster the active nodes according to the geographical location and network latency by applying the K-means algorithm to form node clusters; calculate the affinity between the electricity meter clustering and the node clusters, solve the optimal allocation problem, and determine the mapping relationship from the electricity meter clustering to the node clusters; subdivide the responsibilities of nodes within each node cluster, including the main shard node, the backup node, and the verification node, to generate a spatio-temporally optimized sharding scheme;
[0038] Step S33: Receive the spatio-temporally optimized sharding scheme, design the intra-shard communication protocol, including the data synchronization message format and the consensus message format; design the inter-shard communication protocol, including the cross-shard query message format and the cross-shard transaction message format; establish the shard routing table, including the responsible scope and communication address of each shard; formulate the inter-shard data consistency protocol to handle data update operations involving multiple shards;
[0039] Step S34: Receive the spatio-temporally optimized sharding scheme and the spatio-temporal characteristic model of the electricity meter data, design a data heat evaluation function to evaluate the access frequency of the electricity meter data, formulate a hot data replication strategy based on this function to replicate the frequently accessed data between nodes; according to the time characteristics of the electricity meter data, design a data life cycle management strategy, including data compression, archiving, and cleaning strategies, and update and output the final spatio-temporally optimized sharding scheme.
[0040] Step S4: Adaptive consensus and shard dynamic adjustment mechanism, including:
[0041] Periodically obtain the latest regional real-time load data and substation load data, calculate the current load status, and predict the predicted load status sequence for the next hour; adjust the consensus parameters, including the number of verification nodes and the block time interval, and update the validator selection probability according to the load status, predicted load status sequence, and node weights after load adjustment; evaluate the performance metrics of the current spatio-temporal optimized sharding scheme, including the cross-shard transaction ratio, node load balance, and data access latency, and set the sharding adjustment threshold function; when the performance metrics exceed the threshold, trigger sharding adjustment, redesign the sharding scheme based on the latest spatio-temporal characteristics model of the electricity meter data and the load status, formulate a smooth transition strategy, complete the sharding adjustment, and output the adaptive sharding storage strategy; the specific process is as follows:
[0042] Step S41: Receive the latest regional real-time load data and substation load data every 5 minutes, calculate the current load status using the aforementioned method, and use the time series prediction model to process the historical load data to predict the predicted load status sequence every 5 minutes within the next 1 hour;
[0043] Step S42: Receive the load status, predicted load status sequence, and node weights after load adjustment, adjust the consensus parameters according to the load status, including the number of verification nodes and the block time interval; update the probability of each node being selected as a validator, and output the adjusted consensus parameters;
[0044] Step S43: Receive the spatio-temporal optimized sharding scheme, load status, and predicted load status sequence, calculate the performance metrics of the current sharding scheme, including the cross-shard transaction ratio, node load balance, and data access latency; design a sharding adjustment threshold function based on the load status, and trigger sharding adjustment when the performance metrics exceed the threshold; when the predicted load status indicates a significant change in the future, formulate a pre-adjustment plan;
[0045] Step S44: Receive the sharding adjustment flag, performance metrics, pre-adjustment plan, spatio-temporal optimized sharding scheme, and spatio-temporal characteristics model of the electricity meter data. When sharding adjustment needs to be performed, redesign the sharding scheme based on the latest status; formulate a smooth transition strategy, including a data migration plan, transition time window, and switch checkpoint; execute the transition strategy, complete the sharding adjustment, update and output the adaptive sharding storage strategy, including the updated sharding scheme, dynamic adjustment strategy, and performance monitoring metric definitions.
[0046] Step S5: Data consistency guarantee and verification mechanism, including:
[0047] Obtain the power grid topology and electrical parameters from the power system model library to establish a power flow model; extract the power usage pattern library from historical data to establish an electric energy balance constraint model; based on the power system physical constraint model and the adaptive sharding storage strategy, design data verification rules to form a verification rule set, including basic verification rules, physical constraint-based verification rules, and inter-shard consistency verification rules; construct a hierarchical verification architecture, design a verification task allocation algorithm and a verification result consensus mechanism; establish an abnormal pattern library, implement various anomaly detection algorithms, and design an exception handling process; execute the distributed verification process, collect the results of all verification nodes, apply the verification result consensus mechanism to obtain the final verification result, trigger the exception handling process, and output the verified trusted data and verification report; the specific process is as follows:
[0048] Step S51: Obtain the power grid topology and electrical parameters from the power system model library to establish a power flow model for verifying the physical rationality of the electricity meter data; extract the power usage pattern library containing typical electricity usage patterns of various users from historical data; establish an electric energy balance constraint model expressing the balance relationship between the input electric energy and the consumed electric energy; integrate the above models and output the power system physical constraint model;
[0049] Step S52: Receive the power system physical constraint model and the adaptive sharding storage strategy, design basic verification rules, including data format verification, timestamp verification, and digital signature verification; design physical constraint-based verification rules, including electric energy balance verification, power flow constraint verification, and usage pattern verification; design inter-shard consistency verification rules, including cross-shard data consistency verification and shard boundary verification; set execution conditions and priorities for each verification rule and output the verification rule set;
[0050] Step S53: Receive the verification rule set and the adaptive sharding storage strategy, construct a hierarchical verification architecture including in-node verification, intra-shard verification, and cross-shard verification; design a verification task allocation algorithm for allocating verification tasks according to the verification type and node capabilities; design a verification result consensus mechanism to ensure the consistency of verification results; output the distributed verification process;
[0051] Step S54: Receive the electricity meter data and the distributed verification process, establish an abnormal pattern library containing common abnormal patterns; implement anomaly detection algorithms, including rule-based detection, statistical anomaly detection, and machine learning anomaly detection; design an exception handling process, including anomaly grading, handling strategies, and recovery mechanisms; output the anomaly detection and handling mechanism;
[0052] Step S55: Receive the power meter data, the distributed verification process, and the anomaly detection and handling mechanism, execute the verification process, including in-node verification, intra-shard verification, and cross-shard verification; collect the verification results of all verification nodes, apply the verification result consensus mechanism to obtain the final verification result; trigger the anomaly handling process when anomalies are found; mark the data that passes the verification as trusted data after verification; generate a verification report including the verification process, results, and the handling of abnormal situations.
[0053] According to one aspect of the present application, step S11 is specifically as follows:
[0054] Step S111: Extract the regional real-time load data L(t) of the most recent 24 hours from the real-time database of the SCADA system, including the time stamp t and the corresponding total regional load value; obtain the substation load data LS(i,t) of each substation from the substation monitoring module of the SCADA system, where i is the substation identifier and t is the corresponding time stamp; perform an integrity check on the acquired original data, mark and interpolate the missing data points, and output the complete regional real-time load data L(t) and the substation load data LS(i,t).
[0055] Step S112: Receive the complete regional real-time load data L(t), define the sliding window length T = 15 minutes, for the current moment t, extract the subset of load data within the time range [t - T, t]; calculate the average load within this time window L(t)=(1 / T)·∑L(τ), where τ∈[t - T, t]; repeat the above process to calculate the moving average for each time point in the past 24 hours and form a time series of smoothed load data L(t); store the smoothed load data L(t) in a temporary data cache for subsequent calculations and load status determination.
[0056] Step S113: Receive the regional real-time load data L(t) and the smoothed load data L*(t) generated in step S112, for the current moment t, extract the original load data L(τ) and the corresponding smoothed load value L*(t) within the time range [t - T, t]; calculate the square of the difference between the original load and the smoothed load at each time point (L(τ)-L*(t))²; calculate the mean of these squared differences and take the square root to obtain the load fluctuation index V(t)=√(1 / (T - 1)·∑(L(τ)-L*(t))²), τ∈[t - T, t]; repeat the above process to calculate the load fluctuation index V(t) for each time point in the past 24 hours; output the smoothed load data L*(t) and the load fluctuation index V(t) to step S12 for load status classification and evaluation.
[0057] Step S114: For each key node in the distribution network, obtain the node voltage data V(n, t) and node current data I(n, t) from the SCADA system, where n is the node identifier and t is the timestamp; then the node apparent power S(n, t) = V(n, t)·I(n, t), the active power P(n, t) = S(n, t)·cosφ(n, t), and the reactive power Q(n, t) = S(n, t)·sinφ(n, t), where φ(n, t) is the corresponding power factor angle; apply a sliding window average to the calculated node power data P(n, t) and node reactive power data Q(n, t) to obtain the smoothed node power data P(n, t) and the smoothed node reactive power data Q(n, t); associate these power data with the corresponding substation to form the enhanced substation load data LS+(i, t), which includes the original load and power component information; output the substation load data LS+(i, t) to Step S12 to provide more comprehensive load characteristic information.
[0058] In this way, Step S11 outputs the smoothed load data L*(t), the load fluctuation index V(t), and the enhanced substation load data LS+(i, t) by acquiring and preprocessing the grid load data, providing the necessary data basis for subsequent load state classification and evaluation.
[0059] According to one aspect of the present application, the calculation of the load perception weight in Step S14 is specifically as follows:
[0060] Step S141: Receive the load state S(t), the substation load data LS+(i, t), the adjusted pledged quantity P'(j), and the node area information R(j), and construct the node-substation mapping matrix M(j, i); for each node j, determine its geographical coordinates (x_j, y_j) according to its area information R(j); for each substation i, obtain its geographical coordinates (x_i, y_i) from the grid topology database; calculate the geographical distance d(j, i) = √((x_j - x_i)² + (y_j - y_i)²) from each node j to each substation i; set a threshold d_threshold according to the distance. When d(j, i) ≤ d_threshold, set M(j, i) = 1, indicating that node j is associated with substation i, otherwise set M(j, i) = 0; for nodes whose distances to multiple substations are all less than the threshold, select the substation with the closest distance as the primary association and the second-closest as the secondary association, and assign M(j, i_primary) = 1 and M(j, i_secondary) = 0.5 respectively; output the node-substation mapping matrix M(j, i) for subsequent load state adjustment calculation;
[0061] Step S142: Receive the load status S(t) and the node-substation mapping matrix M(j, i). Query the preset weight adjustment strategy table according to the current load status S(t) to assign a benchmark adjustment factor for different load statuses. When the load status S(t) = HH (high load and high fluctuation), set the benchmark adjustment factor α_base = 0.7; when the load status S(t) = HL (high load and low fluctuation), set the benchmark adjustment factor α_base = 0.9; when the load status S(t) = LH (low load and high fluctuation), set the benchmark adjustment factor α_base = 0.8; when the load status S(t) = LL (low load and low fluctuation), set the benchmark adjustment factor α_base = 1.0. Considering the distribution network topology and security constraints, use advanced power system analysis software to calculate the power supply importance index SI(i) of each substation. Combine the benchmark adjustment factor and the power supply importance index to calculate the substation weight adjustment factor α(i, S(t)) = α_base·(0.8 + 0.2·SI(i)) for each substation i. Output the substation weight adjustment factor α(i, S(t)) for fine adjustment of node weights.
[0062] Step S143: Receive the substation load data LS+(i, t), the substation weight adjustment factor α(i, S(t)), the adjusted pledge amount P'(j), and the node-substation mapping matrix M(j, i). For each substation i, calculate its load proportion factor LP(i, t) = LS+(i, t) / max(LS+(t)), normalized to the interval [0, 1]. Calculate the deviation factor LD(i, t) = |LP(i, t) - 0.5| / 0.5, indicating the degree of load deviation from the median. Calculate the load status adjustment term β(i, t) = 1 - LD(i, t), so that substations with loads closer to the average level obtain higher weights. For each node j, the comprehensive load adjustment coefficient γ(j, t) = ∑[M(j, i)·α(i, S(t))·β(i, t)], accumulating the adjustment contributions of all its associated substations. Finally, calculate the node weight W(j, t) after load adjustment = P'(j)·γ(j, t), comprehensively considering the influence of node pledge amount and load status. Output the node weight W(j, t) after load adjustment for subsequent consensus process and adaptive sharding adjustment in S4.
[0063] According to one aspect of the present application, the time characteristic analysis in step S22 is specifically as follows:
[0064] Step S221: Receive the regularized electricity meter data M''(k,t). For each electricity meter k, extract the data subset of the most recent 7 days. For each hour h ∈ [0, 23], calculate the average electricity consumption of the same hour of each day to form the hourly average load D_avg(k,h); calculate the standard deviation of the load of the same hour of each day to form the hourly load fluctuation D_std(k,h); combine the average value and the standard deviation to generate a daily load curve D(k,h) = {D_avg(k,h), D_std(k,h)} including the mean value and the fluctuation range; apply wavelet transform to denoise the daily load curve D(k,h) to obtain the smoothed daily load curve D_smooth(k,h); calculate the normalization indexes of the daily load curve, including the peak-to-valley ratio PVR(k) = max(D_avg(k,h)) / min(D_avg(k,h)), the load factor LF(k) = avg(D_avg(k,h)) / max(D_avg(k,h)), and the peak time PT(k) = argmax_h(D_avg(k,h)); output the daily load curve D(k,h) and the daily load feature index set DF(k) = {PVR(k), LF(k), PT(k)} for subsequent similarity calculation;
[0065] Step S222: Receive the regularized electricity meter data M''(k,t). For each electricity meter k, extract the data of the most recent 4 weeks. For each day d ∈ [1, 7] of each week, calculate the total electricity consumption of that day to form the daily total electricity consumption W_total(k,d,w), where w represents which week; calculate the average electricity consumption of the same day within 4 weeks to form the intra-week daily average load W_avg(k,d); calculate the standard deviation of the electricity consumption of the same day within 4 weeks to form the intra-week daily load fluctuation W_std(k,d); combine the average value and the standard deviation to generate a weekly load curve W(k,d) = {W_avg(k,d), W_std(k,d)} including the mean value and the fluctuation range; output the weekly load curve W(k,d) and the weekly load feature index set WF(k) = {WWR(k), WCV(k)} for subsequent similarity calculation;
[0066] Step S223: Receive the regularized electricity meter data M''(k,t). For each electricity meter k, apply the Fast Fourier Transform (FFT) to convert the time-domain data into a frequency-domain representation, obtaining the original spectrum data F_raw(k,ω); identify the significant frequency components through power spectral density analysis to obtain the main spectral components F_main(k,ω); extract spectral feature parameters, including the daily cycle intensity DCI(k), weekly cycle intensity WCI(k), monthly cycle intensity MCI(k), and seasonal cycle intensity SCI(k); apply a band-pass filter to separate the fluctuation characteristics at different time scales, including short-term fluctuations (hourly level), medium-term fluctuations (daily level), and long-term fluctuations (weekly level); calculate the fluctuation intensity indicators VS_short(k), VS_medium(k), and VS_long(k) for each time scale; integrate the spectral features to form the electricity meter data spectral feature vector F(k,ω) = {DCI(k), WCI(k), MCI(k), SCI(k), VS_short(k), VS_medium(k), VS_long(k)}; output the electricity meter data spectrum feature F(k,ω) for subsequent similarity calculation.
[0067] In an embodiment of the present application, the time characteristic analysis (correction) in step S22 further includes:
[0068] Step S224: Receive the daily load curve D(k,h), the daily load characteristic index set DF(k), the weekly load curve W(k,d), the weekly load characteristic index set WF(k), and the electricity meter data spectral feature F(k,ω), and calculate the similarity of each dimension for the electricity meter pair (k1,k2); calculate the daily load curve similarity D_sim(k1,k2) = cos(D_smooth(k1,h), D_smooth(k2,h)), representing the cosine similarity of the daily load curves of the two electricity meters; calculate the weekly load curve similarity W_sim(k1,k2) = cos(W_avg(k1,d), W_avg(k2,d)), that is, the cosine similarity of the weekly load curves of the two electricity meters; calculate the daily load similarity DF_sim(k1,k2) based on the Euclidean distance of the characteristic indexes; calculate the weekly load characteristic similarity WF_sim(k1,k2) based on the Euclidean distance of the characteristic indexes; calculate the spectral feature similarity F_sim(k1,k2) = cos(F(k1,ω), F(k2,ω)) based on the cosine similarity of the spectral feature vectors; comprehensively calculate the weighted time similarity T_sim(k1,k2) for each dimension similarity; calculate the time similarity for all electricity meter pairs and construct a complete time similarity matrix T_sim(k1,k2); output T_sim(k1,k2) for subsequent spatio-temporal characteristic model construction.
[0069] According to one aspect of the present application, the construction of the spatio-temporal correlation model in step S24 is specifically as follows:
[0070] Step S241: Receive the time similarity matrix T_sim(k1,k2) and the space similarity matrix S_sim(k1,k2), and initially set the time weight w_t = 0.5 and the space weight w_s = 0.5; through historical data analysis, evaluate the influence degree of time characteristics and space characteristics on the data distribution; apply the cross-validation method to calculate the weighted spatio-temporal similarity ST_sim(k1,k2,w_t,w_s)=T_sim(k1,k2)·w_t + S_sim(k1,k2)·w_s under different weight combinations; evaluate the cohesion and separation of the clustering results under each group of weights, and use the Silhouette Coefficient as the evaluation index; select the weight combination (w_t_opt, w_s_opt) with the highest Silhouette Coefficient as the optimal weight; use the optimal weight to calculate the final spatio-temporal similarity matrix ST_sim(k1,k2)=T_sim(k1,k2)·w_t_opt + S_sim(k1,k2)·w_s_opt; for the meter pairs with similarity lower than the threshold θ_sim, set the similarity value to 0 to form a sparse spatio-temporal similarity matrix ST_sim_sparse(k1,k2); output the spatio-temporal similarity matrix ST_sim(k1,k2) and the sparse spatio-temporal similarity matrix ST_sim_sparse(k1,k2) for subsequent clustering analysis;
[0071] Step S242: Receive the spatio-temporal similarity matrix ST_sim(k1,k2), and construct a similarity graph G between meters, where the meters are nodes and the similarity is the edge weight; apply the hierarchical clustering algorithm to process the graph G. Initially, each meter is an independent cluster; iteratively merge the two most similar clusters, and use the Complete Linkage method to calculate the distance between clusters, that is, the similarity of the least similar meter pair between clusters; through the analysis of the cutting height of the clustering tree, determine the optimal number of clusters C_opt; use the Silhouette Coefficient and the Davies-Bouldin Index to evaluate the effectiveness of different numbers of clusters; fine-tune the number of clusters within the range of C_opt, considering the physical structure of the power grid and the load balancing requirements; finally, divide the meters into C clusters to form a meter clustering set CL={CL(1), CL(2),..., CL(C)}, where CL(c) represents the set of meters belonging to cluster c; assign a cluster label to each meter k, and establish a meter clustering mapping MC(k)=c, indicating that meter k belongs to cluster c; output the meter clustering set CL and the meter clustering mapping MC(k) for subsequent feature vector extraction and sharding strategy design;
[0072] Step S243: Receive the electricity meter clustering set CL. For each cluster c ∈ [1, C], calculate the feature vector CV(c) of this cluster; extract the cluster center electricity meter k_center(c) = argmin_k∈CL(c)∑_k'∈CL(c)(1 - ST_sim(k, k')), and select the electricity meter with the highest sum of similarities with other electricity meters in the cluster as the center; extract the time feature sub-vector CVT(c) of the cluster, which includes the daily load curve D(k_center(c), h), weekly load curve W(k_center(c), d) and spectral feature F(k_center(c), ω) of the cluster center; extract the spatial feature sub-vector CVS(c) of the cluster, which includes the geographical location G(k_center(c)) of the cluster center, the affiliated substation ID and the grid topological location; calculate the statistical feature sub-vector CVStat(c) of the cluster, which includes the cluster size |CL(c)|, the average similarity avg_k,k'∈CL(c)(ST_sim(k, k')) of the electricity meters in the cluster and the cluster radius max_k∈CL(c)(1 - ST_sim(k, k_center(c))); calculate the load feature sub-vector CVL(c) of the cluster, which includes the total average load, peak load and valley load of the cluster; integrate each sub-vector, calculate and form the complete cluster feature vector CV(c) = {CVT(c), CVS(c), CVStat(c), CVL(c)}; output all cluster feature vector sets CV = {CV(1), CV(2), ..., CV(C)} for subsequent model construction and sharding strategy design;
[0073] Step S244: Receive the spatio-temporal similarity matrix ST_sim(k1,k2), the electricity meter clustering set CL, the electricity meter clustering mapping MC(k), and the feature vector set CV, and construct a complete spatio-temporal characteristic model STM of electricity meter data; design a multi-layer model structure, including a similarity layer, a clustering layer, and a feature layer; store the spatio-temporal similarity matrix ST_sim(k1,k2) and the sparse spatio-temporal similarity matrix ST_sim_sparse(k1,k2) in the similarity layer; store the electricity meter clustering set CL, the electricity meter clustering mapping MC(k), and the similarity matrix CS(c1,c2)=avg_k1∈CL(c1),k2∈CL(c2)(ST_sim(k1,k2)) in the clustering layer; store the feature vector set CV and the feature importance weights in the feature layer; analyze the time stability of the clustering structure and verify the persistence of the clustering through historical data; establish a clustering evolution prediction model to predict possible changes in the future clustering structure; design a clustering maintenance strategy, including the timing of regular updates and the incremental update method; integrate all components and output a complete spatio-temporal characteristic model STM of electricity meter data, including the data structures and maintenance strategies of all levels; transfer the spatio-temporal characteristic model STM of electricity meter data to Step S3 for subsequent sharding strategy design.
[0074] According to one aspect of the present application, the spatio-temporal aware sharding division in Step S32 is specifically as follows:
[0075] Step S321: Receive the spatio-temporal characteristic model STM of electricity meter data, the node performance index PI(j), and the node geographical location GL(j), extract the node computing power, storage capacity, and network bandwidth data from the node performance index PI(j); extract the geographical distribution information of the electricity meter clustering from the spatio-temporal characteristic model STM of electricity meter data; construct a node geographical location matrix, where the coordinates of each node j are (x_j, y_j); calculate the geographical distance matrix GD(j1,j2) between nodes; combine ND(j1,j2) and the geographical distance matrix to calculate the comprehensive distance CD(j1,j2)=w_nd·ND(j1,j2) / ND_max+w_gd·GD(j1,j2) / GD_max, where w_nd and w_gd are weight coefficients and w_nd+w_gd=1; use the comprehensive distance matrix CD(j1,j2) as the input, apply the K-means clustering algorithm to divide the network nodes into K node clusters NC(k), k∈[1,K]; output the node clusters NC(k) for subsequent affinity calculation and allocation;
[0076] Step S322: Receive the node cluster NC(k) and the spatio-temporal characteristic model STM of the electricity meter data. Extract the electricity meter clustering set CL and the feature vector set CV from STM. For each electricity meter cluster c and node cluster k, calculate their affinity A(c,k). Considering the geographical location factor, calculate the distance d_geo(c,k) between the geographical center of the electricity meter cluster c and the geographical center of the node cluster k, and convert it to the geographical affinity A_geo(c,k)=1 / (1 + d_geo(c,k)). Considering the network latency factor, calculate the average network latency d_net(c,k) between the electricity meters in the electricity meter cluster c and the nodes in the node cluster k, and convert it to the network affinity A_net(c,k)=1 / (1 + d_net(c,k)). Considering the load matching factor, calculate the matching degree between the data volume and access frequency of the electricity meter cluster c and the computing power and storage capacity of the node cluster k to obtain the load affinity A_load(c,k). Considering all factors, calculate the total affinity A(c,k)=w_geo·A_geo(c,k)+w_net·A_net(c,k)+w_load·A_load(c,k), where w_geo, w_net, and w_load are weight coefficients and w_geo + w_net + w_load = 1. Construct the affinity matrix A, which contains the affinity values between all electricity meter clusters and node clusters. Output the affinity matrix A for solving the optimal allocation problem.
[0077] Step S323: Receive the affinity matrix A and the electricity meter clustering set CL. Model the problem of allocating electricity meter clusters to node clusters as an optimal allocation problem. Define the objective function as maximizing the total affinity ∑_c,k A(c,k)·X(c,k), where X(c,k) is a 0-1 variable indicating whether the electricity meter cluster c is allocated to the node cluster k. Add the constraint conditions: each electricity meter cluster can only be allocated to one node cluster ∑_k X(c,k)=1, c; the load of each node cluster does not exceed its capacity limit ∑_cLoad(c)·X(c,k)≤Cap(k), k. Use the Hungarian algorithm or integer linear programming method to solve the optimal allocation problem. Obtain the optimal allocation scheme X*(c,k), and based on this, establish the mapping MCN(c)=k from the electricity meter cluster to the node cluster, indicating that the electricity meter cluster c is allocated to the node cluster k. Check the load balance, calculate the load ratio of each node cluster, and ensure that the ratio of the maximum load to the minimum load does not exceed the preset threshold. If the load is unbalanced, adjust the weight coefficients and solve again until the balance requirement is met. Output the mapping MCN(c) from the electricity meter cluster to the node cluster for subsequent node responsibility allocation.
[0078] Step S324: Receive the node cluster NC(k), the mapping MCN(c) from electricity meters to node clusters, and the node performance indicator PI(j), and assign responsibilities to the nodes within each node cluster; for each node cluster k, obtain all the electricity meter clusters {c|MCN(c)=k} that the cluster is responsible for; for each electricity meter cluster c, based on the node performance indicators, select the node with the best performance in the node cluster NC(MCN(c)) as the primary shard node P(c,k) to be responsible for the main storage and processing of the data for this cluster; select the N_b nodes with the second-best performance from the remaining nodes as the backup node set B(c,k) for data backup and fault recovery; select the N_v nodes with moderate performance from the remaining nodes as the verification node set V(c,k) to be responsible for data verification; ensure the overall load balance of each node and avoid a single node taking on too many responsibilities; consider the network connection quality between nodes and ensure the minimization of the network latency between the primary shard node and the backup nodes; generate a node responsibility assignment table, which includes all the primary shard nodes P(c,k), the backup node set B(c,k), and the verification node set V(c,k); integrate all the above results and output a complete spatio-temporal optimized sharding scheme SP, which includes the node cluster NC(k), the mapping MCN(c) from electricity meters to node clusters, and the node responsibility assignment.
[0079] According to one aspect of the present application, the data distribution optimization in step S34 is specifically as follows:
[0080] Step S341: Receive the spatio-temporal optimized sharding scheme SP and the spatio-temporal characteristic model STM of electricity meter data, and obtain the electricity meter data access logs for the recent 90 days from the historical database; analyze the access logs and extract the access frequency AF(k,t) of each electricity meter k at different time points t; divide the time axis into typical time periods, such as working hours on weekdays, non-working hours on weekdays, weekends, etc.; calculate the average access frequency of each electricity meter in each time period; consider periodic patterns, such as daily, weekly, and monthly cycles, and apply time series decomposition to extract the trend, seasonal, and random components; combine the above analysis results and design a data heat evaluation function H(k,t)=w_af·AF(k,t)+w_trend·Trend(k,t)+w_season·Season(k,t), where w_af, w_trend, and w_season are weight coefficients; apply the data heat evaluation function H(k,t) to all electricity meters and calculate the data heat for the current time period and the predicted future time periods; sort the electricity meters according to the heat value and identify the hot data set HD={k|H(k,t)>θ_hot} and the cold data set CD={k|H(k,t<θ_cold}, where θ_hot and θ_cold are the hot data and cold data thresholds; output the data heat evaluation function H(k,t), the hot data set HD, and the cold data set CD for subsequent design of the hot data replication strategy;
[0081] Step S342: Receive the hot data set HD, the cold data set CD, the spatio-temporal optimized sharding scheme SP, and the spatio-temporal characteristic model STM of the electricity meter data. According to the node allocation in the spatio-temporal optimized sharding scheme SP, determine the primary node and backup node for each hot data electricity meter k; for each electricity meter k in the hot data set HD, calculate its access geographical distribution GD(k), which represents the access ratio from different regions; analyze the matching degree between the access geographical distribution and the node geographical distribution, and identify the high-access low-latency region HLR(k); select the node with the best performance within each high-access low-latency region HLR(k) as the additional replication node for the hot data; according to the data consistency requirement, design a hot data update strategy, including synchronous replication or asynchronous replication methods; consider the network bandwidth limitation, design a replication scheduling algorithm to avoid excessive network resource occupation by the replication operation; according to the replication node selection and update strategy, generate a hot data replication strategy HDR, which includes the replication node set, the replication method, and the synchronization frequency; output the hot data replication strategy HDR for achieving efficient data access and load balancing;
[0082] Step S343: Receive the hot data set HD, the cold data set CD, and the spatio-temporal characteristic model STM of the electricity meter data. Analyze the time characteristics of the electricity meter data, including the data generation rate, the change of access frequency over time, and the decay of data value over time; according to the regulatory requirements and business needs, determine the retention period for different types of data; design a data hierarchical storage strategy to divide the data into hot, warm, and cold layers according to time and access frequency; for the hot layer data, keep the original format and store it on high-performance storage nodes; for the warm layer data, apply moderate compression and selectively store it on medium-performance nodes; for the cold layer data, apply a high compression ratio algorithm and store it on large-capacity storage nodes; design a data automatic migration mechanism to automatically migrate the data between different levels according to the data age and access pattern; consider the spatio-temporal characteristics of the electricity meter data, design an aggregation storage strategy to appropriately aggregate the data of the same region, the same type of users, or the same time period to reduce the storage space; design a data cleaning strategy to automatically delete the data that exceeds the retention period and is no longer needed; integrate the above strategies to form a complete data life cycle management strategy DLM; combine the hot data replication strategy HDR and the data life cycle management strategy DLM to update the spatio-temporal optimized sharding scheme SP, and output the final spatio-temporal optimized sharding scheme SP'.
[0083] According to one aspect of the present application, the power grid load status monitoring and prediction in step S41 is specifically as follows:
[0084] Step S411: Real-time access the latest regional real-time load data L(t) and substation load data LS(i,t) from the SCADA system, and set the data collection frequency to once every 5 minutes; perform data quality checks on the newly collected load data, including range checks, continuity checks, and reasonableness checks; detect and mark outliers, such as sudden load jumps, sensor fault values, or communication interruption values; for minor anomalies, apply moving median filtering for correction; for severe anomalies, mark them as invalid and exclude them from subsequent analysis; store the processed regional real-time load data L(t) and substation load data LS(i,t) in the real-time data buffer, retaining historical data for the most recent 24 hours; simultaneously obtain the current and forecast weather data W(t) from the weather service, including temperature, humidity, and weather conditions, as auxiliary variables for load forecasting; output the processed regional real-time load data L(t), substation load data LS(i,t), and weather data W(t) for subsequent load status calculation and prediction;
[0085] Step S412: Receive the processed regional real-time load data L(t) and substation load data LS(i,t), and apply the methods described in Steps S11 and S12 to calculate the smoothed load data L*(t) and the load fluctuation index V(t); read the load threshold θL and the fluctuation threshold θV from the system configuration, and if the configured values do not exist, use the default values θL = 0.8·L_max and θV = 0.15·V_max; based on the thresholds, determine the current load status S(t): if L*(t) > θL and V(t) > θV, then S(t) = HH (high load and high fluctuation); if L*(t) > θL and V(t) ≤ θV, then S(t) = HL (high load and low fluctuation); if L*(t) ≤ θL and V(t) > θV, then S(t) = LH (low load and high fluctuation); if L*(t) ≤ θL and V(t) ≤ θV, then S(t) = LL (low load and low fluctuation); record the current load status S(t) with the historical status of the past 24 hours, and calculate the state change frequency and duration; judge the stability of the current load status, and if the status changes frequently within 5 minutes, use the majority voting method to determine the final status; output the current load status S(t) and the state stability index for subsequent load forecasting and adaptive adjustment;
[0086] Step S413: Receive the regional real-time load data L(t), substation load data LS(i,t), weather data W(t), and current load status S(t) for the past 24 hours, and load historical load data under similar date and weather conditions from the historical database as a reference; for the regional total load, establish a combined short-term prediction model, including the time series model ARIMA, regression model, and neural network model; for each substation load, establish an independent prediction model considering its specific load characteristics; use a rolling time window to train the model, with the window length of 7 days and update the model parameters daily; according to the data at the current moment, use the trained model to predict the predicted load data L'(t+Δt) and predicted substation load data LS'(i,t+Δt) every 5 minutes within the next 1 hour, where Δt∈[5,60] and the step size is 5 minutes; apply the method in step S412 to calculate the predicted load status S'(t+Δt) at each future time point according to the predicted load data; evaluate the confidence interval of the prediction and mark the prediction points with low confidence; output the predicted load status sequence S'(t+Δt) and its confidence level for adaptive adjustment of decisions.
[0087] According to one aspect of the present application, the dynamic adjustment of the sharding structure in step S44 is specifically implemented as follows:
[0088] Step S441: Receive the sharding adjustment flag SAF, performance index SPI, pre-adjustment plan PAP, spatio-temporal optimized sharding scheme SP', and the spatio-temporal characteristic model STM of the electricity meter data, and check the value of the sharding adjustment flag SAF; if SAF = 1, it means that the sharding adjustment needs to be executed immediately; if there is an expired pre-adjustment plan PAP, the adjustment also needs to be executed; when the sharding adjustment needs to be executed, extract the latest meter clustering and similarity information from the spatio-temporal characteristic model STM of the electricity meter data; obtain the current network load status NLS from the load monitoring system, including the CPU usage rate, memory usage rate, disk I / O, and network traffic of each node; combine the load status S(t) and the network load status NLS to evaluate the current system pressure level; based on the system pressure level, set the sharding adjustment parameters, including the adjustment intensity, adjustment range, and adjustment priority; determine whether to perform a global adjustment or a local adjustment. The global adjustment involves all shards, and the local adjustment only involves the shards with poor performance indicators; output the adjustment type AT (global / local), adjustment range AR (the set of shards involved), and adjustment parameters AP for subsequent sharding scheme generation;
[0089] Step S442: Receive the adjustment type AT, adjustment range AR, adjustment parameters AP, the spatio-temporal characteristic model STM of the electricity meter data, and the load status S(t). Select the corresponding shard reconstruction method according to the adjustment type AT. If it is a local adjustment, only execute steps S32 to S34 for the shards within the adjustment range AR. If it is a global adjustment, re-execute steps S32 to S34 for all shards. During the execution of shard reconstruction, perform specific optimizations considering the adjustment parameters AP, such as reducing the reallocation ratio and keeping key nodes unchanged. Apply the latest data of the spatio-temporal characteristic model STM of the electricity meter data, combine with the current load status S(t), and recalculate the affinity between the node clusters and the electricity meter clusters. Solve the optimal allocation problem to generate a new mapping of the electricity meter clusters to the node clusters. Reassign the node responsibilities to determine the new primary shard node, backup node, and verification node. Make corresponding adjustments to the hot data replication strategy and the data life cycle management strategy. Integrate all adjustment results to generate a new shard scheme SP_new, including the new node clusters, mapping relationships, and node responsibility assignments. Output the new shard scheme SP_new and the current spatio-temporal optimized shard scheme SP' for subsequent data migration plan formulation.
[0090] Step S443: Receive the new shard scheme SP_new and the current spatio-temporal optimized shard scheme SP'. Compare the differences between the two schemes to identify the data sets that need to be migrated. For each changed electricity meter cluster c, determine its node allocation in the old and new schemes. Calculate the amount of data to be migrated based on the historical data size and growth rate of each electricity meter. Evaluate the network link status and select the optimal migration path to avoid congested links. Assign migration priorities to different data sets according to the data importance and access frequency. Design a migration schedule and arrange the high-impact migrations during the low-load period of the system. Consider the node storage capacity and processing power to ensure that the migration process does not cause node overload. Generate a detailed data migration plan DMP, including the source node, target node, data set, migration path, priority, and planned time. Evaluate the overall impact of the migration plan, including the migration duration, network overhead, and service quality impact. Output the data migration plan DMP for subsequent transition time window design.
[0091] Step S444: Receive the data migration plan DMP, the new sharding scheme SP_new, and the current spatio-temporal optimized sharding scheme SP'. Design an appropriate transition time window TTW during which the old and new sharding schemes coexist. Calculate the shortest required transition time based on the total data volume and available network bandwidth in the data migration plan DMP. Considering the system safety margin, set the transition window to 1.5 to 2 times the shortest required time. Define the switch checkpoint SCP, which includes the conditions that must be met to complete the final switch, such as the migration completion ratio, data consistency verification, and system stability indicators. Design a rollback mechanism that can restore to the original scheme in case of serious problems during the transition. Design a dual-write mechanism that sends write operations to both the old and new nodes simultaneously during the transition to ensure data consistency. Design a read policy to clarify when to read data from the new node and when to read from the old node. Integrate all the above components to form a complete smooth transition strategy STS, which includes the data migration plan DMP, the transition time window TTW, the switch checkpoint SCP, and the rollback mechanism. Execute the smooth transition strategy STS and monitor the migration progress and system status. After meeting the conditions of the switch checkpoint, complete the final switch, update the system metadata to make the new sharding scheme effective. Update and output the adaptive sharding storage strategy ASP, which includes the updated spatio-temporal optimized sharding scheme SP_updated, the dynamic adjustment strategy DAS, and the definition of performance monitoring indicators PMI.
[0092] According to one aspect of the present application, the implementation of the distributed verification process in step S53 is specifically as follows:
[0093] Step S531: Receive the verification rule set VRS and the adaptive sharding storage strategy ASP, and design a three-layer verification architecture, including in-node verification, intra-shard verification, and cross-shard verification. In the in-node verification layer, define the basic verification tasks independently completed by each node, including data format verification, timestamp verification, and digital signature verification. Specify the verification input data format, verification algorithm, and verification result format. Design a parallel processing mechanism for the verification process to improve the verification efficiency of a single node. In the intra-shard verification layer, define the verification tasks that require cooperation among multiple nodes within a shard, including power balance verification and usage pattern verification. Design a verification task allocation mechanism and an intermediate result transmission mechanism among the nodes within a shard. Design a coordinator node selection algorithm for intra-shard verification based on node performance and load conditions. In the cross-shard verification layer, define the verification tasks involving cooperation among multiple shards, including cross-shard data consistency verification and shard boundary verification. Design a coordination mechanism for cross-shard verification to clarify the division of responsibilities of each shard. Design a scheduling mechanism for hierarchical verification to determine the trigger conditions and execution order of each layer of verification. Output a complete hierarchical verification architecture HVA, which includes the verification tasks, coordination mechanisms, and scheduling strategies of each layer for subsequent verification task allocation.
[0094] Step S532: Receive the hierarchical verification architecture HVA and the adaptive sharding storage strategy ASP, extract the performance metrics and current load status of each node from the adaptive sharding storage strategy ASP; define resource requirement models for different types of verification tasks, including computational complexity, memory requirements, and network traffic; design a verification task allocation algorithm based on resource requirements and node capabilities; for in-node verification tasks, balance the allocation of verification work according to the current data distribution and node load; for in-shard verification tasks, consider data location and network topology to minimize data transmission during verification; for cross-shard verification tasks, select nodes with good network connections and sufficient computing resources as coordination nodes; consider the priority of verification tasks to ensure that critical verification tasks obtain sufficient resources; design a dynamic task adjustment mechanism to reallocate tasks during execution according to changes in node load; consider fault tolerance, and be able to automatically reallocate tasks when the assigned node fails; combine the task allocation algorithm and the dynamic adjustment mechanism to form a complete verification task allocation algorithm VTA; output the verification task allocation algorithm VTA for subsequent design of the verification result consensus mechanism;
[0095] Step S533: Receive the verification task allocation algorithm VTA and the hierarchical verification architecture HVA, and design a verification result consensus mechanism to ensure the consistency of results when multiple nodes participate in verification; for in-node verification, no consensus is required, and the single-node verification result is directly used; for in-shard verification, design a consensus mechanism based on majority voting, and reach a consensus when the results of more than 2 / 3 of the verification nodes are the same; calculate the credibility weights of each verification node based on historical verification accuracy and node reputation; use the weighted voting method, with the node credibility as the voting weight; for cross-shard verification, design a two-phase commit protocol to ensure the consistency of verification results for all participating shards; in the first phase, each shard independently verifies and submits preliminary results; in the second phase, the coordination node collects the results of all shards, and after confirming the consistency, submits the final results; design a timeout handling mechanism for how to make decisions when some nodes fail to return results in a timely manner; design a conflict resolution mechanism for how to determine the final result when there are significant differences in verification results; integrate the above mechanisms to form a complete verification result consensus mechanism VRC; output the distributed verification process DVP, including the hierarchical verification architecture HVA, the verification task allocation algorithm VTA, and the verification result consensus mechanism VRC, for subsequent data verification execution.
[0096] According to one aspect of the present application, the data verification execution and result confirmation in step S55 are specifically as follows:
[0097] Step S551: Receive the electricity meter data, the Distributed Verification Process (DVP), and the Anomaly Detection and Handling Mechanism (ADM). According to the Hierarchical Verification Architecture (HVA) in the DVP, plan the verification execution process. First, perform basic in-node verification at the data entry point, including data format verification, timestamp verification, and digital signature verification. Check the format compliance of the electricity meter data using a preset data format template. Verify whether the timestamp is within a reasonable range, and timestamps that are too early or too late will be marked. Use a public key cryptosystem to verify the digital signature of the data to ensure that the data source is authentic and has not been tampered with. For the data that passes the basic verification, mark it as Initially Verified Data (IVD). For the data that fails the verification, record the error type and reason, and mark it as Failed Verification Data (FVD). Pass the initially verified data IVD to the subsequent in-shard verification process, and pass the failed verification data FVD to the anomaly handling component. Output the initially verified data IVD and the failed verification data FVD for subsequent verification and anomaly handling.
[0098] Step S552: Receive the initially verified data IVD and the Distributed Verification Process (DVP). According to the Verification Task Allocation Algorithm (VTA) in the DVP, allocate the in-shard verification tasks to the corresponding nodes. Organize the in-shard nodes to perform power balance verification to check whether the power generation and consumption in a given area are balanced. Apply the power balance constraint model in the Power System Physical Constraint Model (PCM) for verification. Organize the in-shard nodes to perform usage pattern verification to check whether the electricity meter data conforms to the historical usage pattern. Extract the power usage pattern library from the Power System Physical Constraint Model (PCM) for pattern matching. For the data that passes the verification, mark it as In-shard Verified Data (SVD). For the data with doubts in the verification, mark the doubts and record the detailed information to form Suspicious Anomaly Data (SAD). Collect the verification results of all verification nodes, and apply the in-shard consensus method in the Verification Result Consensus Mechanism (VRC) to determine the final verification result. Pass the in-shard verified data SVD to the subsequent cross-shard verification process, and pass the suspicious anomaly data SAD to the anomaly detection component. Output the in-shard verified data SVD and the suspicious anomaly data SAD for subsequent verification and anomaly detection.
[0099] Step S553: Receive the in - shard verification data SVD and the distributed verification process DVP. For data involving multiple shards, coordinate and organize cross - shard verification; according to the cross - shard verification rules in the distributed verification process DVP, determine the data sets that need to undergo cross - shard verification; for these data sets, perform cross - shard data consistency verification to ensure the consistency between related data stored in different shards; perform shard boundary verification to check the data consistency in the shard boundary area; apply a two - stage verification protocol: in the first stage, each shard independently verifies the data it stores; in the second stage, exchange the verification results and necessary proof data for cross - shard consistency checking; for the data that passes the verification, mark it as cross - shard verified data CVD; for the data with inconsistent verification, record the details of the inconsistency to form inconsistent data ICD; apply the cross - shard consensus method in the verification result consensus mechanism VRC to summarize the verification results of each shard and reach a final consensus; merge the data that passes all verification levels into the verified trusted data VTD; transfer the inconsistent data ICD to the exception handling component; output the verified trusted data VTD and the inconsistent data ICD for subsequent result confirmation and exception handling;
[0100] Step S554: Receive the verified trusted data VTD, verification - failed data FVD, suspected abnormal data SAD, inconsistent data ICD, and the anomaly detection and handling mechanism ADM. Apply the anomaly detection and handling mechanism ADM to process the verification - failed data FVD, suspected abnormal data SAD, and inconsistent data ICD; according to the anomaly grading strategy in the anomaly detection and handling mechanism ADM, classify the abnormal data into different severity levels; for minor anomalies, apply automatic repair rules to correct them to form corrected data CD; for medium anomalies, mark and record logs but still retain the data to form flagged abnormal data FAD; for severe anomalies, trigger the alarm mechanism and isolate the data to form isolated abnormal data IAD; re - incorporate the corrected data CD into the verified trusted data VTD; summarize the statistical information of the verification process, including the total number of verifications, pass rate, abnormal type distribution, and handling methods; generate a detailed verification report VRep, including the verification process, result summary, and handled abnormal situations; store the verification report VRep on the blockchain to ensure its immutability; provide the verified trusted data VTD to the business application system for use; at the same time, update the system's anomaly pattern library by incorporating the newly discovered anomaly patterns; output the final verified trusted data VTD and the verification report VRep to complete the entire data verification process.
[0101] In another embodiment of the present application, a distributed storage method for electricity meter data based on improved PoS consensus further includes:
[0102] Periodically obtain the latest regional real-time load data and substation load data, calculate the current load status, and predict the predicted load status sequence for a period of time in the future;
[0103] According to the load status, the predicted load status sequence, and the node weights, adjust the consensus parameters and update the validator selection probability;
[0104] Evaluate the performance metrics of the current spatio-temporal optimized sharding scheme. When the performance metrics exceed the threshold, trigger sharding adjustment;
[0105] Based on the latest spatio-temporal characteristics model of the electricity meter data and the load status, redesign the sharding scheme, formulate a smooth transition strategy, complete the sharding adjustment, and output an adaptive sharding storage strategy;
[0106] Among them, the steps of predicting the predicted load status sequence for a period of time in the future include:
[0107] Periodically receive the latest regional real-time load data and substation load data; apply a preprocessing method to calculate the current load status; use a time series prediction model to process historical load data and predict the predicted load status sequence for a period of time in the future.
[0108] Among them, the steps of evaluating the performance metrics of the current spatio-temporal optimized sharding scheme include:
[0109] Receive the spatio-temporal optimized sharding scheme, the load status, and the predicted load status sequence; calculate the performance metrics of the current sharding scheme, including the cross-shard transaction ratio, the node load balance degree, and the data access latency; design a sharding adjustment threshold function based on the load status; when the performance metrics exceed the threshold, trigger sharding adjustment; when the predicted load status shows a significant change in the future, formulate a pre-adjustment plan;
[0110] Among them, the steps of completing the sharding adjustment include:
[0111] Receive the sharding adjustment flag, the performance metrics, the pre-adjustment plan, the spatio-temporal optimized sharding scheme, and the spatio-temporal characteristics model of the electricity meter data; when sharding adjustment needs to be executed, redesign the sharding scheme based on the latest status; formulate a smooth transition strategy, including a data migration plan, a transition time window, and a switch checkpoint; execute the transition strategy, complete the sharding adjustment, and update and output an adaptive sharding storage strategy.
[0112] In another embodiment of the present application, the improvement of the PoS consensus mechanism for power grid load perception in step S1 includes:
[0113] Step S11, Power grid load data acquisition and preprocessing:
[0114] Obtain the regional real-time load data L(t) and substation load data LS(i,t) (where i is the substation identifier and t is the timestamp) from the SCADA system; apply a moving window average to the regional real-time load data L(t) with a window length of T = 15 minutes to calculate the smoothed load data L*(t); calculate the load fluctuation index V(t) by computing V(t) = √(1 / (T - 1)·∑(L(τ)-L*(t))²), where τ ∈ [t - T, t]; output the smoothed load data L*(t) and the load fluctuation index V(t) for subsequent load state classification;
[0115] S12. Load state classification and evaluation:
[0116] Input the smoothed load data L*(t) and the load fluctuation index V(t); based on the pre-set thresholds θL and θV, define the load state matrix: if L*(t)>θL and V(t)>θV, then the load state S(t) = high load and high fluctuation (HH), if L*(t)>θL and V(t) ≤ θV, then the load state S(t) = high load and low fluctuation (HL), if L*(t) ≤ θL and V(t)>θV, then the load state S(t) = low load and high fluctuation (LH), if L*(t) ≤ θL and V(t) ≤ θV, then the load state S(t) = low load and low fluctuation (LL); output the current load state S(t) and the substation load data LS(i,t) for subsequent weight adjustment;
[0117] S13. Node pledge information processing:
[0118] Obtain the original pledge amount P(j) of each node j from the blockchain network; obtain the regional information R(j) and historical reliability index H(j) of each node j; calculate the adjusted pledge amount P'(j) according to the formula P'(j)=P(j) · (0.8 + 0.2·H(j)); output the adjusted pledge amount P'(j) and the node regional information R(j) for load perception weight calculation;
[0119] S14. Load perception weight calculation:
[0120] Input the load state S(t), substation load data LS(i,t), adjusted pledge amount P'(j), and node regional information R(j); associate the nodes with the corresponding substation i according to the node regional information R(j); design a weight adjustment factor α(S(t)) based on different load states: if S(t)=HH, then α(S(t)) = 0.7, if S(t)=HL, then α(S(t))=0.9, if S(t)=LH, then α(S(t))=0.8, if S(t)=LL, then α(S(t))=1.0;
[0121] Calculate the load status adjustment term β(j,t) = (1 - |LS(i,t) / max(LS(*,t)) - 0.5| / 0.5); finally calculate the node weight W(j,t) after load adjustment as W(j,t) = P'(j) · α(S(t)) · β(j,t); output the node weight W(j,t) after load adjustment for subsequent consensus process and adaptive shard adjustment in S4.
[0122] According to one aspect of the present application, the spatio-temporal characteristic analysis and model construction of S2 electricity meter data include:
[0123] S21. Electricity meter data collection and preprocessing:
[0124] Obtain the historical raw electricity meter data M(k,t) (k is the electricity meter ID, t is the timestamp) of the past 30 days from the electricity meter database; perform data cleaning, detect and remove outliers to obtain the cleaned electricity meter data M'(k,t); perform time regularization on the cleaned electricity meter data M'(k,t) to unify the sampling interval to 15 minutes, obtaining the regularized electricity meter data M''(k,t); output the regularized electricity meter data M''(k,t) for time characteristic analysis;
[0125] S22. Time characteristic analysis:
[0126] Input the regularized electricity meter data M''(k,t); for each electricity meter k, extract the daily load curve D(k,h) (h ∈ [0,23]), representing the electricity consumption pattern within 24 hours of a day; for each electricity meter k, extract the weekly load curve W(k,d) (d ∈ [1,7]), representing the electricity consumption pattern within 7 days of a week; use Fourier transform to extract the spectral characteristics F(k,ω) of the electricity meter data; calculate the time similarity matrix T_sim(k1,k2) = cos(D(k1),D(k2)) · 0.7 + cos(W(k1),W(k2)) · 0.3; output the time similarity matrix T_sim(k1,k2) for subsequent spatio-temporal characteristic model construction;
[0127] S23. Space characteristic analysis: Obtain the electricity meter geographical location information G(k) and the power grid topological structure T from the power grid topology database; based on the power grid topological structure T, calculate the electrical distance E(k1,k2) between electricity meters; based on the electricity meter geographical location information G(k), calculate the geographical distance GD(k1,k2) between electricity meters; comprehensively consider the electrical distance and geographical distance, and calculate the space similarity matrix S_sim(k1,k2) = 1 / (1 + 0.6·E(k1,k2) + 0.4·GD(k1,k2) / GD_max); output the space similarity matrix S_sim(k1,k2) for subsequent spatio-temporal characteristic model construction;
[0128] S24, Spatiotemporal Correlation Model Construction:
[0129] Input the time similarity matrix T_sim(k1,k2) and the space similarity matrix S_sim(k1,k2); calculate the comprehensive spatiotemporal similarity ST_sim(k1,k2) = T_sim(k1,k2) · w_t + S_sim(k1,k2) · w_s, where w_t and w_s are the time and space weights, initially set as w_t = 0.5 and w_s = 0.5; apply the hierarchical clustering algorithm, based on the spatiotemporal similarity ST_sim(k1,k2), divide the electricity meters into C clusters; extract the feature vectors CV(c) (c ∈ [1, C]) of each cluster, including the time and space features of the cluster center; construct the spatiotemporal characteristic model STM of electricity meter data, including: the spatiotemporal similarity matrix ST_sim(k1,k2), the electricity meter clusters C and their feature vectors CV(c), the spatiotemporal clustering mapping relationship MC(k) of the electricity meters, indicating which cluster the electricity meter k belongs to; output the spatiotemporal characteristic model STM of electricity meter data for subsequent sharding strategy design.
[0130] According to one aspect of the present application, the spatiotemporal-aware sharding strategy design in S3 includes:
[0131] S31, Network Node Status Acquisition:
[0132] Obtain the list N of currently active network nodes and the computing power CP(j), storage capacity SC(j), and network bandwidth NB(j) of each node j from the blockchain network monitoring module; calculate the comprehensive performance index of each node
[0133] PI(j) = w_cp·CP(j) / CP_max + w_sc·SC(j) / SC_max + w_nb·NB(j) / NB_max; where w_cp, w_sc, and w_nb are the weights of computing power, storage capacity, and network bandwidth respectively, and satisfy w_cp + w_sc + w_nb = 1; obtain the geographical location information GL(j) of each node j and the network delay matrix ND(j1,j2); output the node performance index PI(j), the node geographical location GL(j), and the network delay matrix ND(j1,j2) for subsequent sharding division;
[0134] S32, Spatiotemporal-Aware Sharding Division:
[0135] Input the spatio-temporal characteristic model STM of electricity meter data, the node performance index PI(j), the geographical location GL(j) of the node, and the network delay matrix ND(j1,j2); Cluster the active nodes N using the K-means algorithm based on the geographical location GL(j) and the network delay matrix ND(j1,j2) to obtain K node clusters NC(k) (k ∈ [1, K]);
[0136] Allocate electricity meter clustering to each node cluster. The optimization principle is: minimize the network delay of electricity meter data access and balance the load of each node cluster. The specific algorithm is as follows:
[0137] Calculate the affinity A(c, k) between each electricity meter clustering c and each node cluster k; Solve the optimal allocation problem based on the affinity matrix A to obtain the mapping MCN(c)=k of the electricity meter clustering to the node cluster; Subdivide the node responsibilities within each node cluster, including the primary shard node P(c, k), the backup node B(c, k), and the verification node V(c, k);
[0138] Output the spatio-temporal optimized sharding scheme SP, including: the node cluster NC(k), the mapping MCN(c) of the electricity meter clustering to the node cluster, and the node responsibility allocation P(c, k), B(c, k), V(c, k) within each node cluster;
[0139] S33. Design of the inter-shard communication protocol:
[0140] Input the spatio-temporal optimized sharding scheme SP; Design the intra-shard communication protocol, including the data synchronization message format DM (internal) and the consensus message format CM (internal); Design the inter-shard communication protocol, including the cross-shard query message format QM (cross-shard) and the cross-shard transaction message format TM (cross-shard); Design the shard routing table SRT, including the responsible scope and communication address of each shard; Design the inter-shard data consistency protocol SICP to handle data update operations involving multiple shards; Output the shard communication protocol set SCP, including all the above communication protocols and the routing table;
[0141] S34. Data distribution optimization:
[0142] Input the spatio-temporal optimized sharding scheme SP and the spatio-temporal characteristic model STM of electricity meter data; According to the electricity meter data access pattern, design the data heat evaluation function H(k, t) to evaluate the access frequency of electricity meter k at time t; According to the data heat evaluation function H(k, t), replicate the frequently accessed data among nodes to form the hot data replication strategy HDR; Based on the time characteristics of electricity meter data, design the data life cycle management strategy DLM, including data compression, archiving, and cleaning strategies; Output the final spatio-temporal optimized sharding scheme SP'.
[0143] The update includes: the original spatio-temporal optimized sharding scheme SP, the hot data replication strategy HDR, and the data life cycle management strategy DLM.
[0144] According to one aspect of the present application, the S4 adaptive consensus and sharding dynamic adjustment mechanism includes:
[0145] S41. Grid load status monitoring and prediction:
[0146] Periodically (every 5 minutes), input the latest regional real-time load data L(t) and substation load data LS(i,t); apply the methods of S11 and S12 to calculate the current load status S(t); use a time series prediction model (such as ARIMA) to process historical load data and predict the predicted load status S'(t+Δt) for the next 1 hour, where Δt∈[5,60] and the step size is 5 minutes; output the current load status S(t) and the predicted load status sequence S'(t+Δt) for adaptive adjustment decision-making;
[0147] S42. Adaptive adjustment of consensus parameters:
[0148] Input the load status S(t), the predicted load status sequence S'(t+Δt), and the node weight W(j,t) after load adjustment; design a consensus parameter adjustment strategy CAP according to the load status: for the high-load and high-fluctuation (HH) state, increase the number of verification nodes and reduce the block generation speed; for the high-load and low-fluctuation (HL) state, maintain the standard number of verification nodes and the standard block generation speed; for the low-load and high-fluctuation (LH) state, slightly increase the number of verification nodes and the standard block generation speed; for the low-load and low-fluctuation (LL) state, reduce the number of verification nodes and increase the block generation speed; specifically calculate the adjusted number of verification nodes VN(t)=VN_base · f_vn(S(t)), where f_vn is an adjustment function based on the load status; calculate the adjusted block time interval BT(t)=BT_base · f_bt(S(t)), where f_bt is an adjustment function based on the load status; update the consensus validator selection probability VS(j,t)=W(j,t) / ∑W(j,t); output the adjusted consensus parameters CP(t), including: the number of verification nodes VN(t), the block time interval BT(t), and the validator selection probability VS(j,t);
[0149] S43. Evaluation of sharding structure dynamic adjustment:
[0150] Input the time-space optimized sharding scheme SP', load status S(t), and predicted load status sequence S'(t+Δt); calculate the performance indicators SPI of the current sharding scheme, including: cross-shard transaction proportion CTP, node load balance NLB, and data access latency DAL; design a sharding adjustment threshold function SAT(S(t)), and set different adjustment thresholds based on the load status; when the performance indicator exceeds the threshold (i.e., SPI>SAT(S(t))), trigger sharding adjustment and set the sharding adjustment flag SAF=1; when the predicted load status indicates significant changes in the future, plan sharding adjustment in advance and set the pre-adjustment plan PAP; output the sharding adjustment flag SAF, performance indicator SPI, and pre-adjustment plan PAP (if any);
[0151] S44. Execute dynamic adjustment of the sharding structure:
[0152] Input the sharding adjustment flag SAF, performance indicator SPI, pre-adjustment plan PAP (if any), time-space optimized sharding scheme SP', and spatio-temporal characteristic model STM of the electricity meter data; if SAF=1 or there is an expired PAP, execute sharding adjustment: based on the latest spatio-temporal characteristic model STM of the electricity meter data and load status S(t), re-execute the steps from S32 to S34; generate a new sharding scheme SP_new; design a smooth transition strategy STS, including: data migration plan DMP, to determine which data needs to be migrated to the new node; transition time window TTW, during which the old and new sharding schemes coexist; switching checkpoint SCP, to complete the switch after meeting specific conditions; execute the smooth transition strategy STS to complete the sharding adjustment; update and output the adaptive sharding storage strategy ASP, including: updated time-space optimized sharding scheme SP_updated; dynamic adjustment strategy DAS, including trigger conditions and adjustment rules; performance monitoring metric definition PMI.
[0153] According to one aspect of the present application, the S5 data consistency guarantee and verification mechanism includes:
[0154] S51. Model the physical constraints of the power system:
[0155] Obtain the power grid topology NT and electrical parameters EP from the power system model library; establish a power flow model PFM to verify the physical rationality of the electricity meter data; extract the electricity usage pattern library EPL from historical data, including typical electricity usage patterns of various users; establish an electricity balance constraint model EBM to express the balance relationship between input electricity and consumed electricity; output the power system physical constraint model PCM, including all the above models, for subsequent data verification;
[0156] S52. Design blockchain data verification rules:
[0157] Input the physical constraint model PCM of the power system and the adaptive sharding storage strategy ASP; design basic verification rules, including data format verification DFV, timestamp verification TSV, and digital signature verification DSV; design verification rules based on physical constraints: power balance verification EBV: check whether the power production and consumption in a given area are balanced; power flow constraint verification PFV: verify whether the power meter data conforms to the power flow constraints; usage pattern verification UPV: verify whether the power usage pattern is consistent with the historical pattern; design inter-shard consistency verification rules: cross-shard data consistency verification CDCV: verify whether the data involving multiple shards is consistent; shard boundary verification SBV: verify the data consistency of the shard boundary area; output the verification rule set VRS, including all verification rules and their execution conditions and priorities;
[0158] S53. Distributed verification process implementation:
[0159] Input the verification rule set VRS and the adaptive sharding storage strategy ASP; design a hierarchical verification architecture HVA: in-node verification NIV: basic verification completed by a single node; intra-shard verification SIV: verification completed by node cooperation within a shard; cross-shard verification CSV: verification involving cooperation of multiple shards; design a verification task allocation algorithm VTA to allocate verification tasks according to the verification type and node capabilities; design a verification result consensus mechanism VRC to ensure the consistency of verification results; output the distributed verification process DVP, including the verification architecture, task allocation, and result consensus mechanism;
[0160] S54. Anomaly detection and handling:
[0161] Input the power meter data and the distributed verification process DVP; design an anomaly pattern library APL, including common anomaly patterns such as data tampering, equipment failures, etc.; implement an anomaly detection algorithm ADA, including: rule-based detection RBD: detect obvious anomalies based on predefined rules; statistical anomaly detection SAD: detect data deviating from the normal distribution based on statistical methods; machine learning anomaly detection MLD: use an anomaly detection model to identify complex anomaly patterns; design an anomaly handling process AHP: anomaly classification AC: classify anomalies into different severity levels; handling strategy HS: handling methods for different levels of anomalies; recovery mechanism RM: a mechanism to recover from the abnormal state; output the anomaly detection and handling mechanism ADM, including the detection algorithm and the handling process;
[0162] S55, Data Verification Execution and Result Confirmation: Input the electricity meter data, the distributed verification process DVP, and the anomaly detection and handling mechanism ADM; Execute the verification process: Perform basic verification using the in-application node verification NIV; Organize the nodes within the shard to execute the in-shard verification SIV; Coordinate multiple shards to execute the cross-shard verification CSV if necessary; Collect the verification results VR(j) of all verification nodes; Apply the verification result consensus mechanism VRC to synthesize the results of each node to obtain the final verification result FVR; If an anomaly is detected, trigger the anomaly handling process AHP; Mark the data that passes the verification as the verified trustworthy data VTD; Generate a verification report VRep, including the verification process, results, and the anomalies handled; Output the verified trustworthy data VTD and the verification report VRep.
[0163] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solution of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A distributed storage method for electric energy meter data based on improved PoS consensus, characterized in that: Methods include: Obtain the electric energy meter data, perform spatiotemporal characteristic analysis, form electric energy meter clusters, and construct a spatiotemporal characteristic model of the electric energy meter data; Obtain blockchain network node information, calculate node performance indicators and network delay between nodes; Based on the spatiotemporal characteristic model of electric energy meter data, node performance indicators and network delay, nodes are grouped into node clusters, and the affinity between meter clusters and node clusters is calculated; Assign meters to node clusters based on affinity and assign responsibilities to nodes; When working, it obtains and formulates hot data replication strategy and data lifecycle management strategy based on the data access characteristics of each meter cluster and the designed data heat evaluation function, and outputs the spatiotemporal optimized sharding solution; The steps of analyzing the spatiotemporal characteristics of the electric energy meter data, forming electric energy meter clusters, and constructing the spatiotemporal characteristic model of the electric energy meter data include: Obtain the original data of historical electric energy meters, perform data cleaning and time regularization, and obtain regularized electric energy meter data; Extract daily load curve and weekly load curve from the regularized meter data and calculate the time similarity matrix; Obtain the geographical location information of the electric meters and the topological structure of the electric grid from the electric grid topology database, calculate the electrical distance and geographical distance between the electric meters, and generate a spatial similarity matrix; Combine the time similarity matrix and the space similarity matrix, calculate the time and space similarity, apply the hierarchical clustering algorithm to form meter clusters, extract cluster feature vectors, and build the time and space characteristic model of electric energy meter data; The steps of grouping nodes to form node clusters and calculating the affinity between the meter clusters and the node clusters include: Receive the spatiotemporal characteristic model of electric energy meter data, node performance indicators, node geographic location and network delay matrix; Based on the node geographic location and network delay matrix, a clustering algorithm is applied to cluster the nodes into node clusters; Calculate the affinity between the meter cluster and the node cluster, which takes into account the geographical location, network delay and load matching factors; The specific method for calculating the affinity between the meter cluster and the node cluster is: Receive the node cluster and the spatiotemporal characteristic model of the electric energy meter data, and extract the electric energy meter cluster set and the feature vector set therefrom; For each meter cluster and node cluster, calculate: The distance between the geographic center of the meter cluster and the geographic center of the node cluster is converted into geographic affinity; The average network delay between the meters in the meter cluster and the nodes in the node cluster is converted into network affinity; The matching degree between the data volume and access frequency of the meter cluster and the computing power and storage capacity of the node cluster is used to obtain the load affinity; Based on geographic affinity, network affinity and load affinity, a total affinity is formed; Based on the affinity values between all meter clusters and node clusters, an affinity matrix is constructed and output; The steps to design a data heat evaluation function and formulate a hot data replication strategy and a data lifecycle management strategy include: Receive the spatiotemporal optimized sharding scheme and spatiotemporal characteristic model of electric energy meter data of the previous cycle; Analyze meter data access logs, design data heat evaluation functions, and evaluate the access frequency of meter data; Based on the data heat evaluation function, identify the hot data set and the cold data set; Select additional replication nodes and formulate a hot data replication strategy based on the geographic distribution of access to the hot data set; Analyze the time characteristics of meter data and form a data lifecycle management strategy based on the constructed data hierarchical storage strategy and data automatic migration mechanism; Update the spatiotemporal optimized sharding scheme and output the final spatiotemporal optimized sharding scheme.
2. According to claim 1, a distributed storage method for electric energy meter data based on improved PoS consensus is characterized in that: The step of calculating the time similarity matrix comprises: Receive the regularized meter data and extract the daily load curve and weekly load curve of each meter; Use Fourier transform to extract the spectrum characteristics of meter data; Based on the daily load curve, weekly load curve and spectrum characteristics, the time similarity matrix between electricity meters is calculated.
3. According to claim 2, a distributed storage method for electric energy meter data based on improved PoS consensus is characterized in that: The method of extracting the spectrum characteristics of the meter data using Fourier transform is specifically as follows: Receive the regularized meter data, and for each meter, apply fast Fourier transform to convert the time domain data into frequency domain representation to obtain the original spectrum data; Through power spectral density analysis, significant frequency components are identified from the original spectrum data to obtain the main spectrum components; Extracting spectrum characteristic parameters, including daily cycle intensity, weekly cycle intensity, monthly cycle intensity and seasonal cycle intensity; Apply bandpass filters to separate the fluctuation characteristics of different time scales, including short-term fluctuations, medium-term fluctuations, and long-term fluctuations; The fluctuation intensity index of each time scale is calculated, and the spectrum characteristic parameters and the fluctuation intensity index are integrated to form the spectrum feature vector of the meter data.
4. According to claim 1, a distributed storage method for electric energy meter data based on improved PoS consensus is characterized in that: The specific method of combining the temporal similarity matrix and the spatial similarity matrix to calculate the temporal and spatial similarity is: Receive the temporal similarity matrix and the spatial similarity matrix, and set the initial temporal weight and spatial weight; The cross-validation method was applied to try different combinations of time weights and space weights, calculate the weighted spatiotemporal similarity, and evaluate the cohesion and separation of the clustering results under each set of weights, using the silhouette coefficient as the evaluation index; Select the weight combination with the highest silhouette coefficient as the optimal weight; Calculate the final spatiotemporal similarity matrix using the optimal weights; For meter pairs whose similarity is lower than the threshold, their similarity values are reset to 0 to form a sparse spatiotemporal similarity matrix.
5. The method for distributed storage of electric energy meter data based on improved PoS consensus according to claim 1 is characterized in that: The method also includes: Obtain regional real-time load data and substation load data, and calculate load fluctuation indicators; Determine the current load status based on regional real-time, substation load data and load fluctuation indicators combined with preset thresholds; Obtain the original staked amount, regional information, and historical reliability indicators of the node, and calculate the adjusted staked amount; Calculate the load-adjusted node weight based on the load status, substation load data, adjusted pledge amount and node area information; Wherein, the step of calculating the load fluctuation index comprises: Obtain regional real-time load data and substation load data from the SCADA system; Apply sliding window averaging to the regional real-time load data to obtain smoothed load data; The load fluctuation index is obtained by calculating the square root of the mean of the squares of the differences between the regional real-time load data and the smoothed load data.
6. The method for distributed storage of electric energy meter data based on improved PoS consensus according to claim 5 is characterized in that: The step of determining the current load state comprises: receiving smoothed load data and load fluctuation indicators, and classifying the grid load status into high load and high fluctuation, high load and low fluctuation, low load and high fluctuation, or low load and low fluctuation according to preset load thresholds and fluctuation thresholds; Output current load status and substation load data; The step of calculating the load-adjusted node weights comprises: Receive load status, substation load data, adjusted stake amount, and node area information; According to the node area information, a mapping matrix between nodes and substations is constructed to associate the nodes with the corresponding substations; Set weight adjustment factors according to different load states and calculate load state adjustment items; Based on the adjusted stake amount and load status adjustment item, the load-adjusted node weight is calculated.
Citation Information
Patent Citations
Household intelligent electric meter data storage system and method based on alliance block chain
CN113468551A
Cloud data placement strategy management system introducing SaaS features
CN116319815A