Dynamic fragmentation block chain consensus method based on depth reinforcement

Through the dynamic sharded blockchain consensus method based on deep reinforcement learning, the problem of insufficient adaptability among cross-slices in the blockchain sharded consensus mechanism is solved, precise quantification of strategy heterogeneity and improvement of consensus performance are achieved, and the stability and efficiency of the blockchain network are enhanced.

CN120455469AInactive Publication Date: 2025-08-08CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510788472.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing blockchain shard consensus mechanism lacks dynamic adaptability across shards, resulting in an intensification of heterogeneity between nodes in network topology, transaction traffic density and reward feedback frequency, and it is impossible to effectively identify and regulate the evolution of strategy between different shards.

Method used

By collecting historical blockchain shard structure data, identifying the timing correlation characteristics of node behavior and consensus rewards, determining the strategy heterogeneity level, performing heterogeneity reconstruction analysis, using deep reinforcement learning to drive strategy iteration and adaptive adjustment of reward functions, calculating cross-shash consensus delay fluctuation indicators, and triggering dynamic shard reorganization to balance the strategy optimization process between shards.

Benefits of technology

The precise quantification of the heterogeneity of inter-shafting strategies is realized, the stability and efficiency of the consensus process are enhanced, and the response coordination and adaptability of the blockchain network in a multi-shafting environment are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455469A_ABST
    Figure CN120455469A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic fragmentation block chain consensus method based on deep reinforcement, and particularly relates to the technical field of block chain consensus. The method comprises the following steps of: acquiring historical block chain fragment structure data, determining policy heterogeneity levels of a plurality of fragments, reconstructing a historical fragment structure, identifying a policy convergence imbalance section and generating a policy imbalance mark fragment set; taking the strategy imbalance mark fragment set as input, and outputting the strategy convergence state difference degree of each fragment node based on deep reinforcement learning; according to the convergence state difference degree, calculating a delay fluctuation index and comparing the delay fluctuation index with a preset threshold value; and when the delay fluctuation index exceeds a threshold value, the dynamic fragmentation structure is triggered to recombine, and a new dynamic fragmentation consensus structure is output, so that the problem of strategy asynchronous convergence caused by fragmentation isomerism can be solved, and consensus response coordination and overall performance stability of the block chain in a multi-fragmentation environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of blockchain consensus technology, and more specifically, to a dynamic sharding blockchain consensus method based on deep reinforcement. Background Art

[0002] In existing blockchain sharding consensus mechanisms, nodes typically participate in the consensus process based on local state or static policies, lacking dynamic adaptability across shards. As blockchain networks scale, significant heterogeneity emerges between nodes in terms of network topology, transaction density, and reward feedback frequency, leading to increasing disparities in the consensus behavior optimization process across shards. Existing technologies fail to effectively identify and regulate the uneven evolution of policies across shards. Summary of the Invention

[0003] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a dynamic sharding blockchain consensus method based on deep enhancement to solve the problems raised in the above-mentioned background technology.

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] A dynamic sharding blockchain consensus method based on deep reinforcement includes the following steps:

[0006] S1: Using historical blockchain sharding structure data, identify the temporal correlation characteristics between node behavior and consensus rewards in each historical shard. Based on these temporal correlation characteristics, determine the strategic heterogeneity level of multiple shards.

[0007] S2: Perform heterogeneity reconstruction analysis on historical shard structure data based on the policy heterogeneity level of each shard to identify policy convergence imbalance segments that cause consensus latency differences between shards and generate a set of policy imbalance-marked shards.

[0008] S3: Based on the policy imbalance-marked shard set, deep reinforcement learning-driven policy iteration and reward function adaptive adjustment are performed on each shard node, and the degree of difference in the policy convergence state of each shard node is output;

[0009] S4: Calculate the cross-shard consensus delay fluctuation index based on the degree of convergence status difference of each shard node strategy, and use the delay fluctuation index to determine whether to trigger dynamic shard reorganization;

[0010] S5: When the latency fluctuation index exceeds the preset threshold, a dynamic reorganization of the shard structure guided by deep reinforcement learning strategy is performed to balance the strategy optimization process between shards and output a stable dynamic shard consensus structure.

[0011] In a preferred embodiment, S1 is specifically:

[0012] Collect historical blockchain sharding structure data, and extract the historical consensus action sequence and historical consensus incentive acquisition sequence of nodes in the shard based on the historical blockchain sharding structure data;

[0013] Establish a temporal association model based on the node's historical consensus action sequence and historical consensus incentive acquisition sequence;

[0014] The temporal association model is used to quantify the strategic similarity of node behaviors in each historical shard, and the quantitative index of strategic heterogeneity of nodes in each historical shard is calculated. The strategic heterogeneity level of multiple shards is determined based on the quantitative index of strategic heterogeneity.

[0015] In a preferred embodiment, S2 is specifically:

[0016] Obtain the policy heterogeneity levels of multiple shards, and perform heterogeneous reconstruction of the node policy optimization process in the historical blockchain shard structure data based on the policy heterogeneity levels, and identify the temporal distribution characteristics of the policy updates of each shard node during the policy optimization process;

[0017] According to the temporal distribution characteristics of the strategy update of each shard node, the strategy imbalance segments with different strategy convergence rates in the node strategy optimization process are located, and the strategy imbalance segments are marked to generate a strategy imbalance marked shard set.

[0018] In a preferred embodiment, S3 is specifically:

[0019] Based on the set of policy imbalance labeled shards, a deep reinforcement learning training environment is constructed to determine the initial structure of the reinforcement learning reward function for policy optimization.

[0020] Perform node policy iterative optimization in the deep reinforcement learning training environment, and obtain the policy update data of each shard node at different times through policy iterative optimization;

[0021] Based on the initial structure of the reinforcement learning reward function, the reward function is adaptively adjusted for the policy update data of each shard node. Based on the policy change trend and fluctuation amplitude of each shard node in the policy update data, the degree of convergence state difference of each shard node strategy is quantitatively calculated.

[0022] In a preferred embodiment, S4 is specifically:

[0023] Based on the degree of convergence difference of each shard node strategy, a cross-shard consensus delay fluctuation assessment model is established;

[0024] The cross-shard consensus delay fluctuation evaluation model is used to calculate the cross-shard consensus delay fluctuation index corresponding to the degree of difference in the policy convergence status between multiple shard nodes;

[0025] Compare the cross-shard consensus delay fluctuation index with the pre-set shard reorganization trigger threshold, and determine whether to trigger dynamic shard reorganization based on the comparison result.

[0026] In a preferred embodiment, S5 is specifically:

[0027] When the cross-shard consensus delay fluctuation index exceeds the preset shard reorganization trigger threshold, the shard structure will be optimized again based on the degree of convergence difference of the strategies of each shard node in the current shard structure;

[0028] Based on the shard structure optimization results, a candidate structure for dynamic shard reorganization is constructed;

[0029] The deep reinforcement learning method is used to evaluate the balance effect of the strategy optimization process of the dynamic sharding reorganization candidate structure, and the dynamic sharding consensus structure obtained after evaluation is output.

[0030] In a preferred embodiment, a deep reinforcement learning method is used to evaluate the balance effect of the strategy optimization process of the dynamic sharding reorganization candidate structure, and the dynamic sharding consensus structure obtained after evaluation is output, specifically:

[0031] For each dynamic sharding reorganization candidate structure, deploy the corresponding sharding node set and strategy distribution state of each candidate structure in a preset simulation environment;

[0032] Run the deep reinforcement learning model in a preset simulation environment to evaluate the node strategy synchronization rate, strategy convergence stability, and cross-shard consensus latency trends of each candidate structure within a given training cycle, and build a composite reward function based on the strategy balance goal.

[0033] The candidate structures are sorted according to the cumulative return value of the composite reward function, and the dynamic sharding reorganization candidate structure with the highest cumulative return value is selected as the output dynamic sharding consensus structure.

[0034] The technical effects and advantages of the present invention's dynamic sharding blockchain consensus method based on deep reinforcement are as follows:

[0035] By extracting the temporal correlation characteristics between node behavior and reward feedback from historical blockchain sharding structure data, the source of strategy differences can be effectively identified, thereby achieving accurate quantification of strategy heterogeneity between shards and improving the controllability of the strategy evolution process; through heterogeneity reconstruction analysis based on strategy heterogeneity levels, potential response bottleneck sections caused by unbalanced strategy convergence can be discovered in a timely manner, and the ability to predict the precursors of consensus performance degradation can be enhanced; combined with the results of strategy imbalance marking, a deep reinforcement learning-driven strategy optimization and reward function adjustment mechanism can be implemented, which can adaptively guide each node to a stable strategy and narrow the behavioral differences between shards from the source; a delay fluctuation evaluation model is constructed through the quantification results of convergence state differences, which realizes dynamic monitoring and precise measurement of cross-shard response coordination; dynamic sharding structure reorganization operations triggered by evaluation indicators can optimize sharding divisions on demand, improve the synchronization of full-network strategy and consensus efficiency, and significantly enhance the stability, efficiency and adaptability of the consensus process in a heterogeneous blockchain environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a schematic diagram of a dynamic sharding blockchain consensus method based on deep enhancement in the present invention. DETAILED DESCRIPTION

[0037] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0038] Example 1

[0039] Figure 1 The present invention provides a dynamic sharding blockchain consensus method based on deep reinforcement, which includes the following steps:

[0040] S1: Using historical blockchain sharding structure data, identify the temporal correlation characteristics between node behavior and consensus rewards in each historical shard. Based on these temporal correlation characteristics, determine the strategic heterogeneity level of multiple shards.

[0041] S2: Perform heterogeneity reconstruction analysis on historical shard structure data based on the policy heterogeneity level of each shard to identify policy convergence imbalance segments that cause consensus latency differences between shards and generate a set of policy imbalance-marked shards.

[0042] S3: Based on the policy imbalance-marked shard set, deep reinforcement learning-driven policy iteration and reward function adaptive adjustment are performed on each shard node, and the degree of difference in the policy convergence state of each shard node is output;

[0043] S4: Calculate the cross-shard consensus delay fluctuation index based on the degree of convergence status difference of each shard node strategy, and use the delay fluctuation index to determine whether to trigger dynamic shard reorganization;

[0044] S5: When the latency fluctuation index exceeds the preset threshold, a dynamic reorganization of the shard structure guided by deep reinforcement learning strategy is performed to balance the strategy optimization process between shards and output a stable dynamic shard consensus structure.

[0045] S1: Identify the temporal correlation characteristics between node behavior and consensus rewards in each historical shard through historical blockchain shard structure data. Based on these temporal correlation characteristics, determine the strategic heterogeneity level of multiple shards, including:

[0046] Collect historical blockchain sharding structure data, and extract the historical consensus action sequence and historical consensus incentive acquisition sequence of nodes in the shard based on the historical blockchain sharding structure data;

[0047] Specifically, first, a blockchain network structure sample set containing multiple historical blockchain operation cycles is constructed. The sample set contains historical blockchain sharding structure data, including the role configuration of nodes in the sharding structure within each time period, the network connection topology between nodes, the consensus initiating node identification, the consensus responding node identification, the consensus completion time, the block broadcast time and other information.

[0048] Based on the historical blockchain sharding data above, we extracted consensus behavior logs for each node in each shard over multiple consecutive time periods and consolidated them into a node consensus behavior history set. This record set contains the sequence of consensus actions executed by each node under different sharding structures. The consensus action sequence consists of multiple consensus events arranged in time sequence. Each consensus event includes the consensus trigger time, the consensus type used (e.g., voting, verification, response), the consensus role type (e.g., initiator, validator), and the corresponding execution result (e.g., whether consensus was reached).

[0049] Synchronously extract the reward information received by nodes after each consensus event is completed to form a historical consensus incentive earning sequence. Reward information is derived from the system incentive details recorded in the ledger at the time of consensus completion, including base rewards, intra-shard delay penalty coefficients, inter-shard response weights, etc., and can be fully restored through the reward distribution records of confirmed transactions in the block structure.

[0050] Establish a temporal association model based on the node's historical consensus action sequence and historical consensus incentive acquisition sequence;

[0051] Specifically, to analyze the causal relationship between node consensus behavior and incentives, a temporal correlation model based on time series behavior analysis was established. The model structure does not rely on a specific model name. Instead, it uses a sliding time window mechanism to align each node's consensus action sequence and incentive acquisition sequence within the same time period. Classification and statistics are then performed by consensus type and node role.

[0052] The constructed temporal association model uses the window length as the basic unit, and sequentially pairs the behavioral characteristics such as the number of consensus actions, consensus delay time, response ratio, failure rate, etc. in each window segment with the incentive value after the end of the window segment to form a time-action-reward mapping set.

[0053] To avoid local contingencies, a sliding overlap mechanism is introduced, sliding each window forward at fixed intervals to generate multiple overlapping sub-window structures, enhancing sample stability. Under the model, a behavioral trend linear fitting method is used to identify the monotonicity and sensitivity between incentive values and behavioral characteristics, thereby quantifying the response deviation and strategy consistency of each node under various behavior-incentive patterns.

[0054] The time series correlation model is used to quantify the strategic similarity of node behaviors in each historical shard, and the quantitative index of strategic heterogeneity of nodes in each historical shard is calculated. The strategic heterogeneity level of multiple shards is determined based on the quantitative index of strategic heterogeneity.

[0055] Specifically, based on the behavior-incentive correspondence pattern extracted from the temporal association model, a strategy feature vector is constructed for each node. The strategy feature vector consists of the following parts:

[0056] The average number of consensus initiations within a unit time window;

[0057] The average response confirmation ratio within the unit time window;

[0058] The average consensus success rate within the unit time window;

[0059] The average excitation value gain rate within the unit time window;

[0060] Strategy volatility between windows.

[0061] After normalization, these eigenvalues are used to construct a strategy feature vector space. Similarity analysis is then performed on all node strategy feature vectors within the same historical shard. This similarity analysis is performed using a vector distance function, where the square root of the sum of the squared differences between the strategy vectors of each pair of nodes is used as the similarity metric.

[0062] For each historical shard, the average strategy heterogeneity index for the shard is calculated by summing the similarity distances between all nodes in the shard. A larger strategy heterogeneity index indicates more differentiated strategy performance between nodes in the shard.

[0063] Based on the historical strategy heterogeneity index values of all shards, thresholds are set according to statistical distribution to divide the strategy heterogeneity into multiple levels. For example, five levels can be defined, each corresponding to a shard optimization difficulty score, which is used to identify strategy imbalance segments and provide input for dynamically reconstructing the shard structure.

[0064] S2: Perform heterogeneity reconstruction analysis on historical shard structure data based on the policy heterogeneity level of each shard to identify policy convergence imbalance segments that cause consensus latency differences between shards. Generate a set of policy imbalance-marked shards, including:

[0065] Obtain the policy heterogeneity levels of multiple shards, and perform heterogeneous reconstruction of the node policy optimization process in the historical blockchain shard structure data based on the policy heterogeneity levels, and identify the temporal distribution characteristics of the policy updates of each shard node during the policy optimization process;

[0066] Specifically, after calculating the policy heterogeneity levels for multiple shards, a mapping table is first established. This table records the policy heterogeneity level values corresponding to each shard in each historical time period and binds them to the shard structure evolution sequence. Higher policy heterogeneity levels indicate greater policy divergence among nodes within the shard, while lower values indicate more consistent policies.

[0067] Based on the policy heterogeneity level, the historical shard structure data is heterogeneously reconstructed. The reconstruction method includes the following operations:

[0068] Divide the entire historical blockchain network operation cycle into continuous time periods and create a snapshot of the sharding structure within each time period;

[0069] Arrange the policy heterogeneity levels of all shards in each snapshot in time to form a shard-level time matrix;

[0070] According to the shard-level time matrix, we mark level mutation events, i.e., shards whose level values change dramatically between consecutive time periods;

[0071] We regard hierarchical mutation events as key nodes of heterogeneity evolution, and the corresponding snapshot structures as anchor points for heterogeneity reconstruction.

[0072] The time window is expanded forward and backward with each anchor point as the center to extract the policy level fluctuation segment in the neighborhood of the anchor point.

[0073] Through the above steps, a heterogeneous reconstructed fragment structure sequence reflecting the dynamic process of strategy evolution is reconstructed, providing a temporal structure basis for identifying strategy imbalance segments.

[0074] Based on heterogeneity, we reconstruct the shard structure sequence, count the policy update records of all nodes in each historical shard, and establish the following structured representation:

[0075] For each node, extract the policy update actions within multiple consecutive time windows, including the policy update time point, update amplitude, and directionality of the policy update;

[0076] Map all policy update actions into time series;

[0077] Parallel normalization of the policy update time series of all nodes in the same shard;

[0078] Based on the normalized results, the mean change rate, variance change rate, and maximum jump rate of the policy update distribution within the shard are calculated;

[0079] The above indicators are combined into the time series distribution feature vector of the sharding strategy update for convergence evaluation.

[0080] In order to enhance the reliability of time series distribution recognition, a sliding time window and redundant segment extension mechanism are introduced to ensure that short-term drastic policy changes are not lost at the boundary time points.

[0081] Based on the temporal distribution characteristics of the policy updates of each shard node, the policy imbalance segments with different policy convergence rates during the node policy optimization process are located, and the policy imbalance segments are marked to generate a policy imbalance marked shard set;

[0082] Specifically, by analyzing the continuity of the time series distribution feature vectors of each sharding strategy update, a time difference recognition function is used to determine whether the node strategy convergence rate is consistent. The judgment method includes the following:

[0083] For each shard, within each sliding time window, calculate the standard deviation of the policy update amplitude between nodes;

[0084] If the standard deviation in multiple consecutive windows is continuously greater than the preset critical value, the shard is determined to be in a policy imbalance state;

[0085] The shard where the time window in the policy imbalance state is located is defined as the policy imbalance segment;

[0086] All policy imbalance segments are aggregated and marked in chronological order to form a policy imbalance marked shard set.

[0087] The label set will serve as the input for subsequent deep reinforcement learning control to guide subsequent strategy synchronization optimization and structural reorganization.

[0088] S3: Based on the policy imbalance-marked shard set, deep reinforcement learning-driven policy iteration and reward function adaptive adjustment are performed on each shard node, and the degree of difference in the policy convergence state of each shard node is output, including:

[0089] Based on the set of policy imbalance labeled shards, a deep reinforcement learning training environment is constructed to determine the initial structure of the reinforcement learning reward function for policy optimization.

[0090] Specifically, based on the set of strategy-imbalanced shards, we extract the behavioral data and reward history of all nodes participating in consensus in each imbalanced shard. This behavioral data includes the type of action taken by each node in consecutive consensus cycles, the time of action execution, the node's role, the consensus triggering event, and the response feedback status. The reward history is the node incentive information recorded in the blockchain ledger after the corresponding action is completed.

[0091] When building a reinforcement learning training environment, set up the following structure:

[0092] State space: It is composed of the action history of nodes in the shard, policy parameters, current shard topology, inter-node communication delay and other dimensions;

[0093] Action space: includes the direction of node policy parameter adjustment, the initiative setting of participating in consensus, and the adjustment range of response strategy;

[0094] The initial reward structure includes a basic consensus completion reward, a cross-shard response penalty, a strategy fluctuation penalty, and a convergence speed reward. The basic consensus completion reward is used to incentivize successful consensus participation, the cross-shard response penalty inhibits frequent cross-shard communication, the strategy fluctuation penalty is used to constrain frequently changing unstable strategy behaviors, and the convergence speed reward provides additional incentives based on the convergence trend of node strategy changes.

[0095] The training environment is initialized using a continuous-time-based multi-period simulation system, allowing nodes to freely evolve strategies within the simulation time to simulate the process of strategy convergence.

[0096] Perform node policy iterative optimization in the deep reinforcement learning training environment, and obtain the policy update data of each shard node at different times through policy iterative optimization;

[0097] Specifically, the reinforcement learning training process is based on a reinforcement learning policy optimization algorithm. It employs a two-stage mechanism that alternates between policy evaluation and policy improvement. During initialization, each node is assigned a basic policy weight configuration, and the behavior state and reward value after each policy adjustment are recorded.

[0098] The training process sets multiple training cycles, each cycle includes the following operations:

[0099] The node runs under the current strategy and completes several consensus actions;

[0100] Record the action-state-reward triplet data of consensus behavior;

[0101] Use the accumulated rewards to calculate the policy update direction and adjust the policy parameters according to the preset learning rate;

[0102] After each training cycle, record the node strategy change range, convergence speed index, and reward value change range;

[0103] After multiple training sessions, the strategy parameters of each node at each time point are output to form a strategy update dataset.

[0104] The policy update dataset records the continuous evolution trajectory of policy parameters in the time dimension, providing a measurable feature basis for quantifying the convergence state differences.

[0105] Based on the initial structure of the reinforcement learning reward function, the reward function is adaptively adjusted for the policy update data of each shard node. Based on the policy change trend and fluctuation range shown in the policy update data of each shard node, the degree of convergence state difference of each shard node strategy is quantitatively calculated;

[0106] Specifically, during the reinforcement learning training cycle, in order to avoid policy drift or extreme convergence, an adaptive adjustment mechanism of the reward function based on the node policy performance is implemented. The specific method is as follows:

[0107] For nodes with significantly higher strategy update frequency but unchanged returns, frequent updates are suppressed by increasing the volatility penalty coefficient;

[0108] For nodes whose strategy updates tend to be stable and convergence effect is good, positive feedback is given by increasing the convergence speed reward coefficient;

[0109] The adjustment of reward function parameters is based on the node's strategy change trend slope, incentive mean growth rate, and strategy gradient change amplitude in the previous cycle;

[0110] After completing several dynamic adjustments to the reward function, calculate the degree of difference in the strategy convergence state based on the strategy change trajectory of each node in all training cycles. The specific calculation method is as follows:

[0111] Construct a strategy change curve for each node and extract the stability characteristics of the change curve, including the maximum jump amplitude, standard deviation, and average slope;

[0112] For all nodes in the same shard, the stability feature vectors are compared pairwise to find the stability difference between each node and other nodes;

[0113] All difference values within each shard are averaged to obtain the overall strategy convergence state difference index of the shard.

[0114] S4: Calculate the cross-shard consensus delay fluctuation index based on the convergence status differences of each shard node strategy. Use the delay fluctuation index to determine whether to trigger dynamic shard reorganization, including:

[0115] Based on the degree of convergence difference of each shard node strategy, a cross-shard consensus delay fluctuation assessment model is established;

[0116] Specifically, the cross-shard consensus delay fluctuation evaluation model is constructed as follows:

[0117] First, we collect actual cross-shard consensus execution logs distributed over multiple historical time periods. The logs include the initiation time, response time, target shard number, and response shard number of each consensus event.

[0118] Based on the log data, the average response delay and the maximum delay difference between any two shards are calculated to form a cross-shard delay feature matrix.

[0119] Using the sharding strategy convergence state difference index as the input variable and the delay feature matrix as the output variable, a function mapping relationship is established;

[0120] To improve modeling accuracy, a sliding time window is used in the function fitting process to extract the covariation trajectory of strategy changes and delayed fluctuations to distinguish long-term trends from short-term disturbances.

[0121] The cross-shard consensus delay fluctuation assessment model adjusts fitting parameters with the goal of minimizing statistical residuals, and ultimately outputs a set of mapping rules for predicting the future cross-shard consensus delay fluctuation based on the current convergence state difference value.

[0122] The cross-shard consensus delay fluctuation assessment model does not rely on preset threshold judgment rules, but forms a decision-making causal structure based on historical co-evolution laws.

[0123] The cross-shard consensus delay fluctuation evaluation model is used to calculate the cross-shard consensus delay fluctuation index corresponding to the degree of difference in the policy convergence status between multiple shard nodes;

[0124] Specifically, after the cross-shard consensus delay fluctuation evaluation model is established, the node strategy convergence state difference index within each shard in the current blockchain is calculated in real time, substituted into the evaluation model as input, and delay fluctuation prediction is performed.

[0125] The following methods are adopted in the forecasting process:

[0126] Unify the clock windows of all shards so that all strategy indicators and forecast values have synchronized timestamps;

[0127] For each target shard, calculate the predicted consensus response delay between it and each of the other shards to form the delay prediction matrix at the current moment;

[0128] In the delay prediction matrix, the difference between the maximum delay value and the minimum delay value is extracted and defined as the cross-shard consensus delay fluctuation index;

[0129] The delayed fluctuation time sliding average function is introduced to smooth the fluctuation indicators in multiple consecutive time periods to avoid amplification of single-moment errors.

[0130] The latency fluctuation index, as a reflection of the strategy consistency and system response stability during the operation of the blockchain, will be used to trigger structural reorganization decisions.

[0131] Compare the cross-shard consensus delay fluctuation index with the pre-set shard reorganization trigger threshold, and determine whether to trigger dynamic shard reorganization based on the comparison result;

[0132] Specifically, after obtaining the latency fluctuation index, set the decision logic for whether the dynamic sharding structure needs to be reorganized:

[0133] Based on the design objectives, a set of shard reassembly trigger thresholds are pre-set, including the minimum allowable delay fluctuation value, the shard quantity ratio threshold, and the number of consecutive exceeding limits.

[0134] Compare the currently calculated latency fluctuation indicators with the set thresholds item by item:

[0135] If the current delay fluctuation index exceeds the minimum allowable delay fluctuation value;

[0136] And the number of shard pairs in abnormal state reaches the minimum reconstruction ratio of the total number of shards in the system;

[0137] and the delayed volatility has not fallen back to the target range for more than three consecutive time periods;

[0138] It is determined that the current state is in a policy-response stability imbalance state, and the dynamic sharding reorganization trigger condition is met;

[0139] Mark the current state as pending for structural adjustment.

[0140] The judgment logic is a combination of conditional constraints, ensuring that frequent triggering due to occasional indicator fluctuations is avoided, thereby enhancing stability and adaptability.

[0141] S5: When the latency fluctuation index exceeds the preset threshold, a dynamic reorganization of the shard structure guided by deep reinforcement learning strategy is performed to balance the strategy optimization process between shards and output a stable dynamic shard consensus structure, including:

[0142] When the cross-shard consensus delay fluctuation index exceeds the preset shard reorganization trigger threshold, the shard structure will be optimized again based on the degree of convergence difference of the strategies of each shard node in the current shard structure;

[0143] Specifically, when the cross-shard consensus delay fluctuation index exceeds the preset threshold, the system immediately calls the shard structure optimization module. The optimization operation depends on the following input information:

[0144] The current shard structure topology of the entire network, including the unique identifiers and communication paths of the nodes contained in each shard;

[0145] The strategy convergence status difference index of each node is a quantitative score of the strategy convergence degree;

[0146] The current consensus latency characteristics between shards are used as a supplementary constraint input.

[0147] The structural optimization process includes the following operation procedures:

[0148] Collect statistics on the policy convergence difference indicators within each shard in the current entire network and extract the average value and standard deviation;

[0149] Measure the gradient of policy heterogeneity between shards to determine whether there is structural policy polarization;

[0150] Based on the principle of minimizing policy heterogeneity, a shard reconstruction allocator is used to migrate nodes with large policy differences to other shards with similar policy convergence levels.

[0151] During each migration, the target shard's policy convergence balance is re-evaluated to ensure that each structural adjustment can improve global consistency.

[0152] The node reallocation operation is iterated until the internal policy convergence difference value of all shards is lower than the maximum acceptable fluctuation threshold or the preset maximum number of iterations is reached.

[0153] The result of structural optimization is a set of candidate structural configurations, each representing several possible new sharding schemes with different policy consistency characteristics and expected consensus efficiency.

[0154] Based on the shard structure optimization results, a candidate structure for dynamic shard reorganization is constructed;

[0155] Specifically, according to several reconstruction schemes, a set of dynamic sharding and reorganization candidate structures is constructed. Each candidate structure contains the following contents:

[0156] Sharding scheme: Each structure corresponds to a set of shards and their corresponding node sets;

[0157] Node historical behavior context: including the past strategy evolution path and consensus response records of nodes assigned to each shard;

[0158] Current strategy parameter snapshot: records the current strategy vector of each node at the time of structure construction;

[0159] Shard communication topology prediction graph: predicts the possible intra-shard and inter-shard communication paths after structural reorganization, which serves as input for latency evaluation.

[0160] To improve the coverage rate during the evaluation phase, the candidate structure set should meet the distribution diversity requirements in three dimensions: number of shards, node distribution balance, and strategy similarity. All candidate structures must be divided without overlap while maintaining full node coverage of the system.

[0161] Use deep reinforcement learning methods to evaluate the balance effect of the strategy optimization process of dynamic sharding reorganization candidate structures, and output the dynamic sharding consensus structure obtained after evaluation;

[0162] Specifically, for each structure in the candidate structure set, a reinforcement learning strategy optimization simulation is performed in a preset evaluation simulation environment:

[0163] Initialize the candidate structure and deploy it in the evaluation environment, setting the initial strategy to the current strategy snapshot of its node;

[0164] Execute the strategy optimization simulation process within a fixed simulation cycle, including each node performing strategy adjustments based on the current state, and recording the strategy convergence path and reward value changes;

[0165] Quantify the policy synchronization rate of each node, the balance of policy update amplitude, and the change trend of delay distribution between shards during the simulation cycle;

[0166] Construct an evaluation function for reinforcement learning scoring, which includes the following three dimensions:

[0167] Intra-shard strategy volatility balance score;

[0168] Inter-shard response latency stability score;

[0169] Node strategy synchronization index score;

[0170] For each candidate structure, the comprehensive scoring index value is calculated, and the one with the highest score is selected as the target structure, that is, the output dynamic sharding consensus structure.

[0171] Use deep reinforcement learning methods to evaluate the balance effect of the strategy optimization process of dynamic sharding reorganization candidate structures, and output the dynamic sharding consensus structure obtained after evaluation, including:

[0172] For each dynamic sharding reorganization candidate structure, deploy the corresponding sharding node set and strategy distribution state of each candidate structure in a preset simulation environment;

[0173] Specifically, for each dynamic sharding and reorganization candidate structure, an independent simulation running instance is created. The simulation environment has the following basic components:

[0174] Shard topology mapping module: used to load the node partitioning scheme, logical connection relationship between nodes, and shard communication path of each shard in each candidate structure;

[0175] Node behavior initialization module: used to read the historical strategy evolution trajectory of each node in the candidate structure and generate a snapshot of the strategy parameters at the current moment;

[0176] State data cache module: used to record each node's strategy changes, sharding delays, response logs, synchronization behavior information, etc. during simulation execution.

[0177] During deployment, the node set corresponding to the candidate architecture is first loaded into the simulator, ensuring that each node's shard and policy state remain consistent with the previous architecture optimization phase. The policy vector for each node is then initialized based on historical node behavior data, ensuring that the policy synchronization rate can be traced back to the actual evolutionary basis. All candidate architectures are executed in independent threads under the same operating environment specifications to ensure fair comparison conditions.

[0178] Run the deep reinforcement learning model in a preset simulation environment to evaluate the node strategy synchronization rate, strategy convergence stability, and cross-shard consensus latency trends of each candidate structure within a given training cycle, and build a composite reward function based on the strategy balance goal.

[0179] Specifically, in the simulation environment, a set of deep reinforcement learning model instances are deployed for each candidate structure to simulate the policy optimization interaction process between nodes. The stage evaluation is carried out using the following technical paths:

[0180] Input feature selection: Select each node's policy update amplitude, policy update frequency, intra-shard consensus response time difference, and inter-shard delay fluctuation value as state space input;

[0181] Action space design: including node strategy parameter fine-tuning direction, consensus participation frequency adjustment factor, and shard response waiting time configuration;

[0182] Reward function construction: Design a three-dimensional composite reward function, which includes the following three sub-items:

[0183] Strategy synchronization reward: The reward value increases when all nodes’ strategy changes tend to be consistent;

[0184] Strategy Stability Bonus: Additional points are awarded when the node strategy maintains a stable change rate over multiple consecutive periods.

[0185] Cross-shard consensus delay suppression item: Provides enhanced feedback when the fluctuation range of cross-shard response time decreases;

[0186] The above three items are combined according to the set weights to form the final composite reward function, which serves as the optimization target of the training model.

[0187] Reinforcement learning training uses a multi-cycle operation mechanism, with each cycle recording the trend of each node's strategy change, reward value response trend, and synchronization level change. After each round of training, the system statistically summarizes the strategy balance index to form a strategy evolution feature spectrum for each candidate structure.

[0188] Sort the candidate structures according to the cumulative return value of the composite reward function, and select the dynamic sharding reorganization candidate structure with the highest cumulative return value as the output dynamic sharding consensus structure;

[0189] Specifically, after the reinforcement learning simulation training of all candidate structures is completed, the system ranks and evaluates each candidate structure based on the cumulative return value of the composite reward function in all training cycles. The operation is as follows:

[0190] For each candidate structure, extract the reward value of each training cycle from its training log;

[0191] Perform sliding integration on the reward value in time series to calculate the cumulative return value of the candidate structure;

[0192] Arrange the cumulative return values of all candidate structures in descending order, and take the one with the highest cumulative return value as the final output structure of this round of sharding reorganization.

[0193] The final output dynamic sharding consensus structure includes the following:

[0194] Shard division details: a list of node identifiers contained in each shard;

[0195] Strategy balance description: Description of the strategy synchronization rate and convergence trend of the target structure during the evaluation cycle;

[0196] Performance estimation report: includes indicators such as predicted cross-shard consensus latency stability and initial strategy performance score after reorganization.

[0197] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.

[0198] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0199] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0200] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0201] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0202] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0203] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0204] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0205] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0206] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A dynamic sharding blockchain consensus method based on deep reinforcement, characterized by: The steps include: S1: Using historical blockchain sharding structure data, identify the temporal correlation characteristics between node behavior and consensus rewards in each historical shard. Based on these temporal correlation characteristics, determine the strategic heterogeneity level of multiple shards. S2: Perform heterogeneity reconstruction analysis on historical shard structure data based on the policy heterogeneity level of each shard to identify policy convergence imbalance segments that cause consensus latency differences between shards and generate a set of policy imbalance-marked shards. S3: Based on the policy imbalance-marked shard set, deep reinforcement learning-driven policy iteration and reward function adaptive adjustment are performed on each shard node, and the degree of difference in the policy convergence state of each shard node is output; S4: Calculate the cross-shard consensus delay fluctuation index based on the degree of convergence status difference of each shard node strategy, and use the delay fluctuation index to determine whether to trigger dynamic shard reorganization; S5: When the latency fluctuation index exceeds the preset threshold, a dynamic reorganization of the shard structure guided by deep reinforcement learning strategy is performed to balance the strategy optimization process between shards and output a stable dynamic shard consensus structure.

2. A dynamic sharding blockchain consensus method based on deep reinforcement according to claim 1, characterized in that: S1, specifically: Collect historical blockchain sharding structure data, and extract the historical consensus action sequence and historical consensus incentive acquisition sequence of nodes in the shard based on the historical blockchain sharding structure data; Establish a temporal association model based on the node's historical consensus action sequence and historical consensus incentive acquisition sequence; The temporal association model is used to quantify the strategic similarity of node behaviors in each historical shard, and the quantitative index of strategic heterogeneity of nodes in each historical shard is calculated. The strategic heterogeneity level of multiple shards is determined based on the quantitative index of strategic heterogeneity.

3. A dynamic sharding blockchain consensus method based on deep reinforcement according to claim 2, characterized in that: S2, specifically: Obtain the policy heterogeneity levels of multiple shards, and perform heterogeneous reconstruction of the node policy optimization process in the historical blockchain shard structure data based on the policy heterogeneity levels, and identify the temporal distribution characteristics of the policy updates of each shard node during the policy optimization process; According to the temporal distribution characteristics of the strategy update of each shard node, the strategy imbalance segments with different strategy convergence rates in the node strategy optimization process are located, and the strategy imbalance segments are marked to generate a strategy imbalance marked shard set.

4. A dynamic sharding blockchain consensus method based on deep reinforcement according to claim 3, characterized in that: S3, specifically: Based on the set of policy imbalance labeled shards, a deep reinforcement learning training environment is constructed to determine the initial structure of the reinforcement learning reward function for policy optimization. Perform node policy iterative optimization in the deep reinforcement learning training environment, and obtain the policy update data of each shard node at different times through policy iterative optimization; Based on the initial structure of the reinforcement learning reward function, the reward function is adaptively adjusted for the policy update data of each shard node. Based on the policy change trend and fluctuation amplitude of each shard node in the policy update data, the degree of convergence state difference of each shard node strategy is quantitatively calculated.

5. A dynamic sharding blockchain consensus method based on deep reinforcement according to claim 4, characterized in that: S4, specifically: Based on the degree of convergence difference of each shard node strategy, a cross-shard consensus delay fluctuation assessment model is established; The cross-shard consensus delay fluctuation evaluation model is used to calculate the cross-shard consensus delay fluctuation index corresponding to the degree of difference in the policy convergence status between multiple shard nodes; Compare the cross-shard consensus delay fluctuation index with the pre-set shard reorganization trigger threshold, and determine whether to trigger dynamic shard reorganization based on the comparison result.

6. A dynamic sharding blockchain consensus method based on deep reinforcement according to claim 5, characterized in that: S5, specifically: When the cross-shard consensus delay fluctuation index exceeds the preset shard reorganization trigger threshold, the shard structure will be optimized again based on the degree of convergence difference of the strategies of each shard node in the current shard structure; Based on the shard structure optimization results, a candidate structure for dynamic shard reorganization is constructed; The deep reinforcement learning method is used to evaluate the balance effect of the strategy optimization process of the dynamic sharding reorganization candidate structure, and the dynamic sharding consensus structure obtained after evaluation is output.

7. A dynamic sharding blockchain consensus method based on deep reinforcement according to claim 6, characterized in that: The deep reinforcement learning method is used to evaluate the balance effect of the strategy optimization process of the dynamic sharding reorganization candidate structure, and the dynamic sharding consensus structure obtained after evaluation is output, specifically: For each dynamic sharding reorganization candidate structure, deploy the corresponding sharding node set and strategy distribution state of each candidate structure in a preset simulation environment; Run the deep reinforcement learning model in a preset simulation environment to evaluate the node strategy synchronization rate, strategy convergence stability, and cross-shard consensus latency trends of each candidate structure within a given training cycle, and build a composite reward function based on the strategy balance goal. The candidate structures are sorted according to the cumulative return value of the composite reward function, and the dynamic sharding reorganization candidate structure with the highest cumulative return value is selected as the output dynamic sharding consensus structure.