Dynamic block chain fragmentation method and system based on composite clustering algorithm

Through the dynamic blockchain sharding method based on composite clustering algorithm, the problem of unbalanced shard load and excessive cross-shash transaction ratio in blockchain sharding technology is solved, and efficient utilization of system resources and performance improvement is achieved.

CN120069870AActive Publication Date: 2025-05-30JINAN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510086223.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

In actual application, blockchain sharding technology faces problems such as unbalanced shard load and excessive cross-shash transaction ratio, resulting in wasted system resources, performance bottlenecks, response time and throughput.

Method used

The dynamic blockchain sharding method based on the composite clustering algorithm is adopted. By determining the transaction frequency matrix and shard allocation matrix, the target optimization transaction model is constructed, and the shard division is used to divide the shard by using the condensed hierarchical clustering and the DBSCAN algorithm. High-frequency correlation accounts are preferred to aggregate high-frequency correlation accounts, and cross-shash transactions are reduced, and load balancing is achieved through the classification processing of noise accounts.

Benefits of technology

Significantly reduce the occurrence of cross-shash transactions, reduce communication costs, improve system throughput, realize dynamic balancing of shard load, and make full use of system resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069870A_ABST
    Figure CN120069870A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic block chain fragmentation method and system based on a composite clustering algorithm, and the method comprises the steps: determining a transaction frequency matrix and a fragmentation distribution matrix based on a block chain fragmentation system, and constructing a target optimization transaction model of the block chain fragmentation system; based on the target optimization transaction model of the block chain fragmentation system, performing preliminary global division processing through an agglomerated hierarchical clustering algorithm; performing division processing on the preliminary block chain fragmentation result through a DBSCAN clustering algorithm to obtain a block chain fragmentation result after secondary division; and performing noise account division processing based on the block chain fragmentation result after secondary division to obtain a final block chain fragmentation result. According to the method, the high-frequency associated accounts can be preferentially aggregated, the low-frequency accounts can be classified, the cross-fragment transaction proportion is minimized, and load balancing is realized. The dynamic block chain fragmentation method and system based on the composite clustering algorithm can be widely applied to the technical field of block chain transaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of blockchain transactions, and in particular, to a dynamic blockchain sharding method and system based on a composite clustering algorithm. Background Art

[0002] As one of the mainstream means to improve the scalability of blockchains, sharding technology has received much attention in recent years. Its core idea is to effectively improve the throughput and performance of blockchains by dividing the computing and storage burdens of the network. However, although sharding technology shows great potential at the theoretical level, it still faces many key challenges in practical applications, especially in two aspects: shard load balancing and an overly large proportion of cross-shard transactions. Unbalanced shard loads may lead to resource shortages in some shards while other shards have idle resources, thus affecting the overall efficiency of the system. Load imbalance not only wastes system resources but also leads to performance bottlenecks. Especially when the network load is high, some shards may become congested due to handling too many requests, which in turn affects the response time and throughput of the entire blockchain system. At the same time, an overly high proportion of cross-shard transactions will incur additional communication and computing overheads. First, the communication overhead increases. Cross-shard transactions require frequent information interaction between shards, increasing the communication cost. Second, the processing flow becomes more complex. The state update of cross-shard transactions involves ledger modifications and consistency guarantees of multiple shards, and usually requires complex protocols (such as two-phase commit or lock-based protocols) to ensure the atomicity of transactions. This not only increases the transaction processing time but may also lead to performance degradation due to operation conflicts, significantly affecting network performance and latency. Summary of the Invention

[0003] To solve the above technical problems, the objective of the present invention is to provide a dynamic blockchain sharding method and system based on a composite clustering algorithm, which can preferentially aggregate high-frequency associated accounts and classify low-frequency accounts, minimizing the proportion of cross-shard transactions and achieving load balance.

[0004] The first technical solution adopted by the present invention is: A dynamic blockchain sharding method based on a composite clustering algorithm, comprising the following steps:

[0005] Based on the blockchain sharding system, determine the transaction frequency matrix and the shard allocation matrix, and construct the target optimization transaction model of the blockchain sharding system;

[0006] Based on the target optimization transaction model of the blockchain sharding system, perform preliminary global partitioning processing through the agglomerative hierarchical clustering algorithm to obtain a preliminary blockchain sharding result;

[0007] Perform partitioning processing on the preliminary blockchain sharding result through the DBSCAN clustering algorithm to obtain the blockchain sharding result after secondary partitioning;

[0008] Based on the blockchain sharding result after secondary partitioning, perform noise account partitioning processing to obtain the final blockchain sharding result.

[0009] Furthermore, the step of determining the transaction frequency matrix and the shard allocation matrix based on the blockchain sharding system and constructing the target optimization transaction model of the blockchain sharding system specifically includes:

[0010] Based on the blockchain sharding system, obtain the historical transaction data between all accounts and construct a transaction frequency matrix;

[0011] Based on the blockchain sharding system, define the total cross-shard transaction frequency and the system load balancing objective function;

[0012] Introduce a preset weight coefficient, combine the total cross-shard transaction frequency and the system load balancing objective function and perform matrix representation to obtain a shard allocation matrix;

[0013] Combine the transaction frequency matrix and the shard allocation matrix and use them to represent the calculation process of cross-shard transactions, and construct the target optimization transaction model of the blockchain sharding system.

[0014] Furthermore, the expression of the target optimization transaction model of the blockchain sharding system is specifically as follows:

[0015]

[0016] In the above formula, λ represents the weight coefficient, M represents the transaction frequency matrix, Z represents the shard allocation matrix, Z T represents the transpose matrix of the shard allocation matrix, i represents the i-th account, j represents the j-th account, Z ip represents whether the account i is assigned to shard p, Z jq represents whether the account j is assigned to shard q, and sum(·) represents the total transaction frequency of all cross-shard transactions.

[0017] Furthermore, the step of performing preliminary global partitioning processing on the target optimization transaction model based on the blockchain sharding system through the agglomerative hierarchical clustering algorithm to obtain a preliminary blockchain sharding result specifically includes:

[0018] Based on the transaction frequency matrix in the target optimization transaction model of the blockchain sharding system, define a distance expression based on the transaction frequency to obtain the similarity degree between accounts;

[0019] Regard an account as a cluster and construct an initial cluster set;

[0020] Based on the similarity degree between accounts, obtain the inter-cluster distance of the initial cluster set through the single-link method, and determine the minimum value of the inter-cluster distance and the number of clusters in the initial cluster set;

[0021] Cluster the initial cluster set through the agglomerative hierarchical clustering algorithm. If the minimum value of the inter-cluster distance is greater than the preset inter-cluster distance threshold or the number of clusters is less than the target value, iterate and merge the two clusters corresponding to the minimum value of the inter-cluster distance, and update the minimum value of the inter-cluster distance and the number of clusters;

[0022] Continue until the updated minimum value of the inter-cluster distance is less than the preset inter-cluster distance threshold or the updated number of clusters reaches the target value, and obtain the preliminary blockchain sharding result.

[0023] Furthermore, the expression of the inter-cluster distance is specifically as follows:

[0024]

[0025] In the above formula, C i , C j represent two clusters, a p and a q represent any accounts belonging to C i and C j , d(a p , a q ) represents the similarity degree between account a p and account a q , and D(C i , C j ) represents the distance between cluster C i and cluster C j .

[0026] Furthermore, the step of dividing the preliminary blockchain sharding result through the DBSCAN clustering algorithm to obtain the blockchain sharding result after secondary division specifically includes:

[0027] Construct a subset data matrix based on the transaction frequency between accounts in the clusters in the preliminary blockchain sharding result, and determine the account distances in the subset data matrix;

[0028] Based on the DBSCAN clustering algorithm, define the neighborhood radius and the minimum number of neighborhood points;

[0029] Traverse each account in the preliminary blockchain sharding result through the DBSCAN clustering algorithm, and construct a neighborhood point set for each account according to the account distances and the neighborhood radius in the subset data matrix;

[0030] Determine the number of neighborhood points of the account according to the neighborhood point set of the account and make a judgment;

[0031] If the number of neighboring points of an account is greater than or equal to the minimum number of neighboring points, mark the account as a core account, and add all the accounts in the neighborhood of the core account to the current cluster until no new accounts can be added, obtaining a density cluster;

[0032] If the number of neighboring points of an account is less than the minimum number of neighboring points and it belongs to a density cluster, mark the account as a boundary account;

[0033] If the number of neighboring points of an account is less than the minimum number of neighboring points and it does not belong to a density cluster, mark the account as a noise account;

[0034] Integrate the density cluster, boundary accounts, and noise accounts to obtain the blockchain sharding result after the secondary partition.

[0035] Furthermore, the expression of the account distance in the subset data matrix is specifically as follows:

[0036]

[0037] f(a p ,a q )=T i [p][q]

[0038] In the above formula, d(a p ,a q ) represents the account distance in the subset data matrix, T i [p][q] represents the transaction frequency between account a p and a q , f(·) represents the transaction frequency function, a r 、a s represent the accounts used for normalization reference.

[0039] Furthermore, the step of performing noise account partitioning processing based on the blockchain sharding result after the secondary partition to obtain the final blockchain sharding result specifically includes:

[0040] According to the noise account set, measure and partition the noise accounts to determine the transaction frequency of the noise accounts;

[0041] Define a transaction frequency threshold, mark the noise accounts with a transaction frequency greater than the corresponding transaction frequency threshold as high-frequency noise accounts, and mark the noise accounts with a transaction frequency less than the transaction frequency threshold as low-frequency noise accounts;

[0042] Obtain the transaction relevance scores between the high-frequency noise accounts and all the shards in the blockchain sharding result after the secondary partition, and allocate the high-frequency noise accounts to the shard corresponding to the highest transaction relevance score;

[0043] The low-frequency noise accounts are evenly distributed to all the shards in the blockchain sharding result after the secondary partitioning through a random hashing algorithm, or the low-frequency noise accounts are concentrated in the cold account shard to obtain the final blockchain sharding result.

[0044] Furthermore, the expression of the transaction correlation score is specifically as follows:

[0045]

[0046] In the above formula, R(a i , C q ) represents the transaction correlation score, a i represents the high-frequency noise account, C q represents the account set of the q-th shard, M ij represents the account a i and a j 's transaction frequency within the previous epoch time window.

[0047] The second technical solution adopted by the present invention is: a dynamic blockchain sharding system based on a composite clustering algorithm, including:

[0048] The first module is used to determine the transaction frequency matrix and the shard allocation matrix based on the blockchain sharding system, and construct the target optimization transaction model of the blockchain sharding system;

[0049] The second module is used to perform preliminary global partitioning processing through an agglomerative hierarchical clustering algorithm based on the target optimization transaction model of the blockchain sharding system to obtain a preliminary blockchain sharding result;

[0050] The third module is used to perform partitioning processing on the preliminary blockchain sharding result through the DBSCAN clustering algorithm to obtain the blockchain sharding result after secondary partitioning;

[0051] The fourth module is used to perform noise account partitioning processing based on the blockchain sharding result after secondary partitioning to obtain the final blockchain sharding result.

[0052] The beneficial effects of the method and system of the present invention are as follows: Based on the blockchain sharding system, the present invention determines the transaction frequency matrix and the shard allocation matrix, constructs the target optimization transaction model of the blockchain sharding system, and can minimize the cross-shard transaction ratio and achieve load balancing. Further, based on the target optimization transaction model of the blockchain sharding system, through the agglomerative hierarchical clustering algorithm for preliminary global partitioning processing, a preliminary blockchain sharding result is obtained, and through the DBSCAN clustering algorithm, the preliminary blockchain sharding result is partitioned to obtain the blockchain sharding result after secondary partitioning. Through the composite clustering algorithm, the high-frequency associated accounts are preferentially aggregated to maximize the in-shard transaction density, significantly reduce the occurrence of cross-shard transactions, reduce the cross-shard communication cost, improve the system throughput, and finally, based on the blockchain sharding result after secondary partitioning, the noise account partitioning process is carried out. Through the classification processing of low-frequency accounts, a flexible allocation strategy is designed. The low-frequency accounts are randomly and evenly allocated or concentrated to a dedicated "cold account shard", thereby realizing the dynamic balance of the load between shards, making full use of the computing power resources of the system, and further realizing the dynamic balance of shard load and the optimization of cross-shard transactions. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is the flowchart of the steps of a dynamic blockchain sharding method based on a composite clustering algorithm of the present invention;

[0054] Figure 2 is the structural block diagram of a dynamic blockchain sharding system based on a composite clustering algorithm of the present invention;

[0055] Figure 3 is the schematic framework diagram of the dynamic sharding partition of accounts by the composite clustering algorithm provided by a specific embodiment of the present invention;

[0056] Figure 4 is the schematic diagram of the proportion of repeated transaction account pairs and random account pairs provided by a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following further describes the present invention in detail with reference to the drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0058] First of all, it should be noted that sharded load balancing is a key issue in blockchain sharding technology. The main challenge lies in how to reasonably allocate computing and storage resources in the network to ensure that the load of each shard can be effectively distributed, thus avoiding over - utilization of resources in some shards while other shards are idle. Load imbalance not only wastes system resources but also leads to performance bottlenecks. Especially when the network load is high, some shards may become congested due to handling too many requests, which in turn affects the response time and throughput of the entire blockchain system.

[0059] Secondly, while sharding technology improves the scalability of the blockchain system, it inevitably introduces an important challenge, namely the cross - shard transaction problem. Since sharding divides accounts or data into different shards, when data interaction involving multiple shards occurs, cross - shard transactions will be generated. Compared with transactions within a single shard, the processing complexity of cross - shard transactions is significantly increased. First, the communication overhead increases. Cross - shard transactions require frequent information interaction between shards, increasing the communication cost. Second, the processing process becomes more complex. The state update of cross - shard transactions involves ledger modifications and consistency guarantees of multiple shards, and usually requires complex protocols (such as two - phase commit or lock - based protocols) to ensure the atomicity of transactions. This not only increases the transaction processing time but may also lead to performance degradation due to operation conflicts.

[0060] To address the two major technical problems of too high cross - shard transaction ratio and load imbalance between shards, the embodiments of the present invention found a key characteristic by observing transaction behaviors in real life: accounts with higher activity levels tend to have stronger dependencies. In other words, accounts that have had transactions are more likely to have transactions again in the future. This characteristic provides a very valuable entry point for the present invention. Based on this characteristic, the historical transaction data between accounts can be used to infer the probability of future transactions, thereby optimizing the shard partitioning strategy of accounts.

[0061] Based on this, referring to Figure 1 , the present invention provides a dynamic blockchain sharding method based on a composite clustering algorithm, and the method includes the following steps:

[0062] S100. Based on the blockchain sharding system, determine the transaction frequency matrix and the shard allocation matrix, and construct the target optimization transaction model of the blockchain sharding system;

[0063] S110. Based on the blockchain sharding system, obtain the historical transaction data between all accounts and construct the transaction frequency matrix;

[0064] Specifically, assume that the blockchain sharding system contains N accounts, denoted as A = {a 1 ,a 2 ,…,a N}。The historical transactions between accounts can be represented by a transaction frequency matrix \(M\in\mathbb{R}\) N×N where \(M\) ij represents the transaction frequency between account \(a\) i and \(a\) j within the last epoch time window. For ease of processing, the transaction matrix needs to satisfy the following properties:

[0065] 1) Symmetry:

[0066]

[0067] The transaction frequency matrix is a symmetric matrix, indicating that the transaction frequency is bidirectional.

[0068] Non-negativity:

[0069]

[0070] The transaction frequency between accounts cannot be negative.

[0071] Diagonal is zero:

[0072]

[0073] S120. Based on the blockchain sharding system, define the total cross-shard transaction frequency and the system load balancing objective function;

[0074] Specifically, the research objective of system modeling: Divide the account set \(A\) into \(K\) clusters (shards) \(S = \{S\) 1 , \(S\) 2 , \(\cdots S\) K} through the clustering algorithm, so that the following two objectives are satisfied simultaneously. One is to minimize cross-shard transactions: allocate accounts with high-frequency transactions to the same shard as much as possible; the other is load balancing, ensuring that the number of account transactions in each shard is roughly uniform to balance computing resources.

[0075] Further, model the objective function under these two objectives.

[0076] First is to minimize cross-shard transactions. The performance overhead brought by cross-shard transactions is the main performance bottleneck of the sharding system. Define \(\delta\) ij as the indicator variable indicating whether account \(a\) i and \(a\) j belong to different shards, and its expression is:

[0077]

[0078] Further define the total cross-shard transaction frequency, and its expression is:

[0079]

[0080] The goal of the embodiment of the present invention is to minimize the total cross-shard transaction frequency, and its expression is:

[0081]

[0082] Secondly, there is load balancing. Shard load balancing requires that the number of account transactions in each shard be as even as possible. Define T p as the total number of account transactions in shard p as |T p |. The goal of load balancing can be expressed as minimizing the maximum deviation of the number of shard account transactions, and its expression is:

[0083]

[0084] Therefore, the load balancing goal is:

[0085]

[0086] In the above formula, B min represents minimizing the load balancing goal.

[0087] S130. Introduce a preset weight coefficient, combine the total cross-shard transaction frequency and the system load balancing objective function and perform matrix representation to obtain a shard allocation matrix;

[0088] Specifically, to optimize the cross-shard transaction ratio and load balancing simultaneously, introduce a weight parameter λ ∈ [0, 1], and the comprehensive objective function is:

[0089] F(S) = λ · f c + (1 - λ) · B

[0090] In the above formula, λ represents controlling the trade-off between cross-shard transactions and load balancing.

[0091] S140. Combine the transaction frequency matrix and the shard allocation matrix and use it to represent the calculation process of cross-shard transactions, and construct an objective optimization transaction model for the blockchain sharding system.

[0092] Specifically, for simplicity of calculation, the objective function can be represented in matrix form. Define the shard allocation matrix Z ∈ {0, 1} N×K , where:

[0093]

[0094] The relationship between the transaction frequency matrix M and the shard allocation matrix Z is used to describe the calculation process of cross-shard transactions as:

[0095] f c = sum(M ⊙ (ZZ T = 0))

[0096] Among them, ⊙ represents element-wise multiplication.

[0097] The load balancing objective can be calculated by the row sum or column sum of Z, and its expression is:

[0098]

[0099] Then, by combining the two, the optimization objective can be modeled as:

[0100]

[0101] In the above formula, λ represents the weight coefficient, M represents the transaction frequency matrix, Z represents the shard allocation matrix, and Z T represents the transpose matrix of the shard allocation matrix, i represents the i-th account, j represents the j-th account, and Z ip = 1 means that account i is allocated to shard p, otherwise Z ip = 0, and sum(·) represents the total transaction frequency of all cross-shard transactions.

[0102] S200. For the target optimization transaction model based on the blockchain sharding system, perform preliminary global partitioning processing through the agglomerative hierarchical clustering algorithm to obtain a preliminary blockchain sharding result;

[0103] Specifically, based on the transaction frequency matrix in the target optimization transaction model of the blockchain sharding system, define a distance expression based on the transaction frequency to obtain the similarity degree between accounts; regard an account as a cluster to construct an initial cluster set; based on the similarity degree between accounts, obtain the inter-cluster distance of the initial cluster set through the single-link method, and determine the minimum value of the inter-cluster distance and the number of clusters in the initial cluster set; perform clustering on the initial cluster set through the agglomerative hierarchical clustering algorithm. If the minimum value of the inter-cluster distance is greater than the preset inter-cluster distance threshold or the number of clusters is less than the target value, iteratively merge the two clusters corresponding to the minimum value of the inter-cluster distance, and update the minimum value of the inter-cluster distance and the number of clusters; until the updated minimum value of the inter-cluster distance is less than the preset inter-cluster distance threshold or the updated number of clusters is equal to the target value, obtain a preliminary blockchain sharding result.

[0104] In this embodiment, for the target optimization transaction model based on the blockchain sharding system, assume that the blockchain sharding system contains N accounts, denoted as A = {a 1 , a 2 , …, a N}. The historical transactions between accounts can be represented by a transaction frequency matrix M ∈ R N×N , where M ij represents the transaction between account a i and a jThe trading frequency within the last epoch time window. To measure the similarity between accounts, the present invention defines a distance formula based on trading frequency, and its expression is:

[0105]

[0106] where max(M) represents the maximum value in the trading frequency matrix, which is used for normalization to ensure that all distance values are within the range of [0, 1]. The smaller this distance value is, the closer the trading relationship between the two accounts is.

[0107] Agglomerative Hierarchical Clustering is adopted. Its characteristic is to start from N single-account clusters and generate a hierarchical clustering structure by gradually merging the closest clusters. Initially, each account is an independent cluster. Let the initial cluster set be C = { {a 1}, {a 2}, …, {a N}}. During the clustering process, the minimum distance between all clusters needs to be calculated each time. The single-linkage method is used to define the distance between clusters, that is:

[0108]

[0109] where C i and C j represent two clusters, and a p and a q are any accounts belonging to C i and C j respectively. After finding the two closest clusters C i and C j , they are merged into a new cluster C k = C i ∪C j , and the cluster set is updated to:

[0110] C ← (C\{C i , C j}) ∪ {C k}

[0111] This process is continuously iterated until the preset stop condition is met. The stop condition can be that the number of clusters reaches the target value k, or the minimum distance between clusters D(C i , C j ) exceeds the preset threshold ρ. After clustering is completed, the system generates k clusters, and each cluster contains a group of accounts with high trading frequencies and closely related to each other.

[0112] The results of hierarchical clustering provide a good initial grouping structure for subsequent refinement processing. Although this stage can effectively identify account groups with strong trading relationships, it insufficiently considers the distribution differences in the density within the clusters and may miss some special noisy accounts. Therefore, it is necessary to further refine in combination with density clustering in the next stage to identify the core accounts in the high-density clusters and the noisy accounts in the low-density regions, thereby optimizing the quality of account division.

[0113] S300. Perform partitioning processing on the preliminary blockchain sharding results through the DBSCAN clustering algorithm to obtain the blockchain sharding results after secondary partitioning;

[0114] Specifically, construct a subset data matrix based on the trading frequency between accounts in the clusters in the preliminary blockchain sharding results, and determine the account distances in the subset data matrix; based on the DBSCAN clustering algorithm, define the neighborhood radius and the minimum number of neighborhood points; traverse each account in the preliminary blockchain sharding results through the DBSCAN clustering algorithm, and construct a neighborhood point set for each account according to the account distances and the neighborhood radius in the subset data matrix; determine the number of neighborhood points of the account and make a judgment based on the neighborhood point set of the account; if the number of neighborhood points of the account is greater than or equal to the minimum number of neighborhood points, mark the account as a core account, and add all the accounts in the neighborhood of the core account to the current cluster until no new accounts can be added, to obtain the density clusters; if the number of neighborhood points of the account is less than the minimum number of neighborhood points and belongs to a density cluster, mark the account as a border account; if the number of neighborhood points of the account is less than the minimum number of neighborhood points and does not belong to a density cluster, mark the account as a noisy account; integrate the density clusters, border accounts, and noisy accounts to obtain the blockchain sharding results after secondary partitioning.

[0115] In this embodiment, the present invention uses the DBSCAN clustering algorithm to further refine the results of hierarchical clustering in the first stage, so as to identify the possible core accounts, border accounts, and noisy accounts in each cluster. The characteristic of DBSCAN is a density-based clustering method that does not require specifying the number of clusters in advance, can effectively identify the clustering in high-density regions through local density information, and at the same time detect low-density noise points, which highly matches the characteristics of the accounts that need to be dynamically processed in the blockchain sharding system.

[0116] In this stage, first take each cluster generated by hierarchical clustering as the input set of DBSCAN. Assume that the results of hierarchical clustering are divided into k initial clusters, denoted as C = { { C 1}, { C 2}, …, { C k}}. For each cluster C i , use the trading frequency between the accounts in the cluster as a feature to construct a subset data matrix M i. During the processing of DBSCAN, two key parameters are defined: the neighborhood radius ε and the minimum number of neighborhood points MinPts. Among them, ε determines the distance threshold at which points are considered "adjacent", while MinPts defines the minimum number of points required to form a high-density region. In the transaction frequency matrix, account a p and a q 's distance can be measured by the following formula, and its expression is:

[0117]

[0118] f(a p , a q ) = T i [p][q]

[0119] where T i [p][q] represents the transaction frequency between account a p and a q . The normalized distance metric d(a p , a q ) ensures that the relative transaction closeness between accounts can effectively reflect their neighborhood relationship. a r and a s represent the accounts used for normalization reference, and they are associated with a p and a q respectively.

[0120] The first step of DBSCAN processing is to traverse all accounts in each cluster C i , calculate the neighborhood point set N p (a ε ), which is defined as the set of accounts that meet the following conditions, and the specific expression is: p N ε

[0121] (a p ) = {a q ∈ C i | d(a p , a q ) ≤ ε}

[0122] If the number of neighborhood points |N ε (a p )| of an account is ≥ MinPts, then the account is marked as a core account (CoreAccount), indicating that it is at the center of a high-density region. Using the core account as a seed, DBSCAN starts to expand the cluster: add all accounts in its neighborhood to the current cluster, and repeat calculating the neighborhood for these newly added accounts to find potential other core accounts for further expansion. This process recurs until no new accounts can be added, forming a complete density cluster.​

[0123] For those accounts with the number of neighborhood points less than MinPts, if they exist in the neighborhood of a core account, they are marked as border accounts. Although border accounts do not belong to the center of the high-density area, they are associated with the high-density area to a certain extent and can therefore be included in the corresponding cluster. Accounts that neither meet the conditions of core accounts nor belong to the neighborhood of any core account are marked as noise accounts. The identification process of noise accounts is an important feature of DBSCAN, and its goal is to discover those accounts that are isolatedly distributed, which may have an important impact on subsequent sharding load balancing and cross-sharding transaction optimization.

[0124] As DBSCAN is gradually executed, the initial cluster C generated by hierarchical clustering is further refined into density clusters with higher resolution. The accounts within each cluster are clearly divided into core accounts, border accounts, and noise accounts, and finally a fine cluster partition structure and the set of noise accounts for each cluster are output. The core of this stage lies in making full use of density information to deeply analyze the account distribution, thereby enhancing the robustness of the system and the flexibility of sharding division.

[0125] S400. Based on the blockchain sharding results after secondary partitioning, perform noise account partitioning processing to obtain the final blockchain sharding results.

[0126] Specifically, according to the set of noise accounts, measure and partition the noise accounts to determine the transaction frequency of the noise accounts; define a transaction frequency threshold, mark the noise accounts with a transaction frequency greater than the transaction frequency threshold as high-frequency noise accounts, and mark the noise accounts with a transaction frequency less than the transaction frequency threshold as low-frequency noise accounts; obtain the transaction correlation scores between the high-frequency noise accounts and all shards in the blockchain sharding results after secondary partitioning, and assign the high-frequency noise accounts to the shard corresponding to the highest transaction correlation score; evenly distribute the low-frequency noise accounts to all shards in the blockchain sharding results after secondary partitioning through a random hashing algorithm to obtain the final blockchain sharding results.

[0127] In this embodiment, after hierarchical clustering and DBSCAN are completed, the accounts marked as noise refer to abnormal accounts that do not belong to any cluster. These accounts often cannot form a tight transaction relationship cluster with other accounts due to the isolation or abnormality of their transaction behaviors. Noise accounts are not completely equivalent to low-transaction-frequency accounts, which may include accounts with high transaction frequencies but discrete transaction distributions, or accounts with low transaction frequencies and sparse transaction counterparts. To more precisely handle these accounts, we need to further classify the noise accounts, clarify the nature of their transaction frequencies, and thus adopt targeted optimization strategies.

[0128] First, measure and divide the noise accounts according to the trading frequency. The trading frequency of each account can be defined as the sum of all its trading times within a specified time window, denoted as S(a i ), and its expression is as follows:

[0129]

[0130] To distinguish high-frequency noise accounts from low-frequency noise accounts, a threshold θ of the trading frequency can be set as needed F . When S(a i ) ≥ θ F , the account a i is considered a high-frequency noise account; conversely, when S(a i ) ≤ θ F , it is regarded as a low-frequency noise account. The setting of the threshold θ F can be dynamically adjusted according to the statistical distribution of the trading frequency. For example, the upper quartile of the trading frequency can be used as the basis for division.

[0131] The characteristic of high-frequency noise accounts is that although their trading frequency is high, their trading distribution is relatively discrete, and there are trading transactions with accounts in multiple shards, and they cannot be effectively attributed to a specific cluster. While low-frequency noise accounts are characterized by a relatively low total trading frequency and usually only have sparse transactions with a small number of accounts. After such division, we can design appropriate processing strategies based on the above characteristics.

[0132] For high-frequency noise accounts, it is necessary to give priority to considering their trading relevance with each shard to reduce the overhead of cross-shard transactions. We define and calculate the trading relevance score R(a i , C q ), and its expression is as follows:

[0133]

[0134] where C q represents the set of accounts in the q-th shard. The high-frequency noise account a i will be assigned to the shard C k with the highest relevance score, that is:

[0135]

[0136] This allocation strategy can effectively reduce the frequency of cross-shard transactions by preferentially meeting the main trading needs of high-frequency noise accounts, thereby significantly improving the trading efficiency and throughput capacity of the system.

[0137] For low-frequency noise accounts, the allocation strategy can be flexibly selected according to actual needs. One method is to use a random hashing algorithm to evenly distribute low-frequency noise accounts to each shard, thereby further achieving load balancing for account transactions between shards. Another method is to centrally allocate these accounts to a dedicated cold shard. Designing such a shard to specifically handle low-frequency account transactions highly matches the situation of uneven node computing power in the actual shard system. By having the shard composed of low-computing-power nodes be responsible for handling low-frequency noise accounts, the computing power resources in the system can be fully utilized, thereby achieving more efficient load balancing.

[0138] Through the classification and refined processing of noise accounts, the algorithm realizes the efficient management of abnormal accounts and avoids the degradation of system performance caused by noise accounts. This strategy can effectively reduce the cross-shard transaction ratio, balance shard loads, and at the same time improve the throughput and response ability of the system.

[0139] In summary, the pseudo-code process of the clustering composite algorithm in the embodiments of the present invention is shown in Table 1. First, the reference committee needs to regard all transaction matrices as a cluster C = { { C 1}, { C 2}, …, { C k}}, and use agglomerative hierarchical clustering to select the two closest clusters for merging each time to generate a specified number of clusters (i.e., Line 1 - 7); then, within each preliminary cluster, further apply DBSCAN with preset parameters ε and MinPts for refined partitioning to capture the local characteristics of the transaction frequency within the cluster and identify noise accounts (i.e., Line 8 - 17); for noise accounts, they are divided into high-frequency and low-frequency categories according to their transaction frequencies. High-frequency accounts are allocated to the optimal shard according to relevance, and low-frequency accounts are randomly allocated or centrally processed (i.e., Line 18 - 22). Finally, the reconstructed account network can be broadcast to all shards, and after each shard receives the broadcast, it updates the local information to the latest status.

[0140] Table 1 Pseudo-code process table of the clustering composite algorithm

[0141]

[0142]

[0143] Finally, referring to Figure 3 , in the embodiments of the present invention, the composite clustering algorithm is adopted to dynamically partition the accounts for shard reconstruction of the blockchain system. The protocol execution is divided into five stages: user transaction request, node screening and preliminary clustering, in-cluster refinement and noise account processing, account partitioning and shard reconstruction, and shard result synchronization.

[0144] 1) User Transaction Request: The user submits a transaction request to the sharded blockchain system via the network. The transaction is assigned to the transaction pool for processing according to the account relationship and shard mapping. Since the propagation time of the transaction in the network is much less than the time required for shard consensus, it is assumed that the transaction is synchronized and regarded as reaching all shards. Each honest node manages the transactions with a unified view based on the global transaction pool maintained locally. The transaction pool design includes a resource allocation mechanism to prevent network resource congestion and flooding attacks on a single shard, and at the same time provides a stable basic data source for subsequent shard division.

[0145] 2) Node Screening and Initial Clustering: The node extracts transaction data from the transaction pool and constructs a transaction frequency matrix between accounts. Subsequently, a new round of account shard division begins. First, hierarchical clustering is used for rough division. Referring to the transaction frequency matrix, the node performs hierarchical clustering on the accounts and initially divides them into larger clusters. The number of clusters does not need to be specified in advance at this stage. If you want to divide into specific shard numbers, a global preliminary shard scheme can be obtained by truncating the clustering tree.

[0146] 3) Intra-cluster Refinement and Noisy Accounts: The DBSCAN clustering method is run separately for each cluster obtained by hierarchical clustering to further refine the cluster structure. The DBSCAN clustering method can not only detect high-density regions in the cluster but also identify low-density isolated accounts (i.e., noisy accounts). For noisy accounts, they can be divided into high-frequency noisy accounts and low-frequency noisy accounts according to the transaction frequency, and then different noise processing methods are used according to the type of noisy accounts to reduce their impact on shard performance and security.

[0147] 4) Account Partitioning and Shard Reconstruction: The shard division results are executed and reconstructed by the reference committee in each epoch. The committee first verifies the node identity, updates the active node list, and adjusts the shard affiliation of some nodes according to the established rules to enhance the system's dynamics and security. Combining the results of hierarchical clustering and DBSCAN, the reference committee balances the shard load based on the transaction frequency and account correlation, optimizes the transaction flow between shards, and completes the formulation of the account shard scheme. After the shard division is completed, the current transaction pool state, account partition information, and shard state are written into the state block to provide data support for the next epoch.

[0148] 5) Shard Result Synchronization: In this stage, peer nodes will update their states according to the information in the state block. The node judges whether it needs to migrate shards according to the account partition result in the state block. For nodes that change shards, they send join requests to the new shards, and at the same time update the ledger information and the neighbor relationship within the shard. After the update is completed, the system starts a new round of consensus, and at the same time the node processes new transaction requests according to the latest shard information to ensure the stable operation of the shard system.

[0149] Specifically, in the embodiments of the present invention, by analyzing the historical transaction relationships between accounts, accounts with a high degree of correlation are classified into the same shard, fundamentally reducing the occurrence of cross-shard transactions. At the same time, to avoid the problem of uneven resource distribution between shards, through a reasonable and uniform partitioning strategy, these strongly correlated accounts are further distributed to each shard, so as to achieve balanced distribution of shard loads while reducing cross-shard transactions. This research motivation provides a theoretical basis and practical direction for solving the key challenges in the practical application of sharding technology, and has important practical significance.

[0150] To verify whether the historical transaction relationships between accounts can reflect the probability of future transactions occurring, the embodiments of the present invention conduct experimental analysis based on 1,000,000 real historical transaction data of Ethereum. The experiment first extracts and analyzes the account pair relationships in the actual transaction data, counts the transaction frequencies of each pair of accounts, and calculates the proportion of repeated transaction pairs in the actual transaction pairs to verify the existence and degree of the correlation between accounts in the real transaction network. Subsequently, by randomly generating a set of account pairs of the same scale, it is checked whether these random account pairs appear in the historical transaction pairs, and their repeated transaction proportion is calculated as a comparison benchmark.

[0151] The experimental results are as Figure 4 shown. The repeated transaction proportion of the actual transaction pairs is as high as 29.32%, while the repeated transaction proportion of the randomly generated account pairs is only 0.92%. This significant difference indicates that the correlation between accounts in the historical transaction pairs is significantly higher than that of the random account pairs, supporting the hypothesis that there is a stronger dependence between highly active accounts. The experiment further verifies the feasibility of inferring the probability of future transactions using historical transaction data, provides empirical support for the subsequent proposed shard partitioning scheme optimized based on historical account transactions, and realizes dynamic balance of shard loads and optimization of cross-shard transactions.

[0152] In summary, the embodiments of the present invention have the following improvement points compared with the prior art:

[0153] 1) Through the composite clustering algorithm, high-frequency associated accounts are preferentially aggregated to maximize the transaction density within the shard, significantly reducing the occurrence of cross-shard transactions. This optimization reduces the cross-shard communication cost and improves the system throughput.

[0154] 2) By classifying and processing low-frequency accounts, a flexible allocation strategy is designed. Low-frequency accounts are randomly and evenly distributed or concentrated into a dedicated "cold account shard", thus achieving dynamic balance of the loads between shards and making full use of the computing power resources of the system.

[0155] 3) It can accurately identify noise accounts with abnormal transactions. High-frequency noise accounts are isolated to the shard with the strongest correlation to prevent abnormal transactions from affecting the whole; low-frequency noise accounts are centrally processed, reducing their interference with the core shards. This mechanism effectively improves the anti-attack ability of the system.

[0156] Refer to Figure 2 , a dynamic blockchain sharding system based on a composite clustering algorithm, including:

[0157] The first module 201 is used to determine the transaction frequency matrix and the shard allocation matrix based on the blockchain sharding system, and construct the target optimized transaction model of the blockchain sharding system;

[0158] The second module 202 is used to perform preliminary global partitioning processing through the agglomerative hierarchical clustering algorithm based on the target optimized transaction model of the blockchain sharding system to obtain the preliminary blockchain sharding result;

[0159] The third module 203 is used to perform partitioning processing on the preliminary blockchain sharding result through the DBSCAN clustering algorithm to obtain the blockchain sharding result after secondary partitioning;

[0160] The fourth module 204 is used to perform noise account partitioning processing based on the blockchain sharding result after secondary partitioning to obtain the final blockchain sharding result.

[0161] The content in the above method embodiments is applicable to this system embodiment. The functions specifically implemented by this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0162] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A dynamic blockchain sharding method based on a composite clustering algorithm, characterized in that: The following steps are involved: Based on the blockchain sharding system, determine the transaction frequency matrix and sharding allocation matrix, and build a target optimization transaction model for the blockchain sharding system; Based on the target optimization transaction model of the blockchain sharding system, a preliminary global partitioning process is performed through the agglomerative hierarchical clustering algorithm to obtain the preliminary blockchain sharding results; The preliminary blockchain sharding results are divided and processed by the DBSCAN clustering algorithm to obtain the blockchain sharding results after secondary division; Based on the blockchain sharding results after secondary partitioning, noise account partitioning processing is performed to obtain the final blockchain sharding results.

2. According to claim 1, a dynamic blockchain sharding method based on a composite clustering algorithm is characterized in that: The step of determining the transaction frequency matrix and the shard allocation matrix based on the blockchain sharding system and constructing a target optimization transaction model for the blockchain sharding system specifically includes: Based on the blockchain sharding system, obtain historical transaction data between all accounts and build a transaction frequency matrix; Based on the blockchain sharding system, define the total frequency of cross-shard transactions and the system load balancing objective function; Introduce the preset weight coefficient, combine the total frequency of cross-shard transactions with the system load balancing objective function and perform matrix representation to obtain the shard allocation matrix; The transaction frequency matrix and the shard allocation matrix are combined to represent the calculation process of cross-shard transactions, and a target optimization transaction model for the blockchain sharding system is constructed.

3. According to claim 2, a dynamic blockchain sharding method based on a composite clustering algorithm is characterized in that: The expression of the target optimization transaction model of the blockchain sharding system is specifically as follows: In the above formula, λ represents the weight coefficient, M represents the transaction frequency matrix, Z represents the shard allocation matrix, and Z T represents the transposed matrix of the shard allocation matrix, i represents the i-th account, j represents the j-th account, and Z ip Indicates whether account i is assigned to shard p, Z jq indicates whether account j is assigned to shard q, and sum(·) represents the total transaction frequency of all cross-shard transactions.

4. According to claim 3, a dynamic blockchain sharding method based on a composite clustering algorithm is characterized in that: The target optimization transaction model based on the blockchain sharding system performs preliminary global partitioning processing through agglomerative hierarchical clustering algorithm to obtain preliminary blockchain sharding results, which specifically includes: Based on the goal optimization of the transaction frequency matrix in the transaction model of the blockchain sharding system, a distance expression based on transaction frequency is defined to obtain the similarity between accounts; Treat an account as a cluster and build an initial cluster set; Based on the similarity between accounts, the inter-cluster distance of the initial cluster set is obtained by the single linkage method, and the minimum value of the inter-cluster distance and the number of clusters in the initial cluster set are determined; The initial cluster set is clustered by the agglomerative hierarchical clustering algorithm. If the minimum value of the inter-cluster distance is greater than the preset inter-cluster distance threshold or the number of clusters is not less than the target value, the two clusters corresponding to the minimum value of the inter-cluster distance are iteratively merged, and the minimum value of the inter-cluster distance and the number of clusters are updated; Until the minimum value of the updated inter-cluster distance is less than the preset inter-cluster distance threshold or the updated number of clusters reaches the target value, the preliminary blockchain sharding result is obtained.

5. According to claim 4, a dynamic blockchain sharding method based on a composite clustering algorithm is characterized in that: The expression of the inter-cluster distance is specifically as follows: In the above formula, C i , C j Represents two clusters, a p with a q Indicates that it belongs to C i and C j Any account of p ,a q ) means a p Account and a q The similarity between accounts, D(C i ,C j ) means C i Cluster and C j The distance between clusters.

6. According to claim 5, a dynamic blockchain sharding method based on a composite clustering algorithm is characterized in that: The step of partitioning the preliminary blockchain sharding result by using the DBSCAN clustering algorithm to obtain the blockchain sharding result after secondary partitioning specifically includes: Construct a subset data matrix based on the transaction frequency between accounts in the cluster in the preliminary blockchain sharding results, and determine the account distance in the subset data matrix; Based on the DBSCAN clustering algorithm, the neighborhood radius and the minimum number of neighborhood points are defined; The DBSCAN clustering algorithm is used to traverse each account in the preliminary blockchain sharding results, and the neighborhood point set of each account is constructed based on the account distance and neighborhood radius in the subset data matrix; Determine the number of neighborhood points of the account and make a judgment based on the neighborhood point set of the account; If the number of neighborhood points of an account is greater than or equal to the minimum number of neighborhood points, the account is marked as a core account, and all accounts in the neighborhood of the core account are added to the current cluster until no new accounts can be added, and a density cluster is obtained; If the number of neighborhood points of an account is less than the minimum number of neighborhood points and belongs to a density cluster, the account is marked as a boundary account; If the number of neighborhood points of an account is less than the minimum number of neighborhood points and does not belong to a density cluster, the account is marked as a noise account; Integrate density clusters, boundary accounts and noise accounts to obtain the blockchain sharding result after secondary partitioning.

7. According to claim 6, a dynamic blockchain sharding method based on a composite clustering algorithm is characterized in that: The expression of the account distance in the subset data matrix is ​​specifically as follows: f(q p ,a q )=T i [p][q] In the above formula, d(a p ,a q ) represents the account distance in the subset data matrix, T i [p][q] indicates account a p and a q The transaction frequency between, f(·) represents the transaction frequency function, a r 、a s Represents the account used for normalization reference.

8. According to claim 7, a dynamic blockchain sharding method based on a composite clustering algorithm is characterized in that: The step of performing noise account partitioning processing based on the blockchain sharding result after the secondary partitioning to obtain the final blockchain sharding result specifically includes: According to the noise account set, the noise accounts are measured and divided to determine the transaction frequency of the noise accounts; Define a transaction frequency threshold, mark the noise account with a transaction frequency greater than the transaction frequency threshold as a high-frequency noise account, and mark the noise account with a transaction frequency less than the transaction frequency threshold as a low-frequency noise account; Obtain the transaction correlation scores of the high-frequency noise account and all shards in the blockchain sharding result after secondary partitioning, and assign the high-frequency noise account to the shard with the highest transaction correlation score; Through the random hash algorithm, low-frequency noise accounts are evenly distributed to all shards in the blockchain sharding result after secondary partitioning, or low-frequency noise accounts are concentrated in the cold account shards to obtain the final blockchain sharding result.

9. According to claim 8, a dynamic blockchain sharding method based on a composite clustering algorithm is characterized in that: The expression of the transaction relevance score is specifically as follows: In the above formula, R(a i ,C q ) represents the transaction relevance score, a i represents the high frequency noise account, C q represents the account set of the qth shard, M ij Indicates account a i with a j The transaction frequency in the last epoch time window.

10. A dynamic blockchain sharding system based on a composite clustering algorithm, characterized in that: Includes the following modules: The first module is used to determine the transaction frequency matrix and the shard allocation matrix based on the blockchain sharding system, and to build a target optimization transaction model for the blockchain sharding system; The second module is used to optimize the transaction model based on the target of the blockchain sharding system, and perform preliminary global partitioning processing through the agglomerative hierarchical clustering algorithm to obtain preliminary blockchain sharding results; The third module is used to partition the preliminary blockchain sharding results through the DBSCAN clustering algorithm to obtain the blockchain sharding results after secondary partitioning; The fourth module is used to perform noise account partitioning processing based on the blockchain sharding result after the secondary partitioning to obtain the final blockchain sharding result.

Citation Information

Patent Citations

  • Fragmentation account adjustment method of block chain fragmentation system and related device

    CN114037531A

  • Fragmentation method and device of block chain system and electronic equipment

    CN116938928A

  • Inductive block chain account distribution method and device based on graph neural network

    CN117391858A

  • Blockchain sharding method combining spectral clustering and reputation value mechanism

    US20240073043A1