A method for constructing a data index of distributed data storage

By building a hypergraph model in a distributed data storage system and applying the minimum cutting theorem, consistent hashing algorithm, redundant coding and game theory model, the problems of uneven load, high data migration cost and insufficient resource utilization in distributed data storage systems are solved, and efficient load balancing, low-cost data migration and efficient resource utilization are achieved.

CN119829551BActive Publication Date: 2025-06-10BEIJING HANXINSHENG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510308247.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-10
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

In existing distributed data storage systems, there are problems such as low data index construction efficiency, uneven load allocation, high data migration costs and insufficient storage resource utilization.

Method used

The hypergraph model of distributed storage system is constructed through advanced graph theory, the minimum cutting theorem is used to optimize load balancing, combined with a consistent hashing algorithm and virtual node mechanism to reduce data migration, redundant coding technology is used to improve fault tolerance, and dynamically adjust load allocation through game theory model and optimal control theory.

Benefits of technology

It significantly improves data retrieval efficiency, ensures system load balancing, reduces data migration overhead, improves the utilization rate of storage resources and the scalability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119829551B_ABST
    Figure CN119829551B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of distributed data storage, and discloses a method for constructing a data index of distributed data storage, including the following steps: constructing a hypergraph model of a distributed storage system through high-order graph theory, the hypergraph model including multiple storage nodes, multiple data blocks, and multiple query requests; performing load balancing optimization on the storage nodes and query requests based on the minimum cut theorem in the high-order graph model; mapping the data blocks and storage nodes to a consistent hashing ring according to the consistent hashing algorithm; reducing data migration through a virtual node mechanism; redundantly storing the data blocks on multiple storage nodes by adopting a redundant coding method; adjusting the load distribution strategy of the storage nodes through a game theory model; and optimizing the load adjustment by adopting an optimal control method. The present invention can improve the efficiency of the data storage system while ensuring the high availability, scalability, and reliability of the system, and is applicable to large-scale distributed data storage systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed data storage, and particularly to a method for constructing a data index of distributed data storage. Background Art

[0002] With the continuous growth of data volume and system scale, the efficient storage and fast access of data have become key issues. Traditional distributed storage methods often rely on simple hash algorithms and static data allocation mechanisms. Although they can ensure the uniform distribution and high availability of data, when facing large-scale data storage and frequent node expansion, the performance and scalability problems of the system gradually emerge.

[0003] Traditional load balancing methods usually configure storage nodes in a static manner and lack the ability to adapt to dynamic load changes in real time. When some nodes become bottlenecks due to excessive access volume, the system often cannot effectively adjust the load, resulting in some nodes being overloaded while others are idle. This not only reduces the utilization efficiency of resources but also affects the speed of data query and the overall performance of the system.

[0004] Current distributed data storage systems have a large data migration overhead during expansion. Although traditional consistent hashing algorithms can solve some data migration problems, when the number of nodes changes drastically, a large amount of data redistribution is still required. Especially in scenarios of large-scale expansion and dynamic node changes, the cost of data migration will be very high, resulting in a slow response speed of the system and even affecting data consistency and stability. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method for constructing a data index of distributed data storage, which solves the problems of low efficiency of data index construction, uneven load distribution, high data migration cost, and insufficient utilization of storage resources in existing distributed data storage systems.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for constructing a data index of distributed data storage includes the following steps:

[0007] Construct a hypergraph model of the distributed storage system through high-order graph theory, where the hypergraph model includes multiple storage nodes, multiple data blocks, and multiple query requests;

[0008] Optimize the load balance of storage nodes and query requests based on the minimum cut theorem in the high-order graph model;

[0009] Map data blocks and storage nodes to the consistent hash ring according to the consistent hashing algorithm, and reduce data migration through the virtual node mechanism;

[0010] The data blocks are redundantly stored on multiple storage nodes by using a redundant coding method;

[0011] The load distribution strategy of the storage nodes is adjusted through a game theory model, and the load adjustment is optimized by using an optimal control method.

[0012] Preferably, each hyperedge in the high-order graph theory model represents the dependency relationship between multiple data blocks and multiple storage nodes or query requests, and the cutting cost of the hyperedge is optimized through the minimum cut theorem to ensure the shortest query path and achieve load balancing.

[0013] Preferably, in the consistent hashing algorithm, the storage nodes and data blocks are mapped to the consistent hash ring through hash values, each storage node is mapped to the hash ring through multiple virtual nodes, and the load of the nodes is optimized through the virtual node mechanism.

[0014] Preferably, the redundant coding adopts Reed - Solomon coding, each data block is stored on multiple storage nodes, and the data storage overhead is further reduced through a data compression method.

[0015] Preferably, the game theory model adjusts the load distribution strategy of the storage nodes through the Nash equilibrium principle, so that each storage node selects an optimal load distribution scheme considering the loads of other nodes.

[0016] Preferably, the load state variable of the system is defined , representing the load states of each storage node at time t. Assuming that the system contains N storage nodes, the load state variable represents the load state of node i at time t, and the update of the load state is based on the number of requests processed by the node, the amount of data stored, and the query load factors. Specifically:

[0017] ;

[0018] Among them, is the computing power of the storage node, is the query request volume of node i, is the amount of data stored by node i;

[0019] The load optimization objective function J of the system is set. This objective function describes the load distribution of all nodes, and the load imbalance of the entire system is minimized by adjusting the node loads. The load optimization objective function J represents the cumulative load cost within the time window T. Specifically:

[0020] ;

[0021] Among them, the integral symbol For the accumulation of the overall system load cost within the time interval [0, T], is the load cost function of the node, reflecting the increasing cost with higher node load, and T is the optimized time window;

[0022] Solve the optimal control strategy of load distribution through the optimal control theory , where is the load adjustment amount of storage node i at time t, and the control objective is to dynamically adjust to minimize the system load within the given time window and satisfy the constraint conditions of load balancing:

[0023] ;

[0024] The constraint conditions include:

[0025] ;

[0026] where, is the total system load, representing the sum of the loads of all storage nodes, is the load upper limit of a single node, and the constraint condition means that the load of each storage node i must be less than or equal to the maximum load upper limit , indicates that the constraint condition applies to each storage node i;

[0027] Based on the Lagrange multiplier method or the variational method, by solving the optimal control problem in the above formula, obtain the optimal load adjustment strategy of each storage node i to achieve load balancing within the global scope of the system and minimize the overall load cost;

[0028] According to the solved optimal control strategy , dynamically adjust the load distribution of each storage node, so that each storage node performs the corresponding load adjustment under the constraints of load balancing and performance optimization, and then achieve the global optimal load distribution of the system.

[0029] Preferably, when the system is expanded, combining the consistent hashing algorithm and the high-order graph model, minimizing data migration includes the following steps:

[0030] Adjust the data mapping on the hash ring through the consistent hashing algorithm, and combine the load balancing strategy of the high-order graph model to minimize the data migration amount between nodes during the expansion process.

[0031] Preferably, the load balancing optimization adopts a graph cutting and recombination strategy based on high-order graph theory. When the node loads are uneven, optimize the load distribution by adjusting the dependency relationship between data blocks and storage nodes.

[0032] The present invention provides a method for constructing a data index for distributed data storage. It has the following beneficial effects:

[0033] 1. The present invention optimizes the load distribution between storage nodes and query requests through a high-order graph model and the minimum cut theorem, significantly reducing the query path length and improving the data retrieval efficiency. At the same time, by combining a game theory model and the optimal control theory, the load of storage nodes is dynamically adjusted to ensure system load balance, avoid overloading or idling phenomena, and improve the overall performance of the system.

[0034] 2. Through the consistent hashing algorithm and the virtual node mechanism, the present invention minimizes the amount of data migration during system expansion, only migrates necessary data blocks, and avoids the high overhead of full-scale data migration. The system can be smoothly expanded, reducing the performance loss caused by expansion, and improving the scalability and sustainability of the system.

[0035] 3. The present invention redundantly stores data blocks on multiple nodes by adopting a redundant coding technique to ensure rapid data recovery in case of a single node failure, enhancing the fault tolerance and reliability of the system. In addition, through the combination of a high-order graph model and the consistent hashing algorithm, the system can continuously optimize resource allocation during dynamic load changes and expansion processes to ensure the stable operation of the system. Description of the Drawings

[0036] Figure 1 It is a schematic diagram of the method flow of the present invention. Detailed Embodiments

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Please refer to the attached Figure 1 , the embodiment of the present invention provides a method for constructing a data index for distributed data storage, including the following steps:

[0039] S1. Construct a hypergraph model of the distributed storage system through high-order graph theory. The hypergraph model includes multiple storage nodes, multiple data blocks, and multiple query requests;

[0040] S2. Optimize the load balance of storage nodes and query requests based on the minimum cut theorem in the high-order graph model;

[0041] S3. Map data blocks and storage nodes to a consistent hashing ring according to the consistent hashing algorithm, and reduce data migration through the virtual node mechanism;

[0042] S4. Redundantly store data blocks on multiple storage nodes using a redundant coding method;

[0043] S5. Adjust the load distribution strategy of storage nodes through a game theory model, and optimize load adjustment using an optimal control method;

[0044] S6. When the system is expanded, combine the consistent hashing algorithm and the high-order graph model to minimize data migration.

[0045] Regarding S1. In this embodiment, a hypergraph model of a distributed storage system is constructed through high-order graph theory. The hypergraph model includes multiple storage nodes, multiple data blocks, and multiple query requests. First, a hypergraph model is constructed through high-order graph theory. The hypergraph model includes multiple storage nodes, multiple data blocks, and multiple query requests. The construction of this hypergraph model aims to provide a theoretical basis for subsequent load balancing optimization, data mapping, and query path optimization. Specifically, the model abstracts and models the relationships between various elements in the system, so that the optimization of data storage and query can be carried out more precisely and efficiently.

[0046] In a possible implementation, the nodes in the hypergraph model are divided into three categories: storage nodes, data blocks, and query requests. Each category of nodes plays a different role in the graph and is connected to other nodes through hyperedges, forming a composite relationship network.

[0047] Storage nodes: Storage nodes are the core components in a distributed system, usually referring to hardware devices or virtualized storage instances that carry out data storage functions. In some embodiments, storage nodes can be independent servers, storage arrays, or cloud storage instances, etc.

[0048] Data blocks: Data blocks are the basic data units stored in the system. Each data block can be power monitoring data, data collected by voltage and current sensors, etc. In the embodiments of the present invention, the granularity of data blocks is flexibly adjustable, and can be either a data point of a single sensor or a multi-dimensional data set.

[0049] Query requests: Query requests represent the access requirements for data in the system, usually initiated by end users, and require obtaining certain data blocks or related information from storage nodes. The types of query requests are diverse, including time range queries, conditional filtering queries, etc.

[0050] In this embodiment, the hyperedges in the hypergraph model represent the dependency relationships between storage nodes, data blocks, and query requests. Specifically, the settings of hyperedges can be defined in the following two ways:

[0051] The hyperedge representing the relationship between the storage node and the data block indicates that the data block is stored on one or more storage nodes.

[0052] The hyperedge representing the relationship between the query request and the data block indicates that a certain query request involves the access of one or more data blocks.

[0053] The hypergraph model can clearly express the storage distribution of data blocks, the access path of query requests, and the dependency relationship between nodes, thus laying a foundation for subsequent optimization steps.

[0054] To better understand this hypergraph model, a specific example can be used to illustrate: Suppose in a power monitoring system, the storage nodes include multiple power module monitoring devices, the data blocks include the real-time voltage and current data of each power module, and the query request is the query operation of the terminal user requesting power data. The hypergraph model abstracts the storage nodes, data blocks, and query requests into three different types of nodes in the graph and connects them through hyperedges to form a complete dependency network.

[0055] In a possible implementation, the construction of the hypergraph model can also be weighted based on the mapping relationship between the storage node and the data block, as well as the priority and access pattern of the query request. Specifically, for each query request , an access weight can be set, indicating the processing priority of this query request; while for each storage node , a load weight can be set, indicating the load situation of this node.

[0056] Based on the above model, the hypergraph model can be formally described by the following mathematical representation:

[0057] ;

[0058] Among them, is the node set of the hypergraph, including storage nodes , data blocks , and query requests . E is the hyperedge set, representing the relationship between nodes:

[0059] Each hyperedge can be defined by the following subset:

[0060] ;

[0061] Among them, e is a relationship that connects storage nodes, data blocks, and query requests.

[0062] To express the load situation between nodes more precisely, the relationship strength between each storage node and the data block can be defined and represented by weights. These weights can be dynamically adjusted based on the storage capacity, processing power of the storage node, and the access frequency of the data block. For example, for the storage node and the data block the relationship strength can be calculated according to the following formula:

[0063] ;

[0064] where represents the access frequency of the data block and is the capacity of the storage node .

[0065] In this way, the hypergraph model not only provides the topological relationships between storage nodes, data blocks, and query requests, but also takes into account factors such as load and capacity, thus being able to provide more accurate information for load balancing optimization.

[0066] The hypergraph model of this embodiment is not limited to describing a single type of data block and storage node. In some embodiments, the hypergraph model can also be extended to support multi-level and multi-dimensional complex data storage structures. For example, storage nodes may be classified according to different data types and access frequencies, and query requests can be further divided into multiple sub-requests according to their query conditions. Such a multi-level and multi-dimensional hypergraph model can more comprehensively express the complex relationships of the storage system.

[0067] In addition, the hypergraph model can also be used in combination with other data structures, such as hash tables, B-trees, etc., to achieve fast data query and efficient access. Under this combination of multi-level data structures, the hypergraph model can better adapt to the dynamic changes in large-scale distributed storage systems.

[0068] The hypergraph model constructed by high-order graph theory in this embodiment provides a basic framework for subsequent storage optimization, query path optimization, and load balancing. Through this method, a more efficient distributed storage data index construction can be achieved, and reliable guarantees can be provided for the scalability, fault tolerance, and query performance of the system.

[0069] Regarding S2, in this embodiment, the minimum cut theorem in the high-order graph model is used to optimize the load balance between storage nodes and query requests. Specifically, the hyperedges in the hypergraph are optimized through the minimum cut theorem, so as to achieve the purpose of load balance, ensure that the query requests are distributed as evenly as possible among the storage nodes, and avoid performance bottlenecks caused by overloading of some storage nodes.

[0070] The minimum cut theorem has extensive applications in graph theory, especially in scenarios such as network traffic control and load balancing, which can help achieve reasonable resource allocation. The core idea of this theorem is to minimize the cost involved by cutting certain edges or hyperedges in the graph, thereby balancing the resource consumption among nodes while ensuring network connectivity.

[0071] In this embodiment, the minimum cut theorem is applied to the hyperedges in the high-order graph model with the aim of optimizing the load distribution between storage nodes and query requests. Specifically, by optimizing the cutting cost of hyperedges, the load of storage nodes can be evenly distributed when responding to query requests, avoiding query delays or system bottlenecks caused by uneven load.

[0072] Specifically, for each query request and storage node , the cutting cost can be defined according to their access frequencies and load conditions. Let the cutting cost represent the load relationship between storage node and query request . The cutting cost can be calculated by the following formula:

[0073] ;

[0074] where is the load weight, representing the processing capacity or request frequency of the storage node or query request. The goal of this cutting cost is to be minimized, thereby achieving load balancing.

[0075] Specifically, for each storage node , its load weight can be defined by the following formula:

[0076] ;

[0077] where Q is the set of query requests involved, is the request load of query request on storage node , and is the capacity of the storage node.

[0078] By optimizing the cutting cost, it can be ensured that query requests are evenly distributed among multiple storage nodes, avoiding some storage nodes being overloaded due to frequent queries. The specific optimization goal can be represented by the following formula:

[0079] ;

[0080] where E is the set of hyperedges, representing the dependency relationship between nodes. The realization of this optimization goal will improve the query efficiency of the system and reduce the response time.

[0081] In a possible implementation, load balancing optimization is completed through step-by-step execution. First, the relationship strength between each storage node and the query request is calculated, and the cutting cost of the hyperedge is determined. Then, based on minimizing the cutting cost, optimization is carried out using a flow algorithm in graph theory (such as the maximum flow-minimum cut algorithm) for cutting, so as to achieve balanced distribution of the load.

[0082] In graph theory, the maximum flow-minimum cut theorem is a classical theorem, which states that in a flow network, the maximum flow is equal to the capacity of the minimum cut. Applying this theorem to the present invention means minimizing the response delay of the query request by cutting the hyperedge, thereby achieving load balancing. Specifically, for each query request , the minimum cut theorem optimization can be achieved through the following process:

[0083] Calculate the flow network: Each storage node and the query request are regarded as nodes in the flow network, and the connection strength (i.e., the flow) between the nodes is defined by the cutting cost.

[0084] Maximum flow calculation: Through the maximum flow algorithm, calculate the maximum flow between the source node (representing the data generation source) and the sink node (representing the final target of data query).

[0085] Minimum cut calculation: Use the minimum cut algorithm to determine the cutting edge that divides the flow network into the source and the sink, and achieve load balancing by minimizing the cost of the cutting edge.

[0086] This method can effectively distribute query requests evenly to multiple storage nodes and avoid overloading of some nodes due to excessive query requests.

[0087] Through load balancing optimization based on the minimum cut theorem, the query efficiency and response speed of the storage system can be significantly improved. The optimized system can evenly distribute query requests among storage nodes, thereby reducing the load difference between nodes and avoiding performance bottlenecks caused by high load on some nodes.

[0088] Specifically, the performance improvement of the system can be reflected in the following aspects:

[0089] Reduced query response time: Since query requests can be evenly distributed to multiple storage nodes, the query response time is effectively controlled, avoiding delays caused by overly concentrated query requests on some storage nodes.

[0090] Improved system throughput: Through load balancing optimization, the overall throughput of the system is improved, and it can handle more query requests simultaneously.

[0091] In the case of load balancing, if a storage node fails, the system can reduce the pressure on other nodes by reallocating query requests, thereby enhancing the fault tolerance of the system.

[0092] In some embodiments, the application of the minimum cut theorem is not limited to load balancing between storage nodes and query requests, but can also be extended to optimizing resource allocation between storage nodes. By optimizing the resource usage between storage nodes, the resource utilization rate of the system can be further improved, energy consumption can be reduced, and storage efficiency can be increased.

[0093] In addition, for different types of query requests and data blocks, more constraint conditions can be introduced into the hypergraph model, such as query priorities, access frequencies of data blocks, etc., to achieve more refined load balancing. In this case, the optimization goal of the minimum cut theorem is not only to minimize query latency, but also to comprehensively consider the priorities of query requests and the access patterns of data blocks, further improving the overall performance of the system.

[0094] In summary, the present invention optimizes the load distribution between storage nodes and query requests through the minimum cut theorem in higher-order graph theory, can effectively achieve system load balancing, improve query efficiency and system throughput, and provides strong theoretical support and technical means for the design and application of distributed storage systems.

[0095] For S3, in this embodiment, the consistent hashing algorithm is used to map data blocks and storage nodes to the consistent hash ring to improve the scalability and fault tolerance of the system. Specifically, each data block and storage node is mapped to a position on the ring through the consistent hashing algorithm, and the virtual node mechanism is introduced to further reduce the need for data migration, ensuring the stability of data and the efficiency of query when the system expands or nodes change.

[0096] The consistent hashing algorithm is a distributed hashing method, aiming to efficiently map storage nodes and data blocks to a ring structure and solve the data migration problem when nodes are added or removed through this structure. In traditional hashing algorithms, when storage nodes change, a large number of data blocks need to be recalculated and remapped, causing a sharp fluctuation in system performance. While the consistent hashing algorithm maps storage nodes and data blocks to the same hash ring, ensuring a more balanced distribution of data blocks and effectively reducing data migration caused by node addition or removal.

[0097] Specifically, assume there are N storage nodes and M data blocks both of which will be mapped to a position on the consistent hash ring through a hash function. Set the hash function as where is a storage node or a data block, and the hash function Map to a point on the ring. The mapping results of each node and data block form a ring, and the mapping relationship between the storage node and the data block is determined according to its position on the hash ring.

[0098] In this mapping method, the storage node will be responsible for storing all data blocks behind it. For example, if the data block is mapped to a position on the hash ring, and the position of is behind the storage node then the data block will be stored on the node

[0099] On the basis of consistent hashing, this embodiment further introduces a virtual node mechanism, aiming to further reduce the impact of node addition and deletion on data distribution. Specifically, in traditional consistent hashing, when the number of nodes is small, some nodes may face overload problems. To alleviate this problem, virtual nodes can be introduced, that is, each physical storage node is mapped to multiple virtual nodes, so that the position of each physical node on the hash ring will be expanded, thus increasing the flexibility of the system.

[0100] The virtual nodes are introduced as follows: Suppose each physical storage node is mapped to v virtual nodes . These virtual nodes are evenly distributed on the hash ring, enhancing the load balancing of the system. When a physical node is added or removed, only the positions of the virtual nodes need to be adjusted, thus reducing large-scale migrations of data blocks.

[0101] Through consistent hashing and the virtual node mechanism, this embodiment can effectively reduce the overhead of data migration, especially when the system is expanded. Specifically, assume that at a certain moment, the number of storage nodes in the system increases from N to N + 1. Traditional hashing algorithms usually require recalculating and remapping all data blocks, resulting in a large amount of data migration. While using the consistent hashing algorithm, only a few data blocks will migrate to the new node, and these data blocks are all the data blocks before the position of the newly added node. Through the further introduction of virtual nodes, the scale of data migration is further optimized.

[0102] Specifically, assume there are N storage nodes. Through the introduction of virtual nodes, the added node will correspond to multiple virtual nodes. When a new node is added, the new virtual nodes will take over some data blocks of the original virtual nodes, while most data blocks will remain unchanged, thus greatly reducing the number of data migrations.

[0103] For example, assume that storage nodes and are mapped to their respective virtual nodes , and , . When a new node joins, its virtual node takes over a portion of the data blocks, and the distribution of other data blocks remains unaffected. Therefore, only a portion of the data blocks are migrated to the new node, thus avoiding large-scale data migration in the system.

[0104] The combination of the consistent hashing algorithm and the virtual node mechanism ensures that the storage system can smoothly perform data migration during the expansion process, avoiding a sharp fluctuation in performance when the system expands. Specifically, by reasonably designing the number and distribution of virtual nodes, it is possible to ensure that the load of data blocks is evenly distributed, and when the system expands, the amount of data block migration remains within the minimum range.

[0105] In some embodiments, the virtual node mechanism is not limited to the addition or deletion of storage nodes and can also be used to improve the fault tolerance of the system. For example, when some storage nodes fail, the availability of data can be quickly restored by adjusting the mapping of virtual nodes, reducing the system downtime. In addition, the number of virtual nodes can be dynamically adjusted according to the load situation to adapt to storage requirements of different scales and loads.

[0106] Through the consistent hashing algorithm and the virtual node mechanism, this embodiment can achieve efficient storage node mapping and data block distribution, optimize the efficiency of data access, and enhance the fault tolerance and scalability of the system. This technical solution provides reliable theoretical support for the construction of distributed storage systems and provides an effective solution for large-scale, highly available data storage systems.

[0107] Regarding S4, in this embodiment, a redundant coding method is adopted to redundantly store data blocks on multiple storage nodes to improve the fault tolerance of the system and the availability of data. Redundant coding can ensure that data is not lost when some storage nodes fail, thereby improving the reliability and stability of the system.

[0108] Redundant coding technology generates multiple copies of data through encoding to ensure that even if some copies are lost or damaged, the system can still recover the data through other copies. In distributed storage systems, redundant coding is widely used to improve the fault tolerance of the system, especially when nodes or data blocks fail.

[0109] Common redundancy encoding methods include copy encoding, RS encoding (Reed-Solomon encoding), etc. In this embodiment, Reed-Solomon encoding is adopted. This encoding method has high fault tolerance and is widely used in distributed storage environments. By splitting data blocks and encoding them into multiple redundant blocks, the risk of data loss can be effectively reduced.

[0110] Reed-Solomon encoding is an encoding method based on finite fields that can effectively correct errors or losses in data. Specifically, Reed-Solomon encoding divides the original data and generates several redundant data blocks. When a data block is lost or damaged, the system can use the remaining redundant data blocks for repair.

[0111] Suppose the data blocks are divided into k data blocks , and m redundant data blocks are generated , forming an encoded group that contains a total of k + m data blocks through Reed-Solomon encoding. These data blocks include the original data blocks and the redundant data blocks. To be able to recover the original data, when the system fails, only any number of data blocks are needed to recover all the original data.

[0112] Specifically, the mathematical formula of Reed-Solomon encoding is:

[0113] ;

[0114] where C represents the encoded data set, which contains k data blocks and m redundant blocks. The generation process of the redundant blocks can be achieved through finite field operations, and its specific method is:

[0115] ;

[0116] where is the value of the original data block, is the pre-computed encoding matrix, specifically depending on the encoding method and the selected finite field.

[0117] According to the principle of Reed-Solomon encoding, in this embodiment, by storing the redundant data blocks on multiple storage nodes, the availability of the data is ensured. When a certain storage node fails, the system can recover the data through the redundant data blocks on other nodes. Specifically, the encoded data blocks , and the redundant data blocks are allocated to multiple storage nodes. Suppose the system contains N storage nodes. Each data block is mapped to N nodes and reasonably allocated according to the load and bandwidth of the nodes.

[0118] For example, assume that storage nodes , data blocks and redundant data blocks are mapped to multiple nodes. To achieve redundant storage of data, the following mapping method can be adopted:

[0119] Allocate the data block to the node .

[0120] Allocate the redundant data block to the node .

[0121] In this way, data blocks and redundant data blocks are distributed among multiple nodes, ensuring that in case of node failure, recovery can be performed through the remaining data blocks. Generally, the number m of redundant blocks is set according to the fault tolerance requirements and failure probability of the system.

[0122] Redundant coding can also improve the scalability of the system. When the system is expanded, new storage nodes can be implemented by introducing new redundant data blocks. By reasonably selecting the number and distribution method of redundant blocks, the fault tolerance of the system can be improved without significantly increasing the amount of data migration. Specifically, when the system is expanded, redundant blocks can be allocated to the newly added nodes, and redundant blocks can be regenerated through Reed-Solomon coding.

[0123] For example, assume that the current system includes N nodes and uses m redundant blocks. When the system needs to be expanded to N +1 nodes, some redundant data blocks can be mapped to the newly added nodes. In this way, data availability during the expansion process can be ensured, and data migration and reallocation can be minimized.

[0124] In some embodiments, redundant coding is not limited to fault tolerance of storage nodes, but can also be used to optimize the efficiency of data transmission. For example, during cross-regional data transmission, redundant coding can effectively reduce the risk of data loss and reduce the delay of data recovery through redundant data blocks. In addition, by adjusting the number of redundant blocks, different loads and storage requirements can be flexibly adapted to improve the overall performance of the system.

[0125] Through the application of redundant coding, especially Reed-Solomon coding, this embodiment can achieve an efficient fault tolerance mechanism and data recovery ability in a distributed data storage system. This technical solution improves the reliability of the system and maintains high data availability in scenarios such as node failure and system expansion, and is an important technical support for building a highly fault-tolerant and highly reliable distributed storage system.

[0126] For S5, in this embodiment, it is proposed to adjust the load distribution strategy of storage nodes through a game theory model and combine it with an optimal control method for load adjustment to achieve load balancing of the distributed storage system and optimal allocation of resources. The goal of this step is to enable the system to dynamically adjust the load distribution among storage nodes according to the load status of nodes and network conditions during operation, thereby improving the overall response speed and throughput of the system and reducing the risk of overload of individual nodes.

[0127] Game theory is a method for analyzing multiple participants and can effectively describe how multiple storage nodes compete for and coordinate resources. In a distributed storage system, each storage node can be regarded as a "player", and the load distribution strategy is the choice of each storage node in the game. Specifically, the game theory model can be used to predict and optimize the strategies of each node during the resource competition process, thereby achieving a globally optimal load distribution.

[0128] In the game theory model, the goal of each storage node is to optimize its own resource allocation as much as possible to maximize its own performance. To achieve load balancing, nodes need to cooperate and play games according to the current load status and request volume at each moment. To quantify the load status of nodes, a utility function for each node can be defined to reflect its load condition. Suppose the load of node is , and its utility function can be expressed as:

[0129] ;

[0130] where is the utility function, usually a monotonically decreasing function, indicating that when the load increases, the utility of the node decreases. For example, a utility function of the following form can be selected:

[0131] ;

[0132] where and are constants that control the impact of node load on utility. Through this utility function, the system can quantify the load pressure of each node and then conduct game analysis.

[0133] In the game theory model, the load distribution decision of storage nodes not only depends on their own load status but also is affected by the load status of other nodes. The goal of each storage node is to minimize its own load while ensuring load balancing of the entire system. This can be described by the Nash equilibrium of the following game:

[0134] ;

[0135] where Represents a storage node And The cost function that affects the load between them. Generally, there is a cost for load transfer between nodes, and when the load of a certain node is too high, it may affect the data interaction efficiency with other nodes. Therefore This mutually influential relationship

[0136] To ensure that the game theory model can effectively optimize the load distribution between nodes in a dynamic environment, the optimal control theory is further introduced to optimize the load adjustment. The optimal control theory aims to adjust the control strategy according to the current system state (such as the load state of each node) and future state prediction, so that the system reaches the global optimum. The optimal control method usually optimizes based on the dynamic system model and the cost function

[0137] In this embodiment, the optimal control method can be used to dynamically adjust the load distribution of the storage nodes to achieve the load balance of the system. Assume that the load state of the system can be represented by the vector The optimal control objective is to minimize the total load of the system and at the same time consider the cost of load transfer

[0138] Specifically, the optimal control problem can be formulated as the following optimization problem

[0139] ;

[0140] Where And Are weight parameters, reflecting the cost of the load of a single node and the load transfer between nodes, and T is the optimization time range. The objective of this optimization problem is to make the load distribution of the system as balanced as possible by adjusting the node load While minimizing the load transfer cost

[0141] To solve this optimization problem, the Lagrange multiplier method can be used to introduce the constraint conditions of the system (for example, the load of a node cannot exceed its maximum bearing capacity), and the optimal strategy of load distribution can be obtained through the dynamic programming method

[0142] In the process of dynamic load adjustment, the system continuously monitors the real-time load state of each node and adjusts through the optimal control method. For example, when the load of a certain node is close to its maximum bearing capacity, the system will dynamically adjust the load distribution of other nodes according to the current load state and the request volume, so as to avoid the overload of a single node. The process of load adjustment can be achieved through the following steps

[0143] Status monitoring: The system regularly obtains the load state of each node .

[0144] Optimal control solution: Based on the load status of nodes, calculate the optimal load distribution strategy through the optimal control method .

[0145] Load distribution adjustment: According to the calculation results, adjust the load distribution between nodes to avoid excessive or insufficient load on a single node.

[0146] Feedback adjustment: After the load distribution is adjusted, continue to monitor the load status of nodes. If the load deviation is too large, re-enter the optimal control optimization process.

[0147] In some embodiments, the load distribution strategy can be adaptively optimized by combining machine learning algorithms to further improve the system's load prediction and regulation capabilities. For example, use historical load data to train a prediction model to predict the future load trend of nodes, so as to perform pre-adjustment before the load changes. In addition, by introducing more complex game models (such as multi-stage game models), the system can perform collaborative adjustments at multiple times to achieve load balancing on a long time scale.

[0148] By combining the game theory model with the optimal control method, this embodiment can achieve dynamic load balancing in a distributed storage system, ensure that the loads of each node are optimally allocated, and improve the operation efficiency and stability of the system.

[0149] Regarding S6, in this embodiment, when the system is expanded, the consistent hashing algorithm and the high-order graph model are combined to minimize data migration. The purpose of this step is to improve the efficiency and stability of the system by reducing the amount of data migration when storage nodes are expanded or reconfigured in a distributed data storage system.

[0150] In a distributed storage system, expansion usually means adding new storage nodes, which will lead to the reallocation of some data blocks. To maintain system performance and avoid the network load and latency caused by data migration, this embodiment combines the high-order graph model and the consistent hashing algorithm to optimize the data migration during system expansion, aiming to minimize the amount of data migration.

[0151] The consistent hashing algorithm is a classic method for solving the problem of data reallocation in a distributed system. Its basic idea is to map storage nodes and data blocks to a hash ring. When the system is expanded and new storage nodes are added, only those data blocks mapped to the newly added nodes need to be migrated. In this way, the amount of data migration can be minimized, full-scale reallocation can be avoided, and the efficiency of the expansion process can be improved.

[0152] In this embodiment, specifically, the new storage node will be mapped to a position on the hash ring through the consistent hashing algorithm. The mapping relationship between data blocks and storage nodes is determined by the hash function, and the new storage node is only responsible for taking over the data blocks in the adjacent area on the hash ring. This mapping method can ensure that as the system expands, only a small amount of data needs to be migrated, while most data blocks remain unchanged, thus reducing the overhead caused by data migration.

[0153] In the consistent hash ring, assuming there are N storage nodes, the data block is mapped to the hash value , if the hash value falls within the responsible interval of the node , then the data block is stored on the node . When a new node joins, it will only take over a part of the data blocks adjacent to its hash value on the hash ring. Therefore, the amount of data migration is:

[0154] ;

[0155] where N is the number of storage nodes in the system, and the original data volume is the data volume stored under the original storage nodes. In this way, the addition of a new node only needs to migrate a small part of the data, thus greatly reducing the migration overhead.

[0156] The high-order graph model is used to describe the relationship between storage nodes and data blocks and the data flow between nodes. For each storage node, the edges in the graph model represent the storage location of data blocks and the transmission of query requests. The load status of nodes, the distribution of storage blocks, and the distribution of query requests can all be modeled through the edges and nodes in the graph.

[0157] During the expansion process, the introduction of the high-order graph model can help analyze how to optimize the distribution of data blocks among storage nodes after the addition of new storage nodes. In particular, the minimum cut theorem is used to analyze the shortest cut paths between different nodes and data blocks, so as to determine which data blocks have the lowest migration cost and which data blocks should be migrated to the new node.

[0158] In practical applications, when new storage nodes are added to the system, the load situation of each storage node can be re-evaluated through the high-order graph model. By calculating the load differences of each storage node, the system can determine the optimal data block migration plan to ensure load balancing and minimize the migration cost.

[0159] Specifically, assume that the set of storage nodes in the system is , the set of data blocks is , and the storage node corresponding to each data block is , the optimal strategy for data block migration can be described by the following minimum cut theorem:

[0160] ;

[0161] where is the cost for data block to migrate from storage node to . By calculating the cut cost between each pair of nodes, the system can select the nodes with the lowest migration cost for data migration, thus ensuring load balance during the expansion process and minimizing the amount of data migration.

[0162] During the system expansion process, in addition to minimizing data migration through the consistent hashing algorithm, it is also necessary to use a high-order graph model to monitor and adjust the load after expansion in real time. Specifically, the system can dynamically adjust the load distribution of each node according to the load status of each node using a load balancing algorithm. By introducing load balancing strategies and optimizing the high-order graph model, resource contention and load imbalance problems during the expansion process can be further reduced.

[0163] For example, the system can use the shortest path algorithm and minimum cut theorem in graph theory to analyze the distribution of data blocks and query requests in real time, calculate the most suitable data migration plan, and ensure the stability and efficiency of data transmission during the expansion process.

[0164] In some embodiments, in addition to combining the consistent hashing algorithm and the high-order graph model, a virtual node mechanism can be introduced to further optimize the data migration process. The introduction of virtual nodes can distribute the load of actual storage nodes to multiple virtual nodes, thereby balancing the load between nodes and preventing a certain node from becoming a bottleneck.

[0165] In addition, as the system continues to expand, the storage scale and complexity of load distribution of the system may increase. Therefore, using machine learning algorithms for load prediction and optimization may further improve the efficiency and accuracy during the expansion process. For example, by training a model to predict the system load status and performing load distribution and data migration in advance, the system pressure caused by sudden loads can be reduced.

[0166] Through the above technical solutions, this embodiment can combine the consistent hashing algorithm and the high-order graph model during the expansion process of the distributed data storage system to minimize data migration, improve the scalability and stability of the system, and thus achieve efficient utilization of resources and optimization of system performance.

[0167] Generally speaking, the present invention provides a method for constructing a data index for distributed data storage. By combining advanced technologies such as high-order graph theory, consistent hashing algorithm, redundant coding, and game theory, the allocation of storage nodes and data blocks, load balancing, and data migration process are optimized. Specifically, the load of storage nodes and query requests is optimized by constructing a hypergraph model and the minimum cut theorem, the data migration is reduced and the scalability is improved by using the consistent hashing algorithm, the data security is ensured by adopting redundant coding, and the load allocation is optimized by combining game theory and optimal control theory. This method can not only effectively handle the data distribution and query efficiency problems in a distributed storage system, but also minimize data migration during the system expansion process, improving the stability and performance of the system.

[0168] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a data index for distributed data storage, characterized in that: The following steps are involved: A hypergraph model of a distributed storage system is constructed by high-order graph theory, wherein the hypergraph model includes multiple storage nodes, multiple data blocks, and multiple query requests; Load balancing optimization of storage nodes and query requests based on the minimum cut theorem in the high-order graph model; Map data blocks and storage nodes to the consistent hash ring according to the consistent hash algorithm, and reduce data migration through the virtual node mechanism; A redundant encoding method is used to redundantly store data blocks on multiple storage nodes; The load distribution strategy of storage nodes is adjusted through game theory model, and the load adjustment is optimized by using optimal control method; When the system is expanded, the consistent hashing algorithm and high-order graph model are combined to minimize data migration; The optimal control method includes the following steps: Define the system load state variable x i (t), represents the load status of each storage node at time t. Assuming that the system contains N storage nodes, the load status variable x i (t) represents the load status of node i at time t. The update of load status depends on the number of requests processed by the node, the amount of data stored, and the query load factor, specifically: x i (t)=f(N i ,Q i ,D i ); Among them, N i is the computing power of the storage node, Q i is the query request volume of node i, D i The amount of data stored for node i; Set the system's load optimization objective function J, which describes the load distribution of all nodes and minimizes the load imbalance of the entire system by adjusting the node load. The load optimization objective function J represents the cumulative load cost within the time window T, specifically: Among them, the integral symbol is the accumulation of the entire system load cost in the time interval [0, T], C(x i (t))dt is the load cost function of the node, reflecting the increase in cost with higher node load, and T is the optimization time window; Solve the optimal control strategy u for load distribution through optimal control theory i (t),u i (t) is the load adjustment of storage node i at time t. The control goal is to dynamically adjust u i (t), so that the system load is minimized within a given time window and the load balancing constraints are met: The constraints include: and Among them, C total is the total system load, which represents the sum of the loads of all storage nodes, L max is the load limit of a single node, and the constraint x i (t)≤L max The load of each storage node i must be less than or equal to the maximum load limit L max , Indicates that the constraint condition applies to every storage node i; Based on the Lagrange multiplier method or the variational method, by solving the optimal control problem in the above formula, the optimal load adjustment strategy u of each storage node i is obtained: i (t) Make the system achieve load balancing on a global scale and minimize the overall load cost; According to the solved optimal control strategy u i (t) Dynamically adjust the load distribution of each storage node so that each storage node performs corresponding load adjustment under the constraints of load balancing and performance optimization, thereby achieving the global optimal load distribution of the system.

2. The method for constructing a data index for distributed data storage according to claim 1, characterized in that: Each hyperedge in the high-order graph model represents the dependency relationship between multiple data blocks and multiple storage nodes or query requests, and the cutting cost of the hyperedge is optimized through the minimum cut theorem to ensure the shortest query path and achieve load balancing.

3. The method for constructing a data index for distributed data storage according to claim 1, characterized in that: In the consistent hashing algorithm, storage nodes and data blocks are mapped to a consistent hashing ring through hash values, each storage node is mapped to the hashing ring through multiple virtual nodes, and the node load is optimized through the virtual node mechanism.

4. The method for constructing a data index for distributed data storage according to claim 1, characterized in that: The redundant coding adopts Reed-Solomon coding, stores each data block in multiple storage nodes, and further reduces the data storage overhead through data compression method.

5. The method for constructing a data index for distributed data storage according to claim 1, characterized in that: The game theory model adjusts the load distribution strategy of the storage nodes through the Nash equilibrium principle, so that each storage node selects the optimal load distribution solution while considering the loads of other nodes.

6. The method for constructing a data index for distributed data storage according to claim 1, characterized in that: When the system is expanded, the consistent hashing algorithm and the high-order graph model are combined to minimize data migration, including the following steps: The data mapping on the hash ring is adjusted through the consistent hashing algorithm, combined with the load balancing strategy of the high-order graph model to minimize the amount of data migration between nodes during the expansion process.

7. The method for constructing a data index for distributed data storage according to claim 6, characterized in that: The load balancing optimization adopts a graph cutting and reorganization strategy based on high-order graph theory. When the node load is uneven, the load distribution is optimized by adjusting the dependency between the data blocks and the storage nodes.

Citation Information

Patent Citations

  • A cloud data processing method and system based on workload

    CN108984308A

  • Power grid load frequency control optimization method, system and device and storage medium

    CN113270872A

  • Distributed file storage method, system and device

    CN117633895A