Data storage method and device, equipment and storage medium

By using jump consistency hashing algorithm and virtual node technology in distributed storage systems, dynamically allocating identifiers of storage nodes and data objects, the problems of data storage and load balancing in distributed environments are solved, and efficient data distribution and load balancing are achieved.

CN119987680APending Publication Date: 2025-05-13SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510189050.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In a dynamically changing distributed environment, how to store data and maintain good load balancing to avoid frequent data migration due to changes in the number of nodes.

Method used

By assigning a first identifier to the nodes in the acquired current storage node list, assigning a second identifier to the target data object, the index of the data object and the target node is determined using a standard hash function and a jump consistency hash algorithm, and selecting a suitable node for storage according to the load situation.

Benefits of technology

It realizes uniform distribution and load balancing of data in a dynamically changing distributed environment, reduces data migration overhead, improves the overall performance and scalability of the system, and solves the problems of load imbalance, data hot spots and limited scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987680A_ABST
    Figure CN119987680A_ABST
Patent Text Reader

Abstract

The invention discloses a data storage method and device, equipment and a storage medium, and relates to the technical field of distributed storage systems.The data storage method comprises the steps that corresponding first identifiers are allocated to nodes in an obtained current storage node list, and corresponding second identifiers are allocated to target data objects; determining a target integer hash value corresponding to the second identifier of the target data object by using a standard hash function, and inputting the target integer hash value and the current number of nodes in the current storage node list into a jump consistency hash algorithm to obtain a target index of the target data object and a target node in the current storage node list; and storing the target data object to the corresponding target node based on the obtained target index, the first identifier corresponding to the target node and the target integer hash value corresponding to the target data object. In this way, data storage can be carried out in a dynamically changing distributed environment, and good load balance is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed storage systems, and in particular to a data storage method, device, equipment and storage medium. Background Art

[0002] With the rapid development of big data and cloud computing technologies, distributed storage systems play an increasingly important role in modern information technology infrastructure. Compared with traditional centralized storage systems, distributed storage systems provide higher scalability, reliability, and availability by distributing data across multiple physical nodes. The core idea of ​​this system is to achieve efficient storage and management of large-scale data through the collaborative work of multiple independent nodes.

[0003] In a distributed storage system, how to reasonably place data on different storage nodes has become a key issue. This process not only affects the access efficiency of data, but is also directly related to the load balancing and fault tolerance of the system. Traditional centralized storage systems usually manage the storage location of data through fixed rules or centralized control, but this approach often creates performance bottlenecks when facing large-scale data and high-concurrency access, and cannot meet the needs of modern distributed systems.

[0004] Currently, existing distributed storage systems usually use DHT (Distributed Hash Table) technology or consistent hashing algorithms based on hash functions. These algorithms map the identifiers of data and storage nodes to a logical ring, so that data can be evenly distributed among nodes and minimize the scope of data migration when nodes change.

[0005] In summary, how to store data in a dynamically changing distributed environment and maintain good load balancing is a problem that needs to be solved urgently. Summary of the invention

[0006] In view of this, the purpose of the present invention is to provide a data storage method, device, equipment and storage medium, which can store data in a dynamically changing distributed environment and maintain good load balancing. The specific scheme is as follows:

[0007] In a first aspect, the present application provides a data storage method, comprising:

[0008] Assigning a corresponding first identifier to a node in the acquired current storage node list, and assigning a corresponding second identifier to a target data object;

[0009] Determine a target integer hash value corresponding to the second identifier of the target data object using a standard hash function, and input the target integer hash value and the current number of nodes in the current storage node list into a jump consistent hashing algorithm to obtain a target index of the target data object and the target node in the current storage node list;

[0010] The target data object is stored in a corresponding target node based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

[0011] Optionally, before storing the target data object to the corresponding target node based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object, the method further includes:

[0012] Detecting the storage capacity and I / O load of the target node;

[0013] Based on the obtained detection result, a target node that meets the preset light load condition is determined, so that the target data object is stored in the target node that meets the preset light load condition based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

[0014] Optionally, determining the target node that meets a preset light load condition based on the obtained detection result includes:

[0015] Determining whether the storage capacity of the target node is within a first preset light load range and whether the I / O load of the target node is within a second preset light load range;

[0016] If the storage capacity of the target node is within a first preset light load range and the I / O load of the target node is within a second preset light load range, it is determined that the target node meets a preset light load condition.

[0017] Optionally, the method further includes:

[0018] Monitoring the current storage node list to obtain corresponding monitoring results;

[0019] If the monitoring result indicates that a new node has been added, then the corresponding new node is added to the current storage node list and the corresponding first identifier is assigned to obtain a new current storage node list;

[0020] If the monitoring result obtained indicates that an existing node has exited, the corresponding existing node and the corresponding first identifier are deleted from the current storage node list to obtain a new current storage node list, and the data objects that meet the preset migration conditions are determined from the exited existing nodes according to the jump consistency hash algorithm, and the preset migration operation is performed on the data objects that meet the preset migration conditions.

[0021] Optionally, the method further includes:

[0022] Monitor the load of each node in the current storage node list;

[0023] If the monitoring result obtained indicates that the load index corresponding to the node in the current storage node list exceeds the preset load index range, a preset data storage adjustment operation is performed on the node exceeding the preset load index range;

[0024] The load condition includes any one or more of the CPU operation condition, memory usage condition and I / O read / write condition.

[0025] Optionally, performing a preset adjustment data storage operation on the nodes exceeding the preset load index range includes:

[0026] Acquire partial data from nodes exceeding the preset load index range;

[0027] The jump consistent hashing algorithm is used to distribute the acquired partial data of the nodes exceeding the preset load index range to the nodes satisfying the preset light load condition.

[0028] In a second aspect, the present application provides a data storage device, comprising:

[0029] An identifier allocation module, used to allocate a corresponding first identifier to a node in the acquired current storage node list, and allocate a corresponding second identifier to a target data object;

[0030] an index determination module, configured to determine a target integer hash value corresponding to the second identifier of the target data object using a standard hash function, and input the target integer hash value and a current number of nodes in the current storage node list into a jump consistent hashing algorithm to obtain a target index of the target data object and a target node in the current storage node list;

[0031] A data storage module is used to store the target data object to a corresponding target node based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

[0032] Optionally, the device further includes:

[0033] a load detection module, configured to detect the storage capacity and I / O load of the target node to determine whether the storage capacity of the target node is within a first preset light load range and whether the I / O load of the target node is within a second preset light load range;

[0034] A node determination module is used to determine whether the target node satisfies a preset light load condition if the storage capacity of the target node is within a first preset light load range and the I / O load of the target node is within a second preset light load range, so as to store the target data object to the target node that satisfies the preset light load condition based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

[0035] In a third aspect, the present application provides an electronic device, including:

[0036] Memory, used to store computer programs;

[0037] The processor is used to execute the computer program to implement the aforementioned data storage method.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the data storage method as described above is implemented.

[0039] In summary, the present application first assigns a corresponding first identifier to the node in the current storage node list obtained, and assigns a corresponding second identifier to the target data object; uses a standard hash function to determine the target integer hash value corresponding to the second identifier of the target data object, and inputs the target integer hash value and the current number of nodes in the current storage node list into the jump consistency hash algorithm to obtain the target index of the target data object and the target node in the current storage node list; based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object, the target data object is stored in the corresponding target node. As can be seen from the above, the present application first assigns a corresponding first identifier to the current node in the current storage node list obtained, and assigns a corresponding second identifier to the target data object. Then, the target integer hash value corresponding to the second identifier is calculated using a standard hash function, and then the target integer hash value and the current number of nodes in the current storage node list are used as inputs and input into the jump consistency hash algorithm to obtain the target index of the target data object and the target node. Finally, according to the target index obtained, the target data object is stored on the corresponding target node, thereby completing the placement operation of the target data object. In this way, the present application can evenly distribute data on multiple storage nodes by introducing an improved consistent hashing algorithm and virtual node technology, avoiding frequent data migration caused by changes in the number of nodes, and can maintain good load balancing in a dynamically changing distributed environment. At the same time, it can reduce data migration overhead, improve the overall performance and scalability of the system, and solve the problems of load imbalance, data hot spots, and limited scalability existing in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0041] Figure 1 A flow chart of a data storage method disclosed in this application;

[0042] Figure 2 A flow chart of a specific data storage method disclosed in this application;

[0043] Figure 3 A flow chart of a specific data storage method disclosed in this application;

[0044] Figure 4A time complexity comparison diagram of an algorithm disclosed in this application;

[0045] Figure 5 A flow chart of a specific data storage method disclosed in this application;

[0046] Figure 6 A schematic diagram of the structure of a data storage device disclosed in this application;

[0047] Figure 7 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] At present, existing distributed storage systems usually adopt DHT technology or consistent hashing algorithm based on hash function. These algorithms map the identifiers of data and storage nodes to a logical ring, so that data can be evenly distributed among nodes and minimize the scope of data migration when nodes change. In order to solve the above technical problems, the present application discloses a data storage method, device, equipment and storage medium, which can store data in a dynamically changing distributed environment and maintain good load balancing.

[0050] See also Figure 1 As shown, an embodiment of the present invention discloses a data storage method, which may include:

[0051] Step S11: assign a corresponding first identifier to a node in the acquired current storage node list, and assign a corresponding second identifier to a target data object.

[0052] In this embodiment, the current storage node list is first obtained, and a first identifier (such as an IP address or a node ID) is assigned to each node. The first identifier of this embodiment is used to identify each node and serves as the basis for subsequent data distribution and search operations. At the same time, all nodes are stored in an ordered list, namely the current storage node list, and the nodes are sorted according to the identifiers in the current storage node list to facilitate subsequent fast search and insertion operations. When constructing the current storage node list, the status information of each node is checked, including the online status, storage capacity and load of the node, to ensure that only available nodes are included in the list.

[0053] In addition, it is also necessary to assign a corresponding second identifier to each target data object that needs to be stored in the node based on the metadata of the target data object (such as file name, size, creation time, etc.). The second identifier can be an IP address or ID, etc.

[0054] Step S12: Use a standard hash function to determine the target integer hash value corresponding to the second identifier of the target data object, and input the target integer hash value and the current number of nodes in the current storage node list into a jump consistent hashing algorithm to obtain the target index of the target data object and the target node in the current storage node list.

[0055] In this embodiment, after a unique second identifier is assigned to each target data object, a corresponding integer hash value is calculated for the second identifier of each target data object using a standard hash function, wherein the standard hash function may be one or more of MurmurHash or CRC32 (A cyclic redundancy check 32).

[0056] The calculated hash value and the current number of nodes are input into the jump consistency hashing algorithm, and multiple "jump" calculations are performed to finally determine the target node where the corresponding target data object should be stored, that is, the storage location for the target data object. Specifically, first determine the first variable j and the second variable k, initialize variable j to 0, and initialize variable k to -1. These two variables are used to control the calculation process in the loop. The loop executes the following steps until the value of k is equal to the number of nodes: Update the value of j to k; Calculate the value corresponding to k according to the formula:

[0057] ;

[0058] Among them, rand() is a pseudo-random number generation function; MAX_RAND is the maximum value of the pseudo-random number generator; k is the second variable; j is the second variable; floor() is a floor function. When the loop ends, the value of k is the target index of the target node where the data should be stored. Through the above method, the target node can be selected in a random manner, thereby achieving uniform distribution of data and load balancing.

[0059] Step S13: store the target data object to the corresponding target node based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

[0060] In this embodiment, according to the target index, the corresponding target node is searched from the storage node list, and the first identifier corresponding to the target node is obtained. The first identifier is used to map the corresponding target node and serve as the target location for data storage. Then, based on the target integer hash value corresponding to the target data object, a storage path or storage block identifier of the target data object is generated. The target data object is transmitted to the target node, and is written to the storage medium of the target node according to the generated storage path or storage block identifier. After the data storage is completed, the mapping relationship between the target data object and the target node is recorded, for example, the identifier of the target data object, the first identifier of the target node, and the storage path and other information, so as to facilitate subsequent rapid retrieval and access.

[0061] As can be seen from the above, the embodiment of the present application will first assign a corresponding first identifier to the current node in the current storage node list obtained, and assign a corresponding second identifier to the target data object. Then, the target integer hash value corresponding to the second identifier is calculated using a standard hash function, and then this target integer hash value and the current number of nodes in the current storage node list are used as inputs and input into the jump consistency hash algorithm to obtain the target index of the target data object and the target node. Finally, according to the target index obtained, the target data object is stored on the corresponding target node, thereby completing the placement operation of the target data object. In this way, the embodiment of the present application can evenly distribute data on multiple storage nodes by introducing an improved consistency hash algorithm and virtual node technology, avoiding frequent data migration caused by changes in the number of nodes, and maintaining good load balancing in a dynamically changing distributed environment. At the same time, it reduces data migration overhead, improves the overall performance and scalability of the system, and solves the problems of load imbalance, data hot spots, and limited scalability in the prior art.

[0062] See also Figure 2 As shown, in order to achieve load balancing when storing target data objects, an embodiment of the present invention discloses a specific data storage method, which may include:

[0063] Step S21: Determine whether the storage capacity of the target node is within a first preset light load range and whether the I / O load of the target node is within a second preset light load range.

[0064] In this embodiment, after determining the target node, the storage capacity and the corresponding I / O load of the target node are found in the current storage node list, and then the storage capacity of the target node is compared with a pre-set first preset light load range, and then the I / O load of the target node is compared with a pre-set second preset light load range to obtain a corresponding comparison result.

[0065] Step S22: If the storage capacity of the target node is within a first preset light load range and the I / O load of the target node is within a second preset light load range, it is determined that the target node meets a preset light load condition.

[0066] In this embodiment, if the comparison result obtained indicates that the storage capacity of the target node is within a first preset light load range and the I / O load of the target node is within a second preset light load range, then the target node whose storage capacity is within the first preset light load range and whose I / O load is within the second preset light load range is determined to meet the preset light load condition, and the target data object can be stored in the corresponding target node.

[0067] Step S23: store the target data object to a target node that meets a preset light load condition based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

[0068] Among them, the specific implementation process of step S23 can refer to the corresponding content disclosed in the above-mentioned embodiment, and will not be repeated here.

[0069] The embodiment of the present application determines whether the storage capacity of the target node is within the first preset light load range and whether the I / O load of the target node is within the second preset light load range, obtains a corresponding comparison result, and determines the target node that meets the preset light load condition according to the comparison result. In this way, the embodiment of the present application can achieve global load balancing and avoid the problem of overloading a single node.

[0070] See also Figure 3 As shown, in order to reduce migration overhead and improve the stability of data storage when the number of nodes changes, an embodiment of the present invention discloses a specific data storage method, which may include:

[0071] Step S31: monitor the current storage node list to obtain corresponding monitoring results.

[0072] Step S32: If the monitoring result indicates that a new node has been added, then the corresponding new node is added to the current storage node list and a corresponding first identifier is allocated to obtain a new current storage node list.

[0073] In this embodiment, after obtaining the corresponding monitoring results, if the monitoring results indicate that a new node has been added, the new node is added to the current storage node list, and a unique first identifier is assigned to the new node. Then, the current storage node list is updated to obtain a new current storage node list.

[0074] Step S33: If the monitoring result obtained indicates that an existing node has exited, the corresponding existing node and the corresponding first identifier are deleted from the current storage node list to obtain a new current storage node list, and the data objects that meet the preset migration conditions are determined from the exited existing nodes according to the jump consistency hash algorithm, and the preset migration operation is performed on the data objects that meet the preset migration conditions.

[0075] In this embodiment, after obtaining the corresponding monitoring results, if the obtained monitoring results indicate that there is an existing node exit, the corresponding existing node to be deleted and the corresponding first identifier are deleted from the current storage node list, and then the current storage node list is updated to obtain a new current storage node list, and then, according to the jump consistency hashing algorithm, the data objects in the existing nodes to be deleted are subjected to the preset migration operation.

[0076] It is important to understand that the jump consistent hashing algorithm uses random numbers to determine whether the data object key should jump to a new node. During the operation of the algorithm, every time a new node is added, the probability of the key moving to the new node is set to 1 / (k+1), in this way to achieve the purpose of load balancing.

[0077] Assume that the consistent hashing algorithm is used to map data objects to nodes. When the total number of nodes n is 1, the key will be assigned to the 0th node, and the probability of being assigned to this node is 1. When a new node is added, so that the total number of nodes n is 2, the probability of the key being assigned to the new node (that is, the first node) is 1 / 2, and the probability of remaining in the original node is also 1 / 2. When another node is added, and the total number of nodes becomes n = 3, the probability of the key being assigned to the new node (the second node) is 1 / 3, and the probability of remaining in the original node (the 0th or first node) is 1 / 3. Further general deductions are made from these basic situations. When the total number of nodes is n = k, the probability of the key being assigned to each node is 1 / n. When a new node is added, the total number of nodes n is k+1, the probability of the key being assigned to the new node (kth node) is 1 / (k+1), and the probability of the key being retained in the original node (0th to k-1th nodes) is 1 / k×k / (k+1)=1 / (k+1). This shows that no matter how the number of nodes changes, the probability of the key being assigned to each node can always remain 1 / n. However, in this process, the time complexity of the algorithm is O(N).

[0078] Furthermore, the consistent hashing algorithm is optimized to obtain an optimized jump consistent hashing algorithm. For example, in the process of node n changing from 1 to 5, the data objects key1 and key2 need to be compared with the target distribution 1 / n based on the value of the random sequence to decide whether to stay in the original node or jump to a new node. Assume that there are currently B+1 nodes. When one more node is added and the total number of nodes becomes B+2, the probability that the key remains in the original node position can be derived by the formula. If the key is assigned to the latest node when it reaches J+1 nodes, then the following formula must be satisfied:

[0079] ;

[0080] in, is the probability that the data object is assigned to the latest node when it reaches J+1 nodes; B is the number of nodes. Therefore, the probability that k does not jump nodes continuously until it reaches J+1 nodes is (B+1) / J. Figure 4 As shown, the time complexity of the optimized jump consistent hashing algorithm is O(logN).

[0081] In this embodiment, if an existing node exits or a new node joins, the current storage node list needs to be updated to obtain a new current storage node list, and the preset migration operation is performed on the data object according to the situation. In addition, the time complexity and computational complexity are reduced by optimizing the jump consistency hashing algorithm. In this way, the scalability of data storage is enhanced. Whether adding a new node or deleting a node, only a small amount of data needs to be redistributed, which meets the needs of large-scale data environments.

[0082] See also Figure 5 As shown, in order to avoid the problem of single node overload, the embodiment of the present invention discloses a specific data storage method, which may include:

[0083] Step S41: monitor the load status of each node in the current storage node list.

[0084] Step S42: If the obtained monitoring result indicates that the load index corresponding to the node in the current storage node list exceeds a preset load index range, a preset data storage adjustment operation is performed on the node exceeding the preset load index range.

[0085] In this embodiment, after real-time monitoring and analysis of the load conditions of each node in the current storage node list, when the monitoring results indicate that the load index corresponding to some nodes in the current storage node list has exceeded the pre-set load index range, this means that these nodes are under excessive working pressure, which will have an adverse effect on the stability and performance of the entire data storage system. In order to ensure smooth data storage and processing, it is necessary to perform a preset data storage adjustment operation on the nodes that exceed the preset load index range. The load conditions include but are not limited to any one or more of the CPU operation conditions, memory usage conditions, and I / O read and write conditions.

[0086] In this embodiment, if it is necessary to perform a preset data storage adjustment operation on nodes that exceed the preset load index range, first, some data from the nodes that exceed the preset load index range can be obtained; and the jump consistency hash algorithm is used to distribute the obtained partial data of the nodes that exceed the preset load index range to the nodes that meet the preset light load condition. Specifically, after determining that the load index of the node exceeds the preset load index range, extract some data from the node that exceeds the preset load index range, and then use the jump consistency hash algorithm to reasonably distribute these data to the nodes that meet the preset light load condition based on the real-time load conditions of each node, so as to balance the overall load of the system and ensure stable and efficient operation of the system.

[0087] The embodiment of the present application migrates part of the data in the overloaded nodes to the lightly loaded nodes, so as to reasonably disperse and optimize the overload on these nodes, so that the load of each node can return to the preset reasonable range, thereby maintaining the balance and stable operation state of data storage, improving the overall performance and reliability of data storage, and avoiding the problem of overloading a single node.

[0088] See also Figure 6 As shown, an embodiment of the present invention discloses a data storage device, including:

[0089] The identifier allocation module 11 is used to allocate a corresponding first identifier to a node in the acquired current storage node list, and allocate a corresponding second identifier to a target data object;

[0090] An index determination module 12 is used to determine a target integer hash value corresponding to the second identifier of the target data object using a standard hash function, and input the target integer hash value and the current number of nodes in the current storage node list into a jump consistent hashing algorithm to obtain a target index of the target data object and the target node in the current storage node list;

[0091] The data storage module 13 is used to store the target data object to the corresponding target node based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

[0092] As can be seen from the above, the present application will first assign a corresponding first identifier to the current node in the current storage node list obtained, and assign a corresponding second identifier to the target data object. Then, the target integer hash value corresponding to the second identifier is calculated using a standard hash function, and then this target integer hash value and the current number of nodes in the current storage node list are used as inputs to the jump consistency hash algorithm to obtain the target index of the target data object and the target node. Finally, according to the target index obtained, the target data object is stored on the corresponding target node, thereby completing the placement operation of the target data object. In this way, the present application can evenly distribute data on multiple storage nodes by introducing an improved consistency hash algorithm and virtual node technology, avoiding frequent data migration caused by changes in the number of nodes, and maintaining good load balancing in a dynamically changing distributed environment. At the same time, it reduces data migration overhead, improves the overall performance and scalability of the system, and solves the problems of load imbalance, data hot spots, and limited scalability in the prior art.

[0093] In some specific embodiments, the data storage device may further include:

[0094] a load detection module, configured to detect the storage capacity and I / O load of the target node to determine whether the storage capacity of the target node is within a first preset light load range and whether the I / O load of the target node is within a second preset light load range;

[0095] A node determination module is used to determine whether the target node satisfies a preset light load condition if the storage capacity of the target node is within a first preset light load range and the I / O load of the target node is within a second preset light load range, so as to store the target data object to the target node that satisfies the preset light load condition based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

[0096] In some specific embodiments, the data storage device further includes:

[0097] A list monitoring module, used to monitor the current storage node list to obtain corresponding monitoring results;

[0098] A first list updating module, configured to add a corresponding new node to the current storage node list and assign a corresponding first identifier to obtain a new current storage node list if the monitoring result indicates that a new node has been added;

[0099] The second list updating module is used to delete the corresponding existing node and the corresponding first identifier from the current storage node list if the monitoring result indicates that an existing node has exited, so as to obtain a new current storage node list, and determine the data objects that meet the preset migration conditions from the exited existing nodes according to the jump consistency hashing algorithm, and perform the preset migration operation on the data objects that meet the preset migration conditions.

[0100] In some specific embodiments, the data storage device may further include:

[0101] A node monitoring module is used to monitor the load of each node in the current storage node list;

[0102] An operation execution module is used to execute a preset data storage adjustment operation on the nodes exceeding the preset load index range if the monitoring result obtained indicates that the load index corresponding to the node in the current storage node list exceeds the preset load index range; wherein the load condition includes any one or more of the CPU operation condition, memory usage condition and I / O read and write condition.

[0103] In some specific embodiments, the data storage device may further include:

[0104] A data acquisition module, used to acquire partial data from nodes exceeding the preset load index range;

[0105] The data distribution module is used to distribute the part of the data of the node exceeding the preset load index range to the node meeting the preset light load condition by using the jump consistent hashing algorithm.

[0106] Furthermore, the present application also discloses an electronic device. Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0107] Figure 7A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the data storage method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0108] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0109] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0110] The operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to computer programs that can be used to complete the data storage method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0111] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the aforementioned disclosed data storage method. The specific steps of the method can refer to the corresponding contents disclosed in the aforementioned embodiments, and will not be repeated here.

[0112] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0113] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0114] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0115] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0116] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A data storage method, characterized in that: include: Assigning a corresponding first identifier to a node in the acquired current storage node list, and assigning a corresponding second identifier to a target data object; Determine a target integer hash value corresponding to the second identifier of the target data object using a standard hash function, and input the target integer hash value and the current number of nodes in the current storage node list into a jump consistent hashing algorithm to obtain a target index of the target data object and the target node in the current storage node list; The target data object is stored in a corresponding target node based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

2. The data storage method according to claim 1, characterized in that: Before storing the target data object to the corresponding target node based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object, the method further includes: Detecting the storage capacity and I / O load of the target node; Based on the obtained detection result, a target node that meets the preset light load condition is determined, so that the target data object is stored in the target node that meets the preset light load condition based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

3. The data storage method according to claim 2, characterized in that: The step of determining a target node that meets a preset light load condition based on the obtained detection result includes: Determining whether the storage capacity of the target node is within a first preset light load range and whether the I / O load of the target node is within a second preset light load range; If the storage capacity of the target node is within a first preset light load range and the I / O load of the target node is within a second preset light load range, it is determined that the target node meets a preset light load condition.

4. The data storage method according to claim 2 or 3, characterized in that: Also includes: Monitoring the current storage node list to obtain corresponding monitoring results; If the monitoring result indicates that a new node has been added, then the corresponding new node is added to the current storage node list and the corresponding first identifier is assigned to obtain a new current storage node list; If the monitoring result obtained indicates that an existing node has exited, the corresponding existing node and the corresponding first identifier are deleted from the current storage node list to obtain a new current storage node list, and the data objects that meet the preset migration conditions are determined from the exited existing nodes according to the jump consistency hash algorithm, and the preset migration operation is performed on the data objects that meet the preset migration conditions.

5. The data storage method according to claim 4, characterized in that: Also includes: Monitor the load of each node in the current storage node list; If the monitoring result obtained indicates that the load index corresponding to the node in the current storage node list exceeds the preset load index range, a preset data storage adjustment operation is performed on the node exceeding the preset load index range; The load condition includes any one or more of the CPU operation condition, memory usage condition and I / O read / write condition.

6. The data storage method according to claim 5, characterized in that: The performing a preset adjustment data storage operation on the nodes exceeding the preset load index range includes: Acquire partial data from nodes exceeding the preset load index range; The jump consistent hashing algorithm is used to distribute the acquired partial data of the nodes exceeding the preset load index range to the nodes satisfying the preset light load condition.

7. A data storage device, characterized in that: include: An identifier allocation module, used to allocate a corresponding first identifier to a node in the acquired current storage node list, and allocate a corresponding second identifier to a target data object; an index determination module, configured to determine a target integer hash value corresponding to the second identifier of the target data object using a standard hash function, and input the target integer hash value and a current number of nodes in the current storage node list into a jump consistent hashing algorithm to obtain a target index of the target data object and a target node in the current storage node list; A data storage module is used to store the target data object to a corresponding target node based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

8. The data storage device according to claim 7, characterized in that: Also includes: a load detection module, configured to detect the storage capacity and I / O load of the target node to determine whether the storage capacity of the target node is within a first preset light load range and whether the I / O load of the target node is within a second preset light load range; A node determination module is used to determine whether the target node satisfies a preset light load condition if the storage capacity of the target node is within a first preset light load range and the I / O load of the target node is within a second preset light load range, so as to store the target data object to the target node that satisfies the preset light load condition based on the obtained target index, the first identifier corresponding to the target node, and the target integer hash value corresponding to the target data object.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the data storage method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the data storage method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Session sharing method and device, storage medium and computer equipment

    CN120915830A