A Metadata Management Method, Apparatus, Electronic Device, and Storage Medium

By determining the target value and migration strategy of the storage node in a distributed storage system, combined with the number of storage groups that have been carried, a higher load balancing is achieved, which solves the problem of insufficient load balancing in the existing technology, and improves the performance and data storage reliability of the storage nodes.

CN115543200BActive Publication Date: 2025-07-22HANGZHOU HIKVISION SYST TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211237230.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2025-07-22
Estimated Expiration
2042-10-08

AI Technical Summary

Technical Problem

In existing distributed storage systems, when load balancing is achieved through a consistent hash algorithm, the load balancing degree is limited, resulting in a degradation in the performance of the storage node.

Method used

By determining the target value of the storage node, comprehensively considering the number of storage groups that have been carried and the number of storage groups that each storage node estimates the number of storage groups that each storage node carries during load distribution balance, and using the estimated migration steps and mapping modules to migrate the target storage group to the target storage node to improve load balancing.

Benefits of technology

The load distribution balance of storage nodes in distributed storage systems is improved, and the problem of the remaining capacity of some storage nodes is too small or the proportion is too small, which improves the reliability and stability of data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115543200B_ABST
    Figure CN115543200B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a metadata management method, apparatus, electronic device, and storage medium, which relate to the technical field of computer networks and can improve the load distribution balance of multiple storage nodes in a distributed storage system. The method includes: determining a target value for each storage node among the multiple storage nodes, where the target value of the storage node is used to represent the number of storage groups that the storage node is estimated to carry when the load distribution of the multiple storage nodes is balanced, and the storage group includes at least one set of metadata; based on the target value of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node, determining N target storage nodes mapped by a target storage group from the multiple storage nodes, where N is a positive integer, and the target storage node is used to store a copy of the target storage group; the target storage group is the storage group to be migrated; and migrating the target storage group to the N target storage nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer networks, and in particular, to a metadata management method, apparatus, electronic device, and storage medium. Background Art

[0002] With the further development of information technology and the large-scale application of networks, there has been an explosive growth of data, which has also brought great development to network storage. Among them, in a distributed storage system, as the amount of system data storage continues to increase, the amount of file metadata borne by the metadata server also grows rapidly. Therefore, it is necessary to evenly disperse the metadata to each storage node in the system as much as possible, that is, to ensure the load balance of each storage node to avoid the performance degradation of nodes caused by uneven load.

[0003] In related technologies, load balancing is achieved through the consistent hashing algorithm. However, this consistent hashing algorithm only achieves balance probabilistically, so the degree of load balance is limited. Summary of the Invention

[0004] Embodiments of this application provide a metadata management method, apparatus, electronic device, and storage medium, which improve the degree of load distribution balance of multiple storage nodes in a distributed storage system.

[0005] In a first aspect, embodiments of this application provide a metadata management method, which is applied to a primary storage node in a distributed storage system including multiple storage nodes. The method includes: determining a target value for each storage node among the multiple storage nodes; the target value of the storage node is used to represent the number of storage groups that the storage node is estimated to bear when the load distribution of the multiple storage nodes is balanced, and the storage group includes at least one set of metadata; based on the target values of each storage node among the multiple storage nodes and the number of storage groups already borne by each storage node, determining N target storage nodes mapped by a target storage group from the multiple storage nodes; N is a positive integer, and the target storage node is used to store a copy of the target storage group; the target storage group is a storage group to be migrated; migrating the target storage group to the N target storage nodes.

[0006] The technical solutions provided by the embodiments of this application at least bring the following beneficial effects: Compared with the related technologies that only rely on probability for load balancing when implementing load balancing based on the hashing algorithm, the technical solutions provided by the embodiments of this application comprehensively consider the number of storage groups already borne by the storage nodes and the number of storage groups that each storage node is estimated to bear when the load distribution is balanced when performing load balancing. Therefore, the degree of load distribution balance of multiple storage nodes is improved.

[0007] In some embodiments, determining the target value of each storage node among the multiple storage nodes includes: if the multiple storage nodes do not meet the load balancing condition, performing a predicted migration step; the predicted migration step includes: predicting and migrating at least one storage group in the second storage node to the first storage node; the load balancing condition is used to measure the degree of uniformity of the load distribution among the multiple storage nodes; repeatedly performing the predicted migration step until the multiple storage nodes meet the load balancing condition, and then determining the number of storage groups that each storage node is predicted to carry when the multiple storage nodes meet the load balancing condition as the target value of the storage node.

[0008] Based on this, the load distribution of the multiple storage nodes can be balanced by performing the predicted migration step, and then the number of storage groups that each storage node is predicted to carry when the multiple storage nodes have a balanced load distribution can be obtained.

[0009] In some embodiments, the above-mentioned first storage node is the storage node with the largest remaining capacity parameter among the multiple storage nodes; the above-mentioned second storage node is the storage node with the smallest remaining capacity parameter among the multiple storage nodes; the above-mentioned load balancing condition includes: the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than a threshold; wherein, the remaining capacity parameter includes: remaining capacity, and / or remaining capacity ratio.

[0010] It should be understood that if the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than the threshold, it means that the range of the remaining capacity parameters among the multiple storage nodes is relatively large, so there is a problem that the remaining capacity of some storage nodes is too small or the remaining capacity ratio is too small. Therefore, the degree of uniformity of the load distribution among the multiple storage nodes can be measured according to the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node, and then the predicted migration step can be performed according to the degree of uniformity of the load distribution to determine the number of storage groups that each storage node is predicted to carry when the load distribution of the multiple storage nodes is balanced.

[0011] In some embodiments, determining N target storage nodes to which a target storage group is mapped from multiple storage nodes based on the target value of each storage node among the multiple storage nodes and the number of storage groups already borne in each storage node includes: determining M under-loaded nodes among the multiple storage nodes based on the target value of each storage node among the multiple storage nodes and the number of storage groups already borne in each storage node; where M is a positive integer, and the under-loaded node is a storage node among the multiple storage nodes whose number of borne storage groups is less than the target value; when M is greater than or equal to N, determining the priority value of each storage node among the multiple storage nodes; the priority value of a storage node is used to represent the probability that the storage node is selected as a target storage node, and the priority value is positively correlated with the probability of being selected as a target storage node; sorting the M under-loaded nodes in descending order of the priority value to obtain the sorted M under-loaded nodes; and selecting the first N under-loaded nodes from the sorted M under-loaded nodes as the N target storage nodes to which the target storage group is mapped.

[0012] Based on this, it can be ensured that each target storage node to which a target storage group is mapped is an under-loaded node, and the number of target storage groups finally borne by each target node is as close as possible and does not exceed the target value of each target storage node, so as to improve the balance degree of load distribution.

[0013] In some embodiments, determining N target storage nodes to which a target storage group is mapped from multiple storage nodes based on the target value of each storage node among the multiple storage nodes and the number of storage groups already borne in each storage node includes: determining M under-loaded nodes among the multiple storage nodes and determining full-loaded nodes among the multiple storage nodes based on the target value of each storage node among the multiple storage nodes and the number of storage groups already borne in each storage node; where M is a positive integer, the under-loaded node is a storage node among the multiple storage nodes whose number of borne storage groups is less than the target value, and the full-loaded node is a storage node among the multiple storage nodes whose number of borne storage groups is greater than or equal to the target value; when M is less than N, obtaining the remaining storage times of the M under-loaded nodes; the remaining storage times is the difference between the target value of an under-loaded node and the number of storage groups borne by the under-loaded node; if the sum of the remaining storage times of the M under-loaded nodes is greater than or equal to N, repeatedly executing the balance migration step until there are N under-loaded nodes among the multiple storage nodes; using the N under-loaded nodes as the N target storage nodes to which the target storage group is mapped; where the balance migration step includes: determining target under-loaded nodes among the M under-loaded nodes whose remaining storage times are greater than 1; and selecting one full-loaded node from the multiple storage nodes to migrate one storage group to a target under-loaded node.

[0014] It should be understood that if the number M of non-full-load nodes among multiple storage nodes is less than N, the target storage group may be stored multiple times on the same storage node. Therefore, to avoid the decline in data storage reliability caused by the multiple storage of the same storage group on the same storage node, a full-load node is selected from multiple storage nodes to migrate a storage group to a non-full-load node with a remaining storage times greater than 1. After the non-full-load node migrates the storage group, it can become a new non-full-load node for the target storage group to store, so that the target storage group has more selection space, avoiding the multiple storage of the target storage group on the same storage node and improving the reliability and stability of data storage.

[0015] In some embodiments, migrating the target storage group to N target storage nodes as described above includes: for each of the N target storage nodes, copying the disk file of the database table corresponding to the target storage group to the target storage node.

[0016] In this way, since the target storage group includes at least one set of metadata, the target storage group can be regarded as a collection of multiple sets of metadata; compared with copying a single piece of metadata to the target storage node, in this application, the collection of metadata is copied to the target storage node through the target storage group, improving the metadata migration efficiency.

[0017] In a second aspect, an embodiment of the present application provides a metadata management device, which includes: an estimation module for determining the target value of each storage node among multiple storage nodes; the target value of the storage node is used to represent the number of storage groups that the storage node is estimated to carry when the load distribution of multiple storage nodes is balanced, and the storage group includes at least one set of metadata; a mapping module for determining N target storage nodes mapped by the target storage group from multiple storage nodes based on the target value of each storage node among multiple storage nodes and the number of storage groups already carried in each storage node; N is a positive integer, and the target storage node is used to store a copy of the target storage group; the target storage group is the storage group to be migrated; a migration module for migrating the target storage group to N target storage nodes.

[0018] In some embodiments, the above-mentioned estimation module is specifically configured to perform an estimation migration step if multiple storage nodes do not meet the load balancing condition; the estimation migration step includes: estimating and migrating at least one storage group in the second storage node to the first storage node; the load balancing condition is used to measure the degree of uniformity of the load distribution among multiple storage nodes; repeatedly execute the estimation migration step until multiple storage nodes meet the load balancing condition, and then determine that the number of storage groups that each storage node is estimated to carry when multiple storage nodes meet the load balancing condition is the target value of the storage node.

[0019] In some embodiments, the above-mentioned first storage node is the storage node with the largest remaining capacity parameter among multiple storage nodes; the above-mentioned second storage node is the storage node with the smallest remaining capacity parameter among multiple storage nodes; the above-mentioned load balancing condition includes: the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than a threshold value; wherein, the remaining capacity parameter includes: remaining capacity, and / or remaining capacity ratio.

[0020] In some embodiments, the above-mentioned mapping module is specifically configured to determine M unloaded nodes among multiple storage nodes based on the target value of each storage node among multiple storage nodes and the number of storage groups already carried in each storage node; wherein, M is a positive integer, and an unloaded node is a storage node among multiple storage nodes whose number of carried storage groups is less than the target value; in the case where M is greater than or equal to N, determine the priority value of each storage node among multiple storage nodes; the priority value of a storage node is used to represent the probability that the storage node is selected as the target storage node, and the priority value is positively correlated with the probability of being selected as the target storage node; sort the M unloaded nodes in descending order of the priority value to obtain the sorted M unloaded nodes; select the first N unloaded nodes from the sorted M unloaded nodes as the N target storage nodes for mapping the target storage group.

[0021] In some embodiments, the above-mentioned mapping module is specifically configured to determine M unloaded nodes among multiple storage nodes and determine the fully loaded nodes among multiple storage nodes based on the target value of each storage node among multiple storage nodes and the number of storage groups already carried in each storage node; wherein, M is a positive integer, an unloaded node is a storage node among multiple storage nodes whose number of carried storage groups is less than the target value, and a fully loaded node is a storage node among multiple storage nodes whose number of carried storage groups is greater than or equal to the target value; in the case where M is less than N, obtain the remaining storage times of the M unloaded nodes; the remaining storage times is the difference between the target value of the unloaded node and the number of storage groups carried by the unloaded node; if the sum of the remaining storage times of the M unloaded nodes is greater than or equal to N, repeat the execution of the equilibrium migration step until there are N unloaded nodes among multiple storage nodes; use the N unloaded nodes as the N target storage nodes for mapping the target storage group; wherein, the equilibrium migration step includes: determining the target unloaded nodes among the M unloaded nodes whose remaining storage times are greater than 1; selecting a fully loaded node from multiple storage nodes to migrate a storage group to the target unloaded node.

[0022] In some embodiments, the above-mentioned migration module is specifically configured to, for each of the N target storage nodes, copy the disk file of the database table corresponding to the target storage group to the target storage node.

[0023] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the processor executes the computer instructions, the electronic device is caused to execute the metadata management method according to the first aspect and any possible design thereof.

[0024] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium includes computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the metadata management method provided in the first aspect and possible implementations.

[0025] In a fifth aspect, an embodiment of the present application provides a computer program product containing computer instructions, and when the computer instructions run on a computer, the computer is caused to execute the metadata management method provided in the above-mentioned first aspect and possible implementations.

[0026] For the specific descriptions of the second aspect to the fifth aspect and their various implementations in the present application, reference may be made to the detailed descriptions in the first aspect and its various implementations. For the beneficial effects of the second aspect to the fifth aspect and their various implementations, reference may be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic structural diagram of a distributed storage system according to some embodiments;

[0028] Figure 2 It is a flowchart of a metadata management method according to some embodiments Figure 1 ;

[0029] Figure 3 It is a flowchart of another metadata management method according to some embodiments Figure 2 ;

[0030] Figure 4 It is a flowchart of yet another metadata management method according to some embodiments Figure 3 ;

[0031] Figure 5 It is a flowchart of yet another metadata management method according to some embodiments Figure 4 ;

[0032] Figure 6 It is a flowchart of yet another metadata management method according to some embodiments Figure 5 ;

[0033] Figure 7Schematic flowchart of a method for determining remaining capacity according to some embodiments;

[0034] Figure 8 Schematic structural diagram of a metadata management device according to some embodiments;

[0035] Figure 9 Schematic structural diagram of an electronic device according to some embodiments. Detailed implementation manners

[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0037] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner. The terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0038] First, a distributed storage system related to the present application will be briefly introduced.

[0039] The distributed storage system uses distributed storage technology. Distributed storage technology is a data storage technology applied to a distributed storage system. It uses the disk space on each machine in the system through a network and constitutes these scattered storage resources into a virtual storage device, and the data is scattered and stored in various parts of the system. Figure 1 A schematic structural diagram of a distributed storage system is shown.

[0040] Refer to Figure 1 , at least one domain (such as the domain 11 in Figure 1 ) is included in the distributed storage system 10, and each domain includes a plurality of storage nodes (such as 111b and 111a shown in Figure 1 ), and there is at least one main storage node among the plurality of storage nodes (such as Figure 1The main storage node 111a) shown and multiple non-main storage nodes (e.g., Figure 1 the non-main storage node 111b) shown. Each storage node includes multiple storage groups (e.g., Figure 1 the multiple storage groups 112) shown, and the storage groups are used to store the metadata of the distributed storage system. In addition, there is a connection between the main storage node 111a and each non-main storage node 111b. Optionally, there can also be a connection between each non-main storage node 111b (not shown in the figure). To further illustrate the structure of the distributed storage system, the following will make a detailed description of Figure 1 each part in the distributed storage system 10 shown.

[0041] Domain 11 is a cluster of multiple storage nodes. Multiple domains can be divided in a distributed storage system 10 to complete different data processing functions. This application does not specifically limit the number of domains in the distributed storage system. There can be at least one main storage node 111a within a domain 11 to maintain and manage the metadata storage of multiple storage nodes within the domain. In some embodiments, before storing multiple pieces of metadata into a domain, the number of storage groups included in the domain can be specified, so that the multiple pieces of metadata are allocated to each storage group in each storage node within the domain according to the consistent hashing algorithm.

[0042] A storage node is a data storage device, and in this application, the storage node is mainly used to store metadata. Referring to Figure 1 , each storage node (such as Figure 1 111b and 111a shown) includes multiple storage groups 112. Optionally, the metadata in each storage node is divided into multiple storage groups within the storage node. This application does not specifically limit the number of storage nodes.

[0043] The main storage node 111a is a storage node elected from multiple storage nodes, and is used to manage and maintain the metadata of each storage node. This application does not specifically limit the election method and the number of main storage nodes. Optionally, the main storage node 111a can read the remaining capacity parameters of the main storage node 111a and / or the non-main storage node 111b, and the remaining capacity parameters include: remaining capacity, and / or remaining capacity ratio. Optionally, the main storage node 111a can also control the migration of the storage groups 112 in the main storage node 111a and / or the non-main storage node 111b. Optionally, the main storage node 111a can also determine the target values of the main storage node 111a and / or the non-main storage node 111b.

[0044] The storage group 112 is a logical grouping in the distributed storage system 10 and is a collection of multiple metadata. The metadata within the same storage group is stored on the same storage node. All the metadata in the same domain is divided into a specified number of storage groups, and the present application does not specifically limit the number of storage groups. In some embodiments, the metadata within a storage group is replicated and stored on different storage nodes, that is, multi-copy mapping, to improve the reliability of data storage.

[0045] The above is the introduction of the related concepts of the distributed storage system involved in the embodiments of the present application, which will not be elaborated below.

[0046] As described in the background art, in a distributed storage system, it is necessary to disperse the metadata as evenly as possible to each storage node in the system, that is, to ensure the load balance of each storage node to avoid the performance degradation of the node caused by unbalanced load. In the related art, the load balance is achieved through the consistent hashing algorithm, but this consistent hashing algorithm only achieves balance probabilistically, so the degree of load balance is limited.

[0047] In response to this, the embodiments of the present application provide a metadata management method. The core idea of this method is to comprehensively consider the number of storage groups already borne by each storage node when performing load balancing, as well as the estimated number of storage groups to be borne by each storage node when the load distribution is balanced, so that when performing load balancing on each storage node storing metadata, it does not rely solely on probabilistic allocation for balancing, thereby improving the degree of load distribution balance of multiple storage nodes in the distributed storage system.

[0048] For ease of understanding, the metadata management method provided by the present application will be specifically introduced below with reference to the accompanying drawings.

[0049] Figure 2 FIG. is a schematic flowchart of a metadata management method proposed by the present application. This metadata management method is applied to the main storage node in a distributed storage system including multiple storage nodes. As Figure 2 shown, the metadata management method includes the following steps S101 to S103:

[0050] S101. Determine the target value of each storage node among the multiple storage nodes.

[0051] Among them, the target value of the storage node is used to represent the estimated number of storage groups to be borne by the storage node when the load distribution of the multiple storage nodes is balanced. The storage group includes at least one group of metadata. Exemplarily, Figure 1 the storage group 112 in includes multiple metadata. The concept of the storage group has been described in detail in the Figure 1 illustrated embodiment and will not be elaborated here.

[0052] In some embodiments, asFigure 3 As shown, the above step S101 can be specifically implemented as the following steps S1011 to S1012:

[0053] S1011. If multiple storage nodes do not meet the load balancing condition, perform an estimated migration step, which includes: estimating the migration of at least one storage group in the second storage node to the first storage node.

[0054] Among them, the load balancing condition is used to measure the degree of uniformity of the load distribution among multiple storage nodes. It should be understood that the above estimated migration step (that is, estimating the migration of at least one storage group in the second storage node to the first storage node) is not an actual migration operation, but an assumed completed migration operation, which belongs to logical operations rather than real events.

[0055] In some embodiments, the first storage node is the storage node with the largest remaining capacity parameter among multiple storage nodes; the second storage node is the storage node with the smallest remaining capacity parameter among multiple storage nodes; optionally, the load balancing condition includes: the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than a threshold; among them, the remaining capacity parameter includes: remaining capacity, and / or remaining capacity ratio. It should be understood that the remaining capacity ratio here is the estimated remaining capacity ratio, and the remaining capacity is the estimated remaining capacity, which will not be elaborated below.

[0056] Optionally, when the remaining capacity parameter is the remaining capacity ratio, the first storage node is the storage node with the largest remaining capacity ratio among multiple storage nodes; the second storage node is the storage node with the smallest remaining capacity ratio among multiple storage nodes; the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node being greater than a threshold is specifically implemented as: the difference between the remaining capacity ratios of the first storage node and the second storage node is greater than a first threshold. The remaining capacity ratio of a storage node is the ratio of the remaining capacity of the storage node to the total capacity of the storage node.

[0057] As a specific example, the method shown in the embodiments of the present application will be specifically described below, where DSD represents a storage node and SG represents a storage group. If a domain of a distributed storage system includes five storage nodes, these five storage nodes are respectively represented by DSD1, DSD2, DSD3, DSD4, and DSD5, and the remaining capacity ratio of DSD1 is 10%, the remaining capacity ratio of DSD2 is 30%, the remaining capacity ratio of DSD3 is 50%, the remaining capacity ratio of DSD4 is 70%, and the remaining capacity ratio of DSD5 is 90%. Then, DSD5 with the largest remaining capacity ratio is used as the first storage node, and DSD1 with the smallest remaining capacity ratio is used as the second storage node. If the first threshold is 15%, since the difference between the remaining capacity parameters of the first storage node and the second storage node is 90% - 10%, that is, 80%, which is greater than 15%, the load in the storage nodes does not meet the balance condition, and it is estimated that a storage group in DSD1 will be migrated to DSD5.

[0058] It should be understood that if the difference in the remaining capacity ratios among multiple storage nodes is large, it means that the load distribution among the multiple storage nodes is quite different. Therefore, the balance of the load can be determined by comparing the difference between the storage node with the largest remaining capacity ratio and the storage node with the smallest remaining capacity ratio among the multiple storage nodes.

[0059] Optionally, when the remaining capacity parameter is the remaining capacity, the first storage node is the storage node with the largest remaining capacity among the multiple storage nodes; the second storage node is the storage node with the smallest remaining capacity among the multiple storage nodes; the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than the threshold, and the specific implementation is: the difference between the remaining capacity of the first storage node and the remaining capacity of the second storage node is greater than the second threshold.

[0060] As a specific example, if a domain of a distributed storage system includes five storage nodes, these five storage nodes are respectively represented by DSD6, DSD7, DSD8, DSD9, and DSD10, and the remaining capacity of DSD6 is 10G, the remaining capacity of DSD7 is 30G, the remaining capacity of DSD8 is 50G, the remaining capacity of DSD9 is 70G, and the remaining capacity of DSD10 is 90G. Then, DSD10 with the largest remaining capacity is used as the first storage node, and DSD6 with the smallest remaining capacity is used as the second storage node. If the first threshold is 15G, since the difference between the remaining capacity parameters of the first storage node and the second storage node is 90G - 10G, that is, 80G, which is greater than 15G, the load in the storage nodes does not meet the balance condition, and it is estimated that a storage group in DSD6 will be migrated to DSD10.

[0061] It should be understood that if the remaining capacity ratios of multiple storage nodes are relatively large, it means that the load distribution of the multiple storage nodes varies greatly. Therefore, the load balance can be determined by comparing the difference between the storage node with the largest remaining capacity and the storage node with the smallest remaining capacity among the multiple storage nodes.

[0062] S1012. Repeatedly execute the above-mentioned estimated migration steps until multiple storage nodes meet the load balance condition, and determine that the number of storage groups estimated to be borne by each storage node when the multiple storage nodes meet the load balance condition is the target value of the storage node.

[0063] It should be understood that since the estimated migration steps in step S1011 are not actual migration operations, correspondingly, the fact that the multiple storage nodes meet the load balance condition described in step S1012 does not mean that the multiple storage nodes actually meet the load balance condition, but rather indicates that it is assumed that the multiple storage nodes meet the load balance condition after the completion of the estimated migration steps.

[0064] It can be seen that the above steps S1011 and S1012 are actually a loop process, aiming to make the load distribution of multiple storage nodes balanced by executing the estimated migration steps, and then obtain the number of storage groups estimated to be borne by each storage node when the multiple storage nodes are load-balanced, that is, the target value of each storage node.

[0065] Exemplarily, still taking the remaining capacity parameter as the remaining capacity ratio, if the first threshold is still 15%, and after one storage group in the above DSD1 is migrated to DSD5, the remaining capacity ratio of DSD1 becomes 40%, and the remaining capacity ratio of DSD5 becomes 60%, then at this time, DSD4 with the largest remaining capacity ratio is taken as the first storage node, and DSD2 with the smallest remaining capacity ratio is taken as the second storage node. Since the difference between the remaining capacity parameters of the first storage node and the second storage node is 70% - 30%, that is, 40%, which is greater than 15%, the load in the storage nodes still does not meet the balance condition. Therefore, the primary storage node estimates to migrate one storage group in DSD2 to DSD4. If the multiple storage nodes do not meet the load balance condition after the estimated migration, then repeat the above steps by analogy until the multiple storage nodes meet the load balance condition, and determine that the number of storage groups estimated to be borne by each storage node when the multiple storage nodes meet the load balance condition is the target value of the storage node.

[0066] Alternatively, still taking the remaining capacity parameter as the remaining capacity as an example, if the second threshold is still 15G, and after one storage group in the above DSD6 is migrated to DSD10, the remaining capacity ratio of DSD6 becomes 40G, and the remaining capacity ratio of DSD10 becomes 60G, then at this time, DSD9 with the largest remaining capacity ratio is taken as the first storage node, and DSD7 with the smallest remaining capacity ratio is taken as the second storage node. Since the difference between the remaining capacity parameters of the first storage node and the second storage node is 70G - 30G, which is 40G and greater than 15G, the load in the storage nodes does not meet the balance condition. The primary storage node migrates one storage group in DSD7 to DSD9. If it is estimated that the multiple storage nodes do not meet the load balance condition after the migration, then the above steps are repeatedly executed by analogy until the multiple storage nodes meet the load balance condition. Then, the number of storage groups that each storage node is estimated to carry when the multiple storage nodes meet the load balance condition is determined as the target value of that storage node.

[0067] It should be understood that if the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than the threshold, it means that the range of the remaining capacity parameters among the multiple storage nodes is relatively large. Therefore, there is a problem that the remaining capacity of some storage nodes is too small or the remaining capacity ratio is too small. Therefore, the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node can be used to measure the uniformity of the load distribution among the multiple storage nodes. Furthermore, the estimated migration steps are executed according to the uniformity of the load distribution to determine the number of storage groups that each storage node is estimated to carry when the load distribution of the multiple storage nodes is balanced.

[0068] S102. Based on the target values of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node, determine N target storage nodes to which the target storage group is mapped from the multiple storage nodes, where N is a positive integer.

[0069] Among them, the target storage node is used to store the replica of the target storage group; the target storage group is the storage group to be migrated.

[0070] In some embodiments, according to the target values of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node, each storage node can be divided into a full-load node or a non-full-load node. Among them, the non-full-load node is a storage node among the multiple storage nodes whose number of carried storage groups is less than the target value. The full-load node is a storage node among the multiple storage nodes whose number of carried storage groups is greater than or equal to the target value.

[0071] In some embodiments, as Figure 4 shown, the above step S102 is specifically implemented as the following steps S1021a to S1024a:

[0072] S1021a. Determine M unloaded nodes among the multiple storage nodes based on the target values of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node.

[0073] Where M is a positive integer.

[0074] S1022a. When M is greater than or equal to N, determine the priority value of each storage node among the multiple storage nodes.

[0075] Where the priority value of a storage node is used to represent the probability that the storage node is selected as a target storage node, and the priority value is positively correlated with the probability of being selected as a target storage node.

[0076] S1023a. Sort the M unloaded nodes in descending order of the priority value to obtain the sorted M unloaded nodes.

[0077] S1024a. Select the first N unloaded nodes from the sorted M unloaded nodes as the N target storage nodes mapped by the target storage group.

[0078] It should be understood that the priority value of each of the N unloaded nodes is greater than or equal to the priority value of the other unloaded nodes among the above M unloaded nodes.

[0079] Exemplarily, when N is 3, if a domain of a distributed storage system includes 10 storage nodes: DSD21, DSD22, DSD23, DSD24, DSD25, DSD26, DSD27, DSD28, DSD29, and DSD30, and the priority values of the above DSD21 to DSD30 calculated according to step S21 are in turn: 45, 134, 267, 44, 115, 78, 168, 24, 423, 580. Then the primary storage node determines that the 3 unloaded nodes with the largest priority values among the above 10 storage nodes are DSD30, DSD29, and DSD23 according to the above priority values, the target values of each storage node, and the number of storage groups actually carried by each storage node.

[0080] Based on this, it can be ensured that each target storage node mapped by the target storage group is an unloaded node, and the number of target storage groups finally carried by each target node is as close as possible and does not exceed the target value of each target storage node, so as to improve the balance degree of load distribution.

[0081] In some other embodiments, as Figure 5 shown, the above step S102 is specifically implemented as the following steps S1021b to S1024b:

[0082] S1021 b. Based on the target values of each storage node among multiple storage nodes and the number of storage groups already carried in each storage node, determine M unloaded nodes among the multiple storage nodes, and determine the fully loaded nodes among the multiple storage nodes; where M is a positive integer.

[0083] S1022b. When M is less than N, obtain the remaining storage times of the M unloaded nodes.

[0084] Where the remaining storage times is the difference between the target value of the unloaded node and the number of storage groups carried by the unloaded node.

[0085] S1023b. If the sum of the remaining storage times of the M unloaded nodes is greater than or equal to N, repeat the balanced migration step until there are N unloaded nodes among the multiple storage nodes.

[0086] Where the balanced migration step includes: determining the target unloaded nodes among the M unloaded nodes with remaining storage times greater than 1; selecting a fully loaded node from the multiple storage nodes to migrate a storage group to the target unloaded node.

[0087] S1024b. Use the N unloaded nodes as the N target storage nodes mapped by the target storage group.

[0088] It should be understood that if the number M of non-fully loaded nodes among the multiple storage nodes is less than N, the target storage group may be stored in the same storage node multiple times. Therefore, to avoid the decline in data storage reliability caused by the same storage group being stored in the same storage node multiple times, select a fully loaded node from the multiple storage nodes to migrate a storage group to a non-fully loaded node with remaining storage times greater than 1. After the non-fully loaded node migrates the storage group, it can become a new non-fully loaded node for the target storage group to store, so that the target storage group has more selection space, avoiding the target storage group being stored in the same storage node multiple times, and improving the reliability and stability of data storage.

[0089] S103. Migrate the target storage group to the N target storage nodes.

[0090] In some embodiments, the above step S103 is specifically implemented as: for each target storage node among the N target storage nodes, copy the disk file of the database table corresponding to the target storage group to the target storage node. In this way, compared with copying single metadata to the target storage node, the present application copies the set of metadata to the target storage node through the target storage group, improving the metadata migration efficiency.

[0091] In some embodiments, the above step S103 is specifically implemented as follows: The primary storage node determines the storage node where the target storage group is located from multiple storage nodes; the primary storage node sends a storage group migration instruction to the storage node where the target storage group is located; after receiving the storage group migration instruction, the storage node where the target storage group is located migrates the target storage group to N target storage nodes according to the storage group migration instruction.

[0092] Figure 2 The provided technical solution at least brings the following beneficial effects: Compared with the related art where load balancing is only based on probability when implementing load balancing based on the hash algorithm, the technical solution provided in the embodiments of the present application comprehensively considers the number of storage groups already borne by the storage nodes and the estimated number of storage groups to be borne by each storage node when the load distribution is balanced during load balancing. Therefore, the balance degree of the load distribution of multiple storage nodes in the distributed storage system is improved.

[0093] In some embodiments, the priority values of the above-mentioned multiple storage nodes in step S1022a can be determined by the primary storage node through the consistent hashing algorithm. For example, as Figure 6 shown, the above step S1022a can be specifically implemented as the following steps S10221a to S10222a:

[0094] S10221a. When M is greater than or equal to N, determine the random value of each storage node according to the identifier of the target storage group and the identifier of each storage node in the multiple storage nodes.

[0095] In some examples, the identifier of the target storage group, the identifier of the storage node, and the random value of the storage node satisfy the following relationship:

[0096] sg_rand(i) = Hash(string(sg_id) + string(DSD_id(i)))

[0097] where string(sg_id) is the identifier of the target storage group, string(DSD_id(i)) is the identifier of the i-th storage node in the multiple storage nodes, and sg_rand(i) is the random value of the i-th storage node in the multiple storage nodes.

[0098] S10222a. Determine the priority value of each storage node according to the random value of each storage node and the weight of each storage node in the multiple storage nodes.

[0099] In some examples, the random value of the storage node, the weight of the storage node, and the priority value of the storage node satisfy the following relationship:

[0100] W(i) = DSD_w(i) * g_rand(i)

[0101] Among them, DSD_w(i) is the weight of the i-th storage node among multiple storage nodes, sg_rand(i) is the random value of the i-th storage node among the multiple storage nodes corresponding to the target storage group, and W(i) is the priority value of the i-th storage node among the multiple storage nodes.

[0102] In some examples, the weight of each storage node is related to the remaining capacity parameter of each storage node.

[0103] Exemplarily, the weight of a storage node is proportional to the remaining capacity ratio of the storage node. For example, if a domain of a distributed storage system includes 5 storage nodes: DSD11, DSD12, DSD13, DSD14, and DSD15. Among them, the remaining capacity ratios of DSD11, DSD12, DSD13, DSD14, and DSD15 are 30%, 36%, 39%, 50%, and 44% in sequence, then the weights of DSD11, DSD12, DSD13, DSD14, and DSD15 are set to: 0.3, 0.36, 0.39, 0.5, 0.44 respectively.

[0104] Alternatively, the weight of a storage node is proportional to the remaining capacity of the storage node. For example, if a domain of a distributed storage system includes 5 storage nodes: DSD16, DSD17, DSD18, DSD19, and DSD20. Among them, the remaining capacities of DSD16, DSD17, DSD18, DSD19, and DSD20 are 30G, 36G, 39G, 50G, and 44G in sequence, then the weights of DSD16, DSD17, DSD18, DSD19, and DSD20 are set to: 30, 36, 39, 50, 44 respectively.

[0105] It should be understood that the relationships between the weights of the above-mentioned storage nodes and the remaining capacity parameters of the storage nodes are only examples. For example, it can also be an exponential relationship, a logarithmic relationship, etc. The present application does not limit this.

[0106] In some embodiments, to determine whether multiple storage nodes meet the load balancing condition after performing the estimated migration step described in step S1011 above, the remaining capacity of the multiple storage nodes should also be determined, so as to judge whether the multiple storage nodes meet the load balancing condition according to the remaining capacity of the multiple storage nodes. Exemplarily, before the first execution of the estimated migration step, the estimated remaining capacity parameter of the storage node is consistent with the actual remaining capacity parameter. Therefore, the remaining capacity of each storage node before the first execution of the estimated migration step can be directly read and recorded from each storage node. In this way, it is convenient to calculate the remaining capacity parameter (such as remaining capacity and / or remaining capacity ratio) of each storage node after the first estimated transfer. Alternatively, after at least one estimated migration step has been executed, the estimated remaining capacity parameter of the storage node has changed relative to the actual remaining capacity parameter, and the primary storage node cannot directly read the remaining capacity of each storage node after at least one estimated migration step from each storage node. To further determine whether the multiple storage nodes meet the load balancing condition after at least one estimated migration step has been executed, Figure 7 shows a method for determining the remaining capacity, which is used to determine the remaining capacity of the first storage node and the second storage node after the most recent execution of the estimated migration step. Refer to Figure 7 , the method includes the following steps S201 to S203:

[0107] S201. Determine the average storage group capacity of the second storage node.

[0108] In some examples, the average storage group capacity of the first storage node is obtained by dividing the stored capacity of the first storage node by the number of storage groups of the second storage node. In this way, the accurate average storage group capacity of the first storage node can be calculated.

[0109] In some examples, the average storage group capacity of the second storage node is obtained by dividing the stored capacity of the second storage node by the number of storage groups of the second storage node. In this way, the accurate average storage group capacity of the second storage node can be calculated.

[0110] In some other examples, the average storage group capacity of the first storage node and the average storage group capacity of the second storage node are the same, and both are the quotient of the total stored capacity in the same domain divided by the total number of storage groups in the domain. It should be understood that all metadata is relatively evenly distributed in each storage group according to the consistent hashing algorithm. Therefore, it can be considered that the capacities of each storage node are the same. In this way, it is not necessary to calculate the average storage group capacity of the storage node every time before estimating the migration, saving the amount of computation.

[0111] S202. Determine the remaining capacity of the first storage node and the remaining capacity of the second storage node before the most recent execution of the estimated migration step.

[0112] In some examples, if the most recent execution of the estimated migration step is the first execution of the estimated migration step by the primary storage node, directly read the remaining capacities of the first storage node and the second storage node from the first storage node and the second storage node. If the most recent execution of the estimated migration step is not the first execution of the estimated migration step by the primary storage node, then use the remaining capacity of the first storage node determined after the previous execution of the estimated migration step as the remaining capacity of the first storage node before the most recent execution of the estimated migration step; and use the remaining capacity of the second storage node determined after the previous execution of the estimated migration step as the remaining capacity of the second storage node before the most recent execution of the estimated migration step.

[0113] S203. Determine the remaining capacities of the first storage node and the second storage node after the most recent execution of the estimated migration step according to the average storage group capacity of the second storage node, and the remaining capacities of the first storage node and the second storage node before the most recent execution of the estimated migration step.

[0114] Among them, the remaining capacity of the first storage node after the most recent execution of the estimated migration step is the sum of the remaining capacity of the first storage node before the most recent execution of the estimated migration step and the average storage group capacity of the second storage node; the remaining capacity of the second storage node after the most recent execution of the estimated migration step is the difference between the remaining capacity of the second storage node before the most recent execution of the estimated migration step and the average storage group capacity of the second storage node. Based on this, it can be determined whether multiple storage nodes after executing at least one estimated migration step meet the load balancing condition.

[0115] In some embodiments, if a new storage node is added to the distributed storage system, the primary storage node executes the above steps S101 to S103 so that the new storage node also stores metadata load, thereby enabling the distributed storage system after adding the new storage node to meet the load balancing condition.

[0116] In other embodiments, the primary storage node in the distributed storage system detects at regular intervals whether multiple storage nodes in the distributed storage system meet the load balancing condition. If not, execute the above steps S101 to S103 to improve the load balancing degree.

[0117] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0118] As Figure 8 shown, the embodiments of the present application provide a metadata management device for executing Figure 2 the metadata management method shown. The metadata management device includes: an estimation module 301, a mapping module 302, and a migration module 303.

[0119] The estimation module 301 is used to determine the target value of each storage node among multiple storage nodes; the target value of the storage node is used to represent the number of storage groups that the storage node is estimated to carry when the load distribution of the multiple storage nodes is balanced.

[0120] The mapping module 302 is used to determine N target storage nodes to which the target storage group is mapped from multiple storage nodes based on the target value of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node; N is a positive integer, and the target storage nodes are used to store copies of the target storage group; the target storage group is the storage group to be migrated, and the storage group includes at least one set of metadata.

[0121] The migration module 303 is used to migrate the target storage group to the N target storage nodes.

[0122] In some embodiments, the above-mentioned estimation module 301 is specifically used to, if the multiple storage nodes do not meet the load balancing condition, execute an estimation migration step, and the estimation migration step includes: estimating and migrating at least one storage group in the second storage node to the first storage node; the load balancing condition is used to measure the degree of uniformity of the load distribution among the multiple storage nodes; repeatedly execute the estimation migration step until the multiple storage nodes meet the load balancing condition, and then determine that the number of storage groups that each storage node is estimated to carry when the multiple storage nodes meet the load balancing condition is the target value of the storage node.

[0123] In some embodiments, the first storage node is the storage node with the largest remaining capacity parameter among multiple storage nodes; the second storage node is the storage node with the smallest remaining capacity parameter among multiple storage nodes; the load balancing condition includes: the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than a threshold; wherein, the remaining capacity parameter includes: remaining capacity, and / or remaining capacity ratio.

[0124] In some embodiments, the mapping module 302 is specifically configured to determine M unloaded nodes among multiple storage nodes based on the target value of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node; where M is a positive integer, and an unloaded node is a storage node among the multiple storage nodes whose number of carried storage groups is less than the target value; when M is greater than or equal to N, determine the priority value of each storage node among the multiple storage nodes; the priority value of a storage node is used to represent the probability that the storage node is selected as the target storage node, and the priority value is positively correlated with the probability of being selected as the target storage node; sort the M unloaded nodes in descending order of the priority value to obtain the sorted M unloaded nodes; select the first N unloaded nodes from the sorted M unloaded nodes as the N target storage nodes for mapping the target storage group.

[0125] In some embodiments, the mapping module 302 is specifically configured to determine M unloaded nodes among multiple storage nodes and determine the fully loaded nodes among the multiple storage nodes based on the target value of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node; where M is a positive integer, an unloaded node is a storage node among the multiple storage nodes whose number of carried storage groups is less than the target value, and a fully loaded node is a storage node among the multiple storage nodes whose number of carried storage groups is greater than or equal to the target value; when M is less than N, obtain the remaining storage times of the M unloaded nodes; the remaining storage times is the difference between the target value of the unloaded node and the number of storage groups carried by the unloaded node; if the sum of the remaining storage times of the M unloaded nodes is greater than or equal to N, repeatedly execute the equilibrium migration step until there are N unloaded nodes among the multiple storage nodes; use the N unloaded nodes as the N target storage nodes for mapping the target storage group; wherein, the equilibrium migration step includes: determining the target unloaded nodes among the M unloaded nodes whose remaining storage times are greater than 1; selecting a fully loaded node from the multiple storage nodes to migrate a storage group to the target unloaded node.

[0126] In some embodiments, the migration module 303 is specifically configured to copy the disk file of the database table corresponding to the target storage group to the target storage node for each target storage node among the N target storage nodes.

[0127] In the case where the functions of the above integrated modules are implemented in the form of hardware, the embodiments of the present application provide a schematic structural diagram of an electronic device involved in the above embodiments. As Figure 9 shown, the electronic device 400 includes: a processor 402, a communication interface 403, and a bus 404. Optionally, the electronic device may further include a memory 401.

[0128] The processor 402 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof that can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 402 may also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0129] The communication interface 403 is used to connect to other devices through a communication network. The communication network may be an Ethernet, a radio access network, a wireless local area network (WLAN), etc.

[0130] The memory 401 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions. It may also be an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0131] As a possible implementation, the memory 401 may exist independently of the processor 402. The memory 401 may be connected to the processor 402 through the bus 404 for storing instructions or program code. When the processor 402 calls and executes the instructions or program code stored in the memory 401, it can implement the metadata management method provided by the embodiments of the present application.

[0132] In another possible implementation, the memory 401 may also be integrated with the processor 402.

[0133] The bus 404 can be an extended industry standard architecture (EISA) bus or the like. The bus 404 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 only a thick line is used to represent it in Figure 9 , but it does not mean that there is only one bus or one type of bus.

[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the electronic device is divided into different functional modules to complete all or part of the functions described above.

[0135] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be instructed by computer instructions to complete the relevant hardware. The program can be stored in the above computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be the memory in any of the foregoing embodiments. The above computer-readable storage medium can also be an external storage device of the above electronic device, such as a plug-in hard disk equipped on the above electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the above computer-readable storage medium can also include both the internal storage unit of the above electronic device and the external storage device. The above computer-readable storage medium is used to store the above computer program and other programs and data required by the above electronic device. The above computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0136] The embodiment of the present application also provides a computer program product. The computer product includes a computer program. When the computer program product runs on a computer, the computer is enabled to execute any one of the metadata management methods provided in the above embodiments.

[0137] Although the present application has been described in connection with various embodiments, those skilled in the art will recognize other variations of the disclosed embodiments upon review of the figures, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce a favorable effect.

[0138] Although the present application has been described in connection with specific features and their embodiments, it will be apparent that various modifications and combinations can be made without departing from the spirit and scope of the application. Accordingly, the specification and drawings are merely exemplary illustrations of the application as defined by the appended claims and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the application. It is obvious that those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

[0139] As described above, the specific implementation manners of the present application are only described, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A metadata management method, characterized in that, A primary storage node applied to a distributed storage system including multiple storage nodes, the method comprising: Determining a target value for each storage node among the multiple storage nodes; the target value of the storage node is used to characterize the number of storage groups that the storage node is estimated to carry when the load distribution of the multiple storage nodes is balanced, and the storage group includes at least one set of metadata; Based on the target value of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node, determining N target storage nodes to which a target storage group is mapped from the multiple storage nodes; N is a positive integer, and the target storage node is used to store a copy of the target storage group; the target storage group is a storage group to be migrated; Migrating the target storage group to the N target storage nodes; Wherein, the N target storage nodes are N unloaded nodes determined according to the priority values of more than N unloaded nodes, and the priority value is used to characterize the probability that the unloaded node is selected as the target storage node, so that the number of storage groups finally carried by the unloaded node after being selected as the target storage node has a minimum difference from its target value and does not exceed its target value.

2. The method according to claim 1, wherein The determining the target value of each storage node among the multiple storage nodes includes: If the multiple storage nodes do not meet the load balancing condition, performing an estimated migration step, the estimated migration step including: estimating and migrating at least one storage group in a second storage node to a first storage node; the load balancing condition is used to measure the degree of uniformity of the load distribution among the multiple storage nodes; Repeatedly performing the estimated migration step until the multiple storage nodes meet the load balancing condition, and determining the number of storage groups that each storage node is estimated to carry when the multiple storage nodes meet the load balancing condition as the target value of the storage node.

3. The method according to claim 2, wherein The first storage node is the storage node with the largest remaining capacity parameter among the multiple storage nodes; The second storage node is the storage node with the smallest remaining capacity parameter among the multiple storage nodes; The load balancing condition includes: the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than a threshold; wherein, the remaining capacity parameter includes: remaining capacity, and / or remaining capacity ratio.

4. The method according to any one of claims 1 to 3, characterized in that, The determining N target storage nodes to which a target storage group is mapped from the multiple storage nodes based on the target value of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node includes: Based on the target value of each storage node among the multiple storage nodes and the number of storage groups already carried in each storage node, determining M unloaded nodes among the multiple storage nodes; wherein M is a positive integer, and the unloaded node is a storage node among the multiple storage nodes whose number of carried storage groups is less than the target value; When M is greater than or equal to N, determine the priority value of each of the multiple storage nodes; the priority value of the storage node is used to characterize the probability that the storage node is selected as the target storage node, and the priority value is positively correlated with the probability of being selected as the target storage node; Sort the M unfull nodes in descending order of the priority value to obtain the sorted M unfull nodes; Select the first N unfull nodes from the sorted M unfull nodes as the N target storage nodes mapped by the target storage group.

5. The method according to claim 4, wherein The determining the N target storage nodes mapped by the target storage group from the multiple storage nodes based on the target value of each storage node in the multiple storage nodes and the number of storage groups already carried in each storage node includes: Based on the target value of each storage node in the multiple storage nodes and the number of storage groups already carried in each storage node, determine M unfull nodes in the multiple storage nodes and determine the full nodes in the multiple storage nodes; where M is a positive integer, the unfull node is a storage node in the multiple storage nodes whose number of carried storage groups is less than the target value, and the full node is a storage node in the multiple storage nodes whose number of carried storage groups is greater than or equal to the target value; When M is less than N, obtain the remaining storage times of the M unfull nodes; the remaining storage times is the difference between the target value of the unfull node and the number of storage groups carried by the unfull node; If the sum of the remaining storage times of the M unfull nodes is greater than or equal to N, repeatedly execute the balanced migration step until there are N unfull nodes in the multiple storage nodes; Wherein, the balanced migration step includes: determining the target unfull nodes among the M unfull nodes whose remaining storage times are greater than 1; selecting a full node from the multiple storage nodes to migrate a storage group to the target unfull node.

6. The method according to claim 1, wherein The migrating the target storage group to the N target storage nodes includes: For each of the N target storage nodes, copy the disk file of the database table corresponding to the target storage group to the target storage node.

7. A metadata management device, characterized in that, including: An estimation module, configured to determine the target value of each storage node in the multiple storage nodes; The target value of the storage node is used to characterize the number of storage groups estimated to be carried by the storage node when the load distribution of the multiple storage nodes is balanced, and the storage group includes at least one set of metadata; A mapping module, configured to determine N target storage nodes mapped by the target storage group from the multiple storage nodes based on the target value of each storage node in the multiple storage nodes and the number of storage groups already carried in each storage node; N is a positive integer, and the target storage node is used to store a copy of the target storage group; The target storage group is the storage group to be migrated; A migration module, configured to migrate the target storage group to the N target storage nodes; Among them, the N target storage nodes are N unloaded nodes determined according to the priority values of each of the more than N unloaded nodes. The priority value is used to represent the probability that the unloaded node is selected as the target storage node, so that the difference between the number of storage groups finally borne by the unloaded node after being selected as the target storage node and its target value reaches the minimum and does not exceed its target value.

8. The apparatus according to claim 7, wherein the estimation module is specifically configured to, if the plurality of storage nodes do not meet the load balancing condition, perform an estimation migration step, and the estimation migration step includes: estimating and migrating at least one storage group in the second storage node to the first storage node; the load balancing condition is used to measure the degree of uniformity of the load distribution among the plurality of storage nodes; repeatedly execute the estimation migration step until the plurality of storage nodes meet the load balancing condition, and determine the number of storage groups estimated to be borne by each storage node when the plurality of storage nodes meet the load balancing condition as the target value of the storage node; the first storage node is the storage node with the largest remaining capacity parameter among the plurality of storage nodes; the second storage node is the storage node with the smallest remaining capacity parameter among the plurality of storage nodes; the load balancing condition includes: the difference between the remaining capacity parameter of the first storage node and the remaining capacity parameter of the second storage node is greater than a threshold; wherein, the remaining capacity parameter includes: remaining capacity, and / or remaining capacity ratio; the mapping module is specifically configured to determine M unloaded nodes among the plurality of storage nodes based on the target value of each storage node among the plurality of storage nodes and the number of storage groups already borne in each storage node; where M is a positive integer, and the unloaded node is a storage node among the plurality of storage nodes whose number of storage groups borne is less than the target value; in the case where M is greater than or equal to N, determine the priority value of each storage node among the plurality of storage nodes; the priority value of the storage node is used to represent the probability that the storage node is selected as the target storage node, and the priority value is positively correlated with the probability of being selected as the target storage node; sort the M unloaded nodes in descending order of priority value to obtain the sorted M unloaded nodes; select the first N unloaded nodes from the sorted M unloaded nodes as the N target storage nodes for target storage group mapping; The mapping module is specifically configured to determine M unloaded nodes among the multiple storage nodes and determine the fully loaded nodes among the multiple storage nodes based on the target values of the respective storage nodes in the multiple storage nodes and the number of storage groups already carried in the respective storage nodes; where M is a positive integer, the unloaded nodes are the storage nodes among the multiple storage nodes whose number of carried storage groups is less than the target value, and the fully loaded nodes are the storage nodes among the multiple storage nodes whose number of carried storage groups is greater than or equal to the target value; when M is less than N, obtain the remaining storage times of the M unloaded nodes; the remaining storage times are the difference between the target value of the unloaded node and the number of storage groups carried by the unloaded node; if the sum of the remaining storage times of the M unloaded nodes is greater than or equal to N, repeatedly execute the equilibrium migration step until there are N unloaded nodes among the multiple storage nodes; where the equilibrium migration step includes: determining the target unloaded nodes among the M unloaded nodes whose remaining storage times are greater than 1; selecting a fully loaded node from the multiple storage nodes to migrate a storage group to the target unloaded node. The migration module is specifically configured to, for each of the N target storage nodes, copy the disk file of the database table corresponding to the target storage group to the target storage node.

9. An electronic device, characterized in that, Including: A memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; Wherein, when the processor executes the computer instructions, the electronic device is caused to execute the metadata management method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions; Wherein, when the computer instructions run on an electronic device, the electronic device is caused to execute the metadata management method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Load adjustment method, management node and storage medium

    CN114143326A

  • Load balancing method and device and computer equipment

    CN114327905A