A distributed file system index calculation method, device, and electronic device
By using the lottery algorithm to calculate the storage locations of virtual nodes and redundant shards in the distributed file system and generate a hash function, the problem of large amount of migration during data migration in the distributed file system is solved, and more uniform file storage is achieved and server burden is reduced.
Patent Information
- Application Number
- CN202111617746.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-27
AI Technical Summary
When migrating data, the existing distributed file system has a large amount of migration, resulting in a large workload on the server.
The lottery algorithm is used to calculate the disk group and disks where each virtual node and redundant shard are located, and a hash function in the form of a two-dimensional array is generated to represent the disk where each redundant shard of each virtual node is located.
Calculating the disk groups and disks to be stored through a two-step lottery algorithm can make the file storage more evenly, significantly reduce the amount of migration of sharded data, and alleviate the problem of server work burden.
Smart Images

Figure CN114265817B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of databases, and in particular, to a method, device, and electronic device for calculating an index of a distributed file system. Background Art
[0002] In the design of a distributed file system, a file is divided into K original data blocks, and then M redundant data blocks are generated using an erasure code algorithm to form K + M sharded data with redundancy. Then, these sharded data are distributed on different disks within a disaster recovery domain to achieve the ability of decentralized disaster recovery. As the capacity of the distributed file system continues to increase, the metadata for file indexing will become more and more, and the cost of managing the huge metadata is also gradually increasing. Therefore, the development direction of the distributed file system is to use the consistent hashing technology to replace the file metadata. The hashing technology can calculate the hash value using the file name and find the server where this file is located through the hash value.
[0003] When a disk is lost or added in the server, data migration needs to be performed on the sharded data. Based on the current hash index algorithm, when performing the above data migration, the amount of migration generated is very large, that is, the amount of data to be migrated is very large, resulting in a large workload of the server. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, and electronic device for calculating an index of a distributed file system, which alleviates the problem of a large workload of the server in the existing distributed file system.
[0005] In a first aspect, the present invention provides a method for calculating an index of a distributed file system, which is applied to a distributed file system. The method includes:
[0006] A parameter acquisition step of acquiring the number of virtual nodes to be allocated, the number of redundant shards, and the system parameters of the distributed file system;
[0007] A virtual node calculation step of calculating the disk group where each virtual node is located using a lottery algorithm;
[0008] A redundant shard calculation step of calculating the disk where each redundant shard is located using a lottery algorithm.
[0009] Further, the virtual node calculation step includes:
[0010] A first lottery step: Based on the current virtual node serial number and system parameters, calculating the disk group where the current virtual node is located using a first lottery function;
[0011] Performing the redundant shard calculation step on the current virtual node;
[0012] Increment the current virtual node sequence number by 1, and return to the first lottery step until the disk group where each virtual node is located is calculated.
[0013] In one embodiment, the redundant shard calculation step includes:
[0014] Second lottery step: Based on the current virtual node sequence number, the current redundant shard sequence number, and system parameters, use the second lottery function to calculate the disk where the current redundant shard of the current virtual node is located;
[0015] Increment the current redundant shard sequence number by 1, and return to the second lottery step until the disk where each redundant shard is located is calculated.
[0016] In another embodiment, the redundant shard calculation step includes:
[0017] Second lottery step: Based on the current virtual node sequence number, the current redundant shard sequence number, and system parameters, use the second lottery function to calculate the server where the current redundant shard of the current virtual node is located;
[0018] Third lottery step: Based on the current virtual node sequence number and system parameters, use the third lottery function to calculate the disk where the current redundant shard of the current virtual node is located;
[0019] Increment the current redundant shard sequence number by 1, and return to the second lottery step until the disk where each redundant shard is located is calculated.
[0020] Further, before incrementing the current redundant shard sequence number by 1 and returning to the second lottery step, it further includes:
[0021] Remove the alternative disks that do not meet the disaster recovery conditions from the system parameters.
[0022] Further, the system parameters include the number of servers, the number of disks of each server, the number of disk groups, and the disk weight values of each disk.
[0023] Further, the method further includes:
[0024] Perform xxhash calculation on the universally unique identifier of the file to be stored, and obtain the virtual node corresponding to the file;
[0025] Determine the disk group to which the file belongs according to the virtual node corresponding to the file, and the disks to which each redundant shard of the file belongs.
[0026] In a second aspect, the present invention further provides a distributed file system index calculation device, which is applied to a distributed file system. The device includes:
[0027] A parameter acquisition module, configured to acquire the number of virtual nodes to be allocated, the number of redundant shards, and the system parameters of the distributed file system;
[0028] A virtual node calculation module, configured to calculate the disk group where each virtual node is located by using a lottery algorithm;
[0029] A redundant shard calculation module, configured to calculate the disk where each redundant shard is located by using a lottery algorithm.
[0030] In a third aspect, the present invention further provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps of the above method are implemented.
[0031] In a fourth aspect, the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions. When the computer-executable instructions are called and run by the processor, the computer-executable instructions cause the processor to run the above method.
[0032] The distributed file system index calculation method provided by the present invention first obtains the number of virtual nodes to be allocated, the number of redundant shards, and the system parameters of the distributed file system through a parameter acquisition step; then, through a virtual node calculation step, calculates the disk group where each virtual node is located by using a lottery algorithm; and finally, through a redundant shard calculation step, calculates the disk where each redundant shard is located by using a lottery algorithm, obtaining a hash function in the form of a two-dimensional array. The number of rows of the array is equal to the number of virtual nodes, and the number of columns is equal to the number of redundant shards, which is used to represent the disk where each redundant shard of each virtual node is located. When a file needs to be stored, the corresponding virtual node can be found from the hash function according to the hash index value of the file, and the disk where each redundant shard of the file should be stored. By calculating the disk group and the disk to be stored through two-step lottery algorithms respectively, the storage of files can be made more uniform. When a disk is lost or added in a certain server, the migration amount of sharded data can be significantly reduced, alleviating the problem of heavy workload of servers in the existing distributed file system.
[0033] Correspondingly, the distributed file system index calculation device, electronic device, and computer-readable storage medium provided by the embodiments of the present invention also have the above technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 Flowchart of the distributed file system index calculation method provided by the embodiment of the present invention;
[0036] Figure 2 Flowchart of the distributed file system index calculation method provided by Embodiment 1 of the present invention;
[0037] Figure 3 Flowchart of the distributed file system index calculation method provided by Embodiment 2 of the present invention;
[0038] Figures 4 to 6 Penalty rate effect chart of the embodiment of the present invention;
[0039] Figure 7 Schematic diagram of the distributed file system index calculation device provided by Embodiment 3 of the present invention. Specific embodiments
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0041] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0042] The embodiment of the present invention provides a distributed file system index calculation method, which can be applied to a distributed file system. As Figure 1 shown, the method includes the following steps:
[0043] S1. Parameter acquisition step: Obtain the number of virtual nodes to be allocated, the number of redundant shards, and the system parameters of the distributed file system.
[0044] S2. Virtual node calculation step: Use the lottery algorithm to calculate the disk group where each virtual node is located.
[0045] S3. Redundant shard calculation step: Use the lottery algorithm to calculate the disk where each redundant shard is located.
[0046] By adopting the distributed file system index calculation method provided by the embodiments of the present invention, a hash function in the form of a two-dimensional array can be finally obtained. The number of rows of this array is equal to the number of virtual nodes, and the number of columns is equal to the number of redundant shards, which is used to represent the disks where each redundant shard of each virtual node is located. When a file needs to be stored, the corresponding virtual node can be found from the hash function according to the hash index value of the file, as well as the disks where each redundant shard of the file should be stored. By calculating the disk group and the disk to be stored through the two-step lottery algorithm respectively, the storage of files can be made more uniform. When a disk is lost or added in a certain server, the migration volume of sharded data can be significantly reduced, alleviating the problem of heavy workload on servers in the existing distributed file system.
[0047] Embodiment 1:
[0048] As Figure 2 shown, in the embodiments of the present invention, the distributed file system index calculation method includes the following steps:
[0049] S11. Parameter acquisition step: Acquire the number of virtual nodes and redundant shards to be allocated, as well as the system parameters of the distributed file system.
[0050] This step can be calculated according to experience: Distribute 40 hash landing points per 1TB of storage space to estimate the rated number of hash landing points of the entire volume. For example, if a user is allocated a maximum of 20TB of space, the number of hash points for this volume is 800, that is, the number of virtual nodes (Vnode) is 800. The number of redundant shards depends on the erasure protection method: P = K + M, where K is the number of original data blocks and M is the number of redundant data blocks.
[0051] This embodiment calculates the distribution of virtual nodes of this volume, and calculates that the hash function is a two-dimensional array of a virtual node placement table. Each row of this two-dimensional array represents a virtual node, and each column represents the serial number of the redundant shard. For example, for a 20TB volume, a total of 800 virtual nodes are required, so the placement table has 800 rows, V = 800; select a stripe with K + M = 10 (each stripe has 10 stripe blocks), P = 10, then the virtual node placement table is an 800×10 two-dimensional table.
[0052] In addition, for the system parameters of the distributed file system, in this embodiment, the disks are considered to be grouped. For example, each server has 48 disks, which are divided into 4 groups. If there are 5 servers, then the first 12 disks of each server are the first group, the next 12 disks are the second group, and so on. That is, each group has 60 disks and is evenly distributed on each server. Then the system parameters include the server parameter server, the grouping parameter group, and the disk parameter disk.
[0053] In some embodiments, the system parameters may further include the weight group_weight of each disk group and the weight disk_weight of each disk. The weight of the disk group is the proportion of the capacity of each disk group in the entire distributed file system, and the weight of the disk is the proportion of the capacity of each disk group in the disk group.
[0054] S12. Loop independent variable vn.
[0055] In this step, vn represents the virtual node number. For example, the initial value of vn is 0 and the maximum value is 800.
[0056] S13. First lottery step: Based on the current virtual node number and system parameters, use the first lottery function to calculate the disk group where the current virtual node is located.
[0057] The first lottery function is gn = straw(group_count, 0, vn, crushmap_id, NULL, group_weight, NULL). It can be seen that the input parameters of the first lottery function include the disk group number (group_count), the server number (0), the current virtual node number (vn), the server id (crushmap_id), the disk id (NULL), the disk group weight (group_weight), and the validity (NULL). The output value is the disk group number (group_count), that is, the drawn disk group is used as the disk group where the current virtual node is located.
[0058] The problem description of the first lottery function is as follows: There is A and a group B. Find which one in group B A should be combined with. So for each combination of A and B, its hash value can be calculated as follows.
[0059] Yn = hash(A, Bn, argments)
[0060] Then, in the set of Yn, select the Bn corresponding to the largest value, and the selected combination for the lottery is found, that is: under the current parameter argmens, A should be combined with Bn.
[0061] In the above problem description, A is vn in this embodiment, B is the disk group in this embodiment, and the parameter argmens is the parameter input to the first lottery function in this embodiment. The mathematical details in the calculation process are that the 64-bit hash value calculated by performing xxhash on the input parameter needs to be converted into a decimal number between [0, 1), then take the natural logarithm, and finally divide by the weight (if necessary) to adjust the scaling factor.
[0062] In this embodiment, a validity array is also constructed to provide a validity mask (valid). The number of elements in this array is equal to the number of disks, representing the validity of each disk to mask item B that is not desired to be selected (the specific validity judgment rule will be described in detail in the subsequent steps).
[0063] After calculating the disk group where the current virtual node is located, calculate the disk where each redundant shard of the current virtual node is located.
[0064] S14. Loop independent variable sn.
[0065] In this step, sn represents the redundant shard serial number. For example, the initial value of sn is 0 and the maximum value is 10.
[0066] S15. Second lottery step: Based on the current virtual node serial number, the current redundant shard serial number, and system parameters, use the second lottery function to calculate the disk where the current redundant shard of the current virtual node is located.
[0067] The second lottery function is disk = straw(disk_count_in_group, 0, P * vn + sn, crushmap_id, disk_id_in_group, disk_weight_in_group, valid_in_group). It can be seen that the parameters input to the second lottery function include the disk serial number (disk_count_in_group), the server serial number (0), the current redundant shard serial number (P * vn + sn), the server id (crushmap_id), the disk id (disk_id_in_group), the disk weight (disk_weight_in_group), and the validity (valid_in_group). The output value is the disk serial number (disk_count_in_group), that is, the selected disk serves as the disk where the current redundant shard is located.
[0068] Assign a value to the virtual node placement table: data[vn][sn] = disk.
[0069] S16. Remove or mask the alternative disks that do not meet the disaster recovery conditions from the system parameters. Specifically, if the number of disks selected on a server exceeds the preset value, all the disks of that server are masked in the validity array; at the same time, the same disk cannot be selected twice by the same virtual node. If a disk has been selected by a redundant shard of the current virtual node, that disk is masked in the validity array.
[0070] Increment the current redundant shard sequence number by 1, sn + 1, and return to step S14 until the disks where each redundant shard is located are calculated (sn = P).
[0071] When sn = P, increment the current virtual node sequence number by 1, vn + 1, and return to step S12 until the disk groups where each virtual node is located are calculated (vn = V).
[0072] When vn = V, end the distributed file system index calculation.
[0073] After that, when a file needs to be stored, perform xxhash calculation on the 16 bytes of the Universally Unique Identifier (UUID) of the file to be stored, obtaining an integer, and then take the remainder of the integer calculated by xxhash with respect to the number of hash points to obtain the index value (an integer less than 800) of the virtual node (Vnode) corresponding to this file.
[0074] Determine the disk group to which the file belongs and the disks to which each redundant shard of the file belongs according to the index value of the virtual node corresponding to the file.
[0075] Embodiment 2:
[0076] As Figure 3 shown, in the embodiment of the present invention, the distributed file system index calculation method includes the following steps:
[0077] S21. Parameter acquisition step: Obtain the number of virtual nodes to be allocated, the number of redundant shards, and the system parameters of the distributed file system.
[0078] For example, the number of virtual nodes V = 800, the number of redundant shards P = 10, and the system parameters include server parameters server, grouping parameters group, disk parameters disk, the weight group_weight of each disk group, and the weight disk_weight of each disk.
[0079] S22. Loop independent variable vn.
[0080] In this step, vn represents the virtual node sequence number. For example, the initial value of vn is 0 and the maximum value is 800.
[0081] S23. First lottery step: Based on the current virtual node number and system parameters, use the first lottery function to calculate the disk group where the current virtual node is located.
[0082] The first lottery function is gn = straw(group_count, 0, vn, crushmap_id, NULL, group_weight, NULL), and its input parameters are the same as those of gn in Embodiment 1.
[0083] Construct validity arrays sec_valid and valid. The number of elements in the array sec_valid is equal to the number of disk groups, and each element corresponds to each disk group; the number of elements in the array valid is equal to the number of disks, and each element corresponds to each disk.
[0084] After calculating the disk group where the current virtual node is located, calculate the disks where each redundant shard of the current virtual node is located.
[0085] S24. Loop independent variable sn.
[0086] In this step, sn represents the redundant shard number. For example, the initial value of sn is 0, and the maximum value is 10.
[0087] S25. Second lottery step: Based on the current virtual node number, the current redundant shard number and system parameters, use the second lottery function to calculate the server where the current redundant shard of the current virtual node is located.
[0088] The second lottery function is sec_i = straw(server_count, vn, disk_number++, "sec", NULL, NULL, sec_valid). It can be seen that the input parameters of the second lottery function include the server number (server_count), the current virtual node number (vn), the number of disks of the server (disk_number++), the server title ("sec"), the server id (NULL), the server weight (NULL), and the server validity (sec_valid). The output value is the server number (server_count), that is, the selected server is used as the server where the current redundant shard is located.
[0089] S26. Third lottery step: Based on the current virtual node number and system parameters, use the third lottery function to calculate the disk where the current redundant shard of the current virtual node is located;
[0090] The third lottery function is disk = straw(disk_count_in_server, P*vn + sn, disk_number, "disk", sec_disks[sec_i], NULL, disk_valid). It can be seen that the input parameters of the third lottery function include the disk serial number in the server (disk_count_in_server), the current redundant shard serial number (P*vn + sn), the number of disks in the server (disk_number), the disk title ("disk"), the disk id in the server (sec_disks[sec_i]), the disk weight (NULL), and the disk validity (disk_valid). The output value is the disk serial number (disk_count_in_server), that is, the selected disk serves as the disk where the current redundant shard is located.
[0091] Assign a value to the virtual node placement table: data[vn][sn] = disk.
[0092] S27. Remove or mask the alternative disks that do not meet the disaster tolerance conditions from the system parameters. Specifically, if the number of disks selected on a server exceeds the preset value, all disks of that server are masked in the validity arrays sec_valid and valid; at the same time, the same disk cannot be selected twice by the same virtual node. If a disk has been selected by a redundant shard of the current virtual node, that disk is masked in the validity arrays sec_valid and valid.
[0093] Increment the current redundant shard serial number by 1, sn + 1, and return to step S24 until the disks where each redundant shard is located are calculated (sn = P).
[0094] When sn = P, increment the current virtual node serial number by 1, vn + 1, and return to step S22 until the disk groups where each virtual node is located are calculated (vn = V).
[0095] When vn = V, end the distributed file system index calculation.
[0096] After that, when a file needs to be stored, perform xxhash calculation on 16 bytes of the universal unique identifier of the file to be stored to obtain an integer, and then take the remainder of the integer calculated by xxhash with respect to the number of hash points to obtain the index value (an integer less than 800) of the virtual node (Vnode) corresponding to this file.
[0097] Determine the disk group to which the file belongs and the disks to which each redundant shard of the file belongs according to the index value of the virtual node corresponding to the file.
[0098] A distributed file system is a cluster service system for storing massive file data. It has the capabilities of data disaster recovery, machine disaster recovery, and self-governance of massive data. It usually provides users with a file system that conforms to the POSIX standard. Users can access the tree-like directory structure it provides through a general operating system and can store, read, and manage their files at any position in the directory tree. In addition, the distributed file system virtualizes a massive storage space (usually provided by the hard disks in a cluster composed of several servers) into a consistent and flat storage space, and can provide a logical storage space domain called a "volume" that far exceeds the capacity of a single disk. On the volume, users can establish their own file directory trees. A distributed file system can support many volumes for different users to use.
[0099] Each volume has a root directory with a fixed UUID. Using this UUID, the shard number, and combining with the disk ID (or first through the server ID) for hash lottery, the disk ID where this file shard is located can be calculated, thus achieving the indexing relationship from the file to the disk.
[0100] The above two embodiments calculate the disk group to be stored and the disk to be stored respectively through a two-step lottery algorithm, which can make the storage of files more uniform. When a disk is lost or added in a certain server, it can significantly reduce the amount of shard data migration and alleviate the problem of heavy workload on servers in the existing distributed file system.
[0101] The hash algorithm proposed in the embodiments of the present invention uses a lottery mechanism to implement consistent hashing, and combines the reality of distributed storage to propose a simplified disaster recovery domain layout diagram, and calculates a hash function (node placement diagram) with less migration amount. The service program can migrate file data according to this hash function. This article also provides a normalized proof of the metrics of this algorithm.
[0102] The performance of the distributed file system indexing calculation method provided by the above two embodiments is examined from six metrics below. The concept of "penalty rate" is used, that is, the ratio of the data migration caused by the loss (or addition) of a disk (or a server) to the disk (or server).
[0103] CNT: The variance of the distribution of hash placement points on the disk. The smaller the variance, the more uniform the distribution of virtual nodes on the disk, and the better the effect.
[0104] T: The time used to calculate the placement table. The smaller the time used, the better.
[0105] D1: Penalty rate for losing (or adding) a disk without considering the shard number. For example, the original distribution of Vnode = 0 is (disk.1, disk.2, disk3, disk.4), and after the change, it is (disk.1, disk.2, disk4, disk.5). Without considering the shard number, only the shard originally on disk.3 needs to be migrated to disk.5. If the shard number is considered, the shard originally on disk.3 needs to be migrated to disk.4, and the shard originally on disk.4 needs to be migrated to disk.5. On this stripe, the penalty rate of the latter (200%) is twice that of the former (100%).
[0106] D2: Penalty rate for losing (or adding) a disk considering the shard number.
[0107] E1: Penalty rate for losing (or adding) a server without considering the shard number.
[0108] E2: Penalty rate for losing (or adding) a server considering the shard number.
[0109] The above indicators can be easily calculated through the placement tables of 2 hash lottery algorithms. Figures 4 to 6 They are some calculated charts. The vertical coordinates of the 6 curves in each attached figure respectively correspond to the above 6 indicators, and the horizontal coordinate is the P value, with a value range between 3 and 24. The dark N line is the curve of Example 1, and the light M line is the curve of Example 2. Figure 4 It is a curve chart of 5 servers and 4 disk groups, Figure 5 It is a curve chart of 7 servers and 4 disk groups, Figure 6 It is a curve chart of 10 servers and 8 disk groups.
[0110] From these charts, it can be found that the calculation method provided by Example 2 performs better under different usage modes.
[0111] Regarding the selection of P (i.e., the erasure protection method: P = K + M), try not to choose an integer multiple of the number of P servers. In this case, the penalty rate is particularly high. When P divided by the number of servers has a remainder of 1 or 2, the penalty rate is particularly small.
[0112] The calculation method provided by Example 2 performs better than that of Example 1. Summarizing according to the phenomenon, it lies in that the less the effectiveness of the lottery algorithm is restricted, the better the lottery effect is. For the calculation method provided by Example 1, for the result of each lottery, compared with Example 2, there are greater restrictions. And Example 2 is divided into two lotteries (excluding determining the disk group). The first time is to determine the server where the shard is located, and the change of the server is less than that of the disk. This lottery shares part of the restrictions, resulting in fewer restrictions in the second lottery.
[0113] Embodiment 3:
[0114] As Figure 7 shown, an embodiment of the present invention provides a distributed file system index calculation device, including:
[0115] A parameter acquisition module 1, configured to acquire the number of virtual nodes to be allocated, the number of redundant shards, and the system parameters of the distributed file system.
[0116] A virtual node calculation module 2, configured to calculate the disk group where each virtual node is located by using a lottery algorithm.
[0117] A redundant shard calculation module 3, configured to calculate the disk where each redundant shard is located by using a lottery algorithm.
[0118] The distributed file system index calculation device provided by the embodiment of the present invention has the same technical features as the distributed file system index calculation method provided by the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.
[0119] Corresponding to the above method, an embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps of the above method are implemented.
[0120] Corresponding to the above method, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions. When the computer-executable instructions are called and run by the processor, the computer-executable instructions cause the processor to run the steps of the above method.
[0121] The device provided by the embodiment of the present invention can be specific hardware on the device or software or firmware installed on the device, etc. The implementation principle and the technical effects generated by the device provided by the embodiment of the present invention are the same as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the foregoing-described systems, devices, and units can all refer to the corresponding processes in the above method embodiment, and will not be repeated here.
[0122] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0123] For another example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0124] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0125] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0126] Finally, it should be noted that the above-mentioned embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention. All should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A distributed file system index calculation method, characterized in that, applied to a distributed file system, the method includes: A parameter acquisition step of acquiring the number of virtual nodes to be allocated, the number of redundant shards, and the system parameters of the distributed file system; A virtual node calculation step of calculating the disk group where each virtual node is located by using a lottery algorithm; A redundant shard calculation step of calculating the disk where each redundant shard is located by using a lottery algorithm, obtaining a hash function in the form of a two-dimensional array, where the number of rows of the two-dimensional array is equal to the number of virtual nodes and the number of columns is equal to the number of redundant shards, and is used to represent the disk where each redundant shard of each virtual node is located, so that when a file to be stored is stored, the disk where each redundant shard of the file to be stored should be deposited can be found from the hash function according to the hash index value of the file to be stored.
2. The method according to claim 1, characterized in that, the virtual node calculation step includes: A first lottery step: calculating the disk group where the current virtual node is located by using a first lottery function based on the current virtual node serial number and system parameters; Performing the redundant shard calculation step on the current virtual node; Adding 1 to the current virtual node serial number and returning to the first lottery step until the disk group where each virtual node is located is calculated.
3. The method according to claim 2, characterized in that, the redundant shard calculation step includes: A second lottery step: calculating the disk where the current redundant shard of the current virtual node is located by using a second lottery function based on the current virtual node serial number, the current redundant shard serial number, and system parameters; Adding 1 to the current redundant shard serial number and returning to the second lottery step until the disk where each redundant shard is located is calculated.
4. The method according to claim 2, characterized in that, the redundant shard calculation step includes: A second lottery step: calculating the server where the current redundant shard of the current virtual node is located by using a second lottery function based on the current virtual node serial number, the current redundant shard serial number, and system parameters; A third lottery step: calculating the disk where the current redundant shard of the current virtual node is located by using a third lottery function based on the current virtual node serial number and system parameters; Adding 1 to the current redundant shard serial number and returning to the second lottery step until the disk where each redundant shard is located is calculated.
5. The method according to claim 3 or 4, characterized in that, before adding 1 to the current redundant shard serial number and returning to the second lottery step, it further includes: Removing alternative disks that do not meet the disaster tolerance condition from the system parameters.
6. The method according to claim 1, characterized in that, the system parameters include the number of servers, the number of disks of each server, the number of disk groups, and the disk weight values of each disk.
7. The method according to claim 1, characterized in that, it further includes: Performing xxhash calculation on the universal unique identifier of the file to be stored to obtain the virtual node corresponding to the file; Determining the disk group to which the file belongs and the disks to which each redundant shard of the file belongs according to the virtual node corresponding to the file.
8. A distributed file system index calculation device, Characterized in that Applied to a distributed file system, the device includes: A parameter acquisition module, configured to acquire the number of virtual nodes to be allocated, the number of redundant shards, and the system parameters of the distributed file system; A virtual node calculation module, configured to calculate the disk group where each virtual node is located by using a lottery algorithm; A redundant shard calculation module, configured to calculate the disk where each redundant shard is located by using a lottery algorithm, and obtain a hash function in the form of a two-dimensional array, where the number of rows of the two-dimensional array is equal to the number of virtual nodes, and the number of columns is equal to the number of redundant shards, and is used to represent the disk where each redundant shard of each virtual node is located, so that when a file to be stored is stored, the disk where each redundant shard of the file to be stored should be stored can be found from the hash function according to the hash index value of the file to be stored.
9. An electronic device, including a memory and a processor, where a computer program that can run on the processor is stored in the memory, Characterized in that When the processor executes the computer program, the steps of the method described in any one of claims 1 to 7 above are implemented.
10. A computer-readable storage medium, Characterized in that The computer-readable storage medium stores machine-executable instructions, and when the computer-executable instructions are called and run by the processor, the computer-executable instructions cause the processor to run the method described in any one of claims 1 to 7.