A method and device for realizing capacity balancing based on disk grouping
By dividing disks according to disk water level in a distributed system and selecting storage disks according to the water level threshold range, the disk capacity imbalance caused by the polling algorithm is solved, and capacity equalization and high-performance storage effects are achieved.
Patent Information
- Application Number
- CN202110937379.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-08-16
AI Technical Summary
In existing distributed systems, the polling algorithm does not depend on the current capacity and response speed of each node when storing data, resulting in unbalanced disk capacity, affecting the availability and performance of the system.
By dividing the disk into several disk groups according to the disk water level, each group corresponds to a water level threshold range. Disk groups are selected in the order of the water level threshold range from low to high for storage, ensuring that the selected disk is in an available state, and recalculate and adjust the disk grouping after the storage request is completed.
It realizes balanced distribution of disk capacity, rationally utilizes storage space, improves system availability and performance, and ensures that the disk water level difference is within the allowable range.
Smart Images

Figure CN113791893B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of storage balancing, and in particular relates to a method and device for realizing capacity balancing based on disk grouping. Background Art
[0002] Disk water level refers to disk usage.
[0003] In a distributed system, in order to achieve high performance, high concurrency, and high availability of the system, capacity balancing is designed in the architecture, which is an important feature of the distributed system. The quality of capacity balancing directly affects the availability of the entire system. A distributed file system has a certain number of disks. How to select disks to store user data requires an algorithm to control the water level difference of each disk within a certain range, so as to achieve capacity balancing of the entire distributed system. If the water level of a disk itself is very high, while other disks are not in use, if the IO request still falls on the disk with the high water level, it will undoubtedly greatly affect the use of the entire storage space.
[0004] Some existing distributed systems use polling algorithms. When a client has an IO request, the files to be stored are stored on each disk in turn, regardless of the current capacity and response speed of each node. At this time, due to the different file sizes, the water level of each disk is different, and the polling algorithm easily leads to unbalanced disk capacity.
[0005] This is a shortcoming of the prior art. Therefore, in view of the above-mentioned defects in the prior art, it is very necessary to provide a method and device for achieving capacity balancing based on disk grouping. Summary of the invention
[0006] The above-mentioned distributed system in the prior art needs to control the water level of each disk to achieve capacity balance, while the existing polling algorithm is not based on the current capacity and response speed of each node, which easily leads to the defect of disk capacity imbalance. The present invention provides a method and device for achieving capacity balance based on disk grouping to solve the above technical problems.
[0007] In a first aspect, the present invention provides a method for achieving capacity balancing based on disk grouping, comprising the following steps:
[0008] S1. Divide the disks into several disk groups according to the disk water level, and each disk group corresponds to a disk water level threshold range;
[0009] S2. After receiving the storage request, the distributed system selects disks in the disk group for storage in the order of the disk water level threshold range from low to high, and ensures that the selected disks are in an available state until the storage request task is completed;
[0010] S3. After completing the storage request task, calculate the water level of each disk again and re-divide the disk groups.
[0011] Furthermore, the specific steps of step S1 are as follows:
[0012] S11. Set the corresponding number of disk water level threshold ranges according to the number of groups to be grouped;
[0013] S12. Obtain the current water level of each disk, and mark the water level status of each disk according to the water level threshold range of each disk;
[0014] S13. Put the disk IDs with the same water level status into the same container data structure. The disk water level is the disk usage rate. The disk grouping is based on the disk water level threshold. The water level threshold of each disk group is a range. When the distributed system is initialized, the water level of each disk is obtained, and the disk IDs in the same group are placed in the same container data structure.
[0015] Furthermore, the specific steps of step S2 are as follows:
[0016] S21. When the distributed system receives an IO request from the client, the storage controller determines the number of required disks, the number of existing disk groups, and the water level threshold range of each disk group;
[0017] S22. Arrange the water level threshold ranges of each disk group in order from low to high, and the water level status of the corresponding disk group in order from low to high;
[0018] S23. Determine whether the disk group with the lowest water level meets the required number of disks;
[0019] If so, select the required number of disks from the disk group with the lowest water level status for storage operation;
[0020] If not, continue to select disks from the disk group with a higher water level until the required number of disks is met and the client's IO request is completed. The disk usage rate in the disk group with a low water level is low, so the disk group with a low water level is selected first for operation. When the number of disks in the disk group with a low water level is insufficient, select disks from the disk group with a high water level.
[0021] Furthermore, the specific steps of step S23 are as follows:
[0022] S231. Locate the disk group with the lowest water level;
[0023] S232. Determine whether the location disk group is empty;
[0024] If yes, go to step S237;
[0025] If not, proceed to step S233;
[0026] S233. Randomly select a disk ID within the disk ID range of the located disk group, and the iterator points to the next disk ID;
[0027] S234. Determine whether the disk corresponding to the selected disk ID is available;
[0028] If yes, the disk corresponding to the selected disk ID is used as the storage disk, and the process goes to step S235;
[0029] If not, the disk corresponding to the selected disk ID is deleted from the positioning disk group, and the process goes to step S236;
[0030] S235. Determine whether the number of storage disks reaches the required number of disks;
[0031] If yes, write and delete data operations are performed on the storage disk according to the IO request of the client, and the process goes to step S3;
[0032] If not, proceed to step S236;
[0033] S236. Determine whether all disk IDs in the disk group have been located;
[0034] If yes, go to step S237;
[0035] If not, select the next disk ID in the located disk group and return to step S234;
[0036] S237. Determine whether the located disk group is the disk group with the highest water level;
[0037] If so, determine that the storage space is insufficient, report an error and return, and end;
[0038] If not, locate the disk group with a higher water level, and return to step S232. Use the iterator to point to the disk ID in the same disk group to ensure that the located disk ID is different to avoid repeated selection of the same disk; when selecting a storage disk, it is necessary to verify the availability of the located disk to prevent the selected storage disk from being unable to store data when in use.
[0039] Furthermore, the method further comprises the following steps:
[0040] S4. The storage controller determines whether a disk has been restored from an unavailable state to an available state;
[0041] If yes, obtain the water level of the disk whose available state is restored, compare it with the disk water level threshold range, add it to the disk group, and return to step S2;
[0042] If not, return to step S2.
[0043] Furthermore, the specific steps of step S3 are as follows:
[0044] S31. After completing the client IO request, the storage controller updates the current water level of each storage disk that has completed the write or delete data operation;
[0045] S32. Determine whether there is a faulty storage disk;
[0046] If yes, go to step S35;
[0047] If not, proceed to step S33;
[0048] S33. Re-mark the water level status of each storage disk according to the water level threshold range of each disk;
[0049] S34. Put the disk IDs of the same water level status into the same container data structure, regroup the storage disks, and end;
[0050] S35. Mark the failed storage disk as unavailable and delete it from the corresponding disk group. Since the water level of the selected storage disk has changed after the storage operation is performed on it, it needs to be updated. After the disk water level is updated, the disk group needs to be re-divided for the next corresponding client IO request.
[0051] In a second aspect, the present invention provides a device for achieving capacity balancing based on disk grouping, comprising:
[0052] The disk grouping module is used to divide the disks into several disk groups according to the disk water level. Each disk group corresponds to a disk water level threshold range.
[0053] The group storage module is used for selecting disks in the disk group for storage in the order of the disk water level threshold range from low to high after the distributed system receives the storage request, and ensures that the selected disks are in an available state until the storage request task is completed;
[0054] The regrouping module is used to recalculate the water level of each disk and re-divide the disk groups after completing the storage request task.
[0055] Furthermore, the disk grouping module includes:
[0056] A water level threshold setting unit, used to set a corresponding number of disk water level threshold ranges according to the number of disks to be grouped;
[0057] A water level status marking unit is used to obtain the current water level of each disk and mark the water level status of each disk according to the water level threshold range of each disk;
[0058] Disk grouping unit, used to put disk IDs with the same water level status into the same container data structure.
[0059] Furthermore, the group storage module includes:
[0060] The storage status acquisition unit is used for the storage controller to determine the required number of disks, the number of existing disk groups, and the water level threshold range of each disk group when the distributed system receives an IO request from the client;
[0061] The disk group arrangement unit is used to arrange the water level threshold ranges of each disk group in order from low to high, and the water level states of the corresponding disk groups are arranged in order from low to high;
[0062] A low water level group disk quantity determination unit is used to determine whether the disk group with the lowest water level status meets the required disk quantity;
[0063] A storage operation unit, configured to select a required number of disks from the disk group with the lowest water level status to perform storage operation when the disk group with the lowest water level status meets the required number of disks;
[0064] The disk continuing selection unit is used to continue selecting disks from the disk group with a higher water level state when the disk group with the lowest water level state does not meet the required number of disks, until the required number of disks is met, thereby completing the client's IO request.
[0065] Furthermore, the regrouping module includes:
[0066] The storage disk water level update unit is used to update the current water level of each storage disk that has completed the write or delete data operation after completing the client IO request;
[0067] A storage disk fault judgment unit, used to judge whether there is a faulty storage disk;
[0068] The storage disk water level status marking unit is used to re-mark the water level status of each storage disk according to the water level threshold range of each disk when there is a faulty storage disk;
[0069] The storage disk regrouping unit is used to put the disk IDs of the same water level status into the same container data structure, regroup the storage disks, and end;
[0070] The storage disk deletion unit is used to mark the failed storage disk as unavailable and delete it from the corresponding disk group when the storage disk fails.
[0071] The beneficial effects of the present invention are:
[0072] The device for achieving capacity balancing based on disk grouping provided by the present invention groups the disk water levels and sets priorities for the groups. Thus, in business usage scenarios, disks in low-water-level disk groups with high priorities are used for storage. This can balance disk usage, make rational use of storage space, and thereby improve system availability and performance, ensure that the water level differences of disks in a cluster are within an allowable range, and achieve capacity balancing.
[0073] In addition, the invention has a reliable design principle, a simple structure and a very broad application prospect.
[0074] It can be seen that compared with the prior art, the present invention has outstanding substantive features and significant progress, and the beneficial effects of its implementation are also obvious. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0076] Figure 1 The method flow diagram of the present invention for realizing capacity balancing based on disk grouping is as follows Figure 1 .
[0077] Figure 2 The method flow diagram of the present invention for realizing capacity balancing based on disk grouping is as follows Figure 2 .
[0078] Figure 3 The method flow diagram of the present invention for realizing capacity balancing based on disk grouping is as follows Figure 3 .
[0079] Figure 4 The figure is a schematic diagram of a device for realizing capacity balancing based on disk grouping according to the present invention.
[0080] In the figure, 1-disk grouping module; 1.1-water level threshold setting unit; 1.2-water level status marking unit; 1.3-disk grouping unit; 2-grouping storage module; 2.1-storage status acquisition unit; 2.2-disk grouping arrangement unit; 2.3-low water level grouping disk quantity judgment unit; 2.4-storage operation unit; 2.5-disk continue selection unit; 3-regrouping module; 3.1-storage disk water level update unit; 3.2-storage disk fault judgment unit; 3.3-storage disk water level status marking unit; 3.4-storage disk regrouping unit; 3.5-storage disk deletion unit. DETAILED DESCRIPTION
[0081] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0082] Embodiment 1:
[0083] like Figure 1 As shown, the present invention provides a method for achieving capacity balancing based on disk grouping, comprising the following steps:
[0084] S1. Divide the disks into several disk groups according to the disk water level, and each disk group corresponds to a disk water level threshold range;
[0085] S2. After receiving the storage request, the distributed system selects disks in the disk group for storage in the order of the disk water level threshold range from low to high, and ensures that the selected disks are in an available state until the storage request task is completed;
[0086] S3. After completing the storage request task, calculate the water level of each disk again and re-divide the disk groups.
[0087] Embodiment 2:
[0088] like Figure 1 and Figure 2 As shown, the present invention provides a method for achieving capacity balancing based on disk grouping, comprising the following steps:
[0089] S1. Divide the disks into several disk groups according to the disk water level. Each disk group corresponds to a disk water level threshold range. The specific steps are as follows:
[0090] S11. Set the corresponding number of disk water level threshold ranges according to the number of groups to be grouped;
[0091] S12. Obtain the current water level of each disk, and mark the water level status of each disk according to the water level threshold range of each disk;
[0092] S13. Put the disk IDs of the same water level status into the same container data structure;
[0093] S2. After receiving the storage request, the distributed system selects disks in the disk group for storage in the order of the disk water level threshold range from low to high, and ensures that the selected disks are in an available state until the storage request task is completed; the specific steps are as follows:
[0094] S21. When the distributed system receives an IO request from the client, the storage controller determines the number of required disks, the number of existing disk groups, and the water level threshold range of each disk group;
[0095] S22. Arrange the water level threshold ranges of each disk group in order from low to high, and the water level status of the corresponding disk group in order from low to high;
[0096] S23. Determine whether the disk group with the lowest water level meets the required number of disks;
[0097] If so, select the required number of disks from the disk group with the lowest water level status for storage operation;
[0098] If not, continue to select disks from the disk group with a higher water level status until the required number of disks is met, completing the client's IO request;
[0099] S3. After completing the storage request task, calculate the water level of each disk again and re-divide the disk groups; the specific steps are as follows:
[0100] S31. After completing the client IO request, the storage controller updates the current water level of each storage disk that has completed the write or delete data operation;
[0101] S32. Determine whether there is a faulty storage disk;
[0102] If yes, go to step S35;
[0103] If not, proceed to step S33;
[0104] S33. Re-mark the water level status of each storage disk according to the water level threshold range of each disk;
[0105] S34. Put the disk IDs of the same water level status into the same container data structure, regroup the storage disks, and end;
[0106] S35. Mark the failed storage disk as unavailable and delete it from the corresponding disk group.
[0107] Embodiment 3:
[0108] like Figure 1 , Figure 2 and Figure 3 As shown, the present invention provides a method for achieving capacity balancing based on disk grouping, comprising the following steps:
[0109] S1. Divide the disks into several disk groups according to the disk water level. Each disk group corresponds to a disk water level threshold range. The specific steps are as follows:
[0110] S11. Set the corresponding number of disk water level threshold ranges according to the number of groups to be grouped;
[0111] S12. Obtain the current water level of each disk, and mark the water level status of each disk according to the water level threshold range of each disk;
[0112] S13. Put the disk IDs of the same water level status into the same container data structure;
[0113] S2. After receiving the storage request, the distributed system selects disks in the disk group for storage in the order of the disk water level threshold range from low to high, and ensures that the selected disks are in an available state until the storage request task is completed; the specific steps are as follows:
[0114] S21. When the distributed system receives an IO request from the client, the storage controller determines the number of required disks, the number of existing disk groups, and the water level threshold range of each disk group;
[0115] S22. Arrange the water level threshold ranges of each disk group in order from low to high, and the water level status of the corresponding disk group in order from low to high;
[0116] S23. Determine whether the disk group with the lowest water level meets the required number of disks;
[0117] If so, select the required number of disks from the disk group with the lowest water level status for storage operation;
[0118] If not, continue to select disks from the disk group with a higher water level status until the required number of disks is met, completing the client's IO request;
[0119] The specific steps of step S23 are as follows:
[0120] S231. Locate the disk group with the lowest water level;
[0121] S232. Determine whether the location disk group is empty;
[0122] If yes, go to step S237;
[0123] If not, proceed to step S233;
[0124] S233. Randomly select a disk ID within the disk ID range of the located disk group, and the iterator points to the next disk ID;
[0125] S234. Determine whether the disk corresponding to the selected disk ID is available;
[0126] If yes, the disk corresponding to the selected disk ID is used as the storage disk, and the process goes to step S235;
[0127] If not, the disk corresponding to the selected disk ID is deleted from the positioning disk group, and the process goes to step S236;
[0128] S235. Determine whether the number of storage disks reaches the required number of disks;
[0129] If yes, write and delete data operations are performed on the storage disk according to the IO request of the client, and the process goes to step S3;
[0130] If not, proceed to step S236;
[0131] S236. Determine whether all disk IDs in the disk group have been located;
[0132] If yes, go to step S237;
[0133] If not, select the next disk ID in the located disk group and return to step S234;
[0134] S237. Determine whether the located disk group is the disk group with the highest water level;
[0135] If so, determine that the storage space is insufficient, report an error and return, and end;
[0136] If not, locate the disk group with a higher water level status and return to step S232;
[0137] S3. After completing the storage request task, calculate the water level of each disk again and re-divide the disk groups; the specific steps are as follows:
[0138] S31. After completing the client IO request, the storage controller updates the current water level of each storage disk that has completed the write or delete data operation;
[0139] S32. Determine whether there is a faulty storage disk;
[0140] If yes, go to step S35;
[0141] If not, proceed to step S33;
[0142] S33. Re-mark the water level status of each storage disk according to the water level threshold range of each disk;
[0143] S34. Put the disk IDs of the same water level status into the same container data structure, regroup the storage disks, and end;
[0144] S35. Mark the failed storage disk as unavailable and delete it from the corresponding disk group.
[0145] In the above embodiment 3, the steps are also included:
[0146] S4. The storage controller determines whether a disk has been restored from an unavailable state to an available state;
[0147] If yes, obtain the water level of the disk whose available state is restored, compare it with the disk water level threshold range, add it to the disk group, and return to step S2;
[0148] If not, return to step S2.
[0149] In the above embodiment 3, taking the number of disks to be grouped as three as an example, the water level states corresponding to the disk groups are low water level state, medium water level state and high water level state;
[0150] When the distributed system receives an IO request from a client, the storage controller determines the number of disks required;
[0151] First, check whether the disk group in the low water level state is empty. If not, select a disk ID from the disk group in the low water level state, and the iterator points to the next disk ID. Then determine whether the positioning disk corresponding to the selected disk ID is available. If available, it means that the positioning disk is working normally in the distributed system cluster. If not available, it means that the positioning disk is not working in the distributed system cluster.
[0152] If the located disk is unavailable, delete the located disk from the disk group in the low water mark state, continue to randomly select the next disk from the disk group in the low water mark state, and move the iterator to point to the next disk ID; if the located disk is available, set the located disk as the storage disk, and determine whether the storage disk reaches the required number of disks. If the required number of disks is not reached, continue to select another disk in the disk group in the low water mark state;
[0153] If the number of disks in the disk group in the low water level state is 0, randomly select a disk ID from the disk group in the medium water level state, and the iterator points to the next disk ID; then determine whether the positioning disk corresponding to the selected disk ID is available; if the positioning disk is not available, delete it from the disk group in the medium water level state, continue to select the next disk from the disk group in the medium water level state, and move the iterator to point to the next disk ID; if the positioning disk is available, set the positioning disk as a storage disk, and determine whether the storage disk has reached the required number of disks. If it has not reached the required number of disks, continue to select another disk in the disk group in the medium water level state;
[0154] If the number of disks in the disk group in the medium water level state is 0, randomly select a disk ID from the disk group in the high water level state, and the iterator points to the next disk ID; then determine whether the positioning disk corresponding to the selected disk ID is available; if the positioning disk is not available, delete it from the disk group in the high water level state, continue to select the next disk from the disk group in the high water level state, and move the iterator to point to the next disk ID; if the positioning disk is available, set the positioning disk as a storage disk, and determine whether the storage disk has reached the required number of disks. If it has not reached the required number of disks, continue to select another disk in the disk group in the high water level state;
[0155] If the number of disks in the high watermark disk group is 0, it means that the storage space is insufficient and an error is returned;
[0156] When the data is written successfully, the real-time water level of the storage disk needs to be compared with the threshold range of the low water level state, the medium water level state, and the high water level state, and the storage disk needs to be re-divided into three groups;
[0157] At the same time, the disks in the distributed system that are in an unavailable state must be monitored. When these disks recover from an unavailable state to an available state, their water level states must be compared with the three water level threshold ranges in a timely manner and classified into corresponding disk groups.
[0158] Embodiment 4:
[0159] like Figure 3 As shown, the present invention provides a device for realizing capacity balancing based on disk grouping, comprising:
[0160] The disk grouping module 1 is used to divide the disk into several disk groups according to the level of the disk water level, and each disk group corresponds to a disk water level threshold range; the disk grouping module 1 includes:
[0161] The water level threshold setting unit 1.1 is used to set a corresponding number of disk water level threshold ranges according to the number of disks to be grouped;
[0162] The water level status marking unit 1.2 is used to obtain the current water level of each disk and mark the water level status of each disk according to the water level threshold range of each disk;
[0163] The disk grouping unit 1.3 is used to put the disk IDs of the same water level status into the same container data structure;
[0164] The group storage module 2 is used for selecting disks in the disk group for storage in the order of the disk water level threshold range from low to high after the distributed system receives the storage request, and ensuring that the selected disks are in an available state until the storage request task is completed; the group storage module 2 includes:
[0165] The storage status acquisition unit 2.1 is used for the storage controller to determine the required number of disks, the number of existing disk groups, and the water level threshold range of each disk group when the distributed system receives an IO request from the client;
[0166] The disk group arrangement unit 2.2 is used to arrange the water level threshold ranges of each disk group in order from low to high, and the water level states of the corresponding disk groups are arranged in order from low to high;
[0167] The low water level group disk quantity judgment unit 2.3 is used to judge whether the disk group with the lowest water level status meets the required disk quantity;
[0168] The storage operation unit 2.4 is used for selecting the required number of disks from the disk group with the lowest water level status to perform storage operation when the disk group with the lowest water level status meets the required number of disks;
[0169] The disk continuing selection unit 2.5 is used to continue selecting disks from the disk group with a higher water level state when the disk group with the lowest water level state does not meet the required number of disks, until the required number of disks is met, thereby completing the client's IO request; the disk continuing selection unit 2.5 includes:
[0170] The lowest water level disk group positioning subunit is used to locate the disk group with the lowest water level;
[0171] The empty disk group determination subunit is used to determine whether the positioning disk group is empty;
[0172] The disk ID selection subunit is used to randomly select a disk ID within the disk ID range of the located disk group when the located disk group is not empty, and the iterator points to the next disk ID;
[0173] The disk availability judgment subunit is used to judge whether the disk corresponding to the selected disk ID is available;
[0174] The storage disk setting subunit is used for, when the disk corresponding to the selected disk ID is available, using the disk corresponding to the selected disk ID as the storage disk;
[0175] The disk deletion subunit is used to delete the disk corresponding to the selected disk ID from the positioning disk group when the disk corresponding to the selected disk ID is unavailable;
[0176] The disk quantity determination subunit is used to determine whether the number of storage disks reaches the required number of disks;
[0177] The storage operation subunit is used to write and delete data on the storage disk according to the IO request of the client when the number of storage disks reaches the required number of disks;
[0178] The disk ID positioning judgment subunit is used to judge whether all disk IDs in the positioning disk group have been positioned when the positioning disk group is empty or the number of storage disks does not reach the required number of disks;
[0179] The next disk positioning subunit is used to select the next disk ID in the positioning disk group when the disk ID in the positioning disk group has not been completely positioned;
[0180] The disk group highest water level judgment subunit is used to judge whether the located disk group is the disk group with the highest water level state when the disk ID in the located disk group is located;
[0181] The storage space insufficient determination subunit is used to determine that the storage space is insufficient when the located disk group is the disk group with the highest water level, report an error and return, and then end;
[0182] The next disk group positioning subunit is used to position the disk group with a higher water level state when the positioned disk group is not the disk group with the highest water level state;
[0183] The regrouping module 3 is used to recalculate the water level of each disk and re-divide the disk groups after completing the storage request task; the regrouping module 3 includes:
[0184] The storage disk water level update unit 3.1 is used for the storage controller to update the current water level of each storage disk that has completed the writing or deleting of data operation after completing the client IO request;
[0185] A storage disk fault judgment unit 3.2, used to judge whether there is a faulty storage disk;
[0186] The storage disk water level status marking unit 3.3 is used to re-mark the water level status of each storage disk according to the water level threshold range of each disk when there is a faulty storage disk;
[0187] The storage disk regrouping unit 3.4 is used to put the disk IDs of the same water level status into the same container data structure, regroup the storage disks, and end;
[0188] The storage disk deletion unit 3.5 is used to mark the failed storage disk as unavailable and delete it from the corresponding disk group when the storage disk fails.
[0189] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all of these shall be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for achieving capacity balancing based on disk grouping, characterized in that: The steps include: S1. Divide the disks into several disk groups according to the disk water level, and each disk group corresponds to a disk water level threshold range; S2. After receiving the storage request, the distributed system selects disks in the disk group for storage in the order of the disk water level threshold range from low to high, and ensures that the selected disks are in an available state until the storage request task is completed; The specific steps of step S2 are as follows: S21. When the distributed system receives an IO request from the client, the storage controller determines the number of required disks, the number of existing disk groups, and the water level threshold range of each disk group; S22. Arrange the water level threshold ranges of each disk group in order from low to high, and the water level status of the corresponding disk group in order from low to high; S23. Determine whether the disk group with the lowest water level meets the required number of disks; If so, select the required number of disks from the disk group with the lowest water level status for storage operation; If not, continue to select disks from the disk group with a higher water level status until the required number of disks is met, completing the client's IO request; The specific steps of step S23 are as follows: S231. Locate the disk group with the lowest water level; S232. Determine whether the location disk group is empty; If yes, go to step S237; If not, proceed to step S233; S233. Randomly select a disk ID within the disk ID range of the located disk group, and the iterator points to the next disk ID; S234. Determine whether the disk corresponding to the selected disk ID is available; If yes, the disk corresponding to the selected disk ID is used as the storage disk, and the process goes to step S235; If not, the disk corresponding to the selected disk ID is deleted from the positioning disk group, and the process goes to step S236; S235. Determine whether the number of storage disks reaches the required number of disks; If yes, write and delete data operations are performed on the storage disk according to the IO request of the client, and the process goes to step S3; If not, proceed to step S236; S236. Determine whether all disk IDs in the disk group have been located; If yes, go to step S237; If not, select the next disk ID in the located disk group and return to step S234; S237. Determine whether the located disk group is the disk group with the highest water level; If so, determine that the storage space is insufficient, report an error and return, and end; If not, locate the disk group with a higher water level status and return to step S232; S3. After completing the storage request task, calculate the water level of each disk again and re-divide the disk groups.
2. The method for realizing capacity balancing based on disk grouping according to claim 1, characterized in that: The specific steps of step S1 are as follows: S11. Set the corresponding number of disk water level threshold ranges according to the number of groups to be grouped; S12. Obtain the current water level of each disk, and mark the water level status of each disk according to the water level threshold range of each disk; S13. Put the disk IDs with the same water level status into the same container data structure.
3. The method for realizing capacity balancing based on disk grouping according to claim 1, characterized in that: The following steps are also included: S4. The storage controller determines whether a disk has been restored from an unavailable state to an available state; If yes, obtain the water level of the disk whose available status is restored, compare it with the disk water level threshold range, add it to the disk group, and return to step S2; If not, return to step S2.
4. The method for realizing capacity balancing based on disk grouping according to claim 1, characterized in that: The specific steps of step S3 are as follows: S31. After completing the client IO request, the storage controller updates the current water level of each storage disk that has completed the write or delete data operation; S32. Determine whether there is a faulty storage disk; If yes, go to step S35; If not, proceed to step S33; S33. Re-mark the water level status of each storage disk according to the water level threshold range of each disk; S34. Put the disk IDs of the same water level status into the same container data structure, regroup the storage disks, and end; S35. Mark the failed storage disk as unavailable and delete it from the corresponding disk group.
5. A device for achieving capacity balancing based on disk grouping, characterized in that: include: A disk grouping module (1) is used to divide the disk into a number of disk groups according to the level of the disk water level, each disk group corresponds to a disk water level threshold range; The group storage module (2) is used for selecting disks in the disk group for storage in the order of the disk water level threshold range from low to high after the distributed system receives the storage request, and ensuring that the selected disks are in an available state until the storage request task is completed; The group storage module (2) comprises: The storage status acquisition unit (2.1) is used for the storage controller to determine the required number of disks, the number of existing disk groups and the water level threshold range of each disk group when the distributed system receives an IO request from the client; The disk group arrangement unit (2.2) is used to arrange the water level threshold ranges of each disk group in order from low to high, and the water level states of the corresponding disk groups are arranged in order from low to high; A low water level group disk quantity judgment unit (2.3), used to judge whether the disk group with the lowest water level status meets the required disk quantity; A storage operation unit (2.4) is used for selecting a required number of disks from the disk group with the lowest water level status to perform a data writing or deleting operation when the disk group with the lowest water level status meets the required number of disks; The disk continuing selection unit (2.5) is used to continue selecting disks from the disk group with a higher water level state when the disk group with the lowest water level state does not meet the required number of disks, until the required number of disks is met, thereby completing the client's IO request; The disk continuation selection unit (2.5) includes: The lowest water level disk group positioning subunit is used to locate the disk group with the lowest water level; The empty disk group determination subunit is used to determine whether the positioning disk group is empty; The disk ID selection subunit is used to randomly select a disk ID within the disk ID range of the located disk group when the located disk group is not empty, and the iterator points to the next disk ID; The disk availability judgment subunit is used to judge whether the disk corresponding to the selected disk ID is available; The storage disk setting subunit is used for, when the disk corresponding to the selected disk ID is available, using the disk corresponding to the selected disk ID as the storage disk; The disk deletion subunit is used to delete the disk corresponding to the selected disk ID from the positioning disk group when the disk corresponding to the selected disk ID is unavailable; The disk quantity determination subunit is used to determine whether the number of storage disks reaches the required number of disks; The storage operation subunit is used to write and delete data on the storage disk according to the IO request of the client when the number of storage disks reaches the required number of disks; The disk ID positioning judgment subunit is used to judge whether all disk IDs in the positioning disk group have been positioned when the positioning disk group is empty or the number of storage disks does not reach the required number of disks; The next disk positioning subunit is used to select the next disk ID in the positioning disk group when the disk ID in the positioning disk group has not been completely positioned; The disk group highest water level judgment subunit is used to judge whether the located disk group is the disk group with the highest water level state when the disk ID in the located disk group is located; The storage space insufficient determination subunit is used to determine that the storage space is insufficient when the located disk group is the disk group with the highest water level, report an error and return, and then end; The next disk group positioning subunit is used to position the disk group with a higher water level state when the positioned disk group is not the disk group with the highest water level state; The regrouping module (3) is used to recalculate the water level of each disk and re-divide the disk groups after completing the storage request task.
6. The device for realizing capacity balancing based on disk grouping according to claim 5, characterized in that: The disk grouping module (1) comprises: A water level threshold setting unit (1.1) is used to set a corresponding number of disk water level threshold ranges according to the number of disks to be grouped; The water level status marking unit (1.2) is used to obtain the current water level of each disk and mark the water level status of each disk according to the water level threshold range of each disk; The disk grouping unit (1.3) is used to put disk IDs with the same water level status into the same container data structure.
7. The device for realizing capacity balancing based on disk grouping according to claim 5, characterized in that: The regrouping module (3) includes: The storage disk water level update unit (3.1) is used for the storage controller to update the current water level of each storage disk that has completed the writing or deleting of data operation after completing the client IO request; A storage disk fault judgment unit (3.2), used for judging whether there is a faulty storage disk; The storage disk water level status marking unit (3.3) is used to re-mark the water level status of each storage disk according to the water level threshold range of each disk when there is a faulty storage disk; The storage disk regrouping unit (3.4) is used to put the disk IDs of the same water level status into the same container data structure, regroup the storage disks, and end; The storage disk deletion unit (3.5) is used to mark the failed storage disk as unavailable and delete it from the corresponding disk group when the storage disk fails.
Citation Information
Patent Citations
Load balancing storage method and system in cloud computing platform
CN103929454A