Hard disk IO flow control method and device of distributed storage system, equipment and medium

By distinguishing service IO types in a distributed storage system and using buffer queues and preset scheduling scale tables for IO scheduling, the problem of backend IO occupies too much resources is solved, and the frontend business IO is priority treatment is achieved to ensure the optimal system performance.

CN119987675AActive Publication Date: 2025-05-13CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510111415.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

When processing background IO, a distributed storage system is prone to occupying too much system resources, resulting in an impact on customer business, and the prior art fails to effectively control the IO flow.

Method used

By obtaining service IO requests and storage pool capacity, the request is stored in the buffer queue according to the IO type identification, and using the preset IO scheduling ratio table, the scheduling ratio of the front-end and back-end service IO is divided into gears according to the storage pool capacity, and the scheduling ratio of the front-end service IO is adjusted to ensure that the front-end service IO has high priority.

Benefits of technology

It effectively reduces the impact of backend IO on the front-end business, ensures that the storage system provides optimal performance under various conditions, avoids system overload and congestion, and improves data read and write efficiency and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987675A_ABST
    Figure CN119987675A_ABST
Patent Text Reader

Abstract

The invention provides a hard disk IO flow control method and device of a distributed storage system, equipment and a medium. The method comprises the steps that multiple service IO requests and the storage pool capacity in the storage system are obtained; the plurality of service IO requests respectively carry respective IO type identifiers; the IO type identifier comprises a foreground service type identifier and a background service type identifier; respectively storing the plurality of service IO requests in corresponding buffer queues according to the plurality of IO type identifiers; according to the storage pool capacity and a preset IO scheduling proportion table, issuing the plurality of service IO requests stored in each buffer queue to a hard disk according to a preset IO scheduling proportion; the preset IO scheduling proportion table comprises the scheduling proportion occupied by each IO type corresponding to the storage pool capacity of different capacity gears, and when the storage pool capacity belongs to different capacity gears, the scheduling proportion occupied by the foreground service type identifier is higher than the scheduling proportion occupied by the background service type identifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of distributed storage technology, and in particular to a hard disk IO flow control method, device, equipment and medium for a distributed storage system. Background Art

[0002] In today's digital age, the amount of data is growing explosively, and enterprises and organizations need to store and process massive amounts of data. For example, Internet companies need to process hundreds of millions of user behavior data and video data every day. A large amount of data needs to be read, written and accessed frequently, resulting in extremely high concurrent IO (Input / Output) request pressure on distributed storage systems. If the IO flow is not controlled, the storage system is prone to overload and congestion, affecting the data read and write efficiency and system stability.

[0003] In related technologies, flow control restrictions are usually imposed on the services of different customers, but there is no control over the background IO of the storage system. For example, in scenarios such as fault reconstruction, high-capacity garbage collection, and data verification, these background flows may occupy too many system resources, causing the background flow to impact the customer's business. Summary of the invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a hard disk IO flow control method, device, equipment and medium for a distributed storage system, which avoids the problem of background traffic occupying too many system resources in different data processing scenarios.

[0005] In order to achieve the above objectives, the embodiments of the present disclosure provide the following technical solutions:

[0006] In a first aspect, an embodiment of the present disclosure provides a hard disk IO flow control method for a distributed storage system, the method comprising:

[0007] Acquire multiple service IO requests and storage pool capacity in the storage system; the multiple service IO requests respectively carry their own IO type identifiers; wherein the IO type identifiers include: a foreground service type identifier and a background service type identifier; the storage pool capacity is used to represent the used capacity of the storage pool;

[0008] According to the multiple IO type identifiers, the multiple service IO requests are stored in corresponding buffer queues respectively;

[0009] According to the storage pool capacity and the preset IO scheduling ratio table, multiple business IO requests stored in each buffer queue are sent to the hard disk according to the preset IO scheduling ratio; the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier.

[0010] As an optional implementation of the embodiment of the present disclosure, before sending multiple service IO requests in the buffer queue to the hard disk according to the preset IO scheduling ratio according to the storage pool capacity and the preset IO scheduling ratio table, the method further includes:

[0011] Divide the storage pool capacity into gears according to preset intervals;

[0012] According to the storage pool capacity of different gears, the scheduling ratio of each IO type is pre-set to build a preset IO scheduling ratio table.

[0013] As an optional implementation of the embodiment of the present disclosure, the scheduling ratio of each IO type is pre-set for the storage pool capacity of different gears, and a preset IO scheduling ratio table is constructed, including:

[0014] According to the storage pool capacity of different gears, determine the IO quantity corresponding to each IO type identifier;

[0015] According to the storage pool capacity of each gear and the IO quantity corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0016] As an optional implementation of the embodiment of the present disclosure, the scheduling ratio of each IO type is pre-set for the storage pool capacity of different gears, and a preset IO scheduling ratio table is constructed, which also includes:

[0017] Determine the IO ratio corresponding to each IO type identifier for storage pool capacities of different gears;

[0018] According to the storage pool capacity of each gear and the IO ratio corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0019] As an optional implementation of the embodiment of the present disclosure, the multiple business IO requests include foreground business IO requests and background business IO requests; the foreground business IO requests are triggered by the front end of the user interface or the application; the background business IO requests are triggered by the system background service or daemon process; the processing priority of the foreground business IO requests is higher than the processing priority of the background business IO requests.

[0020] As an optional implementation of the embodiment of the present disclosure, the background business type identifier includes: a background reconstruction identifier, a background aggregation identifier and a background garbage collection identifier; the scheduling ratio of the background reconstruction identifier is negatively correlated with the capacity level corresponding to the storage pool capacity; the background aggregation identifier is positively correlated with the capacity level corresponding to the storage pool capacity; the background garbage collection identifier is positively correlated with the capacity level corresponding to the storage pool capacity.

[0021] As an optional implementation of the embodiment of the present disclosure, the background service type identifier further includes: a background data verification identifier; and the scheduling ratio of the background data verification identifier is a preset value.

[0022] In a second aspect, an embodiment of the present disclosure provides a hard disk IO flow control device of a distributed storage system, comprising:

[0023] An acquisition module is used to acquire multiple service IO requests and storage pool capacity in the storage system; the multiple service IO requests respectively carry respective IO type identifiers; wherein the IO type identifiers include: a foreground service type identifier and a background service type identifier; the storage pool capacity is used to represent the used capacity of the storage pool;

[0024] A storage module, used for storing the multiple service IO requests in corresponding buffer queues respectively according to the multiple IO type identifiers;

[0025] The flow control module is used to send multiple business IO requests stored in each buffer queue to the hard disk according to the preset IO scheduling ratio based on the storage pool capacity and the preset IO scheduling ratio table; the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier.

[0026] As an optional implementation of the embodiment of the present disclosure, the device further includes a building module, and the building module includes:

[0027] A division unit, used to divide the storage pool capacity into gears according to preset intervals;

[0028] The construction unit is used to pre-set the scheduling ratio of each IO type according to the storage pool capacity of different gears, and construct a preset IO scheduling ratio table.

[0029] As an optional implementation of the embodiment of the present disclosure, the construction unit is used to:

[0030] According to the storage pool capacity of different gears, determine the IO quantity corresponding to each IO type identifier;

[0031] According to the storage pool capacity of each gear and the IO quantity corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0032] As an optional implementation of the embodiment of the present disclosure, the construction unit is also used for:

[0033] Determine the IO ratio corresponding to each IO type identifier for storage pool capacities of different gears;

[0034] According to the storage pool capacity of each gear and the IO ratio corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0035] As an optional implementation of the embodiment of the present disclosure, the multiple business IO requests include foreground business IO requests and background business IO requests; the foreground business IO requests are triggered by the front end of the user interface or the application; the background business IO requests are triggered by the system background service or daemon process; the processing priority of the foreground business IO requests is higher than the processing priority of the background business IO requests.

[0036] As an optional implementation of the embodiment of the present disclosure, the background business type identifier includes: a background reconstruction identifier, a background aggregation identifier and a background garbage collection identifier; the scheduling ratio of the background reconstruction identifier is negatively correlated with the capacity level corresponding to the storage pool capacity; the background aggregation identifier is positively correlated with the capacity level corresponding to the storage pool capacity; the background garbage collection identifier is positively correlated with the capacity level corresponding to the storage pool capacity.

[0037] As an optional implementation of the embodiment of the present disclosure, the background service type identifier further includes: a background data verification identifier; and the scheduling ratio of the background data verification identifier is a preset value.

[0038] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the hard disk IO flow control method of the distributed storage system described in the first aspect or any embodiment of the first aspect is implemented.

[0039] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the hard disk IO flow control method of the distributed storage system described in the first aspect or any embodiment of the first aspect is implemented.

[0040] The hard disk IO flow control method of the distributed storage system provided by the present invention obtains multiple business IO requests and the storage pool capacity in the storage system, the multiple business IO requests respectively carry their own IO type identifiers, and according to the multiple IO type identifiers, the multiple business requests are respectively stored in corresponding buffer queues, and according to the storage pool capacity and a preset IO scheduling ratio table, the multiple IO requests stored in each buffer queue are sent to the hard disk according to the preset IO scheduling ratio, wherein the IO type identifier includes: a foreground business type identifier and a background business type identifier; the storage pool capacity is used to represent the used capacity of the storage pool, the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier. By distinguishing IO types and setting reasonable scheduling ratios in advance, in order not to affect the user experience, the scheduling ratio of the foreground business type identifier is set to be higher than the scheduling ratio of the background business type identifier at any capacity level, so that the performance provided by the storage system to the outside world can always reach the best. In addition, by setting the buffer queue, the problem of excessive load on the storage system in a short period of time is reduced. Through flow control, the impact of background IO on the foreground business is effectively reduced, and under various conditions, the cluster can ensure that it provides users with the best performance in the current state. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0043] Figure 1 A schematic diagram of a flow chart of a hard disk IO flow control method of a distributed storage system in one embodiment;

[0044] Figure 2 A schematic diagram of the structure of a hard disk IO flow control device of a distributed storage system in one embodiment;

[0045] Figure 3 It is a schematic diagram of the structure of the electronic device described in the embodiment of the present disclosure. DETAILED DESCRIPTION

[0046] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0047] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0048] In the embodiments of the present disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present disclosure should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way. In addition, in the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "multiple" refers to two or more.

[0049] Internet companies need to process hundreds of millions of user behavior data, video data, etc. every day. A large amount of data needs to be frequently read, written and accessed, resulting in extremely high concurrent IO (Input / Output) request pressure on distributed storage systems. If the IO flow is not controlled, the storage system is prone to overload and congestion, affecting the data read and write efficiency and system stability.

[0050] In related technologies, flow control restrictions are usually imposed on the services of different customers, but there is no control over the background IO of the storage system. For example, in scenarios such as fault reconstruction, high-capacity garbage collection, and data verification, these background flows may occupy too many system resources, causing the background flow to impact the customer's business.

[0051] Based on the above problems, the present disclosure provides a hard disk IO flow control method for a distributed storage system. Figure 1 As shown, the method comprises the following steps:

[0052] S11. Obtain multiple service IO requests and storage pool capacity in the storage system.

[0053] The multiple service IO requests respectively carry their own IO type identifiers; the IO type identifiers include: a foreground service type identifier and a background service type identifier; the storage pool capacity is used to represent the used capacity of the storage pool.

[0054] An IO request is a request from an application or operating system to a storage device to read or write data.

[0055] Optionally, the multiple business IO requests include foreground business IO requests and background business IO requests; the foreground business IO requests are triggered by the user interface or the front end of the application; the background business IO requests are triggered by the system background service or daemon process; the processing priority of the foreground business IO requests is higher than the processing priority of the background business IO requests.

[0056] Specifically, the local storage module obtains multiple business IO requests and the used capacity of the storage pool of the storage system. The IO requests can be divided into foreground business and background business according to their sources. Foreground business IO requests are usually initiated by the user interface or the front-end part of the application, while background business IO requests are initiated by the system background service or daemon.

[0057] Foreground business IO requests usually need to respond quickly to ensure user experience. For example, when a user clicks a link in a browser, they expect the webpage to load quickly. Foreground business IO requests directly affect the user experience. If the response is slow or there is a freeze, it will lead to a poor user experience. The sources of background business IO requests include system garbage collection and defragmentation, data migration and optimization, data backup and recovery, and log information records generated by the system and applications. Background business IO requests usually do not require immediate response and can be performed when the system load is low, such as at night or on weekends; background business IO requests are often performed in batches, such as batch writing of log data, batch migration of large numbers of files, etc.

[0058] Foreground business IO requests are usually triggered directly by clients, such as operations performed by users in applications (such as opening files, saving documents, browsing web pages, etc.), which have high requirements for response time.

[0059] Background business IO requests are usually automatically executed by the system and do not directly rely on real-time user operations, such as data backup, logging, data migration, garbage collection, etc. They are used to maintain stable system operation, data security and performance optimization. They have relatively low real-time requirements and can be performed when the system load is low to avoid interfering with foreground business.

[0060] In terms of resource scheduling and processing priority, foreground business IO requests are usually given a higher priority to ensure user experience; background business IO requests may be arranged at a low priority to avoid affecting the foreground business.

[0061] The frequency and pattern of foreground business IO requests are related to the user's operating habits and task requirements, and have a high degree of randomness; the frequency and pattern of background business IO requests are usually determined by system configuration and planned tasks, and have a certain regularity.

[0062] S12. According to the multiple IO type identifiers, the multiple service IO requests are stored in corresponding buffer queues respectively.

[0063] Specifically, according to multiple IO type identifiers, multiple service IO requests are stored in corresponding buffer queues respectively. By setting the buffer queue, the problem of excessive load on the storage system in a short period of time is reduced.

[0064] S13. According to the storage pool capacity and the preset IO scheduling ratio table, multiple service IO requests stored in each buffer queue are sent to the hard disk according to the preset IO scheduling ratio.

[0065] Among them, the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier.

[0066] Specifically, the preset IO scheduling ratio table shows the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier. According to the storage pool capacity and the preset IO scheduling ratio table, multiple business IO requests stored in each buffer queue are sent to the hard disk according to the preset IO scheduling ratio.

[0067] The hard disk IO flow control method of the distributed storage system provided by the present invention obtains multiple business IO requests and the storage pool capacity in the storage system, the multiple business IO requests respectively carry their own IO type identifiers, and according to the multiple IO type identifiers, the multiple business requests are respectively stored in corresponding buffer queues, and according to the storage pool capacity and a preset IO scheduling ratio table, the multiple IO requests stored in each buffer queue are sent to the hard disk according to the preset IO scheduling ratio, wherein the IO type identifier includes: a foreground business type identifier and a background business type identifier; the storage pool capacity is used to represent the used capacity of the storage pool, the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier. By distinguishing IO types and setting reasonable scheduling ratios in advance, in order not to affect the user experience, the scheduling ratio of the front-end type identifier is set to be higher than the scheduling ratio of the back-end business type identifier at any capacity level, so that the performance provided by the storage system to the outside world can always reach the best. In addition, by setting the buffer queue, the problem of excessive load on the storage system in a short period of time is reduced. Through flow control, the impact of back-end IO on the front-end business is effectively reduced, and under various conditions, the cluster can ensure that it provides users with the best performance in the current state.

[0068] In one embodiment, before executing the above step S12 (sending multiple service IO requests in the buffer queue to the hard disk according to the preset IO scheduling ratio according to the storage pool capacity and the preset IO scheduling ratio table), the following steps may also be executed:

[0069] a. Divide the storage pool capacity into gears according to preset intervals.

[0070] The preset interval can be set according to actual conditions, for example, the preset interval can be set to 10%, 20%, 25%, etc. It should be noted that the preset interval can be an equal interval or an unequal interval.

[0071] Exemplarily, assuming that the preset interval is an equal interval and the preset interval is 20%, the storage pool capacity can be divided into the following gears according to the preset interval: 0-20%, 20%-40%, 40%-60%, 60%-80%, 80%-100%. Assuming that the preset interval is an unequal interval, the storage pool capacity can be divided into the following gears according to the preset interval: 0-50%, 50%-70%, 70%-80%, 80%-90%, 90%-100%.

[0072] b. According to the storage pool capacity of different gears, pre-set the scheduling ratio of each IO type and build a preset IO scheduling ratio table.

[0073] Specifically, for storage pool capacities of different gears, the scheduling ratio of each IO type is pre-set to construct a preset IO scheduling ratio table.

[0074] Optionally, the above step b (presetting the scheduling ratio of each IO type for storage pool capacities of different gears and constructing a preset IO scheduling ratio table) can be implemented as follows:

[0075] According to the storage pool capacity of different gears, determine the IO quantity corresponding to each IO type identifier;

[0076] According to the storage pool capacity of each gear and the IO quantity corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0077] Specifically, for storage pool capacities of different gears, the IO quantity corresponding to each IO type identifier is determined, and a preset IO scheduling ratio table is constructed according to the storage pool capacity of each gear and the IO quantity corresponding to each IO type identifier.

[0078] Exemplarily, referring to Table 1, Table 1 is a preset IO scheduling ratio table, and the IO type identifiers include: foreground business, background IO type 1, background IO type 2, background IO type 3 and background IO type 4. First, for the storage pool capacity of different gears, the IO quantity corresponding to each IO type identifier is determined, and then according to the storage pool capacity of each gear and the IO quantity corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0079] Table 1

[0080] Capacity gear Front desk business Background IO type 1 Background IO type 2 Background IO type 3 Background IO type 4 <50% 10 4 1 1 1 50%-70% 10 4 1 2 2 70%-80% 10 3 1 3 4 80%-90% 10 2 1 5 6 >90% 10 2 1 6 6

[0081] Optionally, the above step b (pre-setting the scheduling ratio of each IO type for storage pool capacities of different gears and constructing a preset IO scheduling ratio table) can also be implemented in the following way:

[0082] Determine the IO ratio corresponding to each IO type identifier for storage pool capacities of different gears;

[0083] According to the storage pool capacity of each gear and the IO ratio corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0084] Specifically, for storage pool capacities of different gears, the IO ratios corresponding to the IO type identifiers are determined, and a preset IO scheduling ratio table is constructed according to the storage pool capacities of each gear and the IO ratios corresponding to the IO type identifiers.

[0085] Exemplarily, referring to Table 2, Table 2 is another preset IO scheduling ratio table, and the IO type identifiers include: foreground business, background IO type 1, background IO type 2, background IO type 3 and background IO type 4. First, for the storage pool capacity of different gears, the IO ratio corresponding to each IO type identifier is determined, and then according to the storage pool capacity of each gear and the IO ratio corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0086] Table 2

[0087] Capacity gear Front desk business Background IO type 1 Background IO type 2 Background IO type 3 Background IO type 4 <50% 90% 2% 1% 4% 3% 50%-70% 90% 2% 1% 4% 3% 70%-80% 80% 3% 1% 6% 10% 80%-90% 80% 3% 1% 6% 10% >90% 75% 3% 2% 10% 10%

[0088] Optionally, the background business type identifier includes: a background reconstruction identifier, a background aggregation identifier and a background garbage collection identifier; the scheduling ratio of the background reconstruction identifier is negatively correlated with the capacity level corresponding to the storage pool capacity; the background aggregation identifier is positively correlated with the capacity level corresponding to the storage pool capacity; the background garbage collection identifier is positively correlated with the capacity level corresponding to the storage pool capacity.

[0089] Optionally, the background service type identifier further includes: a background data verification identifier; and the scheduling ratio of the background data verification identifier is a preset value.

[0090] Specifically, refer to Table 3, which is another preset IO scheduling ratio table. The IO type identifiers include: foreground business identifier, background reconstruction identifier, background data verification identifier, background aggregation identifier, and background garbage collection identifier. No matter what capacity level is, it is necessary to ensure that the foreground business accounts for the highest proportion. Background reconstruction is to write new data to the hard disk first, and then clean up the old incomplete data. Therefore, the storage pool capacity will increase during the data reconstruction. As the capacity increases, the IO ratio of background reconstruction should be reduced to prevent the storage pool from being filled up due to background reconstruction. Background aggregation and background garbage collection are both for deleting invalid data in the storage pool and freeing up storage pool space. Therefore, as the capacity increases, their proportions should be increased to allow them to release space as quickly as possible. Therefore, their proportions increase with the increase in capacity. Background data verification is to check for silent errors, but the data read from the disk will also be verified in the normal foreground business reading process. Therefore, data verification is not an urgent task compared to other types of IO, and its IO proportion can be maintained as long as it can be maintained.

[0091] Table 3

[0092] Capacity gear Front desk business Backend Refactoring Background data verification Background aggregation Background garbage collection <50% 10 4 1 1 1 50%-70% 10 4 1 2 2 70%-80% 10 3 1 3 4 80%-90% 10 2 1 5 6 >90% 10 2 1 6 6

[0093] For example, assume that the maximum number of IO requests supported by the disk per second is 100, and each IO request will generate an IO for all hard disks in the storage pool. In one case, the used capacity of the storage pool is 40%, and the foreground business is 100 times per second. At this time, due to a sudden failure of a disk, the storage pool is reconstructed, and the reconstruction also generates 100 disk IOs per second. If there is no flow control strategy, due to the emergence of reconstruction, the ratio of the 100 IOs on the disk may change from 100 IOs of foreground business to 50 IOs of foreground business plus 50 IOs of reconstruction, that is, due to the sudden occupation of disk resources by reconstruction, the foreground business drops from 100 per second to 50 per second; with flow control, the IO ratio of the disk will be controlled to 10:4 according to the above preset ratio, that is, there are 100*4 / (10+4)=28.5 reconstruction businesses on the hard disk per second, and the foreground business drops from 100 per second to 71 per second, effectively reducing the impact of reconstruction on the foreground business.

[0094] In another case, the used capacity of the storage pool is 85%, and the foreground business is 100 times per second. To prevent the storage pool capacity from being exhausted, background garbage collection and background aggregation will generate 100 disk IOs per second. If there is no flow control strategy, the IO ratio on the hard disk may be one-third each of the foreground business, background aggregation, and background garbage collection. At this time, the foreground business drops from 100 per second to 33 per second; after flow control, the IO ratio on the hard disk is: foreground business = 100*10 / (10+5+6) = 47.6, which effectively reduces the impact of these two background IOs on the foreground business.

[0095] Traffic control can effectively reduce the impact of background IO on front-end business, and ensure that the cluster provides users with the best performance in the current state under various conditions.

[0096] The hard disk IO flow control method of the distributed storage system provided by the present invention obtains multiple business IO requests and the storage pool capacity in the storage system, the multiple business IO requests respectively carry their own IO type identifiers, and according to the multiple IO type identifiers, the multiple business requests are respectively stored in corresponding buffer queues, and according to the storage pool capacity and a preset IO scheduling ratio table, the multiple IO requests stored in each buffer queue are sent to the hard disk according to the preset IO scheduling ratio, wherein the IO type identifier includes: a foreground business type identifier and a background business type identifier; the storage pool capacity is used to represent the used capacity of the storage pool, the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier. By distinguishing IO types and setting reasonable scheduling ratios in advance, in order not to affect the user experience, the scheduling ratio of the foreground business type identifier is set to be higher than the scheduling ratio of the background business type identifier at any capacity level, so that the performance provided by the storage system to the outside world can always reach the best. In addition, by setting the buffer queue, the problem of excessive load on the storage system in a short period of time is reduced. Through flow control, the impact of background IO on the foreground business is effectively reduced, and under various conditions, the cluster can ensure that it provides users with the best performance in the current state.

[0097] In one embodiment, Figure 2 As shown, a hard disk IO flow control device 200 of a distributed storage system is provided, comprising:

[0098] The acquisition module 210 is used to acquire multiple service IO requests and storage pool capacity in the storage system; the multiple service IO requests respectively carry their own IO type identifiers; wherein the IO type identifiers include: a foreground service type identifier and a background service type identifier; the storage pool capacity is used to represent the used capacity of the storage pool;

[0099] The storage module 220 is used to store the multiple service IO requests in corresponding buffer queues according to the multiple IO type identifiers;

[0100] The flow control module 230 is used to send multiple business IO requests stored in each buffer queue to the hard disk according to the preset IO scheduling ratio based on the storage pool capacity and the preset IO scheduling ratio table; the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier.

[0101] As an optional implementation of the embodiment of the present disclosure, the device further includes a building module, and the building module includes:

[0102] A division unit, used to divide the storage pool capacity into gears according to preset intervals;

[0103] The construction unit is used to pre-set the scheduling ratio of each IO type according to the storage pool capacity of different gears, and construct a preset IO scheduling ratio table.

[0104] As an optional implementation of the embodiment of the present disclosure, the construction unit is used to:

[0105] According to the storage pool capacity of different gears, determine the IO quantity corresponding to each IO type identifier;

[0106] According to the storage pool capacity of each gear and the IO quantity corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0107] As an optional implementation of the embodiment of the present disclosure, the construction unit is also used for:

[0108] Determine the IO ratio corresponding to each IO type identifier for storage pool capacities of different gears;

[0109] According to the storage pool capacity of each gear and the IO ratio corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

[0110] As an optional implementation of the embodiment of the present disclosure, the multiple business IO requests include foreground business IO requests and background business IO requests; the foreground business IO requests are triggered by the front end of the user interface or the application; the background business IO requests are triggered by the system background service or daemon process; the processing priority of the foreground business IO requests is higher than the processing priority of the background business IO requests.

[0111] As an optional implementation of the embodiment of the present disclosure, the background business type identifier includes: a background reconstruction identifier, a background aggregation identifier and a background garbage collection identifier; the scheduling ratio of the background reconstruction identifier is negatively correlated with the capacity level corresponding to the storage pool capacity; the background aggregation identifier is positively correlated with the capacity level corresponding to the storage pool capacity; the background garbage collection identifier is positively correlated with the capacity level corresponding to the storage pool capacity.

[0112] As an optional implementation of the embodiment of the present disclosure, the background service type identifier further includes: a background data verification identifier; and the scheduling ratio of the background data verification identifier is a preset value.

[0113] Applying the embodiments of the present disclosure, the hard disk IO flow control device of the distributed storage system provided by the present disclosure obtains multiple business IO requests and the storage pool capacity in the storage system, the multiple business IO requests respectively carry their own IO type identifiers, and according to the multiple IO type identifiers, the multiple business requests are respectively stored in corresponding buffer queues, and according to the storage pool capacity and a preset IO scheduling ratio table, the multiple IO requests stored in each buffer queue are sent to the hard disk according to the preset IO scheduling ratio, wherein the IO type identifier includes: a foreground business type identifier and a background business type identifier; the storage pool capacity is used to represent the used capacity of the storage pool, and the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and, when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier. By distinguishing IO types and setting reasonable scheduling ratios in advance, in order not to affect the user experience, the scheduling ratio of the foreground business type identifier is set to be higher than the scheduling ratio of the background business type identifier at any capacity level, so that the performance provided by the storage system to the outside world can always reach the best. In addition, by setting the buffer queue, the problem of excessive load on the storage system in a short period of time is reduced. Through flow control, the impact of background IO on the foreground business is effectively reduced, and under various conditions, the cluster can ensure that it provides users with the best performance in the current state.

[0114] For the specific definition of the hard disk IO flow control device of the distributed storage system, please refer to the definition of the hard disk IO flow control method of the distributed storage system mentioned above, which will not be repeated here. Each module in the hard disk IO flow control device of the above-mentioned distributed storage system can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor of the electronic device in the form of hardware, or can be stored in the processor of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0115] The present disclosure also provides an electronic device, Figure 3 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 3As shown, the electronic device provided in this embodiment includes: a memory 31 and a processor 32, the memory 31 is used to store a computer program; the processor 32 is used to execute the steps executed in any embodiment of the hard disk IO flow control method of the distributed storage system provided by the above method embodiment when calling the computer program. The electronic device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the computer program is executed by the processor, a hard disk IO flow control method of a distributed storage system is implemented. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the housing of the computer device, or an external keyboard, touchpad or mouse.

[0116] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present disclosure, and does not constitute a limitation on the computer device to which the scheme of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0117] In one embodiment, the hard disk IO flow control device of the distributed storage system provided by the present disclosure can be implemented in the form of a computer program. Figure 3 The memory of the electronic device may store various program modules of the hard disk IO flow control device constituting the distributed storage system of the electronic device, for example, Figure 2 The acquisition module 210, storage module 220 and flow control module 230 shown in the figure. The computer program composed of each program module enables the processor to execute the steps of the hard disk IO flow control method of the distributed storage system of the electronic device of each embodiment of the present disclosure described in this specification.

[0118] The embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the hard disk IO flow control method of the distributed storage system provided by the above method embodiment is implemented.

[0119] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0120] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0121] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0122] Computer readable media include permanent and non-permanent, removable and non-removable storage media. Storage media can be implemented by any method or technology to store information, and the information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0123] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0124] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A hard disk IO flow control method for a distributed storage system, characterized in that: The method comprises: Acquire multiple service IO requests and storage pool capacity in the storage system; the multiple service IO requests respectively carry their own IO type identifiers; wherein the IO type identifiers include: a foreground service type identifier and a background service type identifier; the storage pool capacity is used to represent the used capacity of the storage pool; According to the multiple IO type identifiers, the multiple service IO requests are stored in corresponding buffer queues respectively; According to the storage pool capacity and the preset IO scheduling ratio table, multiple business IO requests stored in each buffer queue are sent to the hard disk according to the preset IO scheduling ratio; the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier.

2. The method according to claim 1, characterized in that According to the storage pool capacity and the preset IO scheduling ratio table, before sending the multiple service IO requests in the buffer queue to the hard disk according to the preset IO scheduling ratio, the method further includes: Divide the storage pool capacity into gears according to preset intervals; According to the storage pool capacity of different gears, the scheduling ratio of each IO type is pre-set to build a preset IO scheduling ratio table.

3. The method according to claim 2, characterized in that The scheduling ratio of each IO type is pre-set for the storage pool capacity of different gears, and a preset IO scheduling ratio table is constructed, including: According to the storage pool capacity of different gears, determine the IO quantity corresponding to each IO type identifier; According to the storage pool capacity of each gear and the IO quantity corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

4. The method according to claim 2, characterized in that: The method of presetting the scheduling ratio of each IO type for the storage pool capacity of different gears and constructing a preset IO scheduling ratio table also includes: Determine the IO ratio corresponding to each IO type identifier for storage pool capacities of different gears; According to the storage pool capacity of each gear and the IO ratio corresponding to each IO type identifier, a preset IO scheduling ratio table is constructed.

5. The method according to claim 1, characterized in that The multiple business IO requests include foreground business IO requests and background business IO requests; the foreground business IO requests are triggered by the user interface or the front end of the application; the background business IO requests are triggered by the system background service or daemon process; the processing priority of the foreground business IO requests is higher than the processing priority of the background business IO requests.

6. The method according to claim 1, characterized in that The background business type identifier includes: a background reconstruction identifier, a background aggregation identifier and a background garbage collection identifier; the scheduling ratio of the background reconstruction identifier is negatively correlated with the capacity level corresponding to the storage pool capacity; the background aggregation identifier is positively correlated with the capacity level corresponding to the storage pool capacity; the background garbage collection identifier is positively correlated with the capacity level corresponding to the storage pool capacity.

7. The method according to claim 6, characterized in that The background service type identifier also includes: a background data verification identifier; the scheduling ratio of the background data verification identifier is a preset value.

8. A hard disk IO flow control device for a distributed storage system, characterized in that: include: An acquisition module is used to obtain multiple business IO requests and storage pool capacity in the storage system; The multiple service IO requests respectively carry their own IO type identifiers; wherein the IO type identifiers include: a foreground service type identifier and a background service type identifier; the storage pool capacity is used to represent the used capacity of the storage pool; A storage module, used for storing the multiple service IO requests in corresponding buffer queues respectively according to the multiple IO type identifiers; The flow control module is used to send multiple business IO requests stored in each buffer queue to the hard disk according to the preset IO scheduling ratio based on the storage pool capacity and the preset IO scheduling ratio table; the preset IO scheduling ratio table includes the scheduling ratio of each IO type corresponding to the storage pool capacity at different capacity levels, and when the storage pool capacity belongs to different capacity levels, the scheduling ratio of the foreground business type identifier is higher than the scheduling ratio of the background business type identifier.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the hard disk IO flow control method of the distributed storage system described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the hard disk IO flow control method of the distributed storage system described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Flow control method and system, electronic equipment and storage medium

    CN110099012A

  • Data storage method, device and system

    CN114363640A

  • Distributed storage service request processing method and device and distributed storage system

    CN118677913A

  • I / O request scheduling method and device, electronic equipment and storage medium

    CN119088530A

  • System cache processing method and device, electronic equipment and storage medium

    CN119166351A