A QoS Method and System for Data Recovery Based on a Distributed Storage System
By performing IO testing and fitting formula calculations in a distributed storage system, and adjusting the data recovery speed in combination with the monitor, the problem of data recovery having a large impact on IO and long-term risks in the system is solved, and dynamic control and fast and secure recovery are achieved.
Patent Information
- Application Number
- CN202310024104.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-01-09
AI Technical Summary
During the data recovery process, a large number of read and write operations are generated within the distributed storage system and have a great impact on external IO. The existing technology cannot accurately control it, and it still recovers at a fixed speed when there is no IO, resulting in the system being in a risk state for a long time.
By performing read and write tests of various IO sizes in a distributed storage system, compute the fitting formula, and set it to the OSD module, the monitor adjusts the data recovery speed in real time, and dynamically adjusts the recovery speed according to the IO deviation coefficient and the expected value.
It realizes that when the service IO is affected exceeds or is lower than the expected range, the recovery speed is automatically adjusted, the impact on the service IO is reduced, and the system stability is improved.
Smart Images

Figure CN116010168B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data recovery, and in particular to a QoS method and system for data recovery based on a distributed storage system. Background Art
[0002] With the rapid development of information technology, more and more business applications require the support of storage systems. Among them, a distributed storage system that interconnects and associates scattered and independent storage devices through a network and provides storage services as a whole has been widely used;
[0003] Inside a distributed storage system, due to the independence of storage devices, generally multi-copy technology or erasure code technology is used to ensure data security, allowing a certain number of storage devices to be damaged within a certain range without affecting data security. In implementation, when faults that affect data security occur, such as hard disk damage, storage node downtime, and abnormal storage node network, etc., the data stored on the faulty device will be recovered through multi-copy technology or erasure code technology and written to the currently healthy storage device until all data recovery is completed. This process is called data recovery; during the data recovery process, a large number of read and write operations are generated inside the distributed storage system, which will affect the IO outside the distributed storage system. The faster the data recovery speed, the greater the impact on IO. Currently, most methods limit the impact on IO by fixing the data recovery speed. However, the impact of data recovery on IO is not precisely controllable, and data recovery still proceeds at a fixed speed when there is no IO, which will make the data in the distributed storage system remain in a risk state for a long time. Therefore, based on the distributed storage system, the present invention designs a QoS method and system for data recovery based on a distributed storage system, which can control the impact of data recovery on it within the expected range when there is IO, while taking into account the data recovery speed; and recover data as soon as possible when there is no IO, so that the data in the distributed storage system reaches a safe state as early as possible. Summary of the Invention
[0004] The technical problem to be solved by the present invention is that during the data recovery process, a large number of read and write operations are generated inside the distributed storage system, which will affect the I / O outside the distributed storage system. The faster the data recovery speed, the greater the impact on I / O. Currently, most methods limit the impact on I / O by fixing the data recovery speed. However, the impact of data recovery on I / O is not precisely controllable, and data recovery still proceeds at a fixed speed when there is no I / O, which will keep the data in the distributed storage system in a risky state for a long time. The present invention provides a QoS method based on data recovery of a distributed storage system, and the present invention also provides a QoS system based on data recovery of a distributed storage system. Compared with the traditional fixed data recovery speed, when the impact of service I / O on data recovery exceeds the maximum value of the expected impact degree, the data recovery speed can be quickly reduced to minimize the impact of data recovery on service I / O; when the impact of service I / O on data recovery is lower than the minimum value of the expected impact degree or there is no service I / O, the data recovery speed is slowly increased to enable the distributed storage system to reach a safe state as soon as possible; when the impact of service I / O on data recovery is within the expected impact degree range, the data recovery speed is not adjusted, so as to solve the defects caused by the prior art.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] In a first aspect, a QoS method based on data recovery of a distributed storage system, which includes the following steps:
[0007] Step 1: Perform read and write tests of various different I / O sizes outside the distributed storage system and count their read and write time delays to obtain time delay data;
[0008] Step 2: Calculate the obtained time delay data according to the read and write types and I / O sizes to obtain fitting formulas of time delay and I / O size within multiple specific ranges: y = kx + b, where k and b are constants, x is the I / O size, and y is the time delay;
[0009] Step 3: Set the fitting formula into the parameters of the OSD module for reading and writing data from the hard disk. The OSD module is arranged in the distributed storage system and there are several of them;
[0010] Step 4: Calculate the expected time delay of each I / O through the fitting formula in the OSD module, divide the actual time delay counted when the I / O is completed by the expected time delay to obtain the I / O time delay deviation coefficient, and accumulate each I / O time delay deviation coefficient to obtain the deviation coefficient accumulation value;
[0011] Step 5: After reaching the reporting period, report the deviation coefficient accumulation value and the number of I / Os to the monitor;
[0012] Step 6: In the monitor, after reaching the QoS period, if data recovery is detected in the distributed storage system, find the OSD modules that are performing data recovery and the affected OSD modules, and summarize the accumulated deviation coefficient values and the number of I / Os reported by the above OSD modules;
[0013] If data recovery is not detected in the distributed storage system, wait for the start of the next QoS period and execute Step 6 again;
[0014] Step 7: Sum the difference between the accumulated deviation coefficient values reported in the last two reports and the difference between the number of I / Os reported in the last two reports for all the OSD modules in Step 6 to obtain the total I / O latency deviation coefficient and the total number of I / Os in the most recent reporting period;
[0015] Divide the total I / O latency deviation coefficient by the total number of I / Os to obtain the overall latency deviation coefficient;
[0016] Step 8: Compare the overall latency deviation coefficient with the maximum and minimum values of the set expected overall latency deviation coefficient;
[0017] If the overall latency deviation coefficient is greater than the maximum value of the expected overall latency deviation coefficient, send a command to significantly reduce the data recovery speed to the OSD module that calculates the overall latency deviation coefficient;
[0018] If the overall latency deviation coefficient is less than the minimum value of the expected overall latency deviation coefficient, send a command to slightly increase the data recovery speed to the OSD module that calculates the overall latency deviation coefficient;
[0019] If the overall latency deviation coefficient is between the minimum and maximum values of the expected overall latency deviation coefficient, no command is sent.
[0020] In the above QoS method based on data recovery in a distributed storage system: In Step 5: The reporting period is the process in which the OSD module sends the accumulated deviation coefficient value and the number of I / Os to the monitor at fixed time intervals.
[0021] In the above QoS method based on data recovery in a distributed storage system: In Step 6: The QoS period is the process in which the monitor checks whether there is data recovery in the distributed storage system at fixed time intervals and performs corresponding operations.
[0022] In the above QoS method based on data recovery in a distributed storage system: In Step 8: The maximum and minimum values of the expected overall latency deviation coefficient are set according to the expected impact of data recovery on I / O, representing the expected QoS effect.
[0023] In a second aspect, a QoS system for data recovery based on a distributed storage system, comprising: a read-write test module, a fitting module, an OSD module, a monitor, and a comparison module;
[0024] The read-write test module is used to perform read-write tests of multiple different IO sizes and count their read-write delays to obtain delay data;
[0025] The fitting module is used to calculate fitting formulas for delays and IO sizes within multiple specific ranges based on the read-write type and IO size: y = kx + b, where k and b are constants, x is the IO size, and y is the delay;
[0026] The OSD module is used to set the fitting formula into the parameters of the OSD module. The OSD module is arranged in the distributed storage system and there are several of them; it is also used to calculate the expected delay of each IO through the fitting formula, divide the actual delay counted when the IO is completed by the expected delay to obtain the IO delay deviation coefficient, and accumulate the IO delay deviation coefficients of each one to obtain the accumulated deviation coefficient value; it is also used to report the accumulated deviation coefficient value and the number of IOs to the monitor after reaching the reporting period;
[0027] In the monitor, after reaching the QoS period, if it is detected that there is data recovery in the distributed storage system, find the OSD module that is performing data recovery and the OSD modules it affects, and summarize the accumulated deviation coefficient values and the number of IOs reported by the above OSD modules; if it is not detected that the distributed storage system has data recovery, wait for the start of the next QoS period; it is also used to sum the difference between the accumulated deviation coefficient values reported in the last two reports and the difference between the number of IOs reported in the last two reports for all the above OSD modules respectively to obtain the total IO delay deviation coefficient and the total number of IOs in the most recent reporting period; divide the total IO delay deviation coefficient by the total number of IOs to obtain the overall delay deviation coefficient;
[0028] The comparison module is used to compare the overall delay deviation coefficient with the maximum and minimum values of the set expected overall delay deviation coefficient;
[0029] If the overall delay deviation coefficient is greater than the maximum value of the expected overall delay deviation coefficient, send a command to significantly reduce the data recovery speed to the OSD module that calculates the overall delay deviation coefficient;
[0030] If the overall delay deviation coefficient is less than the minimum value of the expected overall delay deviation coefficient, send a command to slightly increase the data recovery speed to the OSD module that calculates the overall delay deviation coefficient;
[0031] If the overall time delay deviation coefficient is between the minimum value and the maximum value of the expected overall time delay deviation coefficient, no instruction is sent.
[0032] In a third aspect, a chip includes: a processor configured to call and run a computer program from a memory, such that a device installed with the chip performs: the method according to any one of the first aspect.
[0033] Compared with the traditional fixed data recovery speed, when the impact of data recovery on service I / O exceeds the maximum value of the expected impact degree, the data recovery speed can be quickly reduced to minimize the impact of data recovery on service I / O; when the impact of data recovery on service I / O is lower than the minimum value of the expected impact degree or there is no service I / O, the data recovery speed can be slowly increased to enable the distributed storage system to reach a safe state as soon as possible; when the impact of data recovery on service I / O is within the expected impact degree range, the data recovery speed is not adjusted.
[0034] According to the technical solutions provided by the QoS method and system for data recovery in a distributed storage system of the present invention, the following technical effects are achieved:
[0035] By pre-testing the distributed storage system to obtain a fitting formula, subsequent QoS control of data recovery can be closer to the actual situation of the distributed storage system, achieving better effects;
[0036] The data recovery speed can be automatically adjusted according to the service I / O situation. When there is service I / O, the quality of service I / O is preferentially guaranteed while taking into account the data recovery speed, and the impact degree on service I / O is controlled within the expected range; when there is no service I / O, data recovery is performed at a relatively fast speed to enable the data in the distributed storage system to reach a safe state as soon as possible;
[0037] The QoS control based on data recovery is implemented in the form of a plugin and can be enabled or disabled as needed without affecting any existing functions;
[0038] The minimum value and the maximum value of the expected overall time delay deviation coefficient can be dynamically modified to adjust the QoS effect at any time. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of a QoS method for data recovery in a distributed storage system according to the present invention. DETAILED DESCRIPTION
[0040] In order to make the technical means, creative features, achieved purposes, and effects of the invention easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to specific drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0041] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0042] It should be noted that the structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have technical essential significance. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention.
[0043] At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" cited in this specification are only for the convenience of clear narration, and are not used to limit the scope in which the present invention can be implemented. The change or adjustment of their relative relationships, without substantial change in the technical content, should also be regarded as the scope in which the present invention can be implemented.
[0044] Glossary:
[0045] QoS: Quality of Service, quality of service;
[0046] IO: Input Output, input and output;
[0047] OSD module: Object Storage Device, object storage device;
[0048] In the first aspect, as Figure 1 shown, a QoS method for data recovery based on a distributed storage system, which includes the following steps:
[0049] Step 1: Perform read and write tests of various different IO sizes outside the distributed storage system and count their read and write time delays to obtain time delay data;
[0050] Step 2: Calculate the time delay data according to the read and write types and IO sizes to obtain fitting formulas for the time delay and IO size within multiple specific ranges: y = kx + b, where k and b are constants, x is the IO size, and y is the time delay;
[0051] Step 3: Set the fitting formula into the parameters of the OSD module for reading and writing data from the hard disk. The OSD module is set in the distributed storage system and there are several of them;
[0052] Step 4: Calculate the expected latency of each IO through the fitting formula in the OSD module, divide the actual latency counted when the IO is completed by the expected latency to obtain the IO latency deviation coefficient, and accumulate each IO latency deviation coefficient to obtain the accumulated deviation coefficient value;
[0053] Step 5: After reaching the reporting period, report the accumulated deviation coefficient value and the number of IOs to the monitor;
[0054] Step 6: In the monitor, after reaching the QoS period, if it is detected that there is data recovery in the distributed storage system, find the OSD module that is performing data recovery and the OSD modules it affects, and summarize the accumulated deviation coefficient values and the number of IOs reported by the above OSD modules;
[0055] If it is not detected that there is data recovery in the distributed storage system, wait for the start of the next QoS period and execute Step 6 again;
[0056] Step 7: Sum the difference between the accumulated deviation coefficient values reported in the last two times and the difference between the number of IOs reported in the last two times for all OSD modules in Step 6 respectively to obtain the total IO latency deviation coefficient and the total number of IOs in the most recent reporting period;
[0057] Divide the total IO latency deviation coefficient by the total number of IOs to obtain the overall latency deviation coefficient;
[0058] The accumulated deviation coefficient value and the number of IOs reported by the OSD module are two continuously accumulating values, and single-time data cannot play a role alone. Therefore, use the difference between the accumulated deviation coefficient values reported in the last two times and the difference between the number of IOs to represent the IO latency deviation coefficient and the number of IOs of a single OSD module in the most recent reporting period. Sum these two items of data for all OSD modules in Step 6 to obtain the total IO latency deviation coefficient and the total number of IOs in the most recent reporting period.
[0059] Step 8: Compare the overall latency deviation coefficient with the maximum and minimum values of the set expected overall latency deviation coefficient;
[0060] If the overall latency deviation coefficient is greater than the maximum value of the expected overall latency deviation coefficient, send a command to significantly reduce the data recovery speed to the OSD module that calculates the overall latency deviation coefficient;
[0061] If the overall latency deviation coefficient is less than the minimum value of the expected overall latency deviation coefficient, send a command to slightly increase the data recovery speed to the OSD module that calculates the overall latency deviation coefficient;
[0062] If the overall latency deviation coefficient is between the minimum and maximum values of the expected overall latency deviation coefficient, no command is sent.
[0063] The above QoS method for data recovery based on a distributed storage system, wherein: in step 5: the reporting period is the process in which the OSD module sends the accumulated deviation coefficient and the number of I / Os to the monitor at fixed time intervals.
[0064] The above QoS method for data recovery based on a distributed storage system, wherein: in step 6: the QoS period is the process in which the monitor checks whether there is data recovery in the distributed storage system at fixed time intervals and performs corresponding operations.
[0065] The above QoS method for data recovery based on a distributed storage system, wherein: in step 8: the maximum and minimum values of the expected overall delay deviation coefficient are set according to the expected impact degree of data recovery on I / Os, indicating the expected QoS effect.
[0066] In a second aspect, a QoS system for data recovery based on a distributed storage system, which includes a read-write test module, a fitting module, an OSD module, a monitor, and a comparison module;
[0067] The read-write test module is used to perform read-write tests of various different I / O sizes and count their read-write delays to obtain delay data;
[0068] The fitting module is used to calculate, according to the read-write type and I / O size, the delay data to obtain fitting formulas for the delay and I / O size within multiple specific ranges: y = kx + b, where k and b are constants, x is the I / O size, and y is the delay;
[0069] The OSD module is used to set the fitting formula into the parameters of the OSD module. The OSD module is set in the distributed storage system and there are several of them; it is also used to calculate the expected delay of each I / O through the fitting formula, divide the actual delay counted when the I / O is completed by the expected delay to obtain the I / O delay deviation coefficient, and accumulate the I / O delay deviation coefficients of each I / O to obtain the accumulated deviation coefficient value; it is also used to report the accumulated deviation coefficient value and the number of I / Os to the monitor after reaching the reporting period;
[0070] In the monitor, after reaching the QoS period, if it is detected that there is data recovery in the distributed storage system, find out the OSD module that is performing data recovery and the OSD modules it affects, and summarize the accumulated deviation coefficient value and the number of I / Os reported by the above OSD modules; if it is not detected that there is data recovery in the distributed storage system, wait for the start of the next QoS period; it is also used to sum the difference between the accumulated deviation coefficient values reported by the above all OSD modules in the most recent two reports and the difference between the number of I / Os reported in the most recent two reports respectively, to obtain the total I / O delay deviation coefficient and the total number of I / Os in the most recent reporting period; divide the total I / O delay deviation coefficient by the total number of I / Os to obtain the overall delay deviation coefficient;
[0071] The comparison module is used to compare the overall delay deviation coefficient with the maximum and minimum values of the set expected overall delay deviation coefficient;
[0072] If the overall delay deviation coefficient is greater than the maximum value of the expected overall delay deviation coefficient, a command to significantly reduce the data recovery speed is sent to the OSD module that calculates the overall delay deviation coefficient;
[0073] If the overall delay deviation coefficient is less than the minimum value of the expected overall delay deviation coefficient, a command to slightly increase the data recovery speed is sent to the OSD module that calculates the overall delay deviation coefficient;
[0074] If the overall delay deviation coefficient is between the minimum and maximum values of the expected overall delay deviation coefficient, no instruction is sent.
[0075] In a third aspect, a chip includes: a processor for calling and running a computer program from a memory, such that a device installed with the chip executes: the method according to any one of the first aspect.
[0076] In summary, a QoS method and system for data recovery based on a distributed storage system according to the present invention, compared with the traditional fixed data recovery speed, can quickly reduce the data recovery speed when the impact of data recovery on service I / O exceeds the maximum value of the expected impact degree, and reduce the impact of data recovery on service I / O; when the impact of data recovery on service I / O is lower than the minimum value of the expected impact degree or there is no service I / O, slowly increase the data recovery speed to enable the distributed storage system to reach a safe state as soon as possible; when the impact of data recovery on service I / O is within the expected impact degree range, the data recovery speed is not adjusted.
[0077] The specific embodiments of the invention have been described above. It should be understood that the invention is not limited to the above specific embodiments, and the devices and structures not described in detail therein should be understood to be implemented in a common manner in the art; those skilled in the art can make various deformations or modifications within the scope of the claims, make several simple deductions, deformations or substitutions, which do not affect the essence of the invention.
Claims
1. A QoS method for data recovery based on a distributed storage system, characterized in that It includes the following steps: Step 1: Perform read and write tests with multiple different IO sizes outside the distributed storage system and count their read and write latency to obtain latency data; Step 2: Calculate the fitting formula of latency and IO size within multiple specific ranges from the latency data according to the read and write types and IO sizes: y = kx + b, where k and b are constants, x is the IO size, and y is the latency; Step 3: Set the fitting formula into the parameters of the OSD module for reading and writing data from the hard disk. The OSD module is set in the distributed storage system and there are several of them; Step 4: Calculate the expected latency of each IO through the fitting formula in the OSD module, divide the actual latency counted when the IO is completed by the expected latency to obtain the IO latency deviation coefficient, and accumulate each IO latency deviation coefficient to obtain the accumulated deviation coefficient value; Step 5: After reaching the reporting period, report the accumulated deviation coefficient value and the number of IOs to the monitor; Step 6: In the monitor, after reaching the QoS period, if it is detected that there is data recovery in the distributed storage system, find out the OSD module that is performing data recovery and the OSD modules it affects, and summarize the accumulated deviation coefficient values and the number of IOs reported by the above OSD modules; If it is not detected that the distributed storage system has data recovery, wait for the start of the next QoS period and execute Step 6 again; Step 7: Sum the difference between the accumulated deviation coefficient values reported by all the OSD modules in Step 6 for the most recent two times and the difference between the number of IOs reported for the most recent two times respectively to obtain the total IO latency deviation coefficient and the total number of IOs in the most recent reporting period; Divide the total IO latency deviation coefficient by the total number of IOs to obtain the overall latency deviation coefficient; Step 8: Compare the overall latency deviation coefficient with the maximum and minimum values of the set expected overall latency deviation coefficient; If the overall latency deviation coefficient is greater than the maximum value of the expected overall latency deviation coefficient, send a command to significantly reduce the data recovery speed to the OSD module that calculates the overall latency deviation coefficient; If the overall latency deviation coefficient is less than the minimum value of the expected overall latency deviation coefficient, send a command to slightly increase the data recovery speed to the OSD module that calculates the overall latency deviation coefficient; If the overall latency deviation coefficient is between the minimum and maximum values of the expected overall latency deviation coefficient, no command is sent.
2. The QoS method for data recovery based on a distributed storage system according to claim 1, wherein: In Step 5: The reporting period is the process in which the OSD module sends the accumulated deviation coefficient value and the number of IOs to the monitor at fixed time intervals.
3. The QoS method for data recovery in a distributed storage system according to claim 2, characterized in that: In Step 6: The QoS period is the process in which the monitor checks whether there is data recovery in the distributed storage system at fixed time intervals and performs corresponding operations.
4. The QoS method for data recovery based on a distributed storage system according to claim 2, characterized in that: In Step 8: The maximum and minimum values of the expected overall latency deviation coefficient are set according to the expected impact degree of data recovery on IOs, indicating the expected QoS effect.
5. A QoS system based on data recovery of a distributed storage system, characterized in that: It includes a read and write test module, a fitting module, an OSD module, a monitor, and a comparison module; The read-write test module is used to perform read-write tests with multiple different I / O sizes and count their read-write latency to obtain latency data; The fitting module is used to calculate the fitting formula of latency and I / O size within multiple specific ranges based on the read-write type and I / O size: y = kx + b, where k and b are constants, x is the I / O size, and y is the latency; The OSD module is used to set the fitting formula into the parameters of the OSD module. There are several OSD modules set in the distributed storage system; it is also used to calculate the expected latency of each I / O through the fitting formula, divide the actual latency counted when the I / O is completed by the expected latency to obtain the I / O latency deviation coefficient, and accumulate each I / O latency deviation coefficient to obtain the accumulated deviation coefficient value; it is also used to report the accumulated deviation coefficient value and the number of I / Os to the monitor after reaching the reporting period; In the monitor, after reaching the QoS period, if it is detected that there is data recovery in the distributed storage system, find the OSD module that is performing data recovery and the OSD modules affected by it, and summarize the accumulated deviation coefficient values and the number of I / Os reported by the above OSD modules; if it is not detected that there is data recovery in the distributed storage system, wait for the start of the next QoS period; it is also used to sum the differences between the accumulated deviation coefficient values reported in the last two reports and the differences between the numbers of I / Os reported in the last two reports for all the above OSD modules respectively, to obtain the total I / O latency deviation coefficient and the total number of I / Os in the most recent reporting period; divide the total I / O latency deviation coefficient by the total number of I / Os to obtain the overall latency deviation coefficient; The comparison module is used to compare the overall latency deviation coefficient with the maximum and minimum values of the set expected overall latency deviation coefficient; If the overall latency deviation coefficient is greater than the maximum value of the expected overall latency deviation coefficient, send a command to significantly reduce the data recovery speed to the OSD module that calculates the overall latency deviation coefficient; If the overall latency deviation coefficient is less than the minimum value of the expected overall latency deviation coefficient, send a command to slightly increase the data recovery speed to the OSD module that calculates the overall latency deviation coefficient; If the overall latency deviation coefficient is between the minimum and maximum values of the expected overall latency deviation coefficient, no command is sent.
6. A chip, characterized in that, It includes: A processor, which is used to call and run a computer program from a memory, so that the device installed with the chip executes: the method according to any one of claims 1-4.
Citation Information
Patent Citations
Adaptive data recovery flow control method and device, electronic equipment and storage medium
CN108804039A
Method and device for monitoring IO time delay of distributed file system and storage medium
CN110287158A