Data recovery method and device, storage node, storage medium and computer program product

By controlling the execution time of the data repair task based on the parameter pair determined by the probability and cumulative frequency curves on the storage node of the distributed storage system, the problem of data repair tasks occupying system resources is solved and the normal operation of the front-end task is ensured.

CN119960696APending Publication Date: 2025-05-09CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510072605.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When the storage nodes in the distributed storage system perform data repair tasks in the background, the system resource utilization rate is too high, affecting the normal operation of the front-end tasks.

Method used

By controlling the execution time of the data repair task based on the first parameter pair determined based on the first probability and the first accumulated frequency curve, ensuring that the data repair task does not run over the front-end task or limiting its performance impact on the front-end task during the superimposed run is within the set performance decay target.

Benefits of technology

Effectively schedule the execution time of data repair tasks, limit their impact on front-end tasks, ensure the normal operation of front-end tasks, and improve the utilization rate of system resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960696A_ABST
    Figure CN119960696A_ABST
Patent Text Reader

Abstract

The invention discloses a data recovery method and device, a storage node, a storage medium and a computer program product, and the method comprises the steps: the storage node in a distributed storage system determines a first parameter pair based on a first probability and a first cumulative frequency curve; the first probability represents the probability that the task completion duration of the front-end task of the storage node is increased when the storage node executes the data recovery task in the background; the first probability is determined based on a set performance degradation target; each coordinate point in the first cumulative frequency curve is used for describing the occurrence frequency of all idle time periods of which the duration is less than or equal to the duration corresponding to the coordinate point in one or more idle time periods in a historical statistical period; based on the first parameter pair, executing a first data recovery task; the first parameter pair is used for controlling the execution time of the first data recovery task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data repair method, device, storage node, storage medium and computer program product. Background Art

[0002] In the related art, storage nodes in a distributed storage system perform data repair tasks in the background to repair lost data. However, the background data repair tasks and the front-end tasks of the storage nodes are superimposed on each other for a long time, resulting in excessive utilization of system resources, which easily affects the normal operation of the front-end tasks. Summary of the invention

[0003] To solve related technical problems, embodiments of the present application provide a data repair method, device, storage node, storage medium and computer program product.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present application provides a data repair method, which is applied to a storage node in a distributed storage system. The method includes:

[0006] Based on the first probability and the first cumulative frequency curve, a first parameter pair is determined; the first probability represents: when the storage node performs a data repair task in the background, the probability of increasing the task completion time of the front-end task of the storage node; the first probability is determined based on the set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: in one or more idle time periods within the historical statistical period, the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point;

[0007] Based on the first parameter pair, a first data repair task is executed; the first parameter pair is used to control the execution time of the first data repair task.

[0008] In the above scheme, the first parameter pair is used to indicate a first duration and a second duration, the first duration represents the duration from the start to the start of execution of the first data repair task, and the second duration represents the duration of execution of the first data repair task.

[0009] In the above solution, determining the first parameter pair based on the first probability and the first cumulative frequency curve includes:

[0010] Based on the first probability, a first coordinate point and a second coordinate point are selected in the first cumulative frequency curve; the vertical coordinate difference between the second coordinate point and the first coordinate point is equal to the first probability;

[0011] The first parameter pair is determined based on the first coordinate point and the second coordinate point; wherein the first duration is equal to the coordinate value of the horizontal coordinate corresponding to the first coordinate point, and the second duration is equal to the horizontal coordinate difference between the second coordinate point and the first coordinate point.

[0012] In the above solution, the performing of the first data repair task based on the first parameter pair includes:

[0013] A first idle time period is selected from one or more idle time periods in the historical statistical period; the duration of the first idle time period is greater than the first duration;

[0014] Determine, based on the first parameter pair and the first idle time period, an execution start time and an execution duration of the first data repair task;

[0015] At the execution start time of the first data repair task, the first data repair task is started to be executed, and / or, when the execution time of the first data repair task is less than or equal to the execution duration, the execution of the first data repair task is stopped.

[0016] In the above solution, determining the execution start time of the first data repair task based on the first parameter pair and the first idle time period includes:

[0017] Based on the first idle time period, determining a first moment in the current statistical cycle that is the same as the start moment of the first idle time period; the first moment represents a start moment of the first data repair task;

[0018] Based on the first moment and the first duration, a start moment for executing the first data repair task is determined.

[0019] In the above scheme, the method further includes:

[0020] The first probability is determined based on the set performance degradation target, the third duration and the fourth duration; wherein,

[0021] The third duration represents the statistical value of the task completion duration of the front-end task of the storage node when the storage node does not perform the data repair task; the fourth duration represents the statistical value of the task completion duration increased by the front-end task of the storage node when the storage node performs the data repair task in the background;

[0022] The third duration and the fourth duration are determined based on the execution status of the tasks in the storage node in a historical statistical period.

[0023] In the above solution, the set performance degradation target is used to indicate: the ratio of the first difference to the third duration; wherein,

[0024] The first difference is equal to: the difference between the fifth duration and the third duration;

[0025] The fifth duration represents a statistical value of the task completion duration of the front-end task of the storage node when the storage node executes the data repair task in the background.

[0026] In the above scheme, the fifth duration is determined based on the sum of the third duration and the first product; and the first product is determined based on the product of the first probability and the fourth duration.

[0027] The embodiment of the present application further provides a data repair device, which is applied to a storage node in a distributed storage system, including:

[0028] A determination unit is used to determine a first parameter pair based on a first probability and a first cumulative frequency curve; the first probability represents: when the storage node performs a data repair task in the background, the probability of an increase in the task completion time of the front-end task of the storage node; the first probability is determined based on a set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: in one or more idle time periods within a historical statistical period, the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point;

[0029] An execution unit is used to execute a first data repair task based on the first parameter pair; the first parameter pair is used to control the execution time of the first data repair task.

[0030] The embodiment of the present application further provides a storage node, which is arranged in a distributed storage system, and includes: a first processor and a first communication interface; wherein,

[0031] The first processor is used to determine a first parameter pair based on a first probability and a first cumulative frequency curve; the first probability represents: when the storage node performs a data repair task in the background, the probability of increasing the task completion time of the front-end task of the storage node; the first probability is determined based on a set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point in one or more idle time periods within a historical statistical period; and,

[0032] Based on the first parameter pair, a first data repair task is executed; the first parameter pair is used to control the execution time of the first data repair task.

[0033] The embodiment of the present application further provides a storage node, which is arranged in a distributed storage system, and includes: a first processor and a first memory for storing a computer program that can be run on the processor,

[0034] Wherein, the first processor is used to execute the steps of any of the above methods when running the computer program.

[0035] An embodiment of the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0036] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps of any of the above methods when executed by a processor.

[0037] In an embodiment of the present application, a storage node in a distributed storage system determines a first parameter pair based on a first probability and a first cumulative frequency curve; wherein the first probability represents: when the storage node executes a data repair task in the background, the probability of an increase in the task completion time of the front-end task of the storage node; the first probability is determined based on a set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point in one or more idle time periods within a historical statistical period; then, the storage node executes the first data repair task based on the first parameter pair, wherein the first parameter pair is used to control the execution time of the first data repair task. That is, the storage node in the distributed storage system determines the first probability based on the performance degradation target, and then determines the first parameter pair based on the first probability and the statistics of the idle time period, so that the storage node can control the execution time of the first data repair task through the first parameter pair, so that the impact of the first data repair task on the front-end task is limited within the set performance degradation target, and compared with the related art, the execution time of the data repair task is reasonably scheduled, thereby ensuring the normal operation of the front-end task. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A schematic diagram of an implementation flow of a data repair method provided in an embodiment of the present application;

[0039] Figure 2 A schematic diagram of the execution of a data repair task provided in an embodiment of the present application;

[0040] Figure 3 A schematic diagram of a first cumulative frequency curve provided in an embodiment of the present application;

[0041] Figure 4 A schematic diagram of the structure of a data repair device provided in an embodiment of the present application;

[0042] Figure 5 A schematic diagram of the hardware composition structure of a storage node provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] In the related technology, the distributed storage system realizes data redundancy and fault tolerance based on the erasure code technology. Specifically, the distributed storage system can divide the stored data into multiple data shards, and generate redundant shards according to the set coding algorithm. The redundant shards can be used to recover the lost or damaged data shards, and can also be expressed as verification shards.

[0044] In a distributed storage system, data shards and check shards are distributed and stored on different storage nodes. When some data shards are lost, the storage nodes in the distributed storage system perform data repair tasks in the background to restore the lost data shards, that is, to restore the lost data. Specifically, the storage node can first determine the storage location of the lost data shards, then perform data repair based on the non-lost data shards and check shards, and then write the repaired data shards into the distributed storage system.

[0045] However, in the related art, the background data repair tasks and the front-end tasks of the storage nodes are superimposed for a long time, resulting in excessive utilization of system resources, which easily affects the normal operation of the front-end tasks. For example, when the storage node performs data repair tasks in the background, it also needs to respond to the user's storage request in real time at the front end. In this case, the background data repair tasks occupy the processing power of the storage node, resulting in high utilization of resources such as the central processing unit (CPU) and memory, which easily affects the read and write performance of the storage system, thereby causing the storage node to time out or retransmit data when responding to the storage request, that is, affecting the normal operation of the front-end tasks.

[0046] Based on this, in an embodiment of the present application, the storage node in the distributed storage system determines the first parameter pair based on the first probability and the first cumulative frequency curve; wherein the first probability represents: when the storage node executes the data repair task in the background, the probability of the task completion time of the front-end task of the storage node increasing; the first probability is determined based on the set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point in one or more idle time periods within the historical statistical period; then, the storage node executes the first data repair task based on the first parameter pair, wherein the first parameter pair is used to control the execution time of the first data repair task. That is to say, the storage node in the distributed storage system determines the first probability based on the performance degradation target, and then determines the first parameter pair based on the first probability and the statistics of the idle time period, so that the storage node can control the execution time of the first data repair task through the first parameter pair, so that the impact of the first data repair task on the front-end task is limited within the set performance degradation target, and compared with the related technology, the execution time of the data repair task is reasonably scheduled, thereby ensuring the normal operation of the front-end task.

[0047] The present application is further described in detail below with reference to the accompanying drawings and embodiments.

[0048] An embodiment of the present application provides a data repair method, which is applied to a storage node in a distributed storage system.

[0049] In practical applications, the storage node may be set in a distributed storage system, and the distributed storage system may include multiple storage nodes. The data stored in the distributed storage system may be distributed on these storage nodes.

[0050] See also Figure 1 , the data repair method provided in the embodiment of the present application includes:

[0051] Step 101: Determine a first parameter pair based on a first probability and a first cumulative frequency curve.

[0052] The first probability represents: when the storage node executes the data repair task in the background, the probability of increasing the task completion time of the front-end task of the storage node. The first probability is determined based on the set performance degradation target.

[0053] Each coordinate point in the first cumulative frequency curve is used to describe: the occurrence frequency of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point in one or more idle time periods within the historical statistical period.

[0054] In practical applications, data repair tasks can be used to repair data to be repaired in a distributed storage system, and the data to be repaired may include: lost data and / or damaged data. The data repair task can be first generated by the distributed storage system. For example, in the process of writing data to a storage node, if there is a data shard that fails to be written, the gateway module in the distributed storage system can record the error information of the data shard in the index table and generate a data repair task based on the error information.

[0055] In actual applications, storage nodes can perform data repair tasks in the background. Data repair tasks can be understood as a background task of storage nodes. Background tasks do not require direct user participation during operation. Front-end tasks can be understood as tasks that users can interact with directly. For example, front-end tasks can include: responding to user storage requests.

[0056] In actual applications, when a storage node executes a data repair task in the background, if the front-end task of the storage node is also running, it may lead to excessive utilization of system resources, thereby affecting the performance of the front-end task, for example, increasing the task completion time of the front-end task. The execution time of the data repair task can be controlled so that: the data repair task does not overlap with the front-end task, or, although the data repair task overlaps with the front-end task, the impact on the performance of the front-end task is within the allowable range. It should be noted that when the impact of the data repair task on the performance of the front-end task is within the allowable range, the front-end task can be regarded as running normally.

[0057] In practical applications, controlling the execution time of a data repair task can also be understood as scheduling the execution time of the data repair task.

[0058] Step 102: Execute a first data repair task based on a first parameter pair.

[0059] The first parameter pair is used to control the execution time of the first data repair task.

[0060] In practical applications, the first probability can be determined based on the set performance degradation target, and then the first parameter pair can be determined based on the first probability and the first cumulative frequency curve. Then, based on the first parameter pair, the execution time of the data repair task to be executed can be determined, and the data repair task can be executed at the determined time, thereby realizing the scheduling of the execution time of the data repair task to be executed.

[0061] Here, the first data repair task can be understood as a data repair task to be executed.

[0062] In actual applications, statistics can be collected for one or more idle time periods in a historical period at set intervals, and then a first cumulative frequency curve can be determined based on the statistical results. Specifically, a first cumulative frequency curve can be constructed based on the statistical results, or the constructed first cumulative frequency curve can be updated. The historical period being counted is also the historical statistical cycle. Exemplarily, the historical statistical cycle can be the past 24 hours. The idle time period can be characterized as a time period in which the resource utilization of the storage node is lower than a set threshold. During the idle time period, it can be considered that the storage node has not executed any front-end tasks.

[0063] In practical applications, during the statistical process, the duration of one or more idle time periods in the historical statistical period can be counted, and then a frequency cumulative curve can be determined based on the duration statistical results. The frequency cumulative curve can also be understood as a frequency cumulative histogram (CFH) in which the horizontal axis interval is divided into infinitely small intervals.

[0064] In the first frequency accumulation curve, the coordinate value of the horizontal coordinate of each coordinate point can represent a time length, and the coordinate value of the vertical coordinate of each coordinate point can represent: the cumulative frequency of the idle time period corresponding to the time length corresponding to the coordinate point, that is, the frequency of occurrence of all idle time periods with a time length less than or equal to the time length corresponding to the coordinate point in one or more idle time periods within the historical statistical period.

[0065] Exemplarily, if there are 10 idle time periods in the historical statistical period used to construct the first frequency curve, wherein the durations corresponding to these 10 idle time periods are: 1 minute (min), 1min, 2min, 2min, 2min, 3min, 3min, 5min, 10min, 12min, respectively, then, for a coordinate point whose horizontal coordinate value in the first frequency curve is 2, the coordinate value of the horizontal coordinate of the coordinate point represents the duration of 2min, and the vertical coordinate of the coordinate point represents: the cumulative frequency of the idle time periods corresponding to the duration of 2min, that is, the frequency of occurrence of all idle time periods with a duration less than or equal to 2min in these 10 idle time periods. It can be understood that the frequency of occurrence of all idle time periods with a duration less than or equal to 2min is 5, therefore, the cumulative frequency of the idle time periods corresponding to the duration of 2min is: 5 / 10, that is, 0.5.

[0066] In practical applications, the first cumulative frequency curve can be regarded as statistics of idle time periods in a historical statistical period, or can be regarded as a description of the duration distribution of idle time periods in a historical statistical period.

[0067] In practical applications, the idle time periods in the current statistical period can be predicted based on the statistics of the idle time periods in the historical statistical period. For example, for an idle time period of 2 minutes in the historical statistical period, if the starting time of the idle time period is 20:00 on the previous day, then the same time in the current statistical period, for example, 20:00 on the current day, can be predicted as the starting time of an idle time period of 2 minutes.

[0068] In practical applications, setting a performance degradation target can be understood as: a quantitative value of the maximum impact of the allowed data repair task on the performance of the front-end task. Setting a performance degradation target can be set by business personnel based on experience, or it can be expressed as a predefined target.

[0069] Here, the first probability is determined based on the set performance degradation target. In actual applications, the first probability can also be regarded as: under the limitation of the set performance degradation target, when the storage node performs data repair tasks in the background, the maximum probability of allowing the task completion time of the front-end task of the storage node to increase.

[0070] Here, the first parameter pair is determined based on the first probability and the first cumulative frequency curve, therefore, the first parameter pair can be considered to be related to the first probability and the first cumulative frequency curve. Furthermore, since the first probability is determined based on the set performance degradation target, and the first cumulative frequency curve is determined based on the statistics of the idle time periods within the historical statistical period, therefore, the first parameter pair can be considered to be related to the set performance degradation target and the duration distribution of the idle time periods within the historical statistical period. On this basis, the storage node in the embodiment of the present application can control the execution time of the first data repair task through the first parameter pair, so that the impact of the first data repair task on the front-end task is limited to the set performance degradation target. Compared with the related art, the execution time of the data repair task is reasonably scheduled, thereby ensuring the normal operation of the front-end task.

[0071] The first parameter pair is further described below.

[0072] In one embodiment, the first parameter pair is used to indicate a first duration and a second duration, wherein the first duration represents the duration from when the first data repair task is started to when it begins to be executed, and the second duration represents the duration of execution of the first data repair task.

[0073] In actual applications, after starting the first data repair task and before executing the first data repair task, the storage node may perform an initialization operation on the first data repair task. Exemplarily, the initialization operation may at least include: creating task description information and applying for resources. At this time, the storage node has not executed the first data repair task, so the first data repair task will not affect the performance of the front-end task. When the startup duration of the first data repair task is equal to the first duration, the storage node may start executing the first data repair task, and the first data repair task being executed may affect the performance of the front-end task.

[0074] In practical applications, the first duration can also be expressed as the execution delay duration.

[0075] In actual applications, the storage node may stop executing the first data repair task when the execution time of the first data repair task is less than or equal to the second time.

[0076] In practical applications, see Figure 2 , executing the data repair task based on the given execution delay duration and execution duration may include three execution situations. For ease of explanation, L is used to represent the duration of the idle time period, I is used to represent the execution delay duration, and T is used to represent the execution duration.

[0077] Case 1: L <I。

[0078] In this case, the idle time period is too short, and the front-end task has already started executing when the start time of the data repair task is less than the execution delay time. If the data repair task is still executed, it will have a significant impact on the performance of the front-end task. Therefore, in this case, the data repair task is not executed.

[0079] Case 2: I≤L≤(I+T).

[0080] In this case, the execution time of the data repair task in the idle time period is: LI. Afterwards, the data repair task is executed in conjunction with the front-end task. When the superimposed execution time reaches W, it can be considered that the utilization rate of system resources is too high. At this time, the storage node stops executing the data repair task and releases system resources to give priority to the front-end task. It can be seen that the actual execution time of the data repair task may be less than the execution duration.

[0081] In actual applications, when the combined execution time of the data repair task and the front-end task reaches W, the unfinished part of the data repair task can be used as a new data repair task and run in the subsequent idle time period.

[0082] In practical applications, the duration W of the superimposed execution can be determined based on the usage of system resources, or the execution of tasks in the storage node in the historical statistical period. Exemplarily, it can be determined based on: when the utilization rate of system resources reaches a set threshold, the statistics of the superimposed execution duration of the data repair task and the front-end task in the historical statistical period; it can also be determined based on the execution duration of the data repair task being executed when the utilization rate of system resources reaches a set threshold.

[0083] Case 3: L>I+T.

[0084] In this case, the idle time period can be considered as being able to provide sufficient time for the data repair task to run, and the data repair task can be uninterrupted during the running process.

[0085] In practical applications, based on the first parameter pair, the execution delay time and execution duration can be reasonably set, and the idle time period for executing the first data repair task can be reasonably selected to avoid the above situation 1, and to ensure as much as possible that the performance impact of the first data repair task on the front-end task does not exceed the set performance degradation target when situation 2 occurs. It can be understood that in this way, the first data repair task and the front-end task can be executed asynchronously, which also increases the flexibility of data repair.

[0086] In one embodiment, based on the first parameter pair, performing a first data repair task includes:

[0087] A first idle time period is selected from one or more idle time periods in a historical statistical period; the duration of the first idle time period is greater than the first duration;

[0088] Determine the execution start time and execution duration of the first data repair task based on the first parameter pair and the first idle time period;

[0089] At the execution start time of the first data repair task, the first data repair task is started to be executed, and / or, when the execution time of the first data repair task is less than or equal to the execution duration, the execution of the first data repair task is stopped.

[0090] In practical applications, the idle time period in the current statistical period can be determined based on the first parameter pair and the first idle time period, and then the execution start time of the first data repair task can be determined based on the start time of the idle time period and the execution delay duration indicated by the first parameter pair. The first parameter indicates the execution duration of the first data repair task, so it can be considered that the execution duration of the first data repair task is determined based on the first parameter pair.

[0091] In one embodiment, determining the execution start time of the first data repair task based on the first parameter pair and the first idle time period includes:

[0092] Based on the first idle time period, determining a first moment in the current statistical cycle that is the same as the start moment of the first idle time period; the first moment represents the start moment of the first data repair task;

[0093] Based on the first moment and the first duration, a start moment of execution of the first data repair task is determined.

[0094] In practical applications, the first moment may be regarded as a start moment of an idle time period in the current statistical period predicted based on the first idle time period.

[0095] Here, the duration of the first idle time period selected from the historical statistical period is greater than the first duration, based on which the duration of the idle time period corresponding to the first moment predicted can also be considered to be greater than the first duration. In this way, the occurrence of the aforementioned execution situation 1 is avoided, thereby ensuring the execution of the first data repair task and avoiding the waste of system resources.

[0096] In actual applications, the storage node may start the first data repair task at a first moment, and start executing the first data repair task when the start duration of the first data repair task is equal to the first duration.

[0097] The method for determining the first parameter pair is further described below.

[0098] In one embodiment, determining the first parameter pair based on the first probability and the first cumulative frequency curve includes:

[0099] Based on the first probability, a first coordinate point and a second coordinate point are selected in the first cumulative frequency curve; the vertical coordinate difference between the second coordinate point and the first coordinate point is equal to the first probability;

[0100] Based on the first coordinate point and the second coordinate point, a first parameter pair is determined; wherein the first duration is equal to the coordinate value of the horizontal coordinate corresponding to the first coordinate point, and the second duration is equal to the horizontal coordinate difference between the second coordinate point and the first coordinate point.

[0101] In practical applications, the coordinate values ​​of two vertical coordinates can be determined based on the first probability, and then the first coordinate point and the second coordinate point can be determined based on the coordinate values ​​of the two vertical coordinates. Then, the first time length and the second time length can be determined based on the coordinate values ​​of the horizontal coordinates of the first coordinate point and the second coordinate point, that is, the first parameter pair is determined.

[0102] For example, based on Figure 3In the process of determining the first parameter pair given the first cumulative frequency curve, coordinate point A and coordinate point B can be selected in the first cumulative frequency curve based on E; wherein E represents the first probability, A represents the first coordinate point, and B represents the second coordinate point. It can be seen that the vertical coordinate difference between A and B is equal to E. Then, based on the coordinate values ​​of the horizontal coordinates of A and B, the first duration and the second duration are determined; wherein the first duration is equal to the coordinate value of the horizontal coordinate corresponding to A, and the second duration is equal to the horizontal coordinate difference between B and A.

[0103] The method for determining the first probability is further described below.

[0104] In one embodiment, the data repair method provided by the embodiment of the present application further includes:

[0105] Based on the set performance degradation target, the third duration and the fourth duration, a first probability is determined; wherein,

[0106] The third duration represents the statistical value of the task completion duration of the front-end task of the storage node when the storage node does not perform the data repair task; the fourth duration represents the statistical value of the task completion duration increased by the front-end task of the storage node when the storage node performs the data repair task in the background;

[0107] The third duration and the fourth duration are determined based on the execution status of the tasks in the storage node during the historical statistical period.

[0108] In practical applications, the execution of tasks in the storage node can be monitored based on a set monitor, and then the third duration and the fourth duration can be determined based on the monitoring result of the set monitor. The set monitor can be set in the distributed storage system, and can also be further set in the storage node.

[0109] Exemplarily, the following can be monitored based on the setting monitor: when the storage node does not execute the data repair task, the task completion time of one or more front-end tasks of the storage node, and then the average task completion time of these front-end tasks is determined as the third time.

[0110] In one embodiment, the performance degradation target is set to indicate: a ratio of the first difference to the third duration; wherein,

[0111] The first difference is equal to: the difference between the fifth duration and the third duration;

[0112] The fifth duration represents a statistical value of the task completion duration of the front-end task of the storage node when the storage node executes the data repair task in the background.

[0113] In actual applications, the fifth duration may be determined based on a monitoring result of a set monitor on the execution status of tasks in the storage node.

[0114] In practical applications, the relationship between the performance degradation target, the third duration, and the fifth duration can be expressed as Formula 1, that is,

[0115]

[0116] Among them, D represents the set performance degradation target, RT represents the fifth duration, and RT FG Represents the third duration.

[0117] In one embodiment, the fifth duration is determined based on the sum of the third duration and the first product; and the first product is determined based on the product of the first probability and the fourth duration.

[0118] In practical applications, the relationship between the third duration, the fourth duration, the fifth duration and the first probability can be expressed as Formula 2, that is,

[0119] RT=RT FG +E×W

[0120] Among them, RT represents the fifth duration, RT FG represents the third duration, E represents the first probability, and W represents the fourth duration.

[0121] In practical applications, based on the above formula 1 and the above formula 2, we can get formula 3, that is,

[0122]

[0123] It can be seen that based on the set performance degradation target, the third duration and the fourth duration, the first probability can be determined.

[0124] In the embodiment of the present application, the storage node in the distributed storage system determines a first probability based on the performance degradation target, and then determines a first parameter pair based on the first probability and the statistics of the idle time period. In this way, the storage node can control the execution time of the first data repair task through the first parameter pair, so that the impact of the first data repair task on the front-end task is limited to the set performance degradation target, so that even in the case of the aforementioned execution situation 2, it can be ensured as much as possible that the performance impact of the first data repair task on the front-end task does not exceed the set performance degradation target. Compared with the related art, the execution time of the data repair task is reasonably scheduled, thereby ensuring the normal operation of the front-end task. Further, in the embodiment of the present application, the parameters used to determine the first parameter pair can be determined based on the execution status of the tasks in the storage node in the historical statistical period. Therefore, in the process of scheduling the data repair task, the execution status of the tasks in the storage node is fully considered, so that the storage node can better adapt to the execution status of the task in the process of scheduling the data repair task, thereby improving the scalability and flexibility of data repair.

[0125] The present application is further described in detail below in conjunction with application examples.

[0126] The application embodiment of the present application provides a data repair method, which is applied to a storage node in a distributed storage system. When the storage node performs data repair, the method mainly includes the following steps:

[0127] Step 1: Determine a first probability and a first cumulative frequency curve.

[0128] In practical applications, the first probability can be determined based on the set performance degradation target, the third duration, and the fourth duration, wherein the third duration and the fourth duration are determined based on the execution status of the tasks in the storage node in the historical statistical period.

[0129] In practical applications, the first cumulative frequency curve may be determined based on statistics of the duration of one or more idle time periods in a historical statistical period.

[0130] Step 2: Determine a first parameter pair based on the first probability and the first cumulative frequency curve.

[0131] The first parameter pair is used to indicate the first duration and the second duration.

[0132] Step 3: Determine the execution time of the first data repair task.

[0133] In actual applications, a first idle time period can be selected from one or more idle time periods in a historical statistical period, and the duration of the first idle time period is greater than the first duration; then, based on the first parameter pair and the first idle time period, the execution start time and execution duration of the first data repair task are determined.

[0134] Step 4: Execute the first data repair task based on the determined execution start time and execution duration.

[0135] In actual applications, the first data repair task can be started at the execution start time of the first data repair task, and / or the first data repair task can be stopped when the execution time of the first data repair task is less than or equal to the execution duration.

[0136] In the application embodiment of the present application, the storage node in the distributed storage system determines a first probability based on the performance degradation target, and then determines a first parameter pair based on the first probability and the statistics of the idle time period. In this way, the storage node can control the execution time of the first data repair task through the first parameter pair, so that the impact of the first data repair task on the front-end task is limited to the set performance degradation target. Compared with the related technology, the execution time of the data repair task is reasonably scheduled, thereby ensuring the normal operation of the front-end task.

[0137] Based on the above embodiments, the present application also provides a data repair device, which is applied to a storage node in a distributed storage system. Figure 4 , the data repair device comprises:

[0138] The determination unit 41 is used to determine a first parameter pair based on a first probability and a first cumulative frequency curve; the first probability represents: when the storage node performs a data repair task in the background, the probability of increasing the task completion time of the front-end task of the storage node; the first probability is determined based on a set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: in one or more idle time periods within a historical statistical period, the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point;

[0139] The execution unit 42 is used to execute a first data repair task based on the first parameter pair; the first parameter pair is used to control the execution time of the first data repair task.

[0140] In one embodiment, the first parameter pair is used to indicate a first duration and a second duration, wherein the first duration represents the duration from the start of the first data repair task to the start of execution, and the second duration represents the duration of execution of the first data repair task.

[0141] In one embodiment, the determining unit 41 determines the first parameter pair based on the first probability and the first cumulative frequency curve, including:

[0142] Based on the first probability, a first coordinate point and a second coordinate point are selected in the first cumulative frequency curve; the vertical coordinate difference between the second coordinate point and the first coordinate point is equal to the first probability;

[0143] The first parameter pair is determined based on the first coordinate point and the second coordinate point; wherein the first duration is equal to the coordinate value of the horizontal coordinate corresponding to the first coordinate point, and the second duration is equal to the horizontal coordinate difference between the second coordinate point and the first coordinate point.

[0144] In one embodiment, the execution unit 42 executes the first data repair task based on the first parameter pair, including:

[0145] A first idle time period is selected from one or more idle time periods in the historical statistical period; the duration of the first idle time period is greater than the first duration;

[0146] Determine, based on the first parameter pair and the first idle time period, an execution start time and an execution duration of the first data repair task;

[0147] At the execution start time of the first data repair task, the first data repair task is started to be executed, and / or, when the execution time of the first data repair task is less than or equal to the execution duration, the execution of the first data repair task is stopped.

[0148] In one embodiment, the determining unit 41 determines the execution start time of the first data repair task based on the first parameter pair and the first idle time period, including:

[0149] Based on the first idle time period, determining a first moment in the current statistical cycle that is the same as the start moment of the first idle time period; the first moment represents a start moment of the first data repair task;

[0150] Based on the first moment and the first duration, a start moment for executing the first data repair task is determined.

[0151] In one embodiment, the determining unit 41 determines the first probability based on the set performance degradation target, the third duration and the fourth duration; wherein,

[0152] The third duration represents the statistical value of the task completion duration of the front-end task of the storage node when the storage node does not perform the data repair task; the fourth duration represents the statistical value of the task completion duration increased by the front-end task of the storage node when the storage node performs the data repair task in the background;

[0153] The third duration and the fourth duration are determined based on the execution status of the tasks in the storage node in a historical statistical period.

[0154] In one embodiment, the set performance degradation target is used to indicate: a ratio of the first difference to the third duration; wherein,

[0155] The first difference is equal to: the difference between the fifth duration and the third duration;

[0156] The fifth duration represents a statistical value of the task completion duration of the front-end task of the storage node when the storage node executes the data repair task in the background.

[0157] In one embodiment, the fifth duration is determined based on the sum of the third duration and the first product; and the first product is determined based on the product of the first probability and the fourth duration.

[0158] In actual application, both the determination unit 41 and the execution unit 42 can be implemented by a processor in the data repair device.

[0159] It should be noted that: the data repair device provided in the above embodiment only uses the division of the above program modules as an example to illustrate when performing data repair. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the data repair device provided in the above embodiment and the data repair method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0160] Based on the hardware implementation of the above program module, and in order to implement the method of the embodiment of the present application, the present application also provides a storage node, which is arranged in a distributed storage system, referring to Figure 5 , the storage node includes:

[0161] The first communication interface 1 is capable of exchanging information with other devices;

[0162] The first processor 2 is connected to the first communication interface 1 to implement information exchange with other devices, and is used to execute the method provided by one or more technical solutions in the above embodiments when running a computer program. The computer program is stored in the first memory 3.

[0163] Specifically, the first processor 2 is used to determine a first parameter pair based on a first probability and a first cumulative frequency curve; the first probability represents: when the storage node performs a data repair task in the background, the probability of an increase in the task completion time of the front-end task of the storage node; the first probability is determined based on a set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point in one or more idle time periods within a historical statistical period; and,

[0164] Based on the first parameter pair, a first data repair task is executed; the first parameter pair is used to control the execution time of the first data repair task.

[0165] In one embodiment, the first parameter pair is used to indicate a first duration and a second duration, wherein the first duration represents the duration from the start of the first data repair task to the start of execution, and the second duration represents the duration of execution of the first data repair task.

[0166] In one embodiment, the first processor 2 determines the first parameter pair based on the first probability and the first cumulative frequency curve, including:

[0167] Based on the first probability, a first coordinate point and a second coordinate point are selected in the first cumulative frequency curve; the vertical coordinate difference between the second coordinate point and the first coordinate point is equal to the first probability;

[0168] The first parameter pair is determined based on the first coordinate point and the second coordinate point; wherein the first duration is equal to the coordinate value of the horizontal coordinate corresponding to the first coordinate point, and the second duration is equal to the horizontal coordinate difference between the second coordinate point and the first coordinate point.

[0169] In one embodiment, the first processor 2 performs a first data repair task based on the first parameter pair, including:

[0170] A first idle time period is selected from one or more idle time periods in the historical statistical period; the duration of the first idle time period is greater than the first duration;

[0171] Determine, based on the first parameter pair and the first idle time period, an execution start time and an execution duration of the first data repair task;

[0172] At the execution start time of the first data repair task, the first data repair task is started to be executed, and / or, when the execution time of the first data repair task is less than or equal to the execution duration, the execution of the first data repair task is stopped.

[0173] In one embodiment, the first processor 2 determines the execution start time of the first data repair task based on the first parameter pair and the first idle time period, including:

[0174] Based on the first idle time period, determining a first moment in the current statistical cycle that is the same as the start moment of the first idle time period; the first moment represents a start moment of the first data repair task;

[0175] Based on the first moment and the first duration, a start moment for executing the first data repair task is determined.

[0176] In one embodiment, the first processor 2 is further configured to determine the first probability based on the set performance degradation target, the third duration, and the fourth duration; wherein,

[0177] The third duration represents the statistical value of the task completion duration of the front-end task of the storage node when the storage node does not perform the data repair task; the fourth duration represents the statistical value of the task completion duration increased by the front-end task of the storage node when the storage node performs the data repair task in the background;

[0178] The third duration and the fourth duration are determined based on the execution status of the tasks in the storage node in a historical statistical period.

[0179] In one embodiment, the set performance degradation target is used to indicate: a ratio of the first difference to the third duration; wherein,

[0180] The first difference is equal to: the difference between the fifth duration and the third duration;

[0181] The fifth duration represents a statistical value of the task completion duration of the front-end task of the storage node when the storage node executes the data repair task in the background.

[0182] In one embodiment, the fifth duration is determined based on the sum of the third duration and the first product; and the first product is determined based on the product of the first probability and the fourth duration.

[0183] It should be noted that the specific processing process of the first communication interface 1 can be understood by referring to the above method.

[0184] Of course, in actual application, the various components in the storage node are coupled together through the bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 4 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5 Various buses are labeled as bus system 4.

[0185] The first memory 3 in the embodiment of the present application is used to store various types of data to support operations in the storage node. Examples of such data include: any computer program used to operate on the storage node.

[0186] The method disclosed in the above embodiment of the present application can be applied to the first processor 2, or implemented by the first processor 2. The first processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the first processor 2. The above-mentioned first processor 2 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The first processor 2 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the first memory 3. The first processor 2 reads the information in the first memory 3 and completes the steps of the above method in combination with its hardware.

[0187] In an exemplary embodiment, the storage node may be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general purpose processors, controllers, MCUs, Microprocessors, or other electronic components to perform the aforementioned methods.

[0188] It can be understood that the first memory 3 of the embodiment of the present application can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAMbus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.

[0189] In an exemplary embodiment, the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a storage node storing a computer program, and the computer program can be executed by the first processor 2 of the storage node to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.

[0190] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program, and the computer program can be executed by the first processor 2 of the storage node to complete the steps of any of the aforementioned methods.

[0191] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0192] The term "and / or" herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the three situations that A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "one or more" herein represents any combination of at least two of any one or more of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set consisting of A, B, and C.

[0193] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.

Claims

1. A data repair method, characterized in that: Applied to a storage node in a distributed storage system, the method comprises: Based on the first probability and the first cumulative frequency curve, a first parameter pair is determined; the first probability represents: when the storage node performs a data repair task in the background, the probability of increasing the task completion time of the front-end task of the storage node; the first probability is determined based on the set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: in one or more idle time periods within the historical statistical period, the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point; Based on the first parameter pair, a first data repair task is executed; the first parameter pair is used to control the execution time of the first data repair task.

2. The method according to claim 1, characterized in that The first parameter pair is used to indicate a first duration and a second duration, wherein the first duration represents the duration from when the first data repair task is started to when it begins to be executed, and the second duration represents the duration of execution of the first data repair task.

3. The method according to claim 2, characterized in that The step of determining the first parameter pair based on the first probability and the first cumulative frequency curve comprises: Based on the first probability, a first coordinate point and a second coordinate point are selected in the first cumulative frequency curve; the vertical coordinate difference between the second coordinate point and the first coordinate point is equal to the first probability; The first parameter pair is determined based on the first coordinate point and the second coordinate point; wherein the first duration is equal to the coordinate value of the horizontal coordinate corresponding to the first coordinate point, and the second duration is equal to the horizontal coordinate difference between the second coordinate point and the first coordinate point.

4. The method according to claim 2, characterized in that: The performing a first data repair task based on the first parameter pair includes: A first idle time period is selected from one or more idle time periods in the historical statistical period; the duration of the first idle time period is greater than the first duration; Determine, based on the first parameter pair and the first idle time period, an execution start time and an execution duration of the first data repair task; At the execution start time of the first data repair task, the first data repair task is started to be executed, and / or, when the execution time of the first data repair task is less than or equal to the execution duration, the execution of the first data repair task is stopped.

5. The method according to claim 4, characterized in that Determining the execution start time of the first data repair task based on the first parameter pair and the first idle time period includes: Based on the first idle time period, determining a first moment in the current statistical cycle that is the same as the start moment of the first idle time period; the first moment represents a start moment of the first data repair task; Based on the first moment and the first duration, a start moment for executing the first data repair task is determined.

6. The method according to claim 1, characterized in that The method further comprises: The first probability is determined based on the set performance degradation target, the third duration and the fourth duration; wherein, The third duration represents the statistical value of the task completion duration of the front-end task of the storage node when the storage node does not perform the data repair task; the fourth duration represents the statistical value of the task completion duration increased by the front-end task of the storage node when the storage node performs the data repair task in the background; The third duration and the fourth duration are determined based on the execution status of the tasks in the storage node in a historical statistical period.

7. The method according to claim 6, characterized in that The set performance degradation target is used to indicate: the ratio of the first difference to the third duration; wherein, The first difference is equal to: the difference between the fifth duration and the third duration; The fifth duration represents a statistical value of the task completion duration of the front-end task of the storage node when the storage node executes the data repair task in the background.

8. The method according to claim 7, characterized in that The fifth duration is determined based on the sum of the third duration and the first product; the first product is determined based on the product of the first probability and the fourth duration.

9. A data repair device, characterized in that: Storage nodes used in distributed storage systems include: A determination unit is used to determine a first parameter pair based on a first probability and a first cumulative frequency curve; the first probability represents: when the storage node performs a data repair task in the background, the probability of an increase in the task completion time of the front-end task of the storage node; the first probability is determined based on a set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: in one or more idle time periods within a historical statistical period, the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point; An execution unit is used to execute a first data repair task based on the first parameter pair; the first parameter pair is used to control the execution time of the first data repair task.

10. A storage node, characterized in that: Set in a distributed storage system, including: a first processor and a first communication interface; wherein, The first processor is used to determine a first parameter pair based on a first probability and a first cumulative frequency curve; the first probability represents: when the storage node performs a data repair task in the background, the probability of increasing the task completion time of the front-end task of the storage node; the first probability is determined based on a set performance degradation target; each coordinate point in the first cumulative frequency curve is used to describe: the frequency of occurrence of all idle time periods whose duration is less than or equal to the duration corresponding to the coordinate point in one or more idle time periods within a historical statistical period; and, Based on the first parameter pair, a first data repair task is executed; the first parameter pair is used to control the execution time of the first data repair task.

11. A storage node, characterized in that: Set up in a distributed storage system, including: a first processor and a first memory for storing a computer program executable on the processor, Wherein, when the first processor is used to run the computer program, the steps of the method described in any one of claims 1 to 8 are executed.

12. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.