Data recovery method, electronic equipment, program product and storage medium

By deleting duplicate information in the lost information recording table of data shards, the duplicate repair problem caused by recording data shards missing information in two independent tables is solved, and resource utilization is improved.

CN119938404APending Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202412000163.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the lost information of data shards will be recorded in two independent repair tables, resulting in duplicate repairs and low resource utilization.

Method used

By determining the shard loss information of data shards of the same storage object recorded in the first data table and the second data table, the shard loss information recorded in one of the data tables is deleted to avoid duplicate repairs.

Benefits of technology

This achieves avoiding duplicate repairs and improves resource utilization, and only requires the loss of shards of the same storage object in a data table.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938404A_ABST
    Figure CN119938404A_ABST
Patent Text Reader

Abstract

The invention discloses a data recovery method, electronic equipment, a program product and a storage medium. The data recovery method comprises the following steps: determining fragment loss information of data fragments of the same storage object recorded in a first data table and a second data table; the first data table comprises fragment loss information of data fragments which are written unsuccessfully when the storage object is subjected to fragment storage, and the second data table comprises fragment loss information of data fragments which are detected when a disk is scanned and cannot be acquired; according to the fragment loss information of the same storage object recorded in the first data table and the second data table, deleting the fragment loss information of the same storage object recorded in the first data table or the second data table; the first data table and the second data table are used for repairing the lost data fragments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data repair method, electronic equipment, program product and storage medium. Background Art

[0002] Data sharding is an important technology for distributed storage. During the data storage process, if a data shard fails to be written, the loss information of the data shard will be recorded in the upload loss table. At the same time, when scanning the storage unit, the uploaded lost data shard can be scanned out, and the loss information of the data shard can be recorded in the storage loss table. That is, the loss of data shards of the same storage object will be recorded in two repair tables. The two repair tables run the repair task independently. When the two repairs are running at the same time, there may be duplicate repairs. Summary of the invention

[0003] In view of this, embodiments of the present invention provide a data repair method, an electronic device, a program product, and a storage medium.

[0004] The technical solution of the embodiment of the present invention is achieved as follows:

[0005] In one aspect, an embodiment of the present invention provides a data repair method, the method comprising:

[0006] Determine the shard loss information of the data shards of the same storage object recorded in the first data table and the second data table; the first data table includes the shard loss information of the data shards that failed to be written when the storage object is stored in shards, and the second data table includes the shard loss information of the data shards that cannot be obtained when the disk is scanned;

[0007] According to the shard loss information of the same storage object recorded in the first data table and the second data table, the shard loss information of the same storage object recorded in the first data table or the second data table is deleted; the first data table and the second data table are used to repair the lost data shards.

[0008] In the above solution, the shard loss information includes an index of the data shard, and the deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table or the second data table includes:

[0009] If the number of indexes of the data shards of the same storage object recorded in the first data table and the second data table is the same, the shard loss information of the same storage object recorded in the first data table or the second data table is deleted.

[0010] In the above solution, the shard loss information includes an index of the data shard, and the deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table or the second data table includes:

[0011] If the index numbers of the data shards of the same storage object recorded in the first data table and the second data table are different, the shard loss information of the same storage object recorded in the first data table is deleted.

[0012] In the above solution, the shard loss information includes an index of the data shard, and the deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table or the second data table includes:

[0013] If the index numbers of the data shards of the same storage object recorded in the first data table and the second data table are the same, the shard loss information of the same storage object recorded in the second data table is deleted.

[0014] In the above solution, the shard loss information includes the index of the data shard and the name of the corresponding storage object, and the determining of the shard loss information of the data shard of the same storage object recorded in the first data table and the second data table includes:

[0015] According to the name of the same storage object, search the first data table and the second data table for indexes of the data slices of the same storage object.

[0016] In the above solution, deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table and the second data table includes:

[0017] Determine the name of the data table to be operated and the deletion entry in the corresponding data table according to the shard loss information of the same storage object recorded in the first data table and the second data table;

[0018] Write deletion information to the message queue according to the name of the data table to be operated and the deletion entry in the corresponding data table;

[0019] Pull a delete message from the message queue, and delete the delete entry in the corresponding data table according to the delete message.

[0020] In the above solution, after deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table and the second data table, the method further includes:

[0021] Generate a plurality of repair tasks according to the first data table and the second data table, wherein the repair tasks are used to repair the lost data shards;

[0022] Determine a scheduling index of each repair node based on the node resource utilization of each repair node among the multiple repair nodes; the scheduling index represents the number of repair tasks that can be performed simultaneously by the repair node;

[0023] The multiple repair tasks are allocated to each repair node according to the scheduling index of each repair node.

[0024] In the above scheme, the node resource utilization includes at least one of CPU utilization, disk utilization and memory utilization; the determining the scheduling index of each repair node based on the node resource utilization of each repair node among the multiple repair nodes includes:

[0025] According to the CPU utilization, disk utilization and memory utilization of each repair node, the scheduling index of the corresponding repair node is determined.

[0026] In the above solution, determining the scheduling index of the corresponding repair node according to the CPU utilization, disk utilization and memory utilization of each repair node includes:

[0027] If any one of the CPU utilization, disk utilization, and memory utilization of the repair node is greater than the first threshold, setting the scheduling index of the corresponding repair node to the first set value;

[0028] If any one of the CPU utilization, disk utilization and memory utilization of the repair node is greater than the second threshold and less than the first threshold, the scheduling index of the corresponding repair node is set to a second set value; the second set value is greater than the first set value;

[0029] If the CPU utilization, disk utilization and memory utilization of the repair node are all less than the second threshold, the scheduling index is calculated based on the CPU utilization, disk utilization and memory utilization, and the scheduling index of the corresponding repair node is greater than the second set value.

[0030] In the above solution, allocating the multiple repair tasks to each repair node according to the scheduling index of each repair node includes:

[0031] If the scheduling index of the repair node is the first set value, no repair task is assigned to the corresponding repair node, and the repair task being executed on the corresponding repair node is stopped;

[0032] If the scheduling index of the repair node is the second set value, no repair task is assigned to the corresponding repair node, and other repair tasks on the corresponding repair node except the repair task being executed are assigned to the repair node whose scheduling index is greater than the second set value;

[0033] If the scheduling index of the repair node is greater than the second set value, a corresponding number of repair tasks are allocated to the corresponding repair node according to the scheduling index.

[0034] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program, which implements the steps of the above-mentioned data repair method when executed by a processor.

[0035] On the other hand, an embodiment of the present invention provides an electronic device, including a processor and a memory, which are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the steps of the data repair method provided by the embodiment of the present invention.

[0036] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, including: the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the data repair method provided in the embodiment of the present invention are implemented.

[0037] The embodiment of the present application determines the shard loss information of the data shards of the same storage object recorded in the first data table and the second data table, wherein the first data table includes the shard loss information of the data shards that failed to be written when the storage object is stored in shards, and the second data table includes the shard loss information of the data shards that cannot be obtained when the disk is scanned. According to the shard loss information of the same storage object recorded in the first data table and the second data table, the shard loss information of the same storage object recorded in the first data table or the second data table is deleted, and the first data table and the second data table are used to repair the lost data shards. The embodiment of the present application can delete the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table and the second data table, and only retain the shard loss information of the same storage object in one data table, so that the repair task will only be executed once, and the repair will not be repeated, thereby improving resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1It is a schematic diagram of an implementation flow of a data repair method provided by an embodiment of the present invention;

[0039] Figure 2 is a schematic diagram of a data repair process provided by an embodiment of the present invention;

[0040] Figure 3 is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] A distributed storage system, in simple terms, is a storage system that stores data in a dispersed manner across multiple independent devices, using multiple storage servers to share the storage load. This not only improves the reliability, availability, and access efficiency of the system, but is also easy to expand. It is currently the most widely used cloud storage architecture.

[0043] The data sharding algorithm is a common distributed storage method and technology. Distributed object storage supports the erasure code mode. The erasure code strategy is K+M, where K is the number of data shards and M is the number of erasure code shards. The shards are written into K+M single-copy storage pools. Under this storage strategy, up to M pieces of data can be lost. During the storage process, the shard writing may fail. When the number of lost shards is less than M, the lost shard information will be recorded in the data table. The lost information record table can be scanned to directly recover the data later. At the same time, there is another situation. After the data is written into K+M single-copy storage pools, the disk may be damaged, causing some shards to be unable to be obtained. In this case, data loss repair requires scanning the storage unit to discover and repair. The tables used by the two repair methods are independent of each other, and the tasks are executed in parallel without direct connection. The start and stop of the other two repair tasks are not affected by external conditions, and they will not be re-planned and scheduled according to node resources. After starting, a new round of repair tasks will not start until the end of the current round of tasks.

[0044] The related technology has the following disadvantages:

[0045] 1. When uploading data, the records of data shard loss will be recorded in the upload loss table. At the same time, when scanning the storage unit, the uploaded lost shard can be scanned and recorded in the storage loss table. That is, if the loss of the same object shard is not repaired, it will exist in two repair task tables. When the two repair methods are run at the same time, there may be repeated repairs. When the two repair tasks are executed separately, no duplicate tasks are removed or optimized, resulting in a waste of system resources.

[0046] 2. The current scheduling algorithm for repair tasks only considers node computing resources and memory resources, but does not consider the resource usage of the disk itself, resulting in the task still running when the disk is under great pressure. There is also a situation where the repair task cannot currently run on the same node. Although this scheduling prevents multiple repair tasks of the same type from consuming node resources, it reduces resource utilization when resources are sufficient.

[0047] In view of the shortcomings of the above-mentioned related technologies, an embodiment of the present invention provides a data repair method, which can improve the efficiency of data repair. In order to illustrate the technical solution of the present invention, a specific embodiment is used for description below.

[0048] refer to Figure 1 , Figure 1 1 is a schematic diagram of an implementation flow of a data repair method provided by an embodiment of the present invention, the data repair method comprising:

[0049] S101, determining the shard loss information of the data shards of the same storage object recorded in the first data table and the second data table; the first data table includes the shard loss information of the data shards that failed to be written when the storage object is stored in shards, and the second data table includes the shard loss information of the data shards that cannot be obtained and are detected during disk scanning.

[0050] For example, a storage object is divided into multiple shards. During the shard storage process, a shard may not be written successfully, that is, the shard is lost. At this time, the loss information of the data shard will be recorded in the upload loss table (corresponding to the first data table).

[0051] When scanning the storage unit, the unavailable shards will be scanned out, including shards that were not successfully written during upload, and situations where some shards cannot be obtained due to disk damage, which is also called shard loss. The scanned lost shards will be recorded in the storage loss table (corresponding to the second data table).

[0052] The first data table and the second data table mentioned above are both data tables existing in the related art. “First”, “Second”, etc. are used to distinguish similar objects, not to describe a specific order or sequence.

[0053] The following situations may occur when the same storage object is stored in shards:

[0054] 1. All shards are written successfully, but some shards cannot be obtained due to disk damage during the storage unit scan. At this time, there are no records in the first data table, but there are records in the second data table.

[0055] 2. The shard writing fails. At this time, there are records in both the first data table and the second data table.

[0056] The corresponding shard loss information can be found in the first data table and the second data table according to the name of the storage object.

[0057] S102: According to the shard loss information of the same storage object recorded in the first data table and the second data table, delete the shard loss information of the same storage object recorded in the first data table or the second data table; the first data table and the second data table are used to repair the lost data shards.

[0058] For the first data table and the second data table, each data table corresponds to a type of repair task, and the repair of lost fragments in the two data tables is performed independently. In order to avoid repeated repairs, this embodiment deduplicates the two repair tasks and only retains one repeated task, thereby improving the effective repair rate of the repair tasks.

[0059] In this regard, this embodiment can delete the shard loss information of the same storage object recorded in the first data table or the second data table based on the shard loss information of the same storage object recorded in the first data table and the second data table, and only retain the shard loss information of the same storage object in one data table. In this way, the repair task will only be executed once and will not be repeated.

[0060] For example, if the shard loss information of the same storage object recorded in the first data table and the second data table is exactly the same, the shard loss information of the same storage object recorded in the first data table or the second data table is randomly deleted.

[0061] If the shard loss information of the same storage object recorded in the first data table and the second data table is not exactly the same, the shard loss information of the storage object recorded in the first data table is deleted. This is because it is possible that the shards are uploaded successfully, but the disk is damaged and some shards cannot be obtained. The information recorded in the second data table is more comprehensive, so the shard loss information of the storage object recorded in the second data table is retained, and the shard loss information of the storage object recorded in the first data table is deleted.

[0062] In the embodiments of the present application, by determining the shard loss information of data shards of the same storage object recorded in the first data table and the second data table, the first data table includes the shard loss information of the data shards that failed to be written during the sharded storage of the storage object, and the second data table includes the shard loss information of the data shards that cannot be obtained detected during the disk scan. According to the shard loss information of the same storage object recorded in the first data table and the second data table, the shard loss information of the same storage object recorded in the first data table or the second data table is deleted. The first data table and the second data table are used to repair the lost data shards. In the embodiments of the present application, according to the shard loss information of the same storage object recorded in the first data table and the second data table, the shard loss information of the same storage object recorded in the first data table or the second data table can be deleted, and the shard loss information of the same storage object is only retained in one data table. In this way, the repair task will only be executed once and will not be repaired repeatedly, improving resource utilization.

[0063] In one embodiment, the shard loss information includes the index of the data shard. According to the shard loss information of the same storage object recorded in the first data table and the second data table, deleting the shard loss information of the same storage object recorded in the first data table or the second data table includes:

[0064] If the number of indexes of the data shards of the same storage object recorded in the first data table and the second data table is the same, then delete the shard loss information of the same storage object recorded in the first data table or the second data table.

[0065] In one embodiment, the shard loss information includes the index of the data shard. According to the shard loss information of the same storage object recorded in the first data table and the second data table, deleting the shard loss information of the same storage object recorded in the first data table or the second data table includes:

[0066] If the number of indexes of the data shards of the same storage object recorded in the first data table and the second data table is different, then delete the shard loss information of the same storage object recorded in the first data table.

[0067] The current data repair of distributed storage is divided into two types. The first type of repair data source is the shards that were not successfully written to the disk when the user uploaded. The second type of repair data source is the missing shards found during the storage unit scan. For example, assume that when a storage object (divided into K + M shards) is uploaded, N shards are not stored successfully (N < M). Then, in the second repair method at this time, at least N shards are repaired. Because without considering disk damage, the storage object will be equal to N shards. If the disk is damaged, it may be more than N shards.

[0068] When storing loss information of lost shards, the embodiment of the present application assigns an index to each lost shard. When deduplicating repair tasks, deduplication can be performed based on the number of indexes of data shards of the same storage object recorded in the first data table and the second data table.

[0069] For example, if the index numbers of data shards of the same storage object recorded in the first data table and the second data table are the same, it means that the shards lost during upload are all scanned during disk scanning, and the shard loss information of the storage object recorded in one of the data tables can be deleted arbitrarily.

[0070] If the index numbers of data shards of the same storage object recorded in the first data table and the second data table are different, it means that more lost shards are found during disk scanning due to disk damage, so only the shard loss information of the storage object recorded in the first data table is deleted.

[0071] In one embodiment, if the index numbers of the data shards of the same storage object recorded in the first data table and the second data table are the same, the shard loss information of the storage object recorded in the second data table is deleted. This is because the second data table is scanned first and then repaired, and the repair efficiency is lower than that of the first data table. Therefore, deleting the records in the second data table and retaining the records in the first data table can improve the repair efficiency.

[0072] In one embodiment, the shard loss information includes an index of the data shard and a name of a corresponding storage object, and determining the shard loss information of the data shard of the same storage object recorded in the first data table and the second data table includes:

[0073] According to the name of the same storage object, search the first data table and the second data table for indexes of the data slices of the same storage object.

[0074] In the first data table and the second data table, the storage object and the shard loss information are associated through the object name and the index of the lost shard, so that the shard loss information of the storage object can be easily found according to the object name.

[0075] In one embodiment, deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table and the second data table includes:

[0076] Determine the name of the data table to be operated and the deletion entry in the corresponding data table according to the shard loss information of the same storage object recorded in the first data table and the second data table;

[0077] Write deletion information to the message queue according to the name of the data table to be operated and the deletion entry in the corresponding data table;

[0078] Pull a delete message from the message queue, and delete the delete entry in the corresponding data table according to the delete message.

[0079] In one embodiment, a checking and sorting unit is introduced, and its function is to continuously scan two types of data tables (first data table and second data table) of information to be repaired. If the number of missing fragments (the number of missing fragment indexes) of the same object is equal in the two table records, the records in any data table are deleted at random, otherwise the data in the first data table is deleted.

[0080] like Figure 2 As shown, the inspection and sorting unit includes a scanning module, a message queue module and a deletion module, wherein the scanning module compares the data in the second data table based on the first data table, and after the comparison, the operation table name and the corresponding deletion entry are packaged and passed to the message queue module, and the next round of scanning is performed again. The deletion module pulls the corresponding deletion information from the message queue module and then operates the corresponding data table. These three scales can be adjusted according to the actual cluster scale and processing rate requirements. If the cluster scale is large or the processing rate needs to be increased, the number of these three modules can be increased to achieve the effect.

[0081] The embodiment of the present application introduces a checking and sorting unit to check the duplication of the two types of repair tasks and decide to retain one of the repair tasks based on the comparison of the number of losses. The two loss modes each handle their own tasks to reduce overlap, which can greatly improve the efficiency of a single repair and reduce the waste of computing node resources caused by repeated repair tasks.

[0082] In one embodiment, after deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table and the second data table, the method further includes:

[0083] Generate multiple repair tasks according to the first data table and the second data table, where the repair tasks are used to repair the lost data shards;

[0084] Determine a scheduling index of each repair node based on the node resource utilization of each repair node among the multiple repair nodes; the scheduling index represents the number of repair tasks that can be performed simultaneously by the repair node;

[0085] The multiple repair tasks are allocated to each repair node according to the scheduling index of each repair node.

[0086] Among them, two types of repair tasks are generated according to the first data table and the second data table. During the repair process of lost shards, the lost shard data will be written. Both types of repair tasks will have a greater impact on the disk. When the disk utilization is high, in order not to affect the performance of online reading, writing and deletion, the data repair tasks must be adjusted.

[0087] Among them, the repair tasks can be assigned to multiple nodes to run in parallel to improve the repair efficiency. The scheduling index of the node is proportional to the node resource utilization rate. The higher the scheduling index, the more repair tasks the repair node can execute simultaneously.

[0088] In one embodiment, the node resource utilization includes at least one of a central processing unit (CPU) utilization, a disk utilization, and a memory utilization.

[0089] The disk repair function supports segmented scheduling, that is, multiple tasks can be executed simultaneously or in a certain order according to the needs. All segments of the first two disk repair functions are executed simultaneously. At this time, you can monitor the node CPU utilization (unit %), memory utilization (unit %) and disk utilization (unit %). These three indicators are used to comprehensively consider whether it is necessary to schedule multiple repair tasks to the node or stop the scheduled tasks on the relevant nodes.

[0090] The step of determining a scheduling index of each repair node based on the node resource utilization rate of each repair node among the multiple repair nodes includes:

[0091] According to the CPU utilization, disk utilization and memory utilization of each repair node, the scheduling index of the corresponding repair node is determined.

[0092] For example, the scheduling index is obtained by performing weighted calculation according to CPU utilization, disk utilization, and memory utilization.

[0093] In one embodiment, determining the scheduling index of the corresponding repair node according to the CPU utilization, disk utilization, and memory utilization of each repair node includes:

[0094] If any one of the CPU utilization, disk utilization, and memory utilization of the repair node is greater than the first threshold, setting the scheduling index of the corresponding repair node to the first set value;

[0095] If any one of the CPU utilization, disk utilization and memory utilization of the repair node is greater than the second threshold and less than the first threshold, the scheduling index of the corresponding repair node is set to a second set value; the second set value is greater than the first set value;

[0096] If the CPU utilization, disk utilization and memory utilization of the repair node are all less than the second threshold, the scheduling index is calculated based on the CPU utilization, disk utilization and memory utilization, and the scheduling index of the corresponding repair node is greater than the second set value.

[0097] For example, disk_util, memory_util and cpu_util are used to represent the percentage values ​​of disk utilization, memory utilization and CPU utilization respectively. The first threshold is the danger threshold (DangerThreshold, DT), the second threshold is the warning threshold (WarningThreshold, WT), and the scheduling index is (Scheduling Index, SI).

[0098] If any one of the parameters of disk_util, memory_util and cpu_util exceeds the danger threshold, that is, disk_util>DT or memory_util>DT or cpu_util>DT, the scheduling index SI=-1. Here, the first setting value is set to -1.

[0099] If any parameter exceeds the alarm threshold, that is, disk_util>WT or memory_util>WT or cpu_util>WT, the scheduling index SI=0. Here, the second setting value is set to 0.

[0100] If the dangerous state or the alarm state is not reached, that is, disk_util, memory_util and cpu_util are all less than WT, the scheduling index SI = int(100-min(disk_util, memory_util, cpu_util)). Where min(disk_util, memory_util, cpu_util) represents the minimum value of the three parameters. At this time, the scheduling index SI is greater than 0.

[0101] In one embodiment, allocating the plurality of repair tasks to each repair node according to the scheduling index of each repair node includes:

[0102] If the scheduling index of the repair node is the first set value, no repair task is assigned to the corresponding repair node, and the repair task being executed on the corresponding repair node is stopped;

[0103] If the scheduling index of the repair node is the second set value, no repair task is assigned to the corresponding repair node, and other repair tasks on the corresponding repair node except the repair task being executed are assigned to the repair node whose scheduling index is greater than the second set value;

[0104] If the scheduling index of the repair node is greater than the second set value, a corresponding number of repair tasks are allocated to the corresponding repair node according to the scheduling index.

[0105] For example, when the scheduling index SI = -1, the execution status of the current repair node is recorded in the database, and the repair task on the current repair node is stopped immediately. No repair task is assigned to the repair node. If there is a repair task to be repaired on the node, the repair task to be repaired can be assigned to other repair nodes.

[0106] If the scheduling index SI=0, one repair task of each type is reserved for this repair node, and the redundant scheduling tasks are distributed to the remaining repair nodes with higher scheduling indexes.

[0107] When the scheduling index SI is greater than 0, the number of repair tasks that can be executed simultaneously by a single repair node will increase according to the size of the scheduling index. In order to avoid drastic changes in the scheduling index value in a short period of time, the scheduling index can be recalculated every 5 minutes. When the number of the same scheduling tasks exceeds the number of schedulable nodes, multiple repair tasks of the same type can be placed on nodes with large scheduling indexes. The larger the scheduling index, the more repair tasks can be executed simultaneously.

[0108] The embodiment of the present application incorporates the disk utilization index into the scheduling algorithm to avoid the situation where the repair task uses excessive disk resources and affects the online read and write business. At the same time, the optimized scheduling algorithm will improve the utilization of nodes with sufficient system resources, dynamically adjust the number of repair tasks executed on a single node according to the usage of node resources, and maximize the utilization of various resources of a single node.

[0109] The embodiment of the present application dynamically regulates the repair task based on a comprehensive algorithm of memory resource utilization rate, memory resource utilization rate and disk utilization rate, which not only prevents the repair task from interfering with the online read, write and delete services, but also ensures that when resources are abundant, multiple scheduling tasks are started on the same node to increase the repair rate while making full use of resources.

[0110] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0111] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0112] It should be noted that the technical solutions described in the embodiments of the present invention can be arbitrarily combined without conflict.

[0113] In addition, in the embodiments of the present invention, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0114] The embodiment of the present application also provides a data repair device, which corresponds to the data repair method at the receiving end mentioned above, and each step in the embodiment of the data repair method mentioned above is also fully applicable to the embodiment of the present device.

[0115] The device includes:

[0116] A determination module, used to determine the shard loss information of the data shards of the same storage object recorded in the first data table and the second data table; the first data table includes the shard loss information of the data shards that failed to be written when the storage object is stored in shards, and the second data table includes the shard loss information of the data shards that cannot be obtained when the disk is scanned;

[0117] A deletion module is used to delete the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table and the second data table; the first data table and the second data table are used to repair the lost data shards.

[0118] In one embodiment, the shard loss information includes an index of the data shard, and the deletion module is specifically configured to:

[0119] If the number of indexes of the data shards of the same storage object recorded in the first data table and the second data table is the same, the shard loss information of the same storage object recorded in the first data table or the second data table is deleted.

[0120] In one embodiment, the shard loss information includes an index of the data shard, and the deletion module is specifically configured to:

[0121] If the index numbers of the data shards of the same storage object recorded in the first data table and the second data table are different, the shard loss information of the same storage object recorded in the first data table is deleted.

[0122] In one embodiment, the shard loss information includes an index of the data shard, and the deletion module is specifically configured to:

[0123] If the index numbers of the data shards of the same storage object recorded in the first data table and the second data table are the same, the shard loss information of the same storage object recorded in the second data table is deleted.

[0124] In one embodiment, the shard loss information includes an index of the data shard and a name of a corresponding storage object, and the determination module is specifically configured to:

[0125] According to the name of the same storage object, search the first data table and the second data table for indexes of the data slices of the same storage object.

[0126] In one embodiment, the deletion module is specifically used to:

[0127] Determine the name of the data table to be operated and the deletion entry in the corresponding data table according to the shard loss information of the same storage object recorded in the first data table and the second data table;

[0128] Write deletion information to the message queue according to the name of the data table to be operated and the deletion entry in the corresponding data table;

[0129] Pull a delete message from the message queue, and delete the delete entry in the corresponding data table according to the delete message.

[0130] In one embodiment, the device further comprises:

[0131] A generating module, configured to generate a plurality of repair tasks according to the first data table and the second data table, wherein the repair tasks are used to repair the lost data slices;

[0132] A scheduling index determination module, used to determine a scheduling index of each repair node based on the node resource utilization rate of each repair node among the multiple repair nodes; the scheduling index represents the number of repair tasks that can be performed simultaneously by the repair node;

[0133] The allocation module is used to allocate the multiple repair tasks to each repair node according to the scheduling index of each repair node.

[0134] In one embodiment, the node resource utilization includes at least one of CPU utilization, disk utilization, and memory utilization; and the scheduling index determination module is specifically used to:

[0135] According to the CPU utilization, disk utilization and memory utilization of each repair node, the scheduling index of the corresponding repair node is determined.

[0136] In one embodiment, the scheduling index determination module is specifically used to:

[0137] If any one of the CPU utilization, disk utilization, and memory utilization of the repair node is greater than the first threshold, setting the scheduling index of the corresponding repair node to the first set value;

[0138] If any one of the CPU utilization, disk utilization and memory utilization of the repair node is greater than the second threshold and less than the first threshold, the scheduling index of the corresponding repair node is set to a second set value; the second set value is greater than the first set value;

[0139] If the CPU utilization, disk utilization and memory utilization of the repair node are all less than the second threshold, the scheduling index is calculated based on the CPU utilization, disk utilization and memory utilization, and the scheduling index of the corresponding repair node is greater than the second set value.

[0140] In one embodiment, the allocation module is specifically used to:

[0141] If the scheduling index of the repair node is the first set value, no repair task is assigned to the corresponding repair node, and the repair task being executed on the corresponding repair node is stopped;

[0142] If the scheduling index of the repair node is the second set value, no repair task is assigned to the corresponding repair node, and other repair tasks on the corresponding repair node except the repair task being executed are assigned to the repair node whose scheduling index is greater than the second set value;

[0143] If the scheduling index of the repair node is greater than the second set value, a corresponding number of repair tasks are allocated to the corresponding repair node according to the scheduling index.

[0144] In actual application, the determination module and deletion module can be implemented by a processor in an electronic device, such as a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU) or a programmable gate array (FPGA).

[0145] It should be noted that: the data repair device provided in the above embodiment only uses the division of the above modules as an example to illustrate when performing data repair. In actual applications, the above processing can be assigned to different modules as needed, that is, the internal structure of the device is divided into different modules to complete all or part of the processing described above. In addition, the data repair device provided in the above embodiment and the data repair method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0146] The above-mentioned data repair device can be in the form of an image file. After the image file is executed, it can be run in the form of a container or a virtual machine to implement the data repair method described in this application. Of course, it is not limited to the image file form. As long as some software forms that can implement the data repair method described in this application are within the scope of protection of this application.

[0147] Based on the hardware implementation of the above program modules and in order to implement the method of the embodiment of the present application, the embodiment of the present application also provides an electronic device. Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the electronic device includes:

[0148] Communication interface 301, capable of exchanging information with other devices such as network devices;

[0149] The processor 302 is connected to the communication interface 301 to implement information exchange with other devices and is used to execute the method provided by one or more technical solutions when running a computer program. The computer program is stored in the memory 303.

[0150] Of course, in actual application, the various components in the electronic device are coupled together through the bus system 304. It can be understood that the bus system 304 is used to realize the connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 3 Various buses are labeled as bus system 304 .

[0151] The memory 303 in the embodiment of the present application is used to store various types of data to support the operation of the computer device. Examples of such data include: any computer program used to operate on the electronic device.

[0152] It can be understood that the memory 303 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), direct memory bus random access memory (DRRAM). The memory described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memory.

[0153] The method disclosed in the above embodiment of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the program in the memory and completes the steps of the above method in combination with its hardware.

[0154] Optionally, when the processor 302 executes the program, it implements the corresponding processes implemented by the electronic device in each method of the embodiments of the present application, which will not be described in detail here for the sake of brevity.

[0155] In an exemplary embodiment, the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a first memory storing a computer program, and the computer program can be executed by a processor of a computer device to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.

[0156] In the several embodiments provided in the present application, it should be understood that the disclosed devices, computer equipment and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0157] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0158] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0159] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, disks or optical disks.

[0160] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, electronic device, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0161] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program, and the computer program can be executed by the processor 302 of the electronic device to complete the steps described in the data repair method in the embodiment of the present application.

[0162] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0163] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.

[0164] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A data repair method, characterized in that: The method comprises: Determine the shard loss information of the data shards of the same storage object recorded in the first data table and the second data table; the first data table includes the shard loss information of the data shards that failed to be written when the storage object is stored in shards, and the second data table includes the shard loss information of the data shards that cannot be obtained when the disk is scanned; According to the shard loss information of the same storage object recorded in the first data table and the second data table, the shard loss information of the same storage object recorded in the first data table or the second data table is deleted; the first data table and the second data table are used to repair the lost data shards.

2. The method according to claim 1, characterized in that The shard loss information includes an index of a data shard, and deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table or the second data table includes: If the number of indexes of the data shards of the same storage object recorded in the first data table and the second data table is the same, the shard loss information of the same storage object recorded in the first data table or the second data table is deleted.

3. The method according to claim 1, characterized in that The shard loss information includes an index of a data shard, and deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table or the second data table includes: If the index numbers of the data shards of the same storage object recorded in the first data table and the second data table are the same, the shard loss information of the same storage object recorded in the second data table is deleted.

4. The method according to claim 1, characterized in that: The shard loss information includes an index of a data shard, and deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table or the second data table includes: If the index numbers of the data shards of the same storage object recorded in the first data table and the second data table are different, the shard loss information of the same storage object recorded in the first data table is deleted.

5. The method according to any one of claims 2 to 4, characterized in that: The shard loss information includes an index of the data shard and a name of a corresponding storage object, and determining the shard loss information of the data shard of the same storage object recorded in the first data table and the second data table includes: According to the name of the same storage object, search the first data table and the second data table for indexes of the data slices of the same storage object.

6. The method according to claim 1, characterized in that The deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table and the second data table includes: Determine the name of the data table to be operated and the deletion entry in the corresponding data table according to the shard loss information of the same storage object recorded in the first data table and the second data table; Write deletion information to the message queue according to the name of the data table to be operated and the deletion entry in the corresponding data table; Pull a delete message from the message queue, and delete the delete entry in the corresponding data table according to the delete message.

7. The method according to claim 1, characterized in that After deleting the shard loss information of the same storage object recorded in the first data table or the second data table according to the shard loss information of the same storage object recorded in the first data table and the second data table, the method further includes: Generate multiple repair tasks according to the first data table and the second data table, where the repair tasks are used to repair the lost data shards; Determine a scheduling index of each repair node based on the node resource utilization of each repair node among the multiple repair nodes; the scheduling index represents the number of repair tasks that can be performed simultaneously by the repair node; The multiple repair tasks are allocated to each repair node according to the scheduling index of each repair node.

8. The method according to claim 7, characterized in that The node resource utilization includes at least one of CPU utilization, disk utilization, and memory utilization; and determining the scheduling index of each repair node based on the node resource utilization of each repair node among the multiple repair nodes includes: According to the CPU utilization, disk utilization and memory utilization of each repair node, the scheduling index of the corresponding repair node is determined.

9. The method according to claim 8, characterized in that The step of determining the scheduling index of the corresponding repair node according to the CPU utilization, disk utilization, and memory utilization of each repair node includes: If any one of the CPU utilization, disk utilization, and memory utilization of the repair node is greater than the first threshold, setting the scheduling index of the corresponding repair node to the first set value; If any one of the CPU utilization, disk utilization and memory utilization of the repair node is greater than the second threshold and less than the first threshold, the scheduling index of the corresponding repair node is set to a second set value; the second set value is greater than the first set value; If the CPU utilization, disk utilization and memory utilization of the repair node are all less than the second threshold, the scheduling index is calculated based on the CPU utilization, disk utilization and memory utilization, and the scheduling index of the corresponding repair node is greater than the second set value.

10. The method according to claim 9, characterized in that The allocating the plurality of repair tasks to each repair node according to the scheduling index of each repair node comprises: If the scheduling index of the repair node is the first set value, no repair task is assigned to the corresponding repair node, and the repair task being executed on the corresponding repair node is stopped; If the scheduling index of the repair node is the second set value, no repair task is assigned to the corresponding repair node, and other repair tasks on the corresponding repair node except the repair task being executed are assigned to the repair node whose scheduling index is greater than the second set value; If the scheduling index of the repair node is greater than the second set value, a corresponding number of repair tasks are allocated to the corresponding repair node according to the scheduling index.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data repair method according to any one of claims 1 to 10 are implemented.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the data repair method according to any one of claims 1 to 10 are implemented.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the steps of the data repair method according to any one of claims 1 to 10.