Disk fault avoidance method and device, equipment and medium

By regularly detecting and marking the partition status in the distributed message system and repairing the disks with abnormal partitions when necessary, the problem of poor operation and maintenance results in the system in response to disk failures is solved, and lower operation and maintenance costs and higher availability are achieved.

CN119938366APending Publication Date: 2025-05-06CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411731202.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When a distributed message system responds to disk failures, the operation and maintenance effect is poor, resulting in high operation and maintenance costs and great business impact.

Method used

By periodically detecting partition status in distributed message systems, marking exceptions and normal partitions, and repairing disk and updating partition status when the amount of abnormal partition data is empty. When a data write request is received, the data is written to the normal partition to avoid the risk of disk failure.

Benefits of technology

It realizes the operation and maintenance of partitions of distributed message systems at the software level, reduces operation and maintenance costs, improves system availability, and effectively avoids the impact of disk failures on business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938366A_ABST
    Figure CN119938366A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed message systems, and discloses a disk fault avoidance method, device and equipment and a medium, the method comprises the following steps: periodically detecting respective partition states of a plurality of partitions under each data unit in a distributed message system based on a first preset period, and marking abnormal partitions and normal partitions under each data unit; periodically detecting the data volume in the abnormal partition based on a second preset period, and when the data volume is empty, repairing the disk and updating the partition state; when a write-in request of target data is received, obtaining a current normal partition list corresponding to the target data unit; the target partition information in the information list is determined according to the preset rule, and the target data is written into the corresponding target partition, the data is written into the normal partition by periodically detecting the partition state under the data unit, meanwhile, disk repair is conducted on the abnormal partition, and the state is updated, so that the risk that the data is stored in a fault disk is avoided in advance; and the data storage security is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage, and in particular to a disk failure avoidance method, device, equipment and medium. Background Art

[0002] In big data business scenarios, distributed messaging systems are an essential technology due to their high throughput and ease of use. However, disk failures often become a major problem, with high operation and maintenance costs and easy impact on business.

[0003] Common solutions to disk failures include disk mirror arrays (raid1) and distributed parity disk arrays (raid5) at the hardware level, but the hardware cost is too high. Of course, there are also some attempts to provide some automated support for operation and maintenance solutions from the software level, but it is inevitable that it will have an impact on the business system, and the high availability effect remains in theory. Summary of the invention

[0004] In view of this, the present invention provides a disk failure avoidance method, device, equipment and medium to solve the problem of poor operation and maintenance effect when the current distributed message system copes with disk failure.

[0005] In a first aspect, the present invention provides a disk failure avoidance method, which is applied to a distributed coordination service of a distributed messaging system, and the method comprises:

[0006] Based on a first preset period, regularly detecting the partition status of each of the multiple partitions under each data unit in the distributed message system, and marking abnormal partitions and normal partitions under each data unit;

[0007] Based on the second preset period, regularly detecting the amount of data in the abnormal partition, when the amount of data is empty, repairing the disk corresponding to the abnormal partition and updating the corresponding partition status;

[0008] When the distributed messaging system receives a write request for target data, it obtains an information list of current normal partitions of a target data unit corresponding to the target data;

[0009] The target partition information under the information list is determined according to a preset writing rule, and the target data is written into the target partition corresponding to the target partition information.

[0010] The method provided in this aspect determines the abnormal partitions and normal partitions under each data unit by regularly detecting the partition status under the distributed message system, and regularly detecting the data volume of the abnormal partition. When the data volume is empty, the disk of the abnormal partition is promptly repaired and the partition status is updated. Therefore, when a data write request is received, the data is written to the normal partition under the corresponding data unit to ensure that the data write partition is a normal partition, avoid the risk of disk failure in advance, and realize the operation and maintenance of the partitions in the distributed message system at the software level, thereby reducing the operation and maintenance costs and having higher availability, and ensuring the operation and maintenance effect of the distributed message system.

[0011] In an optional implementation, the partition includes: a plurality of replicas; and the detecting the partition status of each of the plurality of partitions under each data unit in the distributed message system includes:

[0012] For each partition under each data unit, obtain the health status of each copy under the partition, and determine whether there is an abnormal copy under the partition;

[0013] If so, the partition status of the partition is determined to be abnormal.

[0014] This implementation method determines whether a partition is abnormal by judging whether there are abnormal copies among the copies corresponding to each partition, so as to avoid storing data in a partition with a faulty copy in advance, thereby avoiding risks in advance and ensuring data security.

[0015] In an optional implementation, different replicas correspond to different disk partitions;

[0016] The repairing of the disk corresponding to the abnormal partition includes:

[0017] The cluster management interface of the distributed messaging system is called to replace the disk partition corresponding to the abnormal replica under the abnormal partition, and the corresponding replica function is executed through the replaced normal disk partition.

[0018] This implementation method replaces the disk partition of the abnormal copy, thereby quickly and efficiently repairing the abnormal copy under the abnormal partition, ensuring that the copy under the partition functions normally, and thus improving the security of subsequent data storage.

[0019] In an optional implementation, the health status of each replica under the partition is obtained in the following manner:

[0020] The health status of each replica under the partition is determined based on the ISR mechanism, the heartbeat mechanism or the preset monitoring tool.

[0021] In this implementation, the health status of each replica is determined through a variety of mechanisms or tools, thereby ensuring the accuracy of identifying the health status of the replica, and further ensuring the effective identification of abnormal partitions.

[0022] In an optional implementation manner, the normal partition information list is generated as follows:

[0023] For each data unit, after marking the abnormal partition and the normal partition under the data unit, the partition information of the currently detected normal partition is randomly sorted to generate an information list of the normal partition under the data unit.

[0024] In this implementation, the partition information of each security partition is randomly sorted to ensure that when writing data according to the information list, the data to be written can be evenly written into each partition under the data unit, thereby ensuring that the data in each partition is balanced.

[0025] In an optional implementation, the data in the partition has a data preservation time limit;

[0026] Before repairing the disk corresponding to the abnormal partition and updating the corresponding partition status, the method further includes:

[0027] The expired data in the abnormal partition is cleaned up regularly according to the data preservation period in the abnormal partition to trigger the condition that the data volume is empty.

[0028] This implementation manner periodically cleans up expired data in the partition so that the data in the abnormal partition can be gradually cleared, thereby eliminating the need for data migration during subsequent disk repairs, thereby avoiding excessive system load caused by data migration after disk repairs.

[0029] In an optional implementation manner, determining the target partition information under the information list according to a preset writing rule includes:

[0030] Obtain the historical write record corresponding to the current information list, and determine the partition information of the last write partition corresponding to the last data write request;

[0031] The next partition information of the partition information last written into the partition in the information list is determined as the target partition information.

[0032] In this implementation mode, the partition to which data was last written is determined through historical write records, so that the current data is written into the next adjacent partition in the information list, thereby realizing polling write and ensuring data balance in each partition.

[0033] In a second aspect, the present invention provides a disk failure avoidance device, which is applied to a distributed coordination service of a distributed messaging system, and the device comprises:

[0034] A partition status detection module, used to periodically detect the partition status of each of the multiple partitions under each data unit in the distributed message system based on a first preset period, and mark abnormal partitions and normal partitions under each data unit;

[0035] A partition disk repair module, used for periodically detecting the amount of data in the abnormal partition based on a second preset period, and when the amount of data is empty, repairing the disk corresponding to the abnormal partition and updating the corresponding partition status;

[0036] An information list acquisition module, used for acquiring an information list of a current normal partition of a target data unit corresponding to the target data when the distributed message system receives a write request for the target data;

[0037] The target data writing module is used to determine the target partition information under the information list according to a preset writing rule, and write the target data into the target partition corresponding to the target partition information.

[0038] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the disk failure avoidance method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0039] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the disk failure avoidance method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 is a flowchart of a disk failure avoidance method according to an embodiment of the present invention;

[0042] Figure 2 is a flow chart of another disk failure avoidance method according to an embodiment of the present invention;

[0043] Figure 3 This is an example diagram of a process for avoiding disk failure when writing data according to an embodiment of the present invention;

[0044] Figure 4 is a structural block diagram of a disk failure avoidance device according to an embodiment of the present invention;

[0045] Figure 5 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0047] In big data business scenarios, distributed messaging systems are an essential technology due to their high throughput and ease of use. However, disk failures often become a major problem, with high operation and maintenance costs and easy impact on business.

[0048] Common solutions to disk failures include disk mirror arrays (raid1) and distributed parity disk arrays (raid5) at the hardware level, but the hardware cost is too high. Of course, there are also some attempts to provide some automated support for operation and maintenance solutions from the software level, but it is inevitable that it will have an impact on the business system, and the high availability effect remains in theory.

[0049] To this end, an embodiment of the present invention provides a disk failure avoidance method, which periodically detects the partition status under a distributed message system to determine the abnormal partitions and normal partitions under each data unit, so that when a data write request is received, the data is written to the normal partition under the corresponding data unit to ensure that the data write partition is a normal partition, thereby avoiding the risk of disk failure in advance. At the same time, the data volume of the abnormal partition is periodically detected. When the data volume is empty, the disk of the abnormal partition is repaired in time and the partition status is updated to avoid the system load caused by the need for data migration after the repair.

[0050] According to an embodiment of the present invention, a disk failure avoidance method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0051] In this embodiment, a disk failure avoidance method is provided, which can be used for the distributed coordination service of the above-mentioned distributed messaging system. Figure 1FIG. 1 is a flow chart of a disk failure avoidance method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0052] Step S101: regularly detecting the partition status of each of the multiple partitions under each data unit in the distributed message system based on a first preset period, and marking abnormal partitions and normal partitions under each data unit.

[0053] In a distributed messaging system, Topic, or data unit, is a classification label for messages. Producers publish messages to specific data units, and consumers can subscribe to the data units of interest to receive messages. It is similar to a "channel" for messages, and different types of messages can be divided into different data units according to business logic. Partition, or partition, is a further subdivision of a data unit. A data unit can contain multiple partitions, which is the basic unit for physically storing messages.

[0054] In order to timely understand the partition status of different partitions under each data unit, the partition status of different partitions under each data unit can be detected at intervals of a first preset period, such as 10 minutes, and each partition can be marked, that is, a partition with a normal partition status is marked as a normal partition, and a partition with an abnormal partition status is marked as an abnormal partition. When detecting the partition status, certain measurement indicators can be judged through a pre-set detection tool to determine the status of the partition.

[0055] Step S102: regularly detecting the amount of data in the abnormal partition based on a second preset period, and when the amount of data is empty, repairing the disk corresponding to the abnormal partition and updating the corresponding partition status.

[0056] For distributed systems, users usually pre-set the data retention period in each partition. When the data expires, the data is cleaned up to avoid data accumulation in the partition that makes it impossible to receive new data.

[0057] The method provided by the present invention will write data into the normal partition when writing data. For details, please refer to the subsequent steps and will not be repeated here. Therefore, for the abnormal partition, no new data will be written into it, and other users can still read the data in it that has not expired. However, since no new data is written, as the data continues to expire, the amount of data in the abnormal partition will become less and less, and the final amount of data will be empty.

[0058] Therefore, the amount of data in the abnormal partition can be regularly checked every second preset period. The specific duration of the second preset period can be set according to actual needs. When the amount of data inside is empty, the disk corresponding to the abnormal partition is repaired to avoid excessive system load caused by data migration after repair, so as to ensure the smoothness of system operation. At the same time, after completing the disk repair of the abnormal partition, its partition status will be updated, that is, its partition status will be adjusted to normal, so as to facilitate the subsequent writing of data to the partition that has completed the disk repair.

[0059] Step S103: When the distributed messaging system receives a write request for target data, it obtains an information list of the current normal partitions of the target data unit corresponding to the target data.

[0060] When the distributed messaging system receives a target data write request, that is, the target data needs to be written to a certain partition, for the target data, the specific data unit to be written can be determined according to the specific data type of the data, that is, the target data unit. For the target data unit, it usually has multiple partitions inside. As mentioned in the above step S101, the status of the partitions under each data unit will be checked every first preset period, and the normal partitions and abnormal partitions under each data unit will be marked.

[0061] Therefore, after receiving a write request for the target data, it is necessary to determine which normal partitions are under the target data unit corresponding to the target data, that is, to obtain an information list of the current normal partitions of the target data unit. The information list lists the partition information corresponding to each normal partition, such as partition ID, partition name, etc., so as to facilitate the subsequent determination of the specific partition to which the target data is to be written in the information list.

[0062] Step S104, determining the target partition information in the information list according to a preset writing rule, and writing the target data into the target partition corresponding to the target partition information.

[0063] After obtaining the information list of normal partitions under the current target data unit, the target partition information can be determined in the information list according to a preset rule, and the target data can be written into the partition corresponding to the target partition information.

[0064] The preset rules for determining the target partition information can be set according to actual conditions. For example, a polling writing method is adopted to sequentially determine the target partition information and write data according to the order of the partition information corresponding to each partition in the information list.

[0065] For example, the current information list records the partition IDs of each partition and sorts them, where the specific sorting is: 0001, 0002, 0003, 0004. These numbers represent the partition IDs corresponding to different partitions. Assume that the last data written is the partition corresponding to the partition ID 0002; according to the sorted polling writing method, the target partition information corresponding to the current target data is 0003, that is, the target data needs to be written to the partition corresponding to 0003, and so on. When the partition to which the last data was written is the partition corresponding to the last partition ID in the information list, the partition ID of the target partition is determined from the beginning.

[0066] In another example, for the partition information list, in addition to recording the partition ID of each partition, the current data storage status of each partition can also be recorded: for example: "0001, 60%", "0002, 50%", "0003, 40%", "0004, 50%", where the percentage after the partition ID represents the current data load status of the partition with the corresponding ID. When writing the target data this time, "0003, 40%" can be used as the target partition information, and the target data can be written into the partition corresponding to 0003.

[0067] The specific writing rules can be set according to the actual situation. The above examples are only exemplary implementations and are not limited here.

[0068] The disk failure avoidance method provided in this embodiment detects the partition status under the distributed message system regularly, determines the abnormal partitions and normal partitions under each data unit, and regularly detects the data volume of the abnormal partition. When the data volume is empty, the disk of the abnormal partition is repaired in time and the partition status is updated to avoid the system load caused by the need for data migration after the repair. In addition, when a data write request is received, the data is written to the normal partition under the corresponding data unit to ensure that the data write partition is a normal partition, avoiding the risk of disk failure in advance. The partition operation and maintenance in the distributed message system is realized at the software level, which reduces the operation and maintenance costs and has higher availability, ensuring the operation and maintenance effect of the distributed message system.

[0069] According to an embodiment of the present invention, another disk failure avoidance method embodiment is provided, which can be used for the distributed coordination service of the above-mentioned distributed messaging system. Figure 2 FIG. 4 is a flow chart of another disk failure avoidance method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0070] Step S201: regularly detecting the partition status of each of the multiple partitions under each data unit in the distributed message system based on a first preset period, and marking abnormal partitions and normal partitions under each data unit.

[0071] Specifically, a partition includes: multiple copies.

[0072] It can be understood that in a distributed system, for each partition, it usually has multiple copies, which store the same data, so that when a copy fails, other copies can be used as backup to avoid the partition being directly unusable. These copies correspond to different disk partitions, that is, in a physical sense, different copies correspond to disks in different locations.

[0073] In the above step S201, detecting the partition status of each of the multiple partitions under each data unit in the distributed message system includes:

[0074] Step S201-1, for each partition under each data unit, obtain the health status of each replica under the partition, and determine whether there is an abnormal replica under the partition;

[0075] Step S201 - 2: If yes, the partition status of the partition is determined to be abnormal.

[0076] It can be understood that when the partition status under the data unit is detected every first preset period, for each partition under each data unit, whether the partition is abnormal is mainly determined based on the health status of each replica corresponding to each partition. The health status of the replica can be measured based on the data synchronization status between the replicas. For example, for a partition, it has three replicas, among which there is a leader replica and two follower replicas. When writing data, the leader replica will be written first, and the follower replica will then grab data from the leader replica to achieve data synchronization. The health of each replica can be measured based on whether the time required for the follower replica to synchronize data meets the requirements.

[0077] If there is an abnormal replica under a partition, it means that there is a problem with the data synchronization of the replica under the partition. Therefore, the status of the partition is determined to be abnormal, and no data will be written to it subsequently until its status returns to normal.

[0078] Furthermore, in step S201-1, the health status of each replica under the partition is obtained in the following manner:

[0079] Determine the health status of each replica under the partition based on the ISR mechanism, heartbeat mechanism or preset monitoring tools.

[0080] The ISR mechanism mainly determines whether the data synchronization of the follower replica is normal based on the data lag between each follower replica and the leader replica. Each replica of the heartbeat mechanism periodically sends a heartbeat signal to other nodes. The heartbeat signal contains some basic information about the replica, such as message offset, replica status and other indicators. If the component performing heartbeat detection or other replica nodes does not receive the heartbeat signal within a certain period of time, it is considered that the replica has an abnormality. You can also use some well-written third-party monitoring tools to monitor different indicators of each replica to determine whether the replica is abnormal. When monitoring, you can mainly monitor indicators such as synchronization rate and message lag. The specific monitoring tool can be set according to the actual situation, which will not be repeated here.

[0081] Step S202: for each data unit, after marking the abnormal partitions and normal partitions under the data unit, randomly sort the partition information of the currently detected normal partitions to generate an information list of the normal partitions under the data unit.

[0082] The above step S202 is the method for generating the information list of normal partitions. It can be understood that for each data unit, after completing the detection of its internal partition status at intervals and marking the normal partitions and abnormal partitions therein, all normal partitions under the data unit will be counted and randomly sorted to obtain the information list corresponding to all normal partitions under the data unit. The partition information corresponding to each normal partition is recorded in the information list, such as the partition ID or partition name, and the order between these partition information is randomly sorted, so that the balance of data volume between each partition can be guaranteed when data is written later. When it is necessary to obtain the information list of the current normal partition of the target data unit in the subsequent step S204, it can be directly called.

[0083] Step S203: regularly detecting the amount of data in the abnormal partition based on a second preset period, and when the amount of data is empty, repairing the disk corresponding to the abnormal partition and updating the corresponding partition status.

[0084] Specifically, the data in the partition has a data preservation period;

[0085] Before repairing the disk corresponding to the abnormal partition and updating the corresponding partition status, the method further includes:

[0086] The expired data in the abnormal partition is cleaned up regularly according to the data preservation period in the abnormal partition to trigger the condition that the data volume is empty.

[0087] It can be understood that for a distributed messaging system, the data in each partition is time-sensitive, that is, the data in the partition has a data preservation period, and the data will be deleted when the period expires. Therefore, for an abnormal partition, the data inside it will be continuously cleaned up due to the data's own preservation period, and combined with the mechanism of no longer writing data inside the abnormal partition in the embodiment of the method, the condition that the data volume in the abnormal partition is empty can be triggered to perform disk repair on the abnormal partition.

[0088] Furthermore, different replicas correspond to different disk partitions;

[0089] Repair the disk corresponding to the abnormal partition, including:

[0090] The cluster management interface of the distributed messaging system is called to replace the disk partition corresponding to the abnormal replica under the abnormal partition, and the corresponding replica function is executed through the replaced normal disk partition.

[0091] It can be understood that for a partition, the corresponding multiple copies are used to store the same data, and the corresponding physical carriers, i.e., disk partitions, of these copies are different. Repairing the disk corresponding to an abnormal partition means repairing the disk corresponding to the abnormal copy under the abnormal partition. When repairing, a normal disk partition can be replaced as the physical carrier of the corresponding copy to perform the corresponding copy function.

[0092] For example, a partition 0001 corresponds to three replicas A, B, and C. These replicas are used to store the same data. The disk partitions corresponding to these three replicas are: 0000a, 0000b, and 0000c. If replica C of partition 0001 is abnormal, then find a normal disk partition 0000e, migrate the functions and fixed data corresponding to replica C to the disk partition 0000e, and repair the disk corresponding to the abnormal partition. At this time, the disk partitions corresponding to the three replicas A, B, and C are 0000a, 0000b, and 0000e respectively.

[0093] Step S204: When the distributed messaging system receives a write request for target data, it obtains an information list of the current normal partitions of the target data unit corresponding to the target data.

[0094] When a write request for target data is received, the target data unit corresponding to the target data is first determined, that is, according to the specific type of the target data, it is determined which data unit the target data needs to be stored in, and the information list of the current normal partition of the target data unit is obtained. Based on the information list of normal partitions under different data units generated in the above step S202, it can be directly called when needed.

[0095] Step S205 , determining the target partition information in the information list according to a preset writing rule, and writing the target data into the target partition corresponding to the target partition information.

[0096] Specifically, in step S205, the target partition information under the information list is determined according to the preset writing rule, including:

[0097] Step S205 - 1 , obtaining the historical write record corresponding to the current information list, and determining the partition information of the last write partition corresponding to the last data write request.

[0098] Step S205-2: determine the next partition information of the partition information last written into the partition in the information list as the target partition information.

[0099] It can be understood that when writing data, a polling writing method is adopted, that is, for the current information list, the data to be written is written into the partitions corresponding to different partition information in the information list one by one according to the order in the information list.

[0100] In the specific implementation, the historical write record corresponding to the current information list can be obtained first, which records the partition corresponding to which partition information the data received in the previous few times was written to for the current information list. For example, the current information list is: 0001, 0002, 0003, 0004, 0005, where these numbers represent the partition IDs of different partitions. In the historical write record, it is recorded that the last data was written to the partition with partition ID 0003. Then, for the target data this time, the partition with partition ID 0005 is written, so that the data received each time is written to different partitions in sequence. If the data received last time is written to the partition corresponding to the last partition ID in the information list, then for the target data this time, it is written from the beginning, that is, the partition corresponding to the first partition ID in the information list is written.

[0101] In an example, more information types can be set in the information list according to actual conditions to facilitate the execution of more complex data writing rules. For example, in addition to the partition ID, each row of information can also include the current data volume and current congestion level of the corresponding partition, so as to further set the writing rules based on the data conditions of each partition. The specific setting method can be set based on actual conditions and will not be repeated here.

[0102] The disk failure avoidance method provided by the embodiment of the present invention determines the abnormal partitions and normal partitions under each data unit by regularly detecting the partition status under the distributed message system, so that when a data write request is received, the data is written to the normal partition under the corresponding data unit to ensure that the data write partition is a normal partition, thereby avoiding the risk of disk failure in advance. At the same time, the data volume of the abnormal partition is regularly detected. When the data volume is empty, the disk of the abnormal partition is repaired in time and the partition status is updated to avoid the system load caused by the need for data migration after the repair.

[0103] At the same time, when judging whether there are abnormal copies in the copies corresponding to each partition, the health status of the copies can be detected based on a variety of mechanisms or tools to ensure the accuracy of identifying abnormal partitions, avoid storing data in partitions with faulty copies, and avoid risks in advance to ensure data security.

[0104] After marking the partition status, when generating the information list of the normal partition, the partition information of each security partition is randomly sorted to ensure that when writing data according to the information list, the data to be written can be evenly written to each partition under the data unit to ensure the data balance in each partition.

[0105] It should be noted that this method can be applied to distributed coordination services of a variety of distributed messaging systems, such as Kafka, RabbitMQ, ActiveMQ, and self-developed distributed messaging systems based on specific application scenarios. For these distributed messaging systems, they are optional but not mandatory. There is no restriction on the distributed messaging system for specific applications, and you can choose according to actual needs.

[0106] In order to facilitate the understanding of the above invention embodiments, the Kafka distributed messaging system is taken as an example to assist in understanding the implementation of the above method embodiments. The overall implementation process is as follows:

[0107] First, the simplest and most efficient local broker is used to specify the partition write mechanism to avoid writing to the partition of the faulty disk and ensure that the new data of the application is written normally, that is, the data is written to the normal partition under a certain unit.

[0108] When writing data, the Kafka cluster interface information is read to generate a partition list of the target topic with a normal number of copies of the local broker. Specifically, in order to reduce performance consumption, the partition list under the topic can be initially generated every 1 minute. At the same time, in order to better ensure the balanced effect of partition writing, the order of the partition list is randomly shuffled each time it is generated.

[0109] When writing data, you can use a simple polling method to select one partition in the partition list for writing each time you request; within the same cycle, when the partition list is used up, continue to use it from the beginning in a cyclic manner.

[0110] In order to ensure that each new data is written to a partition with a normal number of replicas and avoid using a partition with a disk failure, the partition with a disk failure problem will be "suspended", that is, no data will be written, and only the read function will be executed. When the data preservation period has passed and the amount of data in the partition drops to zero, the disk corresponding to the partition will be repaired, that is, the disk partition corresponding to the abnormal replica of the partition will be repaired to reduce the system load caused by data migration after the repair.

[0111] In order to repair the disk in time when the data in the abnormal partition is cleared, the amount of partition data that needs to be repaired can be checked periodically. When the number of partitions is zero, for the partition using the faulty disk, the completely normal broker's replica migration plan is used to call the Kafka cluster management interface to automatically repair its replica, and the partition is migrated to a normal available disk at almost zero cost, which will not have any impact on the business. Apply the repaired partition and write data. According to the above mechanism, use the partition with a completely normal replica, independently encapsulate the interface, and call it for all applications, effectively controlling the application transformation cost.

[0112] Specifically, you can combine Figure 3 As shown in the figure, it is an example diagram of a process for avoiding disk failure during data writing according to an embodiment of the present invention. For the Kafka data writing interface, the Kafka data writing module determines the status of different partitions in each unit according to the management interface, that is, the status of different partitions, and determines the partition without a complete copy book as abnormal, and does not serve as a data writing target. Other partitions are polled for writing. Figure 3 As shown, under a certain unit, there are 6 partitions, among which the partition status of partition 4 is abnormal, that is, it is an abnormal partition. The Kafka data writing module does not allocate data to it. The Kafka replica repair module detects the data volume of partitions with insufficient number of replicas. When the data volume within the period is zero, a new broker is selected for the partition according to the management interface to arrange repair processing so that the abnormal partition returns to normal.

[0113] The embodiment of the present invention aims to realize a high-availability system design for business that avoids the impact of disk failures and a low-cost system solution combined with automated operation and maintenance means based on the technical application scenarios of distributed messaging systems. It can ensure a "completely self-healing" effect that does not adversely affect online business systems for common disk failures and operation and maintenance processing operations.

[0114] The above-mentioned embodiment avoids the use of faulty disks from the aspect of software application, which is low-cost compared with the raid solution for disks from the hardware level. It will not cause possible disk IO surge problems due to the need to repair the partition with a certain amount of data due to the loss of copies caused by disk failure in the operation and maintenance solution, which may affect the normal writing of business data. After the amount of partition data caused by the faulty disk is 0, the partition is repaired at a low cost, and the partition can be automatically re-enabled after the repair, which greatly ensures the full utilization of cluster resources. For the impact of disk failure, the operation is controlled to the minimum, the operation and maintenance cost is reduced, and the high availability requirements of the application are maximized, so that high availability is implemented in practice.

[0115] In this embodiment, a disk failure avoidance device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0116] This embodiment provides a disk failure avoidance device, such as Figure 4 As shown, including:

[0117] The partition status detection module 401 is used to periodically detect the partition status of each of the multiple partitions under each data unit in the distributed message system based on a first preset period, and mark abnormal partitions and normal partitions under each data unit.

[0118] The partition disk repair module 402 is used to regularly detect the amount of data in the abnormal partition based on the second preset period, and when the amount of data is empty, repair the disk corresponding to the abnormal partition and update the corresponding partition status;

[0119] The information list acquisition module 403 is used to acquire the information list of the current normal partition of the target data unit corresponding to the target data when the distributed message system receives a write request for the target data.

[0120] The target data writing module 404 is used to determine the target partition information under the information list according to a preset writing rule, and write the target data into the target partition corresponding to the target partition information.

[0121] In some optional implementations, the partition includes: a plurality of replicas;

[0122] The partition status detection module 401, when detecting the partition status of each of the multiple partitions under each data unit in the distributed message system, includes:

[0123] For each partition under each data unit, obtain the health status of each replica under the partition and determine whether there are abnormal replicas under the partition;

[0124] If so, the partition status of the partition is determined to be abnormal.

[0125] In some optional implementations, different replicas correspond to different disk partitions;

[0126] The partition disk repair module 402, when repairing the disk corresponding to the abnormal partition, includes:

[0127] The cluster management interface of the distributed messaging system is called to replace the disk partition corresponding to the abnormal replica under the abnormal partition, and the corresponding replica function is executed through the replaced normal disk partition.

[0128] In some optional implementations, the health status of each replica under a partition is obtained in the following manner:

[0129] Determine the health status of each replica under the partition based on the ISR mechanism, heartbeat mechanism or preset monitoring tools.

[0130] In some optional implementations, the information list of a normal partition is generated in the following manner:

[0131] For each data unit, after the partition state detection module 401 marks the abnormal partitions and normal partitions under the data unit, the partition information of the currently detected normal partitions is randomly sorted to generate an information list of the normal partitions under the data unit.

[0132] In some optional implementations, data in a partition has a data preservation time limit;

[0133] The partition disk repair module 402, before repairing the disk corresponding to the abnormal partition and updating the corresponding partition status, regularly cleans up the expired data in the abnormal partition according to the data preservation time in the abnormal partition to trigger the condition that the data volume is empty.

[0134] In some optional implementations, the target data writing module 404, when determining the target partition information under the information list according to the preset writing rule, includes:

[0135] Obtain the historical write record corresponding to the current information list, and determine the partition information of the last write partition corresponding to the last data write request;

[0136] The next partition information of the partition information last written into the partition in the information list is determined as the target partition information.

[0137] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0138] The disk failure avoidance device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0139] The embodiment of the present invention also provides a computer device having the above Figure 4 The disk failure avoidance device shown.

[0140] See also Figure 5 , Figure 5 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.

[0141] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0142] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0143] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0144] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0145] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.

[0146] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0147] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A disk failure avoidance method, applied to a distributed coordination service of a distributed messaging system, characterized in that: The method comprises: Based on a first preset period, regularly detecting the partition status of each of the multiple partitions under each data unit in the distributed message system, and marking abnormal partitions and normal partitions under each data unit; Based on a second preset period, regularly detecting the amount of data in the abnormal partition, when the amount of data is empty, repairing the disk corresponding to the abnormal partition and updating the corresponding partition status; When the distributed messaging system receives a write request for target data, it obtains an information list of current normal partitions of a target data unit corresponding to the target data; The target partition information under the information list is determined according to a preset writing rule, and the target data is written into the target partition corresponding to the target partition information.

2. The method according to claim 1, characterized in that: The partition includes: a plurality of replicas; the detecting the partition status of each of the plurality of partitions under each data unit in the distributed message system includes: For each partition under each data unit, obtain the health status of each copy under the partition, and determine whether there is an abnormal copy under the partition; If so, the partition status of the partition is determined to be abnormal.

3. The method according to claim 2, characterized in that Different copies correspond to different disk partitions; The repairing of the disk corresponding to the abnormal partition includes: The cluster management interface of the distributed messaging system is called to replace the disk partition corresponding to the abnormal replica under the abnormal partition, and the corresponding replica function is executed through the replaced normal disk partition.

4. The method according to claim 2, characterized in that: The health status of each replica under the partition is obtained as follows: The health status of each replica under the partition is determined based on the ISR mechanism, the heartbeat mechanism or the preset monitoring tool.

5. The method according to claim 1, characterized in that The information list of the normal partition is generated in the following manner: For each data unit, after marking the abnormal partition and the normal partition under the data unit, the partition information of the currently detected normal partition is randomly sorted to generate an information list of the normal partition under the data unit.

6. The method according to claim 1, characterized in that The data in the partition has a data preservation period; Before repairing the disk corresponding to the abnormal partition and updating the corresponding partition status, the method further includes: The expired data in the abnormal partition is cleaned up regularly according to the data preservation period in the abnormal partition to trigger the condition that the data volume is empty.

7. The method according to claim 1, characterized in that The step of determining the target partition information under the information list according to the preset writing rule includes: Obtain the historical write record corresponding to the current information list, and determine the partition information of the last write partition corresponding to the last data write request; The next partition information of the partition information last written into the partition in the information list is determined as the target partition information.

8. A disk failure avoidance device, applied to a distributed coordination service of a distributed messaging system, characterized in that: The device comprises: A partition status detection module, used to periodically detect the partition status of each of the multiple partitions under each data unit in the distributed message system based on a first preset period, and mark abnormal partitions and normal partitions under each data unit; A partition disk repair module, used for periodically detecting the amount of data in the abnormal partition based on a second preset period, and when the amount of data is empty, repairing the disk corresponding to the abnormal partition and updating the corresponding partition status; An information list acquisition module, used for acquiring an information list of a current normal partition of a target data unit corresponding to the target data when the distributed message system receives a write request for the target data; The target data writing module is used to determine the target partition information under the information list according to a preset writing rule, and write the target data into the target partition corresponding to the target partition information.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the disk failure avoidance method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the disk failure avoidance method according to any one of claims 1 to 7.