Multi-magnetic-arm disk fault processing method and device and computing equipment
By identifying the fault type and strategy in multi-arm disk failure handling and flexibly reconstructing using redundant copy data, the problem of data loss caused by multi-arm disk failure is solved, achieving efficient and reliable fault handling and storage system stability.
Patent Information
- Application Number
- CN202511202938.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional fault handling methods are prone to data loss in multi-arm disk environments because one disk in a multi-arm disk can correspond to multiple logical volumes/disks. Traditional methods only perform backup and disk replacement operations on the failed LUN/Disk, resulting in the loss of some data.
This paper provides a method for handling multi-arm disk failures. By determining the failure repair strategy and failure type, it flexibly selects partial reconstruction or full disk reconstruction, uses redundant copy data to reconstruct data to the backup storage unit, avoids data loss caused by blind disk replacement, and ensures the stability of the storage system by monitoring sub-health indicators and capacity expansion warnings.
It effectively avoids data loss, improves the reliability and efficiency of multi-arm disk failure handling, reduces performance loss, and ensures the stable operation of the storage system and efficient utilization of resources.
Smart Images

Figure CN120909831A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of storage, and in particular to a multi-actuator hard disk fault processing method and device and computing equipment. BACKGROUND
[0002] With the rapid development of high-performance computing (HPC), big data, artificial intelligence (AI), etc., the demand for data storage of enterprises is growing explosively, and higher requirements are put forward for the timeliness of data processing. Under this background, storage technology evolves towards large capacity and high performance. Multi-actuator hard disks (Multi-Actuator HDD) are widely used because they can balance storage capacity and read-write performance. Compared with traditional single-actuator hard disks (Single-Actuator HDD), multi-actuator hard disks of the same capacity have better storage performance and significantly reduce the storage cost of high-load scenarios such as HPC and big data analysis.
[0003] However, when a disk fails, the traditional fault processing method is mainly designed for single-actuator hard disks. In a single-actuator hard disk, one disk corresponds to one logical volume / disk (LUN / Disk). In the fault processing process, the traditional algorithm usually only backs up the data of the LUN / Disk that fails, and then performs a disk replacement operation.
[0004] However, when this traditional fault processing method is directly applied to a multi-actuator hard disk environment, many problems will arise. A significant feature of a multi-actuator hard disk is that one disk can correspond to multiple LUNs / Disk (Disk: LUN / Disk = 1: N). Therefore, after performing a disk replacement operation on a multi-actuator hard disk that has failed, the data in some LUNs / Disk will be lost with the replacement of the physical disk, resulting in poor reliability. SUMMARY
[0005] The embodiments of the present application provide a multi-actuator hard disk fault processing method and device and computing equipment to solve the problem that a multi-actuator hard disk fault easily leads to data loss.
[0006] In a first aspect, the embodiments of the present application provide a multi-magnetic-arm disk fault processing method, applied to a computing device, the computing device being in communication connection with a distributed storage system; the distributed storage system comprising a plurality of storage nodes, each of the storage nodes comprising a plurality of multi-magnetic-arm disks, each of the multi-magnetic-arm disks comprising a plurality of storage units; the method comprising: in response to a fault event triggered by a faulty storage unit, determining a fault repair strategy; based on the fault repair strategy and / or a fault type of the faulty storage unit, determining a data reconstruction manner and performing reconstruction; wherein the data reconstruction manner comprises: reconstructing data in the faulty storage unit, or reconstructing all data in a target multi-magnetic-arm disk to which the faulty storage unit belongs; the reconstruction refers to writing data in the storage unit into a standby storage unit by using redundant copy data, the standby storage unit and the faulty storage unit belonging to different multi-magnetic-arm disks.
[0007] The method can flexibly select local reconstruction or full-disk reconstruction when a fault event occurs. If the fault event is caused by local disk failure, the method can only reconstruct the data of the storage unit corresponding to the faulty magnetic arm, reduce the amount of data migration to reduce performance loss, or can reconstruct the entire disk data to improve reliability. If the fault event is caused by global disk failure, the method can reconstruct the entire disk data to ensure reliability. Moreover, the method can prevent data loss caused by blind disk replacement, and eliminate the negative impact of directly performing disk replacement.
[0008] In a possible implementation, the fault type comprises a first fault type and a second fault type; the first fault type refers to that the fault event is caused by local disk failure, and the second fault type refers to that the fault event is caused by global disk failure. In this way, targeted fault processing can be performed based on the fault.
[0009] In a possible implementation, the data reconstruction manner is determined and reconstruction is performed based on the fault repair strategy and / or the fault type of the failed storage unit, including: in a case where the fault repair strategy is a performance priority strategy, real-time performance data of a target storage node is obtained, the target storage node being a storage node to which the failed storage unit belongs, and the real-time performance data including network delay, CPU usage, memory usage, and / or disk usage of all multi-arm disks in the target storage node; in a case where the real-time performance data is less than a preset performance threshold, the fault type is determined based on the fault event; in a case where the fault type is a first fault type, data of the failed storage unit is reconstructed to a standby storage unit at a first reconstruction speed by using redundant copy data, the first reconstruction speed being a data read-write speed; a sub-health flag is set for the failed storage unit, and a no-disk-replacement flag is set for the target multi-arm disk; in a case where the fault type is a second fault type, data in the target multi-arm disk is reconstructed to one or more standby storage units at the first reconstruction speed by using the redundant copy data, and a disk-replacement flag is set for the target multi-arm disk.
[0010] Based on this, reconstruction can be triggered only when the load of the target storage node is lower than the threshold, thereby avoiding aggravating the system burden.
[0011] In a possible implementation, the data reconstruction manner is determined and reconstruction is performed based on the fault repair strategy and / or the fault type of the failed storage unit, including: in a case where the fault repair strategy is a performance priority strategy, real-time performance data of a target storage node is obtained, the target storage node being a storage node to which the failed storage unit belongs, and the real-time performance data including network delay, CPU usage, memory usage, and / or disk usage of all multi-arm disks in the target storage node; in a case where the real-time performance data is less than a preset performance threshold, the fault type is determined based on the fault event; in a case where the fault type is a first fault type, data of the failed storage unit is reconstructed to a standby storage unit at a first reconstruction speed by using redundant copy data, the first reconstruction speed being a data read-write speed; a sub-health flag is set for the failed storage unit, and a no-disk-replacement flag is set for the target multi-arm disk; in a case where the fault type is a second fault type, data in the target multi-arm disk is reconstructed to one or more standby storage units at the first reconstruction speed by using the redundant copy data, and a disk-replacement flag is set for the target multi-arm disk; in a case where the fault repair strategy is a cost priority strategy, the fault type is determined based on the fault event; in a case where the fault type is the first fault type, data of the failed storage unit is reconstructed to a standby storage unit at a second reconstruction speed by using redundant copy data, the second reconstruction speed being a data read-write speed; a sub-health flag is set for the failed storage unit, and a no-disk-replacement flag is set for the target multi-arm disk; in a case where the fault type is the second fault type, data in the target multi-arm disk is reconstructed to one or more standby storage units at the second reconstruction speed by using the redundant copy data, and a disk-replacement flag is set for the target multi-arm disk.
[0012] Under the cost priority strategy, data can be quickly recovered by using the second reconstruction speed. Moreover, the disk-replacement flag can be set only when it is confirmed that the disk is in a whole failure state, thereby avoiding hardware replacement and achieving efficient use of resources.
[0013] In a possible implementation, based on the fault event, a target multi-arm disk to which the failed storage unit belongs is determined; it is checked whether a storage unit other than the failed storage unit in the target multi-arm disk is in a fault state; if the storage unit other than the failed storage unit in the target multi-arm disk is not in the fault state, it is determined that the fault type is a first fault type; if the storage unit other than the failed storage unit in the target multi-arm disk is in the fault state, it is determined that the fault type is a second fault type.
[0014] In this way, whether the target multi-magnetic arm disk has a fault can be accurately determined.
[0015] In a possible implementation, the multi-magnetic arm disks in the plurality of storage nodes are divided into one or more storage pools; the method further includes: traversing each storage pool according to a preset period, calculating a total capacity of all the fault storage units marked with the sub-health flag in each storage pool to obtain a fault storage capacity of the storage pool; and adding an expansion flag to the storage pool when the fault storage capacity exceeds a preset percentage threshold of a total capacity of the storage pool.
[0016] In this way, by periodically monitoring the total capacity of the sub-health storage units, potential risks of the storage pool can be found in time. When the fault storage capacity exceeds the threshold, an expansion warning is triggered, so that insufficient available space or a decrease in data reliability caused by accumulation of sub-health units can be avoided, and stable operation of the storage system can be ensured.
[0017] In a possible implementation, based on the fault repair strategy and / or the fault type of the fault storage unit, a data reconstruction manner is determined and reconstruction is performed, and the method further includes: in a case where the fault repair strategy is a reliability-first strategy, data in the target multi-magnetic arm disk is reconstructed to one or more spare storage units at a third reconstruction speed through redundant copy data, and a disk replacement flag is set for the target multi-magnetic arm disk, the third reconstruction speed being a data read-write speed.
[0018] Under the reliability-first strategy, data in the target multi-magnetic arm disk can be quickly reconstructed at a higher third reconstruction speed, so that the probability of data loss caused by the fault can be reduced.
[0019] In a possible implementation, the method further includes: in response to adding the multi-magnetic arm disk to the first storage pool, balancing data in the first storage pool to the newly added storage unit through a migration operation.
[0020] Based on this, data distribution can be maintained to be balanced, a hot disk risk can be reduced, and overall storage efficiency and stability can be ensured.
[0021] In a second aspect, the embodiments of the present application provide a multi-magnetic-arm disk fault processing apparatus, applied to a computing device, the computing device being in communication connection with a distributed storage system; the distributed storage system comprising one or more storage pools, the storage pool comprising a plurality of multi-magnetic-arm disks, the multi-magnetic-arm disk comprising a plurality of storage units; the apparatus comprising: a determination module configured to determine a fault repair strategy in response to a fault event triggered by a faulty storage unit; and a fault processing module configured to determine a data reconstruction manner and perform reconstruction based on the fault repair strategy and / or a fault type of the faulty storage unit; wherein the data reconstruction manner comprises: reconstructing data in the faulty storage unit, and reconstructing all data in a target multi-magnetic-arm disk to which the faulty storage unit belongs; and the reconstruction refers to writing data in the storage unit into a standby storage unit using redundant copy data, the standby storage unit and the faulty storage unit belonging to different multi-magnetic-arm disks.
[0022] In a third aspect, the embodiments of the present application further provide a computing device, comprising: one or more processors; and a memory configured to store one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of the above first aspect.
[0023] In a fourth aspect, the embodiments of the present application provide a chip, the chip being used to execute the method of any one of the above first aspect.
[0024] In a fifth aspect, the embodiments of the present application provide a computer readable storage medium, the computer readable storage medium storing computer execution instructions, the computer execution instructions being executed by a computer to implement the method of any one of the first aspect.
[0025] In a sixth aspect, the embodiments of the present application provide a program product, comprising a computer program, the computer program being executed by a processor to implement the method of any one of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 A system architecture diagram of a distributed storage system provided by the embodiments of the present application;
[0027] Figure 2 An application scenario diagram of a distributed storage system provided by the embodiments of the present application;
[0028] Figure 3 A schematic diagram of data splitting and replication provided by the embodiments of the present application;
[0029] Figure 4 A first flowchart of a multi-magnetic-arm fault processing method provided by the embodiments of the present application;
[0030] Figure 5A schematic diagram of disk failure and arm failure provided for an embodiment of the present application;
[0031] Figure 6 A schematic diagram of data reconstruction provided for an embodiment of the present application;
[0032] Figure 7 A second flowchart of a multi-arm failure processing method provided for an embodiment of the present application;
[0033] Figure 8 A flowchart of monitoring a storage pool provided for an embodiment of the present application;
[0034] Figure 9 A structural schematic diagram of a multi-arm disk failure processing apparatus provided for an embodiment of the present application;
[0035] Figure 10 A schematic diagram of a computing device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings. In order to clearly describe the technical solutions in the embodiments of the present application, the first, second, etc. descriptions in the embodiments of the present application are only used for indicating and distinguishing the description objects, and do not have the order, nor represent the special limitation of the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.
[0037] Before introducing the technical solutions in the embodiments of the present application, the terms involved in the embodiments of the present application are first described by way of example.
[0038] 1. Distributed storage system: a storage system that stores data in multiple independent storage nodes, which can be physical servers or devices in a cluster.
[0039] 2. Multiple-Copy algorithm: a strategy used in a distributed storage system to ensure data reliability, which is a data slice redundancy protection mechanism. The core idea of the Multiple-Copy algorithm is to store multiple copies of data in different storage locations, so that even if some copy data is lost, other copies can still provide data access services.
[0040] Figure 1 A system architecture diagram of a distributed storage system provided for an embodiment of the present application.
[0041] As Figure 1As shown, the embodiment of the present application provides a distributed storage system, which includes a plurality of storage nodes (Node). The storage node is a basic physical unit in the distributed storage system. Each storage node can be composed of one or more server hardware. From the aspect of form, the server can be a rack server or a whole cabinet server. From the aspect of performance, the server can be a general server, a GPU (graphics processing unit) server, or an AI (artificial intelligence) server.
[0042] Further, the server can be built-in with a plurality of storage medium interfaces (such as SATA, SAS, NVMe, etc.) for connecting hard disk storage devices. The hard disk storage devices can include multi-armed disks and single-armed disks. Specifically, the single-armed disk contains one magnetic arm (also known as a head arm) and a plurality of platters. Among them, the platter is the core storage component of the hard disk storage device and is the physical medium that carries data. In the single-armed disk, all platters share one magnetic arm, and a plurality of heads are installed on the magnetic arm. Each platter has a corresponding head on its upper and lower surfaces. When the magnetic arm moves, all heads will move synchronously in the radial direction (i.e., all heads point to the same track position of each platter at the same time). In this way, the single-armed disk presents a single, indivisible logical storage unit to the upper layer application, i.e., a logical volume / disk (LUN / Disk).
[0043] The multi-armed disk includes a plurality of independent magnetic arms, each corresponding to a group of independent heads and platters. Each magnetic arm can move independently, and the heads also operate independently. That is, the heads of different magnetic arms can access different tracks of different platters at the same time without synchronization, and have higher parallelism. Based on this, the multi-armed disk presents a plurality of independently accessible logical storage units to the upper layer application, i.e., a plurality of logical volumes / disks (LUN / Disk). These logical storage units can be regarded as a plurality of independent "sub-disks" by the upper layer application, and the upper layer application can simultaneously perform read / write operations on these logical storage units through different interfaces or paths. Due to the parallel working characteristics of the multi-armed disk, when the upper layer application processes high-concurrency data requests, the parallel IO capability of the multi-armed disk can be fully utilized, greatly improving the data read / write efficiency. Especially in the random read / write scenario, the performance advantage is more obvious.
[0044] In actual application, the number of storage nodes included in a distributed storage system, and the number of single-armed disks and / or multi-armed disks contained in each storage node, depend on actual configuration, and the embodiment of the present application does not make specific limitation.
[0045] It should be noted that the distributed storage system can implement data redundancy protection in a "multi-copy" manner. Specifically, when data is written, the data can be divided into at least one data stripe, a data stripe contains a plurality of data strips, and then a plurality of groups of data copies are generated based on the data stripe, and each group of data copies contains a plurality of corresponding data strips. The storage of these data strips will be described in detail below, and will not be repeated here.
[0046] Figure 2 An application scenario diagram of the distributed storage system provided by the embodiments of the present application is shown.
[0047] As shown in Figure 2 , the embodiments of the present application also provide a client, which is an entrance for interacting with a user or an application program, and is used to convert the requirements of an upper-layer business (such as writing a file, reading data, deleting a record, etc.) into specific requests (such as API calls, protocol instructions, etc.) for the distributed storage system. Then, the client communicates with the distributed system through a network or a switch. After receiving the request of the client, the distributed storage system cooperates with multiple storage nodes to complete the processing and storage of data, and returns the result to the client. Alternatively, the client can directly generate explicit operation instructions based on the requirements of the upper-layer business. These instructions not only contain data itself, but also carry specific storage strategy parameters, such as the number of copies, the selection of a storage pool (Pool), the size of a data strip, etc. At this time, the distributed storage system only needs to complete the write, read or delete operation based on the operation instruction, and does not need to analyze or infer the requirements of the upper-layer business again.
[0048] Further, the application of the client and the distributed storage system includes a deployment phase and an application phase. The deployment phase includes but is not limited to:
[0049] ① Client environment configuration, such as installing client software, configuring a communication protocol (such as S3, NFS) with the distributed storage system, establishing an identity authentication mechanism (such as a key, a token), and presetting a commonly used storage strategy template.
[0050] ② Hardware deployment of the distributed storage system, including storage node construction, initialization of the storage node, construction of a network switch, creation of a storage pool, and configuration of data strip rules, etc.
[0051] The construction phase of the storage node is a physical preparation process at the hardware level, specifically referring to assembling various hardware components constituting the storage node into a runnable physical unit. For example, a rack-mounted server is selected as the node carrier, and basic computing components such as CPU, memory, and motherboard are installed; a single-arm disk and / or multi-arm disk is inserted, a network interface card (NIC) is connected to ensure that the node can access the external network; the assembled server is deployed to the data center cabinet (i.e., "on the rack"), and the power cord and network cable are connected, and the physical fixation is completed.
[0052] After hardware configuration, each storage node can include a certain number of single-arm disks and / or multi-arm disks. For example, Node 1 includes 12 single-arm disks and 12 multi-arm disks, Node 2 includes 12 single-arm disks and 12 multi-arm disks, and Node M includes 12 single-arm disks and 12 multi-arm disks. It can be seen that the disk configuration of each storage node can be the same, and the number of single-arm disks and multi-arm disks contained in each storage node is the same.
[0053] In the storage node initialization phase, the embodiments of the present application can automatically scan and identify all disks (single-arm disks and / or multi-arm disks) in the storage node and collect hardware information. For example, the disk type (single-arm / multi-arm), the globally unique name (WWN), the serial number (SN), and the arm status (for multi-arm disks, recording the normal and fault status of each arm) are collected, and the LUN / Disk list corresponding to each disk is formed. For example, the LUN / Disk list entry of a single-arm disk can be: [LUN1 (Disk1): capacity 16TB, normal state, corresponding to arm 0 (unique arm)]; the LUN / Disk list entry of a multi-arm disk can be: [LUN1 (Disk1): capacity 8TB, normal state, corresponding to arm 0; LUN2 (Disk2): capacity 8TB, normal state, corresponding to arm 1], wherein "LUN" represents a logical volume.
[0054] After that, the embodiments of the present application can write this information into a certain field of the system metadata table to form a "disk-storage unit-storage node" mapping relationship, which can be used as a basis for judging "the disk to which the storage unit belongs" and "whether the storage unit is faulty".
[0055] After that, the network switch can be built to ensure smooth communication links between the client and the node and between the nodes.
[0056] Further, in the storage pool creation stage, multiple single-arm disks or multiple multi-arm disks can be divided into the same storage pool, one storage pool can contain the disks of multiple storage nodes, and the disks on one storage node can also be divided into different storage pools. For example, according to the business type, Pool-A (multi-arm high-performance pool), Pool-B (single-arm capacity pool), Pool-C (multi-arm high-performance pool), and the like are divided, and the number of copies of the storage pool, the balancing strategy are set. The number of copies is also referred to as the multi-copy protection level, the multi-copy redundancy, or the multi-copy redundancy ratio. For example, Pool-A is set to 3 copies, Pool-B is set to 2 copies, or the number of copies can be set based on different businesses, for example, application A is set to 3 copies, and application B is set to 2 copies, which is not limited in the embodiments of the present application. The embodiments of the present application can write these configurations into the storage pool attribute field of the metadata table.
[0057] Further, the embodiments of the present application can also configure data striping rules. For example, the default slice size is configured to be 4MB, 8MB or 16MB, so that data strips of 4MB, 8MB or 16MB size can be obtained.
[0058] Further, the application stage of the client and the distributed storage system includes but is not limited to: the client receives the business request of the upper application program in real time (such as the batch writing of the data set of the AI training platform, the real-time stream storage of the video monitoring system, the incremental data addition of the log system, etc.), and generates operation instructions dynamically according to the business characteristics. The distributed storage system writes, reads or deletes data based on the operation instructions.
[0059] Further, in order to solve the problem that the multi-arm disk failure is easy to cause data loss, the embodiments of the present application provide a multi-arm disk failure processing method. Based on the multi-arm disk failure processing method, the "arm failure" and the "disk failure" can be clearly distinguished, and the failure repair strategy can be flexibly adjusted based on the failure condition, for example, whether to replace the disk or not. The method can be executed by a computing device. The computing device can be one of the storage nodes of the distributed storage system, or can be a management node in the distributed storage system, in other words, the method provided by the embodiments of the present application can be executed by one of the storage nodes in the distributed storage system, or can be uniformly executed by the management node in the distributed storage system. In some implementation manners, the computing device can also be a server hardware running a client. The server can be a rack server or an entire cabinet server in form, and can be a general server, a GPU server or an artificial intelligence server in performance. At this time, the multi-arm disk failure processing method can be built into the client in the form of a software module. Hereinafter, the computing device is taken as an example of one of the storage nodes or the management node in the distributed storage system for example.
[0060] Specifically, the computing device can be communicatively connected with a distributed storage system, which can include a plurality of storage nodes, each of which is composed of a plurality of multi-armed disks and / or a plurality of single-armed disks, in combination with the foregoing embodiments.
[0061] A multi-armed disk has two or more arms working simultaneously, and the plurality of arms can read or write data in parallel. Since the plurality of arms can independently process input / output (I / O) requests, one physical multi-armed disk can be split into a plurality of logical units (LUNs). For example, DiskA (a multi-armed disk) is divided into two logical units (LUN1 and LUN2). Since these logical units appear to the operating system or application as independent physical disks, that is, the multi-armed disk presents a plurality of logical units or disks (LUN / Disk) to the upper operating system or upper application. Therefore, one multi-armed disk can include a plurality of logical units / disk (LUN / Disk), that is, the correspondence between the multi-armed disk and the LUN / Disk is Disk:LUN / Disk = 1:q (q≥2).
[0062] A single-armed disk includes one LUN / Disk, and the correspondence between the single-armed disk and the LUN / Disk is Disk:LUN / Disk = 1:1.
[0063] It can be understood that "LUN / Disk" refers to a logical storage unit presented by a disk to an upper application. A single-armed disk presents a single, indivisible logical storage unit, that is, one LUN / Disk, to an upper application. A multi-armed disk presents a plurality of independently accessible logical storage units, that is, a plurality of LUNs / Disk, to an upper application.
[0064] It should be noted here that LUN / Disk is collectively referred to as "storage unit" in the following, which is only for the convenience of description and does not constitute any limitation.
[0065] Continuing to refer to Figure 1 For example, the distributed storage system can include M storage nodes composed of multi-armed disks, which are storage node 1 (Node 1), storage node 2 (Node 2), storage node 3 (Node 3),..., and storage node M (Node M), respectively. Each storage node can include two multi-armed disks, and each multi-armed disk can include two storage units. In addition, the distributed storage system can include K storage nodes composed of single-armed disks, which are storage node M+1 (Node (M+1)) to storage node M+K (Node (M+K)), respectively. Each storage node can include four single-armed disks, and each single-armed disk includes one storage unit.
[0066] The multiple multi-platter disks in the multiple storage nodes are divided into one or more storage pools, and / or the multiple single-platter disks in the multiple storage nodes are divided into one or more storage pools. It can be understood that the storage pools can be divided in the configuration stage of the client and the distributed storage system.
[0067] Continuing to refer to Figure 1 For example, the M storage nodes composed of multi-platter disks can all belong to the same storage pool, and the K storage nodes composed of single-platter disks can all belong to the same storage pool.
[0068] It should be further noted that, based on the distributed storage system provided in the embodiments of the present application, the to-be-written data can be stored in a multi-copy easy manner.
[0069] Figure 3 The schematic diagram of data splitting and replication provided in the embodiments of the present application.
[0070] For example, as Figure 3 shown, the to-be-written data can be a 24 MB log file, and then the 24 MB data is split by 6 MB, and 4 data stripes can be generated:
[0071] Data stripe 1 (Stripe_001): contains data stripe 1-data stripe 4 (Strip_001 to Strip_004). It should be noted that in the embodiments of the present application, the data splitting step can be performed by the client or one of the distributed storage nodes.
[0072] The embodiments of the present application can redundantly generate N copies (for example, 2-8 copies) of the data stripe (multiple data stripes) through the multi-copy technology. For example, when the multi-copy redundancy is 3 copies, the embodiments of the present application can construct 2 copies that are completely the same based on the original data, and finally there are 3 copies of the original data+2 copies.
[0073] It should be noted that the first copy can be obtained after data splitting, that is, the copy is created based on the original data slice. Then, through the copying operation, N-1 copies can be obtained, and finally N groups of the same data copies are obtained, and each group of data copies includes multiple data stripes. The multiple data stripes in each group of data copies can completely express the to-be-written data.
[0074] In the embodiments of the present application, the replication step can be performed by the client or one of the distributed storage nodes.
[0075] Hereinafter, a 3-copy example is used for illustration. For example:
[0076] Data stripe 1 (Stripe_001) generates two sets of copies, Copy1 and Copy2, each containing copies of Strip_001 to Strip_004.
[0077] For ease of description, the embodiments of this application collectively refer to the original data and the copied data as multiple sets of data copies. In practical applications, the original data and copied data can be distinguished by different first identifiers, which can be temporarily recorded in the metadata table. The first identifier is, for example, an identifier of "strip ID + copy number + data stripe ID". The copy number of the original data can be "copy0". Then, the second identifier of the first data stripe of the original data of data stripe 1 is: Stripe_001_Copy0_Strip_001, and the first identifier of the first data stripe in the first copy of data stripe 1 is: Stripe_001_Copy1_Strip_001'. In this way, different data stripes in different copies can be distinguished.
[0078] It is understandable that in each data strip, copy0, copy1, and copy2 are redundant copies of each other.
[0079] Furthermore, each data stripe in each data copy of each data stripe can be stored in a storage unit. The specific writing process can be adjusted according to requirements, and this application embodiment does not impose specific limitations on it.
[0080] The multi-magnetic arm fault handling method provided in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0081] Figure 4 This is a first flowchart illustrating the multi-magnetic arm fault handling method provided in an embodiment of this application.
[0082] like Figure 4 As shown, the multi-magnetic arm fault handling method provided in this application embodiment may include the following steps S100-S200.
[0083] S100: In response to a fault event triggered by a faulty storage unit, determine a fault repair strategy.
[0084] Among them, a faulty storage unit refers to a storage unit that cannot provide normal data storage or reading services due to hardware abnormalities (such as read / write failure of the magnetic arm, loose interface, etc.) or logical errors (such as data verification failure). Furthermore, a faulty storage unit can be one of the LUNs / Disks of a multi-magnetic-arm disk or a LUN / Disk of a single-magnetic-arm disk.
[0085] Further, the embodiments of the present application can detect the health status of each storage unit in real time or at a preset frequency. Specifically, the storage node can be used to monitor the health status of the storage unit to discover the storage unit failure in time. Specifically, the storage node can continuously collect disk health status data through self-monitoring, analysis and reporting technology (SMART), and perform threshold monitoring on key indicators of the storage unit, such as read-write response time, error count, read-write error rate, head seek time, temperature, and the like. For example, when the number of consecutive read-write timeouts of a certain storage unit exceeds 5 times, or the CRC check error rate reaches 0.1%, a failure event is generated to trigger a failure warning, and the failure event can carry the disk ID (which can be WWN, SN), LUN / Disk number, timestamp of the failure occurrence, and the like of the failed storage unit.
[0086] In some implementations, the steps of monitoring the failure status and generating the failure event can be performed by a real-time monitoring module in the storage node, which is not specifically limited in the present application.
[0087] Further, the failure repair strategy includes a multi-copy failure handling strategy, including a performance priority strategy, a reliability priority strategy, and / or a cost priority strategy.
[0088] The failure repair strategy can be set in the storage pool construction phase. Specifically, in each storage pool initialization phase, the embodiments of the present application can configure the failure handling strategy according to the business requirements and / or resource conditions of the storage pool, and use it as the default strategy of the storage pool. That is, what preset failure strategy the storage pool corresponds to is already fixed in the configuration phase.
[0089] In the embodiments of the present application, the performance priority strategy can mean that the failure data reconstruction work is performed without affecting the storage system business. The reliability priority can mean that the failure influence is completely eliminated through global reconstruction and triggering of disk replacement actions to ensure the reliability of data storage. The cost priority strategy can mean that the disk replacement is avoided in the case of arm failure, and the additional resource consumption in the failure handling process is reduced.
[0090] S200: determining a data reconstruction manner and performing reconstruction based on the failure repair strategy and / or the failure type of the failed storage unit; wherein the data reconstruction manner includes: reconstructing the data in the failed storage unit, and reconstructing all data in the target multi-arm disk to which the failed storage unit belongs; the reconstruction means writing the data in the storage unit into a standby storage unit using redundant copy data, and the standby storage unit and the failed storage unit belong to different multi-arm disks.
[0091] It can be understood that the target multi-platter disk includes the failed storage unit, that is, the failed storage unit belongs to the target multi-platter disk, is one of the logical storage units in the target multi-platter disk, and the target multi-platter disk is one of the physical disks in the target storage pool. Further, the target storage pool to which the failed storage unit belongs can be pre-configured. For example, during the initialization stage of the client and the distributed storage system, a certain number of multi-platter disks are aggregated into the same storage pool. That is, whether the storage pool uses multi-platter disks or single-platter disks is determined in the configuration stage and is fixed.
[0092] Further, the redundant copy data can be pre-stored in other storage units in the target storage pool, and the other storage units and the failed storage unit belong to different multi-platter disks.
[0093] It should be noted here that which storage unit is the other storage unit is determined by the redundant copy distribution rule. The redundant copy distribution rule is, for example, a cross-disk isolation principle. The cross-disk isolation principle means that different data copies of the same data stripe must be stored on different physical disks.
[0094] For example, the three copy distribution rules of Stripe_001 can be:
[0095] Copy0: stored in storage unit 1 of Disk A (belongs to the disk where the failed storage unit is located);
[0096] Copy1: stored in storage unit 3 of Disk B (other storage unit);
[0097] Copy2: stored in storage unit 6 of Disk C (other storage unit);
[0098] Wherein, Disk A, Disk B and Disk C represent three different multi-platter disks in the target storage pool.
[0099] It should be noted that the above redundant copy distribution rule is only an exemplary introduction, and the specific distribution of the redundant copy data can be adjusted based on actual needs, for example, distributing different data stripes of the same data copy in different multi-platter disks or different storage units, respectively storing different data copies of the same data stripe in different storage nodes, etc.
[0100] Furthermore, reconstruction can refer to the process of restoring data to a secure storage unit based on redundant copies after a multi-arm disk failure. The core is to reconstruct lost, unusable, or security-risk data using existing copies, ensuring data integrity and reliability. Even in the event of hardware failure, data loss can be prevented, maintaining the normal operation of upper-layer applications. In this embodiment, selecting other storage units belonging to different multi-arm disks than the failed storage unit for data reconstruction ensures data accuracy and integrity.
[0101] Furthermore, in the embodiments of this application, when a failure event occurs, the data in the failed storage unit or the target multi-arm disk corresponding to the failed storage unit can be reconstructed to a backup storage unit, where the backup storage unit and the failed storage unit belong to different multi-arm disks. This avoids data loss and prevents reconstruction failures due to failures of other arms on the target multi-arm disk.
[0102] It is worth noting that the embodiments of this application can flexibly choose to reconstruct the data in the faulty storage unit or the target multi-arm disk. This is because the types of faults that cause the fault event are different. For example, the fault event may be caused by a drive failure or a disk failure; drive failure and disk failure are not the same. A drive failure may only affect the read and write operations of a single faulty storage unit, while a disk failure will cause the entire disk to fail. Reconstruction based on the reconstruction method provided in the embodiments of this application can achieve optimization in terms of reliability, cost, performance, and other aspects. If it is a drive failure, only the storage unit data corresponding to the faulty drive can be reconstructed to reduce the amount of data migration and reduce performance loss. Alternatively, the entire disk data can be reconstructed to improve reliability. If it is a disk failure, the entire disk data can be reconstructed to ensure reliability and balance efficiency in fault handling.
[0103] Figure 5 This is a schematic diagram of disk failure and read / write head failure provided in an embodiment of this application.
[0104] like Figure 5 As shown in the embodiments of this application, a magnetic arm failure can refer to a failure in the logical storage unit corresponding to a certain magnetic arm, while other magnetic arms and their associated storage units on the same multi-magnetic arm disk can still function normally. For example, storage unit 1 of Disk A cannot be read or written, but storage unit 2 can still provide stable service; this situation can be called a magnetic arm failure. It should be noted that a magnetic arm failure can refer to a storage unit failure caused by a hardware failure of the magnetic arm, or it can refer to a storage unit failure caused by a hardware failure of the disk surface corresponding to the magnetic arm. In other words, "magnetic arm failure" is only a general term for various "non-full disk failures," and does not mean that a "magnetic arm failure" is necessarily caused by a problem with the magnetic arm.
[0105] In this case, the embodiment of the present application can choose to only reconstruct the failed storage unit, or can choose to reconstruct all the data in the multi-armed disk.
[0106] The disk failure can refer to disk physical damage, interface complete failure, etc. The disk failure can cause all the arms and storage units of the entire multi-armed disk to be unable to work, that is, the disk failure refers to that all the corresponding logical storage units are faulty and cannot work normally. At this time, all the data in the target multi-armed disk must be reconstructed. Continue to refer to Figure 5 For example, Disk B is completely offline due to internal circuit failure, and the storage unit 3 and the storage unit 4 contained in Disk B cannot work. At this time, all the data stored by Disk B needs to be read from the redundant copy and reconstructed to multiple standby storage units to ensure that all the data can be normally accessed through the new storage location, avoiding data loss caused by the complete failure of the disk.
[0107] Further, the embodiment of the present application can set a disk replacement flag for the target multi-armed disk after global reconstruction. In this way, the user can be notified to perform the disk replacement operation.
[0108] In some implementations, step S200 can include steps S201-S202.
[0109] S201: In the case where the target storage pool to which the failed storage unit belongs contains a multi-armed disk, and in the case where the failure event is triggered by the arm failure of the target multi-armed disk, reconstruct the data of the failed storage unit to the standby storage unit through the redundant copy data;
[0110] Or, in the case where the target storage pool to which the failed storage unit belongs contains a multi-armed disk, and in the case where the failure event is triggered by the arm failure of the target multi-armed disk, reconstruct the data in the target multi-armed disk corresponding to the failed storage unit to one or more standby storage units through the redundant copy data, and set a disk replacement flag for the target multi-armed disk.
[0111] In the embodiment of the present application, in the case where the target storage pool to which the failed storage unit belongs adopts a multi-armed disk, if the failure event is triggered by the arm failure of the target multi-armed disk, the embodiment of the present application can perform a local reconstruction operation to only reconstruct the data in the failed storage unit, so as to achieve the purpose of saving cost.
[0112] Figure 6 The schematic diagram of data reconstruction provided by the embodiment of the present application is shown.
[0113] For example, Figure 6As shown in (a), the storage unit 1 in Disk A has a fault, and the storage unit 2 has no fault, and only the data of the storage unit 1 can be reconstructed to the storage node 3.
[0114] Alternatively, in the case where the target storage pool to which the faulty storage unit belongs adopts a multi-arm disk, if the fault event is triggered by the arm fault of the target multi-arm disk, the embodiment of the present application can also perform a global reconstruction operation to reconstruct all the data in the target multi-arm disk corresponding to the faulty storage unit, so as to improve the reliability of data protection.
[0115] Continuing to refer to Figure 6 As shown in (a), the storage unit 5 in Disk C has a fault, and the storage unit 6 has no fault, and the data in Disk C can be reconstructed.
[0116] S202: In the case where the target storage pool to which the faulty storage unit belongs contains a multi-arm disk, and in the case where the fault event is triggered by the disk fault of the target multi-arm disk, the data in the target multi-arm disk corresponding to the faulty storage unit is reconstructed to one or more standby storage units through redundant copy data, and a disk replacement flag is set for the target multi-arm disk.
[0117] For example, as shown in (b), Disk B has a disk fault, and the data in the storage unit 3 and the storage unit 4 (both belong to the faulty storage unit) can be reconstructed to the storage node 3. In this way, all the data in the target multi-arm disk can be reconstructed, and data loss can be completely avoided. Figure 6
[0118] In some implementations, when the multi-arm disk has a fault, the fault event can be triggered by any storage unit in the faulty disk, for example, a certain storage unit is in a write state, and the write failure triggers the fault event, while another storage unit is in a dormant state and cannot trigger the fault event because it does not perform I / O operation.
[0119] In summary, the embodiment of the present application does not directly notify the user to replace the disk when the disk has a fault, that is, the disk replacement instruction is not immediately issued. Instead, it is first determined whether the target storage pool to which the faulty storage unit belongs adopts a multi-arm disk, and then targeted data reconstruction is performed in the case where the target storage pool to which the faulty storage unit belongs adopts a multi-arm disk, such as local reconstruction or full-disk reconstruction. Furthermore, the embodiment of the present application can set the disk replacement flag after all the disks are reconstructed, to notify the user to replace the disk. In this way, after the disk is replaced, the data in the storage unit in the multi-arm disk that has no fault will not be lost. For example, as shown in Figure 6 As shown in (b), the data of the non-faulty storage unit 6 will not be lost after the disk C is replaced. It can be seen that the embodiment of the present application provides a fault processing method for a multi-arm disk, which can timely perform data reconstruction when a fault event occurs, and can prevent the data loss caused by blind disk replacement, and eliminate the negative effects caused by directly performing disk replacement.
[0120] From the above, it can be seen that the embodiment of the present application provides a multi-arm disk fault processing method, which is applied to a computing device, the computing device is in communication connection with a distributed storage system; the distributed storage system includes a plurality of storage nodes, the storage node includes a plurality of multi-arm disks, and the multi-arm disk includes a plurality of storage units; the method includes: in response to a fault event triggered by a faulty storage unit, determining a fault repair strategy; based on the fault repair strategy and / or the fault type of the faulty storage unit, determining a data reconstruction mode and performing reconstruction; wherein the data reconstruction mode includes: reconstructing the data in the faulty storage unit, and reconstructing all data in the target multi-arm disk to which the faulty storage unit belongs; the reconstruction refers to writing the data in the storage unit into a standby storage unit by using the redundant copy data, and the standby storage unit and the faulty storage unit belong to different multi-arm disks. The method can flexibly select local reconstruction or full-disk reconstruction when a fault event occurs, and the method can prevent the data loss caused by blind disk replacement, and eliminate the negative effects (such as data loss) caused by directly performing disk replacement.
[0121] The process of data reconstruction will be described in detail below with reference to the accompanying drawings.
[0122] Figure 7 The second flowchart of the multi-arm fault processing method provided by the embodiment of the present application is shown.
[0123] As Figure 7 shown, step S200 can include steps S203-S207.
[0124] S203: In the case that the fault repair strategy is a performance priority strategy, real-time performance data of the target storage node is acquired.
[0125] The target storage node is the storage node to which the faulty storage unit belongs. It can be understood that the storage node is built based on server hardware, a network interface card, a multi-arm disk and / or a single-arm disk and other hardware facilities. Therefore, in order to evaluate the performance state of the storage node, the embodiment of the present application can count the following real-time performance data: network delay, CPU usage, memory usage and / or disk usage of all multi-arm disks in the target storage node.
[0126] Specifically, the network delay can refer to the communication response time of the storage node with other storage nodes or clients, for example, the network delay can refer to the time interval from sending a data request to receiving a response from the storage node, which can be in milliseconds (ms).
[0127] The CPU usage rate can refer to the proportion of the CPU of the storage node occupied in a unit of time, which can be expressed in percentage (%). The memory usage rate can refer to the proportion of the memory of the storage node occupied, which can be expressed in percentage (%). The disk usage rate can refer to the proportion of the used storage space to the total storage space, which can be expressed in percentage (%). For example, the total capacity of a multi-arm disk is 1 TB, and the amount of stored data is 600 GB, then the disk usage rate is 60%.
[0128] In some implementations, the total usage rate of all disks can be measured uniformly as the disk usage rate, and then subsequent steps are continued to be performed, or the usage rate of each disk can be measured respectively, and then subsequent steps are performed based on the disk usage rate of each disk, which is not limited in the embodiments of the present application.
[0129] It should be noted that the storage node can be a hybrid storage node composed of multiple multi-arm disks and multiple single-arm disks, and the embodiments of the present application can comprehensively evaluate the performance status of the storage node, and count the usage rate of all disks, instead of counting the usage rate of the multi-arm disk only, which can achieve the purpose of considering the usage of all disks comprehensively. This is because: the process of data reconstruction will occupy the resources of the storage node, and if the usage rate of the multi-arm disk is counted only and the usage rate of the single-arm disk is not counted, the situation that the data reconstruction operation affects the read-write efficiency of the single-arm disk is likely to occur.
[0130] S204: In the case that the real-time performance data is less than the preset performance threshold, determining a fault type based on the fault event.
[0131] The fault type includes a first fault type and a second fault type; the first fault type refers to that the fault event is caused by a local disk fault, which corresponds to the aforementioned "arm fault", and the second fault type refers to that the fault event is caused by a global disk fault, which corresponds to the aforementioned "disk fault".
[0132] Based on this step, it can be determined whether the fault event is triggered by a disk fault or an arm fault. If the number of failed storage units in the target multi-arm disk < the number of arms (or the total number of storage units), then the target multi-arm disk does not exist disk fault, but exists arm fault. If the number of failed storage units in the target multi-arm disk = the number of arms (or the total number of storage units), then the target multi-arm disk exists disk fault. The specific judgment steps will be described in detail below, which will not be repeated here.
[0133] Further, the embodiment of the present application can determine whether the real-time performance data is less than the preset performance threshold value through the relationship between each real-time performance index and each preset performance threshold value. Specifically, the CPU usage, the memory usage, and the disk usage of all multi-armed disks in the target storage node can correspond to a first threshold value, for example, 70%, 75%, or 80%, and the network delay can correspond to a second threshold value, for example, 1 ms. The first threshold value and the second threshold value can be set based on actual conditions.
[0134] For example, if the CPU usage, the memory usage, and the disk usage of all multi-armed disks in the target storage node are less than 70% and the network delay is less than 1 ms, it can be determined that the real-time performance data is less than the preset performance threshold value.
[0135] It should be noted that if the condition is not met, that is, the real-time performance data is greater than or equal to the preset performance threshold value, it indicates that the current distributed storage system corresponds to a relatively busy service, at this time, the data reconstruction will affect the service performance, and the service performance should be prioritized, and the data reconstruction condition is not met. If the condition is met, that is, the real-time performance data is less than the preset performance threshold value, it indicates that the distributed storage system is in a service light load state, and the data reconstruction work can be further performed.
[0136] S205: In the case where the fault type is the first fault type, the data of the faulty storage unit is reconstructed to the standby storage unit at a first reconstruction speed through the redundant copy data.
[0137] The embodiment of the present application can reconstruct only the data in the faulty storage unit in the case where the storage unit has a fault but the disk is not completely damaged. For example, as shown in (a) of FIG. 1, the embodiment of the present application can reconstruct only the data in the storage unit 1. Figure 6
[0138] Further, the first reconstruction speed refers to the data read-write speed. The embodiment of the present application can limit the data read-write speed in the data reconstruction process by setting the first reconstruction speed, so as to achieve the purpose of flow control, to avoid occupying too many resources in the reconstruction process and affecting the key service, and to ensure the performance priority.
[0139] In one implementation manner, the first reconstruction speed can be a preset value, for example, the first reconstruction speed is equal to 20 MB / s, 30 MB / s, or 40 MB / s, which is not limited by the embodiment of the present application.
[0140] In some implementation manners, the first reconstruction speed can be equal to 50% of the average data read-write speed of the normal service data of the storage node. For example, in the daily operation of the storage node, the average data read-write speed of the key service is 80 MB / s, and the first reconstruction speed can be preset to 40 MB / s.
[0141] In an implementation, the reconstruction speed can also be dynamically adjusted according to the real-time performance data of the storage node, for example, when the IO requests of the key service surge, the first reconstruction speed is automatically reduced to a preset value, for example, from 100 MB / s to 50 MB / s, and then restored to the normal speed when the service load falls, so as to balance the reconstruction efficiency and service performance.
[0142] S206: Set a sub-health flag for the failed storage unit, and set a no-replacement flag for the target multi-armed disk.
[0143] The sub-health flag and the no-replacement flag can be two fields maintained in the metadata table of the distributed storage system. For example, the sub-health flag is a "SubHealth" field, and the no-replacement flag is a "NeedReplace" field. By setting the value of the field, the corresponding state of the flag can be changed.
[0144] Specifically, when the value of the "SubHealth" field is "N", it can indicate that the failed storage unit is not in a sub-healthy state. When the value of the "SubHealth" field is "Y", it can indicate that the failed storage unit is in a sub-healthy state. When the value of the "NeedReplace" field is "N", it can indicate no replacement, and when the value of the "NeedReplace" field is "Y", it can indicate replacement.
[0145] As can be seen, the embodiments of the present application can set a sub-health flag for the failed storage unit (the value of the "SubHealth" field is "Y") and set a no-replacement flag for the target multi-armed disk (the value of the "NeedReplace" field is "N") by modifying the value of the field.
[0146] It should be noted that the "SubHealth" field and the "NeedReplace" field described above are only an exemplary introduction, and in actual application, the sub-health flag, the replacement flag, and the no-replacement flag can be represented by other fields or means.
[0147] It can be understood that the storage unit not being in a sub-healthy state indicates that the storage unit is not faulty, and correspondingly, the storage unit being in a sub-healthy state indicates that the storage unit is faulty. That is, after setting a sub-health flag for the failed storage unit, the sub-health flag can reflect that the target multi-armed disk is in a sub-healthy state, achieving the purpose of reminding the user that the target multi-armed disk has hidden dangers. Correspondingly, the "no-replacement" flag can be used to remind the user that the multi-armed disk does not need to be replaced, and the "replacement" flag can be used to remind the user that the multi-armed disk is abnormal and needs to be replaced.
[0148] In some implementations, the sub-health indicator can be set on the target multi-arm disk rather than the faulty storage unit to indicate that the target multi-arm disk has a "magnetic arm failure".
[0149] It should be noted that step S206 can be executed during or after the refactoring process.
[0150] S207: In the case of the second type of failure, the data in the target multi-arm disk is reconstructed to one or more spare storage units at a first reconstruction speed using redundant copy data, and a disk swap flag is set for the target multi-arm disk.
[0151] This application embodiment can reconstruct all data in a target multi-arm disk even when the disk is completely damaged. For example... Figure 6 As shown in (b), all the data in storage units 3 and 4 of Disk B can be reconstructed.
[0152] In some implementations, when a disk replacement flag is set for the target multi-arm disk, a sub-health flag may not be set for the faulty storage unit (in which case the value of the "SubHealth" field is "N"), or a sub-health flag may be set for all storage units in the target multi-arm disk (in which case the value of the "SubHealth" field is "Y"). This application does not specifically limit this.
[0153] In some implementations, such as Figure 6 As shown in (b), when the subhealth flag is set on the multi-arm disk rather than the faulty storage unit, this embodiment of the application can set a disk replacement flag for the target multi-arm disk without setting a subhealth flag for the target multi-arm disk (at this time, the value of the "SubHealth" field is "N").
[0154] By performing a full disk reconstruction on a multi-arm disk, the problem of data loss caused by disk failure can be solved.
[0155] Furthermore, the "disk swapping indicator" can be used to remind users to perform disk swapping operations on multi-arm disks, so as to fundamentally solve disk failures and restore the normal operation of the distributed storage system in a timely manner.
[0156] As can be seen, when the fault repair strategy prioritizes performance, the embodiments of this application can first measure the performance status of the storage node, perform the reconstruction operation under light business load, and limit the rate of the reconstruction operation to fully ensure the performance of critical business.
[0157] It should be noted that in the case of disk failure, even if the failure event is triggered by a certain storage unit of the target multi-armed disk, it does not mean that other storage units of the target multi-armed disk do not have failures or will not trigger failure events.
[0158] With reference to the foregoing description of the embodiment of the application, Figure 7 , step S200 can further include the following step S208.
[0159] S208: In the case of a reliability priority strategy for the failure repair strategy, reconstruct the data in the target multi-armed disk to one or more spare storage units at a third reconstruction speed through redundant copy data, and set a disk replacement flag for the target multi-armed disk.
[0160] In the embodiment of the application, if the multi-copy failure processing strategy is configured as a "reliability priority" strategy, whether it is a disk failure or an arm failure, the target multi-armed disk with a failed storage unit is reconstructed, that is, the data in the normal storage unit also needs to be reconstructed.
[0161] In the specific reconstruction process, the reconstruction speed can be high. Specifically, the embodiment of the application can reconstruct at a third reconstruction speed, and the third reconstruction speed specifically refers to the data read / write speed in the data reconstruction process. Moreover, the embodiment of the application can configure the third reconstruction speed to limit the flow.
[0162] Further, the third reconstruction speed can be greater than the first reconstruction speed to achieve the purpose of fast reconstruction and improve the reliability of data recovery. For example, when the first reconstruction speed is equal to 40 MB / s, the third reconstruction speed can be preset to 80 MB / s or 100 MB / s.
[0163] Further, setting the disk replacement flag can mean setting the value of the "SubHealth" field corresponding to the target multi-armed disk to "Y" to notify the user to replace the disk.
[0164] As can be seen, in the case of a performance priority strategy for the failure repair strategy, the embodiment of the application does not distinguish between disk failure and arm failure, but completes full-disk reconstruction as quickly as possible at a faster reconstruction speed to ensure the reliability of data protection, thereby improving the reliability of the operation of the distributed storage system.
[0165] Further, step S200 can further include the following steps S209-S212.
[0166] S209: In the case of a cost priority strategy for the failure repair strategy, determine the failure type based on the failure event.
[0167] Consistent with the judgment logic under the performance priority strategy, the embodiment of the present application can determine the failure type by comparing the number of failed storage units with the number of magnetic arms (or the total number of storage units): if the number of failed storage units < the number of magnetic arms (or the total number of storage units), the target multi-arm disk does not have disk failure, and the failure is triggered by arm failure; if the number of failed storage units = the number of magnetic arms (or the total number of storage units), the target multi-arm disk has disk failure. The specific judgment steps will be described in detail below, and will not be repeated here.
[0168] S210: In the case where the failure type is the first failure type, the data of the failed storage unit is reconstructed to the standby storage unit at a second reconstruction speed through the redundant copy data.
[0169] The core of the cost priority strategy is to achieve the purpose of reducing additional costs through sufficient utilization of resources. Therefore, in the case where the storage unit fails but the disk is not completely damaged, only the failed storage unit can be reconstructed.
[0170] Further, the second reconstruction speed is the data read-write speed in the reconstruction process, and the embodiment of the present application can configure the second reconstruction speed for flow control. Under the cost priority strategy, the second reconstruction speed can be high, so that the reconstruction process can be ended as soon as possible, reducing the time of system resource occupation by the reconstruction operation, thereby reducing the idle cost of resources.
[0171] In some implementations, the second reconstruction speed can be greater than the first reconstruction speed. For example, when the first reconstruction speed is equal to 40 MB / s, the second reconstruction speed can be preset to 80 MB / s or 100 MB / s.
[0172] S211: Set a sub-health flag for the failed storage unit, and set a no-disk replacement flag for the target multi-arm disk.
[0173] The fields of the sub-health flag and the no-disk replacement flag can be consistent with the foregoing embodiment, that is, the sub-health flag is set through the "SubHealth" field in the metadata table, and the no-disk replacement flag is set through the "NeedReplace" field in the metadata table. Specifically, the "SubHealth" field can be set to "Y," indicating that the failed storage unit is in a sub-healthy state; and the "NeedReplace" field can be set to "N," indicating that the target multi-arm disk does not need to be replaced.
[0174] It should be noted that step S211 can be performed during the reconstruction process or after the reconstruction is completed.
[0175] S212: in the case of the fault type being the second fault type, reconstructing the data in the target multi-arm disk to one or more spare storage units at a second reconstruction speed through redundant copy data, and setting a disk replacement flag for the target multi-arm disk.
[0176] In this way, full-disk reconstruction can be performed on the multi-arm disk to solve the data loss problem caused by disk failure. Moreover, the disk replacement flag can prompt the administrator to replace the failed disk in a timely manner.
[0177] To sum up, the embodiments of the present application provide three flexible fault processing strategies when facing disk failure or arm failure of a multi-arm disk, which can be applied to different user needs and have strong flexibility.
[0178] Further, the multi-arm fault processing method provided by the embodiments of the present application can also monitor the overall fault state of the target storage pool. When the fault storage capacity is high, the user is reminded to expand the distributed storage system.
[0179] Figure 8 The flowchart for monitoring the storage pool provided by the embodiments of the present application is shown.
[0180] As Figure 8 shown, the embodiments of the present application further include steps S401-S402.
[0181] S401: traverse each storage pool according to a preset period, calculate the total capacity of all fault storage units marked with a sub-health flag in each storage pool, and obtain the fault storage capacity of the storage pool.
[0182] The preset period can be a day or a week, and can be determined according to the business characteristics of the storage pool and the stability of the sub-health storage unit.
[0183] Further, during the traversal process, the embodiments of the present application can read the information of all fault storage units with a "SubHealth" field value of "Y" from the metadata tag, extract the capacity data of each fault storage unit, and accumulate these capacities to finally obtain the fault storage capacity. For example, there are three sub-health storage units in the storage pool, with capacities of 200GB, 300GB and 500GB respectively, so the fault storage capacity is 1000GB.
[0184] In some implementations, when the sub-health flag is set on the multi-arm disk rather than the fault storage unit, the embodiments of the present application can specifically count and accumulate the capacity of the fault storage unit in the multi-arm disk with a "SubHealth" field value of "Y" to finally obtain the fault storage capacity.
[0185] S402: adding an expansion flag to the storage pool when the failed storage capacity exceeds a preset percentage threshold of the total capacity of the storage pool.
[0186] The preset percentage threshold is, for example, 5% or 10%, and can be flexibly adjusted based on actual conditions. When the preset percentage threshold is 5%, if the failed storage capacity > 5% * total capacity of the storage pool, the expansion flag can be added to the storage pool in time for expansion. The total capacity of the storage pool can refer to the total capacity of all disks in the storage pool.
[0187] In some implementations, the total capacity of the storage pool can also refer to the total capacity of all multi-arm disks in the storage pool, which is not limited in the embodiments of the application.
[0188] Further, the expansion flag can be represented by a "NeedExpand" field corresponding to the target storage pool in the metadata table, and the value "Y" of the "NeedExpand" field can indicate that the storage pool needs to be expanded. Correspondingly, the value "N" of the "NeedExpand" field can indicate that the storage pool does not need to be expanded.
[0189] Further, the expansion means can be disk replacement, disk expansion, and node expansion of the multi-arm disk corresponding to the failed storage unit.
[0190] Specifically, disk replacement of the multi-arm disk with a fault can restore the storage capacity and function of the disk corresponding to the disk, and ensure the normal operation of the storage system.
[0191] Disk expansion refers to synchronous expansion of each storage node in the distributed storage system, and the hardware specifications of the expanded disk need to be consistent with the existing disk. In the actual expansion process, the type of the expanded disk needs to be consistent with the existing disk, such as SaaS type. The capacity of the expanded disk needs to be matched with the existing disk, such as 900G of the original disk, and the capacity of the expanded disk also needs to be 900G. In addition, the number of expanded disks of each storage node needs to be the same.
[0192] Node expansion refers to adding new storage nodes to the distributed storage system, and the minimum number of new nodes is 1. The hardware configuration of the new storage node needs to be consistent with the existing storage node, and the disk type, disk capacity, and disk quantity need to match the specifications of the existing storage node.
[0193] Further, whether it is disk expansion or storage node expansion, after the expansion operation is completed, the embodiment of the application can automatically perform capacity rebalancing process, through data migration and distribution adjustment, to ensure that the capacity utilization rate of each storage node and each disk reaches a balanced state, thereby maintaining the stable operation and good performance of the overall distributed storage system. Specifically, the embodiment of the application further includes the following step S500.
[0194] S500: In response to the first storage pool adding a multi-armed disk, the data in the first storage pool is balanced to the newly added storage unit through a migration operation.
[0195] It can be understood that the first storage pool can be any storage pool in the distributed storage system, and the first storage pool adding a multi-armed disk can be triggered by disk replacement, disk expansion and / or node expansion. In addition, the newly added storage unit is a storage unit in the newly added multi-armed disk.
[0196] Exemplarily, the step of balancing the data in the target storage pool to the newly added storage unit through a migration operation can be: counting the data capacity, load of the existing storage unit and available capacity of the newly added storage unit; according to the "load balancing" principle, preferentially migrating the data in the multi-armed disk with high load (such as the used capacity ratio exceeding 70%) to the newly added storage unit to achieve uniform distribution, while controlling the migration speed to avoid affecting the business.
[0197] In this way, the utilization rate of each storage node and each disk can be balanced.
[0198] It should be noted that the step of determining the fault type based on the fault event in steps S204 and S209 can include the following steps S601-S604.
[0199] S601: Based on the fault event, determine the target multi-armed disk to which the fault storage unit belongs.
[0200] In the embodiment of the application, the fault event can carry the WWN or SN corresponding to the fault storage unit. The WWN is the unique identifier of the target multi-armed disk to which it belongs, and the SN is the serial number of the target multi-armed disk to which it belongs. Based on the WWN or SN, the target multi-armed disk to which the fault storage unit belongs can be determined. The identifier WWN / SN can be obtained from the system metadata table.
[0201] In the distributed storage system, each multi-armed disk can be pre-assigned a WWN (e.g., "5000CCA001234567") and written into the disk firmware and system metadata table. Since a multi-armed disk can include multiple storage units, these storage units can share the WWN of the disk to which they belong. Therefore, by checking the WWN of the failed storage unit, the target multi-armed disk to which the failed storage unit belongs can be directly determined, and then the storage unit under the target multi-armed disk can be determined.
[0202] The SN is a unique serial number assigned by the disk manufacturer to each physical disk (e.g., "ST1234567890") and stored in the disk firmware, which is used to distinguish different disks. All storage units (LUN / Disk) of the same multi-armed disk can be associated with the same SN. In this way, the SN can be used as a basis for determining the "storage unit belonging to the multi-armed disk".
[0203] S602: Check whether the storage units other than the failed storage unit in the target multi-armed disk are in a failure state.
[0204] Specifically, the embodiment of the present application can check the status of all storage units associated with the target multi-armed disk. If the storage units other than the failed storage unit are all in a failure state, it can indicate that the multi-armed disk is completely failed. If the storage units other than the failed storage unit are normal, it indicates that the multi-armed disk is not completely damaged.
[0205] S603: If the storage units other than the failed storage unit in the target multi-armed disk are not in a failure state, determine that the failure type is the first failure type.
[0206] S604: If the storage units other than the failed storage unit in the target multi-armed disk are in a failure state, determine that the failure type is the second failure type.
[0207] In some implementations, the failure type can also be determined as the second failure type when all the storage units other than the failed storage unit are in a failure state. When part of the storage units other than the failed storage unit are in a failure state, the failure type is determined as the first failure type. The determination can be based on actual needs, which is not limited in the embodiment of the present application.
[0208] Through the above steps, it can be determined whether the multi-armed disk has a disk failure.
[0209] It should be noted that after step S100, the following step S701 can also be included.
[0210] S701: in the case that the target storage pool to which the failed storage unit belongs contains single-arm disks, based on the failure event, reconstructing the data in the target single-arm disk to other single-arm disks in the target storage pool through redundant copy data.
[0211] wherein the target single-arm disk is the disk to which the failed storage unit belongs.
[0212] In some implementations, a disk replacement flag can be set for the target single-arm disk to inform the user to replace the disk.
[0213] In this way, the reliability of the target storage pool can be ensured.
[0214] As can be seen from the above, the embodiment of the present application provides a multi-arm disk failure processing method, which introduces arm failure judgment, processes arm failure and disk failure respectively, and eliminates the potential multi-copy degradation or data loss problem in traditional multi-copy failure reconstruction and disk replacement processing. At the same time, in the scenario of configuring multi-arm disks in a distributed storage system, the method can provide multiple strategy choices such as reliability, cost, and performance for different application requirements, and support failure reconstruction and disk replacement based on different strategies. Further, the method provided by the embodiment of the present application is not limited to HPC and big data scenarios. In addition, the method provided by the embodiment of the present application can also be applied to multi-arm disk (≥2 arms), single-arm and multi-arm mixed scenarios, and can be applied to storage node failure, arm failure and disk failure, and has wide application range and high flexibility.
[0215] Figure 9 A structural schematic diagram of a multi-arm disk failure processing device provided by the embodiment of the present application is shown.
[0216] As shown in Figure 9 The embodiment of the present application provides a multi-arm disk failure processing device, applied to a computing device, the computing device being in communication connection with a distributed storage system; the distributed storage system including one or more storage pools, the storage pool including a plurality of multi-arm disks, the multi-arm disk including a plurality of storage units; the multi-arm disks in the plurality of storage nodes being divided into one or more storage pools; the device including: a determination module 1101 configured to determine a failure repair strategy in response to a failure event triggered by a failed storage unit; a failure processing module 1102 configured to determine a data reconstruction manner and perform reconstruction based on the failure repair strategy and / or the failure type of the failed storage unit; wherein the data reconstruction manner includes reconstructing the data in the failed storage unit and reconstructing all data in the target multi-arm disk to which the failed storage unit belongs; the reconstruction refers to writing the data in the storage unit to a standby storage unit using redundant copy data, the standby storage unit and the failed storage unit belonging to different multi-arm disks.
[0217] In a possible implementation, the fault type includes a first fault type and a second fault type; the first fault type refers to that the fault event is caused by a local fault of the disk, and the second fault type refers to that the fault event is caused by a global fault of the disk.
[0218] In a possible implementation, the fault processing module 1102 is specifically configured to: in the case that the fault repair strategy is a performance priority strategy, acquire real-time performance data of a target storage node; the target storage node is a storage node to which the fault storage unit belongs, and the real-time performance data includes network delay, CPU usage, memory usage, and / or disk usage of all multi-arm disks in the target storage node; in the case that the real-time performance data is less than a preset performance threshold, determine a fault type based on the fault event; in the case that the fault type is the first fault type, reconstruct data of the fault storage unit to a standby storage unit at a first reconstruction speed through redundant copy data, where the first reconstruction speed refers to a data read-write speed; set a sub-health flag for the fault storage unit, and set a no-disk-replacement flag for the target multi-arm disk; in the case that the fault type is the second fault type, reconstruct data in the target multi-arm disk to one or more standby storage units at the first reconstruction speed through the redundant copy data, and set a disk-replacement flag for the target multi-arm disk.
[0219] In a possible implementation, the fault processing module 1102 is further configured to: in the case that the fault repair strategy is a cost priority strategy, determine a fault type based on the fault event; in the case that the fault type is the first fault type, reconstruct data of the fault storage unit to a standby storage unit at a second reconstruction speed through redundant copy data, where the second reconstruction speed refers to a data read-write speed; set a sub-health flag for the fault storage unit, and set a no-disk-replacement flag for the target multi-arm disk; in the case that the fault type is the second fault type, reconstruct data in the target multi-arm disk to one or more standby storage units at the second reconstruction speed through the redundant copy data, and set a disk-replacement flag for the target multi-arm disk.
[0220] In a possible implementation, the fault processing module 1102 is further configured to: based on the fault event, determine a target multi-arm disk to which the fault storage unit belongs; check whether a storage unit other than the fault storage unit in the target multi-arm disk is in a fault state; if the storage unit other than the fault storage unit in the target multi-arm disk is not in the fault state, determine that the fault type is the first fault type; if the storage unit other than the fault storage unit in the target multi-arm disk is in the fault state, determine that the fault type is the second fault type.
[0221] In a possible implementation, the multiple-armed magnetic disks in the multiple storage nodes are divided into one or more storage pools; the fault processing module 1102 is further configured to: traverse each storage pool according to a preset period, calculate the total capacity of all the fault storage units marked with the sub-health flag in each storage pool, and obtain the fault storage capacity of the storage pool; and add an expansion flag to the storage pool when the fault storage capacity exceeds a preset percentage threshold of the total capacity of the storage pool.
[0222] In a possible implementation, the fault processing module 1102 is further configured to: in a case where the fault repair strategy is a reliability-first strategy, reconstruct, by using redundant copy data, data in the target multiple-armed magnetic disk to one or more spare storage units at a third reconstruction speed, and set a disk replacement flag for the target multiple-armed magnetic disk, the third reconstruction speed being a data read-write speed.
[0223] In a possible implementation, the apparatus further includes a balancing module 1103 configured to: in response to adding a multiple-armed magnetic disk to the first storage pool, balance, by using a migration operation, data in the first storage pool to the newly added storage unit.
[0224] Figure 10 A schematic diagram of a computing device provided by an embodiment of the present application.
[0225] As shown in Figure 10 , the computing device 1200 can include a server, a terminal, and the like; the computing device 1200 includes one or more processors 1201 and a memory 1202. The memory 1202 is configured to store one or more programs. When the one or more programs are executed by the one or more processors 1201, the one or more processors 1201 implement the multiple-armed magnetic disk fault processing method in the above-described embodiments.
[0226] Continuing to refer to Figure 10 , the computing device 1200 can further include a communications interface 1203 and a communications bus 1204.
[0227] The processor 1201, the memory 1202, and the communications interface 1203 complete communications with each other through the communications bus 1204. The communications interface 1203 is configured to communicate with network elements such as clients or other servers, and the like.
[0228] In some embodiments, the one or more processors 1201 are configured to execute one or more programs 1205, and specifically can execute the related steps in the above-described multiple-armed magnetic disk fault processing method embodiments. Specifically, the program 1205 can include program code including computer-executable instructions.
[0229] Exemplarily, the processor 1201 can be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. The computing device 1200 can include one or more processors, which can be the same type of processor, such as one or more CPUs, or different types of processors, such as one or more CPUs and one or more ASICs.
[0230] In some embodiments, the memory 1202 is configured to store one or more programs 1205. The memory 1202 can include a high-speed RAM memory, and can also include a non-volatile memory (NVM), such as at least one disk memory.
[0231] The program 1205 can be specifically invoked by the processor 1201 to cause the computing device 1200 to perform the multi-armed disk failure processing method operations.
[0232] Some embodiments of the present application provide a computer readable storage medium, which stores at least one executable instruction, which, when executed on the computing device 1200, causes the computing device 1200 to perform the multi-armed disk failure processing method in the above embodiments.
[0233] The executable instruction can be specifically used to cause the computing device 1200 to perform the multi-armed disk failure processing method operations.
[0234] For example, the computer readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0235] The beneficial effects that can be achieved by the readable storage medium provided by some embodiments of the present application can refer to the beneficial effects provided in the corresponding multi-armed disk failure processing method, which will not be described here.
[0236] It is to be understood that the phrases "in one embodiment" or "implementing the embodiment" as used herein does not necessarily refer to the same embodiment, although it may. Further, the terms "comprises", "comprising", or other any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0237] Each of the embodiments described in this specification has at least one aspect. Relatively, the same or similar parts among the embodiments can be mutually referred to, and each embodiment focuses on the difference from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0238] The logic and / or steps represented in the flowchart and / or described herein, for example, can be considered as a sequence of executable instructions for implementing logical functions, and can be specifically embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus or device, such as a computer-based system, a processor-based system, or other system that can fetch the instructions from the instruction execution system, apparatus or device and execute the instructions, or in conjunction with these instruction execution systems, apparatus or devices.
[0239] For the purpose of this specification, a "computer-readable medium" can be any device or apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus or device.
[0240] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electronic devices), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM).
[0241] Additionally, a computer readable medium can be paper or other comparable effectively medium whereupon at least one of the programs is printed, since the programs can be electronically retrieved, for instance by optically scanning the paper or other medium, then
[0242] In the embodiments described above, the steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, implementation can be with any or a combination of the following technologies, which are all well known in the art: a discrete logic circuit having logic gates for implementing logic functions upon an application of data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0243] The above embodiments are only specific embodiments of the present application, and are not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.
Claims
1. A multi-magnetic-arm disk failure handling method, characterized by, The application is applied to a computing device connected with a distributed storage system; the distributed storage system comprises a plurality of storage nodes, and each storage node comprises a plurality of multi-armed disks, and each multi-armed disk comprises a plurality of storage units; The method comprises: In response to a fault event triggered by a faulty storage unit, a fault repair strategy is determined; Based on the fault repair strategy and / or the fault type of the faulty storage unit, a data reconstruction mode is determined and reconstruction is performed; wherein the data reconstruction mode comprises reconstructing data in the faulty storage unit, reconstructing all data in a target multi-armed disk to which the faulty storage unit belongs, and reconstruction refers to writing data in a storage unit to a backup storage unit using redundant copy data, and the backup storage unit and the faulty storage unit belong to different multi-armed disks.
2. The multi-magnetic-arm disk fault handling method of claim 1, wherein, The fault type comprises a first fault type and a second fault type; the first fault type refers to that the fault event is caused by a local disk fault, and the second fault type refers to that the fault event is caused by a global disk fault.
3. The multi-magnetic-arm disk failure handling method of claim 2, wherein, The method comprises: In the case of a performance priority strategy, real-time performance data of a target storage node is obtained; wherein the target storage node is a storage node to which the faulty storage unit belongs, and the real-time performance data comprises network delay, CPU usage, memory usage and / or disk usage of all multi-armed disks in the target storage node; In the case that the real-time performance data is less than a preset performance threshold, the fault type is determined based on the fault event; In the case that the fault type is the first fault type, data of the faulty storage unit is reconstructed to the backup storage unit at a first reconstruction speed through the redundant copy data, wherein the first reconstruction speed refers to the data read / write speed; A sub-health flag is set for the faulty storage unit, and a no-disk-replacement flag is set for the target multi-armed disk; In the case that the fault type is the second fault type, data in the target multi-armed disk is reconstructed to one or more backup storage units at a first reconstruction speed through the redundant copy data, and the disk-replacement flag is set for the target multi-armed disk.
4. The multi-magnetic-arm disk failure handling method of claim 2, wherein, The method further comprises: In the case of a cost priority strategy, the fault type is determined based on the fault event; In the case that the fault type is the first fault type, data of the faulty storage unit is reconstructed to the backup storage unit at a second reconstruction speed through the redundant copy data, wherein the second reconstruction speed refers to the data read / write speed; A sub-health flag is set for the faulty storage unit, and a no-disk-replacement flag is set for the target multi-armed disk; In a case where the fault type is the second fault type, data in a target multi-armed disk is reconstructed to one or more spare storage units at a second reconstruction speed by using the redundant copy data, and a disk replacement flag is set for the target multi-armed disk.
5. The multi-magnetic-arm disk failure handling method according to claim 3 or 4, characterized by, The determining of the fault type based on the fault event comprises: determining, based on the fault event, a target multi-armed disk to which the fault storage unit belongs; checking whether a storage unit other than the fault storage unit in the target multi-armed disk is in a fault state; if the storage unit other than the fault storage unit in the target multi-armed disk is not in the fault state, determining that the fault type is the first fault type; if the storage unit other than the fault storage unit in the target multi-armed disk is in the fault state, determining that the fault type is the second fault type.
6. The multi-magnetic-arm disk failure handling method according to claim 3 or 4, characterized by, The multi-armed disks in the plurality of storage nodes are divided into one or more storage pools; the method further comprises: traversing each of the storage pools at a preset period, calculating a total capacity of all the fault storage units marked with the sub-health flag in each of the storage pools to obtain a fault storage capacity of the storage pool; and when the fault storage capacity exceeds a preset percentage threshold of a total capacity of the storage pool, adding an expansion flag to the storage pool.
7. The multi-magnetic-arm disk fault handling method of claim 1, wherein, The determining of the data reconstruction manner and the reconstruction based on the fault repair strategy and / or the fault type of the fault storage unit further comprises: in a case where the fault repair strategy is a reliability priority strategy, reconstructing data in the target multi-armed disk to one or more spare storage units at a third reconstruction speed by using the redundant copy data, and setting a disk replacement flag for the target multi-armed disk, the third reconstruction speed being a data read-write speed.
8. The multi-magnetic-arm disk fault handling method of claim 1, wherein, The method further comprises: in response to the addition of the multi-armed disk to the first storage pool, balancing data in the first storage pool to the newly added storage unit through a migration operation.
9. A multi-magnetic arm disk failure handling apparatus, characterized by comprising: The application is applied to a computing device, the computing device being in communication connection with a distributed storage system; the distributed storage system comprising one or more storage pools, the storage pool comprising a plurality of multi-armed disks, and the multi-armed disk comprising a plurality of storage units; The apparatus comprises: a determining module configured to determine a fault repair strategy in response to a fault event triggered by a fault storage unit; a fault processing module configured to determine a data reconstruction manner and perform reconstruction based on the fault repair strategy and / or a fault type of the fault storage unit; wherein the data reconstruction manner comprises reconstructing data in the fault storage unit and reconstructing all data in a target multi-armed disk to which the fault storage unit belongs; the reconstruction refers to writing data in a storage unit into a spare storage unit by using redundant copy data, the spare storage unit and the fault storage unit belonging to different multi-armed disks.
10. A computing device, comprising: comprise: one or more processors; and a memory configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-magnetic-arm disk failure processing method according to any one of claims 1-8.