Disk fault processing method and device, equipment and medium
By detecting the disk failure in the RAID array and temporarily storing data and restarting the disk, and adjusting the disk status in combination with the number of failures, the performance degradation and data security risks caused by disk burst failure are solved, and more efficient storage performance and security are achieved.
Patent Information
- Application Number
- CN202510584717.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-01
AI Technical Summary
The handling of disk burst failures in existing RAID arrays results in degraded storage system performance and may cause data security risks.
When a disk read and write failure is detected, the data is temporarily stored in the cache area and restarted the disk. After restarting, the input and output operations are re-executed, and the disk status is adjusted according to the number of failures. The performance loss and data loss are avoided through the intelligent disk restart and IO cache mechanism.
Improves the storage performance and data security of disk arrays, reduces performance degradation and data loss caused by disk failures, and improves system stability and user experience.
Smart Images

Figure CN120407292A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and particularly to a method, apparatus, device and medium for handling disk failures. Background Art
[0002] In a storage system, the Redundant Array of Independent Disks (RAID) technology widely applies to improving the reliability and performance of data storage by dispersing data storage and additionally storing parity blocks. This technology achieves two core goals: on the one hand, by means of concurrent read and write operations among multiple disks, it improves the Input / Output Operations Per Second (IOPS) of input / output (IO); on the other hand, it significantly enhances the reliability of the entire storage system, enabling the array to continue running even when one or more disks fail.
[0003] However, during actual long-term operation, disks may suddenly fail, reducing the performance of the array. There are two common countermeasures: (1) Disk removal and degraded operation: When a disk read / write IO fails, it is removed from the array, and the array enters a degraded state. Write operations skip this disk, and read operations reconstruct data through other disks and parity blocks. (2) IO failure and host retry: Directly inform the host of the IO failure and let the host retry.
[0004] The above two traditional processing methods not only cause a decline in the performance of the storage system, but may also lead to data security risks in the RAID array due to unnecessary degraded operations. Therefore, how to improve the method for handling sudden disk failures in traditional RAID array technology is a technical problem that needs to be urgently solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method, apparatus, device and medium for handling disk failures, which can avoid reducing the storage performance of the disk array and greatly improve the data security of the disk array. The specific solutions are as follows:
[0006] In a first aspect, the present application discloses a method for handling disk failures, including:
[0007] When it is detected that a read / write failure occurs in the input / output operation of the target disk, the target data of the input / output operation is stored in the target cache area, and the target disk is restarted;
[0008] When it is detected that the target disk is in a reset state after restart, the cached target data is read from the target cache area, and the input / output operation is re-executed based on the target data;
[0009] Determine the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjust the operating status of the target disk in the disk array according to the reliability status.
[0010] Optionally, determining the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjusting the operating status of the target disk in the disk array includes:
[0011] Configure a corresponding failure counter for each disk, and set the initial value of the failure counter to 0; where, when a read / write failure occurs on the corresponding disk, increment the value of the corresponding failure counter by one;
[0012] Determine the target failure counter corresponding to the target disk, and compare the value of the target failure counter with a preset number threshold;
[0013] When the value of the target failure counter is not greater than the preset number threshold, determine that the target disk is a reliable disk, and keep the target disk running normally in the disk array;
[0014] When the value of the target failure counter is greater than the preset number threshold, determine that the target disk is an unreliable disk, remove the target disk from the disk array, and then generate multi-dimensional warning information and record the corresponding failure log; where, the multi-dimensional warning information includes device basic information, fault feature information, and maintenance reference information;
[0015] Determine whether a new read / write failure occurs on the target disk during the target monitoring period, and determine whether to update the value of the target failure counter according to the determination result.
[0016] Optionally, determining whether to update the value of the target failure counter according to the determination result includes:
[0017] If no new read / write failure occurs on the target disk during the target monitoring period, gradually decrease the value of the target failure counter according to a preset decay rate until the value of the target failure counter is 0.
[0018] Optionally, the disk failure handling method further includes:
[0019] Determine the aging coefficient of the target disk according to the usage duration of the target disk;
[0020] Determine the environmental coefficient according to the current environmental temperature;
[0021] Determine the fault coefficient of the target disk according to the historical read / write failure frequency of the target disk;
[0022] Determine the preset number threshold based on the aging coefficient, environmental coefficient, and fault coefficient.
[0023] Optionally, re - execute the input / output operation based on the target data, including:
[0024] Determine whether the logical block address of the target data is located in the target addressing space. If the logical block address is located in the target addressing space, it is determined that the target data meets the first verification condition;
[0025] Determine the current check value of the target data based on the encoding rule of the target data. If the current check value matches the preset check value, it is determined that the target data meets the second verification condition;
[0026] Determine whether the operation timestamp of the target data is within the target valid time. If the operation timestamp is within the target valid time, it is determined that the target data meets the third verification condition;
[0027] When the target data meets the first verification condition, the second verification condition, and the third verification condition, re - execute the input / output operation based on the target data.
[0028] Optionally, store the target data of the input / output operation in the target cache area, including:
[0029] Construct a first - in - first - out queue with high priority, a first - in - first - out queue with medium priority, and a first - in - first - out queue with low priority according to the business criticality and response timeliness requirements, and store the target data of the input / output operation in the corresponding first - in - first - out queue.
[0030] Optionally, the first - in - first - out queue with high priority is used to store the immediate response operations in the system - critical operations, the first - in - first - out queue with medium priority is used to store the non - immediate response operations in the system - critical operations or the immediate response operations in other business operations, and the first - in - first - out queue with low priority is used to store the non - immediate response operations in other business operations.
[0031] In a second aspect, the present application discloses a disk failure handling device, including:
[0032] A data cache module, configured to store the target data of the input / output operation in the target cache area and perform a restart process on the target disk when it is detected that a read / write failure occurs in the input / output operation of the target disk;
[0033] A re - execution module, configured to read the cached target data from the target cache area and re - execute the input / output operation based on the target data when it is detected that the target disk is in a reset state after restart;
[0034] A status adjustment module, configured to determine the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjust the running state of the target disk in the disk array according to the reliability status.
[0035] In a third aspect, the present application discloses an electronic device, including:
[0036] a memory for storing a computer program;
[0037] a processor for executing the computer program to implement the aforementioned disk failure handling method.
[0038] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disk failure handling method is implemented.
[0039] It can be seen that the present application proposes a disk failure handling method, including: when it is detected that a read / write failure occurs in an input / output operation of a target disk, storing the target data of the input / output operation in a target cache area, and performing a restart process on the target disk; when it is detected that the target disk is in a reset state after restart, reading the cached target data from the target cache area, and re-executing the input / output operation based on the target data; determining the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjusting the running state of the target disk in the disk array according to the reliability status. It can be seen that in the disk failure handling method involved in the present application, when a read / write failure occurs in the input / output operation of the target disk, the system will automatically temporarily store the target data of the input / output operation in the target cache area, and at the same time automatically trigger the disk restart process. This process is autonomously executed by the system without notifying the host, avoiding bringing additional tasks to the host and effectively preventing the significant decline of the array storage performance due to the occupation of host resources. Further, after the target disk is successfully reset, the system will re-execute the previously temporarily stored input / output operation, effectively avoiding data loss or operation interruption caused by disk failures. In addition, the present application also introduces a more intelligent disk status management mechanism, accurately determining the reliability status of the target disk according to the cumulative number of read / write failures of the target disk. And based on this reliability status, further adjusting the running state of the target disk in the disk array, avoiding directly performing downgrading or removal operations once a disk fails, thereby greatly improving the storage performance and data security of the disk array. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0041] Figure 1 It is a flowchart of a disk failure handling method disclosed in the present application;
[0042] Figure 2 A flowchart for disk I / O failure handling and recovery disclosed in the present application;
[0043] Figure 3 A schematic structural diagram of a disk failure handling device disclosed in the present application;
[0044] Figure 4 A structural diagram of an electronic device disclosed in the present application. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] However, during actual long-term operation, the disk may suddenly fail, reducing the performance of the array. There are two common existing coping methods: (1) Disk removal and degraded operation: When the disk read / write I / O fails, it is removed from the array, and the array enters the degraded state. The write operation skips this disk, and the read operation reconstructs the data through other disks and parity blocks. (2) I / O failure and host retry: Directly inform the host that the I / O fails and let the host retry.
[0047] The above two traditional processing methods will not only cause a decline in the performance of the storage system, but may also lead to data security risks in the RAID array due to unnecessary degraded operations. Therefore, how to improve the processing method for sudden disk failures in traditional RAID array technology is a technical problem that needs to be urgently solved by those skilled in the art.
[0048] For this reason, the embodiments of the present application propose a disk failure handling solution, which can avoid reducing the storage performance of the disk array and greatly improve the data security of the disk array.
[0049] The embodiments of the present application disclose a disk failure handling method. Refer to Figure 1 as shown, this method includes:
[0050] Step S11: When it is detected that a read / write failure occurs in the input / output operation of the target disk, store the target data of the input / output operation in the target cache area, and perform a restart process on the target disk.
[0051] In this embodiment, when a read / write failure occurs in the input / output operation of the target disk, the target data of the input / output operation is stored in the target cache area, and the target disk is restarted. Among them, the target cache area can be a First In First Out (FIFO) queue. In this way, this embodiment can construct a high-priority FIFO queue, a medium-priority FIFO queue, and a low-priority FIFO queue according to the business criticality and response timeliness requirements, and store the target data of the input / output operation in the appropriate FIFO queue. Among them, the high-priority FIFO queue is used to store the immediate response operations in the system critical operations, the medium-priority FIFO queue is used to store the non-immediate response operations in the system critical operations or the immediate response operations in other business operations, and the low-priority FIFO queue is used to store the non-immediate response operations in other business operations.
[0052] Exemplarily, (1) For the immediate response operations in critical operations, such operations play a decisive role in the normal operation of the storage system and the maintenance of data security, and require the disk to complete the response in an extremely short time. For example, in a financial trading system, at the moment when a specific transaction is completed, the system has to immediately write the transaction time, transaction amount, and key information of both parties to the disk. If the disk write delay exceeds 10 milliseconds, it is easy to cause data deviation and trigger transaction disputes; (2) For the non-immediate response operations in critical operations, such operations are crucial for the long-term stable operation of the storage system, data integrity guarantee, etc., but the operation execution time is not required to be completed instantly and can be carried out during relatively idle or appropriate periods of the system. For example, to ensure data reliability, the storage system of a large data center will regularly check the disk array consistency. By comparing redundant data copies, it detects whether there are inconsistencies or damages. This operation is critical for long-term data availability and integrity, but the data center can execute it during the late-night business off-peak period; (3) For the immediate response operations in other business operations, such operations belong to the operations in the non-core business processes supported by the storage system. Although they have relatively little impact on the key functions of the system, when the operations occur, the disk is also required to respond quickly to ensure the smooth operation of the business function and user experience. For example, when a student submits an answer result or asks a question during a real-time classroom interaction, the server has to immediately store these interaction data on the disk. Although it is not a core critical operation, to ensure the smoothness of the classroom and timely feedback, the disk response needs to be fast; (4) For the non-immediate response operations in other business operations, such operations in various services of the storage system belong to the operations that do not have strict requirements for the response time and can be executed slowly according to a specific plan or in the background when system resources permit. For example, an e-commerce platform will regularly re-index and build a product picture library to optimize product search. However, this operation does not affect the real-time transaction process and can be carried out when the system is idle in the early morning. Even if the disk read / write is slow and the operation time is long, it does not affect the normal shopping of users during the day.
[0053] See Figure 2 As shown, when the RAID array performs an IO operation, the IO operation is sliced and calculated. This calculation determines how data is distributed among disks, balances the load on each disk, and plans the operation sequence. After the calculation is completed, each disk (such as Disk 1, Disk 2 to Disk n) starts to perform read and write operations. If an abnormal situation such as a read / write failure occurs on a disk during the operation, the system will start the intelligent disk restart mechanism. At this time, the system will quickly store the IO operation for the faulty disk (here Disk 2) in the first-in, first-out queue. During this period, the host will not notice any abnormality, thus ensuring the continuity of the user experience and the stability of the system, while maintaining data integrity and preventing service interruption caused by disk restart. To avoid unnecessary repeated restarts, the system will check the disk status. If the disk is already in the reset state, new read / write requests will be cached without triggering additional reset operations. Once the disk is successfully reset, the system will automatically resubmit the IO operations stored in the FIFO queue to the disk to ensure that all operations can be correctly executed. The FIFO queue has a pending depth of 1, which determines the number of IO operations that the queue can store. When the queue is full, the system will suspend receiving new IO requests until the disk is successfully reset. When setting the pending depth, it is necessary to comprehensively consider the memory capacity of the RAID controller and the actual IO processing capacity, ensuring that the system can handle sufficient IO requests when a disk fails without affecting the overall performance due to an overly deep queue.
[0054] Step S12: When it is detected that the target disk is in the reset state after restart, read the cached target data from the target cache area and re-execute the input / output operation based on the target data.
[0055] In this embodiment, when it is detected that the target disk is in the reset state after restart, read the cached target data from the target cache area and verify the read target data to ensure data integrity and operation effectiveness. The specific verification process includes a three-level verification mechanism: determine whether the logical block address of the target data is within the target addressing space. If the logical block address is within the target addressing space, it is determined that the target data meets the first verification condition; determine the current check value of the target data based on the encoding rule of the target data. If the current check value matches the preset check value, it is determined that the target data meets the second verification condition; determine whether the operation timestamp of the target data is within the target valid time. If the operation timestamp is within the target valid time, it is determined that the target data meets the third verification condition; when the target data meets the first verification condition, the second verification condition, and the third verification condition, re-execute the input / output operation based on the target data. In this way, the risk of system errors or operation failures caused by data problems is greatly reduced.
[0056] Step S13: Determine the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjust the running status of the target disk in the disk array according to the reliability status.
[0057] When dealing with disk failures, occasional read / write errors and irrecoverable errors caused by hardware damage or storage medium failures are often encountered. Some occasional read / write errors can be resolved by restarting the disk, while irrecoverable errors cannot be repaired even by restarting the disk. To effectively distinguish between these two situations, this application proposes an intelligent processing mechanism based on a failure count threshold. (1) Failure counter setting and counting: Configure a corresponding failure counter for each disk in the disk array and set its initial value to 0. When an IO read / write error occurs on the disk, the value of the corresponding failure counter is incremented by 1. (2) Faulty disk determination and handling: Determine the target failure counter corresponding to the target disk and compare its value with the preset count threshold. If the value of the target failure counter is greater than the preset count threshold, determine that the target disk is an unreliable disk with an irrecoverable error. The system automatically removes it from the disk array, generates multi-dimensional warning information including device basic information, fault feature information, and maintenance reference information, and records a fault log to prompt the user to check and replace the faulty disk to ensure data security and array integrity. (3) Dynamic adjustment of counter value: Determine whether a new read / write failure occurs on the target disk within the target monitoring period (such as 24 hours). If not, gradually decrement the value of the target failure counter according to the preset decay rate (such as decreasing by 1 per day) until it reaches 0. In this way, the system is allowed to self-recover within a certain period of time, reducing the performance loss caused by unnecessary disk removal due to occasional errors and resulting in array degradation. If a new read / write failure occurs on the target disk within the target monitoring period, no decrement operation is performed, and subsequent processing continues based on the comparison result of the value of the failure counter with the preset threshold. (4) Reliable disk determination: When the value of the target failure counter is not greater than the preset count threshold, determine that the target disk is a reliable disk and allow it to operate normally in the disk array.
[0058] In this way, through this intelligent processing solution based on the failure count threshold, this embodiment can effectively distinguish recoverable and irrecoverable disk errors, thereby reducing unnecessary disk restart operations.
[0059] Among them, the determination of the preset number threshold may include the following process: determining the aging coefficient of the target disk according to the usage duration of the target disk; determining the environmental coefficient according to the current ambient temperature; determining the failure coefficient of the target disk according to the historical read / write failure frequency of the target disk; and determining the preset number threshold based on the aging coefficient, the environmental coefficient, and the failure coefficient. Specifically, the usage duration of the target disk is closely related to the aging coefficient. The longer the usage time, the more serious the aging phenomena such as disk hardware wear and performance decline, and the higher the aging coefficient, indicating that the disk is more likely to fail. Therefore, a relatively low preset number threshold is required to detect and handle potential problems in a timely manner. The current ambient temperature affects the stability of the disk. In a high-temperature environment, the performance of the disk's electronic components is easily affected, and the failure rate may increase. By determining the environmental coefficient, environmental factors can be incorporated into the consideration of the preset number threshold. If the environmental coefficient indicates that the current temperature is unfavorable to the disk, the preset number threshold should also be appropriately reduced. The historical read / write failure frequency of the target disk directly reflects the past failure situation of the disk. If the historical read / write failure frequency is high, it means that the disk itself has poor stability and a large failure coefficient. At this time, a lower preset number threshold also needs to be set.
[0060] By comprehensively considering various factors such as the usage duration of the target disk, the current ambient temperature, and the historical read / write failure frequency to determine the preset number threshold, the accuracy of disk failure determination can be significantly improved.
[0061] The intelligent disk restart and IO cache mechanism technology mentioned in this application reduces the storage performance loss caused by occasional disk failures. During the disk restart process, the host is basically not aware, greatly improving the user experience. In addition, the disk failure warning and fault isolation technology proposed in this application can effectively distinguish recoverable and non-recoverable disk errors, reduce unnecessary disk restart operations, improve the stability and reliability of the system, and at the same time send a disk failure warning to the user in a timely manner to ensure data security. In addition, the detailed solution of the raid array disk failure handling technology proposed in this application can be used for reference in the software and hardware design in the storage field. That is, the intelligent disk restart and IO caching technology, as well as the disk failure warning and fault isolation technology proposed in this application. These technologies can not only be applied to the disk failure handling of the raid array, but also the same technical processing can be adopted when disk failures occur in other storage technologies.
[0062] In addition, in order to further improve the performance recovery speed and overall operation efficiency of the storage system after disk failure handling, this application proposes an intelligent scheduling mechanism for cross-disk data migration. When a certain disk is determined to be an unreliable disk and removed from the array, the load distribution of the remaining disks will change. By establishing an intelligent scheduling model, factors such as the storage capacity, read / write performance, current load, and data access frequency of the remaining disks are comprehensively considered to automatically plan the migration path of data among other disks. For example, frequently accessed data is migrated to disks with better read / write performance and lower load to improve the access efficiency of the overall storage system. At the same time, asynchronous migration technology is adopted to avoid affecting ongoing normal read / write operations due to data migration, ensuring the stability and performance of the system during data migration are not significantly disturbed.
[0063] It can be seen that this application proposes a disk failure handling method, including: when it is detected that a read / write failure occurs in the input / output operation of the target disk, storing the target data of the input / output operation in the target cache area, and restarting the target disk; when it is detected that the target disk is in a reset state after restart, reading the cached target data from the target cache area, and re-executing the input / output operation based on the target data; determining the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjusting the running state of the target disk in the disk array according to the reliability status. It can be seen that in the disk failure handling method involved in this application, when a read / write failure occurs in the input / output operation of the target disk, the system will automatically temporarily store the target data of the input / output operation in the target cache area, and at the same time automatically trigger the disk restart process. This process is autonomously executed by the system without notifying the host, avoiding bringing additional tasks to the host and effectively preventing the significant decline of the array storage performance due to the occupation of host resources. Further, after the target disk is successfully reset, the system will re-execute the previously stored input / output operation, effectively avoiding data loss or operation interruption caused by disk failure. In addition, this application also introduces a more intelligent disk status management mechanism, which accurately determines the reliability status of the target disk according to the cumulative number of read / write failures of the target disk. And based on this reliability status, further adjust the running state of the target disk in the disk array to avoid directly performing downgrade or removal operations once the disk fails, thereby greatly improving the storage performance and data security of the disk array.
[0064] Correspondingly, the embodiment of this application also discloses a disk failure handling device, see Figure 3 as shown, the device includes:
[0065] A data cache module 11, configured to store the target data of the input / output operation in the target cache area and restart the target disk when it is detected that a read / write failure occurs in the input / output operation of the target disk;
[0066] A re - execution module 12, configured to, when it is detected that the target disk is in a reset state after a restart, read the cached target data from the target cache area and re - execute the input / output operation based on the target data;
[0067] A status adjustment module 13, configured to determine the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjust the running status of the target disk in the disk array according to the reliability status.
[0068] Among them, for the more specific working processes of the above - mentioned respective modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0069] It can be seen that the present application provides a disk failure handling device, including: a data caching module 11, configured to, when it is detected that a read / write failure occurs in the input / output operation of the target disk, store the target data of the input / output operation in the target cache area and perform a restart process on the target disk; a re - execution module 12, configured to, when it is detected that the target disk is in a reset state after a restart, read the cached target data from the target cache area and re - execute the input / output operation based on the target data; a status adjustment module 13, configured to determine the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjust the running status of the target disk in the disk array according to the reliability status. It can be seen that in the disk failure handling method involved in the present application, when a read / write failure occurs in the input / output operation of the target disk, the system will automatically temporarily store the target data of the input / output operation in the target cache area, and at the same time automatically trigger the disk restart process. This process is autonomously executed by the system without notifying the host, avoiding bringing additional tasks to the host and effectively preventing the significant decline of the array storage performance due to the occupation of host resources. Further, after the target disk is successfully reset, the system will re - execute the previously temporarily stored input / output operation, effectively avoiding data loss or operation interruption caused by disk failures. In addition, the present application also introduces a more intelligent disk status management mechanism, which accurately determines the reliability status of the target disk according to the cumulative number of read / write failures of the target disk. And based on this reliability status, the running status of the target disk in the disk array is further adjusted to avoid directly performing downgrade or removal operations once a disk fails, thereby greatly improving the storage performance and data security of the disk array.
[0070] Further, an embodiment of the present application also provides an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation on the usage scope of the present application.
[0071] Figure 4Schematic diagram of the structure of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. Among them, the memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the disk failure handling method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0072] In this embodiment, the power supply 26 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed here; the input / output interface 24 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitations are made here.
[0073] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include a computer program 221, and the storage method may be short-term storage or permanent storage. Among them, the computer program 221 may further include a computer program capable of performing other specific tasks in addition to the computer program capable of implementing the disk failure handling method executed by the electronic device 20 disclosed in any of the foregoing embodiments.
[0074] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the disk failure handling method disclosed above is implemented.
[0075] For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0076] The various embodiments in this application are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference may be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference may be made to the description in the method part for related parts.
[0077] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0078] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0079] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0080] The above has introduced in detail a disk failure handling method, device, equipment, and storage medium provided by this application. Specific examples are used herein to illustrate the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for handling disk failures, characterized in that, Including: When a read / write failure occurs in the input / output operation of the target disk, store the target data of the input / output operation in the target cache area, and perform a restart process on the target disk; When it is detected that the target disk is in a reset state after restart, read the cached target data from the target cache area, and re-execute the input / output operation based on the target data; Determine the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjust the running state of the target disk in the disk array according to the reliability status.
2. The disk failure handling method according to claim 1, wherein The determining the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjusting the running state of the target disk in the disk array according to the reliability status includes: Configure a corresponding failure counter for each disk, and set the initial value of the failure counter to 0; wherein, when a read / write failure occurs for the corresponding disk, increment the value of the corresponding failure counter by one; Determine the target failure counter corresponding to the target disk, and compare the value of the target failure counter with a preset number threshold; When the value of the target failure counter is not greater than the preset number threshold, determine that the target disk is a reliable disk, and keep the target disk running normally in the disk array; When the value of the target failure counter is greater than the preset number threshold, determine that the target disk is an unreliable disk, remove the target disk from the disk array, and then generate multi-dimensional warning information and record the corresponding failure log; wherein, the multi-dimensional warning information includes device basic information, fault feature information, and maintenance reference information; Judge whether a new read / write failure occurs for the target disk within the target monitoring period, and determine whether to update the value of the target failure counter according to the judgment result.
3. The disk failure handling method according to claim 2, wherein, The determining whether to update the value of the target failure counter according to the judgment result includes: If no new read / write failure occurs for the target disk within the target monitoring period, gradually decrease the value of the target failure counter according to a preset decay rate until the value of the target failure counter is 0.
4. The disk failure handling method according to claim 2, wherein Also including: Determine the aging coefficient of the target disk according to the usage duration of the target disk; Determine the environmental coefficient according to the current environmental temperature; Determine the fault coefficient of the target disk according to the historical read / write failure frequency of the target disk; Determine the preset number threshold based on the aging coefficient, the environmental coefficient, and the fault coefficient.
5. The disk failure handling method according to claim 1, wherein The re-executing the input / output operation based on the target data includes: Judge whether the logical block address of the target data is located in the target addressing space. If the logical block address is located in the target addressing space, determine that the target data meets the first verification condition; Determine the current check value of the target data based on the encoding rule of the target data. If the current check value matches the preset check value, determine that the target data meets the second verification condition; Determine whether the operation timestamp of the target data is within the target valid time. If the operation timestamp is within the target valid time, it is determined that the target data meets the third verification condition; When the target data meets the first verification condition, the second verification condition, and the third verification condition, the input / output operation is re-executed based on the target data.
6. The disk failure handling method according to any one of claims 1 to 5, characterized in that, Storing the target data of the input / output operation in the target cache area includes: Constructing a first-in-first-out queue with high priority, a first-in-first-out queue with medium priority, and a first-in-first-out queue with low priority according to the business criticality and response timeliness requirements, and storing the target data of the input / output operation in the corresponding first-in-first-out queue.
7. The disk failure handling method according to claim 6, wherein The first-in-first-out queue with high priority is used to store the immediate response operations in the system critical operations. The first-in-first-out queue with medium priority is used to store the non-immediate response operations in the system critical operations or the immediate response operations in other business operations. The first-in-first-out queue with low priority is used to store the non-immediate response operations in other business operations.
8. A disk failure handling device, characterized in that, including: A data cache module, which is used to store the target data of the input / output operation in the target cache area and perform a restart process on the target disk when it detects that a read / write failure occurs in the input / output operation of the target disk; A re-execution module, which is used to read the cached target data from the target cache area and re-execute the input / output operation based on the target data when it detects that the target disk is in a reset state after restart; A status adjustment module, which is used to determine the reliability status of the target disk according to the cumulative number of read / write failures of the target disk, and adjust the running status of the target disk in the disk array according to the reliability status.
9. An electronic device, characterized in that, including: A memory for storing a computer program; A processor for executing the computer program to implement the disk failure handling method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, it implements the disk failure handling method according to any one of claims 1 to 7.
Citation Information
Cited By
Data reconstruction method and electronic equipment
CN120929298A