Hard disk health assessment method, hard disk management method and related devices
Through multi-dimensional evaluation of the health status of the hard disk, combined with health indicators, error codes and delay information, the problem of inaccurate hard disk health assessment in the existing technology is solved, and higher evaluation accuracy and user experience is achieved, and it is suitable for the management of single and multi-hard disk systems.
Patent Information
- Application Number
- CN202410206908.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-23
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, the hard disk health assessment method is single and cannot accurately reflect the health status of the hard disk, resulting in poor user experience and inability to deal with the faulty disk in time, increasing the risk of data loss and system paralysis.
The multi-dimensional evaluation method is adopted to comprehensively evaluate the health status of the hard disk through multiple health indicators, error code information and delay information of the target hard disk, including the first evaluation information, the second evaluation information and the third evaluation information, respectively reflecting the overall performance of the hard disk, the accumulation of error codes and the serious delay of the hard disk, and setting preset weights and thresholds to improve the accuracy of the evaluation.
It improves the accuracy of hard disk health assessment, enhances user experience, can handle faulty disks in a timely manner, reduces the risk of data loss, and is suitable for the needs of different application scenarios.
Smart Images

Figure CN120540945A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a hard disk health assessment method, a hard disk management method, and related devices. Background Art
[0002] With the rapid development of big data, cloud computing, and artificial intelligence technologies, the demand for highly reliable storage systems is increasing. Hard drive reliability has become a key factor limiting storage system reliability. Hard drive failures can lead to user data loss and corruption, reduced system read / write performance, and even storage system failure. Methods for estimating hard drive health typically rely on system-provided hard drive assessment information to handle failed drives, resulting in a poor user experience. Summary of the Invention
[0003] In order to solve the above problems, the embodiments of the present application provide a hard disk health assessment method, a hard disk management method and related devices, which can estimate the hard disk health from multiple dimensions based on multiple assessment information, comprehensively evaluate the health status of the hard disk, increase the accuracy of the assessment results, and improve the user experience.
[0004] To this end, the following technical solutions are adopted in the embodiments of the present application:
[0005] In a first aspect, embodiments of the present application provide a hard drive health assessment method, comprising: determining three pieces of assessment information for a target hard drive: first assessment information, second assessment information, and third assessment information; and comprehensively assessing the health status of the target hard drive based on the three pieces of assessment information. The first assessment information is determined based on multiple health indicators of the target hard drive; the second assessment information is determined based on error code information of the target hard drive; and the third assessment information is determined based on latency information of the target hard drive.
[0006] In an embodiment of the present application, the health status of the target hard disk is comprehensively evaluated through multiple dimensions based on three evaluation information, and the evaluation results are more objective and more accurate, thereby improving the user experience. Among them, the first evaluation information is determined based on multiple health indicators of the target hard disk, and therefore, the first evaluation information represents the overall performance of the target hard disk. The second evaluation information is determined based on the error code information of the target hard disk, and therefore, the second evaluation information represents the accumulation of error codes of the target hard disk. The third evaluation information is determined based on the latency information of the target hard disk, and therefore, the third evaluation information represents the severity of the latency of the target hard disk. Generally, when a minor fault occurs in the hard disk, the three evaluation information of the target hard disk will be poor, or when a serious fault occurs in the hard disk, the three evaluation information of the target hard disk will be very poor, and even at least one of the three evaluation information will show an evaluation result of "faulty disk". Therefore, the three evaluation information can be used to understand the degree of health of the hard disk and deal with the faulty disk in a timely manner.
[0007] Furthermore, the first evaluation information can obtain corresponding scores based on multiple health indicators of the target hard disk, thereby reflecting the different stages of the health of the target hard disk, that is, the health of the target hard disk. The second evaluation information can use error code information to more objectively understand the health of the hard disk. For example, when a small number of error codes appear in a short period of time, it may be just a temporary hard disk failure, and the possibility of hard disk failure is not high. Therefore, the second evaluation information can evaluate the health of the hard disk through error code information, thereby improving the accuracy of hard disk evaluation. The third evaluation information can reflect the severity of the target hard disk's latency based on the degree of latency, such as a normal disk, a slightly slow disk, a moderately slow disk, a severely slow disk, etc. In other words, by understanding the different health levels of the target hard disk, faulty disks or slow disks can be handled in a timely manner, which increases the flexibility of hard disk processing and reduces the risk of data loss.
[0008] In one possible implementation, determining the first evaluation information based on multiple health indicators of the target hard disk includes: obtaining indicator data corresponding to the multiple health indicators of the target hard disk; weighting the indicator data according to preset weights to obtain a first evaluation score of the target hard disk; comparing the first evaluation score with the preset evaluation score to determine the first evaluation information; wherein, when the first evaluation score is lower than the preset evaluation score, the first evaluation information indicates that the target hard disk is a faulty disk.
[0009] In this implementation, the preset weights of multiple health indicators of the target hard disk can be determined according to actual needs. The indicator data of multiple health indicators are weighted by the preset weights to obtain a first evaluation score of the target hard disk. By setting a reference preset evaluation score, the first evaluation score and the preset evaluation score are compared to determine the first evaluation information. The comparison method of the first evaluation score and the preset evaluation score can be a ratio method, a subtraction method, etc. Through the comparison results, the overall performance of the target hard disk can be understood through the first evaluation information. Among them, by setting preset weights, the degree of importance attached to different health indicators is expressed, which meets the different needs of users. The preset evaluation score can be a score set by the user, or it can be an expert score or experience score recommended by the system, so that the reference standards of multiple health indicators are diverse and modifiable, meeting the different needs of users and improving the user experience.
[0010] In another possible implementation, the preset weights include a first weight and a second weight, the first weight represents the weight of the first-level health indicator, the second weight represents the weight of the second-level health indicator, the first-level health indicator represents the category of multiple health indicators of the target hard disk, and the second-level health indicator represents multiple health indicators under the first-level health indicator.
[0011] In this implementation, multiple health indicators of the target hard drive are categorized to generate multiple first-level health indicators. Each first-level health indicator also includes multiple second-level health indicators. By assigning different weights to health indicators of different categories and levels, the importance attached to different health indicators is increased. Therefore, the first assessment information can be applied to different application scenarios and meet the diverse needs of users.
[0012] In another possible implementation, determining the second evaluation information based on the error code information of the target hard disk includes: determining the second evaluation information by judging the number of error codes within a preset time range; wherein, within the preset time range, when the number of error codes exceeds a quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk.
[0013] In this implementation, the second evaluation information includes error code information, which can more objectively reflect the health status of the hard disk. By controlling the preset time range and quantity threshold of the error code sliding window, the cumulative duration and number of error codes are changed, the accumulation of certain temporary error codes is avoided, and the accuracy of the hard disk evaluation is improved. In other words, when a small number of error codes appear in a short period of time, it may be just a temporary hard disk failure, and the possibility of hard disk failure is not high, so it is necessary to clear the small number of error codes in time; when a large number of error codes appear in a short period of time, the possibility of hard disk failure is high. By judging the number of error codes within the preset time range, the second evaluation information can clear the small number of error codes within the preset time range in time, thereby improving the accuracy of the hard disk evaluation. In addition, the preset time range and quantity threshold can be controlled to suit different application scenarios.
[0014] In another possible implementation, by judging the number of error codes within a preset time range, the second evaluation information is determined, including: determining the target error code, and recording the time length and the number of error codes starting from the occurrence of the target error code; when the recorded time length reaches the preset time range and the number of recorded error codes exceeds the quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk; when the recorded time length reaches the preset time range and the number of recorded error codes does not exceed the quantity threshold, the second evaluation information indicates that the target hard disk is a normal disk, and the recorded time length and number of error codes are cleared, and the next error code is determined as the target error code.
[0015] In this implementation, the target error code is first determined. The time duration and number of error codes are recorded starting from the occurrence of the target error code. If both the recorded time duration and the number of error codes reach a threshold, the second evaluation information determines that the target hard drive is faulty. If the recorded time duration reaches a preset time range and the number of error codes does not exceed the threshold, the second evaluation information indicates that the target hard drive is functioning properly. The recorded time duration and number of error codes are cleared, the next error code is determined as the target error code, and this step is repeated. Therefore, compared to other error code evaluation methods, the second evaluation information can promptly clear a small number of error codes within the preset time range, thereby improving the accuracy of hard drive evaluation.
[0016] In another possible implementation, based on the delay information of the target hard disk, determining the third evaluation information includes: obtaining the delay time length of the target hard disk; determining the slow disk level of the target hard disk based on the relationship between the delay time length and multiple delay thresholds; and determining the third evaluation information of the target hard disk based on the slow disk level of the target hard disk.
[0017] In this implementation, multiple latency thresholds are set and the target drive's latency is compared with the multiple latency thresholds to determine the target drive's slowness level. Based on the target drive's slowness level, the third evaluation information is used to determine the target drive's evaluation result. This third evaluation information allows the target drive's slowness level to be understood, allowing for different treatments to be applied to the target drive. This increases the diversity of the target drive's evaluation results and allows for targeted, differentiated treatment measures to be taken.
[0018] In another possible implementation, the slow disk level of the target hard disk is determined based on the relationship between the delay time length and multiple delay thresholds, including: when the delay time length is less than the first delay threshold, the target hard disk is judged to be a normal disk; when the delay time length is greater than or equal to the first delay threshold and less than the second delay threshold, the target hard disk is judged to be a mildly slow disk; when the delay time length is greater than or equal to the second delay threshold and less than the third delay threshold, the target hard disk is judged to be a moderately slow disk; when the delay time length is greater than or equal to the third delay threshold and less than the fourth delay threshold, the target hard disk is judged to be a severely slow disk.
[0019] In this implementation, four latency thresholds are set to categorize hard drive slowness into normal, mildly slow, moderately slow, and severely slow levels. Different evaluation results can be obtained based on the target hard drive's slowness level, leading to different handling strategies.
[0020] In a second aspect, an embodiment of the present application provides a hard disk management method, comprising: determining the health status of multiple hard disks according to any one of the above methods; and managing the multiple hard disks according to the health status of the multiple hard disks.
[0021] In an embodiment of the present application, the management of multiple hard disks can be achieved based on the above-mentioned hard disk health assessment method. For example, the multiple hard disks can be a Redundant Array of Inexpensive Disks (RAID). In a multi-hard disk system, the health status of each hard disk can be obtained through the above-mentioned hard disk health assessment method, thereby managing multiple hard disks. The above-mentioned hard disk health assessment method increases the accuracy of the assessment results, thereby improving the reliability of multiple hard disk management when managing multiple hard disks.
[0022] In one possible implementation, managing multiple hard disks based on their health status includes: when a faulty disk occurs among the multiple hard disks, determining management strategies for the multiple hard disks, the management strategies including at least one of a hard disk removal strategy, a cache modification strategy, a hard disk isolation strategy, and a degraded read / write strategy for the remaining hard disks.
[0023] In this implementation, in a multi-disk system, multiple hard disks can be managed based on their health status. For example, when a faulty disk occurs among the multiple hard disks, the management strategy for the multiple hard disks can be changed to promptly handle the faulty disk.
[0024] In a third aspect, an embodiment of the present application provides a hard disk health assessment device, including: an assessment information determination module, used to determine first assessment information based on multiple health indicators of the target hard disk; and to determine second assessment information based on the error code information of the target hard disk; and to determine third assessment information based on the delay information of the target hard disk; a health status assessment module, used to evaluate the health status of the target hard disk based on the first assessment information, the second assessment information and the third assessment information.
[0025] In one possible implementation, the evaluation information determination module is specifically used to: obtain indicator data corresponding to multiple health indicators of the target hard disk; weight the indicator data according to preset weights to obtain a first evaluation score of the target hard disk; compare the first evaluation score with the preset evaluation score to determine first evaluation information; wherein, when the first evaluation score is lower than the preset evaluation score, the first evaluation information indicates that the target hard disk is a faulty disk.
[0026] In another possible implementation, the preset weights include a first weight and a second weight, the first weight represents the weight of the first-level health indicator, the second weight represents the weight of the second-level health indicator, the first-level health indicator represents the category of multiple health indicators of the target hard disk, and the second-level health indicator represents multiple health indicators under the first-level health indicator.
[0027] In another possible implementation, the evaluation information determination module is specifically used to determine the second evaluation information by judging the number of error codes within a preset time range; wherein, within the preset time range, when the number of error codes exceeds a quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk.
[0028] In another possible implementation, the evaluation information determination module is specifically used to: determine the target error code, and record the time length and the number of error codes starting from the occurrence of the target error code; when the recorded time length reaches the preset time range and the number of recorded error codes exceeds the quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk; when the recorded time length reaches the preset time range and the number of recorded error codes does not exceed the quantity threshold, the second evaluation information indicates that the target hard disk is a normal disk, and clears the recorded time length and number of error codes, and determines the next error code as the target error code.
[0029] In another possible implementation, the evaluation information determination module is specifically used to: obtain the delay time length of the target hard disk; determine the slow disk level of the target hard disk based on the relationship between the delay time length and multiple delay thresholds; and determine the third evaluation information of the target hard disk based on the slow disk level of the target hard disk.
[0030] In another possible implementation, the evaluation information determination module is specifically used to: when the delay time length is less than the first delay threshold, judge that the target hard disk is a normal disk; when the delay time length is greater than or equal to the first delay threshold and less than the second delay threshold, judge that the target hard disk is a mildly slow disk; when the delay time length is greater than or equal to the second delay threshold and less than the third delay threshold, judge that the target hard disk is a moderately slow disk; when the delay time length is greater than or equal to the third delay threshold and less than the fourth delay threshold, judge that the target hard disk is a severely slow disk.
[0031] In a fourth aspect, an embodiment of the present application provides a hard disk management device, comprising: a determination module for determining the health status of multiple hard disks from a health status assessment module based on any one of the above-mentioned hard disk health assessment devices; and a management module for managing multiple hard disks based on the health status of the multiple hard disks.
[0032] In one possible implementation, the management module is specifically used to: when a faulty disk occurs among multiple hard disks, determine the management strategy of the multiple hard disks, the management strategy including at least one of a hard disk removal strategy, a cache modification strategy, a hard disk isolation strategy, and a degraded read and write strategy for the remaining hard disks.
[0033] In a fifth aspect, an embodiment of the present application provides an independent disk redundant array RAID card, including a processing circuit and a storage medium, in which computer program code is stored; when the computer program code is executed by the processing circuit, any of the above methods or the algorithm function embodied by any of the above devices is implemented.
[0034] In a sixth aspect, an embodiment of the present application provides a motherboard including the above-mentioned Redundant Array of Independent Disks RAID card, wherein the motherboard is electrically connected to a plurality of hard disks so that the Redundant Array of Independent Disks RAID card manages the plurality of hard disks.
[0035] In a seventh aspect, an embodiment of the present application provides a server comprising the above-mentioned Redundant Array of Independent Disks RAID card or the above-mentioned mainboard and multiple hard disks, wherein the Redundant Array of Independent Disks RAID card is used to manage multiple hard disks.
[0036] In an eighth aspect, a computing device includes a processor and a memory, wherein a program is stored in the memory. When the processor executes the program, it can implement the algorithm function embodied by any of the above methods or any of the above devices.
[0037] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, any of the above methods or the algorithm function embodied in any of the above devices is implemented.
[0038] In a tenth aspect, an embodiment of the present application provides a computer program product, which includes program instructions. When the program instructions are executed by a computer, the computer executes any of the above methods or the algorithmic functions embodied by any of the above devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The following is a brief introduction to the drawings required for use in the embodiments or technical descriptions.
[0040] Figure 1 This is a flowchart of the method for the first solution;
[0041] Figure 2 This is a flow chart of the method for the second solution;
[0042] Figure 3 A flowchart of a hard disk health assessment method provided in an embodiment of the present application;
[0043] Figure 4 This is a schematic diagram of the composition of the first evaluation score in the first evaluation information provided in an embodiment of the present application;
[0044] Figure 5This is a schematic diagram of an error code sliding window in the second evaluation information provided in an embodiment of the present application;
[0045] Figure 6 This is a schematic diagram of different slow disk levels in the third evaluation information provided in an embodiment of the present application;
[0046] Figure 7 A flowchart of a hard disk management method provided in an embodiment of the present application;
[0047] Figure 8 This is a schematic diagram of an example of a process for handling a faulty disk provided in an embodiment of the present application;
[0048] Figure 9 This is a schematic diagram of another example of a process for handling a faulty disk provided in an embodiment of the present application;
[0049] Figure 10 A schematic diagram of the composition of a hard disk health assessment device provided in an embodiment of the present application;
[0050] Figure 11 A schematic diagram of the composition of a hard disk management device provided in an embodiment of the present application;
[0051] Figure 12 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship. For example, A / B means either A or B.
[0053] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.
[0054] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0055] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0056] In order to facilitate understanding of the solution provided by the embodiment of the present application, a brief introduction to some of the terms involved in the solution is first given.
[0057] With the rapid development of big data, cloud computing, and artificial intelligence technologies, the demand for high-reliability storage systems is increasing. The reliability of hard drives in storage systems has become one of the important factors restricting the reliability of storage systems. Hard drive failures can lead to user data loss and damage, reduced system read and write performance, and storage system paralysis.
[0058] To protect important data on a hard drive, you shouldn't rely solely on repair or data recovery. You should also pay attention to the drive's health during use. Monitoring results can be used to predict the condition of the drive. If a drive is in poor condition, back up important files or replace it immediately to prevent data loss and system crashes. Therefore, predicting and assessing the health of the hard drive is crucial for ensuring stable system operation.
[0059] The first solution is to use Self-monitoring Analysis and Reporting Technology (SMART) to determine the health of the hard drive. Figure 1 As shown, the RAID card obtains the hard drive's SMART information and then compares it with the hard drive's Threshold value to determine whether the hard drive is faulty. If any factor in the SMART information falls below the Threshold value, the hard drive is considered faulty and the relevant hard drive failure procedures are executed.
[0060] The second solution is to use the Proactive Disk Health Monitoring (PHDM) method to determine the health of the hard disk. Figure 2 As shown in the figure, the health of the hard drive is predicted by determining the number of bad sectors within the bad sector time window. When the number of bad sectors within the bad sector time window exceeds the threshold, the hard drive is judged to be a failed drive and the relevant process for hard drive failure is executed.
[0061] In the first and second solutions, the evaluation parameters are single and the interplay of multiple factors is not considered. For example, the PHDM analysis method only determines the number of bad sectors within the bad sector time window. The predicted data is not intuitive and cannot accurately determine the current health status of the hard drive. For example, the SMART analysis method determines that the hard drive is faulty and executes the relevant processes for hard drive failure, but cannot provide the current health status of the hard drive. The evaluation parameters are few and the versatility is poor. For example, the PHDM analysis method or the SMART analysis method cannot provide multi-dimensional hard drive health assessment information, and therefore cannot comprehensively analyze the health status of the hard drive, and is only suitable for specific application scenarios.
[0062] In order to solve the problems of the first and second solutions, such as Figure 3 As shown, the embodiment of the present application provides a hard disk health assessment method, which mainly includes the following steps:
[0063] Step S301: Determine first evaluation information based on multiple health indicators of the target hard disk.
[0064] Optionally, determining the first assessment information based on multiple health indicators of the target hard drive includes: obtaining indicator data corresponding to the multiple health indicators of the target hard drive; weighting the indicator data according to preset weights to obtain a first assessment score for the target hard drive; and comparing the first assessment score with a preset assessment score to determine the first assessment information; wherein, when the first assessment score is lower than the preset assessment score, the first assessment information indicates that the target hard drive is a faulty drive. By setting preset weights to indicate the degree of emphasis on different health indicators, different user needs are met. The first assessment information can be applied to different application scenarios, meeting the diverse needs of users.
[0065] Furthermore, the indicator data of multiple health indicators is weighted using preset weights to obtain a first evaluation score for the target hard drive. A reference preset evaluation score is then set and the first evaluation score is compared with the preset evaluation score to determine first evaluation information. The first evaluation score and the preset evaluation score can be compared using a ratio method, a subtraction method, or other methods. The comparison results, based on the first evaluation information, provide an understanding of the overall performance of the target hard drive.
[0066] Optionally, the preset weights include a first weight and a second weight, the first weight represents the weight of the first-level health indicator, the second weight represents the weight of the second-level health indicator, the first-level health indicator represents the category of multiple health indicators of the target hard disk, and the second-level health indicator represents multiple health indicators under the first-level health indicator.
[0067] Furthermore, if Figure 4As shown, the target hard drive's multiple health indicators are categorized into multiple first-level health indicators. Each first-level health indicator also includes multiple second-level health indicators. By assigning different weights to health indicators of different categories and levels, the importance of health indicators of different categories is increased, making it easier to set preset weights.
[0068] For example, in an embodiment of the present application, based on the hard drive sub-health factors, the hard drive sub-health factors can be divided into multiple factor sets, a judgment set can be defined, weights can be set for the divided factors, and a first evaluation score can be obtained through comprehensive calculation. The first evaluation score can be obtained by comparing the factor set and the judgment set. For example, the calculation formula for the first evaluation score can be shown as formula (1) and formula (2).
[0069]
[0070]
[0071] In formula (1) and formula (2), C represents the first evaluation score, U represents the first-level health indicator, W represents the weight of the first-level health indicator, u represents the second-level health indicator, and w represents the weight of the second-level health indicator.
[0072] For example, first-level health indicators can include power-on time, I / O-related information, and disk smart information. Taking I / O-related information as an example, second-level health indicators for I / O-related information can include average I / O response time, I / O access interval, and so on. Second-level health indicators for disk smart information can include multiple factors within disk smart information.
[0073] Step S302: Determine second evaluation information based on the error code information of the target hard disk.
[0074] Optionally, based on the error code information of the target hard disk, determining the second evaluation information includes: determining the second evaluation information by judging the number of error codes within a preset time range; wherein, within the preset time range, when the number of error codes exceeds a quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk.
[0075] Furthermore, the second assessment information includes error code information, which can more objectively reflect the health status of the hard drive. By controlling the preset time range and number threshold of the error code sliding window, the accumulation time and number of error codes are changed, avoiding the accumulation of certain temporary error codes and improving the accuracy of hard drive assessment. In practice, through experimental data analysis, the second assessment information of the embodiment of the present application has a prediction accuracy for faulty disks that is more than double that of the existing SMART information.
[0076] In other words, when a small number of error codes appear within a short period of time, it's likely just a temporary hard drive failure, and the likelihood of a full-blown hard drive failure is low. Therefore, these small numbers of error codes need to be cleared promptly. By determining the number of error codes within a preset time range, the second assessment information can promptly clear these small numbers of error codes within that time range, thereby improving the accuracy of the hard drive assessment. Furthermore, the preset time range and number threshold can be controlled to suit different application scenarios.
[0077] Optionally, by judging the number of error codes within a preset time range, determining the second evaluation information includes: determining the target error code, and recording the time length and the number of error codes starting from the occurrence of the target error code; when the recorded time length reaches the preset time range and the number of recorded error codes exceeds the quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk; when the recorded time length reaches the preset time range and the number of recorded error codes does not exceed the quantity threshold, the second evaluation information indicates that the target hard disk is a normal disk, and clears the recorded time length and number of error codes, and determines the next error code as the target error code.
[0078] Furthermore, the target error code is first determined, and the time length and number of error codes are recorded starting from the occurrence of the target error code. When it is determined that both the recorded time length and the number of error codes reach a threshold, the second evaluation information determines that the target hard drive is a faulty drive. When the recorded time length reaches a preset time range and the number of error codes does not exceed the threshold, the second evaluation information indicates that the target hard drive is normal, clears the recorded time length and number of error codes, and determines the next error code as the target error code, repeating this step. Therefore, compared with other error code evaluation methods, the second evaluation information can promptly clear a small number of error codes within the preset time range, thereby improving the accuracy of hard drive evaluation.
[0079] For example, Figure 5 As shown, assuming that the error code number threshold is 5, when error code 1 appears, the time starts to count, that is, the time sliding window starts, and the statistics start from error 1. When the recorded time reaches the time threshold (preset time range), the error code count is 4. If the error code number threshold is not reached, the error code number will be counted again from the moment the next error code (that is, error code 5) appears. In the second statistical period, if Figure 5 As shown in the figure, when the recorded time reaches the time threshold, the error code count is 7, which exceeds the error code number threshold, then the target hard disk is judged to be a faulty disk and the relevant processing work for the faulty disk is performed. The timing starts at the moment Figure 5 The start time is shown, and the end time is as follows Figure 5 The end time shown.
[0080] For example, the error code may be an IO error code, and the IO error code may be an uncorrectable ECC code (Uncorrectable Error-correcting Code Error, UNC), and the UNC may indicate that a bad block exists in the hard disk.
[0081] Step S303: Determine third evaluation information based on the latency information of the target hard disk.
[0082] Optionally, based on the delay information of the target hard disk, determining the third evaluation information includes: obtaining the delay time length of the target hard disk; determining the slow disk level of the target hard disk based on the relationship between the delay time length and multiple delay thresholds; and determining the third evaluation information of the target hard disk based on the slow disk level of the target hard disk.
[0083] Furthermore, by setting multiple latency thresholds and comparing the target drive's latency length with the multiple latency thresholds, the target drive's slowness level is determined. Based on the target drive's slowness level, the third evaluation information is used to determine the target drive's evaluation result. This third evaluation information can be used to understand the target drive's slowness level and, therefore, to perform different treatments on the target drive. This increases the diversity of the target drive's evaluation results and allows for the implementation of targeted and different treatment measures.
[0084] For example, Figure 6 As shown in the figure, the statistical data of the latency of a hard disk is used for analysis. Figure 6 As shown, when the delay time is less than the first delay threshold, the target hard drive is judged to be normal; when the delay time is greater than or equal to the first delay threshold and less than the second delay threshold, the target hard drive is judged to be mildly slow; when the delay time is greater than or equal to the second delay threshold and less than the third delay threshold, the target hard drive is judged to be moderately slow; when the delay time is greater than or equal to the third delay threshold and less than the fourth delay threshold, the target hard drive is judged to be severely slow. Different evaluation results can be obtained based on the slowness level of the target hard drive, and different processing strategies can be implemented.
[0085] In other embodiments, other factors and indicators can also be used to determine whether a hard disk is a slow disk. For example, if the "number of I / O waits," "average I / O wait time," "service time per I / O," and "utilization" of a hard disk are significantly higher than those of other hard disks over a sustained period of time, then the hard disk is considered a "slow disk."
[0086] Step S304: Evaluate the health status of the target hard disk based on the first evaluation information, the second evaluation information, and the third evaluation information. In this embodiment of the present application, by determining the three evaluation information of the target hard disk: the first evaluation information, the second evaluation information, and the third evaluation information, the health status of the target hard disk can be comprehensively evaluated.
[0087] In an embodiment of the present application, the health status of the target hard disk is comprehensively evaluated through multiple dimensions based on three evaluation information, and the evaluation results are more objective and more accurate, thereby improving the user experience. Among them, the first evaluation information is determined based on multiple health indicators of the target hard disk, and therefore, the first evaluation information represents the overall performance of the target hard disk. The second evaluation information is determined based on the error code information of the target hard disk, and therefore, the second evaluation information represents the accumulation of error codes of the target hard disk. The third evaluation information is determined based on the latency information of the target hard disk, and therefore, the third evaluation information represents the severity of the latency of the target hard disk. Generally, when a minor fault occurs in the hard disk, the three evaluation information of the target hard disk will be poor, or when a serious fault occurs in the hard disk, the three evaluation information of the target hard disk will be very poor, and even at least one of the three evaluation information will show an evaluation result of "faulty disk". Therefore, the three evaluation information can be used to understand the degree of health of the hard disk and deal with the faulty disk in a timely manner.
[0088] Furthermore, the first evaluation information can obtain corresponding scores based on multiple health indicators of the target hard disk, thereby reflecting the different stages of the health of the target hard disk, that is, the health of the target hard disk. The second evaluation information can use error code information to more objectively understand the health of the hard disk. For example, when a small number of error codes appear in a short period of time, it may be just a temporary hard disk failure, and the possibility of hard disk failure is not high. Therefore, the second evaluation information can evaluate the health of the hard disk through error code information, thereby improving the accuracy of hard disk evaluation. The third evaluation information can reflect the severity of the target hard disk's latency based on the degree of latency, such as a normal disk, a slightly slow disk, a moderately slow disk, a severely slow disk, etc. In other words, by understanding the different health levels of the target hard disk, faulty disks or slow disks can be handled in a timely manner, which increases the flexibility of hard disk processing and reduces the risk of data loss.
[0089] One of the uses of the hard disk health assessment method provided in the embodiment of the present application is to manage multiple hard disks based on multiple assessment information. Figure 7 As shown, the embodiment of the present application also provides a hard disk management method, which mainly includes the following steps:
[0090] In step S701 , the health status of multiple hard disks is determined according to the hard disk health assessment method provided in an embodiment of the present application.
[0091] Step S702: managing the multiple hard disks according to the health status of the multiple hard disks.
[0092] In the embodiments of the present application, the hard drive health assessment method described above can be used to manage multiple hard drives. For example, the multiple hard drives can be configured as a Redundant Array of Inexpensive Disks (RAID). In a multi-hard drive system, the health status of each hard drive can be determined using the hard drive health assessment method described above, allowing for management of multiple hard drives.
[0093] Optionally, managing the multiple hard drives based on their health status includes determining a management policy for the multiple hard drives when a faulty hard drive occurs. The management policy includes at least one of a hard drive removal policy, a cache modification policy, a hard drive isolation policy, and a degraded read / write policy for the remaining hard drives. In other words, by changing the management policy for the multiple hard drives, the faulty hard drive can be promptly addressed.
[0094] It should be further explained that the hard disk health assessment method provided in the embodiment of the present application can be applicable to the application scenario of a single hard disk, and can also be applied to the application scenario of multiple hard disks. The embodiment of the present application does not limit this. In the application scenario of a single hard disk, the hard disk health assessment method provided in the embodiment of the present application can be used to evaluate the health status of a certain hard disk and monitor the health status of the hard disk so as to handle the hard disk in a timely manner. In the application scenario of multiple hard disks, the embodiment of the present application is explained based on a RAID card, but the application scenario of multiple hard disks is not limited to the application scenario of a RAID card. For example, it can also be a hard disk array box and other related methods.
[0095] Generally speaking, RAID is a technology that combines multiple independent physical hard drives in various ways to form a single logical drive, thereby providing higher performance and data redundancy than a single hard drive. RAID cards are generally divided into hardware RAID and software RAID. Hardware RAID implements RAID functionality through hardware, including standalone RAID cards and motherboard-integrated RAID chips. RAID cards that utilize software and the CPU perform common RAID calculations. Software RAID consumes a lot of CPU resources, and the vast majority of server equipment uses hardware RAID.
[0096] like Figure 8 As shown, this embodiment of the present application provides an example of a process for handling a failed disk. For example, using a RAID card as an example, the disk array managed by the RAID card includes Hard Drive 1, Hard Drive 2, Hard Drive 3, Hard Drive 4, ..., Hard Drive N. When the RAID card determines that Hard Drive 3 is a failed disk based on the hard drive health assessment method provided in this embodiment of the application, it executes the failed disk handling process.
[0097] The process for handling a failed drive is explained using the third evaluation information (slow drive information) as an example. In addition to defining the slow drive level by setting multiple time thresholds, another example uses the I / O latency relationship between the drives in the RAID group to determine the predicted latency score for the target drive. The slow drive level set based on the latency score is then used to determine the target drive's slow drive level.
[0098] For example, the latency score can be L0, L1, L2, L3, etc. A latency score of L0 indicates that the target hard drive is a normal drive. A latency score of L1 indicates that the average latency of the target hard drive is significantly higher than that of normal drives, and the target hard drive is mildly slow. A latency score of L2 indicates that the average latency of the target hard drive is excessively high, and the target hard drive is moderately slow. A latency score of L3 indicates that the average latency of the target hard drive is several times that of normal drives and more than eight times that of other drives in the same RAID group, and the target hard drive is severely slow.
[0099] For example, the process of handling a failed disk (slow disk) may include: the disk management module of the RAID card scores the IO latency of multiple hard disks in the same RAID group under the RAID card; based on the slow disk level set by the latency score, the slow disk level of multiple hard disks is determined, and the write-back / write-through strategy of switching cache (Cache) and the kick-disk strategy of the redundant RAID level are executed on the slow disk; the disk management module isolates the slow disk, performs fuse processing on the IO sent to the slow disk, and implements a downgraded read and write strategy for other hard disks in the RAID group.
[0100] like Figure 9 As shown, the embodiment of the present application provides another example of the process of handling a faulty disk. For example, according to a hard disk health assessment method provided in an embodiment of the present application, the health of multiple hard disks is determined, assuming that the health of the hard disks is presented in the form of health scores. The health scores of multiple hard disks are displayed on the command-line interface (CLI) of the RAID card; the RAID card defines the hard disk status of the multiple hard disks under the RAID card based on the quantified health scores; based on the hard disk status of the multiple hard disks, the faulty disk is then handled. For example, as Figure 9 As shown, another example of a process for handling a failed disk provided in an embodiment of the present application mainly includes the following steps:
[0101] Step S901: The RAID card periodically scores the health of multiple hard disks to obtain health scores of the multiple hard disks. Figure 9 As shown in the figure, among the normal member disks, the health score of the first hard disk is 1, the health score of the second hard disk is 4, and the health score of the third hard disk is 5. There is at least one idle hot spare disk with a health score of 5 in the disk array.
[0102] In step S902, if the health score of the first hard drive is 1 and the health scores of the remaining hard drives are 2 or greater, pre-copy is initiated directly, and pre-copy is performed on the first hard drive with a health score of 1. The idle hot spare drive with a health score of 5 becomes a diagnostic mode drive. An idle hot spare drive during the pre-copy process is also called a used hot spare drive. The first hard drive with a health score of 1 becomes a pre-copy member drive.
[0103] Step S903: After the pre-copy is successful, the idle hot spare disk with a health score of 5 further becomes a normal member disk with a health score of 5; the first hard disk with a health score of 1 further becomes a faulty disk and is no longer used.
[0104] Step S904: After the failed disk is replaced, the copyback is successful. Among the functioning member disks, the health score of the first hard disk (i.e., the idle hot spare disk after the pre-copy was successful) is 5, the health score of the second hard disk is 4, and the health score of the third hard disk is 5. To ensure redundancy in the disk array, the disk array must have at least one idle hot spare disk with a health score of 5.
[0105] In the embodiment of the present application, by predicting the health of the hard disk under the RAID card and replacing the faulty disk in time based on the RAID preservation strategy, the reliability of the RAID card is improved. The hard disk health evaluation method of the embodiment of the present application is a quantitative scoring mechanism with strong scalability. New health indicators can be added according to actual conditions to make the judgment result more accurate. The hard disk health evaluation method of the embodiment of the present application has a high recognition rate. The fault detection rate (FDR) of the hard disk health prediction can reach more than 66%, and the false alarm rate (FAR) or false alarm rate is less than 0.08%.
[0106] A hard disk management method according to an embodiment of the present application is a RAID group processing strategy and process based on hard disk health prediction, with a high fault detection rate for faulty disks, a low false alarm rate, and a more robust reliability of the RAID card.
[0107] Based on the same concept as the above embodiments, Figure 10 As shown, the embodiment of the present application provides a hard disk health assessment device 1000, which mainly includes:
[0108] The evaluation information determination module 1001 is used to determine first evaluation information based on multiple health indicators of the target hard disk; determine second evaluation information based on error code information of the target hard disk; and determine third evaluation information based on latency information of the target hard disk.
[0109] The health status evaluation module 1002 is configured to evaluate the health status of the target hard disk according to the first evaluation information, the second evaluation information, and the third evaluation information.
[0110] In one possible implementation, the evaluation information determination module 1001 is specifically used to: obtain indicator data corresponding to multiple health indicators of the target hard disk; weight the indicator data according to preset weights to obtain a first evaluation score of the target hard disk; compare the first evaluation score with the preset evaluation score to determine the first evaluation information; wherein, when the first evaluation score is lower than the preset evaluation score, the first evaluation information indicates that the target hard disk is a faulty disk.
[0111] In another possible implementation, the preset weights include a first weight and a second weight, the first weight represents the weight of the first-level health indicator, the second weight represents the weight of the second-level health indicator, the first-level health indicator represents the category of multiple health indicators of the target hard disk, and the second-level health indicator represents multiple health indicators under the first-level health indicator.
[0112] In another possible implementation, the evaluation information determination module 1001 is specifically used to determine the second evaluation information by judging the number of error codes within a preset time range; wherein, within the preset time range, when the number of error codes exceeds a quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk.
[0113] In another possible implementation, the evaluation information determination module 1001 is specifically used to: determine the target error code, and record the time length and the number of error codes starting from the occurrence of the target error code; when the recorded time length reaches the preset time range and the number of recorded error codes exceeds the quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk; when the recorded time length reaches the preset time range and the number of recorded error codes does not exceed the quantity threshold, the second evaluation information indicates that the target hard disk is a normal disk, and clears the recorded time length and number of error codes, and determines the next error code as the target error code.
[0114] In another possible implementation, the evaluation information determination module 1001 is specifically used to: obtain the delay time length of the target hard disk; determine the slow disk level of the target hard disk based on the relationship between the delay time length and multiple delay thresholds; and determine the third evaluation information of the target hard disk based on the slow disk level of the target hard disk.
[0115] In another possible implementation, the evaluation information determination module 1001 is specifically used to: when the delay time length is less than the first delay threshold, judge that the target hard disk is a normal disk; when the delay time length is greater than or equal to the first delay threshold and less than the second delay threshold, judge that the target hard disk is a mildly slow disk; when the delay time length is greater than or equal to the second delay threshold and less than the third delay threshold, judge that the target hard disk is a moderately slow disk; when the delay time length is greater than or equal to the third delay threshold and less than the fourth delay threshold, judge that the target hard disk is a severely slow disk.
[0116] Based on the same concept as the above embodiments, Figure 11 As shown, the embodiment of the present application provides a hard disk management device 1100, which mainly includes:
[0117] The determination module 1101 is configured to determine the health status of multiple hard disks from the health status assessment module according to any one of the hard disk health assessment devices described above.
[0118] The management module 1102 is configured to manage multiple hard disks according to the health status of the multiple hard disks.
[0119] In one possible implementation, the management module 1102 is specifically used to: when a faulty disk occurs among multiple hard disks, determine the management strategy for multiple hard disks, the management strategy including at least one of a hard disk removal strategy, a cache modification strategy, a hard disk isolation strategy, and a degraded read and write strategy for the remaining hard disks.
[0120] Based on the same concept as the aforementioned embodiment, a computing device is also provided in the embodiment of the present application. The computing device includes at least a processor and a memory, and a program is stored in the memory. When the processor executes the program, it can implement the algorithm function embodied by the above method or device.
[0121] Figure 12 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. Figure 12 As shown, the computing device 1200 includes at least one processor 1201, a memory 1202, and a communication interface 1203. The processor 1201, the memory 1202, and the communication interface 1203 are communicatively connected, and the communication connection can be achieved through a wired (e.g., bus) or wireless communication. The communication interface 1203 is used to receive data sent by other devices; the memory 1202 stores computer instructions, and the processor 1201 executes the computer instructions to perform the method in the aforementioned method embodiment.
[0122] It should be understood that in the embodiment of the present application, the processor 1201 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware UI components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0123] The memory 1202 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1201. The memory 1202 may also include a nonvolatile random access memory.
[0124] The memory 1202 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct ram bus RAM (DR RAM).
[0125] It should be understood that the computing device 1200 according to the embodiment of the present application can execute the method mentioned in the embodiment of the present application. The detailed description of the implementation of the method is given above and will not be repeated here for the sake of brevity.
[0126] Based on the same concept as the aforementioned embodiments, an embodiment of the present application provides a redundant array of independent disks (RAID) card, comprising a processing circuit and a storage medium, wherein the storage medium stores computer program code; when the computer program code is executed by the processing circuit, it implements any of the above methods or the algorithmic function embodied in any of the above devices.
[0127] Based on the same concept as the aforementioned embodiment, an embodiment of the present application provides a motherboard including the aforementioned redundant array of independent disks RAID card. The motherboard is electrically connected to multiple hard disks so that the redundant array of independent disks RAID card manages the multiple hard disks.
[0128] Based on the same concept as the aforementioned embodiment, an embodiment of the present application provides a server, including the above-mentioned Redundant Array of Independent Disks RAID card or the above-mentioned mainboard, and multiple hard disks, and the Redundant Array of Independent Disks RAID card is used to manage multiple hard disks.
[0129] Based on the same concept as the aforementioned embodiments, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, any of the above methods or the algorithmic function embodied by any of the above devices is implemented.
[0130] Based on the same concept as the aforementioned embodiments, an embodiment of the present application provides a computer program product, which includes program instructions. When the program instructions are executed by a computer, the computer executes any of the above methods or the algorithmic functions embodied by any of the above devices.
[0131] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0132] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0133] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A hard disk health assessment method, characterized in that: include: Determining first evaluation information based on multiple health indicators of the target hard disk; determining second evaluation information based on the error code information of the target hard disk; Determining third evaluation information based on the latency information of the target hard disk; The health status of the target hard disk is evaluated according to the first evaluation information, the second evaluation information, and the third evaluation information.
2. The method according to claim 1, characterized in that The determining of the first evaluation information based on the multiple health indicators of the target hard disk includes: Obtain indicator data corresponding to multiple health indicators of the target hard disk; Weighting the indicator data according to a preset weight value to obtain a first evaluation score of the target hard disk; The first evaluation score is compared with a preset evaluation score to determine first evaluation information; wherein, when the first evaluation score is lower than the preset evaluation score, the first evaluation information indicates that the target hard disk is a faulty disk.
3. The method according to claim 2, characterized in that The preset weights include a first weight and a second weight, wherein the first weight represents the weight of the first-level health indicator, and the second weight represents the weight of the second-level health indicator. The first-level health indicator represents the category of multiple health indicators of the target hard disk, and the second-level health indicator represents multiple health indicators under the first-level health indicator.
4. The method according to any one of claims 1 to 3, characterized in that The determining of the second evaluation information based on the error code information of the target hard disk includes: The second evaluation information is determined by judging the number of error codes within a preset time range; wherein, within the preset time range, when the number of error codes exceeds a quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk.
5. The method according to claim 4, characterized in that The determining of the second evaluation information by judging the number of error codes within a preset time range includes: Determine a target error code, and record the time length and number of error codes starting from the occurrence of the target error code; When the recorded time length reaches a preset time range and the number of the recorded error codes exceeds a quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk; When the recorded time length reaches the preset time range and the number of error codes recorded does not exceed the quantity threshold, the second evaluation information indicates that the target hard disk is a normal disk, and the recorded time length and the number of error codes are cleared, and the next error code is determined as the target error code.
6. The method according to any one of claims 1 to 5, characterized in that The determining of the third evaluation information based on the latency information of the target hard disk includes: Obtaining the delay time length of the target hard disk; determining a slow disk level of the target hard disk according to a relationship between the delay time length and a plurality of delay thresholds; According to the slow disk level of the target hard disk, third evaluation information of the target hard disk is determined.
7. The method according to claim 6, characterized in that The determining, based on the relationship between the delay time length and the plurality of delay thresholds, of the slow disk level of the target hard disk includes: When the delay time length is less than a first delay threshold, determining that the target hard disk is a normal disk; When the delay time is greater than or equal to the first delay threshold and less than the second delay threshold, the target hard disk is determined to be a mildly slow disk; When the delay time is greater than or equal to the second delay threshold and less than the third delay threshold, the target hard disk is determined to be a moderately slow disk; When the delay time length is greater than or equal to the third delay threshold and less than a fourth delay threshold, it is determined that the target hard disk is a severely slow disk.
8. A hard disk management method, characterized in that: include: The method according to any one of claims 1 to 7, determining the health status of multiple hard disks; The plurality of hard disks are managed according to health status of the plurality of hard disks.
9. The method according to claim 8, characterized in that Managing the multiple hard disks according to the health status of the multiple hard disks includes: When a faulty disk occurs among the multiple hard disks, a management strategy for the multiple hard disks is determined, where the management strategy includes at least one of a hard disk removal strategy, a cache modification strategy, a hard disk isolation strategy, and a degraded read / write strategy for the remaining hard disks.
10. A hard disk health assessment device, characterized in that: include: an evaluation information determining module, configured to determine first evaluation information based on a plurality of health indicators of the target hard disk; and determining second evaluation information based on the error code information of the target hard disk; and determining third evaluation information based on the latency information of the target hard disk; A health status evaluation module is configured to evaluate the health status of the target hard disk according to the first evaluation information, the second evaluation information, and the third evaluation information.
11. The device according to claim 10, characterized in that The evaluation information determination module is specifically configured to: Obtain indicator data corresponding to multiple health indicators of the target hard disk; Weighting the indicator data according to a preset weight value to obtain a first evaluation score of the target hard disk; The first evaluation score is compared with a preset evaluation score to determine first evaluation information; wherein, when the first evaluation score is lower than the preset evaluation score, the first evaluation information indicates that the target hard disk is a faulty disk.
12. The device according to claim 11, characterized in that The preset weights include a first weight and a second weight, wherein the first weight represents the weight of the first-level health indicator, and the second weight represents the weight of the second-level health indicator. The first-level health indicator represents the category of multiple health indicators of the target hard disk, and the second-level health indicator represents multiple health indicators under the first-level health indicator.
13. The device according to any one of claims 10 to 12, characterized in that: The evaluation information determination module is specifically configured to: The second evaluation information is determined by judging the number of error codes within a preset time range; wherein, within the preset time range, when the number of error codes exceeds a quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk.
14. The device according to claim 13, characterized in that The evaluation information determination module is specifically configured to: Determine a target error code, and record the time length and number of error codes starting from the occurrence of the target error code; When the recorded time length reaches a preset time range and the number of the recorded error codes exceeds a quantity threshold, the second evaluation information indicates that the target hard disk is a faulty disk; When the recorded time length reaches the preset time range and the number of error codes recorded does not exceed the quantity threshold, the second evaluation information indicates that the target hard disk is a normal disk, and the recorded time length and the number of error codes are cleared, and the next error code is determined as the target error code.
15. The device according to any one of claims 10 to 14, characterized in that: The evaluation information determination module is specifically configured to: Obtaining the delay time length of the target hard disk; determining a slow disk level of the target hard disk according to a relationship between the delay time length and a plurality of delay thresholds; According to the slow disk level of the target hard disk, third evaluation information of the target hard disk is determined.
16. The device according to claim 15, characterized in that The evaluation information determination module is specifically configured to: When the delay time length is less than a first delay threshold, determining that the target hard disk is a normal disk; When the delay time is greater than or equal to the first delay threshold and less than the second delay threshold, the target hard disk is determined to be a mildly slow disk; When the delay time is greater than or equal to the second delay threshold and less than the third delay threshold, the target hard disk is determined to be a moderately slow disk; When the delay time length is greater than or equal to the third delay threshold and less than a fourth delay threshold, it is determined that the target hard disk is a severely slow disk.
17. A hard disk management device, characterized in that: include: a determination module, configured to determine the health status of a plurality of hard disks from the health status assessment module in the apparatus according to any one of claims 10 to 16; The management module is used to manage the multiple hard disks according to the health status of the multiple hard disks.
18. The device according to claim 17, characterized in that The management module is specifically used to: When a faulty disk occurs among the multiple hard disks, a management strategy for the multiple hard disks is determined, where the management strategy includes at least one of a hard disk removal strategy, a cache modification strategy, a hard disk isolation strategy, and a degraded read / write strategy for the remaining hard disks.
19. A redundant array of independent disks (RAID) card, characterized in that: It includes a processing circuit and a storage medium, wherein the storage medium stores computer program code; when the computer program code is executed by the processing circuit, it implements the algorithm function embodied by the method according to any one of claims 1 to 9 or the device according to any one of claims 10 to 18.
20. A motherboard, characterized in that: The RAID card as claimed in claim 19 is included, wherein the mainboard is electrically connected to a plurality of hard disks so that the RAID card manages the plurality of hard disks.
21. A server, characterized in that: The method comprises the Redundant Array of Independent Disks (RAID) card as claimed in claim 19 or the mainboard as claimed in claim 20, and a plurality of hard disks, wherein the Redundant Array of Independent Disks (RAID) card is used to manage the plurality of hard disks.
22. A computing device, characterized in that The device comprises a processor and a memory, wherein a program is stored in the memory, and when the processor executes the program, the method according to any one of claims 1 to 9 or the algorithm function embodied by the device according to any one of claims 10 to 18 can be realized.
23. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the algorithm function embodied by the method according to any one of claims 1 to 9 or the apparatus according to any one of claims 10 to 18 is implemented.
24. A computer program product, characterized in that The computer program product includes program instructions, and when the program instructions are executed by a computer, the computer executes the method according to any one of claims 1 to 9 or the algorithm function embodied by the apparatus according to any one of claims 10 to 18.
Citation Information
Cited By
Server hard disk replacement method and system
CN121050959A