A method and system for locating a faulty hard disk
By analyzing the change value of the hard disk index and execution time in the RAID array card log, and judging the hard disk failure, the problem of difficulty in locating the server hard disk failure is solved, and the timely detection and location of hard disk failures is achieved.
Patent Information
- Application Number
- CN202110800194.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-15
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2041-07-15
AI Technical Summary
How to detect and locate server hard disk failures in time to ensure the stable operation of the server.
By periodically collecting and parsing the array card logs corresponding to the independent disk redundant array (RAID), obtaining the preset index change values and execution time of the hard disk before and after the patrol (PR) and consistency detection (CC), and determining whether the hard disk meets the preset fault conditions.
Accurate and timely positioning of the faulty hard disk is achieved to ensure the stable operation of the server cluster.
Smart Images

Figure CN113409876B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hard disk fault positioning, and in particular to a method and system for positioning a faulty hard disk. Background Art
[0002] With the development of computer technology, servers have higher and higher computing and storage demands for massive data. As the core component of server storage and computing, the stable operation of hard disks is an important factor in ensuring that servers provide stable services. Therefore, how to promptly determine hard disk failures and locate the failed hard disks in a timely manner is a problem that needs to be solved urgently. Summary of the invention
[0003] In view of this, an embodiment of the present invention provides a method and system for locating a faulty hard disk, so as to timely discover and locate the faulty hard disk.
[0004] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0005] A first aspect of an embodiment of the present invention discloses a method for locating a faulty hard disk, the method comprising:
[0006] Periodically collect the array card logs to be processed corresponding to the redundant array of independent disks RAID to be detected, wherein the RAID to be detected is any RAID in any node server of the server cluster;
[0007] Parse the array card log to be processed, obtain a first change value of a first preset indicator of each hard disk before and after a patrol read PR is performed on the RAID to be detected, and obtain a first execution time required for performing PR on the RAID to be detected, and obtain a second change value of the first preset indicator of each hard disk before and after a consistency check CC is performed on the RAID to be detected, and obtain a second execution time required for performing CC on the RAID to be detected;
[0008] Determine whether there is a faulty hard disk in the RAID to be detected according to the first execution time, the second execution time, the first change value and the second change value corresponding to each hard disk in the RAID to be detected;
[0009] If so, obtain the hard disk information corresponding to the faulty hard disk.
[0010] Preferably, determining whether there is a faulty hard disk in the RAID to be detected according to the first execution time, the second execution time, the first change value and the second change value corresponding to each hard disk in the RAID to be detected includes:
[0011] For each hard disk in the RAID to be detected, judging whether the hard disk meets a preset fault condition according to the first execution time and the second execution time, in combination with the first change value and the second change value corresponding to the hard disk, and if so, determining that the hard disk is a faulty hard disk;
[0012] Among them, the preset fault conditions are: the first change value is greater than or equal to a first threshold, the second change value is greater than or equal to a second threshold, the first execution time is greater than or equal to a third threshold, and the second execution time is greater than or equal to a fourth threshold.
[0013] Preferably, the first preset indicator includes at least: a medium error counter, an expected error counter, other error counters and a hardware error counter.
[0014] Preferably, before the periodic collection of the to-be-processed array card logs corresponding to the to-be-detected redundant array of independent disks RAID, the method further comprises:
[0015] The PR and CC are executed for the RAID to be detected according to a preset execution time and execution cycle, wherein the execution time and the execution cycle are determined based on a second preset indicator and preset information corresponding to the node server to which the RAID to be detected belongs.
[0016] Preferably, the second preset indicator includes at least: CPU usage, memory usage, CUP waiting for IO, total network card traffic per second, swap memory swap utilization, disk busyness and disk IO throughput.
[0017] Preferably, after obtaining the hard disk information corresponding to the faulty hard disk, the method further includes:
[0018] Obtaining application system information and busy / idle time period information of an application system associated with the node server to which the faulty hard disk belongs, and obtaining array card performance information corresponding to the RAID to be detected;
[0019] Formulate a disk replacement strategy based on the application system information, the busy and idle time period information, and the array card performance information in combination with disk replacement rules;
[0020] The disk replacement strategy and the alarm notification are sent to a designated object, wherein the alarm notification at least includes the hard disk information corresponding to the failed hard disk.
[0021] A second aspect of an embodiment of the present invention discloses a system for locating a faulty hard disk, the system comprising:
[0022] A collection unit, used for periodically collecting array card logs to be processed corresponding to a redundant array of independent disks RAID to be detected, wherein the RAID to be detected is any RAID in any node server of the server cluster;
[0023] The parsing unit is used to parse the array card log to be processed, obtain a first change value of a first preset index of each hard disk before and after the patrol read PR is performed on the RAID to be detected, and obtain a first execution time required when the PR is performed on the RAID to be detected, and obtain a second change value of the first preset index of each hard disk before and after the consistency check CC is performed on the RAID to be detected, and obtain a second execution time required when the CC is performed on the RAID to be detected;
[0024] A processing unit is used to determine whether there is a faulty hard disk in the RAID to be detected based on the first execution time, the second execution time, the first change value and the second change value corresponding to each hard disk in the RAID to be detected; if so, obtain the hard disk information corresponding to the faulty hard disk.
[0025] Preferably, the processing unit for determining whether there is a faulty hard disk in the RAID to be detected is specifically used for:
[0026] For each hard disk in the RAID to be detected, judging whether the hard disk meets a preset fault condition according to the first execution time and the second execution time, in combination with the first change value and the second change value corresponding to the hard disk, and if so, determining that the hard disk is a faulty hard disk;
[0027] Among them, the preset fault conditions are: the first change value is greater than or equal to a first threshold, the second change value is greater than or equal to a second threshold, the first execution time is greater than or equal to a third threshold, and the second execution time is greater than or equal to a fourth threshold.
[0028] Preferably, the first preset indicator includes at least: a medium error counter, an expected error counter, other error counters and a hardware error counter.
[0029] Preferably, the system further comprises:
[0030] The execution unit is used to execute PR and CC for the RAID to be detected according to a preset execution time and execution cycle, wherein the execution time and the execution cycle are determined based on a second preset indicator and preset information corresponding to the node server to which the RAID to be detected belongs.
[0031] Based on the above-mentioned embodiment of the present invention, a method and system for locating a faulty hard disk is provided. The method comprises: periodically collecting the log of the array card to be processed corresponding to the RAID to be detected; parsing the log of the array card to be processed, obtaining the first change value of the first preset index of each hard disk before and after the PR is executed on the RAID to be detected, and obtaining the first execution time required when the PR is executed on the RAID to be detected, and obtaining the second change value of the first preset index of each hard disk before and after the CC is executed on the RAID to be detected, and obtaining the second execution time required when the CC is executed on the RAID to be detected; according to the first change value, the first execution time, the second change value and the second execution time corresponding to each hard disk in the RAID to be detected, determining whether there is a faulty hard disk in the RAID to be detected; if there is, obtaining the hard disk information corresponding to the faulty hard disk. According to the change value of the preset index of the hard disk before and after the PR and CC are executed, combined with the execution time corresponding to the hard disk when the PR and CC are executed, the faulty hard disk is determined to accurately and timely locate the faulty hard disk. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0033] Figure 1 A flowchart of a method for locating a faulty hard disk provided by an embodiment of the present invention;
[0034] Figure 2 Another flow chart of a method for locating a faulty hard disk provided by an embodiment of the present invention;
[0035] Figure 3 A structural block diagram of a faulty hard disk positioning system provided by an embodiment of the present invention;
[0036] Figure 4 Another structural block diagram of a system for locating a faulty hard disk provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0039] As can be seen from the background technology, in order to ensure that the server can provide stable services, it is necessary to promptly identify and locate the failed hard disk. Therefore, how to locate the failed hard disk is a problem that needs to be solved urgently.
[0040] Therefore, the embodiment of the present invention provides a method and system for locating a faulty hard disk, which obtains the change value of the preset index of the hard disk before and after executing PR and CC on the RAID by parsing the array card log, and obtains the execution time required when executing PR and CC on the RAID. Combined with the change value and execution time of the preset index corresponding to the hard disk, it is determined whether the hard disk is faulty, so as to accurately and timely locate the faulty hard disk.
[0041] It should be noted that the contents shown in the embodiments of the present invention may involve multiple English abbreviations. The following content will explain the English abbreviations involved in the embodiments of the present invention in advance.
[0042] RAID: Redundant array of independent disks.
[0043] RAID card: array card, a card used to implement RAID mode.
[0044] PR: Patrol Read.
[0045] CC: Consistency Check, that is, consistency check.
[0046] CMDB: Configuration Management Database, that is, configuration management database.
[0047] See also Figure 1 , shows a flow chart of a method for locating a faulty hard disk provided by an embodiment of the present invention, the locating method comprising:
[0048] Step S101: Periodically collect the array card logs to be processed corresponding to the RAID to be detected.
[0049] It should be noted that a server cluster is composed of multiple node servers, each node server contains multiple RAIDs, and each RAID is composed of multiple hard disks. The RAIDs contained in the node servers can be distinguished through the array card log (specifically distinguished by the array number). The array card log also contains at least the hard disk slot number, which can be used to distinguish different hard disks in the RAID.
[0050] It can be understood that the RAID to be detected is any RAID in any node server of the server cluster.
[0051] In the specific implementation of step S101, the array card logs to be processed corresponding to the RAID to be detected are periodically collected.
[0052] Preferably, before collecting the array card logs to be processed corresponding to the RAID to be detected, it is necessary to use the RAID card to execute PR and CC for the RAID to be detected. It is understandable that some computing resources will be occupied when executing PR and CC. In order not to affect the normal operation of the node server, it is usually necessary to formulate the time and cycle for executing PR and CC according to the busyness of the node server (equivalent to formulating a scheduled task). In some embodiments, according to the preset execution time and execution cycle, PR and CC are executed for the RAID to be detected. The execution time and execution cycle are determined based on the second preset indicator and preset information corresponding to the node server to which the RAID to be detected belongs. On this basis, according to the execution time and execution cycle for executing PR and CC, the array card logs to be processed are periodically collected. The array card logs to be processed are the array card logs before and after executing PR and CC.
[0053] It should be noted that the execution time and execution cycle for executing PR and CC can also be adjusted according to actual needs or the failure rate of the RAID to be detected, and the formulation of the execution time and execution cycle is not limited to the above-mentioned method.
[0054] In some specific embodiments, the second preset index at least includes: central processing unit (CPU) utilization, memory utilization, CPU waiting for IO (i.e., input and output), total network card traffic per second, swap memory utilization, disk busyness, and disk IO throughput. The preset information corresponding to the node server to which the RAID to be detected belongs at least includes: application system information corresponding to the application system associated with the node server (such as application system name, importance level information, service level agreement, application manager and application manager contact information, etc.).
[0055] It should be noted that the second preset index mentioned above can be obtained by collecting information of various index items corresponding to the node server to which the RAID to be detected belongs. The index items used to collect the second preset index are shown in Table 1.
[0056] Table 1:
[0057] Indicator Explanation of indicators CpuUtil CPU usage SedMemPerccent Memory usage IOwait CPU Waiting for IO NET_RATE Total network traffic per second SwapUsedPercent Swap utilization DISKPercentBusy Disk busyness DISKIORate Disk IO throughput
[0058] It should be noted that the indicator items for collecting the second preset indicators shown in the above Table 1 are only for illustration. In actual applications, the indicator items for collecting the second preset indicators can be determined according to actual needs. For example, the indicator items for collecting the second preset indicators can be increased or decreased based on Table 1, or other indicator items for collecting the second preset indicators can be selected according to actual needs. No specific limitation is made here.
[0059] Step S102: parsing the array card log to be processed, obtaining a first change value of a first preset indicator of each hard disk before and after PR is performed on the RAID to be detected, and obtaining a first execution time required when PR is performed on the RAID to be detected, and obtaining a second change value of the first preset indicator of each hard disk before and after CC is performed on the RAID to be detected, and obtaining a second execution time required when CC is performed on the RAID to be detected.
[0060] It should be noted that after the RAID card is used to execute PR on the RAID to be detected, the first preset index of each hard disk in the RAID to be detected will change accordingly. Similarly, after the RAID card is used to execute CC on the RAID to be detected, the first preset index of each hard disk in the RAID to be detected will also change accordingly. Before and after executing PR and CC, the value of the first preset index of each hard disk will be recorded in the array card log corresponding to the RAID to be detected. The value of the first preset index of each hard disk before executing PR and CC can be obtained by collecting the array card log before and after executing PR and CC, and the value of the first preset index of each hard disk after executing PR and CC can be obtained.
[0061] It should be further explained that PR and CC are used to repair hard disk errors or to repair inconsistent data. If the execution time of PR or CC is too long, it means that there are many hard disk errors and the hard disk reading and writing capabilities are poor. Therefore, the time required to execute PR and the time required to execute CC can be used as one of the bases for judging whether the hard disk has failed. The time required to execute PR and the time required to execute CC can be obtained from the array card log. Specifically, the start time of executing PR, the end time of executing PR, the start time of executing CC, and the end time of executing CC can be obtained from the array card log, and the time required to execute PR is determined according to the start time and end time of executing PR, and the time required to execute CC is determined according to the start time and end time of executing CC.
[0062] In the process of specifically implementing step S102, the array card log to be processed is parsed to obtain the value of the first preset index of each hard disk (each hard disk in the RAID to be detected) before the PR is executed on the RAID to be detected, and the value of the first preset index of each hard disk after the PR is executed on the RAID to be detected is obtained. By comparing the value of the first preset index of each hard disk before the PR is executed on the RAID to be detected and the value of the first preset index of each hard disk after the PR is executed on the RAID to be detected, a first change value of the first preset index of each hard disk before and after the PR is executed on the RAID to be detected can be determined; after parsing the array card log to be processed, the first execution time required for the PR to be executed on the RAID to be detected (determined according to the start time and the end time of the PR execution) can also be obtained.
[0063] Similarly, the array card log to be processed is parsed to obtain the value of the first preset indicator of each hard disk before CC is executed on the RAID to be detected, and the value of the first preset indicator of each hard disk after CC is executed on the RAID to be detected. The second change value of the first preset indicator of each hard disk before CC is executed on the RAID to be detected and after CC is executed on the RAID to be detected can be determined by the value of the first preset indicator of each hard disk before CC is executed on the RAID to be detected and the value of the first preset indicator of each hard disk after CC is executed on the RAID to be detected. After parsing the array card log to be processed, the second execution time required for CC to be executed on the RAID to be detected can also be obtained (determined according to the start time and end time of CC execution).
[0064] In some specific embodiments, the first preset indicator includes at least: a medium error counter, an expected error counter, other error counters, and a hardware error counter.
[0065] In some embodiments, when obtaining the value of the first preset indicator of each hard disk before and after executing PR and before and after executing CC from the array card log to be processed, it can be obtained by collecting information of each indicator item corresponding to the node server to which the RAID to be detected belongs. The indicator items used to collect the value of the first preset indicator are detailed in the contents shown in Table 2.
[0066] Table 2:
[0067] Indicator Explanation of indicators Slot Number Hard disk slot number Media Error Count Media Error Counters Predictive FailureCount Expected Error Counters Other Error Count Other Error Counters Hardware Error Count Hardware Error Counters
[0068] It can be understood that the hard disk slot number indicator item in Table 2 is used to indicate the corresponding hard disk.
[0069] It should be noted that the indicator items for collecting the value of the first preset indicator shown in the above Table 2 are only for illustrative purposes. In actual applications, the indicator items for collecting the value of the first preset indicator can be determined according to actual needs. For example, the indicator items for collecting the value of the first preset indicator can be increased or decreased based on Table 2, and other indicator items for collecting the value of the first preset indicator can be selected according to actual needs. No specific limitation is made here.
[0070] Combining the above content, it can be known that the first execution duration is determined according to the start time of executing PR and the end time of executing PR for the RAID to be detected, and the second execution duration is determined according to the start time of executing CC and the end time of executing CC for the RAID to be detected. In a specific implementation, the start time of executing PR, the end time of executing PR, the start time of executing CC, and the end time of executing CC can be obtained by matching the keywords from the array card log to be processed. The specific content of the keyword is shown in Table 3.
[0071] Keywords Keyword Definition Patrol Readstarted PR start execution time Consistency Check started CC start execution time Patrol Readcompleted PR end time Consistency Checkdone CC end time
[0072] Step S103: Determine whether there is a faulty hard disk in the RAID to be detected based on the first execution time, the second execution time, the first change value and the second change value corresponding to each hard disk in the RAID to be detected. If yes, execute step S104; if no, return to execute step S101.
[0073] In the process of implementing step S103, for each hard disk in the RAID to be detected, according to the first execution time and the second execution time, combined with the first change value and the second change value of the first preset indicator corresponding to the hard disk, it is determined whether the hard disk meets the preset fault condition. If it does, the hard disk is determined to be a faulty hard disk. Among them, the preset fault conditions are: the first change value is greater than or equal to the first threshold, the second change value is greater than or equal to the second threshold, the first execution time is greater than or equal to the third threshold, and the second execution time is greater than or equal to the fourth threshold. If there is no faulty hard disk in the RAID to be detected, return to step S101, continue to collect new array card logs to be processed, and continue to monitor the hard disks of the RAID to be detected.
[0074] As can be seen from the above content, there are multiple first preset indicators of the hard disk. In the specific implementation, for each hard disk in the RAID to be detected, when the first execution time, the second execution time, the first change value of any one or multiple first preset indicators of the hard disk and the second change indicator value meet the following fault conditions, the hard disk is determined to be a faulty hard disk. The fault conditions are: the first change value is greater than or equal to the first threshold, the second change value is greater than or equal to the second threshold, the first execution time is greater than or equal to the third threshold, and the second execution time is greater than or equal to the fourth threshold. The aforementioned thresholds can be adjusted according to actual conditions.
[0075] Step S104: Obtain hard disk information corresponding to the faulty hard disk.
[0076] In the specific implementation of step S104, if a faulty hard disk is determined from the RAID to be detected, hard disk information of the faulty hard disk is obtained, and the hard disk information indicates which node server, which RAID and which hard disk the faulty hard disk belongs to.
[0077] Preferably, after determining that there is a faulty hard disk in the RAID to be detected and obtaining hard disk information corresponding to the faulty hard disk, the application system information and busy / idle time period information of the application system associated with the node server to which the faulty hard disk belongs are obtained from the CMBD, and the array card performance information corresponding to the RAID to be detected (that is, the array card performance information of the array card used to execute PR and CC for the RAID to be detected) is obtained, the application system information at least includes: the importance level of the application system, the service level grade agreement, the application manager and the contact information of the application manager, and the array card performance information is used to determine the busy / idle time period of the hard disk; according to the application system information, the busy / idle time period information and the array card performance information, in combination with the disk replacement rule, a disk replacement strategy is formulated; the disk replacement strategy and the alarm notification are sent to a designated object (such as an operation and maintenance personnel), so that the designated object can understand the hard disk information of the faulty hard disk according to the alarm notification, and the designated object can process the faulty hard disk according to the disk replacement strategy, and the alarm notification at least includes the hard disk information corresponding to the faulty hard disk.
[0078] Through the contents provided in the above steps S101 to S104, the faulty hard disk is located in each RAID of each node server in the server cluster.
[0079] In the embodiment of the present invention, the change value of the preset index of the hard disk before and after the PR and CC are executed on the RAID is obtained by parsing the array card log, and the execution time required when the PR and CC are executed on the RAID is obtained. Combined with the change value and execution time of the preset index corresponding to the hard disk, it is determined whether the hard disk is faulty, so as to accurately and timely locate the faulty hard disk, thereby ensuring the stable operation of the server cluster.
[0080] To better explain the above embodiments of the present invention Figure 1 The content shown is Figure 2 Give an example.
[0081] See also Figure 2 , shows another flow chart of a method for locating a faulty hard disk provided by an embodiment of the present invention, comprising the following steps:
[0082] Step S201: using CMDB to obtain preset information corresponding to the node server to which the RAID to be detected belongs.
[0083] It should be noted that the preset information corresponding to the node server to which the RAID to be tested belongs includes at least: application system information corresponding to the application system associated with the node server, such as: application system name, importance level information, service level grade agreement, application manager and application manager's contact information and other information.
[0084] Step S202: Obtain a second preset indicator corresponding to the node server to which the RAID to be detected belongs.
[0085] It should be noted that the specific content of the second preset indicator can be found in the above Table 1.
[0086] Step S203: Based on the second preset indicator and preset information, determine the execution time and execution cycle for executing PR and CC.
[0087] Step S204: Periodically collect the array card logs to be processed corresponding to the RAID to be detected.
[0088] Step S205: parsing the array card log to be processed, obtaining the first execution time, the second execution time, the first change value of the first preset indicator of each hard disk in the RAID to be detected, and the second change value of the first preset indicator, and determining the faulty hard disk in the RAID to be detected accordingly.
[0089] Step S206: Determine whether there is a faulty hard disk in the RAID to be detected, if yes, execute step S207, if not, return to execute step S204.
[0090] Step S207: Formulate a disk replacement strategy.
[0091] Step S208: Process the failed hard disk according to the disk replacement strategy.
[0092] Corresponding to the method for locating a faulty hard disk provided in the above embodiment of the present invention, see Figure 3 , the embodiment of the present invention also provides a structural block diagram of a positioning system for a faulty hard disk, the positioning system includes: a collection unit 301, a parsing unit 302 and a processing unit 303;
[0093] The collection unit 301 is used to periodically collect the array card logs to be processed corresponding to the RAID to be detected, and the RAID to be detected is any RAID in any node server of the server cluster.
[0094] The parsing unit 302 is used to parse the array card log to be processed, obtain a first change value of a first preset indicator of each hard disk before and after PR is performed on the RAID to be detected, and obtain a first execution time required when PR is performed on the RAID to be detected, and obtain a second change value of the first preset indicator of each hard disk before and after CC is performed on the RAID to be detected, and obtain a second execution time required when CC is performed on the RAID to be detected.
[0095] In a specific implementation, the first preset indicator includes at least: a medium error counter, an expected error counter, other error counters and a hardware error counter.
[0096] The processing unit 303 is used to determine whether there is a faulty hard disk in the RAID to be detected based on the first execution time, the second execution time, the first change value and the second change value corresponding to each hard disk in the RAID to be detected; if so, obtain the hard disk information corresponding to the faulty hard disk.
[0097] In a specific implementation, the processing unit 303 for determining whether there is a faulty hard disk in the RAID to be detected is specifically used to: for each hard disk in the RAID to be detected, based on the first execution time and the second execution time, combined with the first change value and the second change value corresponding to the hard disk, determine whether the hard disk meets the preset fault condition; if so, determine that the hard disk is a faulty hard disk; wherein the preset fault conditions are: the first change value is greater than or equal to the first threshold, the second change value is greater than or equal to the second threshold, the first execution time is greater than or equal to the third threshold, and the second execution time is greater than or equal to the fourth threshold.
[0098] Preferably, the processing unit 303 is further used to: obtain application system information and busy / idle time period information of an application system associated with a node server to which the faulty hard disk belongs, and obtain array card performance information corresponding to the RAID to be detected; formulate a disk replacement strategy based on the application system information, busy / idle time period information and array card performance information, combined with disk replacement rules; send the disk replacement strategy and an alarm notification to a designated object, wherein the alarm notification at least includes the hard disk information corresponding to the faulty hard disk.
[0099] In the embodiment of the present invention, the change value of the preset index of the hard disk before and after the PR and CC are executed on the RAID is obtained by parsing the array card log, and the execution time required when the PR and CC are executed on the RAID is obtained. Combined with the change value and execution time of the preset index corresponding to the hard disk, it is determined whether the hard disk is faulty, so as to accurately and timely locate the faulty hard disk, thereby ensuring the stable operation of the server cluster.
[0100] Preferably, combined Figure 3 , see Figure 4 , shows a structural block diagram of a faulty hard disk positioning system provided by an embodiment of the present invention, the positioning system also includes:
[0101] The execution unit 304 is used to execute PR and CC on the RAID to be detected according to a preset execution time and execution cycle, and the execution time and execution cycle are determined based on the second preset indicator and preset information corresponding to the node server to which the RAID to be detected belongs.
[0102] In a specific implementation, the second preset indicator includes at least: CPU usage, memory usage, CPU waiting for IO, total network card traffic per second, swap utilization, disk busyness and disk IO throughput.
[0103] In summary, the embodiment of the present invention provides a method and system for locating a faulty hard disk, which obtains the change value of the preset index of the hard disk before and after executing PR and CC on the RAID by parsing the array card log, and obtains the execution time required when executing PR and CC on the RAID. Combined with the change value and execution time of the preset index corresponding to the hard disk, it is determined whether the hard disk has a fault, so as to accurately and timely locate the faulty hard disk.
[0104] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0105] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0106] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for locating a faulty hard disk, characterized in that: The method comprises: Periodically collect the array card logs to be processed corresponding to the redundant array of independent disks RAID to be detected, wherein the RAID to be detected is any RAID in any node server of the server cluster; Parse the array card log to be processed, obtain a first change value of a first preset indicator of each hard disk before and after the patrol read PR is performed on the RAID to be detected, and obtain a first execution time required for performing PR on the RAID to be detected, and obtain a second change value of the first preset indicator of each hard disk before and after the consistency check CC is performed on the RAID to be detected, and obtain a second execution time required for performing CC on the RAID to be detected; the first preset indicator at least includes: a medium error counter, an expected error counter, and a hardware error counter; Determine whether there is a faulty hard disk in the RAID to be detected according to the first execution time, the second execution time, the first change value and the second change value corresponding to each hard disk in the RAID to be detected; If so, obtain the hard disk information corresponding to the faulty hard disk.
2. The method according to claim 1, characterized in that Determining whether there is a faulty hard disk in the RAID to be detected according to the first execution time, the second execution time, the first change value and the second change value corresponding to each hard disk in the RAID to be detected, includes: For each hard disk in the RAID to be detected, judging whether the hard disk meets a preset fault condition according to the first execution time and the second execution time, in combination with the first change value and the second change value corresponding to the hard disk, and if so, determining that the hard disk is a faulty hard disk; Among them, the preset fault conditions are: the first change value is greater than or equal to a first threshold, the second change value is greater than or equal to a second threshold, the first execution time is greater than or equal to a third threshold, and the second execution time is greater than or equal to a fourth threshold.
3. The method according to claim 1, characterized in that The first preset indicator also includes: other error counters.
4. The method according to claim 1, characterized in that Before the periodic collection of the array card logs to be processed corresponding to the redundant array of independent disks RAID to be detected, the method further includes: The PR and CC are executed for the RAID to be detected according to a preset execution time and execution cycle, wherein the execution time and the execution cycle are determined based on a second preset indicator and preset information corresponding to the node server to which the RAID to be detected belongs.
5. The method according to claim 4, characterized in that The second preset indicator at least includes: CPU usage, memory usage, CPU waiting for IO, total network card traffic per second, swap memory swap utilization, disk busyness and disk IO throughput.
6. The method according to claim 1, characterized in that After obtaining the hard disk information corresponding to the faulty hard disk, the method further includes: Obtaining application system information and busy / idle time period information of an application system associated with the node server to which the faulty hard disk belongs, and obtaining array card performance information corresponding to the RAID to be detected; Formulate a disk replacement strategy based on the application system information, the busy and idle time period information, and the array card performance information in combination with disk replacement rules; The disk replacement strategy and the alarm notification are sent to a designated object, wherein the alarm notification at least includes the hard disk information corresponding to the failed hard disk.
7. A system for locating a faulty hard disk, characterized in that: The system comprises: A collection unit, used for periodically collecting array card logs to be processed corresponding to a redundant array of independent disks RAID to be detected, wherein the RAID to be detected is any RAID in any node server of the server cluster; The parsing unit is used to parse the array card log to be processed, obtain a first change value of a first preset indicator of each hard disk before and after the patrol read PR is performed on the RAID to be detected, and obtain a first execution time required for performing PR on the RAID to be detected, and obtain a second change value of the first preset indicator of each hard disk before and after the consistency check CC is performed on the RAID to be detected, and obtain a second execution time required for performing CC on the RAID to be detected; the first preset indicator at least includes: a medium error counter, an expected error counter, and a hardware error counter; A processing unit is used to determine whether there is a faulty hard disk in the RAID to be detected based on the first execution time, the second execution time, the first change value and the second change value corresponding to each hard disk in the RAID to be detected; if so, obtain the hard disk information corresponding to the faulty hard disk.
8. The system according to claim 7, characterized in that The processing unit for determining whether there is a faulty hard disk in the RAID to be detected is specifically used to: For each hard disk in the RAID to be detected, judging whether the hard disk meets a preset fault condition according to the first execution time and the second execution time, in combination with the first change value and the second change value corresponding to the hard disk, and if so, determining that the hard disk is a faulty hard disk; Among them, the preset fault conditions are: the first change value is greater than or equal to a first threshold, the second change value is greater than or equal to a second threshold, the first execution time is greater than or equal to a third threshold, and the second execution time is greater than or equal to a fourth threshold.
9. The system according to claim 7, characterized in that The first preset indicator also includes: other error counters.
10. The system according to claim 7, characterized in that The system further comprises: The execution unit is used to execute PR and CC for the RAID to be detected according to a preset execution time and execution cycle, wherein the execution time and the execution cycle are determined based on a second preset indicator and preset information corresponding to the node server to which the RAID to be detected belongs.
Citation Information
Patent Citations
Disk alarm method and device
CN112084097A
Digital set-top terminal with partitioned hard disk and associated system and method
US20090193486A1