A method, system, device, medium for hard disk auto-reset and self-repair

CN116069559BActive Publication Date: 2026-09-11INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211624995.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2026-09-11
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

[0006]本申请的目的是提供一种硬盘自动复位和自我修复的方法、系统、装置、介质,用于解决故障硬盘的运维成本高的问题

Benefits of technology

[0040]本申请所提供的硬盘自动复位和自我修复的方法,当获取到被监控硬盘的掉线信息后,通过控制被监控硬盘复位进行初步检测,依据SMART健康检测做进一步检测,再对被监控硬盘进行故障问题复判,最后根据故障问题复判的结果对被监控硬盘进行修复,使修复后的被监控硬盘能够重新进行上线使用,无需人工进行硬盘更换,节约了硬盘备件资源,提升了硬盘使用效率,进而降低了硬盘的运维压力及故障硬盘的运维成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116069559B_ABST
    Figure CN116069559B_ABST
Patent Text Reader

Abstract

The application discloses a hard disk automatic reset and self-repairing method, system, device and medium. The hard disk automatic reset and self-repairing method provided by the application performs preliminary detection by controlling the monitored hard disk to reset after obtaining the offline information of the monitored hard disk, performs further detection according to the SMART health detection, performs fault problem rejudgment on the monitored hard disk, and finally repairs the monitored hard disk according to the result of the fault problem rejudgment, so that the repaired monitored hard disk can be used online again, manual hard disk replacement is not needed, hard disk spare part resources are saved, hard disk use efficiency is improved, and then hard disk operation and maintenance pressure and the operation and maintenance cost of the fault hard disk are reduced. The hard disk automatic reset and self-repairing system, device and medium provided by the application have the same beneficial effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of servers, and in particular to a method, system, device, or medium for automatic hard disk reset and self-repair. Background Technology

[0002] With the continuous development of server cloud storage services, the market share of mechanical hard drives is also increasing. Hard drive failures occur frequently, which brings great difficulties to hard drive maintenance and repair work, and at the same time continuously drives up the maintenance costs of failed hard drives.

[0003] Currently, when a client experiences a hard drive failure, the server system is typically taken offline while awaiting centralized system maintenance, where hard drive maintenance personnel will replace the faulty hard drive.

[0004] However, for customers with strong business fault tolerance, taking the client offline from the server system will affect the timeliness of the client's business, waiting for centralized system maintenance will waste server system resources and affect maintenance timeliness; replacing faulty hard drives will waste hard drive hardware resources.

[0005] Therefore, how to reduce the maintenance cost of faulty hard drives is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] The purpose of this application is to provide a method, system, device, or medium for automatic hard disk reset and self-repair, in order to solve the problem of high maintenance costs for faulty hard disks.

[0007] To address the aforementioned technical problems, this application provides a method for automatic hard disk reset and self-repair, comprising:

[0008] Obtain the presence status of the monitored hard drive;

[0009] If the monitored hard drive is offline, then control the monitored hard drive to reset.

[0010] Determine whether to perform a SMART health check on the monitored hard drive based on the reset status of the monitored hard drive;

[0011] Determine whether to reassess the monitored hard drive based on the SMART health check results;

[0012] The monitored hard drive is repaired based on the results of the fault assessment.

[0013] Preferably, determining whether to perform a SMART health check on the monitored hard drive based on the reset status of the monitored hard drive includes:

[0014] If the monitored hard drive comes back online, then a SMART health check will be performed on the monitored hard drive.

[0015] If a DNR error is obtained, it is determined that the monitored hard drive will not be subjected to SMART health check;

[0016] If a SMART health check is performed on the monitored hard drive, it also includes: sending a SMART test command to the monitored hard drive;

[0017] If the monitored hard drive is not subjected to SMART health checks, the following is also included: issuing warranty alerts for faulty hard drives.

[0018] Preferably, determining whether to re-evaluate the fault condition of the monitored hard drive based on the SMART health check results includes:

[0019] If the SMART health check passes, then it is determined that the monitored hard drive will not be re-evaluated for faults.

[0020] If the SMART health check fails, the monitored hard drive will be re-evaluated for faults.

[0021] If the monitored hard drive is to be re-evaluated for faults, it also includes: controlling the monitored hard drive to stop working and re-evaluating the faults of the monitored hard drive.

[0022] Preferably, repairing the monitored hard drive based on the results of fault assessment includes:

[0023] The type of fault of the monitored hard drive is determined based on the results of the fault reassessment.

[0024] Repair the monitored hard drive according to the type of fault.

[0025] Preferably, repairing the monitored hard drive according to the fault type includes:

[0026] If it is a hard drive soft error, then perform error clearing on the monitored hard drive;

[0027] If it is a hard disk error, read the address of the error reported by the monitored hard disk and isolate the disk and head corresponding to the address.

[0028] Preferably, after obtaining the presence status of the monitored hard drive, if the monitored hard drive is offline, the method before controlling the monitored hard drive to reset further includes: recording and displaying the offline status of the monitored hard drive.

[0029] Preferably, if the monitored hard drive is offline, controlling the monitored hard drive to reset includes:

[0030] Send reset commands to P3, S2 / S3, and S5 / S6 of the CPLD chip.

[0031] To address the aforementioned technical problems, this application also provides a system for automatic hard disk reset and self-repair, comprising:

[0032] The acquisition module is used to obtain the presence status of the monitored hard drive;

[0033] The control module is used to reset the monitored hard drive if the monitored hard drive is offline.

[0034] The judgment module is used to determine whether to perform a SMART health check on the monitored hard drive based on the reset status of the monitored hard drive.

[0035] The determination module is used to determine whether to re-evaluate the fault status of the monitored hard drive based on the SMART health check results.

[0036] The repair module is used to repair the monitored hard drive based on the results of the fault assessment.

[0037] To address the aforementioned technical problems, this application also provides a device for automatic hard disk reset and self-repair, including a memory for storing computer programs;

[0038] A processor is used to implement methods for automatic hard disk reset and self-repair when executing computer programs.

[0039] To address the aforementioned technical problems, this application also provides a computer-readable storage medium storing a computer program, and a method for automatically resetting and self-repairing a hard disk when the computer program is executed by a processor.

[0040] The automatic hard drive reset and self-repair method provided in this application, upon receiving information about the offline status of the monitored hard drive, performs a preliminary test by controlling the monitored hard drive to reset, conducts further tests based on SMART health checks, re-evaluates the fault status of the monitored hard drive, and finally repairs the monitored hard drive based on the results of the fault status re-evaluation. This allows the repaired monitored hard drive to be put back online for use without the need for manual hard drive replacement, saving hard drive spare parts resources, improving hard drive utilization efficiency, and thereby reducing the maintenance pressure and maintenance costs of hard drives and faulty hard drives.

[0041] The beneficial effects of the hard disk automatic reset and self-repair system, device, and media provided in this application are the same as those described above. Attached Figure Description

[0042] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 A flowchart illustrating a method for automatic hard disk reset and self-repair provided in this application embodiment;

[0044] Figure 2 A flowchart of a hard disk reset provided in an embodiment of this application;

[0045] Figure 3 This application provides a flowchart for determining whether to re-evaluate the fault status of the monitored hard drive in an embodiment of the present application;

[0046] Figure 4 A flowchart for repairing a monitored hard drive is provided as an embodiment of this application;

[0047] Figure 5 This is an application flowchart of hard disk automatic reset and self-repair provided in an embodiment of this application;

[0048] Figure 6 A schematic diagram of a hard disk automatic reset and self-repair system provided in an embodiment of this application;

[0049] Figure 7 This is a schematic diagram of a hard disk automatic reset and self-repair device provided in an embodiment of this application. Detailed Implementation

[0050] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0051] The core of this application is to provide a method, system, device, and medium for automatic hard disk reset and self-repair, applicable to the server field. With the continuous development of server cloud storage services, the market share of mechanical hard disks is increasing, and hard disk failures are occurring frequently, posing significant challenges to hard disk maintenance and continuously driving up the maintenance costs of faulty hard disks. This application provides functional designs for resetting, health monitoring, and repairing faulty hard disks, addressing the urgency of hard disk maintenance and reducing maintenance costs.

[0052] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] The hard disk automatic reset and self-repair method provided in the embodiments of this application, such as Figure 1 , Figure 1A flowchart illustrating a method for automatic hard disk reset and self-repair provided in this application embodiment, the method comprising:

[0054] S10: Obtain the presence status of the monitored hard drive.

[0055] This application embodiment is based on a Baseboard Management Controller (BMC) and uses a clock signal design to read the presence status of the monitored hard drive at a preset frequency. The presence status of the monitored hard drive includes offline and online states, and the preset frequency is typically once per minute. When the presence status of the monitored hard drive is read as online, the presence status of the monitored hard drive continues to be read at the preset frequency.

[0056] S11: If the monitored hard drive is offline, then control the monitored hard drive to reset.

[0057] After BMC detects that the monitored hard drive has gone offline, it needs to force the monitored hard drive to reset its power supply and signals. For example, Figure 2 A flowchart of a hard disk reset provided in an embodiment of this application is shown below. Figure 2 As shown:

[0058] S20: Hard drive offline connection lost;

[0059] S21: Automatic monitoring and identification of BMC;

[0060] S22: BMC sends a hard disk reset command;

[0061] S23: Hard drive performs power and signal reset;

[0062] S24: Was the hard drive reset successful?

[0063] If the hard drive reset is successful, then execute S25: Hard drive reports DNR failure;

[0064] S26: BMC issues a warranty warning;

[0065] If the hard drive reset fails, execute S27: Hard drive reconnects.

[0066] The system automatically sends reset commands (P3, S2 / S3, S5 / S6) to the Complex Programmable Logic Device (CPLD) chip on the backplane of the monitored hard drive. If the monitored hard drive fails to reset, the BMC will receive a Hard Drive Not Booting (DNR) error. In this case, a faulty disk warranty alarm can be triggered in the BMC management interface. If the monitored hard drive successfully resets, it will be brought back online.

[0067] S12: Determine whether to perform a SMART health check on the monitored hard drive based on the reset status of the monitored hard drive.

[0068] The purpose of SMART health checks is to monitor hard drive reliability, predict disk failures, and perform various types of disk self-tests. The monitored hard drive's reset status includes both failure to reset and successful reset. Based on the reset status, it is determined whether further SMART monitoring and self-testing of the monitored hard drive is necessary. Typically, the following tests are included:

[0069] 01(001) Underlying Data Read Error Rate (Raw_Read_Error_Rate)

[0070] 04(004) Start / Stop Count

[0071] 05(005) Number of remapped sectors (Reallocated_Sector_Ct)

[0072] 09(009) Power-on time cumulative (Power_On_Hours), the total power-on time since leaving the factory.

[0073] 0A(010) Spin_Retry_Count is the number of times the hard disk spindle motor has been restarted.

[0074] 0B(011) Disk calibration retry count (Calibration_Retry_Count)

[0075] 0C(012) Disk Power-On Count

[0076] C2(194) Temperature (Temperature_Celsius)

[0077] C7(199) Parity Error Rate (UDMA_CRC_Error_Count)

[0078] C8(200) Write Error Rate

[0079] F1(241) Total data written to the disk since it left the factory (Total_LBAs_Written), in units of LBAS = 512 bytes.

[0080] F2(242) Total data read by the disk since it was manufactured (Total_LBAs_Read), in units of LBAS = 512Byt.

[0081] It should be noted that the embodiments of this application do not limit the specific items of SMART health detection.

[0082] S13: Determine whether to re-evaluate the fault status of the monitored hard drive based on the SMART health check results.

[0083] The SMART health check can be based on a quantitative score obtained from the results of each test item, or it can be a combination of the test results into two categories: pass and fail. This application's embodiments do not limit this. Taking the pass / fail status of the SMART health check as an example, if the SMART health check passes, it indicates that the monitored hard drive is healthy; if the SMART health check fails, it is determined that the monitored hard drive needs to undergo a fault reassessment.

[0084] S14: Repair the monitored hard drive based on the results of the fault assessment.

[0085] Fault re-judgment refers to re-detecting faults in the monitored hard drive. In this embodiment, as a preferred option, the fault type of the monitored hard drive is determined, and the monitored hard drive is repaired according to the fault type. The repaired monitored hard drive can be brought back online without affecting business operations.

[0086] As can be seen, the hard drive automatic reset and self-repair method provided in this application embodiment, when the BMC obtains the offline information of the monitored hard drive, performs a preliminary test by controlling the monitored hard drive to reset, performs further tests based on SMART health detection, then re-judges the fault of the monitored hard drive, and finally repairs the monitored hard drive based on the result of the fault re-judge, so that the repaired monitored hard drive can be put back online for use without manual hard drive replacement, saving hard drive spare parts resources, improving hard drive utilization efficiency, and thus reducing the maintenance pressure and maintenance cost of hard drives and faulty hard drives.

[0087] Based on the above embodiments, this application embodiment, as a preferred embodiment, limits the determination of whether to perform SMART health checks on the monitored hard drive based on the reset status of the monitored hard drive to include:

[0088] If the monitored hard drive comes back online, then a SMART health check will be performed on the monitored hard drive.

[0089] If a DNR error is obtained, it is determined that the monitored hard drive will not be subjected to SMART health check;

[0090] If a SMART health check is performed on the monitored hard drive, it also includes: sending a SMART test command to the monitored hard drive;

[0091] If the monitored hard drive is not subjected to SMART health checks, the following is also included: issuing warranty alerts for faulty hard drives.

[0092] Specifically, when the monitored hard drive comes back online, the BMC sends a SMART test command to the monitored hard drive to determine whether to perform a SMART health check and monitor the key indicators of the SMART health check. The purpose of the SMART health check is to monitor the reliability of the hard drive, predict disk failures, and perform various types of disk self-tests.

[0093] When a DNR error is received, it indicates that the monitored hard drive cannot boot. If a SMART health check is not performed on the monitored hard drive, the BMC management interface will issue a faulty disk warranty alarm, reminding maintenance personnel to handle the faulty monitored hard drive in a timely manner.

[0094] It should be noted that the specific items for SMART health monitoring in this application embodiment are not limited, and can include underlying data read error rate, disk read / write throughput performance, impact error rate, etc. The underlying data read error rate can be 0 or any value, and the current value should be much greater than the critical value. The underlying data read error rate is the error that occurs when the read / write head reads data from the disk surface. For some hard drives, a value greater than 0 indicates a problem with the disk surface or the read / write head, such as media damage, head contamination, head resonance, etc. Disk read / write throughput performance represents the hard drive's read / write throughput performance; the higher the value, the better. If the current value is low or close to the critical value, it indicates a serious problem with the hard drive. The impact error rate records the frequency of errors caused by mechanical shocks to the hard drive.

[0095] As can be seen, the hard drive automatic reset and self-repair method provided in this application, after the monitored hard drive is successfully reset, sends a SMART test command to perform a SMART health check on the monitored hard drive to confirm its health status. Based on the SMART health check results, it determines whether to re-evaluate the fault condition of the monitored hard drive. Finally, based on the result of the fault condition re-evaluation, the monitored hard drive is repaired, enabling it to be put back online for use without manual replacement, saving hard drive spare parts resources, improving hard drive utilization efficiency, and thus reducing the maintenance pressure and cost of faulty hard drives. When the monitored hard drive reset fails, a BMC warranty alarm is triggered to remind maintenance personnel to promptly maintain the faulty monitored hard drive and prevent prolonged offline operation of the monitored hard drive from affecting business stability.

[0096] This application embodiment, as a preferred embodiment, limits the determination of whether to re-evaluate the fault condition of the monitored hard drive based on the SMART health detection results to include:

[0097] If the SMART health check passes, then it is determined that the monitored hard drive will not be re-evaluated for faults.

[0098] If the SMART health check fails, the monitored hard drive will be re-evaluated for faults.

[0099] If the monitored hard drive is to be re-evaluated for faults, it also includes: controlling the monitored hard drive to stop working and re-evaluating the faults of the monitored hard drive.

[0100] Specifically, after the monitored hard drive is reset and brought back online, a SMART health check is initiated to perform a SMART health self-check on the monitored hard drive. If the self-check passes, it indicates that the hard drive has no faults and no further fault assessment is needed. At this point, it is determined that no fault assessment is required for the monitored hard drive, and the monitored hard drive can continue to be used for normal customer business deployment, reducing hard drive maintenance. If the self-check fails, in order to quickly and accurately repair the monitored hard drive, it is necessary to suspend the use of the monitored hard drive and perform a fault assessment. Figure 3 This application provides a flowchart for determining whether to re-evaluate the fault status of the monitored hard drive. In practical applications, after the hard drive is brought back online, if... Figure 3 As shown:

[0101] S30: Check if the hard drive SMART health check has passed;

[0102] S31: Hard drive failure reassessment;

[0103] S32: Re-upload the service to the hard drive.

[0104] It should be noted that this application embodiment does not limit the specific indicators for whether the SMART health check passes; several key indicators can be selected as the basis for whether the SMART health check passes. The content of the fault re-judgment is not limited. As a preferred embodiment, the fault re-judgment in this application embodiment mainly detects the fault type of the monitored hard drive, so as to repair the faulty monitored hard drive according to the fault type.

[0105] In summary, the hard drive automatic reset and self-repair method provided in this application allows for normal emergency service deployment and resumption of services on the monitored hard drive when the SMART health check passes, reducing hard drive maintenance costs. When the SMART check fails, services on the monitored hard drive are suspended for further fault assessment, facilitating subsequent repair of the monitored hard drive. The repaired monitored hard drive can then be re-enabled without manual replacement, saving hard drive spare parts resources, improving hard drive utilization efficiency, and ultimately reducing hard drive maintenance pressure and costs.

[0106] Based on the above embodiments, this application embodiment, as a preferred embodiment, limits the repair of the monitored hard drive based on the result of fault problem reassessment to include:

[0107] The type of fault of the monitored hard drive is determined based on the results of the fault reassessment.

[0108] Repair the monitored hard drive according to the type of fault.

[0109] Specifically, when the SMART health check of the monitored hard drive fails, the BMC sends a fault dispute command to re-evaluate the fault and determine the type of fault. Different fault types correspond to different repair methods, and the monitored hard drive is repaired according to the repair method corresponding to the fault type. Fault types typically include two types: soft errors and hard errors. Soft errors refer to faults that can be corrected or recovered using software. Specifically, this refers to faults caused by the loss, corruption, or modification of certain important data on the hard drive, resulting in hard drive boot failure or read / write errors, preventing the system from starting and functioning properly. These are generally caused by misoperation, virus damage, etc., and the hard drive platters and disk body are not damaged; only some tools and software are needed for repair. Hard errors refer to physical faults caused by physical damage to the hard drive's mechanical parts or electronic components. To continue using a hard drive with a hard error, the damaged platters and heads need to be isolated.

[0110] As can be seen, the hard drive automatic reset and self-repair method provided in this application, by determining the fault type of the monitored hard drive that fails the SMART health check through fault problem re-judgment, and repairing the monitored hard drive according to the fault type, can improve the efficiency of repairing the monitored hard drive, thereby reducing the service downtime of the monitored hard drive. After repair, there is no need for manual hard drive replacement, saving hard drive spare parts resources, improving hard drive utilization efficiency, avoiding the impact of long-term hard drive offline on services, and reducing the maintenance pressure and maintenance cost of the hard drive.

[0111] As mentioned in the above embodiments, hard drive failure types include soft errors and hard errors. As a preferred embodiment, this application defines the repair of the monitored hard drive based on the failure type as follows:

[0112] If it is a hard drive soft error, then perform error clearing on the monitored hard drive;

[0113] If it is a hard disk error, read the address of the error reported by the monitored hard disk and isolate the disk and head corresponding to the address.

[0114] Specifically, if the monitored hard drive fails the SMART health check, the fault type is determined through fault re-judgment to identify whether it is a soft or hard error. If the fault type is a soft error, the hard drive is cleaned up, and services can be redeployed after the error is cleared. If the fault type is a hard error, the address of the error reported by the monitored hard drive needs to be read, and the corresponding platters and heads need to be isolated. This isolation method masks the hard error, ensuring that the other platters and heads of the monitored hard drive can be used normally without affecting the operation of the monitored hard drive's services. Figure 4 This application provides a flowchart for repairing a monitored hard drive. For example, in practical applications, after reviewing hard drive failure issues, as shown... Figure 4 As shown:

[0115] S40: Fault type determination;

[0116] If the fault type is a hard error, then execute S41: identify the hard disk fault address;

[0117] S42: Isolate the faulty head and disk;

[0118] Return to S32: Re-upload services to the hard drive;

[0119] If the fault type is a soft error, then execute S43: Hard disk error clearing;

[0120] Return to S32: Re-upload the service to the hard drive.

[0121] As can be seen, the hard drive automatic reset and self-repair method provided in this application, when a soft hard drive error occurs, repairs the monitored hard drive by clearing the hard drive error, enabling the faulty monitored hard drive to resume service; when a hard hard drive error occurs, it repairs the monitored hard drive by isolating the faulty read / write head and platter, enabling the faulty monitored hard drive to resume service. By repairing the monitored hard drive instead of manually replacing the hard drive, hard drive spare parts resources are saved, hard drive utilization efficiency is improved, and the maintenance pressure and maintenance cost of the hard drive are reduced.

[0122] In order to alert hard drive maintenance personnel when the monitored hard drive is offline, this application embodiment, as a preferred embodiment, after obtaining the presence status of the monitored hard drive, if the monitored hard drive is offline, before controlling the monitored hard drive to reset, further includes: recording and displaying the offline status of the monitored hard drive.

[0123] Once the BMC detects the offline status of a monitored hard drive, it should record this status. During routine maintenance, hard drives with frequent offline occurrences can be prioritized for inspection, improving operational efficiency. Displaying the offline status of the monitored hard drive serves as a reminder to maintenance personnel. For example, if a monitored hard drive fails to reset subsequently, maintenance personnel can address the issue promptly, preventing disruption to normal business operations.

[0124] This application embodiment is a preferred embodiment. If the monitored hard drive is offline, controlling the monitored hard drive to reset includes:

[0125] Send reset commands to P3, S2 / S3, and S5 / S6 of the CPLD chip.

[0126] This embodiment of the application achieves forced reset of the monitored hard disk by sending reset commands P3, S2 / S3, and S5 / S6 to the CPLD chip.

[0127] Based on the above embodiments, this application provides a preferred embodiment of a method for automatic hard disk reset and self-repair. Figure 5 This application flowchart illustrates an automatic hard disk reset and self-repair mechanism provided in an embodiment of this application. Figure 5As shown, when the monitored hard drive is offline, the BMC automatically monitors and identifies it, sending a hard drive REST command to power and signal the monitored hard drive, forcing it to reset. To determine if the REST operation was successful, if the monitored hard drive fails to REST, the BMC receives a DNR fault report from the monitored hard drive and issues a warranty alarm; if the REST operation is successful, the monitored hard drive reconnects. However, reconnecting does not guarantee the hard drive's health. The SMART health check function needs to be activated to further test various parameters of the monitored hard drive. If the monitored hard drive passes the SMART health check, it is indeed healthy and services can be re-enabled and deployed. If the monitored hard drive fails the SMART health check, it indicates a fault. Further fault assessment is needed to determine the type of fault and select an appropriate repair method. When the monitored hard drive experiences a soft error, a corresponding error clearing command is sent to repair it. When the failure is a hard error, the fault address is identified, and the corresponding read / write head and platter are the faulty components. These are then isolated, allowing the hard drive to resume service. Although the data associated with the isolated head and platter will be lost, isolating the head is significantly more convenient than replacing the entire hard drive for users with backups.

[0128] In summary, the hard drive automatic reset and self-repair method provided in this application determines the fault type of the monitored hard drive that fails the SMART health check through fault problem re-judgment, and repairs the monitored hard drive according to the fault type, thereby improving the efficiency of hard drive repair and reducing the service downtime of the monitored hard drive. When a soft error occurs in the monitored hard drive, the monitored hard drive is repaired by hard drive error clearing, allowing the faulty monitored hard drive to resume service; when a hard error occurs in the monitored hard drive, the monitored hard drive is repaired by isolating the faulty read / write head and platter, allowing the faulty monitored hard drive to resume service. By repairing the monitored hard drive instead of manually replacing the hard drive, hard drive spare parts resources are saved, hard drive utilization efficiency is improved, and the maintenance pressure and maintenance cost of the hard drive are reduced.

[0129] In the above embodiments, the methods for automatic hard disk reset and self-repair have been described in detail. This application also provides embodiments corresponding to the apparatus for automatic hard disk reset and self-repair. It should be noted that this application describes the embodiments of the apparatus from two perspectives: one based on functional modules and the other based on hardware.

[0130] From the perspective of functional modules Figure 6 This is a schematic diagram of a hard disk automatic reset and self-repair system provided in an embodiment of this application, as shown below. Figure 6 As shown in the figure, this application provides a system for automatic hard disk reset and self-repair, including:

[0131] The acquisition module 10 is used to acquire the presence status of the monitored hard drive;

[0132] The control module 11 is used to control the monitored hard drive to reset if the monitored hard drive is offline.

[0133] The judgment module 12 is used to determine whether to perform SMART health check on the monitored hard drive based on the reset status of the monitored hard drive;

[0134] Module 13 is used to determine whether to re-evaluate the fault status of the monitored hard drive based on the SMART health check results.

[0135] Repair module 14 is used to repair the monitored hard drive based on the results of fault reassessment.

[0136] Since the embodiments of the system part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the system part, and they will not be repeated here.

[0137] From a hardware perspective Figure 7 This is a schematic diagram of a hard disk automatic reset and self-repair device provided in an embodiment of this application, as shown below. Figure 7 As shown, the device for automatic hard disk reset and self-repair includes: a memory 20 for storing computer programs;

[0138] The processor 21 is configured to implement the steps of the hard disk automatic reset and self-repair method as described in the above embodiments when executing a computer program.

[0139] The hard drive automatic reset and self-repair device provided in this embodiment may include, but is not limited to, smartphones, tablets, laptops, or desktop computers.

[0140] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one of the following hardware forms: Digital Signal Processor (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.

[0141] The memory 20 may include one or more computer-readable storage media, which may be non-transitory. The memory 20 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 20 is used to store at least the following computer program 201, which, after being loaded and executed by the processor 21, is capable of implementing the relevant steps of the hard disk automatic reset and self-repair method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 20 may also include an operating system 202 and data 203, and the storage method may be temporary or permanent storage. The operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, the results of SMART health checks.

[0142] In some embodiments, the device for automatic hard disk reset and self-repair may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.

[0143] Those skilled in the art will understand that Figure 7 The structure shown does not constitute a limitation on the device for automatic reset and self-repair of the hard disk, and may include more or fewer components than shown.

[0144] The hard disk automatic reset and self-repair apparatus provided in this application includes a memory and a processor. When the processor executes a program stored in the memory, it can implement the following method: a hard disk automatic reset and self-repair method.

[0145] Finally, this application also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps described in the above method embodiments.

[0146] It is understood that if the methods in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0147] The above provides a detailed description of the method, system, apparatus, and medium for automatic hard disk reset and self-repair provided in this application. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

[0148] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A method for automatic reset and self-repair of a hard disk, characterized in that, include: Obtain the presence status of the monitored hard drive; If the monitored hard drive is offline, then control the monitored hard drive to reset. Determine whether to perform a SMART health check on the monitored hard drive based on the reset status of the monitored hard drive; Based on the results of the SMART health check, determine whether to reassess the fault condition of the monitored hard drive; The monitored hard drive is repaired based on the results of the fault assessment. The step of obtaining the presence status of the monitored hard drive includes: The on-site status of the monitored hard drive is read at a preset frequency. The on-site status of the monitored hard drive includes offline status and online status. When the presence status of the monitored hard drive is read as online, the presence status of the monitored hard drive continues to be read at the preset frequency. The step of determining whether to perform a SMART health check on the monitored hard drive based on the reset status of the monitored hard drive includes: If the monitored hard drive comes back online, then a SMART health check will be performed on the monitored hard drive. If a DNR error is obtained, it is determined that no SMART health check will be performed on the monitored hard drive. If a SMART health check is performed on the monitored hard drive, the method further includes: sending a SMARTtest command to the monitored hard drive; If the monitored hard drive is not subjected to SMART health testing, the method also includes: issuing a warranty alert for a faulty hard drive; The process of determining whether to reassess the fault condition of the monitored hard drive based on the SMART health check results includes: If the SMART health check passes, it is determined that no further fault assessment will be performed on the monitored hard drive. If the SMART health check fails, the monitored hard drive will be reassessed for faults. If a fault diagnosis is performed on the monitored hard drive, the method further includes: controlling the monitored hard drive to stop working and performing a fault diagnosis on the monitored hard drive. The repair of the monitored hard drive based on the results of the fault assessment includes: The type of fault of the monitored hard drive is determined based on the results of the fault reassessment. Repair the monitored hard drive according to the fault type; The repair of the monitored hard drive based on the fault type includes: If it is a hard drive soft error, then perform error clearing on the monitored hard drive; If it is a hard disk error, the address of the error reported by the monitored hard disk is read, and the disk and read / write head corresponding to the address are isolated.

2. The method of hard disk auto-reset and self-repair according to claim 1, wherein, After obtaining the presence status of the monitored hard drive, before controlling the monitored hard drive to reset if the monitored hard drive is offline, the method further includes: recording and displaying the offline status of the monitored hard drive.

3. The method for automatic hard disk reset and self-repair according to claim 2, characterized in that, If the monitored hard drive is offline, controlling the monitored hard drive to reset includes: Send reset commands to P3, S2 / S3, and S5 / S6 of the CPLD chip.

4. A system for automatic hard disk reset and self-repair, characterized in that, include: The acquisition module is used to obtain the presence status of the monitored hard drive; The control module is used to control the monitored hard drive to reset if the on-site state of the monitored hard drive is offline. The judgment module is used to determine whether to perform SMART health check on the monitored hard drive based on the reset status of the monitored hard drive; The determination module is used to determine whether to re-evaluate the fault condition of the monitored hard drive based on the SMART health detection results. The repair module is used to repair the monitored hard drive based on the results of the fault reassessment. The step of obtaining the presence status of the monitored hard drive includes: The on-site status of the monitored hard drive is read at a preset frequency. The on-site status of the monitored hard drive includes offline status and online status. When the presence status of the monitored hard drive is read as online, the presence status of the monitored hard drive continues to be read at the preset frequency. The step of determining whether to perform a SMART health check on the monitored hard drive based on the reset status of the monitored hard drive includes: If the monitored hard drive comes back online, then a SMART health check will be performed on the monitored hard drive. If a DNR error is obtained, it is determined that no SMART health check will be performed on the monitored hard drive. If a SMART health check is performed on the monitored hard drive, the method further includes: sending a SMARTtest command to the monitored hard drive; If the monitored hard drive is not subjected to SMART health testing, the method also includes: issuing a warranty alert for a faulty hard drive; The process of determining whether to reassess the fault condition of the monitored hard drive based on the SMART health check results includes: If the SMART health check passes, it is determined that no further fault assessment will be performed on the monitored hard drive. If the SMART health check fails, the monitored hard drive will be reassessed for faults. If a fault diagnosis is performed on the monitored hard drive, the method further includes: controlling the monitored hard drive to stop working and performing a fault diagnosis on the monitored hard drive. The repair of the monitored hard drive based on the results of the fault assessment includes: The type of fault of the monitored hard drive is determined based on the results of the fault reassessment. Repair the monitored hard drive according to the fault type; The repair of the monitored hard drive based on the fault type includes: If it is a hard drive soft error, then perform error clearing on the monitored hard drive; If it is a hard disk error, the address of the error reported by the monitored hard disk is read, and the disk and read / write head corresponding to the address are isolated.

5. A device for automatic hard disk reset and self-repair, characterized in that, Includes memory used to store computer programs; A processor, configured to execute the computer program to implement the steps of the method for automatic hard disk reset and self-repair as described in any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for automatic hard disk reset and self-repair as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Hard disk failure handling method, apparatus, server, and computer-readable medium

    CN109284207A