Server cluster hard disk failure processing method and device, electronic equipment and storage medium

By automating the detection and repair of hard drive failures in server clusters, the problems of high hard drive replacement rates and high risk of data loss have been solved, achieving efficient hard drive failure handling, reducing costs, and improving user experience.

CN111897686BActive Publication Date: 2026-04-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2020-08-05
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies for handling hard drive failures in server clusters result in high hard drive replacement rates, significant data loss risks, and long troubleshooting times, impacting user experience.

Method used

By automatically detecting the type of hard drive failure in the server cluster, a hardware failure detection process is triggered to obtain the detection results and perform repairs based on the results, including voltage reset and backplane slot reseat operations.

Benefits of technology

It reduces hard drive replacement rate, lowers operating costs, improves hard drive maintenance efficiency and data security, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111897686B_ABST
    Figure CN111897686B_ABST
Patent Text Reader

Abstract

The application provides a server cluster hard disk fault processing method and device, electronic equipment and a storage medium, and the method comprises the following steps: obtaining fault warning information of a server cluster hard disk; calling running data of the server cluster hard disk; determining the type of the server cluster hard disk fault based on the running data of the server cluster hard disk; when the type of the server cluster hard disk fault is an unavailable alarm, triggering a matched hardware-free fault detection process, and obtaining a detection result of the hardware-free fault detection process; and repairing the server cluster hard disk fault based on the detection result of the hardware-free fault detection process. The application can automatically detect the fault type of the server cluster hard disk, repair the server cluster hard disk fault, reduce the replacement rate of the server cluster hard disk, reduce the operation cost of the server cluster system, improve the efficiency of server cluster hard disk maintenance, ensure the data security of the server cluster user, and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to hard disk system fault detection and handling technology, and more particularly to server cluster hard disk fault handling methods, devices, electronic equipment and storage media. Background Technology

[0002] With the continuous development of computer technology, server clusters can provide secure and reliable elastic computing services, and can also offer different instance types to meet specific user scenarios. These server cluster instance types consist of different combinations of CPU, memory, storage, and network. However, when a server cluster's hard drive experiences issues such as failure or read-only access, the services of user-deployed sub-machines on that server will also be affected. Related technologies that rely on replacing the faulty hard drive for fault recovery not only increase the waiting time for fault handling but also pose a risk of data loss, impacting user experience. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for handling server cluster hard disk failures, which can automatically detect the types of server cluster hard disk failures, repair server cluster hard disk failures, reduce the replacement rate of server cluster hard disks, reduce the operating cost of server cluster systems, improve the efficiency of server cluster hard disk maintenance, ensure the data security of server cluster users, and improve the user experience.

[0004] The technical solution of this invention is implemented as follows:

[0005] This invention provides a method for handling hard disk failures in a server cluster, the method comprising:

[0006] Obtain fault warning information for the server cluster hard drives;

[0007] In response to the fault warning information of the server cluster hard disk, retrieve the operating data of the server cluster hard disk;

[0008] Based on the operating data of the server cluster hard disks, determine the type of hard disk failure in the server cluster;

[0009] When the type of hard disk failure in the server cluster is an unavailable alarm, a matching hardware failure detection process is triggered, and the detection result of the hardware failure detection process is obtained.

[0010] Based on the detection results of the hardware fault detection process, the hard disk faults of the server cluster are repaired.

[0011] This invention also provides a server cluster hard disk failure handling device, the device comprising:

[0012] The information transmission module is used to obtain fault warning information of the server cluster hard disk;

[0013] The information processing module is used to respond to the fault warning information of the server cluster hard disk and retrieve the operating data of the server cluster hard disk;

[0014] The information processing module is used to determine the type of hard disk failure in the server cluster based on the operating data of the server cluster hard disk;

[0015] The information processing module is used to trigger a matching hardware failure detection process and obtain the detection result of the hardware failure detection process when the type of hard disk failure in the server cluster is an unavailable alarm.

[0016] The information processing module is used to repair the hard disk failure of the server cluster based on the detection results of the hardware failure detection process.

[0017] In the above scheme,

[0018] The information transmission module is used to monitor the operating parameters of the server cluster hard disk. When the operating parameters of the server cluster hard disk reach the warning threshold, it triggers the fault warning information of the server cluster hard disk.

[0019] The information transmission module is used to receive alarm information from server cluster users, and based on the parameters of the server cluster users, determine the server cluster hard disk that matches the server cluster users, and trigger the fault warning information of the server cluster hard disk.

[0020] In the above scheme,

[0021] The information processing module is used to determine the hard disk attribute parameters of the server cluster hard disk, wherein the hard disk attribute parameters include: hard disk model, server model, racking time, version number, and hard disk partition identifier;

[0022] The information processing module is used to retrieve the operating data of the server cluster hard disk stored in the corresponding storage medium based on the hard disk attribute parameters of the server cluster hard disk.

[0023] The information processing module is used to obtain the corresponding fault type characteristics based on the hard disk attribute parameters of the server cluster hard disk;

[0024] The information processing module is used to determine the type of server cluster hard disk failure by cross-comparing the operating data of the server cluster hard disk with the failure type characteristics, wherein the failure types of the server cluster include: unavailability alarm and failure alarm.

[0025] In the above scheme,

[0026] The information processing module is used to obtain detection parameters that match the hard disk of the server cluster when the type of hard disk failure is an unavailable alarm by triggering a matching hardware failure detection process.

[0027] The information processing module is used to perform multi-dimensional processing on the detection parameters that match the hard disk of the server cluster, and determine the detection results of the hard disk of the server cluster in different dimensions.

[0028] The information processing module is used to fuse the detection results of the server cluster hard disk in different dimensions to determine the detection result of the no-hardware-fault detection process.

[0029] In the above scheme,

[0030] The information processing module is used to determine the corresponding parameter health detection results based on the hard disk protection parameters corresponding to the hard disks of the server cluster.

[0031] The information processing module is used to determine the distribution characteristic detection result of the hard disk protection parameters by standardizing the hard disk protection parameters.

[0032] The information processing module is used to determine the deterioration trend of the hard disk protection parameters by processing the dynamic slope of the hard disk protection parameters;

[0033] The information processing module is used to determine the failure probability result corresponding to the hard disk protection parameters by processing the prediction function of the hard disk protection parameters;

[0034] The information processing module is used to fuse the parameter health detection results, distribution characteristic detection results, deterioration trend results, and failure probability results to determine the detection results of the server cluster hard disk in different dimensions.

[0035] In the above scheme,

[0036] The information processing module is used to determine the fault repair method matching the hard disk of the server cluster based on the detection results of the no-hardware-fault detection process.

[0037] The information processing module is used to trigger a corresponding fault repair process based on a fault repair method that matches the hard disk of the server cluster, and to repair the fault of the hard disk of the server cluster through the fault repair process.

[0038] In the above scheme,

[0039] The information processing module is used to determine, when it is determined that the server cluster hard disk supports voltage reset function, the fault repair method matching the server cluster hard disk is voltage reset process;

[0040] The information processing module is used to trigger boot code carrying voltage reset instructions through the voltage reset process, and adjust the loading voltage of the server cluster hard disk through the voltage reset instructions in the boot code to repair the server cluster hard disk failure.

[0041] In the above scheme,

[0042] The information processing module is used to determine, when it is determined that the server cluster hard disk supports the power-on / off function of the hard disk backplane storage, the fault repair method matching the server cluster hard disk is the backplane slot plugging and unplugging process.

[0043] The information processing module is used to trigger boot code carrying backplane slot insertion / removal instructions through the backplane slot insertion / removal process.

[0044] The information processing module is used to insert or remove the slots of the server cluster hard disk backplane storage via the backplane slot insertion / removal instructions in the boot code, so as to repair the server cluster hard disk fault by adjusting the slot position of the server cluster hard disk backplane storage.

[0045] In the above scheme,

[0046] The information processing module is used to detect the drive letter location of the hard disk in the server cluster;

[0047] The information processing module is used to determine the initial configuration information of the server cluster hard disk slot, disk letter, and mount point when it is determined that the disk letter of the server cluster hard disk has shifted position.

[0048] The information processing module is used to adjust the drive letters that have positional offsets based on the initial configuration information of the hard drive slots, drive letters, and mount points of the server cluster.

[0049] In the above scheme, the device further includes:

[0050] The display module is used to display the user interface, which includes a perspective view of the server cluster operating environment from a fixed perspective, and the user interface includes different server cluster identifiers.

[0051] The user interface also includes an editing detection component and a repair component;

[0052] The display module is used to monitor the operating parameters of the server cluster hard disk through the detection component, and trigger a fault warning message for the server cluster hard disk when the operating parameters of the server cluster hard disk reach the warning threshold.

[0053] Based on the repair component and the detection results of the hardware fault detection process, the hard disk faults of the server cluster are repaired.

[0054] In the above scheme,

[0055] The display module is used to call the interface of the target server cluster hard disk through the detection component;

[0056] The display module is used to send a query command through the interface of the target server cluster hard disk based on the repair component, so as to realize the initial configuration information of the slot, drive letter and mount point of the target server cluster hard disk through the repair component.

[0057] This invention also provides an electronic device, the electronic device comprising:

[0058] Memory, used to store executable instructions;

[0059] The processor, when running executable instructions stored in the memory, implements the preceding server cluster hard disk failure handling method.

[0060] This invention also provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement a preceding server cluster hard disk failure handling method.

[0061] The embodiments of the present invention have the following beneficial effects:

[0062] This invention acquires fault warning information from server cluster hard drives; responds to the fault warning information by retrieving the operational data of the server cluster hard drives; determines the type of server cluster hard drive fault based on the operational data; when the fault type is an unavailability alarm, triggers a matching hardware-free fault detection process and acquires the detection results of the hardware-free fault detection process; and repairs the server cluster hard drive fault based on the detection results. This allows for automated detection and repair of server cluster hard drive fault types, reducing the replacement rate of server cluster hard drives, lowering the operating costs of the server cluster system, improving the efficiency of server cluster hard drive maintenance, ensuring data security for server cluster users, and enhancing the user experience. Attached Figure Description

[0063] Figure 1 This is a schematic diagram illustrating a usage scenario of the server cluster hard disk failure handling method provided in an embodiment of the present invention.

[0064] Figure 2 A schematic diagram of the composition structure of an electronic device provided in an embodiment of the present invention;

[0065] Figure 3 This is an optional flowchart illustrating the server cluster hard disk failure handling method provided in an embodiment of the present invention;

[0066] Figure 4 This is an optional flowchart illustrating the server cluster hard disk failure handling method provided in an embodiment of the present invention;

[0067] Figure 5 This is a schematic diagram of voltage reset operation in an embodiment of the present invention;

[0068] Figure 6 This is a schematic diagram of an optional hard drive repair method in an embodiment of the present invention;

[0069] Figure 7 This is a schematic diagram of an optional hard drive repair method in an embodiment of the present invention;

[0070] Figure 8 A front-end display diagram of the server cluster hard disk failure handling method provided in this application;

[0071] Figure 9 A schematic diagram illustrating the process of the server cluster hard disk failure handling method provided in this application;

[0072] Figure 10 This is a front-end display diagram of the server cluster hard disk failure handling method provided in this application. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0074] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0075] In the implementation of this application, the collection and processing of relevant data should be strictly in accordance with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0076] Before providing a further detailed description of the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention will be explained, and the nouns and terms involved in the embodiments of the present invention shall be interpreted as follows.

[0077] 1) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0078] 2) Terminals, including but not limited to: ordinary terminals and dedicated terminals, wherein the ordinary terminals maintain a long connection and / or a short connection with the transmission channel, and the dedicated terminals maintain a long connection with the transmission channel.

[0079] 3) Client: The carrier that implements specific functions in the terminal. For example, a mobile client (APP) is a carrier of specific functions in a mobile terminal, such as performing report creation or report display functions.

[0080] 4) Firmware: This is the code running inside the chip, which is binary code used to implement hard disk failure detection in server clusters.

[0081] 5) A mini program is a type of application developed using a front-end-oriented language (such as JavaScript) and running services within Hyper Text Markup Language (HTML) pages. It is downloaded by a client (such as a browser or any client with an embedded browser engine) via a network (such as the internet) and interpreted and executed within the client's browser environment, saving the step of installation on the client side. For example, mini programs can be downloaded and run on social networking clients to provide various services such as airline ticket purchase, report generation, and data display.

[0082] 6) Runtime environment, the engine used to interpret and execute code. For example, for the runtime environment of a mini-program, it could be JavaScript Core on the iOS platform or X5 JS Core on the Android platform.

[0083] 7) Bootloader code: also known as bootloader, boot mode, startup load, etc., refers to a type of code that runs when the chip starts up. It is usually used to initialize the hardware environment and load the firmware code to run. Usually, it does not need to be updated.

[0084] 8) Components are functional modules of the view in a mini-program, also known as front-end components. They include buttons, titles, tables, sidebars, content, and footers on the page. Components include modular code so that they can be reused in different pages of the mini-program.

[0085] 9) A server cluster refers to a group of servers working together to provide the same service, appearing to the client as a single server. Server clusters can utilize multiple computers for parallel computing to achieve high computing speeds, and can also use multiple computers for backup, ensuring the entire system continues to operate normally even if one machine fails. The server cluster hard disk failure handling method provided in this application can be applied to cloud server and distributed server scenarios, enabling status detection and fault repair of server hard disks in different usage scenarios. Specifically, a cloud server (CVM Cloud Virtual Machine) is a simple, efficient, secure, reliable, and elastically scalable computing service. Its management is simpler and more efficient than traditional single physical servers. Users can quickly create or release any number of cloud servers for their business processes and store user data without purchasing hardware in advance. In a distributed server environment, user data and programs may not reside on a single server but are distributed across multiple servers. Similarly, distributed server environments also require a large number of hard disks, necessitating the server hard disk status detection and fault repair methods provided in this application.

[0086] Figure 1 This is a schematic diagram illustrating a usage scenario of the server cluster hard disk failure handling method provided in this embodiment of the invention. (See attached diagram.) Figure 1With the continuous development of computer technology, cloud servers (Cloud Virtual Machines, CVMs) in server clusters can provide secure and reliable elastic computing services, and can also provide different instance types to meet specific user scenarios. Terminals (including terminals 10-1 and 10-2) are equipped with corresponding clients capable of performing different functions. These clients (including terminals 10-1 and 10-2) obtain different information from the corresponding cloud servers 200 via network 300 and can deploy different services within the server cluster. The terminals connect to the cloud servers 200 via network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both, using a wireless link for data transmission. The instance types provided by the server cluster consist of different combinations of CPU, memory, storage, and network, and store user business data on the server cluster's hard drive. However, if the server cluster's hard drive experiences issues such as disconnection or read-only access, the user's sub-machine services deployed on that server will also be affected. In the embodiments provided by this invention, the server cluster application running on the cloud server 200 can be written in software code environments using different programming languages, and the code objects can be different types of code entities. For example, in C language software code, a code object can be a function. In Java language software code, a code object can be a class, and in iOS Objective-C, it can be a piece of object code. In C++ language software code, a code object can be a class or a function to execute processing instructions from different terminals. This application does not distinguish the source of the compilation environment for the name server cluster.

[0087] The following is a detailed description of the structure of the server cluster hard disk failure handling device according to an embodiment of the present invention. The server cluster hard disk failure handling device can be implemented in various forms, such as a dedicated terminal with server cluster hard disk failure handling device function, or a server equipped with server cluster hard disk failure handling device function, such as the preceding one. Figure 1 The cloud server in the middle is 200. Figure 2 This is a schematic diagram of the composition structure of the server cluster hard disk failure handling device provided in an embodiment of the present invention. It can be understood that... Figure 2 This is only an exemplary structure of the server cluster hard disk failure handling device, not the entire structure; it can be implemented as needed. Figure 2 The structure shown may be part or all of the structure.

[0088] The electronic device provided in this embodiment of the invention includes: at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. The various components in the server cluster hard disk failure handling device are coupled together via a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 205.

[0089] The user interface 203 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.

[0090] It is understood that memory 202 can be volatile memory or non-volatile memory, or both. In this embodiment of the invention, memory 202 is capable of storing data to support the operation of a terminal (such as 10-1). Examples of this data include any computer programs used to operate on the terminal (such as 10-1), such as operating systems and applications. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications.

[0091] In some embodiments, the server cluster hard disk failure handling device provided in this invention can be implemented using a combination of hardware and software. For example, the server cluster hard disk failure handling device provided in this invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the server cluster hard disk failure handling method provided in this invention. For instance, the processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0092] As an example of the server cluster hard disk failure handling device provided in this embodiment of the invention, which is implemented by combining software and hardware, the server cluster hard disk failure handling device provided in this embodiment of the invention can be directly embodied as a combination of software modules executed by processor 201. The software modules can be located in the storage medium, which is located in memory 202. Processor 201 reads the executable instructions included in the software modules in memory 202 and combines them with necessary hardware (e.g., including processor 201 and other components connected to bus system 205) to complete the server cluster hard disk failure handling method provided in this embodiment of the invention.

[0093] As an example, processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0094] As an example of the hardware implementation of the server cluster hard disk failure handling device provided in this embodiment of the invention, the device provided in this embodiment of the invention can be directly executed by a processor 201 in the form of a hardware decoding processor. For example, it can be executed by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components to implement the server cluster hard disk failure handling method provided in this embodiment of the invention.

[0095] In this embodiment of the invention, the memory 202 is used to store various types of data to support the operation of the server cluster hard disk failure handling device. Examples of such data include: any executable instructions for operation on the server cluster hard disk failure handling device, such as executable instructions that can be included in a program implementing the server cluster hard disk failure handling method of this embodiment of the invention.

[0096] In other embodiments, the server cluster hard disk failure handling device provided in this invention can be implemented in software. Figure 2A server cluster hard disk failure handling device stored in memory 202 is shown. This device can be software in the form of programs and plug-ins, and includes a series of modules. As an example of a program stored in memory 202, it may include the server cluster hard disk failure handling device. The server cluster hard disk failure handling device includes the following software modules: an information transmission module 2081 and an information processing module 2082. When the software modules in the server cluster hard disk failure handling device are read into RAM and executed by processor 201, the server cluster hard disk failure handling method provided in this embodiment of the invention will be implemented. The functions of each software module in the server cluster hard disk failure handling device include:

[0097] The information transmission module 2081 is used to obtain fault warning information of the server cluster hard disk;

[0098] The information processing module 2082 is used to retrieve the operating data of the server cluster hard disk in response to the fault warning information of the server cluster hard disk;

[0099] The information processing module 2082 is used to determine the type of hard disk failure of the server cluster based on the operating data of the server cluster hard disk;

[0100] The information processing module 2082 is used to trigger a matching hardware failure detection process and obtain the detection result of the hardware failure detection process when the type of hard disk failure in the server cluster is an unavailable alarm.

[0101] The information processing module 2082 is used to repair the hard disk failure of the server cluster based on the detection results of the hardware failure detection process.

[0102] As described in the preceding embodiments, the related technologies that rely on replacing faulty hard drives for fault recovery not only increase the waiting time for fault handling but also pose a risk of data loss, impacting user experience. In existing methods, when a server cluster experiences hard drive failure, the data center replaces the faulty hard drive for recovery. Specifically, when a hard drive becomes unavailable (read-only, offline, etc.), the options are to replace the hard drive or read individual SMART parameter values ​​of the hard drive for a secondary assessment to determine whether to replace it or reuse it. If reuse is possible, the original hard drive is reconnected by restarting the server. The drawbacks are: 1) If a direct replacement is adopted, since the NTF (no trouble found) ratio of hard drives in a server cluster system is approximately 30%~40%, and may even reach over 50% in special business scenarios, manually replacing these hard drives significantly increases labor and material costs, and also increases unnecessary fault handling time, drastically increasing the risk of business unavailability. 2) Using SMART parameters to determine whether a disk needs to be replaced is insufficient. Relying solely on the current SMART parameter snapshot cannot accurately determine the health of the hard drive, resulting in a high false alarm rate and the risk of repeated failures. Furthermore, in server cluster hard drive environments with non-RAID card topologies (direct HBA / PCH connection), directly plugging and unplugging hard drives may cause system drive letter drift. The usual practice is to restart the server to restore it, which involves multiple steps, is time-consuming, and simultaneously affects services on other hard drives on the entire machine, causing the failure to affect more server cluster users.

[0103] To overcome the above-mentioned shortcomings, refer to Figure 3 This application provides a method for handling hard disk failures in a server cluster. To overcome the aforementioned shortcomings, this invention provides a method for handling hard disk failures in a server cluster. See [link to relevant documentation]. Figure 3 , Figure 3 This is an optional flowchart illustrating the server cluster hard disk failure handling method provided in this embodiment of the invention. It can be understood that... Figure 3 The steps shown can be performed by various electronic devices running server cluster hard disk failure handling equipment, such as mobile phones or tablets with server cluster hard disk failure handling capabilities. The dedicated terminal with server cluster hard disk failure handling equipment can be packaged within... Figure 1 In the terminal 101-1 shown, to execute the preprocessor. Figure 2 The corresponding software modules in the server cluster hard disk failure handling device shown below. The following section addresses... Figure 3 The steps shown are explained.

[0104] Step 301: The server cluster hard disk failure handling device obtains the failure warning information of the server cluster hard disk.

[0105] In some embodiments of the present invention, obtaining fault warning information of the server cluster hard disk can be achieved in the following ways:

[0106] The system monitors the operating parameters of the server cluster hard drives. When the operating parameters of the server cluster hard drives reach a warning threshold, a fault warning message for the server cluster hard drives is triggered. Alternatively, it receives alarm information from server cluster users and, based on the parameters of the server cluster users, determines the server cluster hard drives that match the users and triggers a fault warning message for those hard drives. The server cluster hard drive fault handling method provided in this application can be applied to cloud server and distributed server usage scenarios, enabling status detection and fault repair of server hard drives in different usage scenarios. Therefore, this embodiment of the invention can be implemented in conjunction with cloud technology. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network or local area network to achieve data computation, storage, processing, and sharing. It can also be understood as a general term for network technology, information technology, integration technology, management platform technology, and application technology based on cloud computing business models. Backend services of network systems require a large amount of computing and storage resources, such as video websites, image websites, and many portal websites; therefore, cloud technology needs cloud computing as its support.

[0107] It's important to note that cloud computing is a computing model that distributes computing tasks across a resource pool comprised of numerous computers, enabling various application systems to access computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, resources in the "cloud" are infinitely scalable, readily available, and can be used on demand, expanded at any time, and paid for based on usage. As a provider of fundamental cloud computing capabilities, a cloud resource pool platform, often referred to as a cloud platform or Infrastructure as a Service (IaaS), is established. This platform deploys various types of virtual resources within the resource pool for external customers to choose from. The cloud resource pool primarily includes: computing devices (which can be virtualized machines containing operating systems), storage devices, and network devices. When users use server clusters to store data or deploy different application processes, monitoring the operating parameters of the server cluster's hard drives can promptly detect potential server cluster hard drive failures, preventing data loss due to server cluster hard drive failures that trigger failure warnings.

[0108] Step 302: In response to the fault warning information of the server cluster hard disk, the server cluster hard disk fault handling device retrieves the operating data of the server cluster hard disk.

[0109] Step 303: The server cluster hard disk failure handling device determines the type of server cluster hard disk failure based on the operating data of the server cluster hard disk.

[0110] In some embodiments of the present invention, determining the type of server cluster hard disk failure based on the operational data of the server cluster hard disk can be achieved in the following ways:

[0111] The process involves determining the hard drive attribute parameters of the server cluster hard drives, including: hard drive model, server model, rack installation time, version number, and hard drive partition identifier. Based on these parameters, the process retrieves the operational data of the server cluster hard drives stored in the corresponding storage media. Then, based on the hard drive attribute parameters, the process obtains corresponding fault type characteristics. Finally, by cross-referencing the operational data with the fault type characteristics, the type of server cluster hard drive fault is determined. The fault types include unavailability alarms and failure alarms. Since a large number of hard drives are used in a server cluster system, and these hard drives may come from different hardware manufacturers or be customized by the server cluster operator, cross-referencing the hard drive model, server model, rack installation time, version number, and hard drive partition identifier with the fault type characteristics allows for more accurate fault type identification, avoiding false positives and false negatives due to inconsistent hard drive versions. The threshold values ​​for different types of server cluster hard drives are reliable attribute values ​​specified by the hard drive manufacturers and calculated using a specific formula. If any attribute value exceeds the corresponding threshold, it means the hard drive will become unreliable, and the data stored on it is easily lost. The composition and size of attribute values ​​differ for different types of hard drives, so different thresholds are set for different models of hard drives. During the hard drive failure handling process, cross-checking is required to reduce the probability of false alarms for different failure types.

[0112] Step 304: Server Cluster Hard Disk Failure Handling Device When the type of server cluster hard disk failure is unavailable alarm, the device triggers a matching hardware failure detection process and obtains the detection result of the hardware failure detection process.

[0113] In some embodiments of the present invention, when the server cluster hard disk failure type is an unavailability alarm, a matching hardware failure detection process is triggered to obtain the detection result of the hardware failure detection process; when the server cluster hard disk failure type is an unavailability alarm, a matching hardware failure detection process is triggered to obtain detection parameters matching the server cluster hard disk; the detection parameters matching the server cluster hard disk are processed in multiple dimensions to determine the detection results of the server cluster hard disk in different dimensions; the detection results of the server cluster hard disk in different dimensions are fused to determine the detection result of the hardware failure detection process. Specifically, when the fault type is determined to be a SMART (Self-Monitoring Analysis and Reporting Technology) pre-failure alarm, a hard disk replacement process is triggered to replace the hard disk; when the fault type is a read-only or offline unavailability alarm, a fault detection process is triggered. Here, SMART is an automatic hard disk status detection and early warning system and standard. By monitoring and recording the operation of hard drive hardware components such as read / write heads, platters, motors, and circuitry using detection commands within the hard drive hardware, and comparing this information with preset safety values ​​set by the manufacturer, the system can automatically warn the user and perform minor automatic repairs if the monitored conditions are about to or have exceeded the preset safety range. This proactively protects the hard drive data. SMART uses binary code as its basic instructions, which are written into standard registers to form a specific SMART information table for normal detection and operation. SMART instructions are divided into main commands (Commands) and subcommands (Subcommands). The main command primarily provides information on whether the device supports SMART or ignores certain command features. Subcommands provide detection information supporting SMART devices.

[0114] Continue to refer to Figure 4 , Figure 4 This is an optional flowchart illustrating the server cluster hard disk failure handling method provided in this embodiment of the invention. It can be understood that... Figure 4 The steps shown can be performed by various electronic devices running server cluster hard disk failure handling equipment, such as mobile phones or tablets with server cluster hard disk failure handling capabilities. The dedicated terminal with server cluster hard disk failure handling equipment can be packaged within... Figure 1 In the terminal 101-1 shown, to execute the preprocessor. Figure 2 The corresponding software modules in the server cluster hard disk failure handling device shown below. The following section addresses... Figure 4 The steps shown are explained.

[0115] Step 401: Based on the hard disk protection parameters corresponding to the hard disks in the server cluster, determine the corresponding parameter health test results.

[0116] Unlike conventional SMART assessment methods, the server cluster hard drive failure detection method provided in this application can calculate the hard drive health score from multiple dimensions using different weighted algorithms. These include: the SMART parameter Euclidean distance algorithm (for measuring the health of key SMART parameters); the SMART parameter z-score algorithm (for statistically quantifying the distribution of hard drive parameters within the cluster); the SMART parameter dynamic slope algorithm (for dynamically quantifying the deterioration trend of parameters); and a hard drive underlying parameter machine learning failure prediction algorithm (developing parameters reflecting the health of the hard drive internally in cooperation with hard drive suppliers and performing big data machine learning). The weighted calculation of the hard drive health score refers to Formula 1. In different server cluster usage environments, operations and maintenance personnel can dynamically adjust the different weights in Formula 1 according to the business type and environmental parameters. In a cloud server usage environment, the optional values ​​for the weights are: a0=0.2, a1=0.2, a2=0.3, a3=0.3. Formula 1 is:

[0117] Formula 1

[0118]

[0119] Step 402: Determine the distribution characteristics detection results of the hard disk protection parameters by standardizing the hard disk protection parameters.

[0120] Step 403: Determine the deterioration trend of the hard disk protection parameters by processing the dynamic slope of the hard disk protection parameters.

[0121] Step 404: By processing the prediction function of the hard disk protection parameters, determine the failure probability result corresponding to the hard disk protection parameters.

[0122] Step 405: The health detection results, distribution characteristic detection results, deterioration trend results, and failure probability results of the parameters are fused to determine the detection results of the server cluster hard disk in different dimensions.

[0123] Compared to related technologies that rely on a single dimension to determine server cluster hard drive failures, this application can determine the detection results of server cluster hard drives across different dimensions based on relevant detection parameters. This allows for the calculation of a health score for faulty server cluster hard drives. When the faulty hard drive score is greater than or equal to a preset threshold, the server cluster hard drive is determined to be in an NTF (no trouble found) state (where NTF indicates no hardware failure, no hardware faults found in electronic components, and it is generally reusable), and the faulty hard drive is repaired. If the hard drive detection score is less than the preset threshold, it indicates a hardware failure, and the drive is manually replaced. The server cluster hard drive determination is based on formula 2:

[0124] Formula 2.

[0125] When passing Figure 4 After determining the detection result of the hardware fault detection process as shown in the steps, you can continue to execute step 305.

[0126] Step 305: The server cluster hard disk failure handling device repairs the server cluster hard disk failure based on the detection results of the hardware failure detection process.

[0127] In some embodiments of the present invention, based on the detection results of the hardware fault detection process, the hard disk fault of the server cluster is repaired, including:

[0128] Based on the detection results of the hardware fault detection process, a fault repair method matching the server cluster hard drive is determined. Based on the fault repair method matching the server cluster hard drive, the corresponding fault repair process is triggered, and the faulty hard drive is repaired through this process. Specifically, when the hard drive recovery process is triggered, it can first be determined whether the hard drive / server supports the hard drive pin (PIN3) voltage reset function. If it does, an out-of-band / in-band command is sent to reset the PIN3 voltage, allowing the faulty hard drive to power on and off. If the PIN3 voltage reset function is not supported, it is determined whether the server supports the backplane one-time compiler (efuse) independent slot power-on / off function. If it does, an out-of-band / in-band command is sent to power on and off the faulty drive slot on the backplane. If it is not supported, manual plugging and unplugging is required for repair. This reduces the frequency of manually replacing faulty hard drives and decreases the hardware operating cost of the server cluster hard drives.

[0129] In some embodiments of the present invention, based on a fault repair method matching the server cluster hard disk, a corresponding fault repair process is triggered, and the fault repair process is used to repair the server cluster hard disk fault, including:

[0130] When it is determined that the server cluster hard drive supports voltage reset functionality, the fault repair method matching the server cluster hard drive is determined to be a voltage reset process. Through this voltage reset process, boot code carrying voltage reset instructions is triggered, and the voltage reset instructions in the boot code adjust the loading voltage of the server cluster hard drive to repair the fault. (Refer to...) Figure 5 , Figure 5 This is a schematic diagram illustrating the voltage reset operation in an embodiment of the present invention. Specifically, the third pin (Pin3) of a traditional SATA / SAS hard drive is a reserved pin. In this embodiment, the server cluster hard drives have added a power disable function to Pin3, meaning the host system can control the power-on / off reset of the hard drive through Pin3. Furthermore, it can be designed to control the PIN3 voltage of a single hard drive via the BMC to control the backplane CPLD, enabling independent power-on and power-off of the hard drives. This triggers boot code carrying a voltage reset instruction, which adjusts the load voltage of the server cluster hard drives, reducing the probability of power failure and preventing data loss due to server cluster hard drive power loss.

[0131] In some embodiments of the present invention, the step of triggering a corresponding fault repair process based on a fault repair method matching the server cluster hard disk, and repairing the server cluster hard disk fault through the fault repair process, includes:

[0132] When it is determined that the server cluster hard drive supports the power-on / off function of the hard drive backplane storage, the fault repair method matching the server cluster hard drive is determined to be the backplane slot insertion / removal process. This process triggers boot code carrying backplane slot insertion / removal instructions. The boot code uses these instructions to insert / remove the backplane storage slots of the server cluster hard drive, thereby repairing the server cluster hard drive fault by adjusting the slot positions. (Refer to...) Figure 6 , Figure 6This is an optional hard drive repair schematic diagram in an embodiment of the present invention. Specifically, power-on / off commands are issued out-of-band / in-band to power on and off the faulty drive slots on the backplane. This can be achieved by adding an eFuse circuit to the backplane to enable independent power-on / off of a single hard drive. The BMC controls the backplane CPLD via I2C to control the voltage of the eFuse, thereby resetting the hard drive's level. This triggers boot code carrying backplane slot insertion / removal instructions; the backplane slot insertion / removal instructions in the boot code are used to insert or remove the drive slots in the server cluster hard drive backplane memory, thus repairing the server cluster hard drives. Furthermore, it should be noted that, in conjunction with the previous embodiments, since server cluster systems have many types of hard drives, when a server cluster hard drive does not support the backplane memory power-on / removal function but also supports voltage reset, a prompt message needs to be issued to inform maintenance personnel of the hard drive location identifier in the server cluster, enabling manual insertion / removal.

[0133] In some embodiments of the present invention, the method further includes:

[0134] The drive letter positions of the server cluster hard drives are detected; when a drive letter misalignment is detected, the initial configuration information of the server cluster hard drive's slot, drive letter, and mount point is determined; based on the initial configuration information of the server cluster hard drive's slot, drive letter, and mount point, the misaligned drive letter is adjusted. Among these, [the following is a separate section, likely related to a specific process or feature]... Figure 7 , Figure 7 This is an optional hard drive repair schematic diagram in an embodiment of the present invention. Specifically, after the service is deployed on the server cluster system, a hard drive (slot -> drive letter -> mount point) configuration table can be collected and recorded as an initial snapshot. Specifically, the slot is the physical location information of the hard drive and does not change when the hard drive is replaced; the drive letter is assigned by the kernel according to rules and may change when the hard drive is replaced; the mount point is the mount directory actually used by the upper layer of the service and is not allowed to change. When the kernel detects a new block device, the monitoring system obtains the drive letter and slot information of the new device and verifies and compares it with the snapshot. When it is confirmed that a drive letter drift has occurred, hard drive repair is triggered. Specifically, it includes the following steps:

[0135] Step 701: The server cluster system begins deploying business information.

[0136] Step 702: Determine the server slot, drive letter, and mount point for the initial snapshot.

[0137] Step 703: Has a new block device been detected? If yes, proceed to step 706; otherwise, proceed to step 704.

[0138] Step 704: Obtain the drive letter / slot information for the new block device.

[0139] Step 705: Compare the acquired new block device drive letter / slot information with the snapshot.

[0140] Step 706: Determine if the information is consistent. If yes, end the execution. If no, proceed to step 707.

[0141] Step 707: Drift drive mount point matching.

[0142] Step 708: Remount the drive letter.

[0143] Therefore, based on the slot-drive-mount point correspondence in the configuration snapshot, the mount point corresponding to the drifting drive is obtained. After unmounting the original mount point, it is automatically mounted. If successful, the repair process ends; if unsuccessful, systemctl daemon-reload is issued and the drive is remounted. If it still fails, manual intervention is notified to avoid premature manual intervention in the repair of faulty hard drives in the server cluster and to save the operating costs of the server cluster.

[0144] Continue to refer to Figure 8 The following uses a server cluster as an example of a cloud server environment, illustrating the server cluster hard drive failure handling method provided by this invention through a scenario of alternating use of financial transaction data stored on the cloud server's hard drive. The user then... Figure 1 The terminals shown (including terminals 10-1 and 10-2) obtain the stored financial resources, such as fund and stock transaction data, from the corresponding cloud server 200 via network 300.

[0145] Among them, see Figure 8 , Figure 8 This is a front-end display diagram of the server cluster hard disk failure handling method provided in this application, wherein the terminal (e.g. Figure 1 The terminal 10-1 in the cloud server is equipped with a server cluster client or server cluster runtime plugin that can display relevant financial information. Users can use the corresponding client to store financial data from banks, securities firms, internet finance companies, P2P platforms, and other entities providing payment, lending, and wealth management services on the cloud server. The cloud server's management terminal (e.g., Figure 1 Terminal 10-2 in the middle) through Figure 8The diagram illustrates the front-end display of a cloud server hard drive fault handling method. It detects the operating status of the cloud server hard drive, specifically by displaying a user interface. This user interface includes a fixed-perspective view of the cloud server's operating environment, featuring different cloud server identifiers. The user interface also includes an editing detection component and a repair component. The detection component monitors the operating parameters of the cloud server hard drive and triggers a fault warning message when these parameters reach a warning threshold. Based on the detection results from the hardware fault detection process, the repair component repairs the cloud server hard drive fault.

[0146] Among them, reference Figure 9 , Figure 9 The process diagram for the server cluster hard disk failure handling method provided in this application specifically includes:

[0147] Step 901: Receive hard drive failure warning information.

[0148] Step 902: Determine the fault type.

[0149] Step 903: Identify the serial number of the faulty hard drive.

[0150] Step 904: Comprehensive health assessment of the faulty hard drive.

[0151] Step 905: Has the health threshold of the faulty hard drive reached the threshold? If yes, proceed to step 906; otherwise, manually replace the hard drive.

[0152] Step 906: Trigger online NTF recovery for hard drive.

[0153] Step 907: Determine if PIN3 Reset is supported. If yes, proceed to step 908; otherwise, proceed to step 909.

[0154] Step 908: Repair via hard drive PIN3 Reset.

[0155] Step 909: Determine if eFuseReset is supported. If yes, proceed to step 910; otherwise, proceed to step 911.

[0156] Step 910: Execute the backplane eFuse Reset process.

[0157] Step 911: Perform hard drive unplugging and plugging.

[0158] Step 912: Determine if drive letter drift has occurred. If yes, proceed to step 913; otherwise, proceed to step 914.

[0159] Step 913: Repair the mount point.

[0160] Step 914: Determine whether the faulty hard drive can be repaired. If not, proceed to step 915.

[0161] Step 915: Replace the server hard drive.

[0162] Furthermore, Figure 10 This diagram illustrates the front-end display of the server cluster hard drive failure handling method provided in this application. The detection component calls the interface of the target cloud server hard drive; based on the repair component, a query command is sent through the target cloud server hard drive's interface to obtain the initial configuration information of the target cloud server hard drive's slot, drive letter, and mount point via the repair component. In implementing the cloud server hard drive failure handling process of this application, the information displayed on the interface allows monitoring of the cloud server hard drive repair process, avoiding premature manual intervention in the repair of failed cloud server hard drives, saving cloud server operating costs, and ensuring the security of users' financial data stored on the cloud server, reducing the risk of data loss.

[0163] Beneficial technical effects:

[0164] This invention acquires fault warning information from server cluster hard drives; responds to the fault warning information by retrieving the operational data of the server cluster hard drives; determines the type of server cluster hard drive fault based on the operational data; when the fault type is an unavailability alarm, triggers a matching hardware-free fault detection process and acquires the detection results of the hardware-free fault detection process; and repairs the server cluster hard drive fault based on the detection results. This allows for automated detection and repair of server cluster hard drive fault types, reducing the replacement rate of server cluster hard drives, lowering the operating costs of the server cluster system, improving the efficiency of server cluster hard drive maintenance, ensuring data security for server cluster users, and enhancing the user experience.

[0165] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for handling hard disk failures in a server cluster, characterized in that, The method includes: Obtain fault warning information for the server cluster hard drives; In response to the fault warning information of the server cluster hard disk, retrieve the operating data of the server cluster hard disk; The operating data of the server cluster hard disk is cross-compared with the fault type characteristics corresponding to the hard disk attribute parameters of the server cluster hard disk to obtain the fault type of the server cluster hard disk. The fault types of the server cluster hard disk include: unavailability alarm and failure alarm. When the type of hard disk failure in the server cluster is the unavailable alarm, a matching hardware failure detection process is triggered, and the detection parameters of the hardware failure detection process are obtained. The detection parameters matching the hard drives of the server cluster are processed in multiple dimensions to obtain the parameter health detection results, distribution characteristic detection results, deterioration trend results, and failure probability results of the hard drive protection parameters; The health detection results of the parameters, the distribution feature detection results, the deterioration trend results, and the failure probability results are fused to determine the detection results of the server cluster hard disk in different dimensions. The detection results of the server cluster hard disk in different dimensions are fused to determine the detection result of the no-hardware-fault detection process; Based on the detection results of the hardware fault detection process, a fault repair method matching the hard drive of the server cluster is determined. Based on the determined fault repair method, the boot code is invoked to perform a hardware-level reset repair of the server cluster hard disk failure; the hardware-level reset repair includes at least one of the following: When the fault repair method is a voltage reset process, a boot code carrying a voltage reset instruction is triggered. The boot code is used to adjust the loading voltage of the server cluster hard disk so as to perform a voltage reset of the server cluster hard disk without powering the server back on. When the fault repair method is the backplane slot plugging and unplugging process, a boot code carrying backplane slot plugging and unplugging instructions is triggered. The boot code is used to perform logical power-off and power-on operations on the backplane storage slots of the server cluster hard disk to simulate the physical plugging and unplugging process to repair the fault of the server cluster hard disk. When the server cluster hard drive experiences a drive letter position shift due to the repair, the drive letter with the position shift is adjusted based on the initial configuration information of the server cluster hard drive's slot, drive letter, and mount point to restore the server cluster hard drive's mount.

2. The method according to claim 1, characterized in that, The process of obtaining fault warning information for the server cluster hard disks includes: The system monitors the operating parameters of the server cluster hard drives. When the operating parameters of the server cluster hard drives reach a warning threshold, it triggers a fault warning message for the server cluster hard drives. Alternatively, it receives alarm messages from server cluster users and, based on the parameters of the server cluster users, determines the server cluster hard drives that match the server cluster users and triggers a fault warning message for the server cluster hard drives.

3. The method according to claim 1, characterized in that, The step of cross-comparing the operational data of the server cluster hard disk with the fault type characteristics corresponding to the hard disk attribute parameters of the server cluster hard disk to obtain the type of server cluster hard disk fault includes: Determine the hard drive attribute parameters of the server cluster hard drive, wherein the hard drive attribute parameters include: hard drive model, server model, rack deployment time, version number, and hard drive partition identifier; Based on the hard disk attribute parameters of the server cluster hard disk, retrieve the operating data of the server cluster hard disk stored in the corresponding storage medium. Based on the hard drive attribute parameters of the server cluster hard drives, obtain the corresponding fault type characteristics; The type of hard disk failure in the server cluster is determined by cross-referencing the operating data of the server cluster hard disk with the characteristics of the failure type.

4. The method according to claim 1, characterized in that, The detection parameters matching the hard drives of the server cluster are processed in multiple dimensions to obtain the parameter health detection results, distribution characteristic detection results, deterioration trend results, and failure probability results of the hard drive protection parameters, including: Based on the hard disk protection parameters corresponding to the hard disks in the server cluster, determine the corresponding parameter health detection results; By standardizing the hard disk protection parameters, the distribution characteristics of the hard disk protection parameters are determined. By processing the dynamic slope of the hard disk protection parameters, the deterioration trend of the hard disk protection parameters is determined. By processing the prediction function of the hard disk protection parameters, the failure probability result corresponding to the hard disk protection parameters is determined.

5. The method according to claim 1, characterized in that, The method further includes: The user interface includes a perspective view of the server cluster operating environment from a fixed perspective, and the user interface includes different server cluster identifiers. The user interface also includes an editing detection component and a repair component; The detection component monitors the operating parameters of the server cluster hard disk and triggers a fault warning message for the server cluster hard disk when the operating parameters reach the warning threshold. Based on the repair component and the detection results of the hardware fault detection process, the hard disk faults of the server cluster are repaired.

6. The method according to claim 5, characterized in that, The method further includes: The detection component calls the interface of the target server cluster hard drive; Based on the repair component, a query command is sent through the interface of the target server cluster hard drive to obtain the initial configuration information of the target server cluster hard drive slot, drive letter, and mount point through the repair component.

7. A server cluster hard disk failure handling device, characterized in that, The device includes: The information transmission module is used to obtain fault warning information of the server cluster hard disk; The information processing module is used to respond to the fault warning information of the server cluster hard disk and retrieve the operating data of the server cluster hard disk; The information processing module is used to cross-compare the operating data of the server cluster hard disk with the fault type characteristics corresponding to the hard disk attribute parameters of the server cluster hard disk to obtain the fault type of the server cluster hard disk. The fault types of the server cluster hard disk include: unavailability alarm and failure alarm. The information processing module is used to trigger a matching hardware failure detection process when the type of hard disk failure in the server cluster is the unavailability alarm, and to obtain the detection parameters of the hardware failure detection process; and to perform multi-dimensional processing on the detection parameters matching the hard disk of the server cluster to obtain the parameter health detection results, distribution characteristic detection results, deterioration trend results and failure probability results of the hard disk protection parameters. The health detection results of the parameters, the distribution feature detection results, the deterioration trend results, and the failure probability results are fused to determine the detection results of the server cluster hard disk in different dimensions. The detection results of the server cluster hard disk in different dimensions are fused to determine the detection result of the no-hardware-fault detection process; The information processing module is used to determine the fault repair method matching the hard disk of the server cluster based on the detection results of the no-hardware-fault detection process. Based on the determined fault repair method, the boot code is invoked to perform a hardware-level reset repair of the server cluster hard disk failure; the hardware-level reset repair includes at least one of the following: When the fault repair method is a voltage reset process, a boot code carrying a voltage reset instruction is triggered. The boot code is used to adjust the loading voltage of the server cluster hard disk so as to perform a voltage reset of the server cluster hard disk without powering the server back on. When the fault repair method is the backplane slot plugging and unplugging process, a boot code carrying backplane slot plugging and unplugging instructions is triggered. The boot code is used to perform logical power-off and power-on operations on the backplane storage slots of the server cluster hard disk to simulate the physical plugging and unplugging process to repair the fault of the server cluster hard disk. When the server cluster hard drive experiences a drive letter position shift due to the repair, the drive letter with the position shift is adjusted based on the initial configuration information of the server cluster hard drive's slot, drive letter, and mount point to restore the server cluster hard drive's mount.

8. The apparatus according to claim 7, characterized in that, The device further includes: The display module is used to display the user interface, which includes a perspective view of the server cluster operating environment from a fixed perspective, and the user interface includes different server cluster identifiers. The user interface also includes an editing detection component and a repair component; The display module is used to monitor the operating parameters of the server cluster hard disk through the detection component, and trigger a fault warning message for the server cluster hard disk when the operating parameters of the server cluster hard disk reach the warning threshold. Based on the repair component and the detection results of the hardware fault detection process, the hard disk faults of the server cluster are repaired.

9. The apparatus according to claim 7, characterized in that, The information transmission module is specifically used for: Monitor the operating parameters of the server cluster hard disks, and trigger a fault warning message for the server cluster hard disks when the operating parameters of the server cluster hard disks reach the warning threshold; Alternatively, it can receive alarm information from server cluster users and, based on the parameters of the server cluster users, determine the server cluster hard disk that matches the server cluster users, and trigger fault warning information for the server cluster hard disk.

10. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the server cluster hard disk failure handling method according to any one of claims 1 to 6.

11. A computer-readable storage medium storing executable instructions, characterized in that, When the executable instructions are executed by the processor, they implement the server cluster hard disk failure handling method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Failing disk detection and restoration method and failing disk detection and restoration device

    CN107577545A

  • Ceph-based faulted hard disk processing method and apparatus

    CN107832164A

  • Method, device, and apparatus for hard disk maintenance, and readable storage medium

    CN108845760A

  • Hard disk health degree evaluation method and device

    CN111400122A