Server cluster hard disk status detection methods, devices, electronic equipment, and storage media

CN111897696BActive Publication Date: 2026-08-14TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-05
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

这些实例类型由 CPU、内存、存储和网络组成不同的组合,但是当云服务器的硬盘在运行过程中会发生掉线、只读等问题后,用户部署在该服务器上的子机业务也会受到影响

Benefits of technology

本发明实施例通过监听服务器集群硬盘的运行数据;基于所述服务器集群硬盘的运行数据,通过触发相匹配的状态检测进程,获取与所述服务器集群硬盘相匹配的检测参数;对所述服务器集群硬盘相匹配的检测参数进行多维度处理,确定所述服务器集群硬盘在不同维度中的检测结果;对所述服务器集群硬盘在不同维度中的检测结果进行融合处理,确定所述状态检测进程的检测结果。由此,能够自动化的实时对服务器集群硬盘的故障类型进行检测,减少服务器集群硬盘的更换率,降低云服务器系统的运行成本并提升对服务器集群硬盘维护的效率,保证云服务器用户的数据安全,提高用户的使用体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111897696B_ABST
    Figure CN111897696B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, electronic device, and storage medium for server cluster hard drive status detection. The method includes: monitoring the operational data of the server cluster hard drives; based on the operational data, triggering a matching status detection process to obtain detection parameters matching the server cluster hard drives; performing multi-dimensional processing on the matching detection parameters to determine the detection results of the server cluster hard drives in different dimensions; and fusing the detection results of the server cluster hard drives in different dimensions to determine the detection result of the status detection process. This enables automated, real-time detection and prediction of server cluster hard drive fault types, reducing the replacement rate of server cluster hard drives, lowering the operating costs of cloud server systems, improving the efficiency of server cluster hard drive maintenance, ensuring data security for cloud server users, and enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to hard disk system fault detection and processing technology, and more particularly to server cluster hard disk status detection methods, devices, electronic equipment and storage media. Background Technology

[0002] With the continuous development of computer technology, cloud servers (Cloud Virtual Machines, CVMs) can provide secure and reliable elastic computing services, and offer different instance types to meet specific user scenarios. These instance types consist of different combinations of CPU, memory, storage, and network. However, when a cloud server's hard drive experiences issues such as disconnection or read-only access during operation, the services of the user's sub-machines deployed on that server will also be affected. Current technologies can only determine the hard drive's status based on certain parameters in the hard drive's SMART information when a problem occurs, impacting the data security of users accessing the cloud server. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for detecting the status of server cluster hard drives. These methods can automatically and in real-time detect and predict the fault types of server cluster hard drives, reducing the replacement rate of server cluster hard drives, lowering the operating costs of cloud server systems, improving the efficiency of server cluster hard drive maintenance, ensuring data security for cloud server users, and enhancing the user experience. The technical solution of these embodiments is implemented as follows: This invention provides a method for detecting the hard disk status of a server cluster, the method comprising: Monitor the running data on the server cluster's hard drives; Based on the operating data of the server cluster hard disk, a matching status detection process is triggered to obtain detection parameters that match the server cluster hard disk; The detection parameters matching the hard disks of the server cluster are processed in multiple dimensions to determine the detection results of the hard disks of the server cluster in different dimensions. The detection results of the server cluster hard disks in different dimensions are fused to determine the detection result of the status detection process.

[0004] This invention also provides a server cluster hard disk status detection device, comprising: The information transmission module is used to monitor the operating data of the server cluster hard disk; The information processing module is used to obtain detection parameters that match the server cluster hard disk by triggering a matching status detection process based on the operating data of the server cluster hard disk. The information processing module is used to perform multi-dimensional processing on the detection parameters that match the hard disk of the server cluster, and determine the detection results of the hard disk of the server cluster in different dimensions. The information processing module is used to fuse the detection results of the server cluster hard disk in different dimensions to determine the detection result of the status detection process.

[0005] This invention also provides an electronic device, the electronic device comprising: Memory, used to store executable instructions; The processor, when running executable instructions stored in the memory, implements the preceding server cluster hard disk status detection method.

[0006] This invention also provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement a preceding method for detecting the hard disk status of a server cluster.

[0007] The embodiments of the present invention have the following beneficial effects: This invention, in its embodiments, monitors the operational data of the server cluster hard drives. Based on this data, a matching status detection process is triggered to obtain detection parameters corresponding to the server cluster hard drives. These parameters are then processed in multiple dimensions to determine the detection results across different dimensions. Finally, the detection results from these different dimensions are fused to determine the final result of the status detection process. This allows for automated, real-time detection of server cluster hard drive fault types, reducing hard drive replacement rates, lowering cloud server system operating costs, improving hard drive maintenance efficiency, ensuring data security for cloud server users, and enhancing the user experience. Attached Figure Description

[0008] Figure 1 This is a schematic diagram illustrating a usage scenario of the server cluster hard disk status detection method provided in an embodiment of the present invention. Figure 2 A schematic diagram of the composition structure of an electronic device provided in an embodiment of the present invention; Figure 3 This is an optional flowchart illustrating the server cluster hard disk status detection method provided in an embodiment of the present invention. Figure 4 This is an optional flowchart illustrating the server cluster hard disk status detection method provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of an optional state calculation in an embodiment of the present invention; Figure 6This is a schematic diagram of an optional state calculation in an embodiment of the present invention; Figure 7 This is a schematic diagram of an optional state calculation in an embodiment of the present invention; Figure 8 This is a schematic diagram of an optional state calculation in an embodiment of the present invention; Figure 9 This is a schematic diagram of an optional state calculation in an embodiment of the present invention; Figure 10 A schematic diagram of the front-end display of the server cluster hard disk status detection method provided in this application; Figure 11 A data architecture diagram of the server cluster hard disk status detection method provided in this application; Figure 12 A schematic diagram of the front-end display of the server cluster hard disk status detection method provided in this application; Figure 13 A schematic diagram illustrating the processing effect of the server cluster hard disk status detection method provided in this application; Figure 14 A schematic diagram illustrating the processing effect of the server cluster hard disk status detection method provided in this application; Figure 15 This is a schematic diagram illustrating the processing effect of the server cluster hard disk status detection method provided in this application. Detailed Implementation

[0009] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0010] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0011] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0012] Before providing a further detailed description of the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention will be explained, and the nouns and terms involved in the embodiments of the present invention shall be interpreted as follows.

[0013] 1) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0014] 2) Terminals, including but not limited to: ordinary terminals and dedicated terminals, wherein the ordinary terminals maintain a long connection and / or a short connection with the transmission channel, and the dedicated terminals maintain a long connection with the transmission channel.

[0015] 3) Client: The carrier that implements specific functions in the terminal. For example, a mobile client (APP) is a carrier of specific functions in a mobile terminal, such as performing report creation or report display functions.

[0016] 4) Firmware: This is the code running inside the chip, which is binary code used to implement hard disk failure detection in server clusters.

[0017] 5) A mini program is a type of application developed using a front-end-oriented language (such as JavaScript) and running services within Hyper Text Markup Language (HTML) pages. It is downloaded by a client (such as a browser or any client with an embedded browser engine) via a network (such as the internet) and interpreted and executed within the client's browser environment, saving the step of installation on the client side. For example, mini programs can be downloaded and run on social networking clients to provide various services such as airline ticket purchase, report generation, and data display.

[0018] 6) Runtime environment, the engine used to interpret and execute code. For example, for the runtime environment of a mini-program, it could be JavaScript Core on the iOS platform or X5 JS Core on the Android platform.

[0019] 7) Bootloader code: also known as bootloader, boot mode, startup load, etc., refers to a type of code that runs when the chip starts up. It is usually used to initialize the hardware environment and load the firmware code to run. Usually, it does not need to be updated.

[0020] 8) Components are functional modules of the view in a mini-program, also known as front-end components. They include buttons, titles, tables, sidebars, content, and footers on the page. Components include modular code so that they can be reused in different pages of the mini-program.

[0021] 9) A server cluster refers to a group of servers working together to provide the same service, appearing to the client as a single server. Server clusters can utilize multiple computers for parallel computing to achieve high computing speeds, and can also use multiple computers for backup, ensuring the entire system continues to operate normally even if one machine fails. The server cluster hard disk failure handling method provided in this application can be applied to cloud server and distributed server scenarios, enabling status detection and fault repair of server hard disks in different usage scenarios. Specifically, a cloud server (CVM Cloud Virtual Machine) is a simple, efficient, secure, reliable, and elastically scalable computing service. Its management is simpler and more efficient than traditional single physical servers. Users can quickly create or release any number of cloud servers for their business processes and store user data without purchasing hardware in advance. In a distributed server environment, user data and programs may not reside on a single server but are distributed across multiple servers. Similarly, distributed server environments also require a large number of hard disks, necessitating the server hard disk status detection and fault repair methods provided in this application.

[0022] Figure 1 This is a schematic diagram illustrating a use case of the server cluster hard disk status detection method provided in this embodiment of the invention. (See attached diagram.) Figure 1With the continuous development of computer technology, cloud servers (Cloud Virtual Machines, CVMs) can provide secure and reliable elastic computing services, and can also provide different instance types to meet specific user scenarios. Terminals (including terminals 10-1 and 10-2) are equipped with corresponding clients capable of performing different functions. These clients (including terminals 10-1 and 10-2) obtain different information from the corresponding cloud server 200 via network 300 and can deploy different services on the cloud server. The terminals connect to the cloud server 200 via network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both, using a wireless link for data transmission. The instance types provided by the cloud server consist of different combinations of CPU, memory, storage, and network, and store user business data on the cloud server's hard drive. However, if the cloud server's hard drive experiences issues such as disconnection or read-only access, the user's deployed sub-machine services on that server will also be affected. In the embodiments provided by this invention, the cloud server application running on the cloud server 200 can be written in software code environments using different programming languages, and the code objects can be different types of code entities. For example, in C language software code, a code object can be a function. In Java language software code, a code object can be a class, and in iOS Objective-C, it can be a piece of object code. In C++ language software code, a code object can be a class or a function to execute processing instructions from different terminals. This application does not distinguish the source of the cloud server's compilation environment.

[0023] The structure of the server cluster hard disk status detection device according to an embodiment of the present invention will be described in detail below. The server cluster hard disk status detection device can be implemented in various forms, such as a dedicated terminal with server cluster hard disk status detection device processing function, or a server equipped with server cluster hard disk status detection device processing function, such as the aforementioned... Figure 1 The cloud server in the middle is 200. Figure 2 This is a schematic diagram of the composition structure of an electronic device provided in an embodiment of the present invention. It can be understood that... Figure 2 This is merely an exemplary structure of a server cluster hard drive status detection device in an electronic device, and not the entire structure. Implementation can be carried out as needed. Figure 2 The structure shown may be part or all of the structure.

[0024] The electronic device provided in this embodiment of the invention includes: at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. The various components in the server cluster hard disk status detection device are coupled together via a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 205.

[0025] The user interface 203 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.

[0026] It is understood that memory 202 can be volatile memory or non-volatile memory, or both. In this embodiment of the invention, memory 202 is capable of storing data to support the operation of a terminal (such as 10-1). Examples of this data include any computer programs used to operate on the terminal (such as 10-1), such as operating systems and applications. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications.

[0027] In some embodiments, the server cluster hard disk status detection device provided in this invention can be implemented using a combination of hardware and software. For example, the server cluster hard disk status detection device provided in this invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the server cluster hard disk status detection method provided in this invention. For instance, the hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0028] As an example of the server cluster hard disk status detection device provided in this embodiment of the invention, which is implemented by combining software and hardware, the server cluster hard disk status detection device provided in this embodiment of the invention can be directly embodied as a combination of software modules executed by processor 201. The software modules can be located in the storage medium, which is located in memory 202. Processor 201 reads the executable instructions included in the software modules in memory 202 and combines them with necessary hardware (e.g., including processor 201 and other components connected to bus 205) to complete the server cluster hard disk status detection method provided in this embodiment of the invention.

[0029] As an example, processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0030] As an example of the hardware implementation of the server cluster hard disk status detection device provided in this embodiment of the invention, the device provided in this embodiment of the invention can be directly executed by a processor 201 in the form of a hardware decoding processor. For example, it can be executed by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components to implement the server cluster hard disk status detection method provided in this embodiment of the invention.

[0031] In this embodiment of the invention, the memory 202 is used to store various types of data to support the operation of the server cluster hard disk status detection device. Examples of such data include: any executable instructions for operation on the server cluster hard disk status detection device, such as executable instructions, and a program implementing the server cluster hard disk status detection method of this embodiment of the invention may be included in the executable instructions.

[0032] In other embodiments, the server cluster hard disk status detection device provided in this invention can be implemented in software. Figure 2A server cluster hard disk status detection device stored in memory 202 is shown. This device can be software in the form of programs and plug-ins, and includes a series of modules. As an example of a program stored in memory 202, it may include the server cluster hard disk status detection device. The server cluster hard disk status detection device includes the following software modules: an information transmission module 2081 and an information processing module 2082. When the software modules in the server cluster hard disk status detection device are read into RAM and executed by processor 201, the server cluster hard disk status detection method provided in this embodiment of the invention will be implemented. The functions of each software module in the server cluster hard disk status detection device include: The information transmission module 2081 is used to monitor the operating data of the server cluster hard disk; The information processing module 2082 is used to obtain detection parameters that match the server cluster hard disk by triggering a matching status detection process based on the operating data of the server cluster hard disk. The information processing module 2082 is used to perform multi-dimensional processing on the detection parameters that match the hard disk of the server cluster, and determine the detection results of the hard disk of the server cluster in different dimensions. The information processing module 2082 is used to fuse the detection results of the server cluster hard disk in different dimensions to determine the detection result of the status detection process.

[0033] As described in the preceding embodiments, the related technologies for monitoring server hard drive status only utilize single-dimensional information for calculation, such as calculations based on certain parameters of the hard drive's SMART algorithm. Hard drive health is quantified by setting thresholds based on parameter values; or by performing read / write operations on the hard drive to determine whether it can be read and written normally. However, the approach of using only the current SMART parameters as the basis for judging hard drive status has significant limitations. On the one hand, SMART only contains some shallow parameters and lacks a model for building underlying parameters. Relying solely on SMART parameter snapshots cannot accurately assess the hard drive's status, easily leading to misjudgments. On the other hand, this approach lacks dynamic modeling and horizontal comparison with other hard drives in the same cluster. The algorithm lacks comparison logic (including comparisons between the target cloud server's hard drive and other hard drives within the same cloud server, and also lacks comparisons of changes within the same server cluster itself), resulting in inaccurate evaluation and hindering accurate judgment of the server cluster's hard drive status.

[0034] Furthermore, adding read / write operations to determine disk status is unsuitable for most cloud server business scenarios. Firstly, read / write operations focus on a short area, only determining if the hard drive can execute read / write commands, but not identifying internal problems such as platter scratches or head degradation. Covering the entire disk with read / write operations would be extremely time-consuming and impractical. Secondly, data centers operate under heavy loads, prohibiting background read / write operations, especially write operations, which could lead to data loss or severely impact business performance. This solution is not applicable in cloud server business environments.

[0035] To overcome the above-mentioned shortcomings, refer to Figure 3 This application provides a method for detecting the status of hard disks in a server cluster. To overcome the aforementioned shortcomings, embodiments of this invention provide a method for detecting the status of hard disks in a server cluster. See [link to relevant documentation]. Figure 3 , Figure 3 This is an optional flowchart illustrating the server cluster hard disk status detection method provided in this embodiment of the invention. It can be understood that... Figure 3 The steps shown can be performed by various electronic devices running a server cluster hard drive status detection device, such as mobile phones or tablets with server cluster hard drive status detection capabilities. The dedicated terminal with the server cluster hard drive status detection device can be packaged within... Figure 1 In the terminal 10-1 shown, to execute the preprocessor. Figure 2 The corresponding software module in the server cluster hard disk status detection device of the electronic device shown. The following addresses... Figure 3 The steps shown are explained.

[0036] Step 301: The server cluster hard disk status detection device monitors the operating data of the server cluster hard disks.

[0037] In this invention, embodiments can be implemented using cloud technology. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. It can also be understood as a general term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies based on cloud computing business models. The backend services of network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites; therefore, cloud technology needs cloud computing as its support.

[0038] It's important to note that cloud computing is a computing model that distributes computing tasks across a resource pool comprised of numerous computers, enabling various application systems to access computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, resources in the "cloud" appear infinitely scalable, readily available, on-demand, and expandable, with payment based on usage. As a provider of fundamental cloud computing capabilities, a cloud resource pool platform, often referred to as a cloud platform or Infrastructure as a Service (IaaS), is established. This platform deploys various types of virtual resources within the resource pool for external customers to choose from. The cloud resource pool primarily includes: computing devices (which can be virtualized machines containing operating systems), storage devices, and network devices. When users use cloud servers to store data or deploy different application processes, monitoring the operating parameters of the server cluster hard drives allows for timely detection of potential server cluster hard drive failures, preventing data loss due to server cluster hard drive failures that trigger failure warnings.

[0039] Step 302: The server cluster hard disk status detection device obtains detection parameters that match the server cluster hard disk by triggering a matching status detection process based on the operating data of the server cluster hard disk.

[0040] Among them, SMART (Self-Monitoring Analysis and Reporting Technology) is an automatic hard drive status detection and early warning system and standard. It monitors and records the operation of hard drive hardware components such as read / write heads, platters, motors, and circuitry using detection commands within the hard drive hardware, comparing this information with preset safety values ​​set by the manufacturer. If the monitored status is about to or has exceeded the preset safety range, the host monitoring hardware or software can automatically warn the user and perform minor automatic repairs to proactively protect hard drive data security. SMART uses binary code as its basic instructions, which are written into standard registers to form a specific SMART information table for normal detection and operation. SMART instructions are divided into main instructions (Commands) and subcommands. The main instruction primarily provides information on whether the device supports SMART or to ignore certain instruction characteristics. Subcommands provide detection information supporting SMART devices. By acquiring detection parameters matching the server cluster hard drives, real-time detection of the server cluster hard drive's operating status can promptly identify high-risk server cluster hard drives, allowing for advance deployment of backup hardware and reducing data loss caused by server cluster hard drive failures.

[0041] Step 303: The server cluster hard disk status detection device performs multi-dimensional processing on the detection parameters matching the server cluster hard disk to determine the detection results of the server cluster hard disk in different dimensions.

[0042] Continue to refer to Figure 4 , Figure 4 This is an optional flowchart illustrating the server cluster hard disk status detection method provided in this embodiment of the invention. It can be understood that... Figure 4 The steps shown can be performed by various electronic devices running a server cluster hard drive status detection device, such as mobile phones or tablets with server cluster hard drive status detection capabilities. The dedicated terminal with the server cluster hard drive status detection device can be packaged within... Figure 1 In the terminal 10-1 shown, to execute the preprocessor. Figure 2 The corresponding software module in the server cluster hard disk status detection device of the electronic device shown. The following addresses... Figure 4 The steps shown are explained.

[0043] Step 401: Based on the hard disk protection parameters corresponding to the hard disks in the server cluster, determine the corresponding parameter health detection results.

[0044] Unlike conventional SMART assessment methods, the server cluster hard drive failure detection method provided in this application can calculate the hard drive health score from multiple dimensions using different weighted algorithms. These include: the SMART parameter Euclidean distance algorithm (for measuring the health of key SMART parameters); the SMART parameter z-score algorithm (for statistically quantifying the distribution of hard drive parameters within the cluster); the SMART parameter dynamic slope algorithm (for dynamically quantifying the deterioration trend of parameters); and a hard drive underlying parameter machine learning failure prediction algorithm (developing parameters reflecting the health of the hard drive internally in cooperation with hard drive suppliers and performing big data machine learning). The weighted calculation of the hard drive health score refers to Formula 1. In different server cluster usage environments, operations and maintenance personnel can dynamically adjust the different weights in Formula 1 according to the business type and environmental parameters. In a cloud server usage environment, the optional values ​​for the weights are: a0=0.2, a1=0.2, a2=0.3, a3=0.3. Formula 1 is: Formula 1 Step 402: Determine the distribution characteristics detection results of the hard disk protection parameters by standardizing the hard disk protection parameters.

[0045] In some embodiments of the present invention, the distribution characteristic detection results of the hard disk protection parameters are determined by standardizing the hard disk protection parameters, which can be achieved in the following ways: A target value parameter matching the hard disk protection parameters is determined; the offset of the hard disk protection parameters relative to the target parameter value is determined through standardization of the hard disk protection parameters; based on the offset of the hard disk protection parameters relative to the target parameter value, the distribution characteristic detection result of the hard disk protection parameters is determined. Wherein, reference... Figure 5 , Figure 5 This is an optional state calculation diagram in an embodiment of the present invention. Combining with the preceding formula 1, a0, a1, a2, and a3 are the weighted values ​​for four dimensions. Simultaneously, `basescore` is the basic score for the SMART parameter. The Euclidean distance algorithm is used to calculate the eccentricity value (target parameter value is `targetparam`) of the key SMART basic parameter (`baseparam`). The greater the distance from the target value `targetparam`, the larger the calculated weighted value. The `basescore` calculation references formula 2: Formula 2 Step 403: By processing the dynamic slope of the hard disk protection parameters, determine the deterioration trend of the hard disk protection parameters. Wherein, refer to... Figure 6 , Figure 6 This is an optional state calculation diagram in an embodiment of the present invention; zscore is the statistical score of the SMART parameter distribution. The statistical z-distribution value is used to calculate whether the disk's distribution is offset within the entire cluster. The z-distribution value reflects the offset of the key SMART parameter (param) relative to the target value (u), and is thus quantified as a statistical score. The statistical z-distribution value is as follows... Figure 6 As shown. The zscore calculation is based on formula 3: Formula 3 In some embodiments of the present invention, the deterioration trend of the hard disk protection parameters is determined by dynamic slope processing of the hard disk protection parameters, which can be achieved in the following ways: By processing the dynamic slope of the hard disk protection parameters, the dynamic slope change of the hard disk protection parameters within a single sampling period is determined; based on the dynamic slope change of the hard disk protection parameters within a single sampling period and a matching slope threshold, the deterioration trend of the hard disk protection parameters is determined. Wherein, reference... Figure 7 , Figure 7This is an optional state calculation diagram in an embodiment of the present invention. The dynamicscore is the dynamic trend score of the SMART parameter. A dynamic slope algorithm is used to model the deterioration of key SMART parameters. Within a sampling period t∈[t1, t2], if the parameter dynamic slope k is greater than a predetermined threshold, the average growth amplitude of the parameter within that sampling period is calculated, and the sum of squares and square roots of multiple parameters are taken. The dynamic slope algorithm is as follows: Figure 7 As shown, the dynamicscore calculation is based on formula 4: Formula 4 Step 404: By processing the prediction function of the hard disk protection parameters, determine the failure probability result corresponding to the hard disk protection parameters.

[0046] In some embodiments of the present invention, the failure probability result corresponding to the hard disk protection parameters is determined by processing the prediction function of the hard disk protection parameters, which can be achieved in the following ways: The disk attribute parameters of the server cluster hard drives are determined, including: hard drive model, server model, rack installation time, version number, and hard drive partition identifier. Based on the disk attribute parameters of the server cluster hard drives, a prediction function corresponding to the hard drive protection parameters is determined. Based on the operating data of the server cluster hard drives stored in the storage medium, the failure probability result corresponding to the hard drive protection parameters is determined through the prediction function corresponding to the hard drive protection parameters. Since cloud server systems use a large number of hard drives, these server cluster hard drives may come from different hardware manufacturers or be customized hard drives by cloud server operators. Therefore, by cross-referencing the hard drive model, server model, rack installation time, version number, and hard drive partition identifier with failure type characteristics, the failure type can be more accurately determined, avoiding false positives and false negatives caused by inconsistent hard drive device versions. Specifically, refer to... Figure 8 and Figure 9 , Figure 8 This is a schematic diagram of an optional state calculation in an embodiment of the present invention. Figure 9 This is an optional state calculation diagram in an embodiment of the present invention; wherein, predictionscore is the probability score of hard disk failure. Cloud server operators can develop customized big data machine learning algorithms through in-depth cooperation with hard disk suppliers. By modeling the internal parameters of the hard disk, the probability of hard disk failure within a specific time period in the future can be deduced. By weighted quantization of this probability value, a risk prediction score can be obtained.

[0047] After determining the detection results of the server cluster hard disks in different dimensions, continue to step 304.

[0048] Step 304: The server cluster hard disk status detection device performs fusion processing on the detection results of the server cluster hard disk in different dimensions to determine the detection result of the status detection process.

[0049] Continue to refer to Figure 10 The following describes the server cluster hard disk status detection method provided by this invention, taking the storage of financial transaction data in a cloud server as an example. The user, through... Figure 1 The terminals shown (including terminals 10-1 and 10-2) obtain the stored financial resources, such as fund and stock transaction data, from the corresponding cloud server 200 via network 300.

[0050] Among them, see Figure 10 , Figure 10 This is a front-end display diagram of the server cluster hard disk status detection method provided in this application, wherein the terminal (e.g., Figure 1 The terminal 10-1 in the cloud server is equipped with a cloud server client or cloud server runtime plugin that can display relevant financial information. Users can use the corresponding client to store financial data from banks, securities firms, and internet finance companies providing payment, lending, and wealth management services on the cloud server. The cloud server's management terminal (e.g., Figure 1 Terminal 10-2 in the middle) through Figure 10 The diagram illustrates the front-end display of a server cluster hard drive status detection method. It detects the operating status of the server cluster hard drives, specifically by displaying a user interface. This user interface includes a fixed-perspective view of the cloud server's operating environment, and includes different cloud server identifiers. The user interface also includes a detection component and a display component. The detection component monitors the operating parameters of the server cluster hard drives; the detection component obtains the status detection results of the server cluster hard drives; and the display component presents the status detection results of the server cluster hard drives in the user interface.

[0051] Figure 11 This is a data architecture diagram of the server cluster hard disk status detection method provided in this application. Taking a cloud server environment as an example, the data acquisition module can collect financial data running on the cloud server, report component logs to a unified access layer, and store them in a Kafka message queue for consumption. The real-time calculation and scoring module can consume the raw hardware log data stored in Kafka to calculate the real-time hardware score. The offline scoring module calculates the statistical score of similar components and the dynamic change parameters of each component within a period on a daily basis from the structured data. The API access module can achieve IP-free verification based on SHA512-bit encryption verification.

[0052] Furthermore, Figure 12 This is a front-end display diagram of the server cluster hard disk status detection method provided in this application. Through the display component, the interface of the target server cluster hard disk is presented in the user interface. Based on the detection component, a query command is sent through the interface of the target server cluster hard disk to verify the status of at least one target server.

[0053] Furthermore, Figure 13 This is a schematic diagram illustrating the processing effect of the server cluster hard disk status detection method provided in this application. Figure 14 A schematic diagram illustrating the processing effect of the server cluster hard disk status detection method provided in this application; Figure 15 This is a schematic diagram illustrating the processing effect of the server cluster hard disk status detection method provided in this application. Financial transaction data is stored in the server cluster hard disk. The method can determine the status of the server cluster hard disk in three usage scenarios (A, B, and C) in real time, reducing the need for server cluster hard disk hardware reserves. It also allows for the pre-scheduling of spare parts for different cloud server clusters and data centers based on the determined server cluster hard disk status, thereby improving the flexibility of spare parts reserves.

[0054] Furthermore, the server cluster hard drive status detection method provided in this application can accurately identify the health of hard drives, detect high-risk server cluster hard drives in advance, and effectively avoid the risk of data loss due to server cluster hard drive failures in ultra-large data center applications with millions of servers. At the same time, it reduces manual analysis of server cluster hard drive status, saves cloud server operating costs, and ensures the security of users' financial data stored on cloud servers, reducing the risk of data loss.

[0055] Beneficial technical effects: This invention, in its embodiments, monitors the operational data of the server cluster hard drives. Based on this data, a matching status detection process is triggered to obtain detection parameters corresponding to the server cluster hard drives. These parameters are then processed in multiple dimensions to determine the detection results across different dimensions. Finally, the detection results from these different dimensions are fused to determine the final result of the status detection process. This allows for automated, real-time detection of server cluster hard drive fault types, reducing hard drive replacement rates, lowering cloud server system operating costs, improving hard drive maintenance efficiency, ensuring data security for cloud server users, and enhancing the user experience.

[0056] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting the status of hard disks in a server cluster, characterized in that, The method includes: Monitor the running data on the server cluster's hard drives; Based on the operating data of the server cluster hard disk, when the operating data of the server cluster hard disk exceeds the safety range of the preset safety value, a matching status detection process is triggered to obtain the hard disk protection parameters corresponding to the server cluster hard disk. The centrifugal value of the hard disk protection parameter corresponding to the hard disk of the server cluster is calculated based on the Euclidean distance algorithm, and is used as the corresponding parameter health detection result; wherein, the farther the hard disk protection parameter deviates from the target parameter value, the greater the score of the parameter health detection result obtained by weighted calculation; Determine a target value parameter that matches the hard disk protection parameter; determine the offset of the hard disk protection parameter relative to the target value parameter by standardizing the hard disk protection parameter; and determine the distribution characteristic detection result of the hard disk protection parameter based on the offset of the hard disk protection parameter relative to the target value parameter. By processing the dynamic slope of the hard disk protection parameters, the dynamic slope change of the hard disk protection parameters within a single sampling period is determined; based on the dynamic slope change of the hard disk protection parameters within a single sampling period and a matching slope threshold, the deterioration trend result of the hard disk protection parameters is determined; wherein, if the dynamic slope change of the hard disk protection parameters within a single sampling period is greater than a predetermined threshold, the average growth amplitude of the hard disk protection parameters within the single sampling period is calculated, and the sum of squares and square root of the average growth amplitudes of multiple hard disk protection parameters are processed as the deterioration trend result; Based on the hard drive attribute parameters of the server cluster hard drives, a prediction function corresponding to the hard drive protection parameters is determined. The hard drive attribute parameters include: hard drive model, server model, rack installation time, version number, and hard drive partition identifier. Based on the operating data of the server cluster hard drives, a machine learning algorithm corresponding to the hard drive protection parameters is used to model the underlying parameters reflecting the internal health of the hard drives, which are developed in cooperation with the hard drive supplier, in order to determine the failure probability result corresponding to the hard drive protection parameters. The detection results of the server cluster hard disk in different dimensions are weighted and fused to determine the detection result of the status detection process. The detection results in different dimensions include the parameter health detection result, the distribution feature detection result, the deterioration trend result, and the failure probability result. In the usage environment corresponding to different server clusters, the weight values ​​of the different dimensions are dynamically adjusted according to the business type and environmental parameters.

2. The method according to claim 1, characterized in that, The method further includes: The user interface includes a perspective view of the cloud server operating environment from a fixed perspective, and the user interface includes different cloud server identifiers. The user interface also includes detection components and display components; The detection component monitors the operating parameters of the server cluster hard disks. The detection component is used to obtain the status detection results of the server cluster hard disks; Based on the display component, the status detection results of the server cluster hard disks are presented in the user interface.

3. The method according to claim 2, characterized in that, The method further includes: The display component presents the interface of the target server cluster hard drive in the user interface. Based on the detection component, a query command is sent through the interface of the target server cluster hard disk to verify the status of at least one target server.

4. A server cluster hard disk status detection device, characterized in that, The device includes: The information transmission module is used to monitor the running data of the server cluster hard disk; The information processing module is used to trigger a matching status detection process to obtain the hard disk protection parameters corresponding to the server cluster hard disk when the operating data of the server cluster hard disk exceeds the safety range of the preset safety value, based on the operating data of the server cluster hard disk. The information processing module is used to calculate the centrifugal value of the hard disk protection parameter corresponding to the hard disk of the server cluster based on the Euclidean distance algorithm, so as to serve as the corresponding parameter health detection result; wherein, the farther the hard disk protection parameter deviates from the target parameter value, the greater the score of the parameter health detection result obtained by weighted calculation. The information processing module is used to determine a target value parameter that matches the hard disk protection parameter; determine the offset of the hard disk protection parameter relative to the target value parameter through standardization processing of the hard disk protection parameter; and determine the distribution feature detection result of the hard disk protection parameter based on the offset of the hard disk protection parameter relative to the target value parameter. The information processing module is used to determine the dynamic slope change of the hard disk protection parameters within a single sampling period by processing the dynamic slope of the hard disk protection parameters; based on the dynamic slope change of the hard disk protection parameters within a single sampling period and a matching slope threshold, determine the deterioration trend result of the hard disk protection parameters; wherein, if the dynamic slope change of the hard disk protection parameters within a single sampling period is greater than a predetermined threshold, the average growth amplitude of the hard disk protection parameters within the single sampling period is calculated, and the average growth amplitude of multiple hard disk protection parameters is subjected to square root processing to obtain the deterioration trend result; The information processing module is used to determine the prediction function corresponding to the hard drive protection parameter based on the hard drive attribute parameters of the server cluster hard drive. The hard drive attribute parameters include: hard drive model, server model, rack installation time, version number, and hard drive partition identifier. Based on the operating data of the server cluster hard drive, the module uses a machine learning algorithm corresponding to the hard drive protection parameter to model the underlying parameters reflecting the internal health of the hard drive, which are developed in cooperation with the hard drive supplier, in order to determine the failure probability result corresponding to the hard drive protection parameter. The information processing module is used to perform weighted fusion processing on the detection results of the server cluster hard disk in different dimensions to determine the detection result of the status detection process. The detection results in different dimensions include the parameter health detection result, the distribution feature detection result, the deterioration trend result, and the failure probability result. In different usage environments corresponding to different server clusters, the weighted values ​​of the different dimensions are dynamically adjusted according to the business type and environmental parameters.

5. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the server cluster hard disk status detection method according to any one of claims 1 to 3.

6. A computer-readable storage medium storing executable instructions, characterized in that, When the executable instructions are executed by the processor, they implement the server cluster hard disk status detection method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method, device, and apparatus for hard disk maintenance, and readable storage medium

    CN108845760A

  • A system and method for detect that quality of hard disk

    CN109471765A

  • Hard disk health degree evaluation method and device

    CN111400122A