Hard disk fault warning method and device
By monitoring the hard disk status in real time and comparing it with historical data, alarm information is generated, and the problem of difficulty in detecting hard disk failures is solved, the accuracy of fault detection and operation and maintenance efficiency is improved, and data security is ensured.
Patent Information
- Application Number
- CN202510451765.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, hard disk failures are difficult to detect in time before failure, resulting in data loss or service interruption, and low operation and maintenance efficiency.
By obtaining the real-time status data of the target hard disk and comparing the historical fault status data, using the status reference data matching the target hard disk type and data importance level, an alarm message is generated to remind the operation and maintenance personnel to change the hard disk in time.
It improves the accuracy and operation and maintenance efficiency of hard disk failure detection, reduces the risk of data loss, and ensures data security and business continuity.
Smart Images

Figure CN120276944A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and more particularly to a hard disk failure warning method and device. Background Art
[0002] With the rapid development of cloud computing and big data technologies, the demand for data storage has increased exponentially. As the core storage medium, hard disks store key data such as business systems, databases, and user information. If a hard disk fails, data loss or service interruption may occur. Therefore, it is necessary to maintain hard disks in a timely manner.
[0003] However, during the operation and maintenance of hard disks, it is usually necessary for operation and maintenance personnel to observe the hard disk failure indicator light and the baseboard management controller interface to discover that a hard disk has failed. This is not only inefficient but also unable to perform maintenance in a timely manner before the hard disk fails, resulting in a lag. Summary of the Invention
[0004] In view of the above problems, this application provides a hard disk failure warning method, device, equipment, and medium.
[0005] According to a first aspect of this application, a hard disk failure warning method is provided, including: obtaining target status reference data matching the target hard disk according to the target hard disk type of the target hard disk among a plurality of hard disks and the target importance level of the data stored in the target hard disk, where the target status reference data is determined based on the historical failure status data of other hard disks having the target hard disk type among the plurality of hard disks; determining a real-time failure detection result of the target hard disk according to the real-time status data of the target hard disk and the target status reference data; and in response to the real-time failure detection result indicating that the target hard disk has a risk of failure, generating an alarm message according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located.
[0006] The second aspect of the present application provides a hard disk failure warning device, including: a reference acquisition module, configured to acquire target status reference data matching the target hard disk according to the target hard disk type of the target hard disk among multiple hard disks and the target importance level of the data stored in the target hard disk, where the target status reference data is determined based on the historical failure status data of other hard disks having the target hard disk type among the multiple hard disks; a result determination module, configured to determine the real-time failure detection result of the target hard disk according to the real-time status data of the target hard disk and the target status reference data; and a warning generation module, configured to generate a warning message according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located in response to the real-time failure detection result indicating that the target hard disk has a risk of failure.
[0007] The third aspect of the present application provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, where the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0008] The fourth aspect of the present application further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0009] According to the embodiments of the present application, by comparing the real-time status data of the target hard disk with the status data of the hard disks of the same hard disk type when a failure occurs, it is possible to determine that the target hard disk has a risk of failure before the target hard disk actually fails, so that maintenance can be performed on the target hard disk before a failure occurs, improving security. Moreover, when determining the status reference data, the target status reference data matching the target hard disk type and the target data importance level is used, which not only considers the differences between hard disks of different hard disk types, but also considers the influence of the importance level of the data stored in the hard disk. Therefore, it is possible to perform real-time status detection on the target hard disk specifically, improving the accuracy of real-time failure detection. In addition, by generating a warning message according to the target geographical location, it is convenient for the operation and maintenance personnel to locate the target hard disk, improving the operation and maintenance efficiency. Description of the Drawings
[0010] Through the following description of the embodiments of the present application with reference to the drawings, the above content and other objects, features, and advantages of the present application will become clearer. In the drawings:
[0011] Figure 1 Schematically shows an application scenario diagram of the hard disk failure warning method, device, equipment, and medium according to the embodiments of the present application;
[0012] Figure 2The flowchart of the hard disk failure warning method according to an embodiment of the present application is schematically shown;
[0013] Figure 3 The schematic diagram of determining the target geographical location according to an embodiment of the present application is schematically shown;
[0014] Figure 4 The schematic diagram of determining the critical state reference data according to an embodiment of the present application is schematically shown;
[0015] Figure 5 The flowchart of the hard disk failure warning method according to a specific embodiment of the present application is schematically shown;
[0016] Figure 6 The structural block diagram of the hard disk failure warning device according to an embodiment of the present application is schematically shown; and
[0017] Figure 7 The block diagram of the electronic device suitable for implementing the hard disk failure warning method according to an embodiment of the present application is schematically shown. Detailed Embodiments
[0018] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0019] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0020] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0021] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0022] In server testing, the test environment undertakes key tasks such as development verification, functional testing, and stress simulation. The hard disk status directly affects the reliability of test results, development efficiency, and the security of subsequent production deployment. If data such as automated test cases, logs, and temporary databases are lost due to hard disk failures, it will cause the test process to be interrupted or the results to be irreproducible. Monitoring the health status of the hard disk and replacing faulty disks in advance can prevent the automated test process from stagnating due to hardware problems.
[0023] In traditional testing or operation and maintenance, handling faults is an important task, and hard disk failures are particularly common. If the hard disk is replaced improperly, it will lead to data loss and cause serious consequences. The drawback that the hard disk status is difficult to detect in real time not only affects data security and business continuity but also increases operation and maintenance costs. In a cluster server, it is very difficult to detect a faulty hard disk in time, which adds a lot of difficulties to the replacement and maintenance of the hard disk and poses a great risk to data security. Once a hard disk anomaly or its health status cannot be detected in time, it will lead to data loss or inability to be rebuilt. In traditional testing, to find a hard disk with an abnormal health status, generally, the hard disk status is monitored through the Baseboard Management Controller (BMC) interface to detect the abnormal hard disk. However, the specific health parameters cannot be read in real time. And when the server is equipped with many hard disks, even if an abnormal hard disk is detected, it is impossible to find the abnormal hard disk and read the real-time abnormal parameters immediately, and it is impossible to identify the health status of the hard disk in advance.
[0024] Embodiments of the present application provide a hard disk failure warning method, including: obtaining target status reference data matching the target hard disk according to the target hard disk type of the target hard disk among multiple hard disks and the target importance level of the data stored in the target hard disk, where the target status reference data is determined based on the historical failure status data of other hard disks having the target hard disk type among the multiple hard disks; determining a real-time failure detection result of the target hard disk according to the real-time status data of the target hard disk and the target status reference data; and in response to the real-time failure detection result indicating that the target hard disk has a risk of failure, generating an alarm message according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located.
[0025] Embodiments of the present application dynamically and real-time capture the status data of a hard disk, compare it with the target status reference data, and when it is found that there is a risk of failure in the target hard disk, an alarm message is generated in a timely manner according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located for alarm, reminding the operation and maintenance personnel to replace the hard disk with abnormal health in a timely manner to prevent data loss, thereby protecting the security of hard disk data, greatly improving the discovery efficiency of abnormal hard disks, not only saving a great deal of time and manpower, but also ensuring the security and stability of data during testing or operation and maintenance.
[0026] Figure 1 Schematically shows an application scenario diagram of a hard disk failure alarm method, device, equipment, and medium according to an embodiment of the present application.
[0027] As Figure 1 shown, the application scenario according to this embodiment may include a first server 101, a network 102, and a server cabinet 103, where a plurality of second servers 104 are placed in the server cabinet 103, and each second server 104 may include a plurality of hard disks. The network 102 is a medium for providing a communication link between the first server 101 and the second server 104. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0028] The first server 101 is used to provide operation and maintenance services for the second server 104, such as detecting the status of the hard disk in the second server 104, and generating an alarm message when there is a risk of failure in the hard disk.
[0029] The second server 104 may be a server that provides various services, such as a background management server that supports the website browsed by the terminal device (only an example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0030] It should be noted that the hard disk failure warning method provided by the embodiments of the present application can generally be executed by the first server 101. Correspondingly, the hard disk failure warning device provided by the embodiments of the present application can generally be set in the server 101. The hard disk failure warning method provided by the embodiments of the present application can also be executed by a server or a server cluster different from the first server 101 and capable of communicating with the second server 104. Correspondingly, the hard disk failure warning device provided by the embodiments of the present application can also be set in a server or a server cluster different from the first server 101 and capable of communicating with the second server 104. The hard disk failure warning method provided by the embodiments of the present application can also be executed by a server different from the second server 104. Correspondingly, the hard disk failure warning device provided by the embodiments of the present application can also be set in the second server 104.
[0031] For example, the first server 101 obtains target status reference data matching the target hard disk according to the target hard disk type of the target hard disk among multiple hard disks of the second server 104 and the target importance level of the data stored in the target hard disk, where the target status reference data is determined based on the historical failure status data of other hard disks with the target hard disk type among the multiple hard disks; determines the real-time failure detection result of the target hard disk according to the real-time status data of the target hard disk and the target status reference data; and in response to the real-time failure detection result indicating that the target hard disk has a risk of failure, generates an alarm message according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located.
[0032] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0033] are merely illustrative. According to actual requirements, there can be any number of first servers, networks, server cabinets, and second servers. Figure 1 Based on the Figures 2 to 5 scenario described below, the hard disk failure warning method of the embodiments of the application will be described in detail through
[0034] Figure 2 FIG. schematically shows a flowchart of the hard disk failure warning method according to an embodiment of the present application.
[0035] As Figure 2 shown, the hard disk failure warning method of this embodiment includes operation S210 to operation S230.
[0036] In operation S210, target status reference data matching the target hard disk is obtained according to the target hard disk type of the target hard disk among multiple hard disks and the target importance level of the data stored in the target hard disk.
[0037] Multiple hard drives can be hard drives that provide storage media for a cluster server. For example, multiple servers included in the cluster server can be placed in different cabinets respectively, and different cabinets can be placed in different computer rooms.
[0038] The hard drive type can include information such as the specific model and manufacturer of the hard drive. The importance level can be used to characterize the importance of the data stored in the hard drive.
[0039] The status reference data can be used to determine whether there is a risk of failure for the hard drive. For example, the status reference data can include the status data of the hard drive in the healthy, risky, and faulty conditions.
[0040] Since the hard drive types of multiple hard drives are different, the structures of the hard drives are different, and thus the states when the hard drives fail are different. Therefore, considering the structural differences between different types of hard drives, the status reference data of the hard drives of the same hard drive type can be determined according to the historical failure status data of the hard drives of the same hard drive type, so as to set different status reference data for different hard drive types specifically.
[0041] Further, since the importance levels of the data stored in different hard drives are also different, for example, the impact caused by a hard drive with a higher importance level of stored data when it fails is greater. Therefore, when setting the status reference data, the importance level of the data stored in the hard drive also needs to be considered to improve security.
[0042] The target status reference data is determined based on the historical failure status data of other hard drives with the target hard drive type among multiple hard drives, and is used to determine whether there is a risk of failure for the target hard drive. The historical failure status data can be the status data of other hard drives with the target hard drive type among multiple hard drives when they fail. For example, the status reference data matching the target hard drive type and the target importance level can be obtained as the target status reference data.
[0043] In operation S220, based on the real-time status data and the target status reference data of the target hard drive, determine the real-time failure detection result of the target hard drive.
[0044] The real-time status data can be used to characterize the real-time status of the target hard drive. For example, the real-time status data of the target hard drive can be obtained by means of server background monitoring, and the real-time status data can include the Self-Monitoring Analysis and Reporting Technology (SMART) data of the target hard drive, etc.
[0045] The real-time fault detection result can be determined according to the comparison relationship between the real-time status data and the target status reference data. For example, the real-time fault detection result can be used to characterize whether there is a risk of failure for the target hard disk.
[0046] In operation S230, in response to the real-time fault detection result indicating that there is a risk of failure for the target hard disk, an alarm message is generated according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located.
[0047] Since there are a large number of hard disks in the cluster server, in order for the operation and maintenance personnel to quickly locate the position of the target hard disk, the location where the hard disk is located can be added to the alarm message, such as the target geographical location of the target server where the target hard disk is located. The target geographical location can include location information such as the computer room and cabinet where the target server is located.
[0048] According to the embodiments of the present application, by comparing the real-time status data of the target hard disk with the status data of the hard disks of the same hard disk type when a failure occurs, it is possible to determine the risk of failure of the target hard disk before the target hard disk actually fails, so that maintenance can be performed on the target hard disk before it fails, improving security. And when determining the status reference data, the target status reference data matching the target hard disk type and the target data importance level is used, which not only considers the differences between hard disks of different hard disk types, but also considers the influence of the importance level of the data stored in the hard disk. Therefore, it is possible to perform real-time status detection on the target hard disk specifically, improving the accuracy of real-time fault detection. In addition, by generating an alarm message according to the target geographical location, it is convenient for the operation and maintenance personnel to locate the target hard disk and improve the operation and maintenance efficiency.
[0049] According to the embodiments of the present application, an alarm message is generated according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located, including: determining the target server identifier of the target server according to the target hard disk identifier of the target hard disk by using the first mapping relationship between the hard disk identifier and the server identifier of the server where the hard disk is located; determining the target geographical location where the target server is located according to the target server identifier by using the second mapping relationship between the server identifier and the geographical location where the server is located; and generating an alarm message according to the target geographical location, the real-time status data, and the target status reference data.
[0050] The hard disk identifier can be a drive letter, and the server identifier can include the Internet Protocol (IP) address of the server. For example, when the hard disk is connected to the server, the hard disk identifier of the hard disk can be obtained, and the first mapping relationship between the hard disk identifier and the server identifier of the server where the hard disk is located can be stored.
[0051] The server identifier that has a first mapping relationship with the target hard disk identifier can be determined as the target server identifier, and then the geographical location that has a second mapping relationship with the target server identifier can be determined as the target geographical location. For example, both the first mapping relationship and the second mapping relationship can be stored in the form of a mapping table.
[0052] According to a preset alarm information template, the target geographical location, real-time status data, and target status reference data can be filled into the alarm information template to obtain alarm information, and the alarm information can be sent to the terminal device corresponding to the operation and maintenance personnel, so that the operation and maintenance personnel can determine whether the target hard disk needs to be replaced according to the real-time status data and target status reference data in the alarm information, and find the target hard disk according to the target geographical location when it is determined that the target hard disk needs to be replaced.
[0053] According to the embodiments of the present application, by pre-storing the first mapping relationship and the second mapping relationship, the target hard disk and the target geographical location of the target server where the target hard disk is located can be associated, which is convenient for locating the target hard disk.
[0054] Figure 3 Schematically shows a schematic diagram of determining the target geographical location according to an embodiment of the present application.
[0055] As Figure 3 shown, multiple first mapping relationships between the hard disk identifier and the server identifier of the server where the hard disk is located are stored in the first mapping table 320. For example, the server identifier of the server where the hard disk with the hard disk identifier of hard disk 1 is located is server 1, the server identifier of the server where the hard disk with the hard disk identifier of hard disk 2 is located is server 1, and the server identifier of the server where the hard disk with the hard disk identifier of hard disk 3 is located is server 2.
[0056] As Figure 3 shown, multiple second mapping relationships between the server identifier and the geographical location of the server are stored in the second mapping table 330. For example, the geographical location of the server with the server identifier of server 1 is cabinet 1, the geographical location of the server with the server identifier of server 2 is cabinet 1, and the geographical location of the server with the server identifier of server 3 is cabinet 2.
[0057] As Figure 3 shown, when the target hard disk identifier 310 is hard disk 2, the server 2 that has a first mapping relationship with hard disk 2 can be determined as the target server identifier 330 by using the first mapping table 320, and the cabinet 2 that has a second mapping relationship with server 2 can be determined as the target geographical location 350 by using the second mapping table 340.
[0058] According to an embodiment of the present application, target status reference data matching the target hard disk is obtained according to the target hard disk type of the target hard disk among multiple hard disks and the target importance level of the data stored in the target hard disk, including: determining a set of status reference data corresponding to the target hard disk type according to the target hard disk type; and screening out the status reference data matching the target importance level from the multiple status reference data as the target status reference data.
[0059] The set of status reference data includes multiple status reference data matching the importance level. Since the importance levels of the data stored in hard disks of the same hard disk type can be different, multiple status reference data matching the importance level can be set for the same hard disk type.
[0060] The status reference data may include status critical values of the hard disk health status and the risk status. For example, the status critical values corresponding to hard disks of the same hard disk type with different importance levels are different. For a higher importance level, a more stringent status critical value can be set to ensure timely maintenance before the hard disk fails and avoid a greater impact caused by its failure to ensure security. For a hard disk with a lower importance level of stored data, the impact caused by its failure is smaller. To avoid frequent alarms, a more relaxed status critical value can be set for a lower importance level.
[0061] For example, when the type of the status reference data is hard disk temperature reference data, the status critical value corresponding to hard disk type A and high importance level can be 60°C, and the status critical value corresponding to hard disk type A and low importance level can be 70°C.
[0062] According to an embodiment of the present application, by first determining a set of status reference data corresponding to the target hard disk type, and then determining the target status reference data corresponding to the target importance level from the set of status reference data, the target hard disk can be targeted for real-time fault detection using the target status reference data, improving security.
[0063] According to an embodiment of the present application, the set of status reference data is generated through the following operations: obtaining historical fault status data of other hard disks with the target hard disk type; determining critical status reference data corresponding to the target hard disk type based on the statistical information of the historical fault status data; and generating a set of status reference data based on the critical status reference data and multiple importance levels.
[0064] When a hard disk fails, the status data of the failed hard disk can be stored to obtain historical fault status data.
[0065] The statistical information may include the maximum value, minimum value, etc. of the historical fault status data, or may also include the results obtained after performing statistical operations on the historical fault status data.
[0066] The critical state reference data represents the state threshold at which a hard disk of the target hard disk type has a risk of failure. For example, the critical state reference data for the hard disk temperature can be 60°C. When the hard disk temperature is less than 60°C, it is in a healthy state, and when it is greater than 60°C, there is a risk of failure.
[0067] It is possible to determine that the critical state reference data is the state reference data corresponding to the highest importance level, or it is also possible to determine that the state data close to the critical state reference data is the state reference data corresponding to the highest importance level. For example, it can be determined that the hard disk temperature corresponding to the highest importance level is 60°C, or it can also be determined that the hard disk temperature corresponding to the highest importance level is 58°C for early warning.
[0068] Based on the preset step size, it is possible to determine the state reference data corresponding to other importance levels on the basis of the state reference data corresponding to the highest importance level. For example, the importance levels include a high importance level and a low importance level. The hard disk temperature corresponding to the high importance level is 60°C, and the preset step size is 5°C. On this basis, it is determined that the hard disk temperature corresponding to the low importance level is 65°C.
[0069] According to the embodiments of the present application, by setting the state reference data corresponding to different importance levels on the basis of the same hard disk type, the obtained state reference data set takes into account the influence of the importance level, and thus it is possible to perform targeted real-time fault detection on hard disks with different importance levels of stored data according to the state reference data set, improving security.
[0070] The number of faulty hard disks gradually increases over time, and the historical fault state data is also updated accordingly.
[0071] According to the embodiments of the present application, in response to reaching a predetermined moment, the critical state reference data is updated by using the historical fault data generated between the current predetermined moment and the previous predetermined moment, and the state reference data set is updated based on the critical state reference data, improving the flexibility and accuracy of the state reference data set, and thus improving the accuracy of real-time fault detection.
[0072] According to the embodiments of the present application, based on the statistical information of the historical fault state data, the critical state reference data corresponding to the target hard disk type is determined, including: establishing a probability density function of the fault state data corresponding to the target hard disk type based on the mean and standard deviation of the historical fault state data; determining the confidence interval of the fault state data corresponding to the target hard disk type based on the preset confidence level and the probability density function; determining the critical state reference data based on the extreme values of the confidence interval; or determining the critical state reference data based on the extreme values of the historical fault state data.
[0073] The probability density function can be used to characterize the probability density of the target hard disk corresponding to the fault state data being in the fault state, and then the probability that the target hard disk corresponding to the fault state data is in the fault state can be determined according to the probability density function.
[0074] The preset confidence level can be used to determine the confidence interval of the fault state data from the probability density function. For example, the preset confidence level can be 95%.
[0075] The extreme value of the confidence interval can be the maximum value or the minimum value of the confidence interval, and the extreme value of the confidence interval can be determined as the critical state reference data. For example, if the confidence interval of the faulty hard disk temperature is [60°C, 90°C], 60°C can be determined as the critical state reference data.
[0076] According to the embodiments of the present application, by using the probability density function of the fault state data to determine the critical state reference data, the interference of the error value in the historical fault state data on the critical state reference data can be reduced, and the accuracy of real-time fault detection can be improved.
[0077] It is also possible to directly determine the extreme value of the historical fault data as the critical state reference data. For example, if the minimum value of the historical faulty hard disk temperature is 50°C, 50°C can be determined as the critical state reference data.
[0078] According to the embodiments of the present application, by directly determining the critical state reference data based on the extreme value of the historical fault state data, the extreme state data of the fault can be taken into account, and the safety of real-time fault detection can be improved.
[0079] Figure 4 Schematically shows a schematic diagram of determining the critical state reference data according to the embodiments of the present application.
[0080] As Figure 4 shown, the mean value 420 and the standard deviation 430 of the historical fault state data 410 are determined, and the probability density function 440 of the fault state data is determined according to the mean value 420 and the standard deviation 430. The confidence interval 450 is determined from the probability density function 440 according to the preset confidence level, and the critical state reference data 460 is determined according to the extreme value of the confidence interval 450.
[0081] After the hard disk is connected to the server, the initial state data of the hard disk can be collected, and the initial state data is compared with the state reference data. In the case of passing the comparison, the server background monitoring program is used to obtain the real-time state data of the hard disk, and the state data when the hard disk is in the fault state is stored to obtain the historical fault state data.
[0082] The hard disk sensor can collect real-time status data of the hard disk and store it in a preset storage area. The server background monitoring program can obtain the real-time status data that can determine the hard disk failure status from the preset storage area according to the hard disk type of the hard disk for real-time failure detection.
[0083] In the case where the comparison fails, a prompt message indicating that the initial state of the hard disk is abnormal is generated to prompt the operation and maintenance personnel to replace the hard disk.
[0084] According to an embodiment of the present application, determining the real-time failure detection result of the target hard disk according to the real-time status data and the target status reference data of the target hard disk includes: in response to determining that the real-time status data exceeds the healthy status data range, determining that the real-time failure detection result indicates that the target hard disk has a risk of failure; and in response to determining that the real-time status data is within the healthy status data range, determining that the real-time failure detection result indicates that the target hard disk is in a healthy state.
[0085] According to an embodiment of the present application, the target status reference data has a healthy status data range. For example, the healthy status data range can be used to represent the status data range of the target hard disk in a healthy state.
[0086] For example, when the type of the target status reference data is the hard disk temperature reference data, the healthy status data range can be 30°C - 60°C. When the real-time hard disk temperature of the target hard disk is 40°C, it can be determined that the real-time failure detection result indicates that the target hard disk is in a healthy state. When the real-time hard disk temperature of the target hard disk is 70°C, it can be determined that the real-time failure detection result indicates that the target hard disk has a risk of failure.
[0087] The target status reference data can also have a risk status data range and a failure status data range, etc., which can be specifically set according to actual needs.
[0088] According to an embodiment of the present application, by determining the real-time failure detection result according to the relationship between the real-time status data and the healthy status data range, the failure detection efficiency and detection accuracy can be improved.
[0089] According to an embodiment of the present application, the hard disk failure warning method further includes: in response to the real-time failure detection result indicating that the target hard disk has a risk of failure, controlling the positioning light of the target hard disk to perform a lighting operation based on the target lighting parameter of the target hard disk.
[0090] Since there are multiple servers placed on the cabinet and each server is configured with multiple hard disks, even if the operation and maintenance personnel have determined the cabinet corresponding to the target server where the target hard disk is located, they cannot quickly locate the position of the target hard disk among a large number of hard disks.
[0091] The positioning light is used to locate the target hard disk among multiple hard disks of the target server, and each hard disk can have a positioning light.
[0092] The target lighting parameters can be the parameters required when performing a lighting operation on the positioning light of the target hard disk. For example, the target lighting parameters can include parameters such as the slot where the target hard disk is located.
[0093] According to an embodiment of the present application, by performing a lighting operation on the positioning light of the target hard disk, it is convenient for the operation and maintenance personnel to quickly locate the target hard disk among multiple hard disks of the target server, improving the operation and maintenance efficiency.
[0094] According to an embodiment of the present application, based on the target lighting parameters of the target hard disk, controlling the positioning light of the target hard disk to perform a lighting operation includes: determining the target lighting parameters corresponding to the target hard disk according to the target hard disk type of the target hard disk; and generating a lighting instruction for the target hard disk according to the target lighting parameters to control the positioning light of the target hard disk to perform a lighting operation.
[0095] The configuration of the positioning lights for different hard disk types is different, resulting in different lighting parameters corresponding to different hard disk types. Therefore, when performing a lighting operation on the positioning light of the target hard disk, it is necessary to first determine the target lighting parameters corresponding to the target hard disk according to the target hard disk type of the target hard disk.
[0096] The lighting instruction can be used to control the positioning light of the target hard disk to perform a lighting operation. For example, the server can generate a lighting instruction and send the lighting instruction to the hard disk controller to use the hard disk controller to perform a lighting operation on the positioning light of the target hard disk.
[0097] According to an embodiment of the present application, by controlling the positioning light of the target hard disk to perform a lighting operation according to the target lighting parameters corresponding to the target hard disk type, the accuracy of the lighting operation is improved.
[0098] According to an embodiment of the present application, the hard disk failure warning method further includes: in response to determining that the target importance level is a preset importance level, performing a backup operation on the stored data in the target hard disk, and switching the task currently executed by the target hard disk to be executed by other hard disks in the same storage array as the target hard disk.
[0099] The preset importance level can be a relatively high importance level. For example, data such as business codes can be set as data with a relatively high importance level.
[0100] The stored data in the target hard disk can be written to other hard disks in the same storage array as the target hard disk to back up the stored data in the target hard disk.
[0101] Since there is a risk of failure for the target hard disk and it may fail at any time, the tasks currently executed by the target hard disk can be switched to other hard disks in the same storage array as the target hard disk, so as to avoid the inability to execute tasks due to the sudden failure of the target hard disk, and further avoid the server from being unable to provide services.
[0102] According to the embodiments of the present application, in the case where the target hard disk has a risk of failure and the importance level of the stored data is relatively high, backing up the stored data in the target hard disk in a timely manner and switching the execution tasks of the target hard disk to other hard disks can avoid data loss and service interruption caused by the failure of the target hard disk.
[0103] Figure 5 The flowchart of the hard disk failure warning method according to a specific embodiment of the present application is schematically shown.
[0104] As Figure 5 shown, the hard disk failure warning method includes operations S501 to S514.
[0105] In operation S501, obtain the initial state data of all hard disks of the server.
[0106] In operation S502, determine whether the initial state data is consistent with the state reference data. If they are inconsistent, execute operation S503; otherwise, execute operation S504.
[0107] In operation S503, the initial state of the hard disk is abnormal, and replace the abnormal hard disk.
[0108] In operation S504, start the server background monitoring program.
[0109] In operation S505, perform real-time fault detection based on the real-time state data and the target state reference data.
[0110] In operation S506, determine whether the target hard disk has a risk of failure. If there is a risk of failure, execute operation S507; otherwise, execute operation S505.
[0111] In operation S507, back up the data in the target hard disk.
[0112] In operation S508, obtain the target lighting parameters.
[0113] In operation S509, determine the target geographical location of the target server where the target hard disk is located.
[0114] In operation S510, generate a lighting instruction to control the positioning light of the target hard disk to perform a lighting operation.
[0115] In operation S511, an alarm message is generated based on real-time status data, a target geographical location, and target status reference data.
[0116] In operation S512, the operation and maintenance personnel determine whether the target hard disk needs to be replaced. If the target hard disk needs to be replaced, operation S513 is executed; otherwise, operation S505 is executed.
[0117] In operation S513, the target server and the target hard disk are found according to the target geographical location and the positioning light.
[0118] In operation S514, the target hard disk is replaced and the backup data is imported.
[0119] Based on the above hard disk fault warning method, the present application also provides a hard disk fault warning device. The following will be combined with Figure 6 to describe this device in detail.
[0120] Figure 6 Schematically shows a structural block diagram of a hard disk fault warning device according to an embodiment of the present application.
[0121] As Figure 6 shown, the hard disk fault warning device 600 of this embodiment includes a reference acquisition module 610, a result determination module 620, and an early warning generation module 630.
[0122] The reference acquisition module 610 is configured to obtain target status reference data matching the target hard disk according to the target hard disk type of the target hard disk among multiple hard disks and the target importance level of the data stored in the target hard disk, where the target status reference data is determined based on the historical fault status data of other hard disks with the target hard disk type among multiple hard disks. In one embodiment, the reference acquisition module 610 may be configured to execute operation S210 described above, which will not be elaborated here.
[0123] The result determination module 620 is configured to determine the real-time fault detection result of the target hard disk according to the real-time status data and the target status reference data of the target hard disk. In one embodiment, the result determination module 620 may be configured to execute operation S220 described above, which will not be elaborated here.
[0124] The alarm generation module 630 is configured to generate an alarm message according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located in response to the real-time fault detection result indicating that there is a risk of the target hard disk failing. In one embodiment, the early warning generation module 630 may be configured to execute operation S230 described above, which will not be elaborated here.
[0125] According to an embodiment of the present application, the alarm generation module 630 includes an identifier determination sub-module, a location determination sub-module, and an alarm generation sub-module.
[0126] The identification determination sub-module is configured to determine the target server identification of the target server according to the target hard disk identification of the target hard disk and by using the first mapping relationship between the hard disk identification and the server identification of the server where the hard disk is located.
[0127] The location determination sub-module is configured to determine the target geographical location where the target server is located according to the target server identification and by using the second mapping relationship between the server identification and the geographical location where the server is located.
[0128] The alarm generation sub-module is configured to generate an alarm message according to the target geographical location, the real-time status data, and the target status reference data.
[0129] According to an embodiment of the present application, the reference acquisition module 610 includes a set determination sub-module and a target determination sub-module.
[0130] The set determination sub-module is configured to determine a set of status reference data corresponding to the target hard disk type according to the target hard disk type, where the set of status reference data includes multiple status reference data that match the importance level.
[0131] The target determination sub-module is configured to screen out the status reference data that matches the target importance level from the multiple status reference data as the target status reference data.
[0132] According to an embodiment of the present application, the set of status reference data is generated through the following operations: acquiring the historical failure status data of other hard disks whose hard disk type is the target hard disk type; determining the critical status reference data corresponding to the target hard disk type based on the statistical information of the historical failure status data, where the critical status reference data represents the status critical value at which a hard disk with the target hard disk type has a risk of failure; and generating the set of status reference data based on the critical status reference data and multiple importance levels.
[0133] According to an embodiment of the present application, determining the critical status reference data corresponding to the target hard disk type based on the statistical information of the historical failure status data includes: establishing a probability density function of the failure status data corresponding to the target hard disk type based on the mean and standard deviation of the historical failure status data; determining the confidence interval of the failure status data corresponding to the target hard disk type based on the preset confidence level and the probability density function; determining the critical status reference data based on the extreme values of the confidence interval; or determining the critical status reference data based on the extreme values of the historical failure status data.
[0134] According to an embodiment of the present application, the target status reference data has a healthy status data interval. The result determination module 620 includes a first determination sub-module and a second determination sub-module.
[0135] The first determination sub-module is configured to determine that the real-time fault detection result indicates that there is a risk of a fault occurring in the target hard disk in response to determining that the real-time status data exceeds the healthy status data range.
[0136] The second determination sub-module is configured to determine that the real-time fault detection result indicates that the target hard disk is in a healthy state in response to determining that the real-time status data is within the healthy status data range.
[0137] According to an embodiment of the present application, the hard disk fault warning device 600 further includes a lighting module.
[0138] The lighting module is configured to control the positioning light of the target hard disk to perform a lighting operation based on the target lighting parameters of the target hard disk in response to the real-time fault detection result indicating that there is a risk of a fault occurring in the target hard disk, where the positioning light is used to locate the target hard disk among multiple hard disks of the target server.
[0139] According to an embodiment of the present application, the lighting module includes a parameter determination sub-module and an instruction generation sub-module.
[0140] The parameter determination sub-module is configured to determine the target lighting parameters corresponding to the target hard disk according to the target hard disk type of the target hard disk.
[0141] The instruction generation sub-module is configured to generate a lighting instruction for the target hard disk according to the target lighting parameters to control the positioning light of the target hard disk to perform a lighting operation.
[0142] According to an embodiment of the present application, the hard disk fault warning device 600 further includes a backup switching module.
[0143] The backup switching module is configured to perform a backup operation on the stored data in the target hard disk and switch the task currently executed by the target hard disk to be executed by other hard disks in the same storage array as the target hard disk in response to determining that the target importance level is a preset importance level.
[0144] According to an embodiment of the present application, any of the reference acquisition module 610, the result determination module 620, and the warning generation module 630 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present application, at least one of the reference acquisition module 610, the result determination module 620, and the warning generation module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the reference acquisition module 610, the result determination module 620, and the warning generation module 630 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0145] Figure 7 Schematically shows a block diagram of an electronic device suitable for implementing the hard disk failure warning method according to an embodiment of the present application.
[0146] As Figure 7 shown, the electronic device 700 according to an embodiment of the present application includes a processor 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. The processor 701 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.
[0147] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 702 and / or the RAM 703. It should be noted that the programs can also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in one or more memories.
[0148] According to an embodiment of the present application, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A driver 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the driver 710 as needed so that a computer program read from it can be installed into the storage portion 708 as needed.
[0149] The present application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0150] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.
[0151] An embodiment of the present application also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the hard disk failure warning method provided by the embodiment of the present application.
[0152] When the computer program is executed by the processor 701, it executes the above functions defined in the system / apparatus of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0153] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0154] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or be installed from the removable medium 711. When the computer program is executed by the processor 701, it executes the above functions defined in the system of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0155] In accordance with embodiments of the present application, program code for executing the computer programs provided by the embodiments of the present application can be written in any combination of one or more programming languages. For example, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0157] Those skilled in the art can understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.
[0158] The above describes the embodiments of the present application. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.
Claims
1. A hard disk failure warning method, characterized in that, The method includes: Obtaining target status reference data matching the target hard disk according to the target hard disk type of the target hard disk among multiple hard disks and the target importance level of the data stored in the target hard disk, where the target status reference data is determined based on the historical failure status data of other hard disks with the target hard disk type among the multiple hard disks; Determining a real-time failure detection result of the target hard disk according to the real-time status data of the target hard disk and the target status reference data; and In response to the real-time failure detection result indicating that the target hard disk has a risk of failure, generating an alarm message according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located.
2. The method according to claim 1, characterized in that, The generating an alarm message according to the real-time status data, the target status reference data, and the target geographical location of the target server where the target hard disk is located includes: Determining the target server identifier of the target server according to the target hard disk identifier of the target hard disk and using the first mapping relationship between the hard disk identifier and the server identifier of the server where the hard disk is located; Determining the target geographical location where the target server is located according to the target server identifier and using the second mapping relationship between the server identifier and the geographical location where the server is located; and Generating the alarm message according to the target geographical location, the real-time status data, and the target status reference data.
3. The method according to claim 1, wherein The obtaining target status reference data matching the target hard disk according to the target hard disk type of the target hard disk among multiple hard disks and the target importance level of the data stored in the target hard disk includes: Determining a set of status reference data corresponding to the target hard disk type according to the target hard disk type, where the set of status reference data includes multiple status reference data matching the importance level; and Filtering out the status reference data matching the target importance level from the multiple status reference data as the target status reference data.
4. The method according to claim 3, wherein The set of status reference data is generated through the following operations: Obtaining the historical failure status data of other hard disks with the target hard disk type; Determining critical status reference data corresponding to the target hard disk type based on the statistical information of the historical failure status data, where the critical status reference data represents the status threshold for a hard disk with the target hard disk type to have a risk of failure; and Generating the set of status reference data based on the critical status reference data and multiple importance levels.
5. The method according to claim 4, characterized in that The determining critical status reference data corresponding to the target hard disk type based on the statistical information of the historical failure status data includes: Establishing a probability density function of the failure status data corresponding to the target hard disk type based on the mean and standard deviation of the historical failure status data; Determining a confidence interval of the failure status data corresponding to the target hard disk type based on a preset confidence level and the probability density function; Determining the critical status reference data based on the extreme values of the confidence interval; or Determine the critical state reference data based on the extreme values of the historical fault state data.
6. The method according to claim 1, wherein The target state reference data has a healthy state data range; Determining the real-time fault detection result of the target hard disk according to the real-time state data of the target hard disk and the target state reference data includes: In response to determining that the real-time state data exceeds the healthy state data range, determining that the real-time fault detection result indicates that the target hard disk has a risk of failure; And In response to determining that the real-time state data is within the healthy state data range, determining that the real-time fault detection result indicates that the target hard disk is in a healthy state.
7. The method according to claim 1, characterized in that, The method further includes: In response to the real-time fault detection result indicating that the target hard disk has a risk of failure, based on the target lighting parameters of the target hard disk, controlling the positioning light of the target hard disk to perform a lighting operation, where the positioning light is used to locate the target hard disk among multiple hard disks of the target server.
8. The method according to claim 7, wherein Controlling the positioning light of the target hard disk to perform a lighting operation based on the target lighting parameters of the target hard disk includes: Determining the target lighting parameters corresponding to the target hard disk according to the target hard disk type of the target hard disk; and Generating a lighting instruction for the target hard disk according to the target lighting parameters to control the positioning light of the target hard disk to perform a lighting operation.
9. The method according to claim 1, wherein The method further includes: In response to determining that the target importance level is a preset importance level, performing a backup operation on the stored data in the target hard disk, and switching the task currently executed by the target hard disk to be executed by other hard disks in the same storage array as the target hard disk.
10. A hard disk failure warning device, characterized in that, The device includes: A reference acquisition module, configured to obtain target state reference data matching the target hard disk according to the target hard disk type of the target hard disk among multiple hard disks and the target importance level of the stored data in the target hard disk, where the target state reference data is determined based on the historical fault state data of other hard disks having the target hard disk type among the multiple hard disks; A result determination module, configured to determine the real-time fault detection result of the target hard disk according to the real-time state data of the target hard disk and the target state reference data; and An alarm generation module, configured to generate an alarm message according to the real-time state data, the target state reference data, and the target geographical location of the target server where the target hard disk is located in response to the real-time fault detection result indicating that the target hard disk has a risk of failure.
Citation Information
Cited By
Disk exception processing method and device, equipment and storage medium
CN120523689A
Disk exception processing method, device, equipment and storage medium
CN120523689B