Hard disk fault monitoring method, control device, storage medium and system

By extracting and standardizing hard disk SMART data, the problem of low automation of hard disk data acquisition, management and fault monitoring in the prior art is solved, and more efficient fault prevention and monitoring is achieved, ensuring data integrity and security.

CN120179504APending Publication Date: 2025-06-20WUXI STARS MICRO SYSTEM TECHNOLOGIES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510335084.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has low degree of automation, error prone and poor cross-platform compatibility in hard disk data acquisition, management and fault monitoring.

Method used

By extracting the SMART data of the hard disk, performing standardized processing and persisting storage, obtaining user query instructions, responding to queries and sending SMART data, the client generates and displays hard disk failure indicator data according to user operations.

Benefits of technology

Improves hard disk failure prevention capabilities, ensures data integrity and security, provides a solid foundation for fault monitoring, simplifies user operations and improves cross-platform compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179504A_ABST
    Figure CN120179504A_ABST
Patent Text Reader

Abstract

The invention provides a hard disk fault monitoring method and system, a control device and a storage medium, and relates to the technical field of hard disks. The monitoring method comprises the following steps: extracting SMART data of a hard disk, carrying out standardization processing on the SMART data, and persisting the SMART data in a memory; obtaining a query instruction of a user for the SMART data of the target hard disk; in response to the query instruction, retrieving corresponding SMART query data from the memory; and sending SMART query data, generating corresponding data showing a hard disk fault index by the client, displaying the data to a user, and querying and monitoring the fault of the target hard disk. According to the invention, the integrity and security of all related data of the hard disk are ensured; and the user can quickly monitor the change trend of the hard disk fault index of the selected hard disk along with time, and the user is assisted in monitoring potential problems of the hard disk. Therefore, the fault prevention capability of the hard disk is improved, and support is provided for maintenance of the hard disk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of hard disks, and particularly relates to a hard disk fault monitoring method, a control device, a storage medium, and a system. Background Art

[0002] Whether it is for the manufacture of servers, the supply of storage controllers, or the production of backplanes and hard disks, great importance is attached to the compatibility testing of hard disks. Therefore, during the testing process of hard disks, it is necessary to closely monitor the status and basic information of hard disks, especially the index information related to hard disk faults. By analyzing these data and information, the overall condition of the hard disk can be comprehensively understood. For example, whether there are bad sectors, whether it is about to fail, or whether the lifespan has approached the end, etc. It should be noted that most mainstream manufacturers have adopted SMART (Self-Monitoring, Analysis and Reporting Technology) technology to monitor the health or fault status of hard disks. Using tools such as SMARTCTL, it is possible to easily obtain a detailed report on the hard disk, including but not limited to key indicators such as running time, power-on count, current temperature, the number of detected bad sectors, and the estimated remaining service life. This not only helps to detect potential problems in a timely manner but also effectively extends the service life of the device and reduces the risk of unexpected downtime.

[0003] However, there is still much room for improvement in the data collection, management, and analysis of related technical solutions, especially in aspects such as low automation, error-prone, and poor cross-platform compatibility. Summary of the Invention

[0004] The purpose of this application is to provide a hard disk fault monitoring method, a control device, a storage medium, and a system, aiming to solve the problems existing in the related technical solutions in aspects such as hard disk data collection, management, and fault monitoring.

[0005] According to the first aspect of this application, a hard disk fault monitoring method is provided, which is applied to the server side. The monitoring method includes: extracting the SMART data of the hard disk, performing standardization processing on the SMART data, and then persisting it in the memory; obtaining a query instruction from the user for the SMART data of the target hard disk; in response to the query instruction, retrieving the corresponding SMART query data from the memory; and sending the SMART query data, so that the client generates corresponding data showing the hard disk fault indicators according to the user interaction operation, and displays the corresponding view to the user, for querying and monitoring the faults of the target hard disk.

[0006] This hard disk failure monitoring method is applied to the server side. When collecting the SMART data of the hard disk, it immediately cleans the collected SMART data, extracts the key fields, and performs unified formatting processing. Subsequently, the processed SMART data is persistently stored in the memory. When the user can access the database to query the SMART data, the data is processed, and a chart is rendered and generated on the display interface of the client. This ensures the integrity and security of all relevant data of the hard disk, providing a solid foundation for failure monitoring; it can also retrieve the SMART data of the target hard disk required from the database according to the specific needs of the user, thereby assisting the user in monitoring or identifying potential problems of the hard disk. This not only helps improve the hard disk's failure prevention ability but also provides strong support for maintaining the hard disk's operation.

[0007] In an alternative embodiment, after extracting the SMART data of the hard disk and performing standardization processing on the SMART data, it is persisted in the memory, including: The Docker server uses a preset encapsulation library to send an instruction to query the SMART data to at least one server; in response to the instruction to query the SMART data sent by the Docker server, the at least one server returns the SMART data; the Docker server uses a preset function to convert the returned SMART data into a preset format to complete the standardization processing of the SMART data; and the standardized SMART data is persisted in the database of the memory.

[0008] Ensures the integrity and security of all relevant data, providing a solid foundation for subsequent failure monitoring.

[0009] In an alternative embodiment, the preset function is configured to call a corresponding processing function according to different interface types of the hard disk to convert the SMART data into the preset format, and the preset format includes the JSON format.

[0010] According to different interface types (intf) of the hard disk, call the corresponding processing function, and then convert the SMART data into a preset JSON format, for example. The data structure of the JSON format can not only include the basic information of the hard disk but also detail the key SMART field data. These processed SMART data are persistently stored in the database, providing reliable data support for subsequent failure monitoring. When collecting key fields, it is different for different types of hard disk interfaces. This application also provides methods for extracting key fields for different types of hard disk interfaces and the configuration of key fields.

[0011] In an alternative embodiment, the query instruction includes the interface type of the target hard disk and the time interval to be queried, and also includes the hard disk serial number or the hard disk model. After retrieving the corresponding SMART query data from the memory in response to the query instruction, the monitoring method further includes: determining the target hard disk to be retrieved according to the hard disk serial number or the hard disk model; and / or retrieving the SMART query data of different data types accordingly according to the interface type of the target hard disk.

[0012] It is possible to retrieve the SMART data of the target hard disk required from the database according to the specific needs of the user, and use this data to generate an intuitive and easy-to-understand line chart and other forms of charts. In this way, the user can quickly monitor the change trend of the hard disk failure indicators of the selected hard disk over time, so as to assist the user in monitoring or identifying potential problems of the hard disk. This not only helps to improve the hard disk failure prevention ability, but also provides strong support for maintaining the operation of the hard disk.

[0013] According to a second aspect of the present application, there is provided a hard disk failure monitoring method applied to a client. The monitoring method includes: sending a query instruction of the user for the SMART data of the target hard disk; receiving the corresponding SMART query data retrieved from the memory in response to the query instruction sent by the server side; rendering and displaying the SMART query data; and generating corresponding data showing the hard disk failure indicators by using the SMART query data according to the user interaction operation, and displaying the corresponding view to the user to query and monitor the failure of the target hard disk.

[0014] The hard disk failure monitoring method is applied to a client. When the user can access the database to query the SMART data, the data is processed, and a chart is rendered and generated on the display interface of the client. It is also possible to retrieve the SMART data of the target hard disk required from the database according to the specific needs of the user, so as to assist the user in monitoring or identifying potential problems of the hard disk. This not only helps to improve the hard disk failure prevention ability, but also provides strong support for maintaining the operation of the hard disk.

[0015] In an alternative embodiment, the generating corresponding data showing the hard disk failure indicators by using the SMART query data according to the user operation, and displaying the corresponding view to the user to query and monitor the failure of the target hard disk includes: generating a change trend curve of the corresponding hard disk failure indicator and time according to the hard disk failure indicator selected by the user to assist in predicting potential failures of the target hard disk; rendering and displaying the generated change trend curve of the hard disk failure indicator and time; and when an event corresponding to any data point on the change trend curve of the hard disk failure indicator and time is detected when the mouse hovers over it, a window pops up to display the value corresponding to that moment.

[0016] It can retrieve the SMART data of the target hard disk required according to the specific needs of the user from the database, and use this data to generate an intuitive and easy-to-understand line chart and other forms of charts. In this way, the user can quickly monitor the change trend of the hard disk failure indicators of the selected hard disk over time.

[0017] In an alternative embodiment, according to the user operation, using the SMART query data to generate corresponding data showing the hard disk failure indicators, and displaying the corresponding view to the user to query and monitor the failure of the target hard disk, including: generating a list of target hard disks that meet the filtering conditions according to the filtering conditions input by the user, wherein, for different types of target hard disks, the filtering conditions input by the user are different. Provide multiple ways to query and monitor the hard disk.

[0018] In an alternative embodiment, according to the user operation, using the SMART query data to generate corresponding data showing the hard disk failure indicators, and displaying the corresponding view to the user to query and monitor the failure of the target hard disk, including: obtaining the threshold value of the warning line set by the user for the selected hard disk failure indicator to show the corresponding hard disk failure indicator; and when the measured value of the selected hard disk failure indicator exceeds the threshold value, generating a corresponding view for prompting the failure.

[0019] The function of setting the threshold value is provided. The user can set the threshold value of the warning line for the corresponding hard disk failure indicator in a specific field; when the actual measurement result exceeds the threshold range, it can be marked with a horizontal line on the chart to help the user quickly identify the potential problem.

[0020] According to the third aspect of the present application, a control device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the above-mentioned monitoring method applied to the server side.

[0021] According to the fourth aspect of the present application, a control device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the above-mentioned monitoring method applied to the client side.

[0022] According to the fifth aspect of the present application, a machine-readable storage medium is provided, and instructions are stored on the machine-readable storage medium, and the instructions cause the machine to execute the above-mentioned monitoring method applied to the server side or the above-mentioned monitoring method applied to the client side.

[0023] According to the sixth aspect of the present application, a hard disk failure monitoring system is provided, including a client and a server side. The server side includes the control device applied to the server side as described above, and the client includes the control device applied to the client as described above.

[0024] Through the above technical solution, the hard disk failure monitoring method for the server side provided by the embodiments of the present application extracts the SMART data of the hard disk, and after standardizing the SMART data, persists it in the memory; obtains the user's query instruction for the SMART data of the target hard disk; in response to the query instruction, retrieves the corresponding SMART query data from the memory; and sends the SMART query data, so that the client, according to the user's operation, uses the SMART query data to generate corresponding data showing the hard disk failure indicators and displays them to the user for querying and monitoring the failure of the target hard disk. When collecting the SMART data of the hard disk in the embodiments of the present application, the collected SMART data is cleaned, the key fields are extracted, and unified formatting processing is performed. Subsequently, the processed SMART data is persistently stored in the memory (for example, a MySQL database). When the user can access the database to query the SMART data, the data is processed and rendered and charted on the display interface (for example, a browser) of the client. This ensures the integrity and security of all relevant data of the hard disk and provides a solid foundation for failure monitoring. The embodiments of the present application can also retrieve the SMART data of the target hard disk required from the database according to the specific needs of the user, and use this data to generate intuitive and easy-to-understand line charts and other forms of charts. In this way, the user can quickly monitor the change trend of the hard disk failure indicators of the selected hard disk over time, thereby assisting the user in monitoring or identifying potential problems of the hard disk. This not only helps to improve the hard disk's failure prevention ability but also provides strong support for maintaining the hard disk's operation.

[0025] Other features and advantages of the present application will be described in the subsequent specification, and will be partially obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures and processes pointed out in the specification and the drawings. Description of the Drawings

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1Schematic flow of the hard disk failure monitoring method provided by an exemplary embodiment of the present application Figure 1 .

[0028] Figure 2 Schematic framework diagram of the system architecture applied in an exemplary embodiment of the present application.

[0029] Figure 3 Schematic flow diagram of the example data collection process.

[0030] Figure 4 Schematic flow of the example failure monitoring process Figure 1 .

[0031] Figure 5 Schematic diagram of the interface display of the example failure monitoring Figure 1 .

[0032] Figure 6 Schematic flow of the example failure monitoring process Figure 2 .

[0033] Figure 7 Schematic diagram of the interface display of the example failure monitoring Figure 2 .

[0034] Figure 8 Schematic flow of the hard disk failure monitoring method provided by an exemplary embodiment of the present application Figure 2 .

[0035] Figure 9 Schematic framework diagram of the hard disk failure monitoring system provided by an exemplary embodiment of the present application. Detailed implementation manners

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0037] As described above, the related art uses SMART (Self-Monitoring, Analysis and Reporting Technology) technology to monitor the health or failure status of hard disks. However, the current technical solutions and their existing technical defects include:

[0038] 1) For Smart data recording, it is necessary to manually obtain SMART information. The SMART data of a hard disk is usually obtained manually. The operator needs to input specific commands to collect relevant information and record this information. This method is not only time-consuming and laborious but also completely dependent on manual operation, which is prone to errors or omissions.

[0039] 2) For Smart information management, there are diverse data storage formats. The recorded SMART data often uses Excel spreadsheets or text editors for storage. Since different team members may prefer different tools to process data, this leads to additional format conversion work required during the final aggregation stage, thus increasing unnecessary labor costs and time consumption.

[0040] 3) For Smart data cleaning and analysis, the data processing flow is complex. To prepare data for further analysis, first, it is necessary to clean the information collected from various sources, removing irrelevant items; then, reorganize the data structure according to predetermined criteria. During this process, it is required that the staff can accurately identify which are the key indicators and format them appropriately. Especially for different hard disk interfaces, different key fields need to be extracted. If the original record forms are not unified, additional steps are required to standardize all input materials. The whole process is both time-consuming and error-prone, greatly affecting efficiency.

[0041] In view of the above technical deficiencies, the present application proposes a hard disk failure monitoring method that monitors hard disk failures in real time based on hard disk SMART data.

[0042] Before explaining the embodiments of the present application in detail, the technical terms involved in the present application are explained as shown in Table 1.

[0043] Table 1

[0044]

[0045]

[0046] Figure 1 is a schematic flowchart of the hard disk failure monitoring method provided by an exemplary embodiment of the present application. This monitoring method can be applied to the server side. This monitoring method may include the following steps:

[0047] Step S110: Extract the SMART data of the hard disk, and after standardizing the SMART data, persist it in the memory.

[0048] Figure 2 shows the system architecture applied in an exemplary embodiment of the present application. Please refer to Figure 2Example, during the data collection process, the system architecture can be developed using the Python programming language. During this process, the docker server obtains SMART data from at least one monitored server and persistently stores it in a database. This process ensures the integrity and security of all relevant data, providing a solid foundation for subsequent fault monitoring.

[0049] Preferably, step S110 may include: The Docker server uses a preset encapsulation library to send an instruction to query SMART data to at least one server; in response to the instruction to query SMART data sent by the Docker server, the at least one server returns the SMART data; the Docker server uses a preset function to convert the returned SMART data into a preset format to complete the standardization process of the SMART data; and the standardized SMART data is persisted in the database of the memory.

[0050] Please refer to Figure 3 , for example, in the data collection process, the Docker server can use, for example, the Paramiko encapsulation library to send an instruction to query SMART data to multiple servers. After receiving the response, the multiple servers can pass the obtained SMART data to a preset function (for example, the pre-configured modify_smart_info function). Using this preset function, the obtained SMART data is converted into a preset format to complete the standardization process of the SMART data; and the standardized SMART data is persisted in the database of the memory (for example, the MySQL database), that is, the cleaned data can be pushed to the database and a timestamp can be generated for persistent storage.

[0051] For SMART data records, the embodiments of the present application can achieve all-weather and uninterrupted data collection. Simply set a scheduled task on the Docker server and send a query command to the server at regular intervals to easily complete it.

[0052] Further preferably, the preset function is configured to be able to call corresponding processing functions according to different interface types of hard disks to convert the SMART data into a preset format, where the preset format may include the JSON format.

[0053] Continuing with the above example, the pre-configured modify_smart_info function is configured to call the corresponding processing function according to the different interface types (intf) of the hard disk, and then convert the SMART data into a preset JSON format, for example. The data structure in JSON format can not only include the basic information of the hard disk, but also detail the key SMART field data. These processed SMART data are persistently stored in the database to provide reliable data support for subsequent fault monitoring.

[0054] For different types of hard disk interfaces, there are differences in collecting key fields. The preferred embodiment of the present application also provides methods for extracting key fields for different types of hard disk interfaces and the configuration of key fields.

[0055] For example, for SATA hard disks: mainly focus on SMART Attributes in the SMART data. Since different manufacturers may have different definitions for these attributes, the embodiments of the present application extract some common attributes with general reference value and configure the key fields, as shown in Table 2.

[0056] Table 2 Key fields and their configurations for SATA hard disks

[0057]

[0058]

[0059]

[0060] In Table 2, the parameter with "tpye" being "Pre-fail" can indicate a hard disk failure. When this parameter accumulates to a preset threshold, it can predict that the hard disk is about to fail.

[0061] For example, for SAS hard disks: different from SATA hard disks, SAS hard disks do not directly provide the SMART Attributes attribute set. However, it includes some other important parameters, such as the number of read and write operations, temperature monitoring data, and the number of startups, etc. These information are also very helpful for monitoring the hard disk status, and the relevant parameters and their configurations are shown in Table 3.

[0062] Table 3 Key fields and their configurations for SAS hard disks

[0063]

[0064]

[0065] For example, for NVMe hard drives: Traditional SMART CRL tools cannot obtain SMART data of NVMe hard drives. For this reason, the embodiments of the present application provide a CLI tool to read SMART data. The CLI tool is configured to: by accessing Log Identifier 02h (i.e., log identifier 02), important data throughout the entire controller life cycle can be obtained, and data integrity can be maintained even in the case of power failure. A variety of useful information can be obtained from this source, including but not limited to fault records, current temperature, and total capacity, as shown in Table 4.

[0066] Table 4 Keyword fields of NVMe hard drives and their configurations

[0067]

[0068]

[0069]

[0070]

[0071] Step S120: Obtain a query instruction from the user for the SMART data of the target hard drive.

[0072] Illustrated by way of example, in the fault monitoring phase, the system architecture can also be developed using the Python programming language. The embodiments of the present application can retrieve the required information from the database according to the specific needs of the user, and use this data to generate visual data, such as intuitive and easy-to-understand line charts and other forms of charts.

[0073] Step S130: In response to the query instruction, retrieve the corresponding SMART query data from the memory.

[0074] The preferred query instruction of the embodiments of the present application may include the interface type of the target hard drive and the time interval to be queried, and may also include the hard drive serial number or hard drive model. After retrieving the corresponding SMART query data from the memory in response to the query instruction, the monitoring method may further include: determining the target hard drive to be retrieved according to the hard drive serial number or hard drive model; and / or retrieving SMART query data of different data types accordingly according to the interface type of the target hard drive.

[0075] Please refer to Figure 4For example, the user can select the interface type (intf) of the target hard disk on a web page of the user side, for example. The user can also input the serial number (SN) of the hard disk and the time range in the serial number management interface, and click the "search" button to generate a query instruction for the SMART data of the target hard disk. In response to this query instruction, the server side can retrieve the (historical) SMART query data of the hard disk SN within the corresponding time range from the database of the memory. In a preferred embodiment of the present application, different data processing types can be automatically selected according to the interface type (intf) of the target hard disk, and the SMART query data of the hard disk can be returned (for example, it can include the basic information of the hard disk and the data corresponding to the key fields in Table 2-4). Among them, the basic information of the hard disk includes, for example: serial number, manufacturer name, firmware version, interface type, media type, capacity size, and model, etc. In particular, for SATA hard disks, their pre-fail attribute and the corresponding WHEN_FAILED value can also be additionally displayed.

[0076] Please refer to Figure 6 For example, the user can select the interface type (intf) of the target hard disk on a web page of the user side, for example. The user can also input the hard disk model (Model) of the hard disk and the time range in the hard disk model management interface, and click the "search" button to generate a query instruction for the SMART data of the target hard disk. In response to this query instruction, the server side can retrieve the (historical) SMART query data belonging to this Model of the hard disk Model within the corresponding time range from the database of the memory. In a preferred embodiment of the present application, different data processing types can be automatically selected according to the interface type (intf) of the target hard disk. Among them, the SMART data queried according to the Model indicates which hard disk SNs this hard disk model has; the user can select, for example, the failure type on the right page and click the "Generate SN Table" button, for example, to generate a list of SNs that have this failure type.

[0077] Step S140: Send the SMART query data so that the client side, according to the user interaction operation, generates corresponding data showing the hard disk failure indicators by using the SMART query data and displays the corresponding view to the user to query and monitor the failure of the target hard disk.

[0078] For example, send the SMART query data to the client side for rendering and display, and the user can query and monitor the failure of the target hard disk on the display interface of the client side.

[0079] In a preferred embodiment of the present application, the client, according to user interaction operations, uses the SMART query data to generate corresponding data showing hard disk failure indicators, and displays the corresponding view to the user to query and monitor the failures of the target hard disk, which may include: generating a corresponding change trend curve of the hard disk failure indicator and time according to the selected hard disk failure indicator by the user to assist in predicting potential failures of the target hard disk; rendering and displaying the generated change trend curve of the hard disk failure indicator and time; and when detecting an event corresponding to any data point when the mouse hovers over any data point on the change trend curve of the hard disk failure indicator and time, popping up a window to display the corresponding value at that moment.

[0080] Please refer to Figure 4 and Figure 5 For example, the client can be configured to: in the display interface of the client, after the user completes the search operation, the display interface can display different fields (i.e., the keyword fields shown in Tables 2 - 4) for further selection according to the specific interface type of the selected hard disk. If the user needs to query and monitor one or more fields, just check the corresponding options and click the "Generate Chart" button again. At this time, the system can present the trend of the selected field changing over time (i.e., the change trend curve of the hard disk failure indicator and time) in the form of a line chart, where the horizontal axis can represent time points and the vertical axis can represent the values of the field at different time points. Further, since the SMART data can exhibit linear characteristics, by monitoring the change trend curve of the hard disk failure indicator and time, the possible future change directions of certain indicators can be predicted. In a preferred embodiment of the present application, when the mouse hovers over any data point, when the client detects an event corresponding to any data point on the change trend curve of the hard disk failure indicator and time, a window can be popped up to display the corresponding value at that moment.

[0081] Further preferably, please refer to Figure 5 , the client display interface provided by the exemplary embodiment of the present application is configured to: in addition to being able to display a line chart, it can also list all relevant SMART data records in tabular form, which not only includes the basic information mentioned above, but also the data details of each keyword field. To facilitate browsing a large amount of information, the client display interface can also be configured to: adopt a paging design and allow adjusting the number of items displayed per page.

[0082] Further preferably, the client display interface can also be configured to support the function of viewing multiple fields simultaneously. That is, the user can select multiple fields to be monitored at one time, and then click "Generate Chart" to compare the change trend curves of different hard disk failure indicators and time in the same chart. Among them, for the key attributes in JSON format unique to SATA hard disks (for example, VALUE, WORST, THRESH, RAW_VALUE), the embodiments of the present application can all be parsed and plotted in the same chart, making the comparison more intuitive and convenient.

[0083] In a preferred embodiment of the present application, the client generates corresponding data showing the hard disk failure indicators according to the user interaction operation, using the SMART query data, and displays the corresponding view to the user to query and monitor the failure of the target hard disk. It may further include: obtaining the threshold value set by the user for the selected hard disk failure indicator to show the warning line corresponding to the hard disk failure indicator; and when the measured value of the selected hard disk failure indicator exceeds the threshold value, generating a corresponding view for prompting the failure.

[0084] Continuing with the above example, the embodiments of the present application also provide a function for setting the threshold value. The user can set the threshold value (also called the warning value) for showing the warning line corresponding to the hard disk failure indicator on a specific field; when the actual measurement result exceeds the threshold range, it can be marked with a horizontal line on the chart to help the user quickly identify the potential problem.

[0085] In a preferred embodiment of the present application, the client generates corresponding data showing the hard disk failure indicators according to the user interaction operation, using the SMART query data, and displays the corresponding view to the user to query and monitor the failure of the target hard disk. It may further include: generating a list of target hard disks that meet the filtering conditions according to the filtering conditions input by the user. Among them, for different types of target hard disks, the filtering conditions provided for the user to input are different.

[0086] Please refer to Figure 6 and Figure 7 Example, the client is configured to: a hard disk model (Model) query function, which aims to help users identify whether there are common problems with specific models of hard disks. For example, the user only needs to input the hard disk interface type, model, and time interval, and click the "Query" button, and the system will retrieve all the information of the hard disks under this model and temporarily store it in the cache. Subsequently, the user can set the filtering conditions according to the needs and click the "Generate Table" button. At this time, the hard disk serial numbers (SN) that meet the failure conditions will be listed. Preferably, for hard disks of different interface types, the available filtering conditions can also be different, for example, including:

[0087] 1) For SATA hard disks:

[0088] The overall health status of the corresponding hard disk shown by the SMART data is not "OK";

[0089] There are pre-failure items and their WHEN_FAlLED values are not "-", etc.

[0090] 2) For SAS hard disks:

[0091] The SMART health status of the corresponding hard disk shown by the SMART data is not "OK";

[0092] The read error counter is greater than 0;

[0093] The write error counter is greater than 0;

[0094] The verify error counter is greater than 0;

[0095] The number of non-media related errors is greater than 0, etc.

[0096] 3) For NVMe hard disks:

[0097] The critical warning flag bit is not equal to 0, etc.

[0098] After obtaining the required hard disk SN through the above steps, the user can use the above SN query function to retrieve the detailed SMART data and its key field data of each hard disk. By monitoring these data, the current status of the hard disk and the possible trend changes in the future can be judged more accurately.

[0099] Accordingly, the hard disk failure monitoring method for the server side provided by the embodiments of the present application extracts the SMART data of the hard disk, performs normalization processing on the SMART data, and then persists it in the memory; obtains the query instruction of the user for the SMART data of the target hard disk; in response to the query instruction, retrieves the corresponding SMART query data from the memory; and sends the SMART query data so that the client, according to the user operation, generates the data corresponding to the hard disk failure indicators by using the SMART query data and displays it to the user for querying and monitoring the failure of the target hard disk. When collecting the SMART data of the hard disk, the embodiments of the present application clean the collected SMART data, extract the key fields, and perform unified formatting processing. Subsequently, the processed SMART data is persistently stored in the memory (for example, MySQL database). When the user can access the database to query the SMART data, the data is processed and rendered and charted on the display interface (for example, browser) of the client. This ensures the integrity and security of all relevant data of the hard disk and provides a solid foundation for failure monitoring. The embodiments of the present application can also retrieve the SMART data of the target hard disk required from the database according to the specific needs of the user, and generate intuitive and easy-to-understand line charts and other forms of charts by using these data. In this way, the user can quickly monitor the change trend of the hard disk failure indicators of the selected hard disk over time, thereby assisting the user in monitoring or identifying potential problems of the hard disk. This not only helps to improve the hard disk failure prevention ability, but also provides strong support for maintaining the operation of the hard disk.

[0100] Furthermore, for the SMART data record in the embodiments of the present application, the manual recording is changed to automatic recording, which greatly reduces the error probability caused by human operation. Automatic recording not only reduces the labor input, but also enables all-weather and uninterrupted data collection. It can be easily completed by setting a scheduled task on the Docker server to send a query command to the server at regular intervals. For the SMART information management, it is upgraded from the traditional Excel table to the database storage method, which brings significant advantages. Among them, all information is centrally stored in a secure database, avoiding the risk of data loss and simplifying the unified management process; secondly, the SMART information is standardized by scripts, solving the problems such as inconsistent personal record formats in the past and the need to readjust when summarizing. In addition, the cleaning process of the SMART data is also automated. The scripts configured in the embodiments of the present application can accurately distinguish valid and invalid fields, greatly improving the work efficiency and accuracy. This transformation reduces the need for manual intervention while ensuring the quality of data processing.

[0101] Figure 8It is a schematic flowchart of a hard disk failure monitoring method provided by an exemplary embodiment of the present application. This monitoring method can be applied to a client. The monitoring method may include the following steps:

[0102] Step S810: Send a query instruction for SMART data of the target hard disk by the user.

[0103] Illustrated by an example, in the failure monitoring stage, the client of the system architecture can be developed using the Python programming language. The embodiments of the present application can retrieve the required information from the database according to the specific needs of the user, and use these data to generate visual data, for example, intuitive and easy-to-understand line charts and other forms of charts.

[0104] Please refer to Figure 4 the example. The user can, for example, on the web page of the user side, select the interface type (intf) of the target hard disk. The user can also input the serial number (SN) of the hard disk and the time interval in the serial number management interface, and click the "search" button to generate a query instruction for the SMART data of the target hard disk.

[0105] Please refer to Figure 6 the example. The user can, for example, on the web page of the user side, select the interface type (intf) of the target hard disk. The user can also input the hard disk model (Mode1) of the hard disk and the time interval in the hard disk model management interface, and click the "search" button to generate a query instruction for the SMART data of the target hard disk.

[0106] Step S820: Receive the corresponding SMART query data retrieved from the memory in response to the query instruction sent by the server side.

[0107] Please refer to Figure 4 the example. In response to the query instruction, the server side can retrieve the (historical) SMART query data of the hard disk SN within the corresponding time interval from the database in the memory. In a preferred embodiment of the present application, different data processing types can be automatically selected according to the interface type (intf) of the target hard disk, and the SMART query data of the hard disk is returned (for example, it may include the basic information of the hard disk and the data corresponding to the key fields in Table 2-4). Among them, the basic information of the hard disk includes, for example: serial number, manufacturer name, firmware version, interface type, media type, capacity size, and model, etc. In particular, for SATA hard disks, their pre-fail attribute and its corresponding WHEN_FAILED value can also be additionally displayed.

[0108] Please refer to Figure 6For example, in response to the query instruction, the server can retrieve the (historical) SMART query data of the hard disk Model within the corresponding time range from the database in the memory. In a preferred embodiment of the present application, different data processing types can be automatically selected according to the interface type (intf) of the target hard disk. Among them, the SMART data queried according to the Model indicates which hard disk SNs this hard disk model has; the user can select the fault type on the right page, for example, and click the "Generate SN Table" button to generate a list of SNs with this fault type.

[0109] Step S830: Render and display the SMART query data.

[0110] For example, the client renders and displays the SMART query data, and the user can query and monitor the faults of the target hard disk based on the SMART query data on the display interface of the client.

[0111] Step S840: According to the user operation, generate corresponding data showing the hard disk fault indicators using the SMART query data, and display it to the user for querying and monitoring the faults of the target hard disk.

[0112] Preferably, step S840 may include: generating a corresponding change trend curve of the hard disk fault indicator and time according to the selected hard disk fault indicator by the user to assist in predicting potential faults of the target hard disk; rendering and displaying the generated change trend curve of the hard disk fault indicator and time; and when an event corresponding to any data point on the change trend curve of the hard disk fault indicator and time is detected when the mouse hovers over it, a window pops up to display the corresponding value at that moment.

[0113] Please refer to Figure 4 and Figure 5For example, the client can be configured such that in the display interface of the client, after the user completes a search operation, the display interface can display different fields (i.e., the key fields shown in Table 2-4) for further selection according to the specific interface type of the selected hard disk. If the user needs to query and monitor one or more fields, the user only needs to check the corresponding options and click the "Generate Chart" button again. At this time, the system can present the trend of the selected field changing over time (i.e., the curve of the hard disk failure indicator changing with time) in the form of a line chart, where the horizontal axis can represent time points and the vertical axis can represent the values of the field at different time points. Further, since SMART data can exhibit linear characteristics, by monitoring the curve of the hard disk failure indicator changing with time, the possible future change directions of certain indicators can be predicted. In a preferred embodiment of the present application, when the mouse hovers over any data point, the client can detect the event corresponding to the mouse hovering over any data point on the curve of the hard disk failure indicator changing with time, and a window can pop up to display the value corresponding to that moment.

[0114] Please refer to Figure 5 , the client display interface provided by the exemplary embodiment of the present application is configured such that in addition to being able to display a line chart, it can also list all relevant SMART data records in tabular form, which includes not only the basic information mentioned above but also the data details of each key field. To facilitate browsing a large amount of information, the client display interface can also be configured to adopt a paging design and allow adjustment of the number of items displayed per page.

[0115] Further preferably, the client display interface can also be configured to support the function of viewing multiple fields simultaneously. That is, the user can select multiple fields to be monitored at one time, and then click "Generate Chart" to compare the curves of different hard disk failure indicators changing with time in the same chart. Among them, for the key attributes in JSON format specific to SATA hard disks (such as VALUE, WORST, THRESH, RAW_VALUE), the embodiments of the present application can all be parsed and plotted in the same chart, making the comparison more intuitive and convenient.

[0116] Preferably, step S840 may include: obtaining the threshold value set by the user for the selected hard disk failure indicator to show the warning line corresponding to the hard disk failure indicator; and when the measured value of the selected hard disk failure indicator exceeds the threshold value, generating a corresponding view for prompting a failure.

[0117] Continuing with the above example, the embodiments of the present application also provide a function for setting thresholds. The user can set a threshold (also referred to as a warning value) for the warning line indicating the corresponding hard disk failure index on a specific field; when the actual measurement result exceeds this threshold range, it can be marked with a horizontal line on the chart to help the user quickly identify the location of potential problems.

[0118] Preferably, step S840 may include: generating a list of target hard disks that meet the filtering conditions according to the filtering conditions input by the user. Among them, for different types of target hard disks, the filtering conditions provided for the user to input are different.

[0119] Please refer to Figure 6 and Figure 7 the example. The client can also be configured to: a hard disk model query function, which aims to help the user identify whether there are common problems with specific models of hard disks. For example, the user only needs to input the hard disk interface type, model, and time range, and click the "query" button, and the system will retrieve all the information of the hard disks under this model and temporarily store it in the cache. Subsequently, the user can set the filtering conditions according to the needs and click the "generate table" button. At this time, the hard disk serial numbers (SN) that meet the conditions will be listed. Preferably, for hard disks of different interface types, the available filtering conditions can also be different, for example, including:

[0120] 1) For SATA hard disks:

[0121] The overall health status of the corresponding hard disk shown by the SMART data is not "0K";

[0122] There are pre-failure items and their WHEN_FAILED value is not "-", etc.

[0123] 2) For SAS hard disks:

[0124] The SMART health status of the corresponding hard disk shown by the SMART data is not "0K";

[0125] The read error counter is greater than 0;

[0126] The write error counter is greater than 0;

[0127] The verify error counter is greater than 0;

[0128] The number of non-media-related errors is greater than 0, etc.

[0129] 3) For NVMe hard disks:

[0130] The critical warning flag bit is not equal to 0, etc.

[0131] It should be noted that for the hard disk failure monitoring method applied to the client provided in the embodiments of the present application, the specific implementation manner can refer to the description of the hard disk failure monitoring method applied to the server side in the above embodiments, and will not be elaborated here.

[0132] Accordingly, the hard disk failure monitoring method applied to the client provided in the embodiments of the present application includes: sending a query instruction for SMART data of the target hard disk by the user; receiving the corresponding SMART query data retrieved from the memory in response to the query instruction sent by the server side; and rendering and displaying the SMART query data, so as to generate corresponding data showing the hard disk failure indicators according to the user operation by using the SMART query data and display it to the user for querying and monitoring the failure of the target hard disk. When the user can access the database to query the SMART data, the embodiments of the present application process the data and render and generate a chart on the display interface (such as a browser) of the client. It can also retrieve the SMART data of the target hard disk required from the database according to the specific needs of the user and generate an intuitive and easy-to-understand line chart and other forms of charts by using these data. In this way, the user can quickly monitor the change trend of the hard disk failure indicators of the selected hard disk over time, thereby assisting the user in monitoring or identifying potential problems of the hard disk. This not only helps to improve the hard disk failure prevention ability, but also provides strong support for maintaining the operation of the hard disk.

[0133] Furthermore, the embodiments of the present application adopt a Web-based real-time chart generation method to replace the original Excel charting mode. This method has the following advantages: strong sharing, the generated chart can be immediately shared with other team members; instant response, after the user inputs the query conditions, the system can quickly generate the required chart, which is more efficient and convenient than Excel; trend visualization, intuitively display the change trend of the failure indicators, which helps to quickly locate problems or identify potential risks; safe and convenient, provide a unified access entrance, without worrying about the security of local files, all materials are stored in the cloud (i.e., the server side) and cannot be tampered with; wide applicability, as long as there is a network connection and a modern browser is used, the service can be accessed, and no additional software needs to be installed. The embodiments of the present application also support mainstream hard disk interface types such as SAS, SATA, and NVMe, and are applicable to various server environments equipped with hard disk drives. Its strong compatibility enables any type of hardware device to make full use of the benefits brought by this technology.

[0134] The embodiments of the present application also provide a control device, which may include: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the above-mentioned monitoring method applied to the server side.

[0135] An embodiment of the present application further provides a control device, which may include: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the above-mentioned monitoring method applied to the client.

[0136] An embodiment of the present application further provides a machine-readable storage medium, on which instructions are stored, and the instructions cause the machine to execute the above-mentioned monitoring method applied to the server side or the above-mentioned monitoring method applied to the client side.

[0137] Please refer to Figure 9 , an embodiment of the present application further provides a hard disk failure monitoring system, which may include a client and a server. Among them, the server includes the above-mentioned control device applied to the server side, and the client includes the above-mentioned control device applied to the client side.

[0138] It should be noted that the above-mentioned control device, machine-readable storage medium, and hard disk failure monitoring system can implement the hard disk failure monitoring method provided by the above-mentioned embodiments. The specific implementation manner can refer to the description of the hard disk failure monitoring method in the above-mentioned embodiments, and will not be elaborated here.

[0139] It can be understood that the circuit structures, names, and parameters described in the above embodiments are only examples. Those skilled in the art can also easily combine and adjust the structural features of the above-mentioned multiple embodiments according to the usage needs, and should not limit the concept of the present application to the specific details of the above examples.

[0140] Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A hard disk failure monitoring method, characterized in that: Applied to the server side, the monitoring method includes: Extracting SMART data of the hard disk, and performing standardization on the SMART data before persisting it in a memory; Obtain the user's query command for the SMART data of the target hard disk; In response to the query instruction, retrieving corresponding SMART query data from the memory; and The SMART query data is sent so that the client generates corresponding data showing hard disk failure indicators using the SMART query data according to user interaction operations, and displays the corresponding view to the user to query and monitor the target hard disk for failures.

2. The monitoring method according to claim 1, characterized in that: The extracting of the SMART data of the hard disk and performing standardization on the SMART data before persisting it in the memory includes: The Docker server uses a preset encapsulation library to send a command to query SMART data to at least one server; In response to the instruction sent by the Docker server to query SMART data, the at least one server returns the SMART data; The Docker server converts the returned SMART data into a preset format using a preset function to complete standardized processing of the SMART data; and The standardized SMART data is persisted in the database of the memory.

3. The monitoring method according to claim 2, characterized in that: The preset function is configured to call a corresponding processing function according to different interface types of the hard disk to convert the SMART data into the preset format. The preset format includes JSON format.

4. The monitoring method according to claim 1, characterized in that: The query instruction includes the interface type of the target hard disk and the time interval to be queried, and also includes the hard disk serial number or hard disk model. After the corresponding SMART query data is retrieved from the memory in response to the query instruction, the monitoring method further includes: Determine the target hard disk to be searched according to the hard disk serial number or the hard disk model; and / or According to the interface type of the target hard disk, the SMART query data of different data types are retrieved accordingly.

5. A hard disk failure monitoring method, characterized in that: Applied to the client, the monitoring method includes: Send the user a query command for the SMART data of the target hard disk; Receiving corresponding SMART query data sent by the server in response to the query instruction and retrieved from the memory; Rendering and displaying the SMART query data; and According to the user interaction operation, the SMART query data is used to generate corresponding data showing hard disk failure indicators, and the corresponding view is displayed to the user to query and monitor the failure of the target hard disk.

6. The monitoring method according to claim 5, characterized in that: According to the user operation, the SMART query data is used to generate corresponding data showing hard disk failure indicators, and the corresponding view is displayed to the user to query and monitor the failure of the target hard disk, including: According to the hard disk failure index selected by the user, a corresponding hard disk failure index and time change trend curve is generated to assist in predicting the potential failure of the target hard disk; Rendering and displaying the generated hard disk failure indicator and time trend curve; and When an event corresponding to a mouse hovering over any data point on the hard disk failure indicator and time variation trend curve is detected, a window pops up to display the corresponding value at that moment.

7. The monitoring method according to claim 5, characterized in that: According to the user operation, the SMART query data is used to generate corresponding data showing hard disk failure indicators, and the corresponding view is displayed to the user to query and monitor the failure of the target hard disk, including: According to the filter conditions entered by the user, a list of target hard disks that meet the filter conditions is generated. Among them, for different types of target hard disks, different screening conditions are provided to the user for input.

8. The monitoring method according to claim 5, characterized in that: According to the user operation, the SMART query data is used to generate corresponding data showing hard disk failure indicators, and the corresponding view is displayed to the user to query and monitor the failure of the target hard disk, including: Obtaining a threshold value set by a user for a selected hard disk failure indicator for indicating a warning line of the corresponding hard disk failure indicator; and When the measured value of the selected hard disk failure indicator exceeds the threshold, a corresponding view for prompting the failure is generated.

9. A control device, characterized in that: The control device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the monitoring method according to any one of claims 1 to 4.

10. A control device, characterized in that: The control device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the monitoring method according to any one of claims 5-8.

11. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores instructions, which enable the machine to execute the monitoring method according to any one of claims 1-4 or the monitoring method according to any one of claims 5-8.

12. A hard disk fault monitoring system, characterized in that: The monitoring system includes a client and a server. The server side includes the control device according to claim 9, The client comprises the control device according to claim 10.

Citation Information

Patent Citations

  • Method and device of dynamically diagnosing hard disk failure based on S.M.A.R.T (Self-Monitoring Analysis and Reporting Technology) data

    CN105260279A