Hard disk state monitoring method and device, hard disk bracket and computer equipment

By evaluating the hard disk status and failure risks and generating alarm information, the problem that traditional hard disk monitoring methods cannot predict hard disk failures is solved, real-time and accurate health assessment and fault warning of hard disk status are realized, and data storage security and hard disk stability are ensured.

CN120086094APending Publication Date: 2025-06-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510225731.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional hard disk monitoring methods cannot predict whether the hard disk is on the verge of failure, resulting in the impact of data storage security and server stability.

Method used

By obtaining the hard disk status parameters, determining the first evaluation parameters and the second evaluation parameters, performing logical operations to evaluate the hard disk status and failure risk, and generating alarm information to alert potential failures.

Benefits of technology

Real-time and accurate health assessment of the hard disk status is realized, timely predicting hard disk failures, ensuring data storage security, and enhancing hard disk stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086094A_ABST
    Figure CN120086094A_ABST
Patent Text Reader

Abstract

The invention discloses a hard disk state monitoring method and device, a hard disk bracket and computer equipment, and relates to the technical field of computers, and the method comprises the steps: determining a hard disk state evaluation value through a hard disk state parameter and a first evaluation parameter, and determining a hard disk risk evaluation value through the hard disk state parameter and a second evaluation parameter, and determining a hard disk in a fault risk state by using the hard disk state assessment value and the hard disk risk assessment value, and generating alarm information. The problems that whether the hard disk is endangered to a fault or not cannot be pre-judged, fault warning cannot be conducted in time, and data storage safety is affected are solved. The technical effects of performing real-time and accurate health assessment on the hard disk, monitoring the state of the hard disk, pre-judging the fault of the hard disk in time, guaranteeing the data storage safety and enhancing the stability and reliability of the hard disk are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for monitoring the status of a hard disk, a hard disk tray, and a computer device. Background Art

[0002] In the current data center and server operating environment, the hard disk is a key data storage component, and the stable operation of the hard disk is crucial. It is necessary to detect the status of the hard disk.

[0003] However, in the long-term operation of the server, the traditional hard disk monitoring method can only give an alarm for hard disk failure after the hard disk fails, and the hard disk failure will affect the operation of the data center and the server. It is impossible to predict whether the hard disk is on the verge of failure and give an alarm in time, resulting in a lag in the early warning of potential hard disk failures, seriously affecting the data storage security of the data center and the server, and affecting the stability and reliability of the server.

[0004] Therefore, the related technology has the problem that it is impossible to predict whether the hard disk is on the verge of failure and give an alarm in time, which affects data storage security. Summary of the Invention

[0005] In view of this, the present invention provides a method and device for monitoring the status of a hard disk, a hard disk tray, and a computer device to solve the problem that it is impossible to predict whether the hard disk is on the verge of failure and give an alarm in time, which affects data storage security.

[0006] In a first aspect, the present invention provides a method for monitoring the status of a hard disk. The method includes:

[0007] Obtain hard disk status parameters, and determine a first evaluation parameter and a second evaluation parameter corresponding to the hard disk status parameters, where the first evaluation parameter is used to determine the influence degree of the hard disk status parameters on the working state of the hard disk, and the second evaluation parameter is used to determine the influence degree of the hard disk status parameters on the hard disk failure risk;

[0008] Perform a first preset logical operation on the hard disk status parameters and the first evaluation parameter to obtain a hard disk status evaluation value;

[0009] Perform a second preset logical operation on the hard disk status parameters and the second evaluation parameter to obtain a hard disk risk evaluation value;

[0010] Determine the hard disk in the failure risk state based on the hard disk status evaluation value and the hard disk risk evaluation value, and generate a first alarm message.

[0011] In a second aspect, the present invention provides a hard disk tray, which includes: a display component, a connection line, and an intelligent control and collection component;

[0012] The intelligent control and collection component is used to execute the hard disk status monitoring method of the first aspect or any corresponding implementation manner thereof;

[0013] The intelligent control and collection component is connected to the display component through a connection line. The intelligent control and collection component is used to transmit the hard disk status parameters and the first warning information to the display component through the connection line;

[0014] The display component is used to display the hard disk status parameters and the first warning information.

[0015] In a third aspect, the present invention provides a hard disk status monitoring device, which is deployed in the intelligent control and collection component. The device includes:

[0016] A parameter acquisition module, configured to acquire hard disk status parameters, and determine a first evaluation parameter and a second evaluation parameter corresponding to the hard disk status parameters. Among them, the first evaluation parameter is used to determine the influence degree of the hard disk status parameters on the hard disk working state, and the second evaluation parameter is used to determine the influence degree of the hard disk status parameters on the hard disk failure risk;

[0017] A first parameter processing module, configured to perform a first preset logical operation on the hard disk status parameters and the first evaluation parameter to obtain a hard disk status evaluation value;

[0018] A second parameter processing module, configured to perform a second preset logical operation on the hard disk status parameters and the second evaluation parameter to obtain a hard disk risk evaluation value;

[0019] A status determination module, configured to determine the hard disks in a failure risk state based on the hard disk status evaluation value and the hard disk risk evaluation value, and generate a first warning information.

[0020] In a fourth aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the hard disk status monitoring method of the first aspect or any corresponding implementation manner thereof.

[0021] In a fifth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored. The computer instructions are used to cause a computer to execute the hard disk status monitoring method of the first aspect or any corresponding implementation manner thereof.

[0022] In a sixth aspect, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the hard disk status monitoring method of the first aspect or any corresponding implementation manner thereof.

[0023] Through this application, a hard disk status evaluation value is determined using hard disk status parameters and a first evaluation parameter, a hard disk risk evaluation value is determined using the hard disk status parameters and a second evaluation parameter, a hard disk in a failure risk state is determined using the hard disk status evaluation value and the hard disk risk evaluation value, and an alarm message is generated. This solves the problem that it is impossible to predict whether the hard disk is on the verge of failure and give a timely failure alarm, which affects data storage security. The technical effect is to achieve real-time and accurate health evaluation of the hard disk, monitor the hard disk status, predict hard disk failures in a timely manner, ensure data storage security, and enhance the stability and reliability of the hard disk. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in related technologies, the following will briefly introduce the drawings required for use in the description of the specific embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is a flowchart of a method for monitoring the status of a hard disk according to an embodiment of the present invention;

[0026] Figure 2 is a flowchart of a method for intelligently monitoring the status of a hard disk according to an embodiment of the present invention;

[0027] Figure 3 is a structural diagram of a hard disk tray according to an embodiment of the present invention;

[0028] Figure 4 is a block diagram of a structure of a hard disk status monitoring device according to an embodiment of the present invention;

[0029] Figure 5 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0031] Traditional hard disk monitoring methods often rely on the server system to detect hard disks. When the server is powered on initially or when the server system fails, the server system cannot obtain the hard disk status information in a timely manner, resulting in a lag in the early warning of potential hard disk failures, which seriously affects the stability and reliability of the data center and the server. In addition, during the long-term operation of the server, traditional hard disk monitoring methods do not have a pre-judgment mechanism to determine whether the hard disk is on the verge of failure, which will also lead to a lag in the early warning of potential hard disk failures and affect the stability and reliability of the data center and the server.

[0032] Based on the above, embodiments of the present invention provide a hard disk status monitoring method. When the server is initially powered on, the intelligent management and collection module starts quickly without waiting for the entire machine system to start, and immediately starts the inspection of the hard disk health status, gaining the initiative for subsequent hard disk status monitoring and ensuring the initial stability of the server system. Among them, the intelligent management and collection module is built-in with an intelligent programmable logic core (Integrated Power Loss Cache Controller, IPLCC). During the long-term operation of the server, the intelligent management and collection module regularly performs automatic inspections on the hard disk according to a preset strategy, continuously obtaining dynamic information such as the hard disk temperature, rotation speed, power-on duration, and seek error rate. At the same time, the intelligent management and collection module can adaptively adjust the monitoring focus and key parameter weights for the hard disk according to the change of working conditions, and accurately evaluate the hard disk health. Once a hard disk anomaly is detected, an immediate warning is issued. Through the above method, when the server is initially powered on, the intelligent management and collection module does not need to wait for the server system to be deployed, quickly collects key information of the hard disk and issues a warning, improving the initial monitoring efficiency and accuracy, and ensuring the rapid and stable startup of the system. During the long-term operation stage of the server, the intelligent management and collection module relies on the preset inspection and machine learning algorithms to analyze the hard disk indicators in real time, adaptively adjust the monitoring focus, cope with read and write pressure and environmental interference, detect hidden dangers in advance and issue warnings, enhancing the stability and reliability of the system. To achieve the effect of monitoring the hard disk status and predicting hard disk failures in a timely manner to ensure data storage security.

[0033] According to an embodiment of the present invention, an embodiment of hard disk status monitoring is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer device with data processing capabilities, such as a computer, a server, a mobile terminal, etc. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0034] In this embodiment, a hard disk status monitoring method is provided, which can be used for the above computer device. Figure 1 It is a flowchart of the hard disk status monitoring method according to an embodiment of the present invention, as Figure 1 shown, and the process includes the following steps:

[0035] Step S101: Obtain the hard disk status parameters, and determine the first evaluation parameter and the second evaluation parameter corresponding to the hard disk status parameters. The first evaluation parameter is used to determine the influence degree of the hard disk status parameters on the working status of the hard disk, and the second evaluation parameter is used to determine the influence degree of the hard disk status parameters on the hard disk failure risk.

[0036] Specifically, the intelligent control and collection module is the key core of the hard disk status monitoring method and is used to inspect the health status of the hard disk. During the long-term operation of the server, the intelligent control and collection module performs regular automatic inspections according to a preset strategy, and obtains the hard disk status parameters according to the S.M.A.R.T. (Self-Monitoring Analysis and Reporting Technology) data indicators. The hard disk status parameters are, for example, dynamic information such as the temperature, rotation speed, power-on duration, and seek error rate of the hard disk.

[0037] In addition, in order to determine whether there are potential risks in the hard disk, this embodiment sets a hard disk status evaluation value and a hard disk risk evaluation value. Through the hard disk status evaluation value and the hard disk risk evaluation value, the risk status of the hard disk can be quantitatively reflected. In order to calculate the hard disk status evaluation value of the hard disk, it is necessary to determine the first evaluation parameter corresponding to each hard disk status parameter, such as the data validity coefficient, feature importance factor, sensitivity adjustment coefficient, and first weight of the hard disk status parameter. The first weight represents the influence degree of each hard disk status parameter on the hard disk status evaluation value. In order to calculate the hard disk risk evaluation value of the hard disk, it is necessary to determine the second evaluation parameter, such as the second weight. The first weight represents the influence degree of each hard disk status parameter on the hard disk risk evaluation value.

[0038] Step S102: Perform a first preset logical operation on the hard disk status parameters and the first evaluation parameter to obtain the hard disk status evaluation value.

[0039] Specifically, the hard disk status evaluation value is, for example, the comprehensive evaluation value of key parameters (CPV). The intelligent management and control collection module includes a set of hard disk health evaluation algorithms. This algorithm deeply integrates the concept of machine learning. Based on the various indicators of the hard disk status parameters, it determines the degree of difference in the impact of each hard disk status parameter on the hard disk status, determines the degree of influence of different hard disk status parameters on the hard disk status evaluation value, and then uses the first preset logical operation to calculate the hard disk status evaluation value. For example: calculate the first product of each hard disk status parameter and the corresponding first evaluation parameter, and sum up the first products corresponding to each hard disk status parameter as the hard disk status evaluation value; select some parameters in the first evaluation parameter as coefficients, determine the difference between the hard disk status parameter and the parameter average value, and use the remaining parameters in the first evaluation parameter to adjust the influence degree of the difference on the hard disk status evaluation value, and sum up the difference and coefficients corresponding to each hard disk status parameter as the hard disk status evaluation value. Through the hard disk status evaluation value, the current key parameter status of the hard disk can be comprehensively and intelligently reflected, laying a solid foundation for accurately evaluating the health status of the hard disk in the future.

[0040] Step S103: Perform a second preset logical operation on the hard disk status parameters and the second evaluation parameters to obtain the hard disk risk evaluation value.

[0041] Specifically, the hard disk risk evaluation value is, for example, the risk evaluation index (RI). In order to achieve a more accurate intelligent judgment of the hard disk status, the hard disk risk evaluation value is used to quantitatively reflect the current risk status of the hard disk. Perform a second preset logical operation on the hard disk status parameters and the second evaluation parameters. For example: calculate the second product of each hard disk status parameter and the corresponding second evaluation parameter, and sum up the second products corresponding to each hard disk status parameter as the hard disk risk evaluation value. The hard disk risk evaluation value provides a more forward-looking decision-making basis for the operation and maintenance personnel, effectively reducing the risk of data loss caused by sudden hard disk failures.

[0042] Step S104: Determine the hard disks in the fault risk state based on the hard disk status evaluation value and the hard disk risk evaluation value, and generate the first warning message.

[0043] Specifically, a normal threshold range is preset for the hard disk status evaluation value in this embodiment. When the hard disk status evaluation value exceeds the preset normal threshold range, the system immediately triggers a detailed key parameter analysis process to check for abnormal parameter values. For example: if the hard disk status evaluation value exceeds the preset normal threshold range and there has been no hard disk status evaluation value with the same numerical value before, generate the first warning message; if there has been a hard disk status evaluation value with the same numerical value before, determine whether the hard disk status parameter exceeds the corresponding threshold. If so, generate the first warning message. The first warning message is used to prompt the operation and maintenance personnel to quickly take corresponding measures to eliminate the warning.

[0044] Set a safety threshold and a warning threshold for the hard disk risk assessment value. When the hard disk risk assessment value is lower than the set safety threshold, it is determined that the hard disk is in a low-risk state and runs relatively stably; when the hard disk risk assessment value is between the set safety threshold and the warning threshold, it is determined that the hard disk is in a potential risk state and needs attention; when the hard disk risk assessment value exceeds the warning threshold, it is determined that the hard disk is in a high-risk state, that is, a failure risk state, and is about to face a failure risk, and a first warning message is generated to notify the operation and maintenance personnel. In addition, when encountering complex and changeable hard disk working conditions, the intelligent model can quickly learn and adaptively adjust the judgment strategy, provide more forward-looking decision-making basis for the operation and maintenance personnel, and effectively reduce the risk of data loss caused by sudden hard disk failures.

[0045] The hard disk status monitoring method provided in this embodiment determines the hard disk status assessment value through the hard disk status parameters and the first evaluation parameter, determines the hard disk risk assessment value through the hard disk status parameters and the second evaluation parameter, determines the hard disk in the failure risk state by using the hard disk status assessment value and the hard disk risk assessment value, and generates a warning message. It conducts real-time and accurate health assessment of the hard disk, monitors the hard disk status and predicts hard disk failures in a timely manner, ensures data storage security, and enhances the stability and reliability of the server. It solves the problem that it is impossible to predict whether the hard disk is on the verge of failure and give a failure warning in a timely manner, which affects data storage security.

[0046] In some optional implementation manners, determining the hard disk in the failure risk state based on the hard disk status assessment value and the hard disk risk assessment value includes:

[0047] Determine the current hard disk status parameter and the historical hard disk status parameter in the hard disk status parameters, and determine the parameter range according to the historical hard disk status parameter;

[0048] Determine the current hard disk status assessment value and the historical hard disk status assessment value in the hard disk status assessment value;

[0049] In the case that the current hard disk status assessment value is not within the preset range, determine whether there is at least one historical hard disk status assessment value equal to the current hard disk status assessment value;

[0050] In the case that there is no historical hard disk status assessment value equal to the current hard disk status assessment value, determine that the hard disk corresponding to the current hard disk status assessment value is in the failure risk state, and generate a first warning message;

[0051] In the case that there is at least one historical hard disk status assessment value equal to the current hard disk status assessment value, determine whether the current hard disk status parameter is within the parameter range;

[0052] In the case that there is at least one current hard disk status parameter not within the parameter range, determine that the hard disk corresponding to the current hard disk status assessment value is in the failure risk state, and generate a first warning message.

[0053] Specifically, the intelligent control and collection module includes a health assessment algorithm, which is closely associated with the results of the health assessment algorithm, the key parameters, parameter weight values, and thresholds preset in the initial environment. During the long-term power-on of the hard disk, timely correction will be made based on the long-term learning and optimization results of the health assessment algorithm. The alarm program in the health assessment algorithm quickly extracts key parameters from a large amount of hard disk information obtained, analyzes the details based on the preset parameter weights and thresholds, and combines the hard disk status evaluation value transmitted by the health assessment algorithm to perform real-time and accurate health assessment of the hard disk. During this operation process, once it is found that the key parameters are close to or exceed the threshold, the intelligent control and collection module immediately generates a first alarm message to prompt the operation and maintenance personnel to quickly take corresponding measures to eliminate the warning.

[0054] Determine the current hard disk status parameter collected at the current moment among the hard disk status parameters, and determine the historical hard disk status parameter collected before the current moment. Determine the parameter range according to the historical hard disk status parameter. For example, if the hard disk status parameter is temperature, and the historical hard disk status parameters include 20°C, 60°C, 40°C, 45°C, etc., then the parameter range is 20°C to 60°C.

[0055] Determine the current hard disk status evaluation value calculated at the current moment among the hard disk status evaluation values, and determine the historical hard disk status evaluation value calculated before the current moment. For example, the current hard disk status evaluation value is 1.2, and the historical hard disk status evaluation values include 1.0, 1.1, 1.2, etc.

[0056] The preset range is, for example: 1.0 to 1.2, 1.0 to 1.15… or other numerical ranges that meet the actual requirements. In the case where the current hard disk status evaluation value is not within the preset range, determine whether there is at least one historical hard disk status evaluation value equal to the current hard disk status evaluation value. For example, the current hard disk status evaluation value is 1.3, the preset range is 1.0 to 1.2, the current hard disk status evaluation value is not within the preset range, and the historical hard disk status evaluation values include 1.0, 1.1, 1.2, etc. Therefore, there is at least one historical hard disk status evaluation value equal to the current hard disk status evaluation value.

[0057] If there is no historical hard disk status evaluation value equal to the current hard disk status evaluation value, determine that the hard disk corresponding to the current hard disk status evaluation value is in a fault risk state, that is, it is about to face a fault risk, and generate a first alarm message to notify the operation and maintenance personnel. For example, the first alarm message includes information such as the hard disk has a potential fault and the current hard disk status evaluation value.

[0058] If there is at least one historical hard disk status evaluation value equal to the current hard disk status evaluation value, it is necessary to further determine whether the current hard disk status parameter is within the parameter range. For example, if the current temperature in the current hard disk status parameter is 65°C and the corresponding parameter range is 20°C to 60°C, then the current hard disk status parameter is not within the parameter range. If there is at least one current hard disk status parameter not within the parameter range, it is determined that the hard disk corresponding to the current hard disk status evaluation value is in a fault risk state, and a first warning message is generated.

[0059] In addition, if the power-on duration of the hard disk in the current hard disk status parameter is relatively short, such as 10 hours, 15 hours, etc., it indicates that the hard disk is a newly powered-on hard disk. If the hard disk temperature exceeds the parameter range, no warning needs to be issued.

[0060] In this embodiment, by using the hard disk status parameter, the hard disk status evaluation value, and the preset range, the hard disk in the fault risk state is determined and timely warning is given to remind the administrator that the hard disk may be in the fault risk state. Moreover, by comparing the current data with the historical data, the possibility of false alarms is reduced. Only when the current status evaluation value does not match the historical data and the parameter exceeds the range, a warning message will be generated, improving the accuracy of fault judgment.

[0061] In some optional embodiments, the method further includes:

[0062] When the current hard disk status evaluation value is within the preset range and there is no current hard disk status parameter not within the parameter range, determine whether the hard disk risk evaluation value is higher than the preset warning threshold;

[0063] When the hard disk risk evaluation value is higher than the preset warning threshold, determine that the hard disk corresponding to the hard disk risk evaluation value is in a fault risk state, and generate a first warning message;

[0064] When there is a current hard disk status parameter exceeding the first preset threshold, the hard disk corresponding to the current hard disk status parameter is in a fault risk state, and a first warning message is generated.

[0065] Specifically, in this embodiment, a preset warning threshold and a preset safety threshold are set for judging the hard disk risk evaluation value. If the current hard disk status evaluation value is within the preset range and there is no current hard disk status parameter not within the parameter range, it is determined that the hard disk is operating normally at the current moment. However, the hard disk may have potential faults, which will affect the security of the server and the data center if a fault occurs subsequently. Therefore, it is necessary to use the hard disk risk evaluation value to determine whether the hard disk is in a high-risk state, that is, whether it will face the risk of failure. Determine whether the hard disk risk evaluation value is higher than the preset warning threshold.

[0066] If the hard disk risk assessment value is higher than the preset warning threshold, it is determined that the hard disk corresponding to the hard disk risk assessment value is in a high-risk state, that is, a failure risk state. The hard disk is about to face a failure risk, and a first warning message is generated to notify the operation and maintenance personnel. In addition, when the hard disk risk assessment value is lower than the preset safety threshold, it is determined that the hard disk is in a low-risk state and runs relatively stably; when the hard disk risk assessment value is between the preset safety threshold and the preset warning threshold, it is determined that the hard disk is in a potential risk state, and this hard disk needs to be focused on.

[0067] In this embodiment, a first preset threshold is set for the current hard disk status parameter. For example, the first preset threshold for the temperature in the current hard disk status parameter is 60 °C, and the first preset threshold for the rotation speed is 1000 revolutions per minute. During the above operation process, the current hard disk status parameter is compared with the corresponding first preset threshold. Once it is found that there is a current hard disk status parameter exceeding the first preset threshold, it is determined that the hard disk corresponding to the current hard disk status parameter is in a failure risk state, and the intelligent management and collection module immediately triggers an alarm for the corresponding parameter, generating a first warning message to prompt the operation and maintenance personnel to quickly take corresponding measures to eliminate the warning.

[0068] It should be noted that the first warning message is used to prompt the operation and maintenance personnel to quickly take corresponding measures to eliminate the warning. After the warning is eliminated, the collection module quickly triggers the intelligent collection information function again, seamlessly enters a new round of key parameter extraction and threshold comparison, and so on in a cycle until there is no warning exceeding the threshold, and the hard disk status assessment value recalculated in combination with the health assessment algorithm is within the stable and normal range, then the next step of the disk health status assessment will be steadily promoted.

[0069] In this embodiment, the hard disk risk assessment value, the preset warning threshold, the current hard disk status parameter, and the first preset threshold are used to determine the hard disk in the failure risk state and give an alarm in time, preventing problems caused by disk damage as early as possible, and effectively reducing the risk of data loss caused by sudden hard disk failures.

[0070] In some alternative embodiments, the first evaluation parameter includes a data validity coefficient, a feature importance factor, a sensitivity adjustment coefficient, and a first weight. A first preset logical operation is performed on the hard disk status parameter and the first evaluation parameter to obtain a hard disk status assessment value, including:

[0071] Taking the product of the data validity coefficient, the feature importance factor, and the first weight as the first intermediate parameter;

[0072] Determining the parameter average value of the hard disk status parameter and determining the difference between the hard disk status parameter and the parameter average value;

[0073] According to the difference, the sensitivity adjustment coefficient, and the preset exponent, obtaining a second intermediate parameter;

[0074] Take the ratio of the first intermediate parameter and the second intermediate parameter as the third intermediate parameter, and obtain the hard disk status evaluation value according to the third intermediate parameter corresponding to the hard disk status parameter.

[0075] Specifically, the hard disk status evaluation value is, for example, the Comprehensive Parameter Value (CPV). The first evaluation parameter includes the data validity coefficient, the feature importance factor, the sensitivity adjustment coefficient, and the first weight. The data validity coefficient is, for example, Ei, which represents the original data validity coefficient of the i-th hard disk feature index, and its value range is between 0 and 1. This coefficient is dynamically generated by the intelligent model based on historical data and real-time data quality analysis, and is used to measure the reliability of the currently obtained index data, avoiding incorrect evaluations caused by sensor failures or abnormal data transmissions. For example, if the temperature sensor data fluctuates abnormally at a certain moment, Ei will decrease accordingly, weakening the impact of this abnormal data on the overall evaluation. The feature importance factor is, for example, Fi, which represents the feature importance factor of the i-th hard disk feature index, and is obtained by the intelligent model through learning a large number of failure cases under different hard disk models and operating conditions, reflecting the key degree of this index to the overall health status in the current hard disk operating state. For example, for a hard disk that often performs high-intensity read and write operations, the Fi value corresponding to the read and write error rate will increase significantly during a specific period. The sensitivity adjustment coefficient is, for example, Ki, which represents the sensitivity adjustment coefficient of the i-th hard disk feature index, and is also obtained by the intelligent model based on historical data learning and optimization, and is used to control the impact degree of the index deviation from the mean value on the comprehensive evaluation value, making the algorithm more intelligent and reasonable in responding to changes in key parameters. For example, for a hard disk that is more sensitive to temperature, the Ki value is relatively high, and once the temperature deviates from the mean value, it will have a greater impact on the CPV. The first weight is, for example, Wi, which represents the weight of the i-th hard disk feature index. Similar to the traditional weight, but here, it can be adaptively adjusted not only according to experience and research, but also according to the changes in the hard disk operating environment (such as temperature, humidity, power supply stability, etc.) monitored by the intelligent model in real time and the data rules accumulated during long-term operation, ensuring that the weight always fits the actual situation and accurately reflects the impact degree of each index on the health score.

[0076] In order to achieve more accurate and intelligent acquisition and evaluation of key parameters in this embodiment, a dynamic key parameter acquisition formula is introduced, for example, formula (1).

[0077] CPV = ∑[(Ei × Fi × Wi) ÷ (1 + Exp(-Ki × (Vi - Vi_mean)))] (1)

[0078] Among them, Vi represents the current value of the i-th hard disk status parameter, such as the real-time reading of temperature, the current percentage of read / write error rate, etc. Vi_mean represents the parameter average value of the i-th hard disk status parameter. The intelligent model will continuously record and update the average value of each metric in different time periods (such as the past hour, day, week, etc.), and use this as an important reference for judging whether the current value is abnormal, making the evaluation more dynamic and accurate.

[0079] The process of calculating the hard disk status evaluation value using formula (1) includes: taking the product of the data validity coefficient, the feature importance factor, and the first weight as the first intermediate parameter, for example, Ei×Fi×Wi. Determine the parameter average value Vi_mean of the hard disk status parameter, and determine the difference between the hard disk status parameter and the parameter average value, for example, Vi - Vi_mean.

[0080] According to the difference, the sensitivity adjustment coefficient, and the preset exponent, obtain the second intermediate parameter, for example, 1 + Exp(-Ki×(Vi - Vi_mean)).

[0081] Take the ratio of the first intermediate parameter and the second intermediate parameter as the third intermediate parameter, for example, (Ei×Fi×Wi)÷(1 + Exp(-Ki×(Vi - Vi_mean))). Aggregate the third intermediate parameters corresponding to each hard disk status parameter to obtain the hard disk status evaluation value.

[0082] The comprehensive evaluation value of the key parameters calculated by formula (1) can comprehensively and intelligently reflect the current key parameter status of the hard disk, laying a solid foundation for accurately evaluating the hard disk health status in the future. When the CPV value exceeds the pre-set normal threshold range, the system immediately triggers a detailed key parameter analysis process to check for abnormal parameter values. By comparing the historical parameter values of machine learning with the preset normal values, it is determined whether to trigger an alarm for the key parameter module, and at the same time, the potential risk information is synchronized in real time in the past, providing in-depth data support for alarm triggering to ensure the timeliness and accuracy of hard disk operation and maintenance.

[0083] In this embodiment, the hard disk status evaluation value is calculated according to the data validity coefficient, the feature importance factor, the sensitivity adjustment coefficient, the first weight, and the hard disk status parameter. The current working status of the hard disk is quantitatively reflected through the hard disk status evaluation value, which is convenient for determining the health status of the hard disk and determining the hard disks in the fault risk state.

[0084] In some optional embodiments, the second evaluation parameter includes a second weight. A second preset logical operation is performed on the hard disk status parameter and the second evaluation parameter to obtain the hard disk risk evaluation value, including:

[0085] Generate a fourth intermediate parameter according to the second weight, the hard disk status parameter, and the preset logarithm, and use the sum of the fourth intermediate parameters as the fifth intermediate parameter;

[0086] Use the sum of the second weights as the sixth intermediate parameter;

[0087] Use the ratio of the fifth intermediate parameter to the sixth intermediate parameter as the hard disk risk assessment value.

[0088] Specifically, the hard disk risk assessment value is, for example, the Risk Index (RI). To achieve more accurate intelligent judgment, in this embodiment, an intelligent logic judgment formula is designed in the alarm program of the intelligent control and collection module, such as formula (2). The second weight is, for example, Pi, which represents the weight of the i-th hard disk status parameter for the hard disk risk assessment value. In the above embodiment, the key parameter analysis process is triggered according to the hard disk status assessment value. In this embodiment, the value of the second weight is adjusted according to the evaluation result of the key parameter analysis process. For example, after triggering the key parameter analysis based on the hard disk status assessment value and determining that no alarm is required, the new evaluation result is passed to the key parameter alarm program. The new evaluation result is, for example, when the hard disk temperature is 60 degrees, the power-on duration of a new disk is 1 hour, the hard disk status assessment value is 1.2, and no alarm is required according to historical judgment. The alarm program needs to adjust the weight of the hard disk temperature, which is a hard disk status parameter, based on a series of new information sources such as the transmitted hard disk temperature and the power-on duration of the new disk. Since the power-on duration is short and it is judged as a new disk replacement, its weight in the alarm formula is reduced. This design is equivalent to multiple monitoring of the hard disk health status and early warning.

[0089] RI = v[Pi × Log a (1 + Vi)] ÷ ∑Pi (2)

[0090] Where a is a positive integer, and the specific value of a is set according to actual needs, such as 2, 10, etc. Vi represents the current value of the i-th hard disk status parameter, and the value range of i is from 1 to m (m is the total number of hard disk characteristic indicators participating in the evaluation).

[0091] The process of calculating the hard disk risk assessment value using formula (2) includes: generating a fourth intermediate parameter according to the second weight, the hard disk status parameter, and the preset logarithm, such as Pi × Log a (1 + Vi). Use the sum of the fourth intermediate parameters as the fifth intermediate parameter, and the fifth intermediate parameter is, for example, ∑[Pi × Log a (1 + Vi)]. Use the sum of the second weights as the sixth intermediate parameter, and the sixth intermediate parameter is, for example, ∑Pi. Use the ratio of the fifth intermediate parameter to the sixth intermediate parameter as the hard disk risk assessment value.

[0092] In this embodiment, a hard disk risk assessment value is calculated based on the second weight and the hard disk status parameter, and the risk status of the hard disk at present is quantitatively reflected through the hard disk risk assessment value, which is convenient for predicting potential faults of the hard disk.

[0093] In some alternative embodiments, obtaining the hard disk status parameter includes:

[0094] Obtaining the to-be-processed status data;

[0095] Obtaining intermediate status data and the data type of the intermediate status data from the to-be-processed status data according to the hard disk data protocol parsing library;

[0096] Converting the intermediate status data and the data type into hard disk status parameters according to the physical quantity unit mapping table.

[0097] Specifically, after the tray device loads the hard disk, accesses the hard disk backplane, and powers on, the intelligent management and collection module is activated. The built-in intelligent programmable logic core starts its own firmware and obtains hard disk-related information as the to-be-processed status data according to the S.M.A.R.T. data index.

[0098] The intelligent programmable logic core of the intelligent management and collection module starts the internal parsing program. The parsing program invokes the hard disk data protocol parsing library, which covers the hard disk conventional data format specifications. Based on the hard disk data protocol parsing library, the parsing program can quickly and accurately identify different types of data fields and classify and temporarily store them. For example, numerical information such as temperature and rotation speed and text information such as hard disk bad sectors are classified into numerical values and texts respectively. Intermediate status data and the data type of the intermediate status data are obtained from the to-be-processed status data according to the hard disk data protocol parsing library. The intermediate status data is, for example, temperature, rotation speed, hard disk bad sectors, etc. The data type is, for example, numerical information and text information.

[0099] In addition, the intelligent management module also designs a set of data conversion programs, whose function is to convert the classified and temporarily stored parsed data into easy-to-understand intuitive data. These converted data are stored in the local hard disk information library for display and risk warning. In the data type identification and matching link, the data conversion program is closely associated with the data temporarily stored by the parsing program. The data conversion program designs a set of dynamically updated physical quantity unit mapping tables, which define the data unit conversion rules. The intermediate status data and the data type are converted into hard disk status parameters according to the physical quantity unit mapping table. For example, for text data processing, if a hard disk health status description identified by a specific character sequence, such as an error code like "ERR001", is encountered, the data conversion program connects to the pre-built error code interpretation database and translates the code into an easy-to-understand text message, such as "ERR001 represents a hard disk head seek fault", so that the operation and maintenance personnel can clearly understand the root cause of the problem.

[0100] In some alternative embodiments, before obtaining the hard disk status parameters, the method further includes:

[0101] When the server is initially powered on and the server system has not been started, obtain the initial hard disk status parameters;

[0102] Determine whether there is an abnormal hard disk according to the initial hard disk status parameters and the second preset threshold of the initial hard disk status parameters, where an abnormal hard disk is a hard disk with at least one initial hard disk status parameter exceeding the second preset threshold;

[0103] When there is an abnormal hard disk, generate a second warning message.

[0104] Specifically, when the server is initially powered on and the server system has not been started, the intelligent management and collection module is the key core. The built-in intelligent programmable logic core chip (hereinafter referred to as IPLCC) starts quickly without waiting for the whole machine system to start, immediately starts the inspection of the hard disk health status, seizes the opportunity for subsequent monitoring, and obtains the initial hard disk status parameters, such as: hard disk temperature, rotation speed, power-on duration, seek error rate, etc.

[0105] The core of the hard disk tray device for intelligently monitoring the hard disk status is the intelligent management and collection module. It will not only collect hard disk information in all directions, but also deeply coordinate the intelligent management of key parameters and the precise triggering of risk warnings to comprehensively protect the hard disk operation. In this embodiment, a second preset threshold is set for the initial hard disk status parameters. For example, the second preset threshold corresponding to the temperature in the initial hard disk status parameters is 60 °C, and the second preset threshold corresponding to the rotation speed in the initial hard disk status parameters is 1000 revolutions per minute. Compare the initial hard disk status parameters with the corresponding second preset threshold to determine whether there is an abnormal hard disk with at least one initial hard disk status parameter exceeding the second preset threshold. If so, generate a second warning message. The second warning message may include: the initial hard disk status parameters exceeding the second preset threshold, information about the abnormal hard disk, etc.

[0106] In addition, this embodiment can be combined with a liquid crystal touch screen to display the basic information of the hard disk when the server is initially powered on. During long-term operation, it can be used by operation and maintenance personnel to check history and diagnose tests, and determine whether the warning is lifted. The two cooperate with each other to ensure stability by using the tray body, and build a comprehensive and intelligent hard disk monitoring system to escort the operation of the data center and the server.

[0107] In this embodiment, when the server is initially powered on and the server system has not been started, there is no need for the server system to quickly collect the initial hard disk status parameters of the hard disk and give warnings, which improves the efficiency and accuracy of initial monitoring and ensures the rapid and stable start of the server system.

[0108] In some alternative embodiments, after generating the first warning message or the second warning message, it is necessary to detect and repair the hard disk. The specific process may include steps A1 to A7.

[0109] Step A1: Obtain the information of the hard disk to be processed and the partition information of the hard disk to be processed when it last booted up normally, where the hard disk to be processed is the hard disk corresponding to the first warning message or the second warning message.

[0110] Step A2: According to the information of the hard disk to be processed and the partition information of the hard disk to be processed in the server when it last booted up normally, perform data backup on the hard disk to be processed.

[0111] Step A3: According to the model of the hard disk to be processed included in the information of the hard disk to be processed, obtain the target hard disk firmware that matches the model of the hard disk to be processed.

[0112] Specifically, the hard disk firmware is installed on a small memory chip of the hard disk and is used to boot the hard disk. In the hard disk, the hard disk firmware is responsible for tasks such as driving, controlling, decoding, transmitting, and detecting, such as managing the storage location of data, recording defective sectors that have been damaged, avoiding using these bad defective sectors again during use, recording the temperature of the hard disk during operation or the errors that occur, etc. The hard disk firmware model is related to the hard disk brand, hard disk capacity, interface type, and form factor, etc. Different models of hard disks have different hard disk firmware. According to the model of the hard disk to be processed, obtain the target hard disk firmware that matches the model of the hard disk to be processed. Updating the hard disk firmware is equivalent to updating the software system that boots the hard disk. The hard disk model must be consistent with the hard disk firmware model. If the models are inconsistent, the hard disk will not be able to store data after the firmware update. And since upgrading the hard disk firmware will cause all the original hard disk data to be lost, it is necessary to back up the hard disk data before updating the hard disk firmware.

[0113] The local storage device of the server generally includes server storage hard disks, etc. The database of the server is a database software installed on the server.

[0114] Step A4: Update the hard disk firmware of the hard disk to be processed according to the target hard disk firmware.

[0115] Specifically, updating the hard disk firmware can repair possible vulnerabilities of the hard disk, improve the stability and reliability of hard disk data, and extend the life of the hard disk, etc.

[0116] Step A5: Restore the data of the hard disk to be processed according to the backup data of the hard disk to be processed.

[0117] Specifically, since updating the hard disk firmware may cause all or part of the original hard disk data to be lost, it is necessary to back up the hard disk data before updating the hard disk firmware. After updating the hard disk firmware, it is also necessary to restore the data of the updated hard disk by restoring the backed-up data to the hard disk.

[0118] Step A6: Send a power-down operation command to the complex programmable logic device, so that the complex programmable logic device sends a control command to the hard disk power control module, thereby causing the hard disk to be processed to perform a power-down operation.

[0119] Step A7: Starting from the moment when the hard disk to be processed performs a power-down operation and after a preset time period, send a power-up operation command to the complex programmable logic device, so that the complex programmable logic device sends a control command for controlling the hard disk to be processed to perform a power-up operation to the hard disk power control module corresponding to the hard disk to be processed, thereby causing the hard disk to be processed to perform a power-up operation.

[0120] In this embodiment, the information of the hard disk to be processed and the partition information of the hard disk to be processed when the server last booted up normally can be used to detect the hard disk to be processed in time and perform recovery, without affecting the use of the server. And when performing fault diagnosis, the basic input / output system of the server is used, without the need for additional diagnostic equipment, saving diagnostic costs, and the diagnostic results are reliable, generally improving the diagnostic efficiency of the hard disk to be processed on the server and the reliability of data recovery. Performing power-on and power-off repair on the hard disk to be processed enables the hard disk to be processed that can be repaired not to be replaced anymore.

[0121] In some alternative embodiments, a method for intelligently monitoring the hard disk status is provided, which can solve the same technical problems as those in steps S101 to S104, such as Figure 2 As shown, the method includes:

[0122] The whole machine system is initially powered on; the intelligent management and control collection module is started; the initial hard disk health status is inspected; it is judged whether there is an abnormal hard disk. If so, the liquid crystal screen is triggered to give an alarm and the operation and maintenance are processed. If not, regular automatic inspection (health assessment algorithm) is performed; it is judged whether the threshold is exceeded. If so, key parameter analysis is triggered. If not, key parameter extraction and evaluation are performed; it is judged whether a warning is triggered. If so, the liquid crystal screen is triggered to give an alarm and the operation and maintenance are processed. If not, risk information is synchronized, and key parameter extraction and evaluation are performed; it is judged whether the threshold is exceeded. If so, the liquid crystal screen is triggered to give an alarm and the operation and maintenance are processed. If not, regular automatic inspection (health assessment algorithm) is performed.

[0123] In this embodiment, when the system is initially powered on, the IPLCC chip of the intelligent control and collection module does not need to wait for the operating system to be deployed. Instead, it quickly collects key information of the hard disk and issues a warning, which changes the drawbacks of traditional manual detection that is time-consuming and error-prone, improves the efficiency and accuracy of initial monitoring, and ensures the fast and stable startup of the system. During the long-term operation stage, this module analyzes the hard disk indicators in real time according to the preset inspection and machine learning algorithms, adaptively adjusts the monitoring focus, copes with read / write pressure and environmental interference, discovers potential problems in advance and issues warnings, enhancing the stability and reliability of the system.

[0124] In this embodiment, a hard disk tray is provided, which can be deployed in the above computer device. The hard disk tray includes: a display component, a connecting line, and an intelligent control and collection component;

[0125] The intelligent control and collection component is used to execute the hard disk status monitoring method of steps S101 to S104 or any corresponding embodiment thereof;

[0126] The intelligent control and collection component is connected to the display component through the connecting line. The intelligent control and collection component is used to transmit the hard disk status parameters and the first warning information to the display component through the connecting line;

[0127] The display component is used to display the hard disk status parameters and the first warning information.

[0128] Specifically, as Figure 3 shown, the display component is, for example, a liquid crystal touch screen; the connecting line is, for example, a connecting line between the liquid crystal touch screen and the module; the intelligent control and collection component is, for example, an intelligent control and collection module; the hard disk tray further includes: a hard disk backplane interface. The hard disk tray focuses on intelligent monitoring of the hard disk status, and the core device is the intelligent control and collection module.

[0129] The display component is the human-computer interaction hub in the hard disk tray, and its functions include: displaying the collected hard disk information: after the intelligent control and collection module completes the first round of hard disk information collection, the information is directly sent to the liquid crystal touch screen for querying basic information. The inspection personnel can select the most concerned parameters through the touch screen, such as whether the power-on duration of hard disks such as SSD (Solid State Drive) meets the basic requirements of operation and maintenance. Interactive operation: The operation and maintenance personnel can perform interactive operations through the liquid crystal touch screen, such as setting hard disk parameters, viewing historical data, etc., making the monitoring and management of the hard disk more convenient and flexible. Providing alarm notification display: After the key parameter warning step of the intelligent control and collection module is completed, the analyzed abnormal or fault information is timely displayed as an alarm through the liquid crystal touch screen. Physical connection and working cooperation diagram between the liquid crystal touch screen and the intelligent control and collection module.

[0130] The intelligent control and collection component will execute the hard disk status monitoring method of steps S101 to S104 or any corresponding implementation manner, and transmit the hard disk status parameters and the first warning information to the display component through a connection line, and the display component will display the hard disk status parameters and the first warning information.

[0131] In this embodiment, the hard disk tray is used to perform daily inspection work such as hard disk inspection and abnormal warning, rather than relying on the server system for inspection and warning. At the same time, a health assessment algorithm and a risk warning mechanism integrating machine learning are installed in the intelligent control and collection component, which solves the dependence of traditional hard disk inspection and the singularity of traditional hard disk warning indicators, promotes the progress of hard disk monitoring technology, and ensures the operation of the data center and the server.

[0132] In this embodiment, a hard disk status monitoring device is also provided. This device is used to implement the above-mentioned embodiment and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0133] This embodiment provides a hard disk status monitoring device, which is deployed in the intelligent control and collection component, as Figure 4 shown, including:

[0134] A parameter acquisition module 401, configured to acquire hard disk status parameters, and determine a first evaluation parameter and a second evaluation parameter corresponding to the hard disk status parameters, where the first evaluation parameter is used to determine the influence degree of the hard disk status parameters on the hard disk working state, and the second evaluation parameter is used to determine the influence degree of the hard disk status parameters on the hard disk failure risk;

[0135] A first parameter processing module 402, configured to perform a first preset logical operation on the hard disk status parameters and the first evaluation parameter to obtain a hard disk status evaluation value;

[0136] A second parameter processing module 403, configured to perform a second preset logical operation on the hard disk status parameters and the second evaluation parameter to obtain a hard disk risk evaluation value;

[0137] A status determination module 404, configured to determine the hard disks in the failure risk state based on the hard disk status evaluation value and the hard disk risk evaluation value, and generate a first warning information.

[0138] In some alternative implementation manners, the status determination module 404 includes:

[0139] A first determination unit, configured to determine the current hard disk status parameters and historical hard disk status parameters in the hard disk status parameters, and determine a parameter range according to the historical hard disk status parameters;

[0140] A second determination unit, configured to determine a current hard disk status evaluation value and a historical hard disk status evaluation value from the hard disk status evaluation values;

[0141] A first judgment unit, configured to judge whether there is at least one historical hard disk status evaluation value equal to the current hard disk status evaluation value when the current hard disk status evaluation value is not within a preset range;

[0142] A third determination unit, configured to determine that the hard disk corresponding to the current hard disk status evaluation value is in a failure risk state and generate a first warning message when there is no historical hard disk status evaluation value equal to the current hard disk status evaluation value;

[0143] A second judgment unit, configured to judge whether the current hard disk status parameter is within the parameter range when there is at least one historical hard disk status evaluation value equal to the current hard disk status evaluation value;

[0144] A fourth determination unit, configured to determine that the hard disk corresponding to the current hard disk status evaluation value is in a failure risk state and generate a first warning message when there is at least one current hard disk status parameter not within the parameter range.

[0145] In some alternative embodiments, the status determination module 404 includes:

[0146] A third judgment unit, configured to judge whether the hard disk risk evaluation value is higher than a preset warning threshold when the current hard disk status evaluation value is within the preset range and there is no current hard disk status parameter not within the parameter range;

[0147] A fifth determination unit, configured to determine that the hard disk corresponding to the hard disk risk evaluation value is in a failure risk state and generate a first warning message when the hard disk risk evaluation value is higher than the preset warning threshold;

[0148] A generation unit, configured to determine that the hard disk corresponding to the current hard disk status parameter is in a failure risk state and generate a first warning message when there is a current hard disk status parameter exceeding a first preset threshold.

[0149] In some alternative embodiments, the first evaluation parameter includes a data validity coefficient, a feature importance factor, a sensitivity adjustment coefficient, and a first weight. The first parameter processing module 402 includes:

[0150] A first setting unit, configured to use the product of the data validity coefficient, the feature importance factor, and the first weight as a first intermediate parameter;

[0151] A sixth determination unit, configured to determine the parameter average value of the hard disk status parameter and determine the difference between the hard disk status parameter and the parameter average value;

[0152] A second setting unit, configured to obtain a second intermediate parameter according to the difference value, the sensitivity adjustment coefficient, and the preset exponent;

[0153] A third setting unit, configured to use the ratio of the first intermediate parameter to the second intermediate parameter as a third intermediate parameter, and obtain a hard disk state evaluation value according to the third intermediate parameter corresponding to the hard disk state parameter.

[0154] In some alternative embodiments, the second evaluation parameter includes a second weight, and the second parameter processing module 403 includes:

[0155] A fourth setting unit, configured to generate a fourth intermediate parameter according to the second weight, the hard disk state parameter, and the preset logarithm, and use the sum of the fourth intermediate parameters as a fifth intermediate parameter;

[0156] A fifth setting unit, configured to use the sum of the second weights as a sixth intermediate parameter;

[0157] A sixth setting unit, configured to use the ratio of the fifth intermediate parameter to the sixth intermediate parameter as a hard disk risk evaluation value.

[0158] In some alternative embodiments, the parameter acquisition module 401 includes:

[0159] A first acquisition unit, configured to acquire status data to be processed;

[0160] A second acquisition unit, configured to acquire intermediate status data and the data type of the intermediate status data from the status data to be processed according to the hard disk data protocol parsing library;

[0161] A conversion unit, configured to convert the intermediate status data and the data type into hard disk state parameters according to the physical quantity unit mapping table.

[0162] In some alternative embodiments, the apparatus further includes:

[0163] An acquisition module, configured to acquire initial hard disk state parameters when the server is initially powered on and the server system has not been started;

[0164] An abnormal hard disk determination module, configured to determine whether there is an abnormal hard disk according to the initial hard disk state parameters and a second preset threshold of the initial hard disk state parameters, where an abnormal hard disk is a hard disk for which at least one initial hard disk state parameter exceeds the second preset threshold;

[0165] A generation module, configured to generate a second warning message when there is an abnormal hard disk.

[0166] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above-mentioned embodiments, and will not be elaborated here.

[0167] The hard disk status monitoring device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0168] The embodiment of the present invention further provides a computer device having the above Figure 4 shown hard disk status monitoring device.

[0169] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 5 shown, the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 5 In

[0170] FIG. is taken as an example of one processor 10.

[0171] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0172] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0173] The memory 20 may include volatile memory, for example, random access memory; the memory may also include non-volatile memory, for example, flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.

[0174] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0175] An embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiment is implemented.

[0176] A part of the present invention can be applied as a computer program product, for example, computer program instructions, which, when executed by a computer, can call or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways for computer program instructions to be executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.

[0177] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the present invention.

Claims

1. A hard disk status monitoring method, characterized in that: The method comprises: Obtaining a hard disk status parameter, and determining a first evaluation parameter and a second evaluation parameter corresponding to the hard disk status parameter, wherein the first evaluation parameter is used to determine the degree of influence of the hard disk status parameter on the hard disk working state, and the second evaluation parameter is used to determine the degree of influence of the hard disk status parameter on the hard disk failure risk; Performing a first preset logical operation on the hard disk status parameter and the first evaluation parameter to obtain a hard disk status evaluation value; Performing a second preset logical operation on the hard disk status parameter and the second evaluation parameter to obtain a hard disk risk evaluation value; A hard disk in a failure risk state is determined based on the hard disk state evaluation value and the hard disk risk evaluation value, and a first warning message is generated.

2. The method according to claim 1, characterized in that The determining of the hard disk in a failure risk state based on the hard disk state evaluation value and the hard disk risk evaluation value includes: Determining current hard disk status parameters and historical hard disk status parameters among the hard disk status parameters, and determining a parameter range according to the historical hard disk status parameters; Determining a current hard disk status evaluation value and a historical hard disk status evaluation value from the hard disk status evaluation values; In the case that the current hard disk status evaluation value is not within the preset range, determining whether there is at least one of the historical hard disk status evaluation values ​​that is equal to the current hard disk status evaluation value; In the case that the historical hard disk status evaluation value does not equal the current hard disk status evaluation value, determining that the hard disk corresponding to the current hard disk status evaluation value is in a failure risk state, and generating the first alarm information; In the case that at least one of the historical hard disk status evaluation values ​​is equal to the current hard disk status evaluation value, determining whether the current hard disk status parameter is within the parameter range; In the case that at least one of the current hard disk status parameters is not within the parameter range, it is determined that the hard disk corresponding to the current hard disk status evaluation value is in a failure risk state, and the first alarm information is generated.

3. The method according to claim 2, characterized in that The method further comprises: When the current hard disk status evaluation value is within the preset range and there is no current hard disk status parameter that is not within the parameter range, determining whether the hard disk risk evaluation value is higher than a preset warning threshold; In the case where the hard disk risk assessment value is higher than the preset warning threshold, determining that the hard disk corresponding to the hard disk risk assessment value is in a failure risk state and generating the first warning information; In the case where the current hard disk status parameter exceeds the first preset threshold, the hard disk corresponding to the current hard disk status parameter is in a failure risk state, and the first alarm information is generated.

4. The method according to claim 1, characterized in that: The first evaluation parameter includes a data validity coefficient, a feature importance factor, a sensitivity adjustment coefficient, and a first weight. The first preset logical operation is performed on the hard disk status parameter and the first evaluation parameter to obtain a hard disk status evaluation value, including: Taking the product of the data validity coefficient, the feature importance factor and the first weight as a first intermediate parameter; Determine a parameter average value of the hard disk status parameter, and determine a difference between the hard disk status parameter and the parameter average value; Obtaining a second intermediate parameter according to the difference, the sensitivity adjustment coefficient and a preset index; The ratio of the first intermediate parameter to the second intermediate parameter is used as a third intermediate parameter, and the hard disk status evaluation value is obtained according to the third intermediate parameter corresponding to the hard disk status parameter.

5. The method according to claim 1, characterized in that: The second evaluation parameter includes a second weight, and performing a second preset logical operation on the hard disk status parameter and the second evaluation parameter to obtain a hard disk risk evaluation value includes: generating a fourth intermediate parameter according to the second weight, the hard disk status parameter and a preset logarithm, and taking the sum of the fourth intermediate parameters as a fifth intermediate parameter; taking the sum of the second weights as a sixth intermediate parameter; The ratio of the fifth intermediate parameter to the sixth intermediate parameter is used as the hard disk risk assessment value.

6. The method according to claim 1, characterized in that The step of obtaining hard disk status parameters includes: Get the status data to be processed; Acquire the intermediate state data and the data type of the intermediate state data from the state data to be processed according to the hard disk data protocol parsing library; The intermediate state data and the data type are converted into the hard disk state parameters according to a physical quantity unit mapping table.

7. The method according to claim 1, characterized in that Before obtaining the hard disk status parameters, the method further includes: When the server is initially powered on and the server system is not started, the initial hard disk status parameters are obtained; Determine whether there is an abnormal hard disk according to the initial hard disk status parameter and a second preset threshold value of the initial hard disk status parameter, wherein the abnormal hard disk is a hard disk having at least one of the initial hard disk status parameters exceeding the second preset threshold value; In the case where the abnormal hard disk exists, a second alarm message is generated.

8. A hard disk bracket, characterized in that: The hard disk bracket includes: a display component, a connecting line and an intelligent control and recording component; The intelligent management and recording component is used to execute the hard disk status monitoring method described in any one of claims 1 to 7; The intelligent control and recording component is connected to the display component via the connecting line, and the intelligent control and recording component is used to transmit the hard disk status parameter and the first alarm information to the display component via the connecting line; The display component is used to display the hard disk status parameter and the first alarm information.

9. A hard disk status monitoring device, characterized in that: The device is deployed in the intelligent management and control collection component, and the device includes: A parameter acquisition module, used to acquire hard disk status parameters, and determine a first evaluation parameter and a second evaluation parameter corresponding to the hard disk status parameters, wherein the first evaluation parameter is used to determine the degree of influence of the hard disk status parameters on the hard disk working state, and the second evaluation parameter is used to determine the degree of influence of the hard disk status parameters on the hard disk failure risk; A first parameter processing module, configured to perform a first preset logic operation on the hard disk status parameter and the first evaluation parameter to obtain a hard disk status evaluation value; A second parameter processing module, used for performing a second preset logic operation on the hard disk status parameter and the second evaluation parameter to obtain a hard disk risk evaluation value; The status determination module is used to determine the hard disk in a failure risk state based on the hard disk status evaluation value and the hard disk risk evaluation value, and generate a first alarm message.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the hard disk status monitoring method according to any one of claims 1 to 7 by executing the computer instructions.

Citation Information

Cited By

  • Data reconstruction method and electronic equipment

    CN120929298A