Component fault prediction method and device, computer program product and storage medium

By integrating the numerical level weighting results of multiple associated sensing features, the problem of poor prediction accuracy of server components is solved, and higher prediction accuracy and server stability are achieved.

CN120448179AActive Publication Date: 2025-08-08SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Patent Information

Application Number
CN202510955780.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-08-08
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

In the prior art, the accuracy of server component failure prediction is poor, resulting in a decrease in server reliability.

Method used

By integrating multiple associated sensing features, determine the numerical level of their predicted values and predicted values, and obtain the failure probability prediction result based on the weighted results of the numerical level, avoiding component failure prediction by single sensing data, and making full use of multi-dimensional data for prediction.

Benefits of technology

Improve the accuracy of component failure prediction and ensure the stable operation of the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448179A_ABST
    Figure CN120448179A_ABST
Patent Text Reader

Abstract

The invention discloses a component fault prediction method and device, a computer program product and a storage medium, belongs to the field of servers, is used for predicting a target component fault based on a plurality of associated sensing features, and solves the problem of poor component fault prediction precision. Considering that part faults of a server may be reflected through a plurality of associated sensing features, the method responds to a prediction instruction for target firmware faults at a future target moment, and firstly determines a prediction value and a numerical level of the prediction value by integrating the plurality of associated sensing features; and then a fault probability prediction result is obtained according to a weighted result of the numerical level, so that part fault prediction through single sensing data is avoided, and multi-dimensional data is fully utilized to predict the fault of the target part, so that the prediction accuracy of the part fault can be improved, and stable operation of a server is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of servers, and in particular to a component failure prediction method, device, computer program product and storage medium. Background Art

[0002] A server has multiple components, and the reliability of the components also affects the reliability of the server. Therefore, it is necessary to predict component failures. However, the relevant technology lacks a mature component failure prediction solution, which makes the component failure prediction of the server less accurate, resulting in reduced server reliability.

[0003] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. Summary of the Invention

[0004] The purpose of the present invention is to provide a component failure prediction method, device, computer program product and storage medium. The present invention responds to a prediction instruction for a target firmware failure at a future target moment. By integrating multiple associated sensor features, the predicted value and the numerical level of the predicted value are first determined, and then the failure probability prediction result is obtained based on the weighted result of the numerical level. This avoids component failure prediction through single sensor data and makes full use of multi-dimensional data to realize the prediction of target component failure, thereby improving the prediction accuracy of component failure and ensuring stable operation of the server.

[0005] To solve the above technical problems, the present invention provides a component failure prediction method, comprising: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data, wherein the target component failure has multiple associated sensor features; For any predicted value, determining the numerical level of the predicted value within the reference value range of the associated sensor feature; The failure probability prediction result of the target component failure is determined based on the weighted results of the numerical levels of each prediction value.

[0006] On the other hand, for any predicted value, determining the numerical level of the predicted value within the reference value range of the associated sensing feature includes: For any predicted value, the value span of the associated sensor feature to which the predicted value belongs at a first preset number of sampling points in the past is used as the reference value range of the associated sensor feature to which the predicted value belongs; Dividing the reference value range corresponding to the predicted value into a second preset number of consecutive numerical intervals, wherein the numerical intervals are sorted in ascending order of numerical value; The sequence number of the numerical interval to which the predicted value is closest is used as the numerical level of the predicted value.

[0007] On the other hand, in response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data includes: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a rate of change of the associated sensor feature based on historical data; According to the change rate of the associated sensing feature and the time difference between the target moment and the current moment, a predicted value of the associated sensing feature at the target moment is determined.

[0008] On the other hand, in response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a rate of change of the associated sensor feature according to historical data includes: In response to a prediction instruction for a target component failure at a future target moment, for any associated sensor feature of the target component failure, a rate of change of the associated sensor feature is determined based on sampling values of the associated sensor feature at a first preset number of sampling points in the past.

[0009] On the other hand, dividing the reference value range corresponding to the predicted value into a second preset number of consecutive value intervals includes: Determining the boundaries of each numerical interval based on the binning type, the second preset number, and the reference value range corresponding to the predicted value; According to the boundaries of each numerical interval, the reference value range corresponding to the predicted value is divided into a second preset number of consecutive numerical intervals.

[0010] On the other hand, according to the binning type, the second preset number, and the reference value range corresponding to the predicted value, determining the boundaries of each numerical interval includes: When the binning type is equal frequency binning, counting the total number of data points within the reference value range; Dividing the total by the second preset number to obtain the number of data points expected to be included in each numerical interval; The data points within the reference value range are sorted according to their numerical values, and starting from the minimum value, the data points are divided in sequence so that the actual number of data points contained in each numerical interval is equal to the expected number of data points contained, thereby determining the boundaries of each numerical interval.

[0011] On the other hand, according to the weighted results of the numerical levels of the various prediction values, the failure probability prediction results of the target component failure are determined to include: Weight the weighted results of each predicted value and use the weighted result as the correlation degree; According to the correlation degree and the predicted value of each associated sensor feature, an index is performed in a preset decision tree model of the target component failure to determine the failure probability prediction result of the target component failure.

[0012] On the other hand, the construction of the decision tree model meets the following conditions: The root node is the correlation feature, and the leaf node is the failure probability; When the correlation degree is greater than a preset threshold, selecting a branch path according to the predicted value of at least one associated sensing feature; The failure probability value of a leaf node is generated through a hierarchical weighted voting mechanism.

[0013] On the other hand, in response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data includes: In response to a prediction instruction for a target component failure at a future target time, determining each associated sensing feature corresponding to the target component failure from a preset association analysis matrix, wherein the preset association analysis matrix includes correspondences between component failures and associated sensing features; For any associated sensing feature of the target component failure, a predicted value of the associated sensing feature at the target time is determined based on historical data.

[0014] On the other hand, the correspondence between component failures and associated sensing features includes: The associated sensing features of memory failure include the memory error checking and correction error rate, memory voltage fluctuation amplitude, and memory chip temperature; The associated sensing features of hard disk failure include the hard disk's cyclic redundancy check error rate, power supply ripple amplitude, fan vibration frequency, and hard disk temperature.

[0015] On the other hand, according to the weighted results of the numerical levels of the various prediction values, the failure probability prediction results of the target component failure are determined to include: Determining a weight matrix of a target component failure from a preset correlation analysis matrix, wherein the preset correlation analysis matrix includes a correspondence between component failures and a weight matrix, and the weight matrix includes weights of various associated sensing features of the target component failure; A weighted result of the numerical levels of the predicted values according to the weight matrix of the target component failure is used as a failure probability prediction result of the target component failure.

[0016] On the other hand, the component failure prediction method further includes: Adjusting a weight matrix of the target component failure based on a third preset number of verification data sets of the target component failure in the past in combination with a gradient descent method; The verification data set includes the failure probability prediction results and their corresponding actual failure occurrence results, where the actual failure occurrence results include whether the failure occurred or did not occur.

[0017] On the other hand, adjusting the weight matrix of the target component failure based on a third preset number of verification data sets of the target component failure in the past in combination with a gradient descent method includes: For each set of data in the third preset number of verification data sets, substituting the fault probability prediction result into a preset loss function to calculate a loss value corresponding to the set of data, wherein the loss function is used to measure the degree of difference between the fault probability prediction result and the actual fault occurrence result; Taking the sum of the loss values corresponding to all verification data groups as the total loss, and calculating the gradient of each weight in the weight matrix of the target component fault based on the total loss; According to a preset learning rate, each weight in the weight matrix of the target component failure is updated along the reverse direction of the gradient to complete the adjustment of the weight matrix of the target component failure.

[0018] On the other hand, it is applied to auxiliary processors; The auxiliary processor is arranged on the board where the baseboard management controller is located, and the auxiliary processor is connected to the advanced extensible interface bus on the board where the baseboard management controller is located; The component failure prediction method further includes: Sensor data is collected through various peripheral controllers on the Advanced Scalable Interface bus.

[0019] On the other hand, sensor data is collected through various peripheral controllers on the advanced extensible interface bus, including: Collecting error checking and correction error rate, operating voltage fluctuation and chip temperature of the memory through the integrated circuit bus controller and / or the enhanced integrated circuit bus controller; Collect the fan vibration frequency through the fan speed control interface controller; The voltage ripple amplitude of the power supply is collected through the power management bus controller; The hard disk interface controller collects the hard disk's cyclic redundancy check error rate and hard disk temperature.

[0020] On the other hand, the component failure prediction method further includes: Pushing the failure probability prediction result of the target component failure to the baseboard management controller, so that the baseboard management controller issues an alarm according to the failure probability prediction result and / or pushes the failure probability prediction result to the remote control terminal.

[0021] To solve the above technical problems, the present invention further provides a component failure prediction device, comprising: Memory for storing computer programs; A processor is configured to implement the steps of the component failure prediction method described above when executing the computer program.

[0022] On the other hand, the processor is an auxiliary processor provided on the board where the baseboard management controller is located; The auxiliary processor collects sensor data through various peripheral controllers on the advanced extensible interface bus.

[0023] To solve the above technical problem, the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the component failure prediction method described above when executed by a processor.

[0024] To solve the above technical problems, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the component failure prediction method described above are implemented.

[0025] Beneficial effects: The present invention provides a component failure prediction method. Considering that the component failure of the server may be reflected by multiple associated sensor features, the present invention responds to the prediction instruction of the target firmware failure at the future target moment. By integrating multiple associated sensor features, the predicted value and the numerical level of the predicted value are first determined, and then the failure probability prediction result is obtained according to the weighted result of the numerical level. This avoids the component failure prediction through a single sensor data and makes full use of multi-dimensional data to realize the prediction of the target component failure, thereby improving the prediction accuracy of the component failure and ensuring the stable operation of the server.

[0026] The present invention also provides a component failure prediction device, a computer program product, and a storage medium, which have the same beneficial effects as the above component failure prediction method. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the relevant technologies and the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 A schematic flow chart of a component failure prediction method provided by the present invention; Figure 2 A schematic structural diagram of a baseboard management controller provided by the present invention; Figure 3 A schematic structural diagram of a component failure prediction system provided by the present invention; Figure 4 A schematic flow chart of another component failure prediction method provided by the present invention; Figure 5 A schematic structural diagram of a component failure prediction device provided by the present invention; Figure 6 A schematic structural diagram of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION

[0029] The core of the present invention is to provide a component failure prediction method, device, computer program product and storage medium. The present invention responds to the prediction instruction of the target firmware failure at the future target time. By integrating multiple associated sensor features, the predicted value and the numerical level of the predicted value are first determined, and then the failure probability prediction result is obtained according to the weighted result of the numerical level. It avoids the component failure prediction through a single sensor data and makes full use of multi-dimensional data to realize the prediction of the target component failure, thereby improving the prediction accuracy of the component failure and ensuring the stable operation of the server.

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0031] Please refer to Figure 1 , Figure 1 This is a flow chart of a component failure prediction method provided by the present invention, which includes: S101: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data, wherein the target component failure has multiple associated sensor features; Specifically, taking into account the technical problems in the above background technology, and considering that the component failure of the server may be reflected through multiple associated sensor features, the embodiment of the present invention intends to predict the failure probability of the target component failure at the future target time through "multiple associated sensor features related to the target component failure", and through the predicted values of each associated sensor feature, the component failure probability prediction can be accurately realized based on multivariate data. Therefore, in this step, first, in response to the prediction instruction for the target component failure at the future target time, for any associated sensor feature of the target component failure, the predicted value of the associated sensor feature at the target time can be determined according to historical data.

[0032] S102: For any predicted value, determining the numerical level of the predicted value within the reference value range of the associated sensor feature to which it belongs; Specifically, in order to simplify calculations and improve prediction efficiency, in an embodiment of the present invention, for any predicted value, the numerical level of the predicted value within the reference value range of the associated sensing feature to which it belongs can be determined, so that the numerical level can be used as the data basis to predict the probability of component failure in subsequent steps.

[0033] S103: Determine a failure probability prediction result of a target component failure according to a weighted result of the numerical levels of each prediction value.

[0034] Specifically, after determining the numerical level of each predicted value within the reference value range of the associated sensor feature to which it belongs, the weighted result of the numerical level of each predicted value can be obtained through a simple weighted calculation, and the failure probability prediction result of the target component failure is determined by performing hierarchical weighting on the predicted values of multiple associated sensor features to construct a cross-component associated fault mapping pattern to adapt to the dynamic changes of the server operating environment.

[0035] Specifically, this component fault prediction method can be applied to auxiliary processors (such as neural processing units (NPUs)). In a heterogeneous hardware architecture with a neural processing unit (NPU) embedded in a BMC (Baseboard Management Controller), the NPU collects operating parameters such as the memory's error checking and correction (ECC) error rate in real time through peripheral controllers such as the Advanced eXtensible Interface (AXI) bus and the Inter-Integrated Circuit (I2C) bus controller. For memory failure prediction at a future target moment, the rate of change of memory chip temperature is calculated based on historical data, and the temperature prediction value at the target moment is obtained by combining the time difference between the target moment and the current moment. The numerical span of the predicted value at the first preset number of sampling points in the past is used as the reference value range, and the predicted value is divided into a second preset number of continuous numerical intervals. The serial number of the interval to which the predicted value belongs is determined as the numerical level. Based on the weight matrix corresponding to the memory failure in the preset correlation analysis matrix, the numerical level of each associated sensor feature (such as ECC error rate, memory voltage fluctuation amplitude, and memory chip temperature) is weighted to obtain the failure probability prediction result.

[0036] Specifically, to better illustrate the embodiments of the present invention, please refer to Figure 2 , Figure 2 A schematic diagram of the structure of a baseboard management controller provided by the present invention is shown. Figure 2 The server processor can collect sensor data through the first to fourth peripheral controllers on the high-order scalable interface bus and execute the component failure prediction method without the help of the central processing unit of the baseboard management controller, thereby reducing the burden on the central processing unit.

[0037] The present invention provides a component failure prediction method. Considering that a server component failure may be reflected by multiple associated sensor features, the present invention responds to a prediction instruction for a target firmware failure at a future target moment. By integrating multiple associated sensor features, the method first determines its predicted value and the numerical level of the predicted value, and then obtains a failure probability prediction result based on the weighted result of the numerical level. This avoids component failure prediction based on a single sensor data and makes full use of multi-dimensional data to realize the prediction of target component failure, thereby improving the prediction accuracy of component failure and ensuring the stable operation of the server.

[0038] Based on the above embodiment: As an optional embodiment, for any predicted value, determining the numerical level of the predicted value within the reference value range of the associated sensing feature includes: For any predicted value, the value span of the associated sensor feature to which the predicted value belongs at a first preset number of sampling points in the past is used as the reference value range of the associated sensor feature to which the predicted value belongs; Dividing the reference value range corresponding to the predicted value into a second preset number of consecutive numerical intervals, wherein the numerical intervals are sorted in ascending order of numerical value; The serial number of the numerical interval closest to the predicted value is used as the numerical level of the predicted value.

[0039] Specifically, in order to convert the predicted value into a quantifiable feature level and improve the accuracy of fault prediction, the embodiment of the present invention can dynamically determine the reference value range and divide the numerical interval so that the numerical level can more accurately reflect the position of the predicted value in the historical data distribution.

[0040] Specifically, taking hard drive failure prediction as an example, for the predicted value of the hard drive's cyclic redundancy check (CRC) error rate, the span of the first preset number (e.g., 5) of sampling points in the past is used as the reference value range. Assuming this range is [10, 50] and the second preset number is 2, using equal-frequency binning, the total number of data points within the range is counted and evenly divided into two intervals, with the interval boundary set at 30. If the predicted value is 35, the closest interval is [30, 50], with a sequence number of 2, indicating a numerical level of 2.

[0041] Of course, in addition to this specific form, "for any predicted value, determining the numerical level of the predicted value within the reference value range of the associated sensor feature to which it belongs" can also be in other forms, and the embodiments of the present invention are not limited here.

[0042] As an optional embodiment, in response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data includes: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a rate of change of the associated sensor feature based on historical data; The predicted value of the associated sensor feature at the target time is determined based on the change rate of the associated sensor feature and the time difference between the target time and the current time.

[0043] Specifically, in order to predict future parameter values based on the changing trends of historical data and solve the problem that traditional static thresholds cannot adapt to changes in hardware status, the embodiments of the present invention can make the predicted values more in line with the dynamic trends of component operation by calculating the change rate and time difference.

[0044] For example, when predicting the target time value of a memory chip's temperature, the rate of change is calculated based on the temperature sampled at the first preset number of sampling points (e.g., five) in the past. For example, if the temperature rose from 72°C to 90°C over the past 20 seconds, the rate of change is (90 - 72) / 20 = 0.9°C / second. If the target time is 30 seconds from the current time, the predicted value is 90 + 0.9 × 30 = 117°C (this can be adjusted in conjunction with hardware safety thresholds in actual applications).

[0045] Of course, in addition to this specific form, "in response to a prediction instruction for a target component failure at a future target moment, for any associated sensor feature of the target component failure, determining the predicted value of the associated sensor feature at the target moment based on historical data" can also be in other forms, and the embodiments of the present invention are not limited here.

[0046] As an optional embodiment, in response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a rate of change of the associated sensor feature based on historical data includes: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, a change rate of the associated sensor feature is determined based on sampling values of the associated sensor feature at a first preset number of sampling points in the past.

[0047] Specifically, to ensure the accuracy and reliability of the change rate calculation, in the embodiment of the present invention, statistical analysis can be performed based on sufficient historical sampling data so that the change rate can truly reflect the dynamic change trend of the associated sensing feature.

[0048] Specifically, to calculate the rate of change of the power supply voltage ripple amplitude, obtain the values of the first preset number of sampling points (e.g., 10) in the past, fit the data trend using methods such as linear regression, and calculate the rate of change per unit time. For example, if the ripple amplitude at 10 sampling points increases from 20mV to 40mV within 50 seconds, the rate of change is (40-20) / 50 = 0.4mV / second.

[0049] Of course, in addition to this specific form, "in response to a prediction instruction for a target component failure at a future target moment, for any associated sensor feature of the target component failure, determining the rate of change of the associated sensor feature based on historical data" can also be in other forms, and the embodiments of the present invention are not limited here.

[0050] As an optional embodiment, dividing the reference value range corresponding to the predicted value into a second preset number of consecutive value intervals includes: Determine the boundaries of each numerical interval based on the binning type, the second preset number, and the reference value range corresponding to the predicted value; According to the boundaries of each numerical interval, the reference value range corresponding to the predicted value is divided into a second preset number of consecutive numerical intervals.

[0051] Specifically, to adapt to different data distribution characteristics, the embodiments of the present invention can convert continuous sensor data into discrete interval features through flexible configuration of bin type, quantity and range, which facilitates subsequent association calculation and model reasoning.

[0052] Specifically, for the predicted fan vibration frequency, the reference range is set to [20, 80] Hz, the binning type is equal frequency binning, and the second preset number is 3. The total number of data points within this range is 60, and each interval is expected to contain 20 data points. The data is sorted by value and divided starting from the minimum value to obtain the intervals [20, 35], [35, 55], and [55, 80], with the boundaries of each interval being 35 and 55, respectively.

[0053] Of course, in addition to this specific form, "dividing the reference value range corresponding to the predicted value into a second preset number of continuous numerical intervals" can also be in other forms, and the embodiment of the present invention is not limited here.

[0054] As an optional embodiment, determining the boundaries of each numerical interval according to the binning type, the second preset number, and the reference value range corresponding to the predicted value includes: When the binning type is equal frequency binning, the total number of data points within the reference value range is counted; Divide the total by the second preset number to obtain the number of data points expected to be included in each numerical interval; The data points within the reference value range are sorted according to their numerical values, and starting from the minimum value, the data points are divided in sequence so that the actual number of data points contained in each numerical interval is equal to the expected number of data points contained, thereby determining the boundaries of each numerical interval.

[0055] Specifically, in order to make the binning results more evenly reflect the data distribution density, the embodiment of the present invention uses an equal-frequency binning method to ensure that each interval contains a similar number of data points, thereby avoiding feature quantization deviation caused by uneven data distribution.

[0056] For example, the memory voltage fluctuation amplitude has 100 data points within its reference range, the second preset number is 4, and each interval is expected to contain 25 data points. The data is sorted from smallest to largest, with the maximum value of the first 25 data points being 15mV, which serves as the upper bound of the first interval. The maximum value of the next 25 data points is 25mV, which serves as the upper bound of the second interval. Similarly, the interval boundaries are determined to be 15, 25, and 35mV, forming the interval divisions of [0,15], [15,25], [25,35], and [35,+∞].

[0057] Of course, in addition to this specific form, "determining the boundaries of each numerical interval based on the binning type, the second preset number, and the reference value range corresponding to the predicted value" can also be in other forms, and the embodiments of the present invention are not limited here.

[0058] As an optional embodiment, determining the failure probability prediction result of the target component failure according to the weighted result of the numerical levels of each prediction value includes: Weight the weighted results of each predicted value and use the weighted result as the correlation degree; According to the correlation degree and the predicted value of each associated sensor feature, an index is performed in a preset decision tree model of the target component failure to determine the failure probability prediction result of the target component failure.

[0059] Specifically, in order to combine the correlation degree and the decision tree model to perform fault probability reasoning, the embodiment of the present invention can use the correlation degree to reflect the comprehensive influence of multiple features, and improve the accuracy and interpretability of fault prediction through the hierarchical judgment of the decision tree model.

[0060] For memory failure prediction, the numerical levels of the associated sensor features (ECC error rate, memory voltage fluctuation amplitude, and memory chip temperature) are calculated as 2, 2, and 2, respectively. The weight matrix is [0.4, 0.3, 0.3], and the correlation is 2 × 0.4 + 2 × 0.3 + 2 × 0.3 = 2.0. In the preset decision tree model, the root node is the correlation. When the correlation is greater than 1.5, the temperature is determined to be greater than 80°C. If the predicted temperature is 97°C, the inferred failure probability is 0.9.

[0061] Of course, in addition to this specific form, "determining the failure probability prediction result of the target component failure based on the weighted result of the numerical level of each prediction value" can also be in other forms, which are not limited in the embodiment of the present invention.

[0062] As an optional embodiment, the construction of the decision tree model meets the following conditions: The root node is the correlation feature, and the leaf node is the failure probability; When the correlation degree is greater than a preset threshold, selecting a branch path according to the predicted value of at least one associated sensing feature; The failure probability value of a leaf node is generated through a hierarchical weighted voting mechanism.

[0063] Specifically, in order to build an efficient fault inference model, the embodiment of the present invention uses the correlation as the root node and generates the fault probability through a hierarchical weighted voting mechanism, so that the model can quickly traverse the decision path and adapt to the needs of online real-time inference.

[0064] Specifically, in the decision tree model for memory failures, the root node is the correlation feature. When the correlation is ≤1.5, the leaf node failure probability is 0.1. When the correlation is greater than 1.5, the memory chip temperature is further determined. If it is ≤80°C, the failure probability is 0.4, and if it is greater than 80°C, the failure probability is 0.9. This model uses the LightGBM (Light Gradient Boosting Machine) inference core to achieve hardware-accelerated traversal.

[0065] Of course, in addition to this specific form, the "construction conditions of the decision tree model" can also be in other forms, which are not limited in the embodiment of the present invention.

[0066] As an optional embodiment, in response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data includes: In response to a prediction instruction for a target component failure at a future target time, determining each associated sensing feature corresponding to the target component failure from a preset association analysis matrix, wherein the preset association analysis matrix includes correspondences between component failures and associated sensing features; For any associated sensor feature of the target component failure, a predicted value of the associated sensor feature at the target time is determined based on historical data.

[0067] Specifically, in order to clarify the associated sensing characteristics of the target component, the embodiment of the present invention can establish the correspondence between component failures and characteristics through a preset association analysis matrix, realize the association modeling analysis of multiple component states, and solve the limitations of single component detection.

[0068] For example, the associated sensing features corresponding to memory failures, including memory ECC error rate, memory voltage fluctuation amplitude, and memory chip temperature, are obtained from a preset correlation analysis matrix. Memory ECC error rate data is collected via the I2C bus controller, with each sampling period of 5 seconds. Data is collected at five consecutive time points for subsequent prediction calculations.

[0069] As an optional embodiment, the corresponding relationship between component failure and associated sensing characteristics includes: The associated sensing features of memory failure include the memory error checking and correction error rate, memory voltage fluctuation amplitude, and memory chip temperature; The associated sensing features of hard disk failure include the hard disk's cyclic redundancy check error rate, power supply ripple amplitude, fan vibration frequency, and hard disk temperature.

[0070] Specifically, in order to accurately predict the fault characteristics of different components, the embodiments of the present invention define specific associated sensing features based on the physical characteristics of the hardware, making the fault prediction more targeted and accurate.

[0071] Specifically, the associated sensor features of hard drive failures include the hard drive CRC error rate, power ripple amplitude, fan vibration frequency, and hard drive temperature. The hard drive CRC error rate is collected by the hard drive interface controller, the power ripple amplitude is collected by the Power Management Bus (PMBus) controller, and the fan vibration frequency is collected by the Fantacho controller.

[0072] Of course, in addition to this specific form, the "correspondence between component failure and associated sensor characteristics" can also be in other forms, which are not limited in the embodiments of the present invention.

[0073] As an optional embodiment, determining the failure probability prediction result of the target component failure according to the weighted result of the numerical levels of each prediction value includes: Determining a weight matrix of a target component failure from a preset correlation analysis matrix, wherein the preset correlation analysis matrix includes a correspondence between component failures and a weight matrix, and the weight matrix includes weights of various associated sensing features of the target component failure; The weighted result of the numerical level of each prediction value according to the weight matrix of the target component failure is used as the failure probability prediction result of the target component failure.

[0074] Specifically, to achieve weighted fusion of multiple features, the embodiment of the present invention can weight the numerical level of each feature through a weight matrix in a preset association analysis matrix to reflect the different degrees of influence of different features on the fault and improve prediction accuracy.

[0075] For example, the weight matrix [0.3, 0.2, 0.2, 0.3] corresponding to hard drive failure is obtained from the preset correlation analysis matrix. These weights correspond to the hard drive CRC error rate, power supply ripple amplitude, fan vibration frequency, and hard drive temperature. The numerical levels of these features are 3, 2, 2, and 3, respectively. The weighted result is 3 × 0.3 + 2 × 0.2 + 2 × 0.2 + 3 × 0.3 = 2.6, which is used as the probability prediction result for hard drive failure.

[0076] As an optional embodiment, the component failure prediction method further includes: Adjusting the weight matrix of the target component failure based on a third preset number of verification data sets of the target component failure in the past in combination with a gradient descent method; The verification data set includes the failure probability prediction results and their corresponding actual failure occurrence results, where the actual failure occurrence results include whether the failure occurred or did not occur.

[0077] Specifically, in order to make the weight matrix adapt to the real-time operating environment of the server, the embodiment of the present invention uses historical verification data and gradient descent method to perform dynamic adjustments to improve the model's adaptability to hardware status changes and prediction accuracy.

[0078] For example, for a memory failure weight matrix of [0.4, 0.3, 0.3], a preset number of validation data sets (e.g., 100) are collected over the past three decades. Each set contains a predicted probability of failure and the actual occurrence of a failure. The predicted results are substituted into the cross-entropy loss function. After calculating the total loss, the weight matrix is updated using gradient descent with a learning rate of 0.05. If a data set with a predicted probability of 0.9 actually fails, the adjusted weights might become [0.42, 0.32, 0.26].

[0079] As an optional embodiment, adjusting the weight matrix of the target component failure based on a third preset number of verification data sets of the target component failure in the past in combination with a gradient descent method includes: For each set of data in the third preset number of verification data sets, substituting the failure probability prediction result into a preset loss function to calculate the loss value corresponding to the set of data, wherein the loss function is used to measure the degree of difference between the failure probability prediction result and the actual failure occurrence result; The sum of the loss values corresponding to all validation data sets is taken as the total loss, and the gradient of each weight in the weight matrix of the target component failure is calculated based on the total loss; According to the preset learning rate, each weight in the weight matrix of the target component failure is updated along the reverse direction of the gradient to complete the adjustment of the weight matrix of the target component failure.

[0080] Specifically, to ensure the scientificity and effectiveness of the weight matrix adjustment, the embodiment of the present invention measures the difference between prediction and reality through a loss function, updates the weights based on gradient calculation and learning rate, and realizes the optimization iteration of the weight matrix.

[0081] For example, for a third preset number (e.g., 50 groups) of validation data, the loss value is calculated for each group (e.g., using the mean square error loss function), and the total loss is the sum of all group losses. Calculate the gradient of each weight in the weight matrix, for example, the gradient of weight w1 is , update the weights in the opposite direction of the gradient: , where η is the learning rate 0.01, completing the iterative adjustment of the weight matrix.

[0082] As an optional embodiment, it is applied to an auxiliary processor; The auxiliary processor is arranged on the board where the baseboard management controller is located, and the auxiliary processor is connected to the advanced extensible interface bus on the board where the baseboard management controller is located; Component failure prediction methods also include: Sensor data is collected through various peripheral controllers on the Advanced Scalable Interface bus.

[0083] Specifically, to achieve efficient data collection and fault prediction, the component fault prediction method can be applied to the auxiliary processor in the embodiment of the present invention, and the peripheral controller is connected via the AXI bus to reduce the BMC CPU resource usage and improve the system operation efficiency.

[0084] Specifically, the auxiliary processor can be installed on the board where the BMC is located and connected to peripheral controllers such as the I2C bus controller and PMBus controller via the AXI bus. The I2C bus controller collects memory ECC error rates, operating voltage fluctuations, and chip temperature, while the FANTACHO controller collects fan vibration frequency, enabling real-time collection of sensor data.

[0085] As an optional embodiment, collecting sensor data through each peripheral controller on the advanced extensible interface bus includes: Collecting error checking and correction error rate, operating voltage fluctuation and chip temperature of the memory through the integrated circuit bus controller and / or the enhanced integrated circuit bus controller; Collect the fan vibration frequency through the fan speed control interface controller; The voltage ripple amplitude of the power supply is collected through the power management bus controller; The hard disk interface controller collects the hard disk's cyclic redundancy check error rate and hard disk temperature.

[0086] Specifically, in order to target the sensor data characteristics of different components, the embodiments of the present invention use corresponding peripheral controllers to perform precise data collection, ensure the accuracy and real-time nature of the data, and provide reliable data support for fault prediction.

[0087] The memory ECC error rate (errors / million), operating voltage fluctuation amplitude (mV), and particle temperature (°C) are collected through the I2C bus controller and the enhanced I2C (I3C, Inter-Integrated Circuit 3) bus controller. The fan vibration frequency (Hz) is collected through the FANTACHO controller with a sampling period of 10 seconds. The power supply voltage ripple amplitude (mV) is collected through the PMBus controller. The hard disk CRC error rate (errors / million) and hard disk temperature (°C) are collected through the hard disk interface controller.

[0088] As an optional embodiment, the component failure prediction method further includes: Push the failure probability prediction result of the target component failure to the baseboard management controller, so that the baseboard management controller issues an alarm according to the failure probability prediction result and / or pushes the failure probability prediction result to the remote control terminal.

[0089] Specifically, to achieve timely feedback and remote management of fault prediction results, the embodiments of the present invention push the fault probability prediction results to the BMC for alarm and remote reporting, thereby improving the timeliness and convenience of server hardware reliability management.

[0090] For example, if the NPU calculates and generates a memory failure probability prediction of 0.9, it pushes it to the BMC firmware. The monitoring program in the BMC firmware issues a fault warning based on a preset threshold (such as 0.7) and reports it to the remote control terminal via interfaces such as the web, Redfish, Intelligent Platform Management Interface (IPMI), and Simple Network Management Protocol (SNMP).

[0091] Specifically, to better illustrate the embodiments of the present invention, please refer to Figure 3 and Figure 4 , Figure 3 This is a structural diagram of a component failure prediction system provided by the present invention. Figure 4 A schematic flow chart of another component failure prediction method provided by the present invention, Figure 3 In the system, the remote control terminal is connected to the baseboard management controller through the network interface. Inside the baseboard management controller, the central processing unit, auxiliary processor, and peripheral controller are connected in sequence. The peripheral controller is respectively connected to the memory, hard disk, chassis, fan, and power supply to form a complete system architecture. Figure 4In the process, component operation data is first acquired through data collection, then data preprocessing is used to regularize the data, and then feature generation is used to divide the numerical interval and determine the numerical level. Then, correlation analysis is used to determine the correlation degree, and then model inference outputs the result. The model inference result can be used to update the weight matrix and reversely assist in correlation analysis.

[0092] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of a component failure prediction device provided by the present invention, the component failure prediction device comprising: Memory 51, for storing computer programs; The processor 52 is configured to implement the steps of the component failure prediction method in the aforementioned embodiment when executing the computer program.

[0093] As an optional embodiment, the processor is an auxiliary processor provided on a board where the baseboard management controller is located; The auxiliary processor collects sensor data through various peripheral controllers on the Advanced Scalable Interface bus.

[0094] For an introduction to the component failure prediction device provided by an embodiment of the present invention, please refer to the aforementioned embodiment of the component failure prediction method, and the embodiment of the present invention will not be described in detail here.

[0095] The present invention also provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the component failure prediction method in the aforementioned embodiment.

[0096] For an introduction to the computer program product provided by the embodiment of the present invention, please refer to the aforementioned embodiment of the component failure prediction method, and the embodiment of the present invention will not be described in detail here.

[0097] Please refer to Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer-readable storage medium provided by the present invention. A computer program 62 is stored on the computer-readable storage medium 61. When the computer program 62 is executed by a processor, the steps of the component failure prediction method in the aforementioned embodiment are implemented.

[0098] For an introduction to the computer-readable storage medium provided in an embodiment of the present invention, please refer to the aforementioned embodiment of the component failure prediction method, and the embodiment of the present invention will not be described in detail here.

[0099] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referred to each other. It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method, article or equipment including the element.

[0100] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A component failure prediction method, characterized in that: include: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data, wherein the target component failure has multiple associated sensor features; For any predicted value, determining the numerical level of the predicted value within the reference value range of the associated sensor feature; The failure probability prediction result of the target component failure is determined based on the weighted results of the numerical levels of each prediction value.

2. The component failure prediction method according to claim 1, characterized in that: For any predicted value, determining the numerical level of the predicted value within the reference value range of the associated sensing feature includes: For any predicted value, the value span of the associated sensor feature to which the predicted value belongs at a first preset number of sampling points in the past is used as the reference value range of the associated sensor feature to which the predicted value belongs; Dividing the reference value range corresponding to the predicted value into a second preset number of consecutive numerical intervals, wherein the numerical intervals are sorted in ascending order of numerical value; The sequence number of the numerical interval to which the predicted value is closest is used as the numerical level of the predicted value.

3. The component failure prediction method according to claim 2, characterized in that: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data includes: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a rate of change of the associated sensor feature based on historical data; According to the change rate of the associated sensing feature and the time difference between the target moment and the current moment, a predicted value of the associated sensing feature at the target moment is determined.

4. The component failure prediction method according to claim 3, characterized in that: In response to a prediction instruction for a target component failure at a future target time, determining, for any associated sensor feature of the target component failure based on historical data, a rate of change of the associated sensor feature includes: In response to a prediction instruction for a target component failure at a future target moment, for any associated sensor feature of the target component failure, a rate of change of the associated sensor feature is determined based on sampling values of the associated sensor feature at a first preset number of sampling points in the past.

5. The component failure prediction method according to claim 2, characterized in that: Dividing the reference value range corresponding to the predicted value into a second preset number of consecutive value intervals includes: Determining the boundaries of each numerical interval based on the binning type, the second preset number, and the reference value range corresponding to the predicted value; According to the boundaries of each numerical interval, the reference value range corresponding to the predicted value is divided into a second preset number of consecutive numerical intervals.

6. The component failure prediction method according to claim 5, characterized in that: Determining the boundaries of each numerical interval based on the binning type, the second preset number, and the reference value range corresponding to the predicted value includes: When the binning type is equal frequency binning, counting the total number of data points within the reference value range; Dividing the total by the second preset number to obtain the number of data points expected to be included in each numerical interval; The data points within the reference value range are sorted according to their numerical values, and starting from the minimum value, the data points are divided in sequence so that the actual number of data points contained in each numerical interval is equal to the expected number of data points contained, thereby determining the boundaries of each numerical interval.

7. The component failure prediction method according to claim 1, characterized in that: According to the weighted results of the numerical levels of the predicted values, the failure probability prediction results of the target component failure are determined to include: Weight the weighted results of each predicted value and use the weighted result as the correlation degree; According to the correlation degree and the predicted value of each associated sensor feature, an index is performed in a preset decision tree model of the target component failure to determine the failure probability prediction result of the target component failure.

8. The component failure prediction method according to claim 7, characterized in that: The construction of the decision tree model meets the following conditions: The root node is the correlation feature, and the leaf node is the failure probability; When the correlation degree is greater than a preset threshold, selecting a branch path according to a predicted value of at least one associated sensing feature; The failure probability value of a leaf node is generated through a hierarchical weighted voting mechanism.

9. The component failure prediction method according to claim 1, characterized in that: In response to a prediction instruction for a target component failure at a future target time, for any associated sensor feature of the target component failure, determining a predicted value of the associated sensor feature at the target time based on historical data includes: In response to a prediction instruction for a target component failure at a future target time, determining each associated sensing feature corresponding to the target component failure from a preset association analysis matrix, wherein the preset association analysis matrix includes correspondences between component failures and associated sensing features; For any associated sensing feature of the target component failure, a predicted value of the associated sensing feature at the target time is determined based on historical data.

10. The component failure prediction method according to claim 9, characterized in that: The correspondence between component failures and associated sensing features includes: The associated sensing features of memory failure include the memory error checking and correction error rate, memory voltage fluctuation amplitude, and memory chip temperature; The associated sensing features of hard disk failure include the hard disk's cyclic redundancy check error rate, power supply ripple amplitude, fan vibration frequency, and hard disk temperature.

11. The component failure prediction method according to claim 1, wherein: According to the weighted results of the numerical levels of the predicted values, the failure probability prediction results of the target component failure are determined to include: Determining a weight matrix of a target component failure from a preset correlation analysis matrix, wherein the preset correlation analysis matrix includes a correspondence between component failures and a weight matrix, and the weight matrix includes weights of various associated sensing features of the target component failure; A weighted result of the numerical levels of the predicted values according to the weight matrix of the target component failure is used as a failure probability prediction result of the target component failure.

12. The component failure prediction method according to claim 11, characterized in that: The component failure prediction method further includes: Adjusting a weight matrix of the target component failure based on a third preset number of verification data sets of the target component failure in the past in combination with a gradient descent method; The verification data set includes the failure probability prediction results and their corresponding actual failure occurrence results, where the actual failure occurrence results include whether the failure occurred or did not occur.

13. The component failure prediction method according to claim 12, characterized in that: Adjusting the weight matrix of the target component failure based on a third preset number of verification data sets of the target component failure in the past in combination with a gradient descent method includes: For each set of data in the third preset number of verification data sets, substituting the fault probability prediction result into a preset loss function to calculate a loss value corresponding to the set of data, wherein the loss function is used to measure the degree of difference between the fault probability prediction result and the actual fault occurrence result; Taking the sum of the loss values corresponding to all verification data groups as the total loss, and calculating the gradient of each weight in the weight matrix of the target component failure based on the total loss; According to a preset learning rate, each weight in the weight matrix of the target component failure is updated along the reverse direction of the gradient to complete the adjustment of the weight matrix of the target component failure.

14. The component failure prediction method according to any one of claims 1 to 13, characterized in that: Applied to auxiliary processors; The auxiliary processor is arranged on the board where the baseboard management controller is located, and the auxiliary processor is connected to the advanced extensible interface bus on the board where the baseboard management controller is located; The component failure prediction method further includes: Sensor data is collected through various peripheral controllers on the Advanced Scalable Interface bus.

15. The component failure prediction method according to claim 14, characterized in that: Sensor data is collected through various peripheral controllers on the Advanced Scalable Interface bus, including: Collecting error checking and correction error rate, operating voltage fluctuation and chip temperature of the memory through the integrated circuit bus controller and / or the enhanced integrated circuit bus controller; Collect the fan vibration frequency through the fan speed control interface controller; The voltage ripple amplitude of the power supply is collected through the power management bus controller; The hard disk interface controller collects the hard disk cyclic redundancy check error rate and hard disk temperature.

16. The component failure prediction method according to claim 14, characterized in that: The component failure prediction method further includes: Pushing the failure probability prediction result of the target component failure to the baseboard management controller, so that the baseboard management controller issues an alarm according to the failure probability prediction result and / or pushes the failure probability prediction result to the remote control terminal.

17. A component failure prediction device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the component failure prediction method according to any one of claims 1 to 16 when executing the computer program.

18. The component failure prediction device according to claim 17, characterized in that The processor is an auxiliary processor provided on the board where the baseboard management controller is located; The auxiliary processor collects sensor data through various peripheral controllers on the advanced extensible interface bus.

19. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the component failure prediction method according to any one of claims 1 to 16 are implemented.

20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the component failure prediction method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Server fault prediction method, system, device and medium

    CN115168168A

  • Dynamic threshold generation method and device, equipment and storage medium

    CN115563180A

  • Method, system and equipment for predicting memory fault and storage medium

    CN116483675A

  • Fault prediction method, device and equipment and readable storage medium

    CN116684306A

  • Monitoring method and system for ship shafting supporting structure

    CN117990043A

Cited By

  • Fault information processing method and system of intelligent capsule fresh extraction beverage machine

    CN120632664A

  • A fault information processing method and system of an intelligent capsule fresh extract beverage machine

    CN120632664B