Fault prediction method and device, computing equipment and storage medium
By filtering and sorting the base model pool, the appropriate base model is selected for fault prediction for different data distribution types, which solves the problems of degradation of prediction accuracy and data loss caused by changes in data distribution in the prior art, and achieves more efficient and accurate fault prediction.
Patent Information
- Application Number
- CN202510435731.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing node failure prediction technology is affected when the data distribution changes, which is prone to failure and lead to data loss, affecting the user experience.
The target base model pool is filtered based on the numerical change law type of performance parameters, and dynamically sort the base model with the highest prediction accuracy for numerical prediction. The appropriate base model is used for processing for different data distributions to avoid the impact of data distribution changes on the fault prediction process.
Improves the accuracy and stability of fault prediction, avoids data loss, and improves user experience.
Smart Images

Figure CN120386658A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and particularly to a fault prediction method, apparatus, computing device, and storage medium. Background Art
[0002] With the development of the supercomputing Internet, in order to improve computing power, each enterprise and organization connects to the supercomputing Internet in its respective business, and massively connects the supercomputing Internet to clusters and nodes. During the use of the supercomputing Internet, it is necessary to quickly identify faulty nodes and repair the faulty nodes to prevent data loss and improve the user experience.
[0003] The existing node fault prediction technology mainly relies on the monitoring and analysis of the performance parameters of physical machines, and compares the values of the obtained performance parameters with static thresholds. Once the data exceeds the threshold, it can be considered that the node has failed. However, when the data distribution changes, the prediction accuracy will be affected by using this method, and it is easy to have failures and data loss, which affects the user experience. Summary of the Invention
[0004] To solve the above problems in the prior art, embodiments of the present application provide a fault prediction method, apparatus, computing device, and storage medium, which are used to predict whether a node will fail in different ways under different data distributions, and can ensure the stability during the fault prediction process, avoid data loss, and ensure the user experience.
[0005] In a first aspect, an embodiment of the present application provides a fault prediction method, including:
[0006] Based on the generation time of multiple values of the same performance parameter and the magnitude relationship of the multiple values, determine the type of change rule of the multiple values, and select the target base model pool corresponding to the determined type of change rule from multiple base model pools;
[0007] Input the first part of the multiple values into each target base model in the target base model pool, and based on the second part of the multiple values and the values output by each target base model, determine the prediction accuracy of each target base model; wherein the generation time of the first part of the values is earlier than the generation time of the second part of the values;
[0008] Based on the prediction accuracy of each target base model, select a set number of target base models;
[0009] Input the second part of the values into the set number of target base models to obtain at least one numerical prediction result, and determine the fault prediction result based on the at least one numerical prediction result.
[0010] By this method, the influence of data distribution changes on the fault prediction process can be avoided. For different data distributions, that is, different types of data change rules, different basic model pools can be used for processing. The basic models in different basic model pools have different processing capabilities for different data change rules. Selecting a basic model with stronger processing capabilities to process data of the corresponding data change rule type can ensure the accuracy of prediction, avoid the occurrence of faults and data loss, and improve the user experience.
[0011] In a possible embodiment, before determining the type of change rule of the multiple values based on the generation time of the multiple values of the same performance parameter obtained and the magnitude relationship of the multiple values, the method further includes:
[0012] Periodically obtain the values of the performance parameter;
[0013] For each obtained value of the performance parameter, determine the relationship between the value of the performance parameter and the value range; the value range is determined based on the obtained values of the performance parameter.
[0014] If the value of the performance parameter is not within the value range, then use the target value determined based on the value range as the value of the performance parameter.
[0015] During the process of obtaining the values of the performance parameter, the obtained performance parameter can be preprocessed to avoid the influence on the subsequent fault prediction process due to the failure to obtain some performance parameters or the error in obtaining some performance parameters, which can improve the accuracy of the fault prediction process to a certain extent.
[0016] In a possible embodiment, the value range is determined based on the obtained values of the performance parameter in the following manner:
[0017] For every N obtained values of the performance parameter, determine the value range based on some or all of the N obtained values of the performance parameter, a preset upward floating range, and a preset downward floating range; where N is a positive integer.
[0018] By determining the corresponding value range in the above manner before obtaining the values of the performance parameter and using different numbers of performance parameters to determine the value range, the value range can be determined more accurately.
[0019] In a possible embodiment, each basic model in the basic model pool is obtained by the following method:
[0020] Classify the multiple sets of training data according to each type of change rule, so that each type of change rule corresponds to at least one set of training data; each set of training data includes data to be processed and verification data;
[0021] For any base model, input at least one set of training data corresponding to each type of change rule into the any base model, and based on the output training results and the verification data, determine the fitting results corresponding to each type of change rule;
[0022] Based on the multiple fitting results, the change rule types corresponding to the any base model;
[0023] Save at least one base model corresponding to each type of change rule into the base model pool corresponding to the change rule type.
[0024] Classify multiple base models in the above manner, determine the performance parameters of the change rule types that each base model is more suitable for processing, and set a corresponding base model pool for each change rule type to save the corresponding base model. By classifying the base models, it is possible to avoid comparing the processing capabilities of each base model for performance parameters during the fault prediction process, reduce the time required during the fault prediction process, and improve the efficiency of the fault prediction process.
[0025] In a possible embodiment, the determining the fault prediction result based on the at least one numerical prediction result includes:
[0026] If the number of the obtained numerical prediction results is one, determine the fault prediction result based on the final numerical prediction result and the change rule type corresponding to the numerical prediction result, where the final numerical prediction result is the numerical prediction result;
[0027] If the number of the obtained numerical prediction results is greater than one, linearly fit the multiple prediction results to obtain the final numerical prediction result, and determine the fault prediction result based on the final numerical prediction result and the corresponding change rule type.
[0028] Determine different processing methods for different numbers of numerical prediction results. If there are multiple numerical prediction results, linear fitting can be used to determine the most accurate numerical prediction result based on the multiple numerical prediction results, improving the efficiency and accuracy of the fault prediction process.
[0029] In a possible embodiment, the determining the fault prediction result based on the final numerical prediction result and the corresponding change rule type includes:
[0030] If the change pattern type of the multiple numerical values is the first type, and the difference between the final numerical value prediction result and the maximum value among the multiple performance parameters is greater than the first threshold, then determine that the fault prediction result is a fault occurrence;
[0031] If the change pattern type of the multiple numerical values is the second type, and the final numerical value prediction result is greater than the second threshold, then determine that the fault prediction result is a fault occurrence;
[0032] If the change pattern type of the multiple numerical values is the third type, and the final numerical value prediction result is not within the target confidence interval; the target confidence interval is determined based on the maximum and minimum values among the multiple numerical values of the same performance parameter.
[0033] Determining different fault prediction methods for different change pattern types of numerical values can more specifically determine the optimal fault prediction method corresponding to each change pattern type, and can also improve the efficiency and accuracy of the fault prediction process.
[0034] In a possible embodiment, after determining the fault prediction result based on the at least one numerical value prediction result, the method further includes:
[0035] If it is determined that the fault prediction result is a fault occurrence, then prompt the user with fault information; the fault information includes at least one of the following: the performance parameter corresponding to the fault, the change pattern type corresponding to the performance parameter corresponding to the fault, the numerical value of the performance parameter corresponding to the fault, and the repair measure for the fault.
[0036] After determining a fault occurrence, the user can be prompted with fault information, enabling the user to repair the fault based on the repair measures included in the fault information, or to determine more suitable repair measures for the fault based on other content, thereby improving the efficiency of the fault repair process.
[0037] In a second aspect, an embodiment of the present application provides a fault prediction device, including:
[0038] A target base model pool determination unit, configured to determine the change pattern type of the multiple numerical values based on the generation time of the multiple numerical values of the same performance parameter obtained and the magnitude relationship of the multiple numerical values, and select the target base model pool corresponding to the determined change pattern type from multiple base model pools;
[0039] A prediction accuracy determination unit, configured to respectively input a first part of the multiple numerical values into each target base model in the target base model pool, and determine the prediction accuracy of each target base model based on a second part of the multiple numerical values and the numerical values output by each target base model; wherein the generation time of the first part of the numerical values is earlier than the generation time of the second part of the numerical values;
[0040] A target base model determination unit, configured to select a set number of target base models based on the prediction accuracy of each target base model;
[0041] A fault prediction result determination unit, configured to respectively input the second part of the numerical values into the set number of target base models to obtain at least one numerical prediction result, and determine a fault prediction result based on the at least one numerical prediction result.
[0042] In a third aspect, an embodiment of the present application provides a computing device, including:
[0043] A memory, configured to store program instructions;
[0044] A processor, configured to call the program instructions stored in the memory and execute the steps included in the method described in the first aspect according to the obtained program instructions.
[0045] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0046] The embodiment of the present application provides a method, which can screen the target base model pool according to the numerical change law types corresponding to different performance parameters, dynamically sort the base models in the target base model pool, select a set number of target base models with the highest prediction accuracy, use the target base models to perform numerical prediction to obtain numerical prediction results, and then use the numerical prediction results to determine the corresponding fault prediction results. Through this method, the influence of data distribution changes on the fault prediction process can be avoided. The base models in different base model pools have different processing capabilities for different data change laws. For different data distributions, that is, different data change law types, base models with stronger processing capabilities can be selected to process data of the corresponding data change law types, which can ensure the accuracy of prediction, avoid the occurrence of faults and data loss, and improve the user experience. Description of the Drawings
[0047] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 Schematic diagram of an application scenario of a fault prediction method provided by an embodiment of the present application;
[0049] Figure 2 Schematic flowchart of a fault prediction method provided by an embodiment of the present application;
[0050] Figure 3 Schematic flowchart of a specific process of a fault prediction method provided by an embodiment of the present application;
[0051] Figure 4 Schematic diagram of the structure of a fault prediction device provided by an embodiment of the present application;
[0052] Figure 5 Schematic diagram of the structure of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following clearly and completely describes the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Among them, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0054] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the text is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone.
[0055] Specifically, in the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation to the present application. Moreover, the "connection" and "coupling" mentioned in the present application, unless otherwise specified, both include direct and indirect connection (coupling).
[0056] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0057] The present application will be further described in detail below with reference to the drawings and specific embodiments.
[0058] In one possible embodiment, Figure 1 FIG. shows a schematic diagram of an application scenario of a fault prediction method provided by an embodiment of the present application. Refer to Figure 1 As shown, in the computing device cluster 100, there are multiple computing devices, such as computing device 110, computing device 120, computing device 130, etc. A fault prediction method provided by an embodiment of the present application can be applied to each computing device in the computing device cluster 100. Among them, each computing device in the computing device cluster 100 can be a Graphics Processing Unit (GPU) server, or a terminal device, or other devices capable of realizing data processing and computing. The present application does not make any limitations here.
[0059] Figure 1 The application scenario of... is only an example of an application scenario for implementing the embodiments of the present application. The embodiments of the present application are not limited to the above Figure 1 described application scenario. Below, in combination with the above-described application scenario, the fault prediction method provided by the exemplary embodiment of the present application will be described with reference to the drawings. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0060] Figure 2The flowchart of a fault prediction method provided by an embodiment of the present application is shown. As Figure 2 shown, the method may include the following steps:
[0061] Step S201: Based on the generation time of multiple values of the same performance parameter obtained and the magnitude relationship of the multiple values, determine the type of change rule of the multiple values, and select the target base model pool corresponding to the determined type of change rule from multiple base model pools.
[0062] In a possible embodiment, before performing step S201, the values of the performance parameter may be obtained periodically. Wherein, the period may be once per second or once per 0.1 second. The present application does not limit this here. After each value of the performance parameter is obtained, the relationship between the value of the performance parameter and the corresponding value range may be determined. If the value of the performance parameter is within the value range, the obtained value of the performance parameter may be directly used in the subsequent steps; if the value of the performance parameter is not within the value range, the target value determined based on the value range may be used as the value of the performance parameter. Wherein, the value range is determined based on the values of the performance parameter that have been obtained. For every N values of the performance parameter obtained, the value range may be determined based on some or all of the N values of the performance parameter obtained, a preset upward floating range, and a preset downward floating range.
[0063] In the process of determining the above value range, N may be a positive integer. For example, if 10 values of the performance parameter have been obtained, the value range may be determined according to the last 5 values of the performance parameter obtained, or according to the last 3 values of the performance parameter obtained, or the value range may be determined according to all 10 values of the performance parameter. Specifically, if the value range is determined by the last 5 values of the performance parameter, the average value of the last 5 values of the 10 performance parameters may be determined, and then the preset upward floating range and the preset downward floating range may be added respectively above and below the average value; it may also be the median of the last 5 values of the 10 performance parameters, and then the preset upward floating range and the preset downward floating range may be added respectively above and below the median.
[0064] In another possible embodiment, after all the data is obtained, the preprocessing process of the values of the performance parameter may also be completed in the following manner, including resampling and filling missing values for all the data. The 3-σ algorithm may also be used to detect outliers, and the outliers may be replaced with the average value of the four points adjacent to the outliers. The Butterworth low-pass filter may also be used to perform a filtering operation on the values of the performance parameter to achieve the effect of highlighting the data trend and reducing the influence of noise.
[0065] In a possible embodiment, the type of change pattern of multiple values can be determined based on the generation time of multiple values of the same performance parameter and the magnitude relationship of the multiple values. Among them, the type of change pattern of the values can include periodic type, stable type, and trend type. The values corresponding to the periodic data can be 88, 90, 92, 88, 90, 92 within a period of time, with every three values forming a cycle. The values corresponding to the stable type can be 90, 90, 90, 90, 90 within a period of time, and the values are stable, so it is stable type data. The values corresponding to the trend type data can be 80, 81, 82, 83, 84, 85 within a period of time. It should be noted that in the actual process of obtaining values, the values may not be as accurate as in the example, but within the allowable tolerance range, the type of change pattern of multiple values can be determined based on the generation time of multiple values of the same performance parameter and the magnitude relationship of the multiple values.
[0066] In another possible embodiment, the STL decomposition method can be used to decompose multiple values of the same performance parameter into three components, namely, periodic component, stable component, and trend component. Then, the Spearman rank correlation coefficient can be used to calculate the correlation degree of each of the three components with the original data, and the same performance parameter can be classified into the component type with the highest correlation coefficient. For example, if the periodic component of a certain performance parameter is 80%, the stable component is 90%, and the trend component is 99%, then it can be considered that this performance parameter is a performance parameter of the trend type.
[0067] In a possible embodiment, before selecting the target base model pool corresponding to the determined change pattern type from multiple base model pools, each base model can be classified first to determine multiple base model pools. First, multiple base models can be prepared, such as ARIMA, SARIMAX, Prophet, XGBoost, MLPLSTM, KNN, GBDT, linear regression, and random forest, etc. The specific number of base models is not limited in this application. Then, multiple groups of training data can be classified according to each change pattern type, that is, three groups of training data can be obtained, namely, periodic type training data, stable type training data, and trend type training data. It should be noted that each change pattern type corresponds to at least one group of training data, and each group of training data should include the data to be processed and the verification data. For any one of the multiple base models, at least one group of training data corresponding to each change pattern type can be input, and based on the output training results and the verification data, the fitting results corresponding to each change pattern type can be determined. Based on the multiple fitting results, the change pattern type corresponding to any one base model can be determined. Finally, at least one base model corresponding to each change pattern type can be saved into the base model corresponding to the change pattern type respectively to obtain multiple base model pools.
[0068] For example, for the SARIMAX base model, the SARIMAX base model is trained using periodic training data, stable training data, and trend training data respectively. The fitting results of the SARIMAX base model for the periodic training data are 95%, for the stable training data are 80%, and for the trend training data are 60%. Then, it can be determined that the type of change pattern corresponding to the SARIMAX base model should be periodic. Therefore, the SARIMAX base model can be saved to the base model pool corresponding to the periodic type.
[0069] In a possible embodiment, after determining multiple base model pools, the target base model pool corresponding to the determined type of change pattern can be selected from the multiple base model pools.
[0070] Step S202: Input the first part of the multiple values into each target base model in the target base model pool, and determine the prediction accuracy of each target base model based on the second part of the multiple values and the values output by each target base model.
[0071] In a possible embodiment, after obtaining the values of multiple performance parameters, the obtained values of the multiple performance parameters can be divided into two parts. The first part of the values is used for processing and prediction by each target base model in the target base model pool. Each target base model in the target base model pool makes a prediction based on the first part of the values, and each target base model outputs a value. The value output by each target base model is compared with the second part of the values to determine the prediction accuracy of each target base model. Usually, the number of the first part of the values is greater than 1, and the number of the second part of the values can be 1 or any number greater than 1. For example, the obtained values of the multiple performance parameters are 90, 91, 92, 93, 94, 95, 96, a total of 7 values. If the first part of the values includes the first 6 values, 5 target base models make predictions based on the first part of the values 90, 91, 92, 93, 94, 95. The value output by target base model A is 90, the value output by target base model B is 91, the value output by target base model C is 93, the value output by target base model D is 96, and the value output by target base model E is 99. At this time, the prediction accuracy of each target base model is determined according to the second part of the value 96. If the first part of the values includes the first 5 values, 5 target base models make predictions based on the first part of the values 90, 91, 92, 93, 94. The values output by target base model A are 90, 95, the values output by target base model B are 91, 93, the values output by target base model C are 93, 94, the values output by target base model D are 96, 97, and the values output by target base model E are 99, 101. Similarly, the output of each target base model is compared with the second part of the values to determine the prediction accuracy of each target base model.
[0072] Step S203: Select a set number of target base models based on the prediction accuracy of each target base model.
[0073] In one possible embodiment, the prediction accuracies of the respective target base models in the target base model pool for the obtained performance parameters can be sorted, and a set number of target base models can be selected therefrom. For example, the prediction accuracies of the respective target base models for the obtained performance parameters are: the prediction accuracy of model A is 70%, the prediction accuracy of model B is 80%, the prediction accuracy of model C is 90%, the prediction accuracy of model D is 99%, and the prediction accuracy of model E is 88%. If the set number is 3, that is, 3 target base models are selected therefrom, model C, model D, and model E can be selected for subsequent steps.
[0074] In another possible embodiment, the Symmetric Mean Absolute Percentage Error (SMAPE) of each model can be evaluated to achieve the dynamic ranking of each base model in the target base model pool, and the method of selecting a set number of target base models from the target base model pool can also be realized.
[0075] In another possible embodiment, training data with different fitting degrees for different change pattern types can be prepared in advance. For example, the fitting degree of periodic training data A to the periodic pattern is 60%, the fitting degree of periodic training data B to the periodic pattern is 70%, the fitting degree of periodic training data C to the periodic pattern is 80%, and the fitting degree of periodic training data D to the periodic pattern is 90%. Through multiple sets of training data with different fitting degrees, determine the data with the fitting degree that each target base model in the target base model pool corresponding to the periodic pattern is best at processing. For example, target base model A is best at processing data with a fitting degree of 90% to the periodic pattern, target base model B is best at processing data with a fitting degree of 60% to the periodic pattern, target base model C is best at processing data with a fitting degree of 80% to the periodic pattern, and target base model D is best at processing data with a fitting degree of 70% to the periodic pattern. Subsequently, based on multiple values of the obtained performance parameters, determine the fitting degree of the obtained values to each change pattern type. For example, the fitting degree of multiple values of the obtained performance parameters to periodic data is 95%. If only 1 target base model needs to be selected, then only target base model A can be selected to implement subsequent steps.
[0076] Step S204: Input the second part of the values into the set number of target base models respectively to obtain at least one numerical prediction result, and determine the fault prediction result based on at least one numerical prediction result.
[0077] In a possible embodiment, after selecting a set number of target base models, the second part of the numerical values that have not been input into the target base models can be input into the target base models to predict subsequent numerical values. After receiving the second part of the numerical values, the target base models can obtain at least one numerical prediction result. Among them, if the set number of selected target base models is 1, the obtained numerical prediction result is 1. If the set number of selected target base models is multiple, the obtained numerical prediction results are also multiple, and each target base model can generate a numerical prediction result.
[0078] In a possible embodiment, if the obtained numerical prediction result is 1, the fault prediction result can be determined based on the type of change law corresponding to the 1 numerical prediction result; if the obtained numerical prediction results are multiple, the multiple prediction results can be linearly fitted to obtain the final numerical prediction result, and the fault prediction result can be determined based on the final numerical prediction result.
[0079] In another possible embodiment, for each of the multiple selected target base models, the second part of the data can be used to perform K-fold cross-validation to generate multiple prediction results, and the average value of the multiple prediction results can be taken, and then the average value can be used as the final numerical prediction result. That is, multiple numerical prediction results are obtained through multiple target base models, and then the obtained multiple numerical prediction results can be input into the meta-model XGBoost to determine the final fault prediction result. This method is particularly suitable for dealing with the situations of high model diversity and large dataset scale, and can effectively improve the generalization ability of the model and the accuracy of prediction.
[0080] In a possible embodiment, the fault prediction result can be determined in the following manner. If the type of change law of multiple numerical values is the first type, and the difference between the final numerical prediction result and the maximum value among multiple performance parameters is greater than the first threshold, it can be determined that the fault prediction result is a fault. Among them, the first type of numerical change law type can include periodic and stable data; if the type of change law of multiple numerical values is the second type, and the final numerical prediction result is greater than the second threshold, it can be determined that the fault prediction result is a fault. Among them, the second type of numerical change law type can include stable and trend data; if the type of change law of multiple numerical values is the third type, and the final numerical prediction result is not within the target confidence interval, it can be determined that the fault prediction result is a fault, where the target confidence interval is determined based on the maximum and minimum values among multiple numerical values of the same performance parameter, and the third type of numerical change law type can include periodic and trend. It should be noted that different fault prediction methods can be determined according to different types of numerical change laws, and different fault prediction methods can also be determined according to different performance parameters. This application does not limit this here.
[0081] In a possible embodiment, if it is determined that the fault prediction result is a fault occurrence, relevant information about the fault can be prompted to the user. Among them, at least one of the performance parameters corresponding to the fault, the type of change rule corresponding to the performance parameter of the fault, the value of the performance parameter corresponding to the fault, and the repair measure of the fault can be prompted. The repair measure of the fault can be to implant a large intelligent question-and-answer model inside the computing device, input the relevant content of the fault, such as the performance parameter corresponding to the fault, the type of change rule corresponding to the performance parameter of the fault, the value of the performance parameter corresponding to the fault, etc. into the large intelligent question-and-answer model, and then display the output of the large intelligent question-and-answer model for the input content to the user to solve the fault.
[0082] Through a fault prediction method, device, computing device, and storage medium provided by an embodiment of the present application, a target base model pool can be screened according to the type of numerical change rule corresponding to different performance parameters, and the base models in the target base model pool can be dynamically sorted to select a set number of target base models with the highest prediction accuracy. The numerical prediction results are obtained by using the target base models for numerical prediction, and then the corresponding fault prediction results are determined. Through this method, the influence of data distribution changes on the fault prediction process can be avoided. The base models in different base model pools have different processing capabilities for different data change rules. For different data distributions, that is, different types of data change rules, base models with stronger processing capabilities can be selected to process the data corresponding to the data change rule types, which can ensure the accuracy of prediction, avoid the occurrence of faults and data loss, and improve the user experience.
[0083] In a specific embodiment, Figure 3 shows a specific flowchart of a fault prediction method provided by an embodiment of the present application, as Figure 3 shown, the method may include the following steps:
[0084] Step S301, periodically obtain the value of the performance parameter.
[0085] Step S302, for each obtained value of the performance parameter, determine the relationship between the value of the performance parameter and the value range.
[0086] Step S303, if the value of the performance parameter is not within the value range, use the target value determined based on the value range as the value of the performance parameter.
[0087] In a possible embodiment, the value range is determined based on the values of the acquired performance parameters in the following manner: for every N values of the performance parameters acquired, the value range is determined based on some or all of the N values of the acquired performance parameters, a preset upward floating range, and a preset downward floating range; where N is a positive integer.
[0088] Step S304: Classify multiple groups of training data according to each type of change rule, so that each type of change rule corresponds to at least one group of training data.
[0089] Step S305: For any base model, input at least one group of training data corresponding to each type of change rule into any base model, and determine the fitting result corresponding to each type of change rule based on the output training result and the validation data.
[0090] Step S306: Determine the type of change rule corresponding to any base model based on multiple fitting results.
[0091] Step S307: Save at least one base model corresponding to each type of change rule into the base model pool corresponding to the type of change rule.
[0092] Step S308: Based on the generation time of multiple values of the same performance parameter acquired and the magnitude relationship of the multiple values, determine the type of change rule of the multiple values, and select the target base model pool corresponding to the determined type of change rule from multiple base model pools.
[0093] Step S309: Input the first part of the multiple values into each target base model in the target base model pool, and determine the prediction accuracy of each target base model based on the second part of the multiple values and the values output by each target base model.
[0094] Step S310: Select a set number of target base models based on the prediction accuracy of each target base model.
[0095] Step S311: Input the second part of the values into the set number of target base models to obtain at least one numerical prediction result.
[0096] Step S312: Determine whether the number of numerical prediction results is one. If so, execute Step S313; if not, execute Step S314.
[0097] Step S313: Determine the fault prediction result based on the final numerical prediction result and the type of change rule corresponding to the numerical prediction result.
[0098] In a possible embodiment, the final numerical prediction result is the numerical prediction result.
[0099] Step S314: Linearly fit multiple prediction results to obtain a final numerical prediction result, and determine a fault prediction result based on the final numerical prediction result and the corresponding change law type.
[0100] In a possible embodiment, if the change law type of multiple numerical values is the first type, and the difference between the final numerical prediction result and the maximum value among multiple performance parameters is greater than the first threshold, then determine that the fault prediction result is a fault occurrence.
[0101] If the change law type of multiple numerical values is the second type, and the final numerical prediction result is greater than the second threshold, then determine that the fault prediction result is a fault occurrence.
[0102] If the change law type of multiple numerical values is the third type, and the final numerical prediction result is not within the target confidence interval, then determine that the fault prediction result is a fault occurrence; the target confidence interval is determined based on the maximum and minimum values among multiple numerical values of the same performance parameter.
[0103] Step S315: If it is determined that the fault prediction result is a fault occurrence, then prompt the user with fault information; the fault information includes at least one of the following: the performance parameter corresponding to the fault, the change law type corresponding to the performance parameter corresponding to the fault, the numerical value of the performance parameter corresponding to the fault, and the repair measure for the fault.
[0104] Based on the same inventive concept, Figure 4 The following is a structural block diagram of a fault prediction device provided by an embodiment of the present application. As Figure 4 shown, the fault prediction device 400 may include:
[0105] A target base model pool determination unit 401, configured to determine the change law type of multiple numerical values of the same performance parameter based on the generation time of the multiple numerical values and the magnitude relationship of the multiple numerical values, and select a target base model pool corresponding to the determined change law type from multiple base model pools.
[0106] A prediction accuracy determination unit 402, configured to respectively input the first part of the multiple numerical values into each target base model in the target base model pool, and determine the prediction accuracy of each target base model based on the second part of the multiple numerical values and the numerical values output by each target base model; wherein the generation time of the first part of the numerical values is earlier than the generation time of the second part of the numerical values.
[0107] A target base model determination unit 403, configured to select a set number of target base models based on the prediction accuracy of each target base model.
[0108] A fault prediction result determination unit 404 is configured to input the second part of the numerical values into the set number of target base models respectively to obtain at least one numerical prediction result, and determine a fault prediction result based on the at least one numerical prediction result.
[0109] In a possible implementation manner, the target base model pool determination unit 401 may further be configured to periodically obtain the numerical values of performance parameters;
[0110] For each obtained numerical value of a performance parameter, determine the relationship between the numerical value of the performance parameter and the value range; the value range is determined based on the obtained numerical values of the performance parameters;
[0111] If the numerical value of the performance parameter is not within the value range, use the target value determined based on the value range as the numerical value of the performance parameter.
[0112] In a possible implementation manner, the target base model pool determination unit 401 may further be configured to, for each obtained N numerical values of performance parameters, determine the value range based on some or all of the obtained N numerical values of performance parameters, a preset upward floating range, and a preset downward floating range; where N is a positive integer.
[0113] In a possible implementation manner, the target base model pool determination unit 401 may further be configured to classify the multiple groups of training data according to each change rule type, so that each change rule type corresponds to at least one group of training data; each group of the training data includes data to be processed and verification data;
[0114] For any base model, input the at least one group of training data corresponding to each change rule type into the any base model, and determine the fitting result corresponding to each change rule type based on the output training result and the verification data;
[0115] Determine the change rule type corresponding to the any base model based on multiple fitting results;
[0116] Save the at least one base model corresponding to each change rule type into the base model pool corresponding to the change rule type respectively.
[0117] In a possible implementation manner, the fault prediction result determination unit 404 is specifically configured to, if the number of obtained numerical prediction results is one, determine the fault prediction result based on the final numerical prediction result and the change rule type corresponding to the numerical prediction result, where the final numerical prediction result is the numerical prediction result;
[0118] If the number of the obtained numerical prediction results is greater than one, linearly fit the multiple prediction results to obtain a final numerical prediction result, and determine the fault prediction result based on the final numerical prediction result and the corresponding change law type.
[0119] In a possible implementation manner, the fault prediction result determination unit 404 is specifically configured to, if the change law type of the multiple numerical values is the first type, and the difference between the final numerical prediction result and the maximum value among the multiple performance parameters is greater than a first threshold, determine that the fault prediction result is a fault occurrence;
[0120] If the change law type of the multiple numerical values is the second type, and the final numerical prediction result is greater than a second threshold, determine that the fault prediction result is a fault occurrence;
[0121] If the change law type of the multiple numerical values is the third type, and the final numerical prediction result is not within the target confidence interval, determine that the fault prediction result is a fault occurrence; the target confidence interval is determined based on the maximum value and the minimum value among the multiple numerical values of the same performance parameter.
[0122] In a possible implementation manner, the fault prediction result determination unit 404 may also be configured to, if it is determined that the fault prediction result is a fault occurrence, prompt the user with fault information; the fault information includes at least one of the following: the performance parameter corresponding to the fault, the change law type corresponding to the performance parameter corresponding to the fault, the numerical value of the performance parameter corresponding to the fault, and the repair measure for the fault.
[0123] Based on the same inventive concept, an embodiment of the present application provides a computing device, which can implement the functions of the fault prediction method described above. Please refer to Figure 5 This computing device 500 includes a memory 501, a processor 502, and a bus 503.
[0124] The memory 501 is used to store a computer program executed by the processor 502. The memory 501 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and programs required to run an instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0125] The memory 501 can be a volatile memory, such as a random-access memory (RAM); the memory 501 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or the memory 501 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 501 can be a combination of the above memories.
[0126] The processor 502 can include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 502 is used to implement the fault prediction method in the above embodiments when calling the computer program stored in the memory 501.
[0127] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 501 and the processor 502 is not limited. In the embodiments of the present application Figure 5 it is shown that the memory 501 and the processor 502 are connected through a bus 503. The bus 503 is shown as a thick line in Figure 5 The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 503 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 only a thick line is used to represent it in
[0128] Based on the same inventive concept, the embodiments of the present application provide a computer-readable storage medium. The computer program product includes: computer program code. When the computer program code runs on a computer, it causes the computer to execute the fault prediction method described in any of the foregoing. Since the principle of solving the problem by the above computer-readable storage medium is similar to that of the fault prediction method, the implementation of the above computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be described again.
[0129] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0130] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0131] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0133] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A fault prediction method, characterized in that, Including: Determine the change rule type of the multiple values based on the generation time of the multiple values of the same performance parameter obtained and the magnitude relationship of the multiple values, and select the target base model pool corresponding to the determined change rule type from multiple base model pools; Input the first part of the multiple values into each target base model in the target base model pool respectively, and determine the prediction accuracy of each target base model based on the second part of the multiple values and the values output by each target base model; wherein the generation time of the first part of the values is earlier than the generation time of the second part of the values; Select a set number of target base models based on the prediction accuracy of each target base model; Input the second part of the values into the set number of target base models respectively to obtain at least one value prediction result, and determine the fault prediction result based on the at least one value prediction result.
2. The method according to claim 1, characterized in that, Before determining the change rule type of the multiple values based on the generation time of the multiple values of the same performance parameter obtained and the magnitude relationship of the multiple values, the method further includes: Periodically obtain the values of the performance parameter; For each obtained value of the performance parameter, determine the relationship between the value of the performance parameter and the value range; the value range is determined based on the obtained values of the performance parameter; If the value of the performance parameter is not within the value range, use the target value determined based on the value range as the value of the performance parameter.
3. The method according to claim 2, wherein The value range is determined based on the obtained values of the performance parameter in the following manner: For every N obtained values of the performance parameter, determine the value range based on some or all of the N obtained values of the performance parameter, a preset upward floating range, and a preset downward floating range; wherein, N is a positive integer.
4. The method according to claim 1, characterized in that Each base model in the base model pool is obtained by the following method: Classify the multiple groups of training data according to each change rule type, so that each change rule type corresponds to at least one group of training data; each group of training data contains data to be processed and verification data; For any base model, input at least one group of training data corresponding to each change rule type into the any base model, and determine the fitting result corresponding to each change rule type based on the output training result and the verification data; Determine the change rule type corresponding to the any base model based on multiple fitting results; Save at least one base model corresponding to each change rule type into the base model pool corresponding to the change rule type respectively.
5. The method according to claim 1, characterized in that, The determining the fault prediction result based on the at least one value prediction result includes: If the number of obtained value prediction results is one, determine the fault prediction result based on the final value prediction result and the change rule type corresponding to the value prediction result, wherein the final value prediction result is the value prediction result; If the number of the obtained numerical prediction results is greater than one, linearly fit the multiple prediction results to obtain a final numerical prediction result, and determine the fault prediction result based on the final numerical prediction result and the corresponding change law type.
6. The method according to claim 5, characterized in that, The determining the fault prediction result based on the final numerical prediction result and the corresponding change law type includes: If the change law type of the multiple numerical values is the first type, and the difference between the final numerical prediction result and the maximum value among the multiple performance parameters is greater than a first threshold, determine that the fault prediction result is a fault occurs; If the change law type of the multiple numerical values is the second type, and the final numerical prediction result is greater than a second threshold, determine that the fault prediction result is a fault occurs; If the change law type of the multiple numerical values is the third type, and the final numerical prediction result is not within the target confidence interval, determine that the fault prediction result is a fault occurs; the target confidence interval is determined based on the maximum value and the minimum value among the multiple numerical values of the same performance parameter.
7. The method according to claim 1, wherein After determining the fault prediction result based on the at least one numerical prediction result, the method further includes: If it is determined that the fault prediction result is a fault occurs, prompt the user with fault information; the fault information includes at least one of the following: the performance parameter corresponding to the fault, the change law type corresponding to the performance parameter corresponding to the fault, the numerical value of the performance parameter corresponding to the fault, and the repair measure for the fault.
8. A fault prediction device, characterized in that, Includes: A target base model pool determination unit, configured to determine the change law type of the multiple numerical values of the same performance parameter based on the generation time of the multiple numerical values and the magnitude relationship of the multiple numerical values, and select the target base model pool corresponding to the determined change law type from multiple base model pools; A prediction accuracy determination unit, configured to input the first part of the multiple numerical values into each target base model in the target base model pool respectively, and determine the prediction accuracy of each target base model based on the second part of the multiple numerical values and the numerical values output by each target base model; wherein the generation time of the first part of the numerical values is earlier than the generation time of the second part of the numerical values; A target base model determination unit, configured to select a set number of target base models based on the prediction accuracy of each target base model; A fault prediction result determination unit, configured to input the second part of the numerical values into the set number of target base models respectively to obtain at least one numerical prediction result, and determine the fault prediction result based on the at least one numerical prediction result.
9. A computing device, characterized in that, Includes: A memory, configured to store program instructions; A processor, configured to call the program instructions stored in the memory and execute the steps included in the method according to any one of claims 1-7 according to the obtained program instructions.
10. A computer-readable storage medium storing a computer program therein, characterized in that: When the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Hard disk fault prediction model establishing method based on model fusion and application thereof
CN112214369A
Equipment fault prediction method and device and server
CN113971777A
Training method and device of user prediction model, prediction method and device and storage medium
CN114118192A
Fault testing method and device, server and storage medium
CN119668957A
Method, Apparatus, and Device for Updating Hard Disk Prediction Model, and Medium
US20230004824A1