Database performance prediction method based on machine learning

Through machine learning-based methods, feature extraction and important parameter recognition of database configuration parameter space, combined with Bayesian optimized XGBoost model, considering the impact of the database running environment, the problem of manual configuration parameters in the existing technology is solved, and more efficient and accurate database performance prediction is achieved.

CN119961110APending Publication Date: 2025-05-09ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411666414.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-20
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing database performance prediction methods rely on manual identification of important configuration parameters, are costly and error-prone, and fail to fully consider the impact of the database operating environment on performance, resulting in insufficient accuracy and reliability of the prediction results.

Method used

Using a machine learning-based method, feature extraction is performed on the database configuration parameter space, important configuration parameters are identified through random forest model and SHAP value analysis, and combined with Bayesian optimized XGBoost model, the performance prediction model is built to consider the impact of the database running environment.

Benefits of technology

It reduces the cost and time of manually identifying configuration parameters, improves the accuracy and reliability of database performance prediction, enables more scientific selection of features and fully considers the impact of the operating environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961110A_ABST
    Figure CN119961110A_ABST
Patent Text Reader

Abstract

The invention discloses a database performance prediction method based on machine learning. The method is characterized in that a configuration parameter space is obtained, and the value range of each configuration parameter is determined; performing feature extraction based on machine learning to obtain an optimal feature subset; performing random sampling on the optimal feature subset for many times to obtain a plurality of second samples; using YCSB to obtain database response delay of each second sample as a database performance label; obtaining a plurality of database operation environments, and performing a pressure test based on the performance benchmark test tool to obtain a performance evaluation index in each database operation environment; all the second samples, all the performance evaluation indexes and database performance labels are input into an XGBoost model based on Bayesian optimization for fitting, and a performance prediction model is obtained; inputting the configuration parameter space and the performance evaluation index of the to-be-tested database into the performance prediction model to complete performance prediction; the method has the advantage that the accuracy and reliability of database performance prediction results are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a performance prediction method, in particular to a database performance prediction method based on machine learning. Background Art

[0002] With the rapid development of enterprise applications, databases, as the core module for data storage and processing, play a vital role. In order to cope with tasks in different scenarios, database systems provide various configuration parameters that can optimize database performance, improve data processing efficiency, and ensure data security. With the continuous upgrade of database versions, the number and complexity of configuration parameters are also increasing. Therefore, how to determine configuration parameters and predict database performance has become an important research topic.

[0003] At present, the mainstream database performance prediction method is: first manually select several configuration parameters that have a greater impact on database performance based on expert experience, then sample the values ​​of the configuration parameters to obtain a training set, then fit a prediction model based on the training set based on a regression algorithm, and finally predict the performance of the database through the prediction model. However, the following problems remain:

[0004] 1. Identifying important configuration parameters manually is an extremely complex, time-consuming and labor-intensive task, which often requires a lot of resources and high costs. This process usually relies on the expertise and experience of domain experts, and these experts are paid at a high salary, which significantly increases the cost of manually identifying important configuration parameters. In addition, during the manual screening and verification process, experts need to test and evaluate the impact of each configuration option on system performance one by one, which is not only time-consuming but also prone to errors.

[0005] 2. In the process of fitting the prediction model, current research only considers database configuration parameters as input, and does not consider the impact of the database operating environment on performance. This is because the operating environment is usually difficult to quantify. It is difficult to characterize the actual operating environment of the database only by CPU specifications, memory specifications, and CPU utilization, and thus cannot be applied to the prediction model. This limitation causes the prediction model to be unable to fully reflect the complexity and variability of the actual operating environment when predicting database performance, affecting the accuracy and reliability of the prediction results. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a database performance prediction method based on machine learning, which not only reduces the cost but also improves the accuracy and reliability of the database performance prediction results.

[0007] The technical solution adopted by the present invention to solve the above technical problem is: a database performance prediction method based on machine learning, comprising the following steps:

[0008] Step ①, manually filter out the configuration parameters that are irrelevant to the performance in the database to be tested, and form the remaining configuration parameters into a configuration parameter space;

[0009] Step ②, determining the value range of each configuration parameter in the configuration parameter space according to the database description document, and obtaining the configuration parameter space after the value range is determined;

[0010] Step ③, based on machine learning, feature extraction is performed on the configuration parameter space after the value range is determined to obtain the optimal feature subset. The specific operation process is as follows:

[0011] Step ③-1, the configuration parameter space after the value range is determined is recorded as C;

[0012] Step ③-2, according to the value range of each configuration parameter, perform k random sampling on each configuration parameter in C to obtain k first samples, and form all the first samples into a sample set;

[0013] Step ③-3, label the sample set to obtain a labeled sample set, specifically: use YCSB to perform a performance stress test on each first sample in the sample set, obtain the response delay of each first sample, and use it as the label value of each first sample, and form all the first samples with label values ​​into a labeled sample set;

[0014] Step ③-4, build a random forest model:

[0015] Step ③-4-1, use the bootstrap method to sample the first sample in the labeled sample set multiple times, and form a sub-dataset from the first sample obtained from each sampling;

[0016] Step ③-4-2, randomly select A configuration parameters in each sub-dataset as candidate features of the sub-dataset;

[0017] Step ③-4-3, build a corresponding decision tree based on the candidate features of each sub-dataset, and use the Gini index as the judgment criterion for splitting attributes when building the decision tree;

[0018] Step ③-4-4, combine all constructed decision trees to obtain a random forest model;

[0019] Step ③-5, obtain the SHAP value of each configuration parameter in the random forest model. The SHAP value is used to quantify the marginal contribution of each configuration parameter to the prediction result of the random forest model. Specifically: the contribution of the configuration parameter g to the prediction result of the first sample x, that is, the SHAP value is recorded as SHAP g (f, x), Among them, T represents the total number of decision trees in the random forest, Pathst (x) represents the set of all paths that the first sample x passes through on the t-th decision tree, p represents the p-th path in the set of all paths, and f t represents the predicted value on the tth decision tree, Δf t (p|g) represents the change in the predicted value caused by the configuration parameter g on the pth path, 1≤x≤k, 1≤g≤n, and n represents the total number of configuration parameters in C;

[0020] Step ③-6, select the configuration parameter with the smallest SHAP value from the random forest model through the sequential backward search algorithm, and delete the configuration parameter from C to obtain a new configuration parameter space, which is recorded as C';

[0021] Step ③-7, determine whether the total number of configuration parameters in C' is less than or equal to a preset threshold. If so, C' is taken as the optimal feature subset and output; if not, C' is taken as the new C and return to step ③-2;

[0022] Step ④, according to the value range of each configuration parameter, perform D random sampling on each configuration parameter in the optimal feature subset to obtain D second samples;

[0023] Step ⑤, use YCSB to obtain the database response delay of each second sample, and group the database response delays of all second samples into a set as a database performance label;

[0024] Step ⑥, select multiple different database operating environments, perform stress testing on the content of each database operating environment based on a performance benchmark testing tool, and obtain performance evaluation indicators for each database operating environment;

[0025] Step 7: input all the second samples, the performance evaluation indicators of all database operating environments, and the database performance labels into the XGBoost model based on Bayesian optimization for fitting, and obtain a performance prediction model;

[0026] Step ⑧, input the configuration parameter space and performance evaluation index of the database to be tested after the value range is determined into the performance prediction model, obtain the database response delay, and complete the performance prediction.

[0027] Compared with the prior art, the advantages of the present invention are that it extracts features from the configuration parameter space after determining the value range based on machine learning to obtain the optimal feature subset, which can efficiently identify the configuration options that have a greater impact on database performance, reducing the time cost and labor cost required to be invested. By calculating the SHAP value, it can more accurately measure the contribution value of each configuration parameter to the performance, thereby identifying the most important configuration parameters in the feature extraction process. This method ensures the scientificity and objectivity of the feature selection process by evaluating the marginal contribution of each configuration parameter to the output of the random forest model. Based on the performance benchmarking tool, a performance stress test is performed on the operating environment to obtain performance indicators, and then the predicted feature space is expanded, which fully considers the impact of the database operating environment on the performance and solves the problem that the operating environment is usually difficult to quantify, so that the database operating environment information can be effectively quantified and used for more accurate prediction, thereby improving the accuracy and reliability of the database performance prediction results.

[0028] Further, in step ③-4-2, Where B represents the total number of configuration parameters in C.

[0029] Furthermore, in step ④, D ≥ 50.

[0030] Furthermore, in step ⑥, the contents of the database operating environment include CPU, memory, physical hard disk, files and network; the performance evaluation indicators include CPU score, memory bandwidth, physical hard disk IO rate, file IO rate and network bandwidth; the CPU score is obtained based on sysbench testing CPU performance, the memory bandwidth is obtained based on stream testing memory, the physical hard disk IO rate is obtained based on fio testing physical hard disk IO, the file IO rate is obtained based on fio testing file IO, and the network bandwidth is obtained based on netperf testing network. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0032] Figure 2 It is a specific flow diagram of step ③ in the present invention. DETAILED DESCRIPTION

[0033] The present invention is further described in detail below with reference to the accompanying drawings.

[0034] like Figure 1 As shown, a database performance prediction method based on machine learning is characterized by comprising the following steps:

[0035] Step ①, manually filter out the configuration parameters that are irrelevant to the performance in the database to be tested, and form the remaining configuration parameters into a configuration parameter space;

[0036] Step ②, determining the value range of each configuration parameter in the configuration parameter space according to the database description document, and obtaining the configuration parameter space after the value range is determined;

[0037] Configuration parameters in a database can usually be divided into two types: discrete and continuous. For discrete configuration parameters, optional values ​​are usually given in the database description document, such as 0, 1, and 2; for continuous configuration parameters, the database description document usually gives the maximum and minimum values. The value range of each configuration parameter can be determined by combining the database description document and the actual operating environment of the current database.

[0038] Step ③, based on machine learning, feature extraction is performed on the configuration parameter space after the value range is determined to obtain the optimal feature subset;

[0039] like Figure 2 As shown, the specific operation process of step ③ is as follows:

[0040] Step ③-1, the configuration parameter space after the value range is determined is recorded as C;

[0041] Step ③-2, according to the value range of each configuration parameter, perform k random sampling on each configuration parameter in C to obtain k first samples, and form all the first samples into a sample set; k ≥ 50;

[0042] For example, C includes specifying the InnoDB buffer pool size, the upper limit of the number of concurrent threads, the thread pool cache size, the memory size that can be allocated for sorting operations, the memory size requested and released for operations, and the memory cache size of temporary tables. Some examples of the first sample are shown in Table 1.

[0043] Table 1 Example of the first sample of part

[0044]

[0045]

[0046] Step ③-3, label the sample set to obtain a labeled sample set, specifically: use YCSB to perform a performance stress test on each first sample in the sample set, obtain the response delay of each first sample, and use it as the label value of each first sample, and form all the first samples with label values ​​into a labeled sample set;

[0047] Step ③-4, build a random forest model:

[0048] Step ③-4-1, use the bootstrap method to sample the first sample in the labeled sample set multiple times, and form a sub-dataset from the first sample obtained from each sampling;

[0049] Step ③-4-2, randomly select A configuration parameters in each sub-dataset as candidate features of the sub-dataset;

[0050] in, Where B represents the total number of configuration parameters in C;

[0051] Step ③-4-3, build a corresponding decision tree based on the candidate features of each sub-dataset, and use the Gini index as the judgment criterion for splitting attributes when building the decision tree;

[0052] Step ③-4-4, combine all constructed decision trees to obtain a random forest model;

[0053] Step ③-5, obtain the SHAP value of each configuration parameter in the random forest model. The SHAP value is used to quantify the marginal contribution of each configuration parameter to the prediction result of the random forest model. Specifically: the contribution of the configuration parameter g to the prediction result of the first sample x, that is, the SHAP value is recorded as SHAP g (f, x), Among them, T represents the total number of decision trees in the random forest, Paths t (x) represents the set of all paths that the first sample x passes through on the t-th decision tree, p represents the p-th path in the set of all paths, and f t represents the predicted value on the tth decision tree, Δf t (p|g) represents the change in the predicted value caused by the configuration parameter g on the pth path, 1≤x≤k, 1≤g≤n, and n represents the total number of configuration parameters in C;

[0054] Step ③-6, select the configuration parameter with the smallest SHAP value from the random forest model through the sequential backward search algorithm, and delete the configuration parameter from C to obtain a new configuration parameter space, which is recorded as C';

[0055] Step ③-7, determine whether the total number of configuration parameters in C' is less than or equal to a preset threshold. If so, C' is used as the optimal feature subset and output; if not, C' is used as the new C and returns to step ③-2. The preset threshold is defined by the user according to the version of the database and is generally set to 10.

[0056] Step ④, according to the value range of each configuration parameter, perform D random sampling on each configuration parameter in the optimal feature subset to obtain D second samples; where D≥≥50;

[0057] Step ⑤, use YCSB to obtain the database response delay of each second sample, and group the database response delays of all second samples into a set as a database performance label;

[0058] For example: Assuming that after the previous steps, the optimal feature subset obtained is {specified InnoDB buffer pool size, thread pool cache size}, the D second samples obtained are shown in Table 2, and the database performance labels are shown in Table 3;

[0059] Table 2 Example of the second sample of part

[0060]

[0061] Table 3 Examples of performance labels for some databases

[0062] Second sample number Database Performance Tags 1 0.6s 2 0.7s 3 1.1s 4 0.3s ... ... D 0.7s

[0063] Step ⑥, select multiple different database operating environments, perform stress tests on the content of each database operating environment based on the performance benchmark test tool, and obtain the performance evaluation index of each database operating environment; wherein, a database operating environment is a computer, and the content of the database operating environment includes CPU, memory, physical hard disk, file and network; the performance evaluation index includes CPU running score, memory bandwidth, physical hard disk IO rate, file IO rate and network bandwidth; CPU running score is obtained based on sysbench test CPU performance, memory bandwidth is obtained based on stream test memory, physical hard disk IO rate is obtained based on fio test physical hard disk IO, file IO rate is obtained based on fio test file IO, and network bandwidth is obtained based on netperf test network; some performance evaluation indicators are shown in Table 4;

[0064] Table 4 Examples of some performance evaluation indicators

[0065] Database operating environment number CPU Benchmarks Memory bandwidth Physical hard disk IO rate Network bandwidth 1 0.10 0.10 0.10 0.10 2 0.20 0.20 0.20 0.20 3 0.32 0.32 0.32 0.32 4 0.33 0.33 0.33 0.33

[0066] Step 7: input all the second samples, the performance evaluation indicators of all database operating environments, and the database performance labels into the XGBoost model based on Bayesian optimization for fitting, and obtain a performance prediction model;

[0067] Step ⑧, input the configuration parameter space and performance evaluation index of the database to be tested after the value range is determined into the performance prediction model, obtain the database response delay, and complete the performance prediction.

[0068] Explanation of terms in this patent:

[0069] YCSB (Yahoo! Cloud Serving Benchmark) is a benchmark tool for evaluating the performance of NoSQL databases;

[0070] Bootstrap Method (Bootstrapping or self-service sampling method) is a uniform sampling with replacement from a given training set, that is, every time a sample is selected, it is equally likely to be selected again and added to the training set again;

[0071] Sysbench is a modular, cross-platform, multi-threaded benchmark tool that can perform performance tests in many aspects such as CPU / memory / threads / IO / database.

[0072] Stream is a benchmark tool used to evaluate the memory bandwidth performance of a computer system;

[0073] fio is an open source IO (input and output) stress testing tool, mainly used to test the IO performance of the disk;

[0074] Netperf is a widely used network performance measurement tool, mainly used to evaluate and analyze network transmission performance based on TCP or UDP protocols.

[0075] References for XGBoost model based on Bayesian optimization: Su J, Wang Y, Niu X, et al. Prediction of ground surface settlement by shield tunneling using XGBoost and Bayesian Optimization [J]. Engineering Applications of Artificial Intelligence, 2022, 114: 105020.

Claims

1. A database performance prediction method based on machine learning, characterized in that The following steps are involved: Step ①, manually filter out the configuration parameters that are irrelevant to the performance in the database to be tested, and form the remaining configuration parameters into a configuration parameter space; Step ②, determining the value range of each configuration parameter in the configuration parameter space according to the database description document, and obtaining the configuration parameter space after the value range is determined; Step ③, based on machine learning, feature extraction is performed on the configuration parameter space after the value range is determined to obtain the optimal feature subset. The specific operation process is as follows: Step ③-1, the configuration parameter space after the value range is determined is recorded as C; Step ③-2, according to the value range of each configuration parameter, perform k random sampling on each configuration parameter in C to obtain k first samples, and form all the first samples into a sample set; Step ③-3, label the sample set to obtain a labeled sample set, specifically: use YCSB to perform a performance stress test on each first sample in the sample set, obtain the response delay of each first sample, and use it as the label value of each first sample, and form all the first samples with label values ​​into a labeled sample set; Step ③-4, build a random forest model: Step ③-4-1, use the bootstrap method to sample the first sample in the labeled sample set multiple times, and form a sub-dataset from the first sample obtained from each sampling; Step ③-4-2, randomly select A configuration parameters in each sub-dataset as candidate features of the sub-dataset; Step ③-4-3, build a corresponding decision tree based on the candidate features of each sub-dataset, and use the Gini index as the judgment criterion for splitting attributes when building the decision tree; Step ③-4-4, combine all constructed decision trees to obtain a random forest model; Step ③-5, obtain the SHAP value of each configuration parameter in the random forest model. The SHAP value is used to quantify the marginal contribution of each configuration parameter to the prediction result of the random forest model. Specifically: the contribution of the configuration parameter g to the prediction result of the first sample x, that is, the SHAP value is recorded as SHAP g (f, x), Among them, T represents the total number of decision trees in the random forest, Paths t (x) represents the set of all paths that the first sample x passes through on the t-th decision tree, p represents the p-th path in the set of all paths, and f t represents the predicted value on the tth decision tree, Δf t (p|g) represents the change in the predicted value caused by the configuration parameter g on the pth path, 1≤x≤k, 1≤g≤n, and n represents the total number of configuration parameters in C; Step ③-6, select the configuration parameter with the smallest SHAP value from the random forest model through the sequential backward search algorithm, and delete the configuration parameter from C to obtain a new configuration parameter space, which is recorded as C'; Step ③-7, determine whether the total number of configuration parameters in C' is less than or equal to a preset threshold. If so, C' is taken as the optimal feature subset and output; if not, C' is taken as the new C and return to step ③-2; Step ④, according to the value range of each configuration parameter, perform D random sampling on each configuration parameter in the optimal feature subset to obtain D second samples; Step ⑤, use YCSB to obtain the database response delay of each second sample, and group the database response delays of all second samples into a set as a database performance label; Step ⑥, select multiple different database operating environments, perform stress testing on the content of each database operating environment based on a performance benchmark testing tool, and obtain performance evaluation indicators for each database operating environment; Step 7: input all the second samples, the performance evaluation indicators of all database operating environments, and the database performance labels into the XGBoost model based on Bayesian optimization for fitting, and obtain a performance prediction model; Step ⑧, input the configuration parameter space and performance evaluation index of the database to be tested after the value range is determined into the performance prediction model, obtain the database response delay, and complete the performance prediction.

2. The method for predicting database performance based on machine learning according to claim 1, characterized in that In step ③-4-2, Where B represents the total number of configuration parameters in C.

3. The method for predicting database performance based on machine learning according to claim 1, characterized in that In step ④, D≥50.

4. The method for predicting database performance based on machine learning according to claim 1, characterized in that In step ⑥, the content of the database operating environment includes CPU, memory, physical hard disk, files and network; the performance evaluation indicators include CPU score, memory bandwidth, physical hard disk IO rate, file IO rate and network bandwidth; the CPU score is obtained based on sysbench testing CPU performance, the memory bandwidth is obtained based on stream testing memory, the physical hard disk IO rate is obtained based on fio testing physical hard disk IO, the file IO rate is obtained based on fio testing file IO, and the network bandwidth is obtained based on netperf testing network.