Hybrid department data center online service performance interference quantitative evaluation method
By collecting and analyzing micro-architecture-level indicators in the mixed-part scenario of the data center, and combining deep learning technology, a quantitative evaluation model for online service performance interference is built, which solves the problem that online service performance interference cannot be accurately evaluated in the data center, and achieves more accurate performance interference evaluation.
Patent Information
- Application Number
- CN202510141902.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-03
AI Technical Summary
In the new mixed-industry scenario of data centers, the performance interference of online services cannot be accurately quantified, resulting in a decline in service quality of online services and the failure to meet service level goals (SLO).
A hybrid online service performance interference quantitative evaluation model based on micro-architecture-level indicators is proposed. By collecting indicators related to the processor micro-architecture, the information entropy is calculated, and a quantitative evaluation model for online service performance interference is constructed by combining the autoencoder, multi-head attention mechanism and multi-layer perception machine.
Through this model, the performance interference of online services can be accurately quantified and evaluated. The system entropy index has a strong correlation with the application-level indicator tail latency of online services, and the performance interference situation can be more accurately evaluated.
Smart Images

Figure CN120086104A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer architecture, and particularly relates to a method for quantitatively evaluating the interference of online service performance in a hybrid data center. Background Art
[0002] Under the influence of the cloud computing technology trend, data centers in the Internet and related industries have developed rapidly in terms of quantity, scale, and scalability. In a data center, loads can generally be divided into two categories: online services and offline batch processing jobs. Online services refer to services with strict requirements for real-time performance and service stability, such as web services, streaming computing, etc. Such services usually have two characteristics: 1) Since they receive and process user requests in real time, they are sensitive to latency responses, and users have high requirements for real-time performance; 2) Since the intensity of user requests arriving is different, this results in different resource requirements at different time points. The most important indicator for evaluating the quality of online services is the tail latency, that is, the latency at the 95th percentile among all latency indicators. Offline batch processing jobs are not driven by user requests, but by large-scale computing requirements. Batch processing jobs can be split into multiple tasks and processed in parallel on different nodes, such as big data analysis services, machine learning training tasks, etc. Such tasks generally do not require real-time performance, but require a higher overall throughput.
[0003] In traditional data centers, in order to ensure the quality of online services, online services and offline batch processing jobs are usually separately deployed on independent servers. However, this independent deployment mode has a very low utilization rate of node resources. Therefore, a hybrid deployment mode (hybrid) of online and offline services has emerged in data centers. Hybrid means that online services and offline batch processing jobs are simultaneously deployed on a single node. Offline batch processing jobs can make full use of the resource fragments generated during the operation of online services. Through this deployment mode, the utilization rate of node resources can be significantly improved. For example, for large cloud service providers such as Alibaba and Google, it has become a popular practice to deploy different types of services together for operation.
[0004] Since the emergence of the hybrid deployment mode in data centers, its deployment method has also changed with the development of data centers. Under the traditional hybrid deployment mode of data centers, there are fixed and proprietary online services within the enterprise cloud. At the same time, when the online service requests and resource demands are low, offline batch processing jobs are allowed to be accepted to improve the overall resource utilization rate. In this deployment mode, since the online services are proprietary to the enterprise, the source code and related evaluation metrics of the online services can be obtained, and the quantity and types of the enterprise's proprietary online services are relatively fixed. With the development of data centers, the hybrid deployment mode has changed. In the new hybrid scenario, the data center can externally receive the deployment and operation of online services, and at the same time, it will have fixed offline batch processing jobs. Since the online services are external, the source code and evaluation metrics of the online services cannot be obtained at this time.
[0005] Although the hybrid deployment of online and offline services in data centers improves the utilization rate of node resources, it also brings serious resource competition problems. Deploying online services and offline jobs on the same physical node at the same time will compete for limited hardware resources such as CPU, memory, and cache. When the resource competition is relatively serious, it will cause performance interference to the online services, resulting in a decline in service quality and an inability to meet the service level objective (SLO). Take the widely representative benchmark suites Tailbench and BigDataBench as examples. Mixing the online service Imgdnn in TailBench and the offline batch processing job Wordcount in BigDataBench for deployment, the tail latency of the online service in the hybrid deployment mode is 10.27 times that in the independent deployment mode. To optimize the performance interference of online services, it depends on a reasonable deployment plan, and the premise of designing a reasonable deployment plan is to be able to accurately quantify and evaluate the performance interference of online services.
[0006] Current performance interference analysis mainly relies on application-level metrics (such as the latency and throughput of the load). However, application-level methods usually have certain limitations. First, application-level metrics are often closely related to specific application scenarios and are difficult to be effectively promoted across different application scenarios. For example, tail latency is a key metric for the performance of online services. However, different applications have significantly different tolerances for latency. For example, the 95th percentile latency of the online service Imgdnn is between 2 ms and 3 ms, while the 95th percentile latency of Sphinx can reach about 2 seconds. Therefore, even if Sphinx with a relatively large tail latency performs poorly in some scenarios, it cannot be simply concluded that the performance of the Sphinx service is poor. Second, in the new co-location mode, users often cannot obtain the source code of online services and application-level evaluation metrics, and application-level metrics are difficult to comprehensively reflect the impact of underlying resource contention. Compared with application-level metrics, system-level metrics can more comprehensively capture the dynamic changes in the use of hardware resources and provide more accurate support for the quantification of performance interference. However, there are also some problems in existing system-level research. For example, the granularity of metric features is too coarse to capture the subtle changes in system behavior. System-level macroscopic metrics cannot accurately reflect the performance fluctuations of applications and cannot be mapped to the quantitative evaluation of degraded interference performance. In addition, most existing methods are limited to simple linear correlation analysis and fail to deeply explore the high-order dependence relationships between metrics, which may cause some important information to be missed when quantifying performance interference.
[0007] The present invention aims at the new co-location scenario of online and offline services in the data center. According to the characteristics of the fluctuations in online service requests, it analyzes the fluctuations in the use of system resources, and then quantitatively evaluates the performance interference of online services. Summary of the Invention
[0008] Based on the above problems, the present invention proposes a co-located online service performance interference quantification evaluation model based on micro-architecture-level metrics, which is used to quantitatively evaluate the performance interference of online services in the new co-location scenario of the data center. Based on the micro-architecture-level metrics collected during the operation of co-located loads, the entropy of all metrics is calculated according to the information entropy calculation formula. An online service performance interference quantification evaluation model is constructed by combining an autoencoder, a multi-head attention mechanism, and a multi-layer perceptron. Using the model constructed by this method, a system entropy metric can be finally obtained. This metric has a strong correlation with the tail latency of the online service application-level metric. According to the size of this metric, the performance interference of online services can be more accurately quantified and evaluated.
[0009] The co-located data center online service performance interference quantification evaluation method proposed by this method mainly consists of six steps: metric selection, metric collection, metric entropy calculation, metric dimensionality reduction, feature weighting, and system entropy calculation.
[0010] 1. Metric Selection
[0011] In the index selection step, starting from the composition of the processor microarchitecture, representative indicators related to resource competition are selected. Table 1 shows the component modules and functions of the processor microarchitecture. The microarchitecture of the processor consists of a front-end module and a back-end module. The back-end can be further divided into an execution domain and a memory domain. The selected microarchitecture-level indicators include instruction cache-related indicators and branch prediction hit rate of the front-end module, instructions per cycle (IPC), context switch per second (context_switch, cs), CPU core migration per second (cpu_migration, cm), cache hit rate, etc. of the back-end module, as well as indicators such as Page Fault and I / O of Off-Core resources. The specific selection of these microarchitecture-level indicators needs to be further determined according to the processor used. Finally, the selected indicators form a set of resource competition observation indicators, providing a data basis for subsequent steps.
[0012] Table 1 Component Modules and Functions of Processor Microarchitecture
[0013]
[0014] 2. Index Collection
[0015] Set 20 different load request intensities in ascending order of the online service QPS, and run them in a mixed mode with offline analysis jobs respectively. Use the Perf tool and the Proc file system provided by Linux for data collection. After the online service preheats and runs for 10 s, start collecting and analyzing performance indicators, which ensures that the online service load reaches a stable running state. During the index collection process, the offline batch processing job is always in the running process to ensure that the collected are the index data when the online service and the offline analysis job are running together. The index collection time interval is 1 s, which ensures that the index collection work does not occupy too much system resources.
[0016] 3. Index Entropy Calculation
[0017] To calculate the entropy value of the indicators under each online service load request intensity, first discretize the time series data. For the set of competing resource indicators M, m i represents the sample of indicator i on the time series. Divide the sample of each indicator i into 10 intervals according to size respectively, and calculate the probability distribution of the indicator value falling into each interval j. The probability p(j) of each interval can be calculated by formula (1):
[0018]
[0019] where count(m iThe number of times the index value of index i falls within interval j is denoted as ∈intervalj), where intervalj represents interval j and n represents the sample size.
[0020] Next, according to the definition of information entropy, the entropy H(m of index i under a certain load intensity i ) can be expressed by formula (2):
[0021]
[0022] By calculating the entropy values of each index, the index entropy set H of all indexes in the set of observed indexes during load operation can be obtained ik , where i represents index i and k represents different request intensities of the online service. These entropy values represent the uncertainty or fluctuation degree of the load on this index. A higher entropy value indicates that the index fluctuates greatly, indicating that the interference of the load on this resource may be stronger; on the contrary, a lower entropy value indicates that this index is relatively stable.
[0023] 4. Index Dimensionality Reduction
[0024] After calculating the index entropy, the indexes are grouped according to the Pearson correlation coefficient p between the index entropy and the online service load intensity (QPS), and then each group of indexes is compressed by an autoencoder. The specific method is divided into the following steps:
[0025] (1) Index grouping: First, calculate the Pearson correlation coefficient p between the index entropy H of each index under different load request intensities i and the online service load intensity (QPS) according to formula (3) i , where H i represents a set of index entropies calculated for index i under different load request intensities, QPS represents the set of load request intensities, cov(H i , QPS) represents the covariance of H i and QPS, and σ QPS respectively represent the standard deviations of H i and QPS, E[·] is used to calculate the mathematical expectation, and μ QPS respectively represent the means of H i and QPS.
[0026]
[0027] This p iThe value represents the correlation between the index entropy of index i and the load intensity. According to the magnitude of the correlation, the indexes are divided into three groups: the indexes with Pearson correlation coefficient greater than 0.4 are positively correlated indexes, the indexes with Pearson correlation coefficient less than -0.4 are negatively correlated indexes, and the indexes with Pearson coefficient between -0.4 and 0.4 are weakly correlated indexes.
[0028] (2) Autoencoder dimensionality reduction: First, construct the encoder part, using a two-layer fully connected neural network. The first layer of the encoder maps the dimension of the input data to the intermediate layer, and the second layer further compresses the output of the intermediate layer to the target dimension for output. For positively correlated and negatively correlated indexes, the output dimension of the encoder is set to 4; for weakly correlated indexes, the output dimension of the encoder is set to 2. Then, construct the decoder part symmetric to the encoder. Finally, after being processed by the autoencoder, the positively correlated and negatively correlated indexes output 4-dimensional feature representations, and the weakly correlated indexes output 2-dimensional feature representations. Through this dimensionality reduction process, data can be effectively compressed and key features can be retained.
[0029] 5. Feature weighting
[0030] Use the multi-head attention mechanism to assign weights to the output features of the autoencoder. The present invention uses the features output by the autoencoder for the generation of Query, Key, and Value in the multi-head attention mechanism. Take the feature data after dimensionality reduction by the autoencoder as the input. Apply three independent linear transformation operations to this input data to generate Query, Key, and Value respectively. Then adjust the data dimension according to the feature dimension and the number of attention heads, so that the processed Query (Q), Key (K), and Value (V) adapt to the calculation requirements of each attention head, and perform parallel calculations on these projections. In the present invention, the feature dimension is 10 and the number of attention heads is 5. Each attention head independently captures information in different dimensions, and then integrates this information through concatenation and projection to generate a richer feature representation. For each head, calculate the similarity score between Q and K according to formula (4), and use a scaling factor to adjust the score to stabilize the gradient:
[0031]
[0032] where d k is the dimension of Q and K, which is used to scale the size of the attention. The calculation for a single attention head can be represented by formula (5):
[0033]
[0034] where W i Q , W i K and W iV They are linear mapping matrices for Q, K, and V respectively. Finally, the outputs of all heads are concatenated and the final output is generated through a linear projection, as shown in Equation (6):
[0035] MHO = MultiHead(Q, K, V) = Concat(head 1 ,..., head h )W o (6)
[0036] where MHO represents the output feature, MultiHead(Q, K, V) represents the use of the multi-head attention mechanism, Q, K, and V represent the query vector, key vector, and value vector respectively, Concat(head 1 , …, head h ) represents concatenating the outputs of h attention heads, head i represents the output of the i-th attention head, h is the number of attention heads, and W o is the output mapping matrix.
[0037] 6. System Entropy Calculation
[0038] A multi-layer perceptron (MLP) is used to characterize the non-linear relationship between these features, and finally the output system entropy is calculated. The multi-layer perceptron is a fully connected neural network structure. The MLP model of the present invention includes an input layer, two hidden layers, and an output layer. The Dropout regularization method is adopted in the design to prevent overfitting, and at the same time, the ReLU activation function is used to enhance the non-linear expression ability of the model.
[0039] The input layer receives a feature vector of size 10 dimensions and is mapped to a 64-dimensional space through the first fully connected layer. Subsequently, non-linearity is introduced through the ReLU activation function, and a Dropout layer is added thereafter, with a dropout rate set to 0.5, aiming to randomly discard some neurons to avoid network overfitting. Then, the model enters the second hidden layer, which maps the 64-dimensional output to 32 dimensions, and the ReLU activation function is used to continue enhancing the non-linear expression ability of the model. To further improve the generalization ability, a Dropout layer with a dropout rate of 0.1 is used after the second hidden layer. Finally, the model maps the 32-dimensional output to 1 dimension through a fully connected layer as the final output result. To ensure that the output value is positive, the ReLU activation function is used after the output layer. Equation (7) gives the calculation process of finally calculating the system entropy SE through the MLP:
[0040] SE = MLP(MHO) = W n ·σ(W n-1 (...σ(W 1 ·MHO + b1 )...) + b n-1 ) + b n (7)
[0041] Among them, MHO is the output feature of the multi-head attention mechanism, W i and b i are the weight matrix and bias term of the i-th layer respectively, and σ is the activation function.
[0042] The present invention customizes the loss function of the model. Since the present invention is not a prediction task, the common mean square error is not used as the loss function. The present invention hopes that the calculated system entropy can have a strong correlation with the tail latency, an application-level metric of the online service. Therefore, 1 minus the Pearson correlation coefficient between the output entropy of the data and its tail latency is used as the loss function. To solve the problem of different dimensions of the tail latency of different online services and unify the performance comparison, the present invention proposes the tail mean ratio metric, that is, the ratio of the 95th percentile latency to the average latency. Finally, the tail mean ratio is used instead of the tail latency in the loss function, and the loss function is defined according to formula (7):
[0043]
[0044] where, e i and tar i are the sample entropy and the corresponding tail mean ratio calculated for each sample, and are the means of the sample entropy and the sample tail mean ratio respectively. To minimize the loss function, step 6 is repeated 1000 times during model training, and the iteration result with the minimum loss function value is selected as the final model, and the trained model is output. Given a set of metric data, after metric dimensionality reduction and feature weighting, the trained model can be directly used to calculate the system entropy. Description of the Drawings
[0045] Figure 1 is the cluster platform to which the online service performance interference quantification evaluation method for the hybrid deployment data center adheres.
[0046] Figure 2 is the architecture diagram of the present invention.
[0047] Figure 3 is the flow chart of the present invention.
[0048] Figure 4 is the flow chart for calculating the metric entropy.
[0049] Figure 5 is the flow chart for metric dimensionality reduction. Detailed Implementation Manner
[0050] The present invention will be described below in conjunction with the drawings and specific implementation manners.
[0051] The online service performance interference quantification and evaluation method proposed by the present invention is built on multiple connected servers. Figure 1 It is the deployment diagram of the platform built by this method. The platform consists of multiple computer servers (platform nodes), which are connected through a network to store data and execute tasks distributively. The platform nodes are divided into two categories: including one management node and multiple computing nodes. The platform built by the method of the present invention includes three core software modules: a data processing module, a load operation module, and a resource metric collection module. Among them, the load operation module runs online services and offline batch processing jobs, and the metric collection module is responsible for collecting micro-architecture level metrics during the load operation and transmitting the metric data to the data processing module of the management node. The data processing module is responsible for processing the metric values and calculating the system entropy.
[0052] The online service performance interference quantification and evaluation method proposed by the present invention is divided into two main stages: an offline training stage and an online analysis stage, which are used for model training and system entropy calculation respectively. Figure 2 It is the overall framework diagram of this method.
[0053] In a data center environment, offline analysis jobs are usually resident loads, while new online services arrive randomly, forming a mixed load with the offline jobs. In the offline stage, the loads in the TailBench benchmark suite are used as online services, and they are run together with the offline analysis jobs in the server to form a mixed load, and the performance metrics in the set of observed metrics under different load intensities are collected. These metrics cover aspects such as computing, memory access, and control, and are used to comprehensively reflect the resource competition situation. Then, the set of observed metrics is screened according to the relationships between the metrics and the relationships between the metrics and the load performance to form a training set for the system entropy calculation model. Then, a deep learning method is used to train the system entropy calculation model on this basis, so that the model can capture the non-linear relationship between the load contention characteristics and the system entropy.
[0054] In the online stage, when a new online service load enters the system, the system will first judge the load type and collect relevant observed metrics based on the judgment result. Subsequently, these real-time data will be input into the system entropy calculation model trained in the offline stage to calculate the system entropy value under the current mixed load environment. This entropy value provides a basis for quantifying the impact of resource competition on performance and helps to optimize resource allocation and improve system performance.
[0055] The following is combined with Figure 3The overall flowchart of the invention content illustrates the specific usage process of this method. The hardware platform used in the present invention is based on the Intel Xeon E5-2670 v processor, which has 20 cores and is suitable for multi-threaded parallel processing and high-performance computing. The experimental system has good parallel processing capabilities and storage support, and can meet the high-load experimental requirements. The data processing environment is Windows 10, and the programming language is selected as Python 3, which is used for tasks such as data preprocessing, metric entropy calculation, and model training, ensuring the flexibility and development efficiency of experimental data processing.
[0056] The experiments of the present invention use the online services and BigDataBench offline analysis jobs in the hybrid workload benchmark test suite Tailbench as benchmark test programs. This benchmark test suite covers the fields of artificial intelligence, big data, interactive databases, and HPC applications. The present invention selects WordCount as the offline analysis job, Img-dnn and Sphinx as the compute-intensive online services, Masstree and Silo as the memory-intensive online services. In addition, the present invention selects a memory-intensive service VoltDB outside the Tailbench test suite to verify the generalization of the method of the present invention.
[0057] The experiments of the present invention are divided into an offline training stage and an online verification stage. The online verification stage uses a new load combination different from that in the offline training stage for verification. The offline analysis job load and the online service load are mixed and deployed on a single machine node for experiments. 16 cores are allocated for the offline analysis job, 4 cores are allocated for the online service, and it is ensured that one load thread runs on each core. The Perf tool and the Proc file system provided by Linux are used for data collection. After the online service pre-warms and runs for 10 s, the collection and analysis of performance metrics start, which ensures that the online service load reaches a stable running state. During the metric collection process, the offline analysis job Wordcount is always running, ensuring that the collected metric data is for the co-running of the online service and the offline analysis job. The metric collection time interval is 1 s, which ensures that the metric collection work does not occupy too much system resources.
[0058] The present invention selects two types of online services, memory-intensive and compute-intensive, for experiments respectively to illustrate the reliability of the method of the present invention. In the memory-intensive load combination, the present invention uses Masstree-Wordcount as the training load combination, and Silo-Wordcount and VoltDB-Wordcount as the verification load combinations; in the compute-intensive load combination, the present invention uses Imgdnn-Wordcount as the training load combination, and Sphinx-Wordcount as the verification load combination.
[0059] 1. Index Selection
[0060] First, micro-architecture level indicators are selected as the overall indicator set for the mixed deployment scenario of online and offline services in the data center. According to the composition structure of the micro-architecture, the present invention selects a total of 48 indicators at the micro-architecture layer as the observation indicator set, and these indicators cover all components of the processor micro-architecture. Table 2 shows all the micro-architecture level indicators selected by the present invention and their description information.
[0061] Table 2 Micro-architecture Level Indicator Set
[0062]
[0063]
[0064] 2. Indicator Collection
[0065] The Perf tool and the Proc file system provided by Linux are used for data collection. After the online service preheats and runs for 10 s, the performance indicators are collected and analyzed. The collection time interval is 1 s, and the total collection duration is about 70 s.
[0066] 3. Indicator Entropy Calculation
[0067] After collecting all the indicators, first calculate the entropy value of all the indicators, that is, the indicator entropy. The calculation steps of the indicator entropy are as Figure 4 shown.
[0068] 3.1) Indicator Discretization.
[0069] 3.1.1) Use the indicator data, request intensity, and tail latency of the online service collected under different load intensities as the input.
[0070] 3.1.2) Discretize the indicator data under a certain load intensity. Divide a certain indicator into 10 intervals according to its size.
[0071] 3.2) Calculate the probability distribution of each interval according to the reduction of the indicator values included in each indicator interval.
[0072] 3.3) Calculate the indicator entropy under this load intensity based on the probability distribution of the indicator values falling in each interval.
[0073] For the data of different indicators under different load intensities, repeat the process from 3.1.2) to 3.3) to obtain the indicator entropy set of all indicators under different load intensities.
[0074] 4. Indicator Dimensionality Reduction
[0075] The specific steps are as Figure 5 shown.
[0076] 4.1) Calculate the Pearson correlation coefficient between the metric entropy of different metrics under different request intensities and the online service load intensity QPS according to the metric entropy set.
[0077] 4.2) Divide the metrics into three groups according to their correlation with the request intensity. If the correlation coefficient is greater than 0.4, it is a positively correlated metric; if the correlation coefficient is less than -0.4, it is a negatively correlated metric; if the correlation coefficient is between -0.4 and 0.4, it is a weakly correlated metric.
[0078] 4.3) Perform dimensionality reduction on the metrics through an autoencoder, and adjust the output feature dimensions of the positively correlated metrics and negatively correlated metrics to 4, and the output feature dimension of the weakly correlated metrics to 2.
[0079] 5. Feature Weighting
[0080] Use the output features of the autoencoder as the input of the multi-head attention mechanism, and use the multi-head attention mechanism to assign different weights to the features. The output features of this step will be used as the input for the final MLP to calculate the system entropy.
[0081] 6. System Entropy Calculation
[0082] Finally, calculate the system entropy through a multi-layer perceptron. The calculated system entropy should have a strong correlation with the application-level metrics of the online service.
[0083] Table 3 QPS Ranges of Different Online Services
[0084]
[0085] Table 3 shows the load intensity ranges of different online services in the experiments of the present invention. In the offline training phase, according to the maximum QPS that different loads can support, the load intensity is evenly divided into 20 groups from small to large, and the performance metric data under the QPS of each group is collected respectively. And the experiment is repeated 15 times under the QPS of each group, reducing the contingency of the experimental results. All the collected data is used as the initial data of the experiment. In the online verification phase, the metric data under 20 different QPS are collected respectively, and the data is collected 5 times for each group. The average value of the 5 metric data is taken as the metric data under this QPS.
[0086] For the memory-intensive online service load combination, the present invention uses the mixture of Masstree and Wordcount in the TailBench benchmark as the training load combination, and uses Silo in the TailBench benchmark and VoltDB not belonging to the TailBench benchmark to form test load combinations with Wordcount respectively. Use the Pearson correlation coefficient between the calculated system entropy and the online service as the evaluation criterion for quantifying the performance interference of the system entropy.
[0087] The experimental results of the present invention prove that whether it is the workload in the TailBench benchmark or other workloads, there is a strong correlation between the system entropy calculated by the method of the present invention and the tail average ratio. When using the Silo-Wordcount workload combination as the verification, the Pearson correlation coefficient between its system entropy and the tail average ratio is 0.87. When using VoltDB-Wordcount as the verification, the Pearson correlation coefficient between its system entropy and the tail average ratio is 0.78. This shows that the method of the present invention has good generalization.
[0088] For the compute-intensive online service workload combination, similar to the memory-intensive experiment, the present invention uses the mixture of Imgdnn and Wordcount in the TailBench benchmark as the training workload combination, and uses the mixture of Sphinx and Wordcount as the verification workload combination. The Pearson correlation coefficient between the system entropy and the tail average ratio of the compute-intensive type is 0.85, showing a strong correlation. This also proves the rationality of the method of the present invention in the workload combinations composed of different types of online services.
[0089] Finally, it should be noted that the above examples are only used to illustrate the present invention and do not limit the technology described in the present invention. All technical solutions and their improvements that do not depart from the spirit and scope of the invention should be covered by the scope of the claims of the present invention.
Claims
1. A method for quantitatively evaluating the performance interference of online services in a co-location data center, characterized in that It consists of six steps: indicator selection, indicator collection, indicator entropy calculation, indicator dimension reduction, feature weighting, and system entropy calculation; 1) Indicator selection In the indicator selection step, representative indicators related to resource competition are selected based on the composition of the processor microarchitecture; the processor microarchitecture composition includes the front-end module and the back-end module, and the back-end can be divided into the execution domain and the memory domain; the selected microarchitecture-level indicators include the instruction cache-related indicators and branch prediction hit rate of the front-end module, the number of instructions per cycle, the number of context switches per second, the number of CPU core migrations per second, and the cache hit rate of the back-end module; According to the online service QPS, set multiple groups of different load request intensities from small to large, and run them together with the offline analysis job. Use the Perf tool and the Proc file system provided by Linux to collect data. After the online service is preheated for 10 seconds, start collecting and analyzing performance indicators. During the indicator collection process, the offline batch processing job is always in the process of running to ensure that the indicator data collected are the indicator data when the online service and the offline analysis job are running together. The indicator collection interval is 1 second. 2) Index entropy calculation Discretize the time series data; for the competitive resource indicator set M, m i Represents the sample of indicator i in the time series. Each sample of indicator i is divided into 10 intervals according to its size, and the probability distribution of the indicator value falling in each interval j is calculated; the probability p(j) of each interval can be calculated by formula (1): Among them, count(m i ∈interval j) represents the number of indicator values of indicator i falling within interval j, intervalj represents interval j, and n represents the number of samples; Next, according to the definition of information entropy, the entropy H(m i ) can be expressed by formula (2): By calculating the entropy value of each indicator, we can obtain the indicator entropy set H of all indicators in the observed indicator set during load operation. ik ,i represents the index i, and k represents the different request intensities of online services; 3) Dimensionality reduction of indicators After the indicator entropy is calculated, the indicators are grouped according to the Pearson correlation coefficient p between the indicator entropy and the online service load intensity (QPS), and then each group of indicators is compressed through the autoencoder; the specific method is divided into the following steps: (1) Index grouping: First, the index entropy H of each index under different load request intensities is calculated according to formula (3): i The Pearson correlation coefficient between the online service load intensity (QPS) and i , where H i represents a set of index entropies calculated under different load request intensities for index i, QPS represents the load request intensity set, cov(H i ,QPS) represents H i Covariance with QPS, and σ QPS They represent H i The standard deviation of QPS, E[·] is used to calculate the mathematical expectation, and μ QPS They represent H i and the average value of QPS; The p i The value represents the correlation between the indicator entropy and the load intensity of indicator i. According to the size of the correlation, the indicators are divided into three groups: the indicators with a Pearson correlation coefficient greater than 0.4 are positively correlated indicators, the indicators with a Pearson correlation coefficient less than -0.4 are negatively correlated indicators, and the indicators with a Pearson coefficient between -0.4 and 0.4 are weakly correlated indicators; (2) Autoencoder dimensionality reduction: First, construct the encoder part, using a two-layer fully connected neural network; the first layer of the encoder maps the dimension of the input data to the middle layer, and the second layer further compresses the output of the middle layer to the target dimension output; for positive and negative correlation indicators, the dimension of the encoder output is set to 4; for weak correlation indicators, the dimension of the encoder output is set to 2; then, construct a decoder part that is symmetrical to the encoder; finally, after processing by the autoencoder, the positive and negative correlation indicators output 4-dimensional feature representations, and the weak correlation indicators output 2-dimensional feature representations; 4) Feature weighting Using a multi-head attention mechanism, weights are assigned to the output features of the autoencoder; the features output by the autoencoder are used to generate queries, keys, and values in the multi-head attention mechanism; the feature data after dimensionality reduction by the autoencoder is used as input; three independent linear transformation operations are applied to the input data to generate queries, keys, and values respectively; then the data dimension is adjusted according to the feature dimension and the number of attention heads, so that the processed queries (Q), keys (K), and values (V) meet the computational requirements of each attention head, and these projections are calculated in parallel, with a feature dimension of 10 and a number of attention heads of 5; each attention head independently captures information of different dimensions, and then integrates this information together through splicing and projection; for each head, the similarity score of Q and K is calculated according to formula (4), and the score is adjusted using a scaling factor to stabilize the gradient: Among them, d k is the dimension of Q, K, which is used to scale the size of attention. The calculation of a single attention head can be expressed by formula (5): Among them, W i Q , W i K and W i V are the linear mapping matrices of Q, K and V respectively; finally, the outputs of all heads are concatenated and the final output is generated by linear projection, as shown in formula (6): MHO=MultiHead(Q,K,V)=Concat(head1,...,head h )W o (6) Among them, MHO represents the output feature, MultiHead(Q,K,V) represents the use of multi-head attention mechanism, Q, K and V represent the query vector, key vector and value vector respectively, Concat(head1,…,head h ) means concatenating the outputs of h attention heads. i represents the output of the i-th attention head, h is the number of attention heads, and W o is the output mapping matrix; 5) System entropy calculation A multi-layer perceptron (MLP) is used to characterize the nonlinear relationship between these features, and finally the output system entropy is calculated; the MLP model consists of an input layer, two hidden layers, and an output layer. The Dropout regularization method is used in the design to prevent overfitting, and the ReLU activation function is used to enhance the nonlinear expression ability of the model; The input layer receives a 10-dimensional feature vector, which is mapped to a 64-dimensional space through the first fully connected layer. Subsequently, nonlinearity is introduced through the ReLU activation function, and a Dropout layer is added thereafter with a loss rate set to 0.
5. Next, the model enters the second hidden layer, which maps the 64-dimensional output to 32 dimensions, and uses the ReLU activation function to further enhance the nonlinear expression ability of the model. A Dropout layer with a loss rate of 0.1 is used after the second hidden layer. Finally, the model maps the 32-dimensional output to 1 dimension through a fully connected layer as the final output result. The ReLU activation function is used after the output layer. Formula (7) gives the calculation process of the system entropy SE calculated by MLP: SE=MLP(MHO)=W n ·σ(W n-1 (...σ(W1·MHO+b1)...)+b n-1 )+b n (7) Among them, MHO is the output feature of the multi-head attention mechanism, W i and b i are the weight matrix and bias term of the i-th layer, and σ is the activation function; The Pearson correlation coefficient between the output entropy of the data minus 1 and its tail delay is used as the loss function. The tail-mean ratio indicator is proposed, which is the ratio of the 95th percentile delay to the average delay. Finally, the tail-mean ratio is used instead of the tail delay in the loss function. The loss function is defined according to formula (7): Among them, e i and tar i The sample entropy calculated for each sample and the corresponding tail mean ratio, and are the mean of sample entropy and sample tail mean ratio respectively; in order to minimize the loss function, repeat step 6 when training the model, iterate 1000 times, select the iteration result with the smallest loss function value as the final model, and output the trained model; given a set of indicator data, after indicator dimensionality reduction and feature weighting, use the trained model to calculate the system entropy.