Computer performance volatility prediction
Through machine learning models, the underlying performance indicators of hardware and software are utilized to establish an empirical distribution, which solves the problem of difficult prediction of computer system performance volatility and achieves more accurate volatility prediction and performance optimization.
Patent Information
- Application Number
- CN202410665604.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-05-27
- Publication Date
- 2025-05-23
AI Technical Summary
Modern computer systems show high uncertainty due to configuration differences between hardware, software and operating systems and multi-process interference, which makes performance and power consumption volatility difficult to predict and explain.
Using machine learning models, by collecting underlying performance indicators of hardware and software as indicator identifiers, an empirical distribution is established to model and predict the performance volatility of computer systems.
More accurate predictions of computer system performance volatility is achieved, helping to optimize performance, reduce volatility, and providing automatic recommendations to improve system resource management efficiency.
Smart Images

Figure CN120029864A_ABST
Abstract
Description
Background Art
[0001] Modern computer systems are highly nondeterministic, arising from various configuration differences between hardware, middleware, operating systems, and interference from multiple processes that can be executed simultaneously. This uncertainty can make it difficult to predict and explain the observed volatility in performance and power consumption. Some techniques attempt to predict a single point of performance for a given application under known conditions; however, the volatility of performance cannot be adequately described by a single point identifier of performance. Instead, the volatility of performance can be better represented by a distribution of identifiers of performance. Performance can be considered to include a degree of randomness, similar to a random variable.
[0002] Viewing the performance of computer systems in general as random variables can lead to new methods and models for describing the relationship between hardware and software, for example. One of the goals of better understanding performance volatility is the idea of "explainable performance," or the ability to decompose and better understand in an actionable way the various factors that influence the complex performance behavior observed during operation, as manifested in the volatility benchmarks associated with the execution of an application. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description in conjunction with the accompanying drawings, in which:
[0004] Figure 1 is a description of a system for computer volatility forecasting, in one embodiment.
[0005] Figure 2 is a description of a machine learning (ML) model in one embodiment;
[0006] Figure 3 is a flow chart describing a method of training an ML model in one embodiment;
[0007] Figure 4 is a flow chart describing a method of using an ML model in one embodiment;
[0008] Figure 5 is a description of a computing node in one embodiment;
[0009] Figure 6 is a description of a high performance computing (HPC) cluster in one embodiment;
[0010] Figure 7 is a flow chart describing a method for predicting computer performance volatility in one embodiment;
[0011] Figure 8 is a flow chart describing a method for predicting computer performance volatility in one embodiment;
[0012] Fig. 9 is a flowchart depicting a method for predicting computer performance volatility in one embodiment; and
[0013] Fig.10 is a flowchart depicting a method for identifying computer performance volatility in one embodiment. DETAILED DESCRIPTION
[0014] In the following description, details will be set forth by way of example in order to facilitate discussion of the disclosed subject matter. However, it will be apparent to those of ordinary skill in the art that the disclosed embodiments are exemplary and not an exhaustive list of all possible embodiments.
[0015] Throughout this disclosure, hyphenated forms of reference numerals refer to particular instances of elements, and non-hyphenated forms of reference numerals refer to generic or collective references to elements. Thus, by way of example (not shown in the figures), device "12-1" refers to an instance of a class of devices that can be collectively referred to as device "12", and any one of them can be collectively referred to as device "12". In the drawings and the description, like reference numerals are intended to denote like elements.
[0016] In the early days of digital computing, computer systems were relatively simple, and thus performance modeling and evaluation of computer systems were also relatively simple, such as in the realm of 8-bit processor architectures running assembly code. In such early computer systems, in one example of a volatility benchmark, estimating the run time of a loop involved only multiplying the number of assembly instructions in the loop by the number of iterations and dividing the product by the clock speed.
[0017] The performance of modern computer technology has become very advanced, more complex, and more uncertain. There are many factors contributing to computer system performance volatility, such as heterogeneous accelerators, multi-level networks, parallel and concurrent architectures, operating system (OS) heuristics, hierarchical software abstractions, and potential interference from various simultaneously executing processes from modular multi-tenant systems such as cloud computing systems.
[0018] Some root causes of performance volatility in modern computer systems have been postulated or identified. In some scenarios, when simultaneously executing processes share computer resources such as local hardware resources or network resources, contention and interference may occur between these shared resources, resulting in performance volatility. For example, certain background OS processes may cause contention and interference, which may affect performance volatility and may be observed in volatility benchmarks. In another example, functions on a computer system for energy management may cause performance volatility. In some cases, memory management that may reside in software or hardware, such as so-called "garbage collection" program routines, may cause performance volatility. In another example, certain global resources (such as network switches) may cause contention, such as contention that may be caused by network traffic between other sources, thereby causing performance volatility. Other sources of performance volatility may include processor differences in multi-processor systems and cache limitations with certain processors, etc. When a computer system is undergoing maintenance activities, the performance of certain resources may be constrained and cause performance volatility. For example, when data or task processing needs to be queued, delays that cause performance volatility may be observed. In some cases, available power may be constrained under certain conditions or at certain points in time, resulting in performance variability.
[0019] Difficulties may exist in explaining and predicting performance volatility because it may not be easy to limit the performance volatility of a computer system to its range or to model it with high accuracy. However, for various reasons, the expected impact of accurate performance volatility estimates is enormous. The ability to accurately estimate performance volatility can be beneficial in terms of resource efficiency of computer systems, such as for optimal job and process scheduling and power management. The ability to accurately estimate performance volatility can affect the feasibility of certain purchasing decisions for computer systems, for example, by accurately predicting that a given application may have better performance (as observed in a volatility benchmark), lower price-performance ratio, lower performance volatility, lower tail latency, or a combination thereof when executing the given application on one type of computer system relative to executing the given application on another type of computer system. In the development process of computer systems and applications, having accurate performance volatility estimates reflected in volatility benchmarks can play a role in the design, verification, and regression testing of high-performance software and hardware components, including in the work of optimizing both software and hardware components.
[0020] Certain embodiments of the present disclosure provide a method for modeling and predicting volatility in the performance of a computer system, as observed in a volatility benchmark. The modeling method can identify and describe the volatility of certain explanatory variables, referred to as "telltale indicators," and the telltale indicators themselves, which putatively explain the empirical distribution of the observed volatility benchmark. The telltale indicators can represent various types of computer system performance indicators, as will be described in further detail, and can include telltale indicators collected from hardware using hardware performance counters, and telltale indicators collected from the OS using externally accessible metrics included with the OS.
[0021] In some embodiments, the method and system for computer performance volatility prediction can use ubiquitous machine learning (ML) models to predict the distribution of volatility benchmarks, rather than using the actual value of the volatility benchmark in any single instance. Certain embodiments may include so-called "black box" predictions, such as those derived from ML models composed of trained neural networks. Certain embodiments may use statistical models and knowledge to derive insights and defined explanations of the relationship between hardware and software. Certain embodiments may include a computing mechanism to automatically collect underlying performance metrics closely related to hardware as indicator identifiers, as well as the empirical distribution of such indicator identifiers, to model the distribution of user-level or high-level computer system performance. The underlying indicator identifiers can be collected from hardware and OS, and can be analyzed using ML models and statistical models. ML models can be used to predict the shape and properties of high-level volatility benchmark distributions. Statistical models can be used to link specific aspects of high-level volatility benchmark distributions with specific behaviors of OS and hardware to better understand the relationship between hardware and software in operation. Certain embodiments include statistically analyzing these relationships to provide insights on software, architectural features, and performance bottlenecks, and can be configured to provide automatic suggestions, such as changes in configuration parameters, for potential optimization of performance and reduction of performance volatility.
[0022] Certain embodiments may include an input that is a value of an indicator identifier and a measure of volatility, such as a computer hardware counter. The measurement may be in aggregate form or a time series to help link performance volatility to different application stages or portions of the actual code of the application. Certain embodiments may use automated ML techniques to estimate a volatility benchmark as a response variable based on the indicator identifier as an explanatory variable. Certain embodiments may use automated statistical techniques to model the distribution of the volatility benchmark and extract potentially profound relationships between the volatility benchmark and the indicator identifier. Certain embodiments may provide automatic recommendations to a user on how to optimize a given application to reduce performance volatility and tail latency. Certain embodiments may interact with a system's resource manager (e.g., an OS scheduler) to optimize resource allocation for reduced performance volatility and tail latency.
[0023] Certain embodiments may be used to simultaneously measure performance volatility (e.g., a response variable or volatility benchmark) and the volatility of an identified indicator identifier, such as a processor metric, an OS metric (e.g., an explanatory variable or an indicator identifier). Certain embodiments may be used to analytically link explanatory variables and response variables together to build a predictive model of application performance volatility. Certain embodiments may be used to formulate relationships between different modes in a performance volatility distribution and different explanatory variables to provide insights and quantitative explanations of the sensitivity of an application to different system factors. Certain embodiments may be used to more accurately predict performance and performance volatility on unseen systems and verify accuracy using robust contextual confidence intervals. Certain embodiments may be used to link performance volatility to application stages to assist in debugging performance volatility. Certain embodiments may be used to provide automatic recommendations for optimizing performance volatility and tail latency, and / or improving deadline compliance in real-time applications.
[0024] The disclosed method and system for predicting computer performance volatility can provide various benefits associated with eliminating or reducing performance volatility. For example, the quality of service (QoS) standards of computing services, such as those specified in a service level agreement (SLA) or other contract for computing services, can be optimized, and the quality of service (QoS) standards of computing services can be made less sensitive to performance volatility, which is desirable for both suppliers and buyers of such outsourced computing services. In multi-node application execution such as batch synchronous processing (BSP), the idle time waiting for nodes to complete is reduced by allowing nodes to synchronize their respective execution cycles and by reducing performance volatility (as shown in volatility benchmarks and their distributions) so that nodes complete execution within a shorter time window, thereby improving overall performance. For example, when volatility benchmarks are more accurately identified, mismatched hardware configurations can be identified, and optimal hardware configuration recommendations can be made, such as cloud and software as a service solutions, which can provide broad market access to supercomputing at different price points. In another example, when the close relationship between node sensors and application execution is determined by analyzing volatility benchmarks, scheduler data processing can be performed more energy-aware and more simply.
[0025] Certain embodiments will now be described with reference to the accompanying drawings.
[0026] Reference now Figure 1 , a volatility prediction engine (VPE) 100 is depicted as a schematic block diagram. As described herein, VPE 100 represents data and functional elements that can be implemented by a computer. As shown, VPE 100 includes a machine learning (ML) system 110, which can train and derive an ML model 120, which can be used to generate a volatility output 130, as will be described in further detail. Volatility output 130 may include output data from VPE 100, which describes, predicts, identifies and diagnoses the distribution of performance volatility in a computer system executing an application. It is noted that the present disclosure is not primarily concerned with predicting the variance of computer performance itself in a specific context, but rather, VPE 100 is capable of predicting the behavior of a computer system executing an application relative to the performance volatility in the form of a distribution of performance volatility.
[0027] As described above, modern hardware, operating systems (OS), and software applications (referred to herein simply as "applications") are typically non-deterministic due to design or implementation factors, such that their various empirical characteristics can be considered random variables. Therefore, the measured or empirical performance distribution of an application is an accumulation or combination of such random variables. This combination relationship can be additive (normal distribution), multiplicative (lognormal distribution), compound (multimodal distribution), or various combinations thereof. This combination relationship may sometimes be too complex or too variable to predict a single volatility benchmark, such as the running time of a given application in a given execution context. However, it is generally possible to predict the range or distribution of performance volatility (e.g., the range or distribution of volatility benchmarks) and the dependence of such distribution on the indicator identifiers of the underlying hardware and software performance. The observed volatility benchmarks and the empirical distribution of the indicator identifiers, combined with sensitivity testing, can be used to link the behavioral patterns of the indicator identifiers to the specific true distribution of the volatility benchmarks. This method, as presented in the embodiments disclosed herein, can be used to identify and analyze the role of a given computer configuration in the performance volatility associated with executing an application.
[0028] As described herein, a volatility benchmark associated with executing an application may be selected from at least one of: runtime, response latency, response latency probability, data throughput rate, interval time, end timestamp, start timestamp, or data throughput capability.
[0029] As described herein, the true distribution is characterized by at least one of the following:
[0030] Statistical variables, including mean, median, mode position, mode number, and mode size;
[0031] Distribution variables, including standard deviation, standard error, and variance;
[0032] Confidence intervals;
[0033] High density areas; or
[0034] • At least one parameter for a curve fit, such as for a normal (Gaussian) curve fit, a bimodal curve fit, a multimodal curve fit, or a log-modal curve fit.
[0035] As described herein, the configuration of a computer system may specify at least one of the following:
[0036] Central processing unit (CPU) parameters, including base clock frequency, cache size, number of cores, number of logical processors, and peripheral bus clock speed;
[0037] Graphics processing unit (GPU) parameters, including GPU version, GPU clock speed, GPU cache memory;
[0038] Memory parameters, including physical memory size, number of memory cards, memory card size, memory interface, nominal memory write speed, nominal memory read speed, nominal memory latency, number of channels, memory clock speed, memory access control, memory allocation control;
[0039] OS parameters, including version number, update count, update list, registry content, directory content;
[0040] Network parameters, including network capacity, number of physical ports, type of physical ports, media type, power management mode; or
[0041] Local storage parameters, including the number of physical volumes, the number of logical volumes, the size of the volumes, the capacity / volume, the file system identifier, the file system version, the storage media type, redundant volumes, the file system write speed, and the file system read speed.
[0042] As described herein, the indication identifier may include at least one of the following:
[0043] Central processing unit (CPU) metrics, including CPU utilization percentage, actual clock frequency, number of processes, number of threads, number of handles, cache events, CPU events, cycle counts, instruction counts, IO events, operating system events, kernel identifiers;
[0044] Graphics processing unit (GPU) metrics, including GPU utilization percentage, GPU memory usage, and GPU shared memory;
[0045] Memory metrics, including memory usage, available memory, committed memory, cached memory, paged memory, non-paged memory;
[0046] OS metrics, including application runtime, code segment runtime, virtual memory size, application CPU time, application end timestamp, code segment end timestamp;
[0047] Network metrics, including throughput rate, sent data rate, received data rate, network capacity rate, packet error rate, application network data usage, application network data rate; or
[0048] Local storage metrics, including response time, average response time, percentage of active time, transfer rate, file system latency, write speed, read speed, and access time.
[0049] like Figure 1 As shown, in VPE 100, ML system 110 may receive training data 102 to train ML model 120 for a particular implementation, such as for a particular application executing on a given target computer platform (see also Figure 5 and 6 ). In some embodiments, the training data 102 may be collected directly from an application executing concurrently with the ML system 110, such as on a large number of computer systems. For example, the same application may be executed on multiple computer systems to generate a baseline volatility benchmark distribution that may represent or approximate a "true distribution" of a volatility benchmark. In various cases, the ML system 110 may use the training data 102 to train the ML model 120 until a certain condition is met, such as a confidence interval for the output of the ML model 120 is within a certain range. In addition to the training data 102, the ML system 110 may also have access to validation data 104, which may represent reference data that is known or expected to produce a desired result for comparison with the training data 102. In this way, the validation data 104 may be used to validate the performance volatility of the ML model 120 trained using the training data 102, such as within a certain confidence level.
[0050] In the VPE 100, after the ML model 120 has been sufficiently or acceptably trained using the ML system 110, the ML model 120 may be extracted for operation or use as a prediction engine for performance volatility. In operation or use as a prediction engine for performance volatility, the ML model 120 may receive input data 106, which may be any new data for performance volatility analysis, and the ML model 120 may accordingly generate a volatility output 130 as a result output. The volatility output 130 may include information that identifies or isolates a causal relationship between the performance volatility of an application and an "indicator identifier," wherein the performance volatility of the application is measured using a "volatility benchmark" associated with the execution of the application in a given context, the "indicator identifier" including computer system software and hardware metrics that may exhibit a causal relationship with the volatility benchmark. The volatility output 130 may also include information on potential causes of observed volatility behavior of the volatility benchmark inferred or suggested based on the causal relationship, and suggestions on how the observed volatility in the form of an "empirical distribution" may be optimized or may be able to reach or approach a desired true distribution. Thus, the volatility output 130 may also include information describing a prediction of a true distribution of a volatility benchmark for an application in a given execution environment based on the input data 106, which includes an empirical distribution of volatility benchmarks and indicative identifiers for a relatively small number or size of executions of the application, such as based on a relatively small sampling of executions of the application, e.g., a small number of runs of the application. The volatility output 130 may also include information describing a prediction of a true distribution of a volatility benchmark for an application based on the input data 106, which includes an empirical distribution of volatility benchmarks and indicative identifiers from different execution environments, such as different configurations of computer systems executing the application other than those used for the training data 102. The volatility output 130 may also include information describing a statistically derived relationship between the indicative identifiers and the volatility benchmarks that the ML model 120 is able to automatically generate.
[0051] In the operation of VPE 100, the following five use cases are disclosed for illustration purposes. It should be noted that other use cases or combinations of use cases may also be implemented using VPE 100.
[0052] Use Case 1: Pre-Execution Prediction. An application is executed multiple times, such as a large number of execution instances, on a given computer system having a first configuration. In the multiple executions, training data 102 including selected volatility benchmarks is recorded along with some or all available indicator identifiers. An ML model 120 is trained using the volatility benchmarks and indicator identifiers recorded as training data 102. In some embodiments, the training may be a first training of the ML model 120. In other embodiments, the training may be augmented to a previous training of the ML model 120, such as by using different training data 102. The ML model 120 is then extracted after sufficient or desired training based on the first configuration. The ML model 120 is used with input data 106, wherein the input data 106 includes the volatility benchmarks and indicator identifiers recorded during execution of the application using a second configuration, the second configuration being different from the first configuration of the computer system used for training. A volatility output 130 is generated by the ML model 120 as output data, which includes a prediction of a true distribution of at least one of the volatility benchmarks for the application using the second configuration. For each different subsequent configuration of the computer system, the ML model 120 may be reused with different input data 106 to generate corresponding volatility outputs 130 .
[0053] In another example, different applications may be executed on a given computer system having a first configuration. The ML model 120 may be used to predict the true distribution of certain volatility benchmarks for the different applications. In some cases, the ML model 120 may provide such predictions without requiring execution of a large number of different applications on the first configuration, such as by using as input data 106 an indicator identifier and a single run of volatility benchmarks for the different applications.
[0054] Use Case 2: On-the-fly Predictions. As in Use Case 1, the ML model 120 is trained for an application executing on a first configuration of a computer system. Then, execution of the application is initiated on the first configuration. During execution of the application, before execution is completed, input data 106 including volatility benchmarks and indicative identifiers are recorded and used by the ML model 120 to generate certain characteristics of the true distribution of at least one of the volatility benchmarks for successive executions of the application on the first configuration as volatility outputs 130. For example, the characteristics may include predictions about the location of certain modes in the true distribution.
[0055] Use Case 3: Combined Pre-Execution / In-Execution Prediction. As in Use Case 1, the ML model 120 is trained for an application executing on a first configuration of a computer system. Volatility outputs 130 are generated as in Use Case 1. The volatility outputs 130, which include predicted true distributions of volatility benchmarks, are used to schedule application workloads. Then, execution of the application is initiated on the first configuration, and additional volatility outputs 130 are generated using new input data 106 generated as the application executes. During execution of the application, before execution is completed, the new input data 106, which includes the volatility benchmark and an indicator identifier, is recorded and used by the ML model 120 to generate new volatility outputs 130. The new volatility outputs 130 are used to modify the true distribution of at least one of the volatility benchmarks generated in Use Case 1.
[0056] Use Case 4: Run Time Analysis. As in Use Case 1 or Use Case 3, an ML model 120 is trained for an application executed on a first configuration of a computer system. A volatility benchmark includes runtimes of the application, for which a true distribution is obtained as a prediction from a volatility output 130. For an application having a true distribution that includes long-tail runtimes, first input data 106 of multiple executions of the application is collected or obtained. Using a statistical model on the first input data 106, indicator identifiers that contribute to the long-tail runtimes are correlated. The indicator identifiers with the strongest causal relationships are then correlated with larger long-tail values of runtime. It is then recommended to modify certain aspects of the first configuration, such as certain hardware or software settings or parameters, to reduce the long-tail runtimes. The best improved true distribution of runtimes is used in conjunction with the modified indicator identifiers to predict an upper bound on the long-tail runtime of the application.
[0057] Use Case 5: Indicator Analysis. As in Use Case 1 or Use Case 3, train an ML model 120 for an application executing on a first configuration of a computer system. Use the ML model 120 to predict true distributions of different volatility benchmarks. Based on the true distribution of a given volatility benchmark, use a statistical model on the first input data 106 to identify indicator identifiers that contribute to an observed mode of the true distribution of the given volatility benchmark. Generate multiple associations between the modes of the given volatility benchmark and different indicator identifiers. Propose modifications to certain aspects of the first configuration, such as modifications to hardware or software settings or parameters, to reduce the volatility of the given volatility benchmark, or determine upper / lower bounds for the given volatility benchmark. Repeat the process for different or related volatility benchmarks to define bounds on the performance volatility of the application.
[0058] In addition to the specific use cases 1-5 described above, the VPE 100 may be used for a variety of additional functions. In one embodiment, the VPE 100 may be used to predict statistical values associated with various volatility benchmarks, such as standard deviation, kurtosis, etc., such as by using regression techniques. In one embodiment, the VPE 100 may be used to classify applications based on volatility / sensitivity categories, such as network sensitive (e.g., a volatility benchmark distribution associated with network synchronization events), CPU-clock sensitive (a volatility benchmark distribution associated with volatility in clock speed attributed to power management), OS sensitive (a volatility benchmark distribution associated with operating system scheduling decisions), insensitive, etc. In one embodiment, the VPE 100 may be used to predict a volatility benchmark distribution (e.g., an "exponential distribution" with λ=0.5, a "lognormal distribution with μ=0.13, σ=1.1"), such as by using a Bayesian posterior distribution from a priori measurements. In one embodiment, the application of interest is not isolated from other applications executing concurrently on the computer system. In this case, by specifically correlating indicator identifiers in the time domain with features such as the originating process, timestamps, and OS context switch events, the VPE 100 can measure and quantify the volatility baseline caused by interference from other applications. In one embodiment, when the volatility output 130 indicates a multi-modal volatility baseline distribution, the ML model 120 can automatically identify which indicator identifier contributes to which mode. Such predictions can be accomplished using linear regression models, decision trees, support vector decomposition, neural networks and deep learning, as well as various ensemble methods. In one embodiment, the VPE 100 is applied to real-time applications to identify sources of volatility in order to improve compliance with scheduling deadlines and reduce tail latency. In one embodiment, the VPE 100 can predict power volatility (separate from performance volatility), where power volatility can help the power manager comply with a specific power budget. In a specific embodiment, the VPE 100 can be used to optimize bulk synchronous processing (BSP) applications, such as those executed on the HPC cluster 600 (see Figure 6 ). BSP applications can be tightly coupled applications where a delay in one process or node can cause delays in the entire application. For example, the ML model 130 can point out slow processes and indicator identifiers associated with delays, while proposing specific remedial actions (such as replacing / upgrading hardware components).
[0059] The operations and functions performed by the ML system 110 can be summarized as data collection and data generation steps to extract the ML model 120. Therefore, the ML system 110 can first collect training data 102 from a system under test (not shown), and then extract the corresponding ML model 120 to analyze the system under test, such as by using the input data 106 to describe an operating scenario of the system under test different from that embodied in the training data 102. The collection and processing of the training data 102 can be performed by the following about Figure 3 300. In certain embodiments, the collection and processing of training data 102 may include identifying features of interest in the form of indicator identifiers, which are explanatory variables. The indicator identifiers may include low-level hardware features, such as cache events or I / O events, and OS operations, such as determining when a core identifier of a process / thread has changed (e.g., context switch) or an OS event (e.g., OS system call). Then, a training environment may be established, such as making a computer system with a given configuration available to execute an application of interest. The computer system for training may be prepared to a defined initial state or condition, such as by restarting or refreshing certain caches, queues, and memory buffers. The OS associated with a given configuration for training may be configured to access the indicator identifiers to be recorded, such as by enabling certain OS system calls to access hardware metrics and counters indicating performance features. Then, the application of interest may be executed on the computer system, such as N times (where N may be a very large number), referred to as "sampling" or collecting training data. In some embodiments, a smaller value of N is used and the N runs are repeated as a block, and certain reset actions may be performed before each block is repeated. Sampling may be repeated until a stopping condition for training is achieved, such as by indicating that the number of samples collected is a representative sample size. In various embodiments, the stopping condition may rely on a Geleman-Rubin statistic or a Monte-Carlo standard error for evaluation. When the stopping condition is not met, further sampling may be performed to generate more training data 102. When the stopping condition for sampling is reached, the ML system 110 may extract the ML model 120 for independent use.
[0060] When VPE 100 is used for distribution prediction (such as for predicting the distribution of a volatility benchmark), sampling can be performed to identify the volatility benchmark of interest and the associated indication identifiers, especially to identify the characteristics of the distribution of the volatility benchmark. The sampled data can be divided into training data 102 and validation data 104, and the ML system 110 can use this data to train the ML model 120 for distribution prediction. For example, training can be performed until certain confidence levels indicate an acceptable degree of convergence or a desired level of accuracy is achieved. Then, the ML model 120 can be used for the distribution prediction of the volatility benchmark.
[0061] When VPE 100 is used for statistical regression analysis, such as determining which indication identifiers are associated with which specific characteristics of the distribution of a particular volatility benchmark (e.g., having a causal relationship), the volatility output 130 and / or the input data 106 can be used. Statistical analysis can be performed, such as performing statistical analysis on the empirical distribution of the volatility benchmark included in the input data 106 or on the predicted true distribution of the volatility benchmark included in the volatility output 130, to identify those statistical characteristics associated with the mode of the corresponding distribution, such as multiple modes, the relative positions of the modes, the relative sizes of the modes, and outliers or other non-modal characteristics. When the modes have been identified and characterized in this way, further statistical analysis can be applied to relate specific indication identifiers to the identified modes. For example, when the distribution is identified as having a normal mode (Gaussian) indicating the sum of random variables, various indication identifiers presumably associated with the corresponding volatility benchmark can be analyzed to determine which sum of indication identifiers can match this normal distribution. In another example, when the distribution is identified as having a lognormal mode indicating the product of random variables, various indication identifiers presumably associated with the corresponding volatility benchmark can be analyzed to determine which product combination of indication identifiers can match this lognormal distribution. For the observed combinations in the previous two examples, the sampled distribution can be divided into mode subsets, and the above analysis can be repeated on the mode subsets.
[0062] Figure 2 Shows the ML model 120-1 in one embodiment. The ML model 120-1 is depicted as a neural network architecture having an input layer 210, inner layers 212, 214, and an output layer 216. The ML model 120-1 is a general representation that can be used to receive the input data 106 and generate the volatility output 130, as described above with reference to Figure 1As described. Thus, in various embodiments, the input data 106 may be provided to the ML model 120-1 as an input layer 210, and the volatility output 130 may be received from the ML model 120-1 as an output layer 216. It is noted that although the ML model 120-1 is depicted as having small groups of nodes or artificial neurons (referred to herein as simply "neurons"), the dimensions and structure of the ML model 120-1 may be adjusted for various specific types of data. For example, as shown, the ML model 120-1 may be expanded to a input neurons, y input layers (each having b...x neurons), and z output neurons. It is noted that in other values in various embodiments, a, b...x, y, and z may each have different dimensions, such as 10 3 , 10 6 , 10 9 , 10 12 .
[0063] exist Figure 2 In the mathematical processing of the ML model 120 - 1 , the processing at each layer can be represented by an activation expression, which can be summarized by Expression 1.
[0064] ∑ i (w i x i )+b expression 1
[0065] In Expression 1, i represents the index variable or dimension of each layer input, such as Figure 2 a, b...x and z in ; x represents the input value at each neuron, such as from another neuron; w represents the weighting coefficient applied at each neuron; and b represents a constant for each neuron. The output of each neuron can be represented by an activation function, which takes the result of expression 1 (e.g., the activation expression) as a parameter. Specifically, the ML model 120-1 can use deep learning (DL), which can be used to determine higher-level highly complex data abstractions using a hierarchical layered neural network architecture for learning. The ML model 120-1 can learn by stating, describing, and implementing higher-level more abstract features on top of lower-level less abstract features. In this way, the ML model 120-1 can use DL to analyze and learn from large amounts of unstructured data, which can be unlabeled and unclassified.
[0066] Figure 3 is a flowchart of a method 300 for training an ML model 120, such as described above with respect to Figure 1 Various method steps in method 300 may be omitted or rearranged in different embodiments.
[0067] Figure 3 The method 300 in FIG. 3 begins at step 302 by initializing the ML model. The initialization in step 302 may be associated with or dependent on the structure and content of the training data 102, validation data 104, input data 106, and volatility output 130 in the VPE 100. At step 304, the ML model is trained using the training data. At step 306, the ML model is validated using the validation data. At step 308, a determination is made as to whether the ML model output is correct within an acceptable confidence level. When the result of step 308 is no, the method 300 loops back to step 304. When the result of step 308 is yes, the method 300 proceeds to step 310 by extracting the ML model for use with different input data. The different input data may be the input data 106.
[0068] Figure 4 is a flowchart of a method 400 for using an ML model 120, such as described above with respect to Figure 1 Various method steps in method 400 may be omitted or rearranged in different embodiments.
[0069] Figure 4 The method 400 in begins at step 402 by reading input data. The input data may be input data 106. At step 404, the input data is preprocessed. At step 404, preprocessing the input data may involve filtering or smoothing or otherwise preparing the input data 106 for use with the ML model 130. At step 406, features may be extracted from the input data. Certain features or attributes of certain input data 106 extracted in step 406 may be used as other parts of the input data 106. At step 408, the ML model is applied to the features and the input data to generate output data. The output data in step 408 may be a volatility output 130. At step 410, the output data is provided. At step 412, contextual recommendations are provided based on the output data. At step 414, a statistical regression analysis is performed to identify explanatory variables. The explanatory variables in step 414 may be indicator identifiers or configuration parameters, etc., and may be associated with certain volatility benchmarks.
[0070] Figure 5 A block diagram depicting a computing node 500 according to one or more embodiments of the present disclosure is shown. The embodiments described herein may be implemented using computing nodes, such as computing nodes 500 in a standalone manner, or computing nodes 500 in a cluster of multiple computing nodes, such as an HPC cluster 600 including multiple computing nodes 500, as described below with respect to Figure 6Therefore, computing node 500 may represent any of a variety of computer systems or computing devices, such as a personal computer, a desktop computer, a laptop computer, a server, a blade computer, a modular computer, and an HPC computing node.
[0071] like Figure 5 As shown, computing node 500 includes a processor subsystem 520, a memory 530, a local storage resource 550, a network interface 560, an input / output (I / O) subsystem 540, and a local system bus 522 for interconnecting various local components with processor subsystem 520. Network interface 560 can implement connection to network 570, which will be described in further detail below.
[0072] like Figure 5 As shown, the processor subsystem 520 may include an integrated circuit, such as in the form of a chip, for interpreting and executing program instructions and process data. The processor subsystem 520 may include a general-purpose processor configured to execute program codes accessible to the processor subsystem 520. The processor subsystem 520 may include a special-purpose processor in which certain instructions are integrated into the processor subsystem 520. The processor subsystem 520 may represent a single processor or multiple processors working together in the computing node 500. The processor subsystem 520 may also represent a plurality of different kinds of processors, such as processors for different types of tasks, including CPUs and GPUs used in the computing node 500. In addition, the processor subsystem 520 may include multiple cores or micro-cores (not shown) for executing program codes or processing different processes. In some embodiments, the processor subsystem 520 may interpret and execute program instructions and process data stored locally (e.g., in the memory 530). In a specific embodiment, the processor subsystem 520 may interpret and execute program instructions and process data stored remotely (e.g., in a network storage resource accessible using the network interface 560).
[0073] exist Figure 5 In the drawings, system bus 522 may represent various suitable types of bus structures, such as a memory bus, a peripheral bus or a local bus using various bus architectures in selected embodiments.
[0074] Also in Figure 5In the embodiment of the present invention, the memory 530 may include a system, device or apparatus (e.g., a computer-readable medium) that can be used to retain and retrieve program instructions and data over a period of time. The memory 530 may include volatile memory such as random access memory (RAM), cache memory, magnetic memory, etc. In some embodiments, the memory 530 includes any of various non-volatile memories that retain data after power is removed, such as a hard disk, an optical drive such as a compact disk (CD) drive or a digital versatile disk (DVD) drive, flash memory, electrically erasable programmable read-only memory (EEPROM), a memory card, a magnetic memory, an optical magnetic memory, etc. The memory 530 may also include or represent a computer-readable medium (not shown), which includes but is not limited to portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. The computer-readable medium may include a non-temporary medium that stores data, excluding carrier waves and / or temporary electronic signals that are transmitted wirelessly or through a wired connection. Computer readable media may store code and / or processor executable instructions, where the code and / or processor executable instructions may represent a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements, etc.
[0075] exist Figure 5 5, a memory 530 is shown including an operating system (OS) 532, which may represent an execution environment for various program codes executed on the computing node 500. The operating system 532 may be any of a variety of standard or customized operating systems, such as, but not limited to, Microsoft operating system, UNIX or UNIX-based operating system, mobile device operating system (e.g., Google Android(TM) platform, iOS, etc.), MacOS operating system, embedded operating system, etc. Also shown is a memory 530 including VPE 100-1, wherein VPE 100-1 represents VPE 100 (with respect to Figure 1 ) to enable the processor subsystem 520 of the computing node 100 to execute at least some portions of the VPE 100.
[0076] In computing node 500, I / O subsystem 540 may include systems, devices or apparatuses that are generally used to receive data from computing node 500, or transmit data to computing node 500, or receive and transmit data within computing node 500. In different embodiments, I / O subsystem 540 may be used to support various peripheral devices, such as touch panel, display adapter, keyboard, touch pad or camera, etc. I / O subsystem 540 may represent, for example, various communication interfaces, graphic interfaces, video interfaces, user input interfaces and peripheral interfaces. For example, I / O subsystem 540 may support various output or display devices, such as screen, monitor, general display device, liquid crystal display (LCD), plasma display, touch screen, projector, printer, external storage device or other output device. In some instances, I / O subsystem 540 may support multi-mode system, which allows users to provide multiple types of I / O to communicate with computing node 500.
[0077] exist Figure 5 In the embodiment of the present invention, the local storage resource 550 may include non-volatile or persistent computer-readable media, such as hard disk drives, CD-ROMs, and other types of rotating storage media, flash memory, EEPROMs, or other types of solid-state storage media, and the local storage resource 550 may generally be used to store instructions and data and allow access to the stored instructions and data as needed. In some embodiments, the local storage resource 550 may include a storage device or storage subsystem (not shown) having one or more storage device arrays, such as for supporting redundancy, mirroring, and / or real-time data error correction and recovery.
[0078] In addition, Figure 5 500 to a network 570, which may represent a local area network (LAN), a wide area network (WAN) such as the Internet, a mobile network, or other type of network. Network interface 560 may provide communication with another device (such as another computing node). Network interface 560 may include or support a wireless network or a wired network. Wired network media supported by network interface 560 may include analog media, a universal serial bus (USB), Ethernet, optical fiber, dedicated wired media, public switched telephone network (PSTN), integrated services digital network (ISDN), and ad hoc network media, etc. The wireless network media supported by the network interface 560 may include visible light communication (VLC), world interoperability for microwave access (WiMAX), Wireless signal transmission, BLE wireless signal transmission, Wireless signal transmission, RFID wireless signal transmission, Near Field Communication (NFC) wireless signal transmission, Dedicated Short Range Communication (DSRC) wireless signal transmission, 802.11 WiFi wireless signal transmission, WLAN signal transmission, IR communication wireless signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, etc.
[0079] Such as Figure 5 As shown, network interface 560 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of computing node 500 based on receiving one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based GPS, the Russian-based Global Navigation Satellite System (GLONASS), the Chinese-based BeiDou Navigation Satellite System (BDS), and the European-based Galileo GNSS. In certain embodiments, network interface 560 may support the expansion or addition of new network interfaces or media.
[0080] At least some portions of computing node 500 may be implemented in circuitry. For example, the components of computing node 500 may include electronic circuits or other electronic hardware, which may include programmable electronic circuits, microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and other suitable electronic circuits. Certain functions integrated into computing node 500 may be provided using executable code accessible by the electronic circuits to perform the methods and operations described herein. As described above, the executable code includes computer software, firmware, program code, or various combinations thereof. When specified, non-transitory media explicitly excludes transitory media such as energy, carrier signals, light beams, and electromagnetic waves.
[0081] Figure 6 A block diagram description of an HPC cluster according to one or more embodiments of the present disclosure is shown. The embodiments described herein may be implemented using an HPC cluster, such as the HPC cluster 600 including multiple computing nodes 500 shown (see Figure 5 ). Although four computing nodes 500-1, 500-2, 500-3, 500-4 are shown in Figure 6 for purposes of description, it should be noted that any number of computing nodes 500 may be used. In certain embodiments, a large number of computing nodes 500 may be aggregated in HPC cluster 600 to provide greater computing power and may be used to implement a supercomputer in some embodiments. Thus, workloads (such as those regarding Figure 5 The VPE 100 - 1 shown and described, for example, enables multi-node application execution so that the computing nodes 500 share the processing of a workload, which can be executed in parallel or simultaneously among the computing nodes 500 .
[0082] HPC cluster 600 can be generally described as each comprising a local processor and a local memory and connected via a dedicated high-bandwidth, low-latency network (on Figure 6 The HPC cluster 600 is a collection of computing nodes 500 interconnected by a high-speed local network 622 (shown in FIG. 6 ). The HPC cluster 600 can aggregate and combine the computing power of multiple computing nodes 500 accordingly to perform large-scale workloads. The HPC cluster 600 can provide the flexibility and scalability of HPC resources so that the computing power can be well matched to the current and evolving workload requirements in an economical and seamless manner. The HPC cluster 600 can also provide great flexibility in cluster configuration to handle task parallelization, data distribution, parallel execution, cluster monitoring and control, and the output of combined parallelized calculations. Applications can be executed on the HPC cluster 600 in a local or distributed manner, for example, on a single HPC computing node 500-1 or on multiple HPC computing nodes 500-2, 500-3, 500-4.
[0083] like Figure 6 As shown, the HPC cluster 600 is shown to include storage nodes 650, which may represent storage devices compatible with a high-speed local network 622. The high-speed local network 622 may be a dedicated local bus, such as InfiniBand, 40Gb Ethernet, or a high-speed peripheral connection interface (PCIe). Therefore, the storage node 650 can use the low-latency high-speed local network 622 to provide access to storage resources to support HPC workloads processed using the HPC cluster 600. It is also noted that the HPC cluster 600 may include a dedicated network interface ( Figure 6 ), such as by using computing node 500.
[0084] Figure 7 is a flow chart of a method 700 for predicting computer performance volatility. Figure 1 The VPE 100 in the embodiment of the present invention may be used to perform at least some portions of the method 700, as described herein. Various method steps in the method 700 may be omitted or rearranged in different embodiments.
[0085] exist Figure 7In the method 700, the method 700 may begin at step 702 by executing a first application on a first computer system having a first configuration for a first time. At step 704, during the execution of the first application for the first time, a first value is recorded, the first value comprising at least one first indicator identifier associated with the first computer system. At step 706, a second value is recorded, the second value comprising at least one first volatility benchmark associated with the execution of the first application for the first time. At step 708, using a machine learning (ML) model, the machine learning (ML) model is configured to predict a true distribution of a volatility benchmark based on an empirical distribution of learned indicator identifiers and an empirical distribution of volatility benchmarks associated with the execution of the application, the first value and the second value are input to the ML model. At step 710, output information indicating the true distribution of the first volatility benchmark for the first configuration is received from the ML model.
[0086] Figure 8 is a flow chart of a method 800 for predicting computer performance volatility. Figure 1 The VPE 100 in the embodiment of the present invention may be used to perform at least some portions of the method 800, as described herein. Various method steps in the method 800 may be omitted or rearranged in different embodiments.
[0087] exist Figure 8 In the method 800, at step 802, method 800 may begin by initiating at least a partial execution of a first application on a first computer system having a first configuration. At step 804, during the partial execution of the first application, a first value is recorded, the first value comprising at least one first indicator identifier associated with the first computer system. At step 808, using a machine learning (ML) model that is trained to predict a true distribution of a volatility benchmark based on an empirical distribution of the learning indicator identifiers and an empirical distribution of the volatility benchmark associated with the execution of the application, the first value is input to the ML model. At step 810, output information indicating a true distribution of the first volatility benchmark for the first configuration is received from the ML model, including information indicating a confidence interval of the true distribution.
[0088] Fig. 9 is a flow chart of a method 900 for predicting computer performance volatility. Figure 1 The VPE 100 in the embodiment of the present invention may be used to perform at least some portions of the method 900, as described herein. In different embodiments, various method steps in the method 900 may be omitted or rearranged.
[0089] exist Fig. 9In step 902, method 900 may begin by providing a machine learning (ML) model that is trained to predict a true distribution of a first volatility benchmark for a first application based on an empirical distribution of learning indicator identifiers and an empirical distribution of volatility benchmarks associated with executing the first application on a first computer system having a first configuration. At step 904, output information indicating the true distribution of the first volatility benchmark for the first configuration is received from the ML model, including information indicating a confidence interval for the output information. At step 906, output information indicating the true distribution of the first volatility benchmark for a second configuration different from the first configuration is received from the ML model. At step 908, output information indicating the true distribution of the first volatility benchmark associated with executing the second application different from the first application is received from the ML model.
[0090] Fig.10 is a flow chart of a method 1000 for predicting computer performance volatility. Figure 1 The VPE 100 in the embodiment of the present invention may be used to perform at least some parts of the method 1000, as described herein. Various method steps in the method 1000 may be omitted or rearranged in different embodiments.
[0091] exist Fig.10 In the method 1000, at step 1002, a machine learning (ML) model is provided, the machine learning (ML) model being trained to learn a relationship between an empirical distribution of indicator identifiers and an empirical distribution of a volatility benchmark associated with executing a first application on a first computer system having a first configuration. At step 1004, at least one empirical distribution of a first volatility benchmark associated with the first application is assigned to the ML model. At step 1006, output information indicating at least one first indicator identifier associated with the first volatility benchmark is received from the ML model. At step 1008, at least one configuration parameter of the first configuration that exhibits a causal correlation with the empirical distribution of the first volatility benchmark is predicted based on the first indicator identifier and the first configuration. At step 1010, at least one first modification in the configuration parameter that may affect the empirical distribution of the first volatility benchmark is output based on the causal correlation. At step 1012, a statistical relationship that explains the causal correlation is determined using the statistical model, including a confidence factor for the statistical relationship. At step 1014, at least one second modification in the configuration parameter that may affect the empirical distribution of the first volatility benchmark is output based on the statistical relationship.
[0092] Certain implementation examples of the methods and systems disclosed herein for computer volatility forecasting are described below.
[0093] Example 1
[0094] Identifying applications that are particularly sensitive to dynamic random access memory (DRAM) latency. Using two different computer systems each having two different configurations, a volatility output 130 of a VPE 100 indicates that certain volatility benchmarks, such as application runtime, exhibit higher performance on a first computer system having an indicator identifier whose value is associated with lower DRAM latency than on a second computer system having an indicator identifier whose value is associated with higher DRAM latency. However, other differences in configuration parameters, such as a difference in CPU clock speed between the first computer system and the second computer system, may obscure or mask the correlation with DRAM latency that is apparent from the volatility output 130.
[0095] In other examples, on the first computer system, where other indicator identifiers remain the same, the volatility output 130 of the VPE 100 indicates a CPU migration that migrates a process of an application from one non-uniform memory access (NUMA) node to another non-uniform memory access (NUMA) node. Other indicator identifiers associated with higher performance may be indicated by the volatility output 130 of the VPE 100, such as an indicator identifier for multiple L3 cache misses.
[0096] Example 2
[0097] It can be observed that large HPC clusters configured as supercomputers (see Figure 6 The computing resources of the HPC cluster 600 in the example are typically managed by a batch job scheduler, such as the Simple Linux Resource Management Tool (SLURM). When a new job (such as an application) is queued for execution, the scheduler decides when and on which HPC nodes to execute the job. The scheduler uses the job's runtime estimate to schedule multiple short jobs that are expected to complete execution before the resources consumed by the short job are used for another subsequent job of a larger or higher priority. The runtime estimate for the short job is calculated or obtained externally. However, during the execution of the short job, the actual runtime of the short job does not correspond to the runtime estimate, which is partly due to the performance volatility of some HPC nodes. Therefore, the job scheduler may take undesirable actions such that jobs that exceed their allocated runtimes are killed (causing waste) or jobs with higher priority are delayed.
[0098] In contrast, the variable output 130 of the VPE 100 is used to provide a more accurate estimate of the run time, for example, by predicting the true distribution of the run time as a variable baseline. The prediction made by the VPE 100 provides an upper bound on the expected range of the run time with an estimated confidence, which is incorporated into the scheduling decision to avoid killing or delaying jobs. More details of an example implementation will be described below.
[0099] At time 14:00, a high-priority job A queues up and requests 3,000 processors from the job scheduler of the HPC cluster to execute on the HPC cluster. Based on the variable output 130 of job A at time 14:00, the job scheduler determines that 2,000 processors are available, while job B is currently executing using 1,000 processors of the HPC cluster. The job scheduler decides to schedule job A at time 14:30 when job B is expected to complete.
[0100] At time 14:10, a low-priority job C queues up and requests 1,000 processors from the job scheduler to execute on the HPC cluster. In the absence of the variable output 130, the job scheduler initially estimates that job C will have a run time of 15 minutes. Since it is expected that job C will not affect the execution time of job A, the job scheduler determines that job C can be backfilled ahead of job A in the execution queue. However, in reality, the true distribution of the run time of job C, which is correctly predicted by the variable output 130 as the variable baseline, indicates that there is a 20% probability that job C will exceed 15 minutes and a probability of less than 1% that it will exceed 20 minutes. Based on the true distribution provided by the variable output 130, the job scheduler makes a decision on whether to execute job C at time 14:10 based on a predefined policy. Since there is a probability greater than 99% that the pending low-priority job C in the queue will complete execution before 14:30, the job scheduler backfills job C ahead of job A, and job A is still scheduled at 14:30. In this way, the job scheduler can improve the utilization of the HPC cluster and avoid unused idle time of the HPC cluster, which is economically desirable and made possible by the VPE 100.
[0101] Example 3
[0102] During procurement of a CPU for a given configuration of a new computer system, there are often different CPU options available for purchase for the same CPU model, which may involve selecting from a number of cores, cache sizes, clock frequencies, and other CPU features. Because the business impact of such selections in a given enterprise may be undefined or unknown, the basis for making selections on such CPU-related features during procurement may be difficult or unclear. As a result, a more or less powerful CPU configuration may ultimately be selected based on other criteria (such as available procurement budget) rather than on the actual performance impact on the enterprise, and such selection is undesirable and may adversely affect the enterprise financially or in terms of user productivity when an incompatible CPU is selected.
[0103] By using the volatility output 130 of the VPE 100, the true distribution of volatility benchmarks associated with existing computer systems can be predicted and used in predictive analysis of new configurations of new computer systems to be purchased. Such predictive analysis enabled by the volatility output 130 can better color decisions on the value or utility of certain CPU features to include in new configurations by showing the impact of various CPU characteristics on the true distribution of volatility benchmarks for applications and conditions actually experienced by users in the enterprise. In this way, CPU options that provide overall higher performance and lower volatility can be selected for the enterprise, which is desirable.
[0104] For example, volatility output 130 may indicate that an existing configuration of an enterprise server typically exhibits a high level of cache misses when the CPU is running under a typical application workload of the enterprise. This information may guide purchasing decisions for the next generation of enterprise servers to include a larger cache size when selecting a new CPU.
[0105] Similarly, when volatility output 130 indicates that the page miss rate on the second enterprise server is higher than the page miss rate on the other servers, VPE 100 may also predict that a DRAM module with a larger size will improve performance and reduce volatility. Specifically, VPE 100 may be used to predict the relationship between the page miss rate and the size of the DRAM module on various servers, thereby enabling comparison of this relationship between servers.
[0106] Example 4
[0107] It can be observed that some applications can exhibit very long tail performance, which causes some execution instances of the application to have excessively long runtimes. In certain large enterprise environments that process billions of transactions per day, the cost impact of such extended runtimes of application workloads can be very large and economically significant due to the high cost of operating such large data centers.
[0108] By using VPE 100, volatility output 130 and statistical modeling using volatility output 130 can associate a volatility benchmark having a long-tail distribution of run times with a composite indicator identifier that points to a specific aspect of the configuration of the computer system being used. The empirical distribution of indicator identifiers predicted by volatility output 130 can indicate abnormal conditions (e.g., high CPU temperatures) or normal but infrequent operating conditions (such as OS service wakeups for routine maintenance) that explain the observed behavior. In addition, as noted, volatility output 130 can include confidence levels and magnitude levels for the relationships found or predicted between volatility benchmarks and indicator identifiers. Using this valuable insight provided by VPE 100, remedial actions to reduce observed volatility can be focused on high-impact or high-likelihood indicator identifiers that have a causal relationship with the observed volatility benchmark.
[0109] Specifically, it has been observed that query latency for large language model (LLM) generative artificial intelligence (AI) applications includes very large latency that results in very long runtimes, which is undesirable. Using VPE 100, the true distribution of runtimes for queries predicted by volatility output 130 (volatility benchmark) indicates a log-normal distribution with 98% confidence, indicating that the true distribution is a product of multipliers (e.g., representing a composite of two or more indicator identifiers that behave as random variables).
[0110] Bayesian linear regression is used to statistically model fit the logarithmic relationship of query runtime (volatility benchmark) to various indicator identifiers and their combinations or permutations. It was found that most indicator identifiers had a negligible effect on query runtime, but six indicator identifiers were found to have p-values less than 0.05, indicating statistical significance for these six indicator identifiers. Stepwise regression was further performed to identify the single most influential indicator identifier, and the interaction between the indicator identifiers "query length" and "available RAM" had p < 0.001, which indicates a high degree of statistical significance. This result is used to make modifications in the load balancer, namely sending long queries to be executed by compute nodes with additional available RAM, and sending short queries to be executed by compute nodes with less available RAM.
[0111] Example 5
[0112] The image resizing service performs an image reduction job that reduces an input image of size 640×640 pixels to an output image of size 320×320 pixels. The true distribution of the volatility benchmark for runtime is predicted by the VPE 100 and the results show that the true distribution exhibits a bimodal distribution: 70% of the service executions have runtimes very close to 1 second and 30% of the service executions have runtimes very close to 3 seconds. Using the VPE 100 to identify the relationship between runtime (volatility benchmark) and various indicator identifiers, the volatility output 130 indicates that the indicator identifiers associated with the input data do not show little or no correlation with runtime. In other words, the VPE 100 predicts that the attributes of the input image will not have any meaningful impact on runtime. However, the volatility output 130 does predict strong correlations with the indicator identifiers for multiple core migrations and processor identifiers, with a statistically significant confidence level. The volatility output 130 also predicts that the strong correlation is valid for CPUs from a first manufacturer, but not for CPUs from a second manufacturer. For a configuration including CPUs from a first manufacturer, volatility output 130 provides an indication that cores located further from system memory cause different execution patterns for the indicated identifiers (run times) that are clearly visible in the real distribution, such that the run times are dominated by the latency of such memory operations, which is three times longer than the latency of accessing memory closer to the cores. Therefore, VPE 100 provides a prescriptive recommendation to bind each process of a service to a single core closest to the memory where the input data is stored, thereby eliminating the three-times long execution pattern, which is desirable.
[0113] In summary, the methods and systems for predicting computer volatility disclosed herein can be used to explain the performance volatility of applications executed on existing computer systems or nodes, in order to characterize, debug, and improve performance. Certain embodiments can predict the performance volatility of applications to be executed on new computer systems or nodes that are planned or in design. Certain embodiments can measure, control, and reduce volatility in HPC clusters and AI applications that are highly sensitive to synchronization mismatches. Certain embodiments can optimize the training / inference performance of HPC clusters and AI applications by identifying some of the causes of anomalous performance and subsequently eliminating or mitigating such causes. Certain embodiments can improve resource efficiency and resource management on supercomputers and HPC clusters. Certain embodiments can predict or recommend specific CPU and associated peripheral device combinations that can minimize volatility (such as that observed in the empirical distribution of volatility benchmarks). Certain embodiments can identify potential cross-application interference exhibited in the volatility benchmarks of application execution. Certain embodiments can provide feedback, discovery, and insights into the true distribution of volatility benchmarks for improving scheduling decisions made by job schedulers. Certain embodiments can identify hardware resources to be up / downregulated based on the impact on the empirical distribution of volatility benchmarks.
[0114] As disclosed herein, an ML model can be trained using an observed volatility benchmark and an indication identifier of an application executed on a computer system having a given configuration to predict the true distribution of the volatility benchmark. The volatility benchmark can be correlated with the indication identifier to determine the mode in the true distribution associated with a particular indication identifier. A causal relationship between certain indication identifiers and the volatility benchmark can be determined. Normative measures can be provided to improve the performance observed in the volatility benchmark, such as modifying the indication identifier.
[0115] As disclosed herein, an ML model can be trained to learn the relationship between the empirical distribution of indication identifiers and the empirical distribution of a volatility benchmark associated with executing an application on a first computer system having a first configuration. At least one empirical distribution of the first volatility benchmark associated with the application is specified to the ML model. Output information indicating at least one indication identifier associated with the first volatility benchmark is received from the ML model.
[0116] Various aspects may be described as a process or method, which is depicted as a flow chart, a flowchart, a data flow diagram, a structure diagram, or a block diagram. Although a flow chart may describe an operation as a sequential process, many operations may be performed in parallel or simultaneously. In addition, the order of the operations may be rearranged. When the operation of a process or method is completed, the process or method may be terminated, but there may be additional steps not included in the flow chart. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to the return of the function to the calling function or the main function.
[0117] Computer executable instructions stored or otherwise obtained from a computer readable medium can be used to implement the processes and methods according to the above examples. Such instructions may include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a processing device to perform a specific function or function group or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a specific function or function group. A portion of the computer resources used may be accessed over a network. Computer executable instructions may be, for example, binary code, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer readable media that can be used to store instructions, information, and / or information created during the method according to the example include a disk or optical disk, a flash memory, a USB device with a non-volatile memory, a networked storage device, etc.
[0118] In the above-mentioned figure description, any component described with respect to the figure can be equivalent to one or more components of the same or similar name and / or number described with respect to any other figure in various embodiments described herein. For the sake of brevity, the description of these components is not repeated for each figure.
[0119] Throughout the application, ordinal numbers (e.g., first, second, third, etc.) may be used as adjectives of elements (i.e., any nouns in the application). The use of ordinal numbers does not imply or create any particular sequence of elements, nor does it limit any element to being only a single element, unless expressly indicated, such as by the use of the terms "before," "after," "single," and other such terms. Rather, the use of ordinal numbers is to distinguish between such elements, such as for classification purposes. As an example, a first element is different from a second element, and a first element may include more than one element and is after (or before) a second element in the ordering of the elements.
[0120] Although the present disclosure has been described with reference to illustrative embodiments, the description is not intended to be construed in a limiting sense. Various modifications and combinations of the illustrative embodiments, as well as other embodiments, will be apparent with reference to the description.
Claims
1. A method for predicting computer performance volatility, the method comprising: Initiating at least partial execution of a first application on a first computer system having a first configuration; during the portion of execution of the first application, recording a first value, the first value comprising at least one first indicator identifier associated with the first computer system; inputting the first value into a machine learning (ML) model using the ML model, the machine learning (ML) model being trained to predict a true distribution of the at least one first volatility benchmark based on an empirical distribution of learning indicator identifiers and an empirical distribution of volatility benchmarks associated with executing the first application; as well as Output information indicative of the true distribution of the first volatility benchmark for the first configuration is received from the ML model, including information indicative of a confidence interval for the true distribution.
2. The method according to claim 1, wherein: The first volatility benchmark is selected from at least one of the following: running time, response delay, response delay probability, data throughput rate, interval time, end timestamp, start timestamp or data throughput capacity.
3. The method according to claim 1, wherein: The true distribution is characterized by at least one of the following: statistical variables, including mean, median, mode position, mode number, mode size; Distribution variables, including standard deviation, standard error, and variance; Confidence interval; High density area; or At least one parameter for curve fitting, including for normal (Gaussian) curve fitting, bimodal curve fitting, multimodal curve fitting, or logarithmic modal curve fitting.
4. The method according to claim 1, wherein: The first configuration of the first computer system specifies at least one of the following: Central processing unit (CPU) parameters, including base clock frequency, cache size, number of cores, number of logical processors, and peripheral bus clock speed; Graphics processing unit (GPU) parameters, including GPU version, GPU clock speed, GPU cache; Memory parameters, including physical memory size, number of memory cards, memory card size, memory interface, nominal memory write speed, nominal memory read speed, nominal memory latency, number of channels, memory clock speed, memory access control, memory allocation control; Operating system (OS) parameters, including version number, update count, update list, registry content, directory content, and power management mode; Network parameters, including network capacity, number of physical ports, type of physical ports, media type; or Local storage parameters, including the number of physical volumes, the number of logical volumes, the size of the volume, capacity / volume, file system identifier, file system version, storage medium type, redundant volumes, file system write speed, and file system read speed.
5. The method according to claim 1, wherein: The indication identifier includes at least one of the following: Central processing unit (CPU) metrics, including CPU utilization percentage, actual clock frequency, number of processes, number of threads, number of handles, cache events, CPU events, cycle counts, instruction counts, I / O events, operating system events, kernel identifiers; Graphics Processing Unit (GPU) metrics, including GPU utilization percentage, GPU memory usage, and GPU shared memory; Memory metrics, including memory usage, available memory, committed memory, cached memory, paged memory, nonpaged memory; Operating system (OS) metrics, including application runtime, code segment runtime, virtual memory size, application CPU time, application end timestamp, code segment end timestamp; Network metrics, including throughput rate, sent data rate, received data rate, network capacity rate, packet error rate, application network data usage, application network data rate; or Local storage metrics, including response time, average response time, percentage of active time, transfer rate, file system latency, write speed, read speed, access time.
6. A method comprising: providing a machine learning (ML) model trained to learn a relationship between an empirical distribution of an indicator identifier and an empirical distribution of a volatility benchmark associated with executing an application on a first computer system having a first configuration; assigning to the ML model at least one empirical distribution of a first volatility benchmark associated with the application; as well as Output information indicative of at least one identifier associated with the first volatility benchmark is received from the ML model.
7. The method according to claim 6, further comprising: Output information indicative of a true distribution of the first volatility benchmark for a second configuration different from the first configuration is received from the ML model.
8. The method according to claim 6, wherein: The first volatility benchmark is selected from at least one of the following: running time, response delay, response delay probability, data throughput rate, interval time, end timestamp, start timestamp or data throughput capacity; and where the true distribution is characterized by at least one of the following: Statistical variables, including mean, median, mode position, mode number, and mode size; Distribution variables, including standard deviation, standard error, and variance; Confidence interval; High density area; or At least one parameter for curve fitting, including for normal (Gaussian) curve fitting, bimodal curve fitting, multimodal curve fitting, or logarithmic modal curve fitting.
9. The method according to claim 6, wherein: The first configuration of the first computer system specifies at least one of the following: Central processing unit (CPU) parameters, including base clock frequency, cache size, number of cores, number of logical processors, and peripheral bus clock speed; Graphics processing unit (GPU) parameters, including GPU version, GPU clock speed, GPU cache; Memory parameters, including physical memory size, number of memory cards, memory card size, memory interface, nominal memory write speed, nominal memory read speed, nominal memory latency, number of channels, memory clock speed, memory access control, memory allocation control; Operating system (OS) parameters, including version number, update count, update list, registry content, directory content, and power management mode; Network parameters, including network capacity, number of physical ports, type of physical ports, media type; or Local storage parameters, including the number of physical volumes, the number of logical volumes, the size of the volume, capacity / volume, file system identifier, file system version, storage medium type, redundant volumes, file system write speed, and file system read speed.
10. The method according to claim 6, wherein: The at least one indication identifier includes at least one of the following: Central Processing Unit (CPU) metrics, including CPU utilization percentage, actual clock frequency, number of processes, number of threads, number of handles, cache events, CPU events, cycle counts, instruction counts, IO events, operating system events, kernel identifiers; Graphics Processing Unit (GPU) metrics, including GPU utilization percentage, GPU memory usage, and GPU shared memory; Memory metrics, including memory usage, available memory, committed memory, cached memory, paged memory, nonpaged memory; Operating system (OS) metrics, including application runtime, code segment runtime, virtual memory size, application CPU time, application end timestamp, code segment end timestamp; Network metrics, including throughput rate, sent data rate, received data rate, network capacity rate, packet error rate, application network data usage, application network data rate; or Local storage metrics, including response time, average response time, percentage of active time, transfer rate, file system latency, write speed, read speed, access time.
11. The method according to claim 6, wherein: A plurality of indicator identifiers are associated with the empirical distribution of the first volatility benchmark, and wherein each of the indicator identifiers represents a mode in the empirical distribution of the first volatility benchmark.
12. The method according to claim 6, further comprising: predicting, based on the indicative identifier and the first configuration, at least one configuration parameter of the first configuration that exhibits a causal correlation with the empirical distribution of the first volatility benchmark; as well as Based on the causal correlation, at least one first modification in the configuration parameter that may affect the empirical distribution of the first volatility benchmark is output.
13. The method according to claim 12, further comprising: using a statistical model to determine a statistical relationship that explains the causal association, including a confidence coefficient for the statistical relationship; as well as Based on the statistical relationship, at least one second modification in the configuration parameter that may affect the empirical distribution of the first volatility benchmark is output.
14. The method according to claim 13, wherein: The statistical model includes at least one of the following: Regression curve fitting; Bayesian inference; statistical correlation; and analysis of variance (ANOVA); and decision tree statistical models.
15. A computer system for predicting computer performance volatility, the computer system comprising at least one processor, the processor being configured to: providing a machine learning (ML) model trained to predict a true distribution of a first volatility benchmark for a first application based on an empirical distribution of learning indicator identifiers and an empirical distribution of volatility benchmarks associated with executing the first application on a first computer system having a first configuration; and Output information indicative of the true distribution of the first volatility benchmark for the first configuration is received from the ML model, including information indicative of a confidence interval for the output information.
16. The computer system of claim 15, further comprising: Output information indicative of the true distribution of the first volatility benchmark for a second configuration different from the first configuration is received from the ML model.
17. The computer system of claim 15, further comprising: Output information indicative of the true distribution of the first volatility benchmark associated with executing a second application different from the first application is received from the ML model.
18. The method of claim 15, wherein: The first volatility benchmark is selected from at least one of the following: running time, response delay, response delay probability, data throughput rate, interval time, end timestamp, start timestamp or data throughput capacity.
19. The method according to claim 15, wherein: The first configuration of the first computer system specifies at least one of the following: Central processing unit (CPU) parameters, including base clock frequency, cache size, number of cores, number of logical processors, and peripheral bus clock speed; Graphics processing unit (GPU) parameters, including GPU version, GPU clock speed, GPU cache; Memory parameters, including physical memory size, number of memory cards, memory card size, memory interface, nominal memory write speed, nominal memory read speed, nominal memory latency, number of channels, memory clock speed, memory access control, memory allocation control; Operating system (OS) parameters, including version number, update count, update list, registry content, directory content, and power management mode; Network parameters, including network capacity, number of physical ports, type of physical ports, media type; or Local storage parameters, including the number of physical volumes, the number of logical volumes, the size of the volume, capacity / volume, file system identifier, file system version, storage medium type, redundant volumes, file system write speed, and file system read speed.
20. The method according to claim 15, wherein: The indication identifier includes at least one of the following: Central Processing Unit (CPU) metrics, including CPU utilization percentage, actual clock frequency, number of processes, number of threads, number of handles, cache events, CPU events, cycle counts, instruction counts, IO events, operating system events, kernel identifiers; Graphics Processing Unit (GPU) metrics, including GPU utilization percentage, GPU memory usage, and GPU shared memory; Memory metrics, including memory usage, available memory, committed memory, cached memory, paged memory, nonpaged memory; Operating system (OS) metrics, including application runtime, code segment runtime, virtual memory size, application CPU time, application end timestamp, code segment end timestamp; Network metrics, including throughput rate, sent data rate, received data rate, network capacity rate, packet error rate, application network data usage, application network data rate; or Local storage metrics, including response time, average response time, percentage of active time, transfer rate, file system latency, write speed, read speed, access time.