Pattern-recognition enabled autonomous configuration optimization for data center

By simulating environmental conditions in a chamber and training models like MSET or MSET2, the method optimizes data center configurations efficiently, addressing performance issues due to temperature and altitude, and ensuring SLA compliance.

JP2025108431AActive Publication Date: 2025-07-23ORACLE INT CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025042459
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-02-26
Filing Date
2025-03-17
Publication Date
2025-07-23
Estimated Expiration
2041-02-19

AI Technical Summary

Technical Problem

Determining an optimal data center configuration is complex and time-consuming due to the dependency on environmental factors like temperature and altitude, leading to performance degradation and SLA violations, especially with CPU-intensive and I/O-intensive workloads.

Method used

Utilizing an environmental chamber to simulate various temperature and altitude conditions, training a model like MSET or MSET2 with telemetry data, and selecting a pre-trained model for the new data center configuration to optimize performance metrics.

Benefits of technology

Reduces the time required to set up optimal data center configurations from days to hours by using a library of pre-trained models, ensuring compliance with SLA requirements and minimizing performance degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108431000001_ABST
    Figure 2025108431000001_ABST
Patent Text Reader

Abstract

To provide a method for determining an optimal configuration for a data center, a system including a processor and a memory device, and a non-transitory computer-readable medium.SOLUTION: A method includes steps of: receiving or measuring a first performance metric from a data center; determining whether the first performance metric meets an SLA requirement; measuring one or more environmental characteristics of the data center as the data center executes a workload; identifying a pre-trained model in a model library with similar environmental characteristics; generating performance metrics for the pre-trained model; comparing the performance metrics from a data center configuration with performance metrics generated by the pre-trained model; determining a new configuration for the data center from the model; and using the configuration from the model for the installed data center.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cross - reference to Related Applications This application claims the benefit and priority of U.S. Patent Application No. 16 / 801,590, filed on February 26, 2020, entitled "PATTERN - RECOGNITION ENABLED AUTONOMOUS CONFIGURATION OPTIMIZATION FOR DATA CENTERS", which is hereby incorporated by reference in its entirety.

[0002] Background A data center may include any collection of computer systems, as well as related components such as processing systems, telecommunications systems, and data storage systems. A data center may include several servers or CPU cores for performing processing operations, network equipment for communication, memory devices for storing data, and / or redundant or backup components for infrastructure, power, data storage, and / or data communication. Modern data centers may be private / on - premise data centers installed in a customer's on - premise facility, while other data centers may be public - facing cloud - based data centers provided by a service provider. Cloud - based data centers provide hardware and / or software that may be used by cloud tenants to perform cloud tenant data operations.

[0003] A data center configuration may include any configuration parameters that define the configuration and operation of a data center. The configuration of a data center can be beneficial for installing and operating various hardware and software components to meet the operating requirements. Generally, determining the optimal data center configuration is a very complex process, but selecting the appropriate configuration is important for ensuring satisfactory performance of the data center. A sub-optimal configuration can lead to long latency, low throughput, multiple I / O timeouts, and ultimately a violation of the service level agreement. Due to the complexity of the interactions between various software / hardware components within a data center, determining the optimal configuration is currently considered more of an intuition rather than a science supported by empirical results. This often leads to a trial-and-error process that can take days or weeks to complete. Therefore, an improvement in the technology for determining the optimal data center configuration can be beneficial.

Summary of the Invention

[0004] Summary The performance of computer hardware in a data center is increasingly dependent on environmental factors such as ambient temperature and altitude. At higher temperatures, the fan speed has to be increased to cool the data center during CPU-intensive and memory I / O-intensive workloads. At higher altitudes, the fan speed has to be increased to compensate for the thinner air. The increase in fan speed leads to more vibrations and acoustic noise that interfere with the mechanical operation of the hard disk drive. The elevated temperature also causes more throttling in terms of CPU speed and voltage. Without properly considering these environmental factors, a data center configuration cannot be systematically optimized with respect to performance.

[0005] The embodiments described herein solve these and other technical problems by using parameterized machine learning pattern recognition techniques to optimize data center configurations. An environmental chamber that cycles through various environmental conditions such as various altitudes and temperature combinations can have various data center configurations placed within it. A stress workload can be run on the data center within the environmental chamber, and a complete telemetry reading can be provided for each set of environmental conditions. The telemetry data can then be used to train a model such as the MSET model or the MSET2 model at each altitude / temperature combination.

[0006] When installing a new data center, a stress workload can be run using the default data center configuration, and telemetry signals can be captured in real time. Different performance metrics, including CPU efficiency, I / O efficiency and throughput, as well as other quality of service (QoS) metrics, can be calculated for the default conditions. If the performance of the default configuration is not satisfactory, the operating conditions of the new data center can be used to select the closest fitting model from a library of pre-trained models. The selected, pre-trained model may have been trained using a configuration that operates under similar environmental conditions within the environmental chamber prior to the installation of the new data center. For example, the altitude and / or temperature of the new data center can be used to select the closest fitting model that was trained at a similar altitude and / or temperature. That model can then be used to determine the optimal configuration to be used for the new data center. Instead of repeatedly performing trial-and-error adjustments to the default configuration, this method provides the optimal configuration using a maximum of two different configuration tests, thereby reducing the time required to set up a new set of configurations from days to hours.

[0007] Brief Description of the Drawings A further understanding of the nature and advantages of the various embodiments can be realized by reference to the remainder of the specification and the drawings, in which like reference numerals are used throughout to refer to like components. In some cases, a sub-label is associated with the reference numeral to indicate one of a plurality of like components. When referring to a reference numeral without specifying an existing sub-label, it is intended to refer to all such like components.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Mode for Carrying Out the Invention

[0009] Detailed Description Optimizing the configuration of modern data centers is becoming an important process for delivering predictable Quality of Service (QoS) to customers of both on-premises and cloud-based data centers. Conventionally, data centers have been configured with little regard for environmental characteristics such as temperature and altitude. A data center is configured to optimize the CPU and I / O performance of the data center so that the configuration meets the customer's Service Level Agreement (SLA). This configuration was then assumed to be valid over a range of environmental conditions. For example, 10 years ago, temperature variations of 20°C to 30°C did not have a significant impact on the performance of the data center. Similarly, altitude variations between sea level and 5,000 feet did not dramatically affect the operation of the data center hardware / software. Generally, the nature of the computer hardware used over the first 25 - 30 years of data center computing was such that CPU and I / O performance were minimally affected by changes in ambient temperature or the altitude of the data center. Therefore, the configurations used to design data centers could be easily implemented in multiple data centers with different operating temperatures and / or altitudes.

[0010] However, the progress of CPU and memory technologies has led to a situation where data center performance now highly depends on temperature and altitude, along with the increasing prevalence of CPU-intensive and I / O-intensive workloads. This dependency can have the greatest impact on modern data center configurations that include memory storage using spinning hard disk drives (HDDs) rather than solid state drives (SSDs) and / or aggressive CPU designs that use dynamic voltage and frequency scaling (DVFS). This impact of temperature and altitude on performance applies to all information technology (IT) systems, regardless of the manufacturer. Temperature and altitude also affect both on-premises data centers and cloud data centers, as well as external storage configurations such as internal HDDs and network attached storage (NAS). Furthermore, this is an issue that will continue to grow in the future as the dependency on ambient temperature / altitude increases exponentially with each new generation of CPUs and HDDs.

[0011] FIG. 1 shows a schematic diagram of a cooling system for a data center according to some embodiments. The majority of server and storage systems use fans 100 to air-cool internal components. Additionally, many storage systems may include an internal power supply with a cooling fan. These fans 100 may have variable speeds to adapt to the variable temperatures within the data center. As the ambient temperature within the data center rises, the fan speed may also increase proportionally to cool the internal components.

[0012] Modern HDDs 102 have been found to be extremely sensitive to even slight vibrations. For example, mechanical and / or acoustic vibrations can significantly reduce the I / O rate of the HDD 102. As the ambient temperature rises and the speed of the fan 100 increases to compensate, the mechanical vibrations from the fan 100 also increase. These mechanical vibrations can interfere with the read / write operations of the HDD 102 and thus potentially degrade the performance of the HDD 102. In addition, the acoustic noise from the fan can also increase as the fan speed increases. This acoustic noise also affects the HDD and further degrades the I / O performance. Therefore, as data centers operate at higher temperatures, the mechanical and acoustic vibrations due to the increase in fan speed can measurably reduce the I / O performance of the HDD 102. Data centers around the world are currently raising the ambient temperature from above 30°C to 35°C in an attempt to save energy used for air conditioning, so this effect resulting from higher temperatures is likely to increase in the future. However, the corresponding degradation in performance at higher temperatures increases the time taken to execute customer workloads, which in turn offsets the potential energy savings from the increase in operating temperature. The I / O performance of the HDD 102 is measurably reduced. Data centers around the world are currently raising the ambient temperature from above 30°C to 35°C in an attempt to save energy used for air conditioning, so this effect resulting from higher temperatures is likely to increase in the future. However, the corresponding degradation in performance at higher temperatures increases the time taken to execute customer workloads, which in turn offsets the potential energy savings from the increase in operating temperature.

[0013] In addition to the effects of temperature rise, it has also been found that altitude changes can cause further degradation of the performance of the HDD 102. For example, data centers at high altitudes are located in areas with thinner air than data centers located at sea level. This thinner air at these higher altitudes has less cooling capacity. This requires the fan 100 to further increase its speed to increase the airflow through the server. This increased fan speed, due to the warmer temperature, can be added to the increases described above and can further generate mechanical and / or acoustic vibrations that interfere with the performance of the HDD 102. For example, the fan typically operates 13% faster in a high-altitude environment (e.g., Denver, Colorado) compared to a similar data center configuration at sea level.

[0014] FIG. 1 also shows a CPU that can be part of a data center configuration. In addition to affecting the I / O performance of various memory components, when using DVFS, an increase in temperature can also have an adverse effect on the performance of CPU 106. Systems that use DVFS generally adjust the power and / or speed settings on the CPU and other peripheral devices to optimize resource consumption. DVFS enables CPU 106 to execute the CPU workload using the minimum amount of power required for the CPU workload, maximizing power savings and improving the lifespan of CPU 106. The system monitors the workload and dynamically adjusts the operating voltage and / or clock frequency to match the performance of the CPU to the requirements of the workload.

[0015] Modern CPUs and graphics processing units (GPUs) use very sensitive and aggressive DVFS schemes to continuously throttle the frequency and / or operating voltage. However, higher temperatures generally cause more leakage current within CPU 106, and thus DVFS frequency throttling increases exponentially with each degree increase in temperature. For example, an increase from 15°C to 35°C shows a degradation of more than 50% in CPU performance as the altitude of the data center increases from 10,000 feet above sea level. This combination of CPU and I / O performance degradation due to temperature and altitude poses a challenging problem for any data center installed in a new operating environment.

[0016] Previously, the process for configuring a data center was an iterative process involving multiple trial-and-error cycles to adjust various aspects of the configuration to meet the SLA requirements of the data center. Once the data center was installed, a default configuration was used, and then the performance of the data center was evaluated to determine whether the default configuration met the SLA requirements. The default configuration was often created in an environment with different operating temperatures and / or altitudes compared to the operating temperature and / or altitude of the installed data center. Due to the effects of the higher temperature and / or altitude described above, the performance of the new data center will often not match the performance of the data center for which the default configuration was originally designed. If the performance of the new data center is unsatisfactory, various parameters of the data center configuration are adjusted, the new configuration is implemented for the data center, and tested against the SLA. This process will be repeated and continued until a satisfactory data center configuration is identified. This requires multiple iterations, which added a significant delay to the setup time for the data center.

[0017] In addition to the multiple iterations required to initially configure a new data center, changes in temperature and altitude after installation and during operation will typically cause a configuration that was initially satisfactory to fail to meet SLA requirements during I / O-intensive or CPU-intensive workloads. Even if the default configuration is ultimately optimized for the new data center, as the temperature rises within the new data center, for the reasons described above, the performance of CPU 106 and the performance of HDD 102 may degrade. This performance degradation causes customer timeout errors, increases the time required to complete the customer's workload, and as a result, causes customer dissatisfaction and a progressive increase in incidents if the SLA requirements are not met. Diagnosing the root cause of this performance degradation is a very difficult process, especially when caused by fluctuations in temperature and / or altitude differences. To resolve this situation after a new data center has been configured, data center administrators often had to provide a completely new configuration. In that case, the new configuration will often face the same timeout errors, performance degradation, and consistency issues observed in the original configuration of the data center.

[0018] The embodiments described herein solve these and other technical problems by determining an optimal data center configuration, using an automated configuration optimization framework, based on the workloads of memory, CPU, and I / O performance that may be required for any customer use case scenario. These embodiments may utilize an environmental chamber to control the environmental characteristics of various different data center configurations. For example, a thermal / altitude chamber may be used to cycle through a wide range of temperatures and / or altitudes so that the performance of a data center configuration can be measured against different combinations of environmental characteristics. Since various data center configurations are tested with different combinations of temperature and altitude, a library of models may be trained corresponding to each altitude / temperature combination. After the library of models is trained, this library may be used to determine the optimal configuration of a new data center when the new data center is installed and configured. The environmental characteristics of the installed data center may be used to select the nearest neighbor model within the library. Then, similar operating requirements may be provided to the model to predict various performance metrics such as CPU performance, memory I / O performance, etc. If the performance metrics generated by the model are better than the performance of the installed data center, the data center may be configured using the configuration provided by the model. This may require installing and testing up to two different configurations for the data center to ensure that the data center is optimally configured.

[0019] The embodiments described herein are compatible with any type of data center. For example, some data centers may include on-premises data centers that are privately owned and / or managed by a customer and are typically installed at the customer's location. Some data centers may also include cloud-based data centers that are hosted by a service provider and include centrally located hardware / software that is available for customer use. Some data centers may include internal storage where disk drives are located internally on a server. Other data centers may include external data storage devices communicatively coupled to a server. Object-based storage may also be used by some data centers to manage data as objects. The configuration methods described herein are compatible with any of these types of data centers or storage technologies.

[0020] FIG. 2 shows a simplified diagram of an object storage system 200 according to some embodiments. This object storage system 200 is provided as just one example of a type of data center that may use the autonomous configuration optimization techniques described herein. However, this object storage system 200 is not intended to be limiting, and these techniques may be used with any type of data center. Those skilled in the art will recognize that these techniques are not limited by data center type and, instead, may be applied to any data center in light of the present disclosure.

[0021] Customers may use the storage cloud service as a storage solution to store their data / files in the public cloud. The object storage system 200 may include an on-premises server 204 that contains customer data stored in a file directory 206. The server 204 may be communicatively coupled to additional server instances 202a, 220b as part of a data center that uses the object storage system 200. The server 204 may include a cloud storage software application (CSSA) installed at the customer site. The CSSA may function as a cloud storage gateway that connects on-premises applications and workflows to the public cloud. For example, the CSSA may be implemented using the Oracle Cloud Infrastructure Storage Software Appliance (registered trademark). The CSSA may communicate with a storage cloud service 212 in the public cloud where the customer's object data may be stored. The CSSA may manage I / O traffic between the customer's data and / or applications and the public cloud storage device.

[0022] An efficient configuration of the CSSA can be important for achieving maximum storage I / O throughput and / or computing performance. Any sub-optimal configuration of the CSSA can result in unsatisfactory functionality of the CSSA characterized by low throughput or computing performance. This can lead to a "timeout" failure that is immediately recognizable by the customer. Low performance due to the CSSA configuration can result in a violation of the SLA requirements provided by the cloud service. This can be particularly true for CPU-intensive and I / O-intensive workloads.

[0023] The CSSA may include one or more configuration parameters 210. These parameters should be configured to meet the customer's performance specifications. These performance specifications are often expressed as I / O throughput requirements based on, for example, the number of concurrent users, the number and / or size of stored files, the directory hierarchy structure, the number of I / O operations (reads, writes, appends, etc.). The performance of the data center in meeting these performance specifications may be managed by the configurable parameters 210 of the CSSA. These are parameters that can be adjusted to address the customer's performance requirements. Configuration parameters may include any hardware and / or software parameters that define the setup and / or operation of the data center. Configuration parameters may include the number of CPUs in the data center; the speed and / or type of CPUs in the data center; the number of HDDs used in the data center; the type of hardware used in the data center, including network cards, disk controllers, and databases; cache size; and / or any other parameters. Configuration parameters may also include software settings used to control the operation of the hardware. As used herein, configuration or configuration parameters may include any aspect of the data center, as understood by one of ordinary skill in the art.

[0024] Similar to any data center, the object storage system 200 may provide many different configurations that may be optimized during installation and operation. The interaction between the customer workload mix and the CSSA can often be very complex and recommending an optimal configuration solution can be the key to achieving satisfactory performance. Any sub-optimal combination of the configurations described above is likely to result in long latencies, low throughput, multiple I / O timeouts, and / or SLA violations. As described above, for the object storage system 200 or any other data center recommending an optimal configuration is becoming increasingly difficult as the dependence of performance on altitude and temperature increases.

[0025] To quickly and efficiently provide an optimal configuration for a data center, the embodiments described herein may use a two-stage process. In the first stage, a library of models may be trained for various configurations at many different combinations of temperature and altitude. This library of pre-trained models may then be provided to a second stage that is used when a new or existing data center is installed or reconfigured. The first stage for training the library of models may be performed offline, regardless of a particular data center installation for a customer. Alternatively, the models may be trained using an environmental chamber that circulates through various combinations of temperature and altitude for different data center configurations. This process may be performed without the pressure of any particular data center installation for a customer and, thus, may develop a very comprehensive and robust library of trained models over time. During the second stage, access to this library may be made and used to recommend an optimal configuration for a given data center in real time without significant delay.

[0026] FIG. 3 illustrates a system that may be used to train a library of models at various temperatures and / or altitudes, according to some embodiments. The system may include an environmental chamber 300 configured to adjust and control one or more environmental characteristics of an operating environment for a data center configuration under evaluation. Many different environmental conditions, such as temperature, altitude, humidity, pressure, etc., may be monitored and / or controlled by the environmental chamber 300. As an example, embodiments described herein may use a thermal-altitude chamber that is configured to control the ambient temperature inside the environmental chamber 300 and simulate various altitudes within the environmental chamber 300. For the reasons of performance discussed above, temperature and altitude may be the environmental characteristics most relevant to data center performance, but this is not meant to be limiting. Other embodiments may use other environmental characteristics when training, monitoring, and using models for a data center.

[0027] Inside the environmental chamber 300, various configurations of different data centers may be installed. For example, the data center configuration 300 within the environmental chamber 300 depicted in FIG. 3 may include several servers, HDDs, and / or other hardware that may be used to build a data center. Configuration 300 may also include various configuration parameters that affect or control the operation of the data center. In some embodiments, the types of data center configurations may be divided into a few simplified categories, such as low, medium, and / or high-performance data centers. For example, a first (e.g., low-performance) category of data center may include a limited number of CPUs and / or HDDs. A second (e.g., medium-performance) category of data center may include a greater number of multi-core CPUs and / or GPUs, along with a greater number of HDDs. A third (e.g., high-performance) category of data center may be referred to as an "engineered system" and may use the fastest CPUs / GPUs and provide many petabytes of memory. Within each of these categories of data centers, additional configuration parameters may vary the number of hardware components and / or software used within the data center.

[0028] When the data center configuration 300 is installed in the environmental chamber 300, the controller 302 may provide the multivariable workload script 304 to the data center configuration 310. The data center may then execute the multivariable workload script 304 within the environmental chamber 300. As described below, the workload script 304 may be executed with different combinations of environmental characteristics such as temperature and altitude. The multivariable workload script 304 may include extreme stress workloads that ramp the CPU workload from idle to maximum load and sinusoidally cycle over time to fully evaluate the range of CPU performance. In addition Further, the multivariate workload script 304 may be configured to similarly cycle the operation of the HDD from idle to maximum load. The multivariate workload script 304 may be designed to cover a range of various workloads that may be experienced by a data center in various customer installations.

[0029] The controller 302 may include a circulation algorithm 308 that circulates the environmental chamber 300 through various permutations of different environmental characteristics. In this example, the circulation algorithm may cycle through a range of temperatures in selected increments of simulated altitude. In some embodiments, this process may be performed over multiple time intervals (e.g., over 24 hours).

[0030] At each different combination of temperature and altitude, the multivariate workload script 304 may be executed by a data center within the environmental chamber 300, and telemetry data from the configuration may be collected by the controller 302. In some embodiments, a telemetry harness such as Oracle's Continuous System Telemetry Harness (CSTH) is used When used, real-time telemetry data may be collected when the data center executes the multivariate workload script 304. The telemetry data may include any electrical and / or environmental characteristics of the operating data center, including distributed temperature, ambient temperature, power usage, and the like. The telemetry data may also generate performance metrics for the data center. For example, the performance metrics may include the performance metrics of each CPU during configuration 310. For each CPU, the telemetry harness 312 may capture data indicating the percentage of the maximum performance of that CPU. (This may also indicate the reciprocal of the amount of throttling that occurs.). For each HDD, performance metrics may be generated that represent the I / O rate recorded as a percentage of the maximum I / O rate. The HDD performance metrics may also include an indication of the number of Mb / s of I / O throughput. These performance metrics may be used to determine whether the data center configuration 310 meets the customer's SLA requirements for each combination of temperature and altitude.

[0031] The telemetry data captured by the telemetry harness 312 may be provided to a model training process 306 that uses the telemetry data to train a model specific to the current combination of temperature and altitude. The model training process 306 may load models corresponding to each increment of temperature and / or altitude. The telemetry data received when the data center executes the multivariate workload script 304 at that combination of temperature and altitude may be provided to the controller 302 such that each combination of temperature / altitude generates a complete telemetry set of signals for training the corresponding model. After a complete cycle through each temperature and altitude increment has been executed, the controller 302 may have a complete telemetry time series signal library over all combinations of temperature and altitude.

[0032] Figure 4 illustrates a flowchart of a circulation algorithm 308 for controlling the time and / or altitude increments of an environmental chamber, according to some embodiments. The algorithm may begin by initializing the simulated altitude within the environmental chamber to an initial value (402). Throughout the present disclosure, mathematical representations of various altitudes (both simulated and measured) may be referred to using the variable H. The initial altitude (H L ) may be set to any value. In some embodiments, the initial altitude may be set to sea level (i.e., 0 feet altitude). Next, the algorithm may determine whether the current simulated altitude H is less than the maximum simulated altitude (404). The maximum altitude (H H ) may also be set to any value, including 5000 feet, 7000 feet, 10,000 feet, etc.

[0033] Next, the algorithm may initialize the temperature within the environmental chamber to an initial value (410). Throughout the present disclosure, mathematical representations of temperature may be referred to using the variable T. The initial temperature (T L ) may be set to any value. In some embodiments, the initial temperature may be set to 0°C, 5°C, 15°C, 20°C, etc. Next, the algorithm may determine whether the current temperature T is less than the maximum temperature (412). The maximum temperature (T H ) may be set to any value, such as 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, etc.

[0034] Next, the algorithm may cause the data center to execute a multivariate workload script at the specified temperature and altitude combination. In some embodiments, the algorithm may wait for the altitude and / or temperature within the environmental chamber to stabilize as these values are initially set or incremented. Once the multivariate workload script is executed at the specified altitude and temperature, the algorithm may collect a complete set of telemetry for the signals at that altitude and temperature to be used to train the corresponding model, as described in more detail below (418).

[0035] After the script is executed and telemetry signals are captured at the current temperature / altitude combination, the algorithm may then increment the temperature (416). The temperature may be incremented by any value such as 0.5 °C, 1.0 °C, 2.0 °C, 2.5 °C, 5.0 °C, etc. The algorithm may then determine whether the incremented temperature is higher than the maximum temperature at which this data center configuration should be tested. If the maximum temperature has not yet been reached, the algorithm may execute the multivariate workload script again (414) and collect the complete set of telemetry for the signals at that new temperature / altitude combination (418). This cycle may continue collecting telemetry signals over the temperature range at the current altitude until the maximum temperature is reached.

[0036] After reaching the maximum temperature (412), the algorithm may then increment the altitude simulated by the environmental chamber (406). The altitude may be incremented by any value such as 100 feet, 200 feet, 250 feet, 500 feet, 1000 feet, etc. After the altitude is incremented, the algorithm may determine whether the new altitude is greater than the maximum altitude at which this data center configuration is to be evaluated (404). If the current altitude is still less than the maximum altitude, the algorithm may increment over the full range of temperatures at the new altitude as described above. However, when the maximum altitude is reached, the temperature / altitude characterization for a particular data center configuration may be complete. At this point, the system may provide a complete library of telemetry time series signals over all combinations of temperature and altitude for training the model at each of those combinations.

[0037] Figure 5 illustrates a process for training a model with each temperature / altitude combination, according to some embodiments. A set of telemetry data 502 collected during the execution of the multivariate script at each time / altitude combination may be provided to process 306 to train the model. Different embodiments may use different models. For example, some embodiments may use a pattern recognition model such as a multivariate state estimation technique (MSET) modeling technique, an MSET2 modeling technique, or any other non-linear non-parametric modeling method. The model may be trained using any of the telemetry data and learn the correlation patterns among all of the configuration parameters that affect the performance of the data center. For example, MSET2 (which has been used for prediction and cyber security use cases in other applications) may learn the interaction of independent variables such as CPU performance and I / O throughput as a function of the environmental conditions of a particular data center configuration. Additionally, at each temperature / altitude combination, multiple configurations may be tested to train the model. For example, the process shown in Figure 3 characterizes a single configuration, but in practice multiple configurations may be characterized, and each configuration provides a telemetry signal throughout the full range of temperature / altitude combinations shown in Figure 4.

[0038] The process for training a library of models can be a continuous cumulative calculation that is performed offline using data accumulated from various data centers. The above example depends on telemetry data provided from an environmental chamber in a controlled characterization environment, but this is not intended to be limiting. Other embodiments may use telemetry data accumulated from other data center environments, including live data center installations used by customers. The result of this offline training process is a library of pre-trained models that can be used to generate performance metrics such as quality of service (QoS) metrics for different configurations at a given temperature / altitude combination. The library of pre-trained models may also be used to generate recommended configurations.

[0039] The second stage of the autonomous configuration process may include real-time recommendations. When a new data center is installed for a customer or cloud environment, the pattern recognition algorithm may transition from the training mode to the real-time "recommendation" mode. For example, a particular MSET model 506 closest to the measured temperature / altitude combination of the data center may be selected, and the model 506 may be used to output a new configuration to be used for a given environmental condition based on all other correlated variables. For example, the prediction method may provide a set 508 of data center requirements for the new data center. The telemetry input may be provided to the MSET model 506 to generate performance metrics such as QoS metrics at a given temperature / altitude combination. These QoS metrics may then be compared to the measured QoS metrics of the current configuration of the data center, and if the current configuration is not optimal, a recommendation 510 for a new configuration may be provided.

[0040] This predictive machine learning technology may use a model trained with a data center configuration at various temperatures / altitudes. When a customer use case is provided to this method as an inquiry, the method may recommend an optimal data center configuration option for the customer, using a library of pre-trained MSET models, based on the customer's workload requirements and the environmental conditions measured at the new data center installation. In some embodiments, a technician installing a data center may access a library of pre-trained models using an application or "app" on a smartphone, tablet computer, or other computing device. The app may receive performance data including the measured temperature / altitude from the data center and may identify the closest matching model within the model library. Once the model is identified, MSET may be used to output the lowest cost configuration that provides optimal performance metrics meeting the SLA requirements. Some embodiments may also provide the maximum performance achievable given the altitude and ambient temperature of the data center. Conversely, some embodiments may provide a set of recommended environmental conditions given the data center workload and SLA requirements.

[0041] Figure 6 shows a model library 602 that may be generated from the model training process 306, according to some embodiments. As described above, the model library 602 may include entries for each combination of temperature and altitude generated by the temperature / altitude cycling algorithm of FIG. 4. The model 506 may be trained at each temperature / altitude combination using a plurality of different data center configurations in the environmental chamber illustrated in FIG. 3. The number of models 506 populating the model library 602 may depend on the range of temperature / altitude generated by the cycling algorithm in FIG. 4. The number of models 506 may also depend on the temperature / altitude increment used between each execution of the workload script by the data center. The smaller the temperature / altitude increment, the finer the granularity of the model library 602 This may be the case. This increases the likelihood that a new data center installation can be closely matched with a pre-trained model within the model library 602.

[0042] It should be emphasized that this second stage of the process for configuring a new data center installation may be executed in a matter of hours or even minutes. Since the model library is pre-populated with pre-trained models, the model inputs from the new data center installation can be quickly provided to the existing models to generate corresponding performance metrics that may be compared with the performance metrics of the current configuration of the new data center. If the closest model generates better QoS results than the current configuration, the configuration provided by the model may be used instead of the current configuration of the new data center installation. This may be compared with previous processes used to configure a data center. Generally, there has been no systematic parameterized method for determining whether the current configuration was optimal. Instead, the various parameters of the data center configuration would have been adjusted until the SLA requirements could be met. This was often a trial-and-error process that required many iterations and could take days or weeks. This autonomous data center configuration process described herein dramatically reduces this time by limiting the number of different configurations that need to be tested against the SLA requirements to a maximum of two.

[0043] FIG. 7 illustrates a flowchart 700 of a method for autonomously determining an optimal configuration for a data center using a library of models pre-trained with various combinations of environmental characteristics, according to some embodiments. The method may include receiving or measuring a first performance metric from the data center (702). The data center may include a new data center installation or an existing data center whose configuration has been updated. For example, an existing data center may begin to fail to meet SLA requirements at a higher workload or various temperatures, and this method may be executed to reconfigure this data center in a more optimal way. The data center may have an existing configuration, referred to herein as the "first configuration," to distinguish this configuration from other configurations. Subsequent configurations may simply be referred to as the "second configuration" to distinguish it from the first configuration. The terms first / second do not mean to indicate an order, priority, importance, or any other comparative quality between the configurations. Similarly, the "first performance metric" may be so named to distinguish the performance metric generated by the data center from a "second" performance metric generated by one of the pre-trained models. The first performance metric may include any QoS metric such as CPU performance as a percentage of maximum performance, I / O throughput, I / O rate, latency, etc. The first performance metric may be measured and / or generated in real time when the data center executes a workload such as the multivariate workload script described above.

[0044] Optionally, the method may include determining (704) whether a first performance metric meets the SLA requirements. For example, the number of timeout errors may be compared to the SLA requirements. The measured bandwidth or latency may be compared to the SLA requirements. The CPU performance or power consumption may be compared to the SLA requirements. In some cases, if the first performance metric meets or exceeds the SLA requirements, the method may simply determine that the current configuration of the data center is sufficient, and the data center may operate using the current configuration (716). Note that generally, there will be multiple SLA requirements and multiple performance metrics to be compared to those SLA requirements. Although this method describes only a single performance metric and SLA requirement as an example, it should be understood that other embodiments may require each performance metric to meet the corresponding SLA requirement before the data center can continue to operate in its current configuration without any adjustment.

[0045] The method may further include receiving or measuring one or more environmental characteristics of the data center when the data center executes a workload (706). A telemetry system similar to that described above in connection with FIG. 3 may be used in the data center to collect telemetry data during the execution of a multivariate workload script. The telemetry data may include environmental conditions such as temperature and / or altitude.

[0046] The method may further include identifying pre-trained models in the model library with similar environmental characteristics (708). The models may be pre-trained using data from one or more data center configurations operating in a controlled environment such as the environmental chamber described above in connection with FIG. 3. The method may receive environmental characteristics and autonomously select a corresponding model trained with a similar temperature / altitude combination from the model library. In some cases, the current environmental conditions of the data center may match the environmental conditions of the pre-trained model. For example, in the case of a data center operating at 25° C. at the interface level, the model library may include corresponding models trained with those corresponding environmental conditions. However, it is often possible that the current operating conditions of the data center may fall somewhere between the available temperature / altitude combinations in the model library. If an exact match is not available, the method may select a model having environmental characteristics most similar to the environmental characteristics of the data center. For example, some embodiments may use a nearest neighbor algorithm to select a pre-trained model, as described in detail below in connection with FIG. 8.

[0047] The method may also include generating a performance metric for the pre-trained model (710). This performance metric may be referred to herein as the "second performance metric" to distinguish it from the first performance metric described above. This performance metric may be of the same type as the first performance metric so that they may be compared. This performance metric may include any QoS metric described herein. Again, the method may include generating multiple performance metrics, each of which may be compared to a performance metric from the data center configuration and each of which may be required to meet SLA requirements. This step of the method may be repeated for each performance metric under consideration.

[0048] This method may further include comparing performance metrics from the data center configuration with performance metrics generated by a pre-trained model (712). If the performance metrics from the data center are better than the performance metrics from the model, the configuration of the data center may be considered optimal, and the data center may continue to operate with the existing configuration (716). However, if the performance metrics from the pre-trained model are better than the performance metrics from the data center configuration, the method may include determining a new configuration for the data center from the model and using the configuration from the model for the installed data center (714). Since the model was trained using various different configuration parameters, the model may output a data center configuration that produces the best QoS results for its altitude / temperature combination.

[0049] If a more optimal data center configuration is generated by the model, the new data center configuration can be implemented in the data center (716). The new configuration can then be tested to ensure that it meets the SLA requirements, and the configuration process may be completed. Note that only two different configurations may need to be tested to determine the optimal configuration. If the first configuration is better than the configuration generated by the corresponding pre-trained model, the first configuration may be considered optimal. However, if a second configuration generated by the pre-trained model produces better QoS metrics, the second configuration may be considered optimal without any further configuration testing.

[0050] ​It should be understood that the specific steps shown in FIG. 7 provide a specific way to optimize the configuration of a data center according to various embodiments. Sequences of other steps may also be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Further, the individual steps shown in FIG. 7 may include multiple sub-steps that may be performed in various sequences appropriate for the individual steps. Additionally, additional steps may be added or removed depending on the specific application. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0051] FIG. 8 shows a flowchart 800 of a method for selecting a pre-trained model having environmental characteristics similar to those of a data center according to some embodiments. As described above, an exact match between the environmental characteristics of a data center and those of a pre-trained model may not always be available. Depending on the temperature / altitude increments used in generating the model library, some data centers may operate at a temperature / altitude between the temperatures / altitudes corresponding to the pre-trained models. Accordingly, some embodiments may select a model by identifying the model having the most similar environment of characteristics. This may include using a nearest neighbor algorithm to identify the closest temperature / altitude match within the model library. The method of flowchart 800 shows one exemplary way of identifying a nearest neighbor model.

[0052] This method may include receiving (802) the ambient temperature and altitude around the installed data center. The temperature / altitude may be received as part of the telemetry data used when testing the first data center configured as described above. The method may also include accessing (804) the temperature and altitude used to generate the model library. For example, a sequence of M temperature values may be combined with a sequence of N altitude values to index an entry in the model library. The method may then include calculating a distance, such as the Euclidean distance, between each combination of M temperature values and N altitude values in the library and the ambient temperature and altitude of the data center (806). In some embodiments, the method may first calculate and minimize the distance between altitudes and then minimize the distance between ambient temperatures. The method may then identify (808) the temperature / altitude pair in the library having the minimum distance from the temperature / altitude of the data center. This minimization of the distance may be calculated using the following formula:

[0053] [Number]

[0054] In this formula, T0 and H0 represent the ambient temperature and altitude of the data center, and T i and H j represent the ranges of temperature values and altitude values in the model library. When the values of T i and H j that minimize this equation are identified, these values may be used to index the model library and retrieve the corresponding nearest neighbor model to be used.

[0055] The embodiments disclosed in this specification may use various computing environments, including networked computing environments, cloud service environments, and computer architectures with processors, memories, instruction sets, and the like. The following figures illustrate hardware / software systems that may be used to implement the various embodiments described in this specification. These hardware / software systems may also be used to interact with the embodiments described in this specification.

[0056] In addition, each step of these methods may be automatically executed by a computer system and / or input / output involving the user may be provided. For example, the user may provide input for each step of a method, and each of these inputs may respond to a specific output that requires such input, and that output is generated by the computer system. Each input may be received in response to the corresponding required output. Further, the input may be received from the user, as a data stream from another computer system, retrieved from a memory location, retrieved via a network, or requested from a web service. Similarly, the output may be provided to the user, as a data stream to another computer system, stored in a memory location, transmitted via a network, or provided to a web service. In short, each step of the methods described herein may be implemented by a computer system, with or without user involvement, and may include any number of inputs, outputs, and / or requests to and from the computer system. Steps without user involvement may be said to be automatically executed by the computer system without human intervention. Thus, in light of the present disclosure, it will be understood that each step of each method described herein may be modified to include input from the user and output to the user, or may be automatically performed by the computer system without human intervention if any decision is made by the processor. Further, some embodiments of each of the methods described herein may be implemented as a set of instructions stored on a tangible non-transitory storage medium to form a tangible software product.

[0057] FIG. 9 shows a simplified diagram of a distributed system 900 for implementing one embodiment. In the illustrated embodiment, the distributed system 900 includes one or more client computing devices 902, 904, 906, and 908, which are configured to execute and operate client applications such as web browsers, proprietary clients (e.g., Oracle Forms), etc. via one or more networks 910. The server 912 may be communicatively coupled to the remote client computing devices 902, 904, 906, and 908 via the network 910.

[0058] In various embodiments, the server 912 may be adapted to execute one or more services or software applications provided by one or more of the components of the system. In some embodiments, these services may be provided to the users of the client computing devices 902, 904, 906, and / or 908 as web-based services or cloud services, or under a software-as-a-service (SaaS) model. The users operating the client computing devices 902, 904, 906, and / or 908 may then interact with the server 912 using one or more client applications to utilize the services provided by these components.

[0059] In the configuration shown in the figure, software components 918, 920, and 922 of system 900 are shown as being implemented on server 912. In other embodiments, one or more of the components of system 900 and / or the services provided by these components may be implemented by one or more of client computing devices 902, 904, 906, and / or 908. A user operating a client computing device may then use one or more client applications to utilize the services provided by these components. These components may be implemented in hardware, firmware, software, or a combination thereof. It should be understood that various different system configurations are possible that are different from the distributed system 900. It should be understood that various different system configurations are possible. The embodiment shown in the figure is, therefore, an example of a distributed system for implementing the system of the embodiment and is not intended to be limiting.

[0060] Client computing devices 902, 904, 906, and / or 908 may be portable handheld devices (e.g., iPhone (registered trademark), cellular phone, iPad (registered trademark), computing tablet, personal digital assistant (PDA)) or may be wearable devices (e.g., Google Glass (registered trademark) head-mounted display), software such as Microsoft Windows Mobile (registered trademark), and / or various operating systems such as iOS, Windows Phone, Android, BlackBerry 10, Palm OS, etc. It runs a byte operating system and is an effective communication protocol such as the Internet, email, Short Message Service (SMS), Blackberry (registered trademark), or others. The client computing device can be a general-purpose personal computer, for example, various versions of Microsoft Windows (registered trademark), Apple Macintosh (registered trademark), and / or Linux (registered trademark) operating systems, and / or a personal computer and / or laptop computer that runs them. The client computing device can be a workstation computer that runs any of various commercially available UNIX (registered trademark) or UNIX-like operating systems, including but not limited to various GNU / Linux (registered trademark) operating systems such as Google (registered trademark) Chrome OS. Alternatively, or in addition, the client computing devices 902, 904, 906, and 908 can be any other electronic device such as a thin client computer, an Internet-enabled game system (e.g., a Microsoft Xbox game console with or without a Kinect (registered trademark) gesture input device), and / or a personal messaging device that can communicate via the network 910. including personal computers and / or laptop computers running various versions thereof, and / or a workstation computer running any of various commercially available UNIX (registered trademark) or UNIX-like operating systems, including but not limited to various GNU / Linux (registered trademark) operating systems such as Google (registered trademark) Chrome OS. Alternatively, or in addition, the client computing devices 902, 904, 906, and 908 can be any other electronic device such as a thin client computer, an Internet-enabled game system (e.g., a Microsoft Xbox game console with or without a Kinect (registered trademark) gesture input device), and / or a personal messaging device that can communicate via the network 910.

[0061] The exemplary distributed system 900 is shown with four client computing devices, but any number of client computing devices may be supported. Other devices, such as devices with sensors, may interact with the server 912.

[0062] The network 910 within the distributed system 900 includes, but is not limited to, various ones such as TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Packet Exchange), AppleTalk, etc. It may be any type of network well known to those skilled in the art that can support data communication using any of the commercially available protocols. By way of example only, network 910 may be a local area network (LAN) such as one based on Ethernet®, Token Ring, etc. Network 910 may be a wide area network and the Internet. It may be a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a network operating under any of the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite, Bluetooth®, and / or any other wireless protocol), including but not limited to virtual networks; and / or any combination of these and / or other networks.

[0063] Server 912 may be composed of one or more general-purpose computers, dedicated server computers (including, for example, PC (personal computer) servers, UNIX® servers, midrange servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or any other suitable configuration and / or combination. In various embodiments, server 912 may be adapted to execute one or more of the services or software applications described in the foregoing disclosure. For example, server 912 may correspond to a server for executing the processes described above in accordance with certain embodiments of the present disclosure.

[0064] Server 912 may execute an operating system including any of the above, and any server operating system available on the market. Server 912 may also execute any of a variety of other server applications and / or middle-tier applications, including an HTTP (Hypertext Transfer Protocol) server, an FTP (File Transfer Protocol) server, a CGI (Common Gateway Interface) server, a JAVA (registered trademark) server, a database server, etc. Exemplary database servers include, but are not limited to, those available on the market from Oracle, Microsoft, Sybase, IBM (registered trademark) (International Business Machines), etc.

[0065] In some embodiments, Server 912 may include one or more applications for analyzing and collating data feeds and / or event updates received from users of client computing devices 902, 904, 906, and 908. As an example, the data feeds and / or event updates may include real-time events related to sensor data applications, financial stock market dashboards, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc., received from one or more third-party information sources and continuous data streams, and may include Twitter (registered trademark) feeds, Facebook (registered trademark) updates or real-time updates, but are not limited thereto. Server 912 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 902, 904, 906, and 908. It is not limited thereto. Server 912 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 902, 904, 906, and 908.

[0066] The distributed system 900 may also include one or more databases 914 and 916. The databases 914 and 916 may be located in various locations. By way of example, one or more of the databases 914 and 916 may reside on a non-transitory storage medium that is local to (and / or resident on) the server 912. Alternatively, the databases 914 and 916 may be remote from the server 912 and communicate with the server 912 via a network-based connection or a dedicated connection. In a set of embodiments, the databases 914 and 916 may reside within a storage area network (SAN). Similarly, any necessary files for performing functions attributable to the server 912 may be stored locally on the server 912 and / or remotely, as appropriate. In a set of embodiments, the databases 914 and 916 may include a relational database, such as a database provided by Oracle, that is adapted to store, update, and retrieve data in response to SQL-formatted commands.

[0067] FIG. 10 is a simplified block diagram of one or more components of a system environment 1000 that may provide services provided by one or more components of an embodiment system as cloud services, according to one embodiment of the present disclosure. In the illustrated embodiment, the system environment 1000 includes one or more client computing devices 1004, 1006, and 1008 that may be used by a user to interact with a cloud infrastructure system 1002 that provides cloud services. The client computing devices may be configured to operate a client application, such as a web browser, a client application under intellectual property rights (e.g., Oracle Forms), or some other application, that may be used by a user of the client computing device to interact with the cloud infrastructure system 10 02 to use the services provided by the cloud infrastructure system 1002.

[0068] It should be understood that the illustrated cloud infrastructure system 1002 may have components other than those illustrated. Further, the embodiments shown in the figures are merely examples of cloud infrastructure systems that may incorporate embodiments of the present invention. In some other embodiments, the cloud infrastructure system 1002 may have more or fewer components than those shown in the figures, may combine two or more components, or may have different configurations or arrangements of components.

[0069] Client computing devices 1004, 1006, and 1008 may be devices similar to those described above for 902, 904, 906, and 908.

[0070] The exemplary system environment 1000 is shown with three client computing devices, but any number of client computing devices may be supported. Other devices, such as devices with sensors, may interact with the cloud infrastructure system 1002.

[0071] Network 1010 may facilitate data communication and exchange between clients 1004, 1006, and 1008 and the cloud infrastructure system 1002. Each network may be any type of network well known to those skilled in the art that can support data communication using any of a variety of commercially available protocols, including those described above for network 910.

[0072] The cloud infrastructure system 1002 may include one or more computers and / or servers that may include those described above for server 912.

[0073] In one embodiment, the services provided by a cloud infrastructure system may include hosting services that are made available on demand to users of the cloud infrastructure system, such as online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database processing, managed technical support services, etc. The services provided by the cloud infrastructure system can scale dynamically to meet the needs of its users. A particular instantiation of a service provided by the cloud infrastructure system is referred to herein as a "service instance". Generally, any service made available to a user via a communication network such as the Internet from a cloud service provider's system is referred to as a "cloud service". Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own on-premises servers and systems. For example, the cloud service provider's system may host an application, and the user may order and use that application on demand via a communication network such as the Internet.

[0074] In some examples, services within a computer network cloud infrastructure may include storage, hosted databases, hosted web servers, software It may include protected computer network access to an air application, or other services provided to a user by a cloud vendor or known in the art. For example, the service can include password-protected access to remote storage on the cloud via the Internet. As another example, the service can include a web-service-based hosted relational database and a scripting language middleware engine for private use by networked developers. As another example, the service can include access to an email software application hosted on a cloud vendor's website.

[0075] In one embodiment, the cloud infrastructure system 1002 may include a set of application, middleware, and database service offerings that are self-service, subscription-based, elastically scalable, reliable, highly available, and delivered to customers in a secure-protected manner. An example of such a cloud infrastructure system is the Oracle Public Cloud provided by the assignee.

[0076] In various embodiments, the cloud infrastructure system 1002 may be adapted to automatically provision, manage, and track customer subscriptions to services provided by the cloud infrastructure system 1002. The cloud infrastructure system 1002 may provide cloud services via different deployment models. For example, the service may be provided under a public cloud model where the cloud infrastructure system 1002 sells cloud services (e.g., owned by Oracle) owned by an organization and the service is made available to the general public or different industry enterprises. As another example, the service may be provided under a private cloud model where the cloud infrastructure system 1002 operates only for a single organization and provides the service to one or more entities within that organization. The cloud service may also be provided under a community cloud model where the cloud infrastructure system 1002 and the services provided by the cloud infrastructure system 1002 are shared by several organizations within the relevant community. The cloud service may also be provided under a hybrid cloud model that is a combination of two or more different models.

[0077] In some embodiments, the services provided by the cloud infrastructure system 1002 may include one or more services provided under a service as software (SaaS) category, a platform as a service (PaaS) category, an infrastructure as a service (IaaS) category, or other categories of services including hybrid services. A customer may order one or more services provided by the cloud infrastructure system 1002 via a subscription order. The cloud infrastructure system 1002 then executes a process for providing the services in the customer's subscription order.

[0078] In some embodiments, the services provided by the cloud infrastructure system 1002 may include, but are not limited to, application services, platform services, and infrastructure services. In some examples, the application services may be provided by the cloud infrastructure system via a SaaS platform. The SaaS platform may be configured to provide cloud services corresponding to the SaaS category. For example, the SaaS platform may provide the ability to build and deliver a suite of on-demand applications on an integrated development and deployment platform. The SaaS platform may manage and control the underlying software and infrastructure for providing the SaaS services. By utilizing the services provided by the SaaS platform, customers can utilize applications running on the cloud infrastructure system. Customers can obtain application services without the need to purchase separate licenses and support. Various different SaaS services may be provided. Examples include, but are not limited to, services that provide solutions for sales performance management, enterprise integration, and business agility for large organizations. Infrastructure may be managed and controlled. By utilizing the services provided by the SaaS platform, customers can utilize applications running on the cloud infrastructure system. Customers can obtain application services without the need to purchase separate licenses and support. Various different SaaS services may be provided. Examples include, but are not limited to, services that provide solutions for sales performance management, enterprise integration, and business agility for large organizations.

[0079] In some embodiments, the platform service may be provided by a cloud infrastructure system via a PaaS platform. The PaaS platform may be configured to provide cloud services corresponding to the PaaS category. Examples of platform services may include, but are not limited to, services that enable an organization (such as Oracle) to integrate existing applications on a shared common architecture, and the ability to build new applications that utilize shared services provided by the platform. The SaaS platform may manage and control the underlying software and infrastructure for providing SaaS services. Customers can obtain PaaS services provided by the cloud infrastructure system without the need for the customers to separately purchase licenses and support. Examples of platform services include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), etc.

[0080] By using the services provided by the PaaS platform, customers can adopt programming languages and tools supported by the cloud infrastructure system and can also control the deployed services. In some embodiments, the platform services provided by the cloud infrastructure system may include database cloud services, middleware cloud services (e.g., Oracle Fusion Middleware services), and Java cloud services. In one embodiment, the database cloud service may support a shared service deployment model that enables organizations to pool database resources and provide databases as services to customers in the form of a database cloud. The middleware cloud service may provide a platform for customers to develop and deploy various business applications, and the Java cloud service may provide a platform for customers to deploy Java applications in the cloud infrastructure system.

[0081] Various different infrastructure services may be provided by the IaaS platform in the cloud infrastructure system. The infrastructure services facilitate the management and control of underlying computing resources such as storage, network, and other basic computing resources for customers who use the services provided by the SaaS platform and the PaaS platform.

[0082] In certain embodiments, the cloud infrastructure system 1002 may also include infrastructure resources 1030 for providing resources used to provide various services to customers of the cloud infrastructure system. In one embodiment, the infrastructure resources 1030 may include a pre-integrated and optimized combination of hardware such as servers, storage, and networking resources to execute services provided by the PaaS platform and the SaaS platform.

[0083] In some embodiments, the resources in the cloud infrastructure system 1002 may be shared by multiple users and dynamically reallocated as needed. Additionally, the resources may be allocated to users at different times. For example, the cloud infrastructure system 1030 may enable a first set of users in a first time period to utilize the resources of the cloud infrastructure system for a specified number of hours and then reallocate the same resources to another set of users in a different time period, thereby maximizing resource utilization.

[0084] In certain embodiments, some internal shared services 1032 may be provided that are shared by different components or modules of the cloud infrastructure system 1002 and by the services provided by the cloud infrastructure system 1002. These internal shared services may include, but are not limited to, security and identity services, integration services, corporate repository services, corporate manager services, virus scanning and whitelisting services, high availability, backup and recovery services, services to enable cloud support, email services, notification services, file transfer services, etc.

[0085] In certain embodiments, the cloud infrastructure system 1002 may provide comprehensive management of cloud services (e.g., SaaS, PaaS, and IaaS services) in the cloud infrastructure system. In one embodiment, the cloud management functionality may include the ability to provision, manage, and track customer subscriptions received by the cloud infrastructure system 1002, among other things.

[0086] In one embodiment, as shown in the figure, the cloud management functionality may be provided by one or more modules such as an order management module 1020, an order orchestration module 1022, an order provisioning module 1024, an order management and monitoring module 1026, and an identity management module 1028. These modules may be provided using or may include one or more computers and / or servers that may be a general-purpose computer, a dedicated server computer, a server farm, a server cluster, or any other suitable configuration and / or combination.

[0087] In an exemplary operation 1034, a customer using a client device, such as client devices 1004, 1006, or 1008, may interact with the cloud infrastructure system 1002 by requesting one or more services provided by the cloud infrastructure system 1002 and placing an order for a subscription to one or more services provided by the cloud infrastructure system 1002. In certain embodiments, the customer may access a cloud user interface (UI), cloud UI 1012, cloud UI 1014, and / or cloud UI 1016 and place a subscription order through these UIs. Order information received by the cloud infrastructure system 1002 in response to the customer placing an order may include information identifying the customer and the one or more services provided by the cloud infrastructure system 1002 that the customer is attempting to subscribe to.

[0088] After an order is placed by a customer, order information is received via the cloud UIs 1012, 1014, and / or 1016.

[0089] In operation 1036, the order is stored in the order database 1018. The order database 1018 can be one of several databases that are operated by the cloud infrastructure system 1018 and are associated with other system elements and operated in relation to them.

[0090] In operation 1038, the order information is transferred to the order management module 1020. In some examples, the order management module 1020 may be configured to perform billing and accounting functions related to the order, such as verifying the order and invoicing the order if it is verified.

[0091] In operation 1040, information about the order is communicated to the order orchestration module 1022. The order orchestration module 1022 may use the order information to orchestrate the provisioning of services and resources for the order placed by the customer. In some examples, the order orchestration module 1022 may use the services of the order provisioning module 1024 to orchestrate the provisioning of resources to support the subscribed services.

[0092] In certain embodiments, the order orchestration module 1022 enables the management of the business processes associated with each order, and applies business logic to determine whether an order should proceed to provisioning. In operation 1042, upon receiving an order for a new subscription, the order orchestration module 1022 allocates resources and sends a request to the order provisioning module 1024 to configure the resources required to fulfill the subscription order. The order provisioning module 1024 enables the allocation of resources for the services ordered by the customer. The order provisioning module 1024 provides a level of abstraction between the cloud services provided by the cloud infrastructure system 1000 and the physical implementation layer used to provision the resources for providing the requested services. Thus, the order orchestration module 1022 may be isolated from the implementation details, such as whether the services and resources are actually provisioned on-the-fly or pre-provisioned and only allocated / assigned upon request.

[0093] In operation 1044, upon provisioning of the services and resources, a notification of the provided services may be sent to the customer on the client devices 1004, 1006, and / or 1008 by the order provisioning module 1024 of the cloud infrastructure system 1002.

[0094] In operation 1046, the customer's subscription order may be managed and tracked by the order management and monitoring module 1026. In some cases, the order management and monitoring module 1026 may be configured to collect usage statistics for the services in the subscription order, such as the amount of storage used, the amount of data transferred, the number of users, and the amount of system up-time and system down-time.

[0095] In one embodiment, the cloud infrastructure system 1000 may include an identity management module 1028. The identity management module 1028 may be configured to provide identity services such as access management and authorization services in the cloud infrastructure system 1000. In some embodiments, the identity management module 1028 may control information about customers who desire to utilize services provided by the cloud infrastructure system 1002. Such information may include information for authenticating the nature of such customers, and information describing what actions they are permitted to perform on various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). The identity management module 1028 may also include management of descriptive information about each customer, as well as information about how and who can access and modify that descriptive information. It can include information that describes what actions they are permitted to perform on various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). The identity management module 1028 may also include management of descriptive information about each customer, as well as information about how and who can access and modify that descriptive information.

[0096] FIG. 11 shows an exemplary computer system 1100 in which various embodiments of the present invention may be implemented. The system 1100 may be used to implement any of the computer systems described above. As shown in the figure, the computer system 1100 includes a processing unit 1104 that communicates with several peripheral subsystems via a bus subsystem 1102. These peripheral subsystems may include a processing acceleration unit 1106, an I / O subsystem 1108, a storage subsystem 1118, and a communication subsystem 1124. The storage subsystem 1118 includes a tangible computer-readable storage medium 1122 and a system memory 1110.

[0097] The bus subsystem 1102 provides a mechanism for the various components and subsystems of the computer system 1100 to communicate with each other as intended. The bus subsystem 1102 is shown schematically as a single bus, although alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 1102 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus, using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Extended ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus which can be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.

[0098] The processing unit 1104 can be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers) and controls the operation of the computer system 1100. One or more processors may be included in the processing unit 1104. These processors may include single-core processors or multi-core processors. In certain embodiments, the processing unit 1104 may be implemented as one or more independent processing units 1132 and / or 1134 in which a single-core processor or multi-core processor is included in each processing unit. In other embodiments, the processing unit 1104 may also be implemented as a quad-core processing unit formed by integrating two dual-core processors on a single chip.

[0099] In various embodiments, the processing unit 1104 can execute various programs in response to program code and can maintain multiple simultaneously-executed programs or processes. At any given time, some or all of the program code to be executed can reside in the processor 1104 and / or the storage subsystem 1118. Through suitable programming, the processor 1104 can provide the various functionalities described above. The computer system 1100 may further include a processing acceleration unit 1106 that can include a digital signal processor (DSP), a special-purpose processor, and the like.

[0100] The I / O subsystem 1108 may include user interface input devices and user interface output devices. User interface input devices may include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touch screen incorporated in a display, a scroll wheel, a click wheel, a dial, a button, a switch, a keypad, a voice input device with a voice command recognition system, a microphone, and other types of input devices. The user interface input devices can enable a user, for example, to control and interact with input devices such as a Microsoft Xbox (registered trademark) 360 game controller through a natural user interface using gestures and spoken commands, and can include motion sensing and / or gesture recognition devices such as a Microsoft Kinect (registered trademark) motion sensor. The user interface input devices can detect eye movements from the user (e.g., "blinks" while taking a photo and / or while making a menu selection) and input eye gestures to the input device (e.g., Google Eye gesture recognition devices such as a Google Glass (registered trademark) blink detector that converts input to Glass (registered trademark) may also be included. Additionally, the user interface input device may include an audio recognition sensing device that enables the user to interact with an audio recognition system (e.g., Siri (registered trademark) navigator) via voice commands.

[0101] The user interface input device may include, but is not limited to, a three-dimensional (3D) mouse, joystick or pointing stick, game pad and graphic tablet, and audio / visual devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser range finders, and eye tracking devices. Additionally, the user interface input device may include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, medical ultrasound examination devices, etc. The user interface input device may also include, for example, audio input devices such as MIDI keyboards, digital musical instruments, etc.

[0102] The user interface output device may include a non-visual display such as a display subsystem, an indicator light, or an audio output device. The display subsystem may be a flat panel device such as one using a cathode ray tube (CRT), a liquid crystal display (LCD), or a plasma display, a projection device, a touch screen, etc. Generally, the use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from the computer system 1100 to the user or another computer. For example, the user interface output device may include, but is not limited to, various display devices for visually conveying text, graphics, and audio / video information, such as a monitor, a printer, a speaker, headphones, an automotive navigation system, a plotter, an audio output device, and a modem.

[0103] The computer system 1100 may include a storage subsystem 1118 that includes software elements shown as being currently located within the system memory 1110. The system memory 1110 may store program instructions that are loadable and executable on the processing unit 1104, as well as data generated during the execution of these programs.

[0104] Depending on the configuration and type of computer system 1100, system memory 1110 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that are immediately accessible to processing unit 1104 and / or are currently being manipulated and executed by processing unit 1104. In some implementations, system memory 1110 may include multiple different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS) that includes basic routines that help transfer information between elements within computer system 1100 during startup, etc. may typically be stored in ROM. By way of example and not limitation, system memory 1110 also shows application programs 1112, which may include client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc., program data 1114, and operating system 1116. By way of example, operating system 1116 may include various versions of Microsoft Windows (registered trademark), Apple Macintosh (registered trademark), and / or Linux (registered trademark) operating systems, various commercially available UNIX (registered trademark) or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome (registered trademark) OS, etc.), and / or mobile operating systems such as iOS, Windows (registered trademark) Phone, Android (registered trademark) OS, BlackBerry (registered trademark) 10 OS, and Palm (registered trademark) OS.

[0105] The memory subsystem 1118 may also provide a tangible computer-readable storage medium for storing the basic programming and data structures that provide the functionality of some embodiments. Software (programs, code modules, instructions) that provides the above-described functionality when executed by a processor may be stored in the storage subsystem 1118. These software modules or instructions may be executed by the processing unit 1104. The storage subsystem 1118 may also provide a repository for storing data used in accordance with the present invention.

[0106] The storage subsystem 1100 may also include a computer-readable storage medium reader 1120 that may be further connected to a computer-readable storage medium 1122. Together with the system memory 1110 and, optionally, in combination with the system memory 1110, the computer-readable storage medium 1122 may comprehensively represent a storage medium added to a remote, local, fixed, and / or removable storage device for temporarily and / or more permanently accommodating, storing, transmitting, and retrieving computer-readable information.

[0107] A computer-readable storage medium 1122 that includes code or a portion of code can also include any suitable medium known in or used in the art, including, but not limited to, volatile and non-volatile, removable and non-removable media such as storage media and communication media implemented in any method or technology for the storage and / or transmission of information. This can include tangible computer-readable storage media such as RAM, ROM, electronically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or other tangible computer-readable media. This can also include non-tangible computer-readable media such as any other medium that can be used to transmit a data signal, data transmission, or desired information and can be accessed by computing system 1100.

[0108] By way of example, computer-readable storage medium 1122 can include a hard disk drive that reads from and writes to a non-removable non-volatile magnetic medium, a magnetic disk drive that reads from and writes to a removable non-volatile magnetic disk, a CD ROM, a DVD, and an optical disk drive that reads from and writes to a removable non-volatile optical disk such as a Blu-Ray (registered trademark ) disk, or other optical media. Computer-readable storage medium 1122 can include a Zip (registered trademark) drive, a flash memory card, a universal serial bus (USB) flash drive, a secure digital (SD) card, a DVD disk, a digital It may include, but is not limited to, video tapes, etc. The computer-readable storage medium 1122 may also be a solid-state drive (SSD) based on non-volatile memory such as a flash memory-based SSD, an enterprise flash drive, a solid-state ROM, an SSD based on volatile memory such as solid-state RAM, dynamic RAM, static RAM, a DRAM-based SSD, a magnetoresistive RAM (MRAM) SSD, and a hybrid SSD that uses a combination of DRAM and a flash memory-based SSD. Disk drives and the computer-readable media associated therewith may provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data to the computer system 1100.

[0109] The communication subsystem 1124 provides an interface to other computer systems and networks. The communication subsystem 1124 serves as an interface for sending and receiving data between other systems and the computer system 1100. For example, the communication subsystem 1124 may enable the computer system 1100 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1124 may include radio frequency (RF) transceiver components, global positioning system (GPS) receiver components, and / or other components for accessing a wireless voice and / or data network (e.g., using cellular phone technology, advanced data network technologies such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), WiFi (IEEE802.11 family of standards, or other mobile communication technologies, or any combination thereof)). In some embodiments, the communication subsystem 1124 can provide wired network connectivity (e.g., Ethernet (registered trademark)) in addition to, or instead of, a wireless interface.

[0110] In some embodiments, communication subsystem 1124 may also receive input communications in the form of structured and / or unstructured data feeds 1126, event streams 1128, event updates 1130, etc., on behalf of one or more users who may also use computer system 1100.

[0111] By way of example, communication subsystem 1124 may be configured to receive data feeds 1126 in real time from users of social networks and / or other communication services, such as Twitter (registered trademark) feeds, Facebook ( registered trademark) updates, web feeds such as Rich Site Summary (RSS) feeds, and / or / or real-time updates from one or more third-party information sources.

[0112] In addition, communication subsystem 1124 may also be configured to receive data in the form of continuous data streams, which may include event streams 1128 and / or event updates 1130 of real-time events that are essentially continuous or infinite and without an explicit termination. Examples of applications that generate continuous data may include, for example, sensor data applications, financial stock market dashboards, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc.

[0113] Communication subsystem 1124 may also be configured to output structured and / or unstructured data feeds 1126, event streams 1128, event updates 1130, etc., to one or more databases that communicate with one or more streaming data source computers coupled to computer system 1100.

[0114] Computer system 1100 may be a handheld portable device (e.g., an iPhone It can be any one of various types, including (registered trademark) mobile phones, iPad (registered trademark) computing tablets, PDAs, wearable devices (e.g., Google Glass (registered trademark) head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or any other data processing system.

[0115] Due to the ever-changing nature of computers and networks, the description of the computer system 1100 shown in the figure is intended merely as a specific example. Many other configurations are possible with more or fewer components than the system depicted in the figure. For example, customized hardware may also be used and / or certain elements may be implemented in hardware, firmware, software (including applets), or a combination. Further, connections to other computing devices such as network input / output devices may be employed. Based on the disclosure and teachings provided herein, those skilled in the art will understand other aspects and / or methods for implementing various embodiments.

[0116] In the foregoing description, for purposes of explanation, numerous specific details have been set forth in order to provide a thorough understanding of the various embodiments of the present invention. However, it will be apparent to those skilled in the art that some of these specific details may not be required to practice embodiments of the present invention. In other instances, well-known structures and devices are shown in block diagram form.

[0117] The foregoing description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the foregoing description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes can be made to the functions and configurations of the elements without departing from the spirit and scope of the invention as recited in the claims.

[0118] In the above description, specific details are provided to give a complete understanding of the embodiments. However, it will be understood by those skilled in the art that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may sometimes be shown as components in block diagram form in order not to obscure the embodiments with unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

[0119] It should also be noted that individual embodiments may sometimes be described as processes depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe an operation as a sequential process, many of the operations can be executed in parallel or simultaneously. In addition, the order of the operations may be rearranged. A process is terminated when its operations are completed, but it may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0120] The term "computer-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and various other media that can store, contain, or carry instructions and / or data. A code segment or machine-executable instruction may represent any combination of a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or instructions, data structures, or program statements. A code segment may transfer and / or by receiving, may be coupled to another code segment or hardware circuit. Information, arguments, parameters, data, etc. may be passed, transferred, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0121] Furthermore, embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks may be stored in a machine-readable medium. The necessary tasks may be executed by a processor.

[0122] In the foregoing specification, aspects of the invention have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that the invention is not limited thereto. The various features and aspects of the above invention may be used individually or together. Furthermore, embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

[0123] Furthermore, for purposes of illustration, the method has been described in a particular order. It should be understood that in alternative embodiments, the method may be performed in an order different from that described. Also, the above method may be performed by hardware components or may be embodied as a sequence of machine-executable instructions that, when used, cause a machine such as a general or special purpose processor or logic circuit programmed with such instructions to perform the above method. It should also be understood that these machine-executable instructions may be stored on one or more machine-readable media such as a CD-ROM or other type of optical disk, a floppy disk, ROM, RAM, EPROM, EEPROM, magnetic or optical card, flash memory, or other type of machine-readable medium suitable for storing electronic instructions. Alternatively, the method may be performed by a combination of hardware and software.

Claims

1. A non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: Receiving a first performance metric from a data center, The data center having a first configuration, The first performance metric being generated when the data center executes a workload, and the operations further including: Receiving one or more environmental characteristics of the data center when the data center executes the workload; Identifying a model trained using data from one or more data center configurations operating in an environment having environmental characteristics similar to the one or more environmental characteristics of the data center; Generating a second performance metric from the model; Comparing the second performance metric to the first performance metric; and Determining a second configuration for the data center from the model.

2. The non-transitory computer-readable medium of claim 1, wherein the model is identified from a plurality of models, each of the plurality of models being trained using data from one or more data center configurations operating in an environment having different environmental characteristics.

3. The non-transitory computer-readable medium of claim 1, wherein the one or more data center configurations are disposed in an environmental chamber to control the environmental characteristics while the one or more data center configurations execute the workload.

4. The non-transitory computer-readable medium of claim 3, wherein the environmental chamber controls the temperature around the one or more data center configurations and a simulated altitude.

5. The non-transitory computer-readable medium of claim 4, wherein the environmental chamber is configured to simulate an altitude and collect a complete set of telemetry of signals at a plurality of temperatures at the altitude.

6. The non-transitory computer-readable medium of claim 5, wherein the plurality of temperatures includes temperatures from about 15°C to 35°C.

7. The non-transitory computer-readable medium of claim 4, wherein the environmental chamber simulates a plurality of altitudes between approximately sea level and 5000 feet.

8. The non-transitory computer-readable medium of claim 7, wherein the plurality of temperatures are incremented at intervals of about 1°C.

9. The non-transitory computer-readable medium according to claim 1, wherein the one or more environmental characteristics include the ambient temperature around the data center and the altitude at which the data center is installed.

10. Identifying the model trained using data from one or more data center configurations operating in the environment having environmental characteristics similar to the one or more environmental characteristics, The non-transitory computer-readable medium according to claim 1, comprising executing a nearest neighbor algorithm to identify environmental characteristics most similar to the one or more environmental characteristics.

11. The non-transitory computer-readable medium according to claim 1, wherein the nearest neighbor algorithm minimizes the difference between the temperature of the model and the temperature of the data center and minimizes the difference between the simulated altitude of the model and the altitude of the data center.

12. The non-transitory computer-readable medium according to claim 1, wherein the first configuration includes the number and type of processors within the data center.

13. The non-transitory computer-readable medium according to claim 1, wherein the first configuration includes the number and type of hard disk drives within the data center.

14. The non-transitory computer-readable medium according to claim 1, wherein the first configuration includes a data cache size.

15. The non-transitory computer-readable medium according to claim 1, wherein the data center includes a cloud data center.

16. The operation further includes comparing the first performance metric with a service level agreement (SLA), determining that the first performance metric does not meet the SLA, and identifying the model in response to determining that the first performance metric does not meet the SLA. The non-transitory computer-readable medium according to claim 1.

17. The operation further includes determining that the second performance metric exceeds the first performance metric, providing the second configuration to be embedded by the data center. The non-transitory computer-readable medium according to claim 1.

18. The non-transitory computer-readable medium according to claim 1, wherein the first performance metric includes the processor performance of each processor within the data center and the I / O performance of each hard disk drive within the data center.

19. A system, one or more processors, and one or more memory devices including instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: receiving a first performance metric from a data center, the data center having a first configuration, the first performance metric being generated when the data center executes a workload, and the operations further including: receiving one or more environmental characteristics of the data center when the data center executes the workload; identifying a model trained using data from one or more data center configurations operating in an environment having environmental characteristics similar to the one or more environmental characteristics of the data center; generating a second performance metric from the model; comparing the second performance metric to the first performance metric; determining, from the model, a second configuration for the data center. Claim 20 A method for autonomously determining an optimal configuration of a data center using a library of models pre-trained with various combinations of environmental characteristics, the method including: receiving a first performance metric from a data center, the data center having a first configuration, the first performance metric being generated when the data center executes a workload, and the method further including: receiving one or more environmental characteristics of the data center when the data center executes the workload; identifying a model trained using data from one or more data center configurations operating in an environment having environmental characteristics similar to the one or more environmental characteristics of the data center; generating a second performance metric from the model; comparing the second performance metric to the first performance metric; determining, from the model, a second configuration for the data center.

Citation Information

Patent Citations

  • Monitoring control system, monitoring control method, monitoring control server, and monitoring control program

    JP2010237901A

  • Data center heat removal system and method

    JP2018509675A

  • System and method for dynamically modeling data center partitions

    US20100082309A1

  • Inference of altitude using pairwise comparison of telemetry signals

    US20110258157A1

  • Characterizing the i / o-performance-per-watt of a computing device across a range of vibrational operating environments

    US20180058976A1