Detecting large-scale datacenter faults using near real-time / offline data with ML models

JP2024521357A5Pending Publication Date: 2025-05-30ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023574401
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-03
Filing Date
2022-05-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Detecting large-scale failures in data centers is challenging due to the complexity of the infrastructure, making it difficult to efficiently identify the source of failures and perform timely repairs, which degrades user experience.

Method used

Utilizing machine learning models to analyze near real-time and offline data from data centers to detect and estimate probable sources of failures by processing input data with a set of rules that identify correlations between devices and applications, generating fault notification messages.

Benefits of technology

Facilitates rapid identification of failure sources, reducing the time to detect and resolve issues by providing actionable insights and solutions, thereby enhancing data center reliability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

FIELD OF THE DISCLOSURE Embodiments of the present disclosure relate to detecting data center faults and generating alerts. The fault detection service described herein may process near real-time data from various sources in a data center and determine one or more suspected sources of a detected fault by processing the data with a model. The models described herein may include one or more machine learning models that process the near real-time data and offline data and incorporate a set of rules for determining one or more suspected sources of the fault. An alert message may be generated to provide the suspected source of the fault and other data related to the fault.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Nonprovisional Patent Application No. 17 / 338,478, entitled “DETECTING DATACENTER MASS OUTAGE WITH NEAR REAL-TIME / OFFLINE DATA USING ML MODELS,” filed June 3, 2021, the entire contents of which are incorporated by reference into this disclosure for all purposes.

[0002] Field The disclosed technology relates to detecting catastrophic failures in data centers, and more particularly, to detecting catastrophic failures in data centers using machine learning models that analyze real-time and / or offline data. [Background technology]

[0003] background A data center may include multiple computing devices (e.g., servers) configured to perform various processing tasks and associated devices for powering the computing devices and connecting the computing devices to external devices. The servers may be arranged in racks containing several servers, and power for the servers in the rack is provided from a rack power supply. The internal conditions (e.g., temperature, humidity) of the data center may be controlled and monitored (e.g., using sensors and air conditioning devices) to prevent overheating or loss of functionality of the servers in the data center. Summary of the Invention [Problem to be solved by the invention]

[0004] However, failures may occur in data centers for various reasons. A failure may include loss of any functionality of any computing device in the data center, for example, loss of functionality of applications running on a server or overheating and shutdown of a server. Such failures may degrade the experience of users interacting with the devices in the data center and / or the applications running on the devices. Thus, operators maintaining data centers desire to efficiently identify the source of the failure and resolve the causes that cause the failure. However, as many devices and applications are implemented in data centers, efficiently identifying the source of the failure may become increasingly difficult. [Means for solving the problem]

[0005] overview The present embodiment relates to detecting large-scale datacenter failures with near-real-time data using one or more models. A first exemplary embodiment provides a method executed by a cloud infrastructure node for estimating one or more suspected sources of datacenter failures. The method may include obtaining a set of input data providing various parameters related to the datacenter, a set of devices in the datacenter, and applications running on the devices. The method may also include detecting a failure of at least one function of the datacenter. The failure may be due to a loss of (e.g., application) functionality or a loss of computing resources (e.g., loss of connectivity to a server, loss of power to a server).

[0006] The method may also include estimating one or more suspected sources of the failure by processing the set of input data with the model. The model may incorporate a set of rules that identify a correlation between the set of input data and the device or an application executing on the device as the one or more suspected sources of the failure. The method may also include generating a failure notification message providing the one or more suspected sources of the failure.

[0007] A second exemplary embodiment relates to a cloud infrastructure node. The cloud infrastructure node may include a processor and a non-transitory computer-readable medium. The non-transitory computer-readable medium includes instructions that, when executed by the processor, cause the processor to obtain a set of input data providing various parameters related to a data center, a set of devices in the data center, and applications running on the devices. The instructions also cause the processor to detect a failure of a function of the data center.

[0008] The instructions may further cause the processor to estimate one or more suspected sources of the fault using the model using the set of input data. Estimating the one or more suspected sources of the fault may include generating an estimated level of each parameter using a set of rules available to the model and historical data associated with each parameter included in the set of input data. In some examples, the historical data includes previously received data from the same source (e.g., for each parameter). Estimating the one or more suspected sources of the fault may include identifying one or more anomalous parameters including actual levels having a threshold deviation from each corresponding estimated level by comparing the estimated level of each parameter with an actual level of each parameter included in the set of input data. Estimating the one or more suspected sources of the fault may include identifying one or more devices and / or applications corresponding to each of the identified anomalous parameters. Each of the identified one or more devices and / or applications is included as one or more suspected sources of the fault. The instructions may further cause the processor to generate a fault notification message providing the one or more suspected sources of the fault.

[0009] A third embodiment relates to a non-transitory computer-readable medium. The non-transitory computer-readable medium may include a set of instructions that, when executed by a processor, causes the processor to perform a process. The process may include obtaining a set of input data providing various parameters associated with the data center. The process may also include detecting a fault in the data center. Additionally, the process may include estimating one or more suspected sources of the fault using a model that uses the set of input data.

[0010] Estimating the one or more suspected sources of the fault may include identifying one or more anomalous parameters having a threshold deviation from a corresponding respective estimated level by comparing the estimated level of each parameter with an estimated level of each parameter included in the set of input data. Estimating the one or more suspected sources of the fault may also include identifying one or more devices and / or applications corresponding to each of the identified anomalous parameters. Each of the identified one or more devices and / or applications is included as one or more suspected sources of the fault. The process may also include generating a fault notification message providing the one or more suspected sources of the fault.

[0011] Furthermore, the embodiments may be implemented using a computer program product comprising computer programs / instructions which, when executed by a processor, cause the processor to perform any of the methods / techniques described in the present disclosure. [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary data center in accordance with at least one embodiment. [Diagram 2] 1 is a flow diagram illustrating a method for generating an alert for a fault in accordance with at least one embodiment. [Diagram 3] FIG. 2 is a block diagram illustrating an example fault detection service in accordance with at least one embodiment. [Figure 4] FIG. 1 illustrates an example alert in accordance with at least one embodiment. [Diagram 5] FIG. 1 is a block diagram illustrating an example method for estimating one or more suspected sources of a fault in a data center, in accordance with at least one embodiment. [Figure 6] FIG. 1 is a block diagram illustrating one pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 7]FIG. 2 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 8] FIG. 2 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 9] FIG. 2 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 10] FIG. 1 is a block diagram illustrating an example computer system in accordance with at least one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Detailed Description A data center may include multiple devices, such as computing devices, power supplies that power the computing devices, network devices that communicate data to and from the computing devices, and / or multiple sensors that monitor / control the environment of the data center. Often, for example, a failure in a data center may result in an inability to access the computing devices or associated processes performed by the computing devices, or an inability to transfer data between devices in the data center.

[0014] Data center failures can result from a variety of causes, such as a power failure, a failure of one or more devices in the data center, overheating of devices in the data center, or failure to run an application. Particularly because data centers contain many devices, processing resources, and applications / services, it can be difficult to efficiently identify the source of a failure and perform a repair action. Longer times to repair a failure can degrade the experience of users interacting with the devices / applications in the data center.

[0015] The present embodiment relates to detecting faults and generating alerts in a data center. In particular, the fault detection service described herein may process near real-time data from various sources in a data center and determine one or more suspected sources of a detected fault by processing the data with a model.

[0016] For example, a failure may be caused by a malfunction of a rack power supply (e.g., powering a rack of servers), leading to a loss of functionality of the corresponding server. In this case, the fault detection system described herein may identify one or more anomalous parameters by processing the near real-time data with a model. In this example, the model may identify that the power level of the rack power supply near the time of detecting the failure has fallen below a threshold level or that the power level of a server in the rack has fallen below a threshold level. The model may determine that the rack power supply and / or the server in the rack are the suspected source of the failure by processing the near real-time input data using a set of rules.

[0017] The warning message may be generated to provide a suspected source of the failure and other data related to the failure. In the above example, the warning message may identify the rack power supply and / or the servers in the rack as the suspected sources of the failure, may identify the anomaly parameters identified by the model, a confidence value for each of the suspected sources of the failure, etc. The warning message may provide insight into the failure and may efficiently repair the failure.

[0018] The near real-time data may include environmental data from equipment in the data center. Exemplary near real-time data may include server temperatures, server / rack power usage, tickets acquired, sensor data, etc., and is stored with a timestamp indicating the time the near real-time data was acquired. In response to the occurrence of a fault, the fault detection service may identify one or more suspected sources of the fault by running a model using the near real-time data and the offline data as inputs.

[0019] The models described herein may include one or more machine learning models incorporating a set of rules to process the near real-time data and the offline data and determine one or more inferred sources of failure. For example, the model may identify one or more anomalous parameters of equipment in a data center that are likely to cause a failure. The model may output one or more inferred sources or causes of the failure, such as an identification of equipment, power sources, applications, etc. that are likely to cause the failure, and a confidence value that provides an estimated confidence of each inferred source causing the failure. The inferred sources of failure may establish correlations of failures at scale that may be used to inform failure recovery by providing patterns detected from the near real-time data. For example, the inferred sources of failure may provide a blueprint of how the failure (and any associated causes) are spreading across components / applications in a data center. Utilizing the inferred sources of failure as a blueprint to recover from failures may reduce the overall time to detect and resolve the failure.

[0020] As an illustrative example, a fault may be detected from the data center by an instruction from an operator, or may be automatically detected from the data center by a fault detection system (e.g., by detecting an abnormal parameter, by detecting a number of incoming tickets identifying the fault). For example, a fault may be caused by an abnormal increase in power from a rack power supply in a server rack causing the server rack to lose functionality and overheating the servers in the server rack.

[0021] In this example, the fault detection service may obtain data (e.g., near real-time data 202) related to server temperatures (e.g., 206), rack power usage (e.g., 210), ticket data (e.g., 212), etc., and prepare the data according to timestamps for processing by the model. The model may process the obtained data to identify anomalous parameters that may indicate a cause of a fault. For example, multiple received tickets may identify a fault that occurred at a first time point. Additionally, a rack power metric for a rack power supply may abnormally increase at a first time point, and fan speeds for servers in the rack (indicating a core server temperature increase) may increase at a first time point.

[0022] In this example, the model may identify that a primary cause of the failure may include a power surge in a rack power supply that causes the servers in the rack to overheat (and function limited). Using a set of rules, the model may identify the likelihood that the cause of the failure is a power surge in the rack power supply, and may assign a confidence value (e.g., as a percentage) to the cause of the failure.

[0023] In this example, the fault detection service may provide solution data that provides one or more steps to resolve the cause of the fault. Exemplary solution data may identify resetting or replacing a rack power supply or resetting a server in the rack. The fault detection service may output a warning that includes the characteristics of the fault, a probable cause of the fault, the solution data, and / or one or more graphs showing abnormal parameters identified from data obtained from the data center.

[0024] 1 is a block diagram illustrating an exemplary data center 100. Data center 100 may include an environment (e.g., a room, a building) that houses multiple computing devices and associated devices that provide power to the computing devices and facilitate data communication between the computing devices and devices external to data center 100. Data center 100 may provide a controlled environment to maintain threshold environmental conditions (e.g., temperature, humidity) within data center 100.

[0025] The data center 100 may include multiple server stacks 102a-n including computing devices (e.g., servers 104a-l). A server stack (e.g., 102a-n) may include a rack in which a set of servers 104a-l are located. Each server stack 102a-n may include one or more power sources (e.g., power units 106a-n) and network devices (e.g., 108a-n) that enable data transmission between devices within the data center 100. In some embodiments, the fault detection services described herein may be implemented on one or more computing devices (e.g., servers 104a-l) or on computing devices external to the data center 100.

[0026] Each server 104a-l in data center 100 may implement applications / plug-ins / add-ons / virtual machines, etc. configured to perform various processing tasks, such as, for example, maintaining and updating databases. The servers 104a-l may include multiple sensors configured to capture data regarding each server, such as, for example, core temperature, power usage, fan speeds, status, etc. of each server 104a-l.

[0027] Each server 104a-l may be connected to one or more power units 106a-n. Each power unit may be associated with a server stack 102a-n and provide power to the servers 104a-l. The power units 106a-n may monitor multiple power parameters (e.g., voltage, current) provided by each power unit 106a-n, which may be provided to the fault detection service as near real-time data.

[0028] The servers 104a-l may communicate data via network devices 108a-n. The network devices 108a-n may include network switches, routers, etc. that may forward data between the servers 104a-l and receiving devices. In some cases, the network devices 108a-n may implement a streaming service that provides low latency data communication between the servers 104a-l and a fault detection service running on a cloud infrastructure node as described herein.

[0029] The data center 100 may include multiple sensors 110a-n. The sensors 110a-n may monitor / control the environment of the data center 100. Exemplary sensors may include temperature sensors, humidity sensors, pressure sensors, etc. Data captured by the sensors 110a-n may be provided to a fault detection service as near real-time data.

[0030] 2 is a flow diagram 200 illustrating a method for generating an alert for a fault. As described below, the alert may include a notification (e.g., a message, email, text notification) that is provided to a data center operator and provides insight into the fault and the potential source of the fault. The method for generating an alert for a fault may be performed by a fault detection service described herein.

[0031] The fault detection service may obtain near real-time data 202 and offline data 204 from various sources. The near real-time data 202 and offline data 204 may be processed as input data to be processed using the models described herein. The near real-time data 202 may include various types of data, such as server temperature data 206, server power usage data 208, rack power usage data 210, ticket data 212, and any type of other data 214. The fault detection service may obtain the near real-time data 202 from sources in the data center via a streaming service that provides low latency data communication to the fault detection service.

[0032] Server temperature data 206 may include data regarding internal temperatures of servers in a data center provided by sensors (e.g., 110a-n) or servers (e.g., 104a-n). Server temperature data 206 may identify server temperatures at a particular point in time, which may allow trends in server temperatures to be monitored over time. As described herein, an increase in server temperature of one or more servers may indicate increased power usage or overheating of the servers, which may cause failures. In some cases, server temperature data 206 may include fan speed data identifying fan speeds of servers in a data center, which may indicate the temperature of the servers.

[0033] The server power usage data 208 may identify the power consumption of each server in a data center. Exemplary parameters associated with the server power usage data 208 may include the voltage, current, power consumption, production load, etc. of each server during a period of time. The server power usage data 208 may be obtained by placing various sensors inside or near the servers.

[0034] The rack power usage data 210 may provide power usage of the servers (and / or attached devices) in a rack (e.g., server stack 102a-n). The rack power usage may be provided by the rack's power source (e.g., power units 106a-n). The power units 106a-n may measure multiple power parameters (e.g., voltage, current, and power consumption of the rack and each device in the rack).

[0035] The ticket data 212 may include a set of tickets obtained by a ticket node (e.g., an application running on a computing device to obtain tickets from devices in the data center or devices external to the data center). Tickets may be received for detected problems / warnings with devices or applications running on devices in the data center. Tickets may be provided automatically by devices communicating with devices in the data center or manually by operators interacting with devices in the data center.

[0036] As one example, a device may automatically generate a ticket when it is unable to obtain data from an application running on a first server in a data set. As another example, a client may generate a ticket when it is unable to access a database maintained by a second server in a data center via a client device. As described herein, the ticket may be associated with a timestamp and may be used to identify the cause of the failure or outage.

[0037] Other data 214 may include network data identifying data transmission characteristics of devices within the data center, application data parameters of applications running on devices within the data center, change records of applications / devices within the data center, and the like.

[0038] The fault detection service may also obtain offline data 204 from data sources, such as one or more databases containing static data center information. Examples of offline data 204 may include equipment data 216, location data 218, and other data 220. Equipment data 216 may identify the number of equipment in the data center, and location data 218 may include the location of each equipment in the data center. Equipment data 216 and location data 218 may identify equipment groups in the data center, such as servers grouped into racks. Other data 220 may identify applications running on each server, the capabilities of each equipment in the data center, the software version of each equipment, the types of equipment in the data center (e.g., sensors, network equipment, power equipment), etc.

[0039] At 222, the near real-time data 202 and the offline data 204 may be combined. This may include ordering the data according to data type and storing the data in a data source (e.g., database, table) based on a timestamp associated with the data. As data is acquired over time, the data may be stored in a database / table by data type according to the time the data was received. For example, server temperature data for a first server in a data center may be stored according to the time the data was acquired to provide the temperature of the first server over a period of time. As another example, rack power usage may be stored to provide trends in rack power usage over time. Trends and variations in parameters in the received data may provide insight into abnormal parameters in the data center and potential causes of failures in the data center.

[0040] In some embodiments, an estimated level of a parameter measured from a data center may be generated based on the data center's historical levels. For example, historical server temperature data may be captured over time and an estimated temperature at each point in time may be generated. Historical server temperatures are just one example of a historical level. By comparing the estimated level with the corresponding parameter, deviations from the estimated level may be detected, which may indicate an anomaly that may be a source of failure.

[0041] The model 224 may be executed using the combined data to determine one or more probable causes of the failure. In some cases, the model may be executed in response to detection of a failure (e.g., manual indication by an operator, automated detection by inspection of ticket data).

[0042] The model 224 may process the combined data as input parameters that may be used to detect anomalous behavior that may indicate a source of a fault. The model 224 may include a machine learning model that may incorporate a number of rules 226 to process the combined data (e.g., the data combined at 222) and detect one or more suspected sources of a fault.

[0043] The rules 226 may be generated from previously identified failures and known solutions to the failures. For example, if a previous failure was due to a power surge in a rack power supply, the new rule may include instructions to monitor for similar power surges in any power supply and similar characteristics of failures detected due to the power surge. The rules 226 may also be generated based on past data center data or feedback data provided in response to the solution of the failure. In some examples, the past data center data is data center data and / or metadata obtained at a past time.

[0044] For example, the rules 226 may include instructions for processing input parameters to determine whether the parameters have anomalous characteristics at any point in time, and may include instructions for correlating the anomalous parameters with one or more devices as a suspected source of a fault and identifying one or more devices affected by the anomalous parameters.

[0045] In a first example, the model may use a first rule to determine whether the server power data of a first server includes an anomalous characteristic. For example, rule 226 may provide instructions for the model to detect deviations between the actual power level and the estimated power level by comparing the server power data of the first server to the estimated power level. The rule may specify that when the actual power level exceeds a threshold deviation from the estimated power level at a particular time, the model 224 identifies the server power level of the first server as an anomalous parameter.

[0046] As another example, the rules 226 may include instructions for identifying any changes made to the applications / software of the equipment in the data center within a threshold time of detecting the failure (e.g., receiving a ticket indicating the failure). For example, if an add-on caused the failure, the rules may identify any changes made to the software in the data center within the threshold time of detecting the failure. In this example, the rules may identify add-ons that were implemented at a similar time to the time of detecting the failure and that constitute a potential source of the failure.

[0047] Subsequent rules may determine one or more suspected sources of failure by processing the abnormal characteristics. Rules 226 may include instructions for correlating the abnormal characteristics with a corresponding device / set of devices / set of applications / etc. For example, if the power level of a first rack power supply spikes above an estimated level, the rules may specify that servers connected to the first rack power supply are more likely to cause failures because the increased power level may result in loss of functionality. As another example, if a threshold number of tickets identify a failure of a first application, the rules may specify that all servers implementing the first application (or virtual machines implementing the first application) may include suspected sources of failure. The model may determine suspected sources of failure by combining and executing multiple rules.

[0048] Often, a model may be used to determine the likelihood that each inferred source is an actual source of the failure by combining multiple rules. The likelihood that each inferred source is an actual source of the failure may be expressed as a confidence level. The confidence level may, for example, but is not limited to, identify a strength of correlation between each inferred source of the failure and the actual source of the failure, a low false positive rate, or a high true positive rate based on near real-time data. For example, the confidence level of each inferred source of the failure may be estimated based on the number of rules that identify each source as an inferred source of the failure. The confidence level may be estimated based on the number of rules that have been executed to identify each device / application as an inferred source of the failure.

[0049] For example, a first suspected source of the failure may include a server, and a second suspected source of the failure may include a network switch for communicating data to and from the server. In this example, two rules implemented by the model (e.g., a rule for identifying an abnormal temperature level of the server and a rule for identifying a loss of functionality of an application running on the server) may identify the server as the first suspected source of the failure. Further, in this example, one rule (e.g., a rule for identifying that data communication throughput from a port corresponding to the server is lower than an estimated level) may identify the network switch as the second suspected source of the failure. In this example, the first suspected source of the failure may have a higher confidence level than the second suspected source of the failure.

[0050] At 228, a failure may be detected. A failure may include any identified loss of functionality implemented by any device in the data center. Exemplary failures may result from an overheating server, unavailability of an application running on the server, lack of data communication with the server / application running on the server, etc.

[0051] In some embodiments, the failure may be detected manually by an operator indicating the failure occurred. In other embodiments, the failure may be detected automatically, for example, by processing tickets or other near real-time data to detect loss of functionality or data communication of equipment / applications in the data center. The model may be configured to detect the failure by processing the input data. A process for estimating one or more suspected sources of the failure may be executed in response to detecting the failure.

[0052] At 230, an alert may be generated. The alert may provide a notification to an operator to identify the fault, one or more potential sources of the fault, and known solutions to the fault. For example, the alert may provide a description of the fault, one or more potential sources of the fault (e.g., inferred from the model 224), any solution data for resolving the fault, a depiction of one or more parameters to attest to the potential source of the fault, etc. Alerts are described in more detail with reference to FIG.

[0053] 3 is a block diagram illustrating an example fault detection service 314. As described above, the fault detection service 314 may be implemented on one or more interconnected computing devices outside of a data center. As described herein, the fault detection service 314 may estimate one or more suspected sources of a fault by obtaining input parameters (e.g., near real-time data, offline data) and processing the input parameters with a model.

[0054] The fault detection service 314 may obtain near real-time data from the problem detection service 302. The problem detection service 302 may obtain near real-time data (e.g., server temperature data, rack power usage data, data center sensor data). As described with reference to FIG. 2, the problem detection service 302 may provide any near real-time data 202. In some cases, the problem detection service 302 may classify the near real-time data according to data type for subsequent storage and processing by the fault detection service 314.

[0055] The near real-time data sent by the problem detection service may be forwarded to the fault detection service 314 via the streaming service 304. The streaming service 304 may enable low-latency data transmission between the problem detection service 302 and the fault detection service 314. For example, the streaming service 304 may include an API that provides a low-latency connection between the problem detection service 302 and the fault detection service 314.

[0056] The telemetry service 306 may generate and provide a set of power-related parameters of the power sources in the data center to the fault detection service 314. For example, the telemetry service 306 may provide multiple power parameters (e.g., voltage, current, resistance, power) of each power source (e.g., rack power units 106a-n).

[0057] The resource management service 308 may monitor and track components within a data center and the location of each component within the data center. For example, the resource management service 308 may maintain a list of the location and identifier of each server within each rack in the data center and a list of all power sources that power the corresponding servers. The resource management service 308 may maintain a list of the location of any device within the data center, the applications running on each server within the data center, and all devices that are directly connected to other devices in the data center.

[0058] The ticket data service 310 may obtain and process tickets received in association with the data center. For example, in response to a failure to execute an application or provide data to an external device, a ticket may be generated to identify the failure. As another example, a client may request the generation of a ticket in response to a failure of an application or device to provide a particular function. The ticket data service 310 may aggregate and identify characteristics of each ticket received. As described herein, the ticket data service 310 may identify a particular application / device, etc., by analyzing characteristics from each ticket received, and may provide insight into a probable cause of the failure. Data obtained from the telemetry service 306, the resource management service 308, and the ticket data 310 may be stored in the object storage 312. The object storage 312 may include a database that arranges the received data according to data type and time the data was obtained.

[0059] The fault detection service 314 may obtain near real-time data (e.g., temperature data 316, device power data 318, rack power data 320, location data 322) and arrange the data according to data type. For example, the near real-time data may be processed to determine characteristics associated with each piece of data, such as the data type (e.g., temperature, power), the device / component associated with each piece of data, the time the data was obtained, etc.

[0060] The fault detection service 314 may move and transform the received data (e.g., temperature data 316, device power data 318, rack power data 320, location data 322) by performing an extract, transform, and load (ETL) process 324. For example, the ETL process 324 may take the near real-time data and identify the data type associated with each portion of the data. The ETL process 324 may also associate devices / components with the various portions of the data (e.g., with a set of devices from the resource management service 308) and assign timestamps to the various portions of the data. The processed data may be stored in a database 326 that provides associated temperature and power data.

[0061] The fault detection service 314 may infer one or more probable sources of the fault by processing stored data (e.g., data stored in database 326) with model 328. In some embodiments, model 328 may process the input data and determine whether a fault has occurred and / or identify characteristics associated with the fault (by processing ticket data or by identifying anomalous parameters in near real-time data).

[0062] The model 328 may be retrieved from a model store 332, which may store various types of machine learning models. The model 328 may incorporate various rules from the rule store 330, which are executed by the model. For example, the rules may identify an abnormal parameter (e.g., a threshold deviation of the parameter from an estimated level at a particular time), identify a device corresponding to the abnormal parameter, and a device / associated application corresponding to a received ticket, etc. The model 328 may output one or more inferred sources of the failure and a confidence level that identifies a confidence that the inferred source of the failure corresponds to a failure. For example, the model may identify a power surge in a first power source by processing near real-time data and determine, using rules from the rule store 330, that the inferred source of the failure includes the first power source. In some cases, data from previous failures (e.g., abnormal parameters from the failure, known solutions to the failure) may be fed back to the model store 332 / rule store 330 to incrementally add rules to identify the source of the failure.

[0063] In some embodiments, the fault detection service 314 may identify solution data corresponding to the suspected source of the fault. For example, if the model 328 identifies a first power supply as the suspected source of the fault, the fault detection service 314 may obtain the corresponding solution data (e.g., resetting the power supply, replacing the power supply) by retrieving the solution data from the solution data database 334. As another example, if the model 328 identifies a newly modified application running on a set of services as the suspected source of the fault, the fault detection service 314 may obtain the corresponding solution data (e.g., reverting the application changes) by retrieving the solution data from the solution data database 334.

[0064] The fault detection service may generate alerts via the alert service 336. The alerts may provide a message describing the fault, the suspected sources of the fault, a confidence value associated with each suspected source of the fault, solution data related to each suspected source of the fault, etc.

[0065] 4 illustrates an example alert 400. The alert 400 may include a message (e.g., an email message, a text message, or a graphical output on a device associated with an operator). The alert 400 may provide multiple sources of data related to the failure. For example, the alert 400 may provide failure data 402 to identify characteristics of the failure (e.g., time of detecting the failure, device / application affected by the failure). The failure data 402 may include data provided by a client or operator, data inferred from a model identifying characteristics of the failure, etc.

[0066] The alert 400 may provide inferred fault sources 404 that identify one or more inferred sources of the fault. For example, the alert may provide a list of inferred sources of the fault (e.g., 404) and data points (e.g., anomaly parameters, ticket data) for identifying each source as a inferred source of the fault. The alert 400 may also include one or more confidence values ​​406 associated with each inferred source of the fault and that identify an estimated likelihood that each inferred source of the fault actually includes a source of the fault.

[0067] The alert 400 may provide a graphical representation 408 of one or more parameters corresponding to the suspected source of the fault. For example, if the suspected source of the fault is a rack power supply, the alert 400 may include a graphical representation 408 of the power parameter. In this example, the graphical representation 408 may provide an actual power level 410 compared to an estimated power level 412 (e.g., estimated from past power levels) during a period of time (e.g., time points T1-T6). In some examples, the past power level includes a previous level received from a previous time. Also, in this example, the graphical representation 408 may show multiple power level anomaly deviations 414a-b from the estimated power level 412. The multiple power level anomaly deviations 414a-b may provide insight that the power supply associated with the power level 410 was the source of the fault.

[0068] 5 is a block diagram 500 illustrating an example method for estimating one or more suspected sources of a fault in a data center. Cloud infrastructure nodes may implement a fault detection service configured to perform the methods described herein.

[0069] At block 502, the method may include obtaining a set of input data providing various parameters related to a data center, a set of devices in the data center, and applications running on the devices. As described above with reference to FIG. 2, the set of input data may include near real-time data (e.g., 202) and offline data (e.g., 204). In some embodiments, the set of input data may identify any one of the following: a temperature of each server in the data center, a power level of each power supply in each rack in the data center, climate data for the data center obtained from a set of sensors in the data center, ticket data identifying any function of the data center obtained, a set of devices in the data center, and a location of all devices in the data center.

[0070] In some embodiments, the set of input data includes a location of each device in the data center and a device type for each device in the data center. The method may include identifying a data type and one or more associated devices associated with each portion of the set of input data by processing the set of input data. Exemplary data types may include server temperature data (e.g., 206), server power usage data (e.g., 208), device data (e.g., 216), etc. The method may also assign a timestamp to each portion of the set of input data indicating a time at which each portion of the set of input data was acquired. By arranging the data of a particular type according to the timestamp, the fault detection service may estimate a trend of a parameter over a period of time (e.g., identify a change in a parameter over time). The method may also include storing the set of input data in a database (e.g., database 326) according to the data type and the assigned timestamp. The fault detection service may use the stored data as input for a model to estimate a probable source of the fault.

[0071] At block 504, the method may include detecting a failure of at least one function of the data center. The failure may be due to a loss of functionality (e.g., an application) or a loss of computing resources (e.g., a loss of connection to a server, a loss of power to a server). Detecting the failure may include obtaining a failure notification from an external computing device identifying that a failure has occurred or detecting that a threshold number of tickets have been received identifying the loss of at least one functionality or loss of some computing resources of the data center.

[0072] At block 506, the method may include estimating one or more suspected sources of the failure by processing the set of input data with the model. The model may incorporate a number of rules that identify a correlation between the set of input data and a device or an application running on the device as the one or more suspected sources of the failure.

[0073] In block 508, estimating one or more suspected sources of the failure may include generating an estimated level for each parameter using the set of rules available to the model and historical data associated with each parameter included in the set of input data. For example, an estimated temperature level for the first server may be determined by processing historical server temperature data for the first server. Comparing the estimated level for each parameter to a detection level may identify whether any parameter deviates from the estimated level. Such deviations may indicate that a device or application is a suspected source of the failure.

[0074] In block 510, estimating one or more suspected sources of the failure may include identifying one or more anomalous parameters having actual levels that have a threshold deviation from the corresponding estimated levels by comparing the estimated levels of each parameter with the actual levels of each parameter included in the set of input data. For example, a parameter may be anomalous if its actual level has a threshold deviation from the estimated level of that parameter. A parameter having a threshold deviation from an estimated level may indicate an overheating server, a power surge in a power supply, lost network packets, etc.

[0075] In block 512, estimating one or more suspected sources of the failure may include identifying one or more devices and / or applications corresponding to each of the identified anomalous parameters. Each of the identified one or more devices and / or applications may be included as one or more suspected sources of the failure. For example, in response to determining that a server temperature level of a first server has suddenly increased above an estimated level, the model may identify the first server as a suspected source of the failure.

[0076] In some embodiments, the set of rules is estimated based at least in part on a correlation between previously resolved failures and the identified source of each failure of the previously resolved failures. In these embodiments, the method may include using the model to identify, using a first rule of the set of rules, a first abnormal parameter associated with a first application running on a portion of the servers in the data center. For example, a change to an application running on the set of servers may cause an update. In this example, the model may identify an abnormal parameter associated with the application, e.g., an increase in server temperature level, a loss of data packet transmission.

[0077] In these embodiments, the method may include using the model using a second rule of the set of rules to identify that a change to the execution of the first application occurred within a threshold period of time of detecting the failure. For example, the model may identify that a change to the application occurred within the time of detecting the failure (e.g., detect that the change occurred within 5 minutes of detecting the failure). This may indicate that the change to the application is a suspected source of the failure. The first application may be included in the failure notification message as a first suspected source of the failure. The failure notification message may further provide solution data identifying instructions for reverting the first application to a previous version to remove the change to the execution of the first application.

[0078] At block 514, the method may include generating a fault notification message providing one or more suspected sources of the fault. In some embodiments, the fault notification message includes a graphical representation of the first anomaly parameter and an estimated level of the first anomaly parameter.

[0079] In some embodiments, the method may include estimating, for each inferred source of the one or more inferred sources of the fault, a confidence level based at least in part on a number of rules correlating to parameters associated with each inferred source of the fault, and the fault notification message includes the confidence level.

[0080] In some embodiments, the method may include retrieving, for each inferred source of the one or more inferred sources of the fault, solution data associated with each of the one or more inferred sources of the fault, the solution data providing known methods for resolving the fault specific to each inferred source of the fault, and the fault notification message including the solution data.

[0081] As mentioned above, Infrastructure as a Service (IaaS) is one particular type of cloud computing. IaaS may be configured to provide virtualized computing resources over a public network (e.g., the Internet). In an IaaS model, a cloud computing provider may host infrastructure elements (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, an IaaS provider may provide various services (e.g., billing, monitoring, logging, load balancing, clustering, etc.) that are attached to the infrastructure elements. Thus, because these services may be policy-driven, an IaaS user may implement policies to drive load balancing to maintain application availability and performance.

[0082] In some examples, an IaaS customer may access resources and services over a wide area network (WAN), such as the Internet, and may use the cloud provider's services to install the remaining elements of the application stack. For example, a user may log into an IaaS platform to create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and install enterprise software on the VMs. The customer may use the provider's services to perform a variety of functions, including balancing network traffic, troubleshooting applications, monitoring performance, managing disaster recovery, etc.

[0083] In most cases, the cloud computing model requires the participation of a cloud provider, which may be, but does not have to be, a third-party service dedicated to providing (e.g., offering, renting, selling) IaaS. Enterprises can also deploy private clouds and become providers of infrastructure services.

[0084] In some examples, IaaS deployment is the process of deploying a new application or a new version of an application onto provisioned application servers, etc. IaaS deployment may include the process of provisioning a server (e.g., installing libraries, daemons, etc.). IaaS deployment is often managed below the hypervisor layer (e.g., servers, storage, network hardware, and virtualization) by the cloud provider. Thus, the customer may perform OS, middleware, and / or application deployment (e.g., self-service virtual machines (e.g., that can be spun up on demand), etc.

[0085] In some instances, IaaS provisioning may include obtaining a computer or virtual host to be used and installing the necessary libraries or services on the computer or virtual host. In most cases, deployment does not include provisioning, which must be performed first.

[0086] In some cases, there are two different challenges in IaaS provisioning. First, there is the challenge of provisioning an initial set of infrastructure before doing anything. Second, there is the challenge of evolving the existing infrastructure (e.g., adding new services, modifying services, removing services) after everything is provisioned. In some cases, these two challenges may be addressed by allowing the configuration of the infrastructure to be defined declaratively. In other words, the infrastructure (e.g., which elements are needed and how these elements interact) may be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., which resources depend on which, how they work together) may be described declaratively. In some instances, once the topology is defined, workflows may be generated to create and / or manage the different elements described in the configuration files.

[0087] In some examples, the infrastructure may include many interconnected elements. For example, there may be one or more virtual private clouds (VPCs) (e.g., potentially on-demand pools of configurable and / or shared computing resources), also known as a core network. In some examples, there may be one or more inbound and / or outbound traffic group rules and one or more virtual machines (VMs) that are provisioned to define how to configure the inbound and / or outbound traffic of the network. Other infrastructure elements such as load balancers, databases, etc. may also be provisioned. The infrastructure may evolve incrementally as more and more infrastructure elements are desired and / or added.

[0088] In some examples, continuous deployment techniques may be employed to enable deployment of infrastructure code across various virtual computing environments. The described techniques may also enable infrastructure management within these environments. In some examples, a service team may write code that is desired to be deployed to one or more, typically many, different production environments (e.g., across a variety of different geographic locations, sometimes spanning the globe). However, in some examples, the infrastructure for deploying the code must first be set up. In some examples, provisioning may be done manually, resources may be provisioned using a provisioning tool, and / or the code may be deployed using a deployment tool after the infrastructure has been provisioned.

[0089] 6 is a block diagram 600 illustrating an example pattern of an IaaS architecture according to at least one embodiment. A service operator 602 may be communicatively connected to a secure host tenancy 604, which may include a virtual cloud network (VCN) 606 and a secure host subnet 608. In some examples, the service operator 602 may use one or more client computing devices. The one or more client computing devices may run software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android®, BlackBerry 8, and Palm OS, and may be Internet, email, short message service (SMS), BlackBerry®, or other communication protocol enabled handheld mobile devices (e.g., iPhone®, cell phone, iPad®, tablet, personal digital assistant (PDA) or wearable devices (Google Glass® head mounted display). The client computing devices may be general purpose personal computers including, by way of example, personal computers and / or laptop computers running various versions of Microsoft Windows® operating systems, Apple Macintosh® operating systems, and / or Linux® operating systems. Alternatively, the client computing devices may be workstation computers running various commercially available UNIX® or UNIX-like operating systems including, but not limited to, various GNU / Linux operating systems, e.g., Google Chrome® OS.Alternatively or additionally, the client computing devices may be other electronic devices capable of communicating over VCN 606 and / or an Internet-accessible network, such as thin-client computers, Internet-enabled gaming systems (e.g., Microsoft Xbox® game consoles with or without Kinect® gesture input devices), and / or personal messaging devices.

[0090] The VCN 606 may include a local peering gateway (LPG) 610, which may be communicatively connected to a secure shell (SSH) VCN 612 via an LPG 610 included in the SSH VCN 612. The SSH VCN 612 may include an SSH subnet 614, which may be communicatively connected to a control plane VCN 616 via an LPG 610 included in the control plane VCN 616. The SSH VCN 612 may also be communicatively connected to a data plane VCN 618 via the LPG 610. The control plane VCN 616 and the data plane VCN 618 may be included in a service tenancy 619, which may be owned and / or operated by the IaaS provider.

[0091] The control plane VCN 616 may include a control plane demilitarized zone (DMZ) layer 620 that functions as a perimeter network (e.g., a portion of an enterprise network between an enterprise intranet and an external network). DMZ-based servers have a certain reliability and can contain security breaches. In addition, the DMZ layer 620 may include one or more load balancer (LB) subnets 622, a control plane app layer 624 that may include an app subnet 626, and a control plane data layer 628 that may include a database (DB) subnet 630 (e.g., a front-end DB subnet and / or a back-end DB subnet). The LB subnet 622 included in the control plane DMZ layer 620 may be communicatively connected to the app subnet 626 included in the control plane app layer 624 and an Internet gateway 634 that may be included in the control plane VCN 616, and the app subnet 626 may be communicatively connected to the DB subnet 630, a service gateway 636, and a network address translation (NAT) gateway 638 included in the control plane data layer 628. The control plane VCN 616 may include a service gateway 636 and a NAT gateway 638.

[0092] The control plane VCN 616 can include a data plane mirror app layer 640, which can include an app subnet 626. The app subnet 626 included in the data plane mirror app layer 640 can include a virtual network interface controller (VNIC) 642 that can run a compute instance 644. The compute instance 644 can communicatively connect the app subnet 626 of the data plane mirror app layer 640 to the app subnet 626 that can be included in the data plane app layer 646.

[0093] The data plane VCN 618 may include a data plane app layer 646, a data plane DMZ layer 648, and a data plane data layer 650. The data plane DMZ layer 648 may include a LB subnet 622 that may be communicatively connected to an app subnet 626 of the data plane app layer 646 and an Internet gateway 634 of the data plane VCN 618. The app subnet 626 may be communicatively connected to a service gateway 636 of the data plane VCN 618 and a NAT gateway 638 of the data plane VCN 618. Additionally, the data plane data layer 650 may include a DB subnet 630 that may be communicatively connected to the app subnet 626 of the data plane app layer 646.

[0094] The Internet gateway 634 of the control plane VCN 616 and the Internet gateway 634 of the data plane VCN 618 may be communicatively connected to a metadata management service 652, which may be communicatively connected to the public Internet 654. The public Internet 654 may be communicatively connected to a NAT gateway 638 of the control plane VCN 616 and the NAT gateway 638 of the data plane VCN 618. The service gateway 636 of the control plane VCN 616 and the service gateway 636 of the data plane VCN 618 may be communicatively connected to cloud services 656.

[0095] In some examples, a service gateway 636 of a control plane VCN 616 or a data plane VCN 618 can make application programming interface (API) calls to cloud services 656 without traversing the public Internet 654. API calls from the service gateway 636 to the cloud services 656 can be one-way. The service gateway 636 can make API calls to the cloud services 656, and the cloud services 656 can send request data to the service gateway 636. However, the cloud services 656 may not initiate the API calls to the service gateway 636.

[0096] In some examples, secure host tenancy 604 may be directly connected to service tenancy 619, which may otherwise be separate. Secure host subnet 608 may communicate with SSH subnet 614 through LPG 610, which allows bidirectional communication with otherwise separate systems. By connecting secure host subnet 608 to SSH subnet 614, secure host subnet 608 may access other entities in service tenancy 619.

[0097] The control plane VCN 616 allows users of the service tenancy 619 to configure or provision desired resources. The desired resources provisioned in the control plane VCN 616 may be deployed or used in the data plane VCN 618. In some examples, the control plane VCN 616 may be isolated from the data plane VCN 618, and the data plane mirror app layer 640 of the control plane VCN 616 may communicate with the data plane app layer 646 of the data plane VCN 618 via a VNIC 642, which may be included in the data plane mirror app layer 640 and the data plane app layer 646.

[0098] In some examples, a user or customer of the system may make a request, such as, for example, a create, read, update, or delete (CRUD) operation, via the public Internet 654, which may communicate the request to a metadata management service 652. The metadata management service 652 may communicate the request to the control plane VCN 616 via an Internet gateway 634. The request may be received by a LB subnet 622 included in the control plane DMZ layer 620. The LB subnet 622 may determine that the request is valid, and in response to this determination, the LB subnet 622 may send the request to an app subnet 626 included in the control plane app layer 624. If the request is validated and requires a call to the public Internet 654, it may send the call to the public Internet 654 to a NAT gateway 638, which may make the call to the public Internet 654. The metadata to be stored by the request may be stored in the DB subnet 630.

[0099] In some examples, the data plane mirror app layer 640 may facilitate direct communication between the control plane VCN 616 and the data plane VCN 618. For example, it may be desirable for changes, updates, or other suitable modifications to a configuration to be applied to resources included in the data plane VCN 618. The control plane VCN 616 may communicate directly with resources included in the data plane VCN 618 via the VNIC 642 and thus may effect the changes, updates, or other suitable modifications to the configuration.

[0100] In some embodiments, the control plane VCN 616 and the data plane VCN 618 may be included in the service tenancy 619. In this case, a user or customer of the system may not own or operate either the control plane VCN 616 or the data plane VCN 618. Instead, the IaaS provider may own or operate the control plane VCN 616 and the data plane VCN 618, both of which may be included in the service tenancy 619. This embodiment may prevent a user or customer from interacting with other users' resources or other customers' resources by enabling network isolation. This embodiment may also allow a user or customer of the system to store databases privately without having to rely on the public Internet 654, which may not have the desired level of threat protection for storage.

[0101] In another embodiment, the LB subnet 622 included in the control plane VCN 616 may be configured to receive signals from the service gateway 636. In this embodiment, the control plane VCN 616 and the data plane VCN 618 may be configured to be called by the IaaS provider's customers without calling the public Internet 654. The IaaS provider's customers may desire this embodiment because databases used by the customers may be stored in the service tenancy 619, which may be controlled by the IaaS provider and isolated from the public Internet 654.

[0102] 7 is a block diagram 700 illustrating another example parameter of an IaaS architecture according to at least one embodiment. A service operator 702 (e.g., service operator 602 of FIG. 6) may be communicatively connected to a secure host tenancy 704 (e.g., secure host tenancy 604 of FIG. 6), which may include a virtual cloud network (VCN) 706 (e.g., VCN 606 of FIG. 6) and a secure host subnet 708 (e.g., secure host subnet 608 of FIG. 6). The VCN 706 may include a local peering gateway (LPG) 710 (e.g., LPG 610 of FIG. 6), which may be communicatively connected to a secure shell (SSH) VCN 712 (e.g., SSH VCN 612 of FIG. 6) via an LPG 610 included in the SSH VCN 712. SSH VCN 712 can include an SSH subnet 714 (e.g., SSH subnet 614 in FIG. 6), which may be communicatively connected to a control plane VCN 716 (e.g., control plane VCN 616 in FIG. 6) via an LPG 710 included in the control plane VCN 716. The control plane VCN 716 may be included in a service tenancy 719 (e.g., service tenancy 619 in FIG. 6), and the data plane VCN 718 (e.g., data plane VCN 618 in FIG. 6) may be included in a customer tenancy 721, which may be owned or operated by a user or customer of the system.

[0103] The control plane VCN 716 may include a control plane DMZ layer 720 (e.g., the control plane DMZ layer 620 of FIG. 6 ) that may include a LB subnet 722 (e.g., the LB subnet 622 of FIG. 6 ), a control plane app layer 716 (e.g., the control plane app layer 624 of FIG. 6 ) that may include an app subnet 726 (e.g., the app subnet 626 of FIG. 6 ), and a control plane data layer 728 (e.g., the control plane data layer 628 of FIG. 6 ) that may include a database (DB) subnet 730 (e.g., similar to the DB subnet 630 of FIG. 6 ). The LB subnet 722 included in the control plane DMZ layer 720 may be communicatively connected to the app subnet 726 included in the control plane app layer 716 and to an Internet gateway 734 (e.g., the Internet gateway 634 of FIG. 6 ) that may be included in the control plane VCN 716. The app subnet 726 may be communicatively connected to a DB subnet 730, a service gateway 736 (e.g., the service gateway of FIG. 6), and a network address translation (NAT) gateway 738 (e.g., the NAT gateway 638 of FIG. 6), which are included in the control plane data layer 728. The control plane VCN 716 may include the service gateway 736 and the NAT gateway 738.

[0104] The control plane VCN 716 may include a data plane mirror app layer 740 (e.g., data plane mirror app layer 640 of FIG. 6 ), which may include an app subnet 726. The app subnet 726 included in the data plane mirror app layer 740 may include a virtual network interface controller (VNIC) 742 (e.g., VNIC 642) that may run a compute instance 744 (e.g., similar to compute instance 644 of FIG. 6 ). The compute instance 744 may facilitate communication between the app subnet 726 of the data plane mirror app layer 740 and the app subnet 726 included in the data plane app layer 746 (e.g., data plane app layer 646 of FIG. 6 ) via the VNIC 742 included in the data plane mirror app layer 740 and the VNIC 742 included in the data plane app layer 746.

[0105] The Internet gateway 734 included in the control plane VCN 716 may be communicatively connected to a metadata management service 752 (e.g., metadata management service 652 of FIG. 6), which may be communicatively connected to a public Internet 754 (e.g., public Internet 654 of FIG. 6). The public Internet 754 may be communicatively connected to a NAT gateway 738 included in the control plane VCN 716. The service gateway 736 included in the control plane VCN 716 may be communicatively connected to cloud services 756 (e.g., cloud services 656 of FIG. 6).

[0106] In some examples, the data plane VCN 718 may be included in the customer tenancy 721. In this case, the IaaS provider may provide a control plane VCN 716 for each customer, and the IaaS provider may configure a unique compute instance 744 for each customer, which is included in the service tenancy 719. Each compute instance 744 may allow communication between the control plane VCN 716 included in the service tenancy 719 and the data plane VCN 718 included in the customer tenancy 721. The compute instance 744 may allow resources provisioned in the control plane VCN 716 included in the service tenancy 719 to be deployed or used in the data plane VCN 718 included in the customer tenancy 721.

[0107] In another example, an IaaS provider customer may have a database that resides in customer tenancy 721. In this example, control plane VCN 716 may include data plane minor app tier 740, which may include app subnet 726. Data plane mirror app tier 740 may reside in data plane VCN 718, but data plane mirror app tier 740 may not reside in data plane VCN 718. That is, data plane mirror app tier 740 may access customer tenancy 721, but data plane mirror app tier 740 may not reside in data plane VCN 718 and may not be owned or operated by the IaaS provider customer. Data plane mirror app tier 740 may be configured to make calls to data plane VCN 718, but may not be configured to make calls to any entities included in control plane VCN 716. A customer may desire to deploy or use resources in the data plane VCN 718 provisioned to the control plane VCN 716, and the data plane mirror application tier 740 may facilitate the desired deployment or other use of the customer's resources.

[0108] In some embodiments, the IaaS provider's customer may apply filters to the data plane VCN 718. In this embodiment, the customer can determine what the data plane VCN 718 can access, and the customer may restrict access from the data plane VCN 718 to the public Internet 754. The IaaS provider may not be able to apply filters or control access from the data plane VCN 718 to any external networks or databases. The customer applying filters and controls to the data plane VCN 718 included in the customer tenancy 721 may help isolate the data plane VCN 718 from other customers and the public Internet 754.

[0109] In some embodiments, cloud services 756 may be called by service gateway 736 to access services that may not be on the public Internet 754, on the control plane VCN 716, or on the data plane VCN 718. The connection between cloud services 756 and the control plane VCN 716 or data plane VCN 718 may not be live or continuous. Cloud services 756 may be on another network owned or operated by the IaaS provider. Cloud services 756 may be configured to receive calls from service gateway 736 and may be configured not to receive calls from the public Internet 754. Some cloud services 756 may be isolated from other cloud services 756, and control plane VCN 716 may be isolated from cloud services 756 that may not be located in the same region as control plane VCN 716. For example, control plane VCN 716 may be located in “Region 1” and cloud service “Deployment 6” may be located in “Region 1” and “Region 2”. If a call to deployment 6 is made by a service gateway 736 included in a control plane VCN 716 located in region 1, the call may be sent to deployment 6 in region 1. In this example, the control plane VCN 716 or deployment 6 in region 1 may not be communicatively connected to deployment 6 in region 2.

[0110] 8 is a block diagram 800 illustrating another example pattern of an IaaS architecture according to at least one embodiment. A service operator 802 (e.g., service operator 602 of FIG. 6) may be communicatively connected to a secure host tenancy 804 (e.g., secure host tenancy 604 of FIG. 6), which may include a virtual cloud network (VCN) 806 (e.g., VCN 606 of FIG. 6) and a secure host subnet 808 (e.g., secure host subnet 608 of FIG. 6). VCN 806 may include an LPG 810 (e.g., LPG 610 of FIG. 6), which may be communicatively connected to an SSH VCN 812 (e.g., SSH VCN 612 of FIG. 6) via an LPG 810 included in SSH VCN 812. SSH VCN 812 can include an SSH subnet 814 (e.g., SSH subnet 614 in FIG. 6), which may be communicatively connected to a control plane VCN 816 (e.g., control plane VCN 616 in FIG. 6) via an LPG 810 included in the control plane VCN 816, and may be communicatively connected to a data plane VCN 818 (e.g., data plane 618 in FIG. 6) via an LPG 810 included in the data plane VCN 818. The control plane VCN 816 and the data plane VCN 818 may be included in a service tenancy 819 (e.g., service tenancy 619 in FIG. 6).

[0111] The control plane VCN 816 may include a control plane DMZ layer 820 (e.g., the control plane DMZ layer 620 of FIG. 6 ) that may include a load balancer (LB) subnet 822 (e.g., the LB subnet 622 of FIG. 6 ), a control plane app layer 824 (e.g., the control plane app layer 624 of FIG. 6 ) that may include an app subnet 826 (e.g., similar to the app subnet 626 of FIG. 6 ), and a control plane data layer 828 (e.g., the control plane data layer 628 of FIG. 6 ) that may include a DB subnet 830. The LB subnet 822 included in the control plane DMZ layer 820 may be communicatively connected to the app subnet 826 included in the control plane app layer 824 and to an Internet gateway 834 (e.g., the Internet gateway 634 of FIG. 6 ) that may be included in the control plane VCN 816. The app subnet 826 may be communicatively connected to a DB subnet 830 included in the control plane data layer 828, a service gateway 836 (e.g., the service gateway in FIG. 6) and a network address translation (NAT) gateway 838 (e.g., the NAT gateway 638 in FIG. 6). The control plane VCN 816 may include the service gateway 836 and the NAT gateway 838.

[0112] The data plane VCN 818 may include a data plane app layer 846 (e.g., data plane app layer 646 of FIG. 6), a data plane DMZ layer 848 (e.g., data plane DMZ layer 648 of FIG. 6), and a data plane data layer 850 (e.g., data plane data layer 650 of FIG. 6). The data plane DMZ layer 848 may include a LB subnetwork 822 that may be communicatively connected to the data plane app layer 846 included in the data plane VCN 818 and a trusted app subnetwork 860 and an untrusted app subnetwork 862 of an Internet gateway 834. The trusted app subnetwork 860 may be communicatively connected to a service gateway 836 included in the data plane VCN 818, a NAT gateway 838 included in the data plane VCN 818, and a DB subnetwork 830 included in the data plane data layer 850. The untrusted app subnetwork 862 may be communicatively connected to the service gateway 836 included in the data plane VCN 818 and a DB subnetwork 830 included in the data plane data layer 850. The data plane data layer 850 may include a DB subnet 830 that may be communicatively connected to a service gateway 836 included in the data plane VCN 818.

[0113] The untrusted app subnet 862 may include one or more primary VNICs 864(1)-(N) that may be communicatively connected to tenant virtual machines (VMs) 866(1)-(N). Each tenant VM 866(1)-(N) may be communicatively connected to a respective app subnet 867(1)-(N) that may be included in a respective container egress VCN 868(1)-(N) that may be included in a respective customer tenancy 870(1)-(N). Each secondary VNIC 872(1)-(N) may facilitate communication between the untrusted app subnet 862 included in the data plane VCN 818 and the app subnet included in the container egress VCN 868(1)-(N). Each container egress VCN 868(1)-(N) may include a NAT gateway 838 that may be communicatively connected to the public Internet 854 (e.g., public Internet 654 of FIG. 6).

[0114] The internet gateway 834 included in the control plane VCN 816 and the internet gateway 834 included in the data plane VCN 818 may be communicatively connected to a metadata management service 852 (e.g., metadata management system 652 of FIG. 6 ), which may be communicatively connected to the public internet 854. The public internet 854 may be communicatively connected to a NAT gateway 838 included in the control plane VCN 816 and the NAT gateway 838 included in the data plane VCN 818. The service gateway 836 included in the control plane VCN 816 and the service gateway 836 included in the data plane VCN 818 may be communicatively connected to cloud services 856.

[0115] In some embodiments, data plane VCN 818 may be consolidated into customer tenancies 870. This consolidation may be useful or desirable for an IaaS provider's customers in some cases, such as when they may want support when executing code. Customers may provide code that, when executed, may be disruptive, may communicate with other customer resources, or may cause undesirable effects. Thus, the IaaS provider may determine whether to execute code that a customer has provided to the IaaS provider.

[0116] In some examples, an IaaS provider's customer may grant temporary network access to the IaaS provider and may request a function to be added to the data plane app layer 846. The code to perform the function may be executed in VMs 866(1)-(N) but may not be configured to run elsewhere on the data plane VCN 818. Each VM 866(1)-(N) may be connected to one customer tenancy 870. Each container 871(1)-(N) included in a VM 866(1)-(N) may be configured to execute code. In this case, there may be double isolation (e.g., containers 871(1)-(N) may execute code, and containers 871(1)-(N) may be included in at least VMs 866(1)-(N) included in the untrusted app subnet 862), which may help prevent erroneous or unwanted code from damaging the IaaS provider's network or from damaging a different customer's network. Containers 871(1)-(N) may be communicatively connected to customer tenancy 870 and may be configured to send or receive data from customer tenancy 870. Containers 871(1)-(N) may not be configured to send or receive data from any other entity in data plane VCN 818. Once code execution is complete, the IaaS provider may kill or discard containers 871(I)-(N).

[0117] In some embodiments, trusted app subnet 860 may execute code that may be owned or operated by the IaaS provider. In this embodiment, trusted app subnet 860 may be communicatively connected to DB subnet 830 and configured to perform CRUD operations in DB subnet 830. Untrusted app subnet 862 may be communicatively connected to DB subnet 830, but in this embodiment, untrusted app subnet 862 may be configured to perform read operations in DB subnet 830. Containers 871(1)-(N) included in each customer's VMs 866(1)-(N) that may execute code from the customer may not be communicatively connected to DB subnet 830.

[0118] In other embodiments, the control plane VCN 816 and the data plane VCN 818 may not be directly communicatively coupled. In this embodiment, there may be no direct communication between the control plane VCN 816 and the data plane VCN 818. However, there may be indirect communication by at least one method. An LPG 810 may be established by an IaaS provider that may facilitate communication between the control plane VCN 816 and the data plane VCN 818. In another example, the control plane VCN 816 or the data plane VCN 818 may make a call to a cloud service 856 through a service gateway 836. For example, a call from the control plane VCN 816 to the cloud service 856 may include a request for a service that may communicate with the data plane VCN 818.

[0119] 9 is a block diagram 900 illustrating another example parameter of an IaaS architecture according to at least one embodiment. A service operator 902 (e.g., service operator 602 of FIG. 6) may be communicatively connected to a secure host tenancy 904 (e.g., secure host tenancy 604 of FIG. 6), which may include a virtual cloud network (VCN) 906 (e.g., VCN 606 of FIG. 6) and a secure host subnet 908 (e.g., secure host subnet 608 of FIG. 6). VCN 906 may include an LPG 910 (e.g., LPG 610 of FIG. 6), which may be communicatively connected to an SSH VCN 912 (e.g., SSH VCN 612 of FIG. 6) via an LPG 910 included in SSH VCN 912. SSH VCN 912 can include an SSH subnet 914 (e.g., SSH subnet 614 in FIG. 6), which may be communicatively connected to a control plane VCN 916 (e.g., control plane VCN 616 in FIG. 6) via an LPG 910 included in the control plane VCN 916, and to a data plane VCN 918 (e.g., data plane 618 in FIG. 6) via an LPG 910 included in the data plane VCN 918. The control plane VCN 916 and the data plane VCN 918 may be included in a service tenancy 919 (e.g., service tenancy 619 in FIG. 6).

[0120] The control plane VCN 916 may include a control plane DMZ layer 920 (e.g., the control plane DMZ layer 620 of FIG. 6 ) that may include a LB subnet 922 (e.g., the LB subnet 622 of FIG. 6 ), a control plane app layer 924 (e.g., the control plane app layer 624 of FIG. 6 ) that may include an app subnet 926 (e.g., the app subnet 626 of FIG. 6 ), and a control plane data layer 928 (e.g., the control plane data layer 628 of FIG. 6 ) that may include a DB subnet 930 (e.g., the DB subnet 830 of FIG. 8 ). The LB subnet 922 included in the control plane DMZ layer 920 may be communicatively connected to the app subnet 926 included in the control plane app layer 924 and to an Internet gateway 934 (e.g., the Internet gateway 634 of FIG. 6 ) that may be included in the control plane VCN 916. The app subnet 926 may be communicatively connected to a DB subnet 930 included in the control plane data layer 928, a service gateway 936 (e.g., the service gateway of FIG. 6) and a network address translation (NAT) gateway 938 (e.g., the NAT gateway 638 of FIG. 6). The control plane VCN 916 may include the service gateway 936 and the NAT gateway 938.

[0121] The data plane VCN 918 may include a data plane app layer 946 (e.g., data plane app layer 646 of FIG. 6), a data plane DMZ layer 948 (e.g., data plane DMZ layer 648 of FIG. 6), and a data plane data layer 950 (e.g., data plane data layer 650 of FIG. 6). The data plane DMZ layer 948 may include a trusted app subnet 960 (e.g., trusted app subnet 860 of FIG. 8) and an untrusted app subnet 962 (e.g., untrusted app subnet 862 of FIG. 8) of the data plane app layer 946 and a LB subnet 922 that may be communicatively connected to an Internet gateway 934 included in the data plane VCN 918. The trusted app subnet 960 may be communicatively connected to a service gateway 936 included in the data plane VCN 918, a NAT gateway 938 included in the data plane VCN 918, and a DB subnet 930 included in the data plane data layer 950. The untrusted app subnet 962 may be communicatively connected to a service gateway 936 included in the data plane VCN 918 and to a DB subnet 930 included in the data plane data layer 950. The data plane data layer 950 may include a DB subnet 930 that may be communicatively connected to a service gateway 936 included in the data plane VCN 918.

[0122] The untrusted app subnet 962 may include primary YNICs 964(1)-(N) that may be communicatively connected to tenant virtual machines (VMs) 966(1)-(N) that reside in the untrusted app subnet 962. Each tenant VM 966(1)-(N) may execute code in a respective container 967(1)-(N) and may be communicatively connected to an app subnet 926 that may be included in a data plane app layer 946 that may be included in a container egress VCN 968. Each secondary VNIC 972(1)-(N) may facilitate communication between the untrusted app subnet 962 included in the data plane VCN 918 and the app subnet included in the container egress VCN 968. The container egress VCN may include a NAT gateway 938 that may be communicatively connected to the public Internet 954 (e.g., public Internet 654 of FIG. 6).

[0123] The internet gateway 934 included in the control plane VCN 916 and the internet gateway 934 included in the data plane VCN 918 may be communicatively connected to a metadata management service 952 (e.g., metadata management system 652 of FIG. 6 ), which may be communicatively connected to the public internet 954. The public internet 954 may be communicatively connected to the internet gateway 934 included in the control plane VCN 916 and the NAT gateway 938 included in the data plane VCN 918. The internet gateway 934 included in the control plane VCN 916 and the service gateway 936 included in the data plane VCN 918 may be communicatively connected to cloud services 956.

[0124] In some examples, the pattern illustrated by the architecture of block diagram 900 of FIG. 9 may be considered an exception to the pattern illustrated by the architecture of block diagram 800 of FIG. 8 and may be desirable for an IaaS provider's customer when the IaaS provider cannot communicate directly with the customer (e.g., in a non-connected region). The customer may access each container 967(1)-(N) included in each customer's VM 966(1)-(N) in real time. The containers 967(1)-(N) may be configured to call each secondary VNIC 972(1)-(N) included in the app subnet 926 of the data plane app tier 946, which may be included in a container egress VCN 968. The secondary VNICs 972(1)-(N) may send the call to a NAT gateway 938, which may send the call to the public Internet 954. In this example, the containers 967(1)-(N) that the customer can access in real time may be isolated from the control plane VCN 916 and may be isolated from other entities included in the data plane VCN 918. Additionally, containers 967(1)-(N) may be isolated from resources of other customers.

[0125] In another example, a customer can invoke cloud service 956 using containers 967(1)-(N). In this example, the customer can execute code in containers 967(1)-(N) that requests a service from cloud service 956. Containers 967(1)-(N) can send the request to secondary VNICs 972(1)-(N), which can send the request to a NAT gateway, which can send the request to public internet 954. Public internet 954 can send the request to LB subnet 922, which is included in control plane VCN 916, via internet gateway 934. In response to determining that the request is valid, the LB subnet can send the request to app subnet 926, which can send the request to cloud service 956 via service gateway 936.

[0126] It should be noted that the illustrated IaaS architectures 600, 700, 800, and 900 may include elements other than those illustrated, and the illustrated embodiments are only some examples of cloud infrastructure systems that may incorporate embodiments of the present disclosure. In other embodiments, the IaaS system may have more or fewer elements than those illustrated, may combine two or more elements, or may have a different configuration or arrangement of elements.

[0127] In certain embodiments, the IaaS system described in this disclosure may include a suite of application, middleware, and database services that are delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. One example of such an IaaS system is the Oracle® Cloud Infrastructure (OCI) offered by the present applicant.

[0128] 10 illustrates an exemplary computer system 1000 in which various embodiments may be implemented. System 1000 may be used to implement any of the computer systems described above. As illustrated, computer system 1000 includes a processing unit 1004 that communicates with a number of peripheral subsystems via a bus subsystem 1002. These peripheral subsystems may include a processing acceleration unit 1006, an I / O subsystem 1008, a storage subsystem 1018, and a communication subsystem 1024. The storage subsystem 1018 includes a tangible computer-readable storage medium 1022 and a system memory 1010.

[0129] Bus subsystem 1002 provides a mechanism for allowing the various components and subsystems of computer system 1000 to communicate with each other as intended. Although bus subsystem 1002 is shown diagrammatically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 1002 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include a Peripheral Component Interconnect (PCI) bus, which may be implemented as an Industry Standard Architecture (ISA) bus, a MicroChannel Architecture (MCA), a Bus Extended ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard, among others.

[0130] Processing unit 1004, which may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of computer system 1000. Processing unit 1004 may include one or more processors. These processors may include single-core or multi-core processors. In some embodiments, processing unit 1004 may be implemented as one or more independent processing units 1032 and / or 1034 with single-core or multi-core processors included in each processing unit. In other embodiments, processing unit 1004 may be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.

[0131] In various embodiments, the processing unit 1004 may execute various programs in response to program code and may maintain multiple programs or processes executing simultaneously. At any given time, some or all of the program code being executed may reside in the processor 1004 and / or the storage subsystem 1018. The processor 1004 may provide various functionality as discussed above through appropriate programming. The computer system 1000 may further include a processing acceleration unit 1006, which may include a digital signal processor (DSP), a special purpose processor, and / or the like.

[0132] The I / O subsystem 1008 may include user interface input devices and user interface output devices. User interface input devices may include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touch screen integrated into a display, a scroll wheel, a click wheel, a dial, a button, a switch, a keypad, a voice input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognizers, such as, for example, a Microsoft Kinect® motion sensor. The Microsoft Kinect® motion sensor may control and interact with input devices, such as a Microsoft Xbox® 360 game controller, through a natural user interface (NUI) that utilizes gestures and voice commands. User interface input devices may also include an eye gesture recognizer, such as a Google Glass® blink detector. The Google Glass® blink detector detects a user's eye activity (e.g., "blinking" when taking a photo and / or selecting a menu) and converts the eye activity into input that is entered into an input device (e.g., Google Glass®). Additionally, the user interface input device may include a voice recognition detector that enables a user to interact with a voice recognition system (e.g., the Siri® navigator) via voice commands.

[0133] User interface input devices may also include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, game pads, graphics tablets, audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser range finders, and eye-tracking devices. User interface input devices may also include medical imaging input devices such as, for example, computed tomography devices, magnetic resonance imaging devices, ultrasound emission tomography devices, and medical ultrasound devices. User interface input devices may also include audio input devices such as, for example, MIDI keyboards and electronic musical instruments.

[0134] User interface output devices may include non-visual displays such as a display subsystem, indicator lights, or audio output devices. The display subsystem may be, for example, a flat panel device using a cathode ray tube (CRT), liquid crystal display (LCD) or plasma display, a projection device, or a touch screen. In general, when the term "output device" is used, it is intended to include all possible types of devices and mechanisms for outputting information from computer system 1000 to a user or to another computer. For example, user interface output devices include, but are not limited to, various display devices that visually convey text, images, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, voice output devices, and modems.

[0135] Computer system 1000 may include a storage subsystem 1018. The storage subsystem 1018 comprises software elements, which are illustratively located in system memory 1010. The system memory 1010 may store program instructions loadable and executable by the processing unit 1004, and data generated by the execution of these programs.

[0136] Depending on the configuration and type of computer system 1000, the system memory 1010 may be volatile memory (e.g., random access memory (RAM)) and / or non-volatile memory (e.g., read-only memory (ROM), flash memory). Typically, RAM contains data and / or program modules that are immediately accessible to and / or currently being operated on and executed by the processing unit 1004. In some implementations, the system memory 1010 may include multiple different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS), containing the basic routines that help to transfer information between elements within the computer system 1000, such as during start-up, may typically be stored in ROM. By way of example and not limitation, system memory 1010 also illustrates application programs 1012, program data 1014, and an operating system 1016, which may include client applications, a web browser, a mid-tier application, a relational database management system (RDBMS), and the like.As an example, operating system 1016 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® 10 OS, and Palm® OS operating systems.

[0137] The storage subsystem 1018 may also provide a tangible computer-readable storage medium for storing the basic programming and data constructs that provide the functionality of some embodiments. Software (programs, code modules, instructions) that, when executed by a processor, provide the above-described functionality may be stored in the storage subsystem 1018. These software modules or instructions may be executed by the processing unit 1004. The storage subsystem 1018 may also provide a repository for storing data used in accordance with the present disclosure.

[0138] The storage subsystem 1000 may also include a computer readable storage medium reader 1020 further connectable to a computer readable storage medium 1022. The computer readable storage medium 1022 may comprehensively represent remote, local, fixed and / or removable storage devices, in addition to storage media for temporarily and / or permanently containing, storing, transmitting and retrieving computer readable information together with, or in combination with, the system memory 1010 as appropriate.

[0139] Additionally, the computer readable storage medium 1022 containing the code or portions of code may include any suitable medium known or used in the art, including storage and communication media, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any manner or technology for storing and / or transmitting information. This may include tangible computer readable storage media, such as RAM, ROM, electronically erasable programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage, or other tangible computer readable media. This may also include intangible computer readable media, such as data signals, data transmissions, or other media usable to transmit the desired information and accessible by computer system 1000.

[0140] By way of example, the computer readable storage medium 1022 may include hard disk drives that read from or write to non-removable, non-volatile magnetic media, magnetic disk drives that read from or write to removable, non-volatile magnetic disks, and optical disk drives that read from or write to removable, non-volatile optical disks, such as CD ROMs, DVDs and Blu-ray disks or other optical media. The computer readable storage medium 1022 may include, but is not limited to, Zip drives, flash memory cards, universal serial bus (USB) flash drives, secure digital (SD) cards, DVD disks, digital video tapes, and the like. The computer-readable storage media 1022 may also include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory such as solid-state ROM, SSDs based on volatile memory such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs using a combination of DRAM and flash memory-based SSDs. Disk drives and their associated computer-readable media may provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer system 1000.

[0141] The communications subsystem 1024 provides an interface with other computer systems and networks. The communications subsystem 1024 serves as an interface for receiving data from other systems and transmitting data from the computer system 1000 to other systems. For example, the communications subsystem 1024 may enable the computer system 1000 to connect to one or more devices via the Internet. In some embodiments, the communications subsystem 1024 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular technology, advanced data network technologies such as 3G, 4G, or EDGE (enhanced data rates for global evolution)), WiFi (IEEE 802.11 family of standards or other mobile communications technologies or any combination thereof), global positioning system (GPS) receiver components, and / or other components. In some embodiments, the communications subsystem 1024 may provide a wired network connection (e.g., Ethernet) in addition to or instead of a wireless interface.

[0142] Additionally, in some embodiments, the communications subsystem 1024 may receive incoming communications in the form of structured and / or unstructured data feeds 1026, event streams 1028, event updates 1030, etc. on behalf of one or more users that may be using the computer system 1000.

[0143] As an example, the communications subsystem 1024 may be configured to receive data feeds 1026, such as web feeds, such as Twitter feeds, Facebook updates, Rich Site Summary (RSS) feeds, etc., in real time from users of social networks and / or other communications services, and / or receive real-time updates from one or more third party sources.

[0144] The communications subsystem 1024 may also be configured to receive data in the form of a continuous data stream, which may include an event stream 1028 of real-time events that may be continuous or may be essentially unbounded with no clear ends and / or event updates 1030. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, and the like.

[0145] The communications subsystem 1024 may also be configured to output structured and / or unstructured data feeds 1026, event streams 1028, event updates 1030, etc. to one or more databases that may be in communication with one or more streaming data source computers coupled to the computer system 1000.

[0146] Computer system 1000 may be one of a variety of types, including a handheld mobile device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google Glass® head mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack or other data processing system.

[0147] Because computers and networks are constantly evolving, the illustrated description of the computer system 1000 is intended only as a specific example. Many other configurations having more or fewer components than the illustrated system are possible. For example, customized hardware may also be used and / or specific elements may be implemented in hardware, firmware, software (including applets), or a combination. Additionally, connections to other computing devices, such as network input / output devices, may be utilized. Based on the disclosure and teachings provided in this disclosure, one of ordinary skill in the art will appreciate other means and / or methods for implementing various embodiments.

[0148] Although specific embodiments of the present disclosure have been described, various changes, modifications, alternative configurations, and equivalents are also encompassed within the scope of the present disclosure. The embodiments of the present disclosure are not limited to operating in a specific data processing environment, but may freely operate in multiple data processing environments. Furthermore, while the embodiments of the present disclosure have been described using a specific sequence of actions and steps, it will be apparent to those skilled in the art that the scope of the present disclosure is not limited to the sequence of actions and steps described. Various features and aspects of the above-described embodiments may be used individually or jointly.

[0149] Furthermore, while embodiments of the present disclosure have been described using a particular combination of hardware and software, it should be appreciated that other combinations of hardware and software are within the scope of the present disclosure. Embodiments of the present disclosure may be implemented using only hardware, only software, or a combination thereof. Various processes described in this disclosure may run on the same processor or any combination of different processors. Thus, when a component or module is described as being configured to perform a particular process, the configuration may be achieved, for example, by designing an electronic circuit to perform the process, by programming a programmable electronic circuit (such as a microprocessor) to perform the process, or a combination thereof. Processes may communicate using a variety of techniques, including but not limited to conventional techniques for communicating between processes. Different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.

[0150] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will be apparent, however, that additions, subtractions, deletions and other modifications and changes may be made without departing from the broad spirit and scope defined by the appended claims. Thus, although specific embodiments of the disclosure have been described, these embodiments are not intended to be limiting. Various modifications and equivalents thereof are intended to be encompassed within the scope of the appended claims.

[0151] The indefinite article "a" / "an", the definite article "the" and similar references used in the context of describing this disclosure (especially in the context of the claims) should be construed to include both the singular and the plural, unless otherwise stated in this disclosure or the context clearly indicates otherwise. The terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (i.e., meaning "including but not limited to"), unless otherwise stated. The term "connected" should be construed as being partly or entirely contained within, attached to, or joined together, even if there is something intervening. In this disclosure, the recitation of ranges of values ​​is intended merely as a shorthand method of referring to each individual value contained within the range, and each individual value is incorporated into this disclosure as if it were set forth individually in this disclosure, unless otherwise stated in this disclosure. All methods described in this disclosure can be performed in any suitable order, unless otherwise stated in this disclosure or the context clearly indicates otherwise. In this disclosure, the use of any and all examples or exemplary language (e.g., "such as") is intended to better clarify the embodiments of the disclosure and does not limit the scope of the disclosure, unless otherwise specified. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the disclosure.

[0152] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is intended to be understood in context as being used generally to indicate that an item, term, etc. may be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless otherwise noted. Thus, such disjunctive language is not generally intended to or implies that a particular embodiment requires that at least one of X, at least one of Y, or at least one of Z be present.

[0153] Preferred embodiments of the present disclosure are described herein, including the best mode known for carrying out the present disclosure. Variations of these preferred embodiments will become apparent to those skilled in the art upon reading the foregoing description. Those skilled in the art may adopt such variations as appropriate, and the present disclosure may be carried out in ways other than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by this disclosure unless otherwise indicated herein.

[0154] All references cited in this specification, including publications, patent applications, and patents, are incorporated by reference to the same extent as if each individual reference was individually and specifically indicated to be incorporated by reference and was set forth in its entirety herein.

[0155] In the foregoing specification, aspects of the disclosure have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that the disclosure is not limited thereto. Various features and aspects of the above-described disclosure may be used individually or jointly. Moreover, the embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Thus, the specification and drawings should be regarded as illustrative rather than restrictive.

Claims

1. A method for estimating one or more presumed sources of failure within a data center, the method comprising: obtaining a set of input data providing various parameters related to the data center, a series of devices within the data center, and applications running on the devices within the data center; detecting a failure of at least one function of the data center, at least partially based on the obtained set of input data; estimating one or more presumed sources of the failure by processing the set of input data using a model in response to detecting the failure, the model incorporating a set of rules that identify a correlation between the set of input data and the device or the application running on the device as the one or more presumed sources of the failure; estimating the one or more presumed sources of the failure comprises: generating an estimated level for each parameter included in the set of input data using the set of rules available to the model and past data related to each parameter; identifying one or more abnormal parameters including actual levels having a threshold deviation from the corresponding estimated levels by comparing the estimated level of each parameter with the actual level of each parameter included in the set of input data; identifying one or more devices and / or applications corresponding to each of the identified abnormal parameters, each of the identified one or more devices and / or the applications being included as the one or more presumed sources of the failure; the method further comprising generating a failure notification message providing the one or more presumed sources of the failure, at least partially based on processing the set of input data.

2. The method of claim 1, wherein the set of input data identifies any one of a temperature of each server within the data center, a power level of each power supply within each rack of the data center, climate data of the data center obtained from a series of sensors within the data center, ticket data identifying any function of the obtained data center, a series of devices within the data center, and a location of all devices within the data center.

3. The set of input data includes the location of each device in the data center and the device type of each device in the data center, The method is, by processing the set of input data, identifying the data types and one or more related devices associated with each part of the set of input data; assigning a timestamp indicating the time when each part of the set of input data was obtained to each part of the set of input data; The method according to claim 1, further comprising storing the set of input data in a database according to the identified data types and the assigned timestamps.

4. Detecting the failure includes obtaining a failure notification specifying that the failure has occurred from an external computing device, or The method according to claim 1, further comprising detecting that a threshold number of tickets are received that identify a loss of at least one function of the obtained data center or a loss of some computing resources.

5. The failure notification message includes a graphical representation of a first abnormal parameter and an estimated level of the first abnormal parameter, according to the method of claim 1.

6. The method further includes estimating a confidence level for each of the one or more presumed sources of the failure, based at least in part on some rules that correlate with the parameters associated with each of the one or more presumed sources of the failure, The failure notification message includes the confidence level, according to the method of claim 1.

7. For each of the one or more presumed sources of the failure and the confidence level for each of the one or more presumed sources of the failure, the abnormal characteristics of the set of input data are correlated with devices or applications in the data center so as to provide insights regarding the actual source of the failure. The method according to claim 6.

8. The method further includes retrieving, for each of the one or more presumed sources of the failure, resolution data associated with each of the one or more presumed sources of the failure, The resolution data provides known methods for resolving the failures specific to each of the presumed sources of the failure. The method according to claim 1, wherein the failure notification message includes the solution data.

9. A cloud infrastructure node comprising: a processor; and a non-transitory computer-readable medium containing instructions that, when executed by the processor, cause the processor to perform the following operations, the operations including: obtaining a set of input data providing various parameters related to a data center, a series of devices within the data center, and applications running on the devices within the data center; detecting a failure of a function of the data center based at least in part on the obtained set of input data; in response to detecting the failure, estimating one or more presumed sources of the failure by processing the set of input data using a model, the model incorporating a set of rules specifying a correlation between the set of input data and the device or the application running on the device as the one or more presumed sources of the failure, the operations including: generating a failure notification message providing the one or more presumed sources of the failure based at least in part on the estimation.

10. The non-transitory computer-readable medium further causes the processor to: using the model that uses a first rule among the set of rules, identify that a first abnormal parameter is related to a first application running on some servers within the data center; using the model that uses a second rule among the set of rules, further identify that a change to the execution of the first application occurred within a threshold period of the time when the failure was detected; the first application is included in the failure notification message as a first presumed source of the failure; The cloud infrastructure node according to claim 9, wherein the failure notification message further provides solution data specifying instructions for reverting the first application to a previous version so as to remove the change to the execution of the first application.

11. The set of input data includes the positions of the devices within the data center and the device types of the devices within the data center. The non-transitory computer-readable medium causes the processor to by processing the set of input data, identify the data types and one or more associated devices related to each part of the set of input data; assign a timestamp indicating the time when each part of the set of input data was obtained to each part of the set of input data; further cause the set of input data to be stored in a database according to the identified data types and the assigned timestamps, the cloud infrastructure node according to claim 9 or 10.

12. Detecting the failure includes obtaining a failure notification from an external computing device identifying that the failure has occurred, or further detecting that a threshold number of tickets identifying a loss of at least one function of the obtained data center or a loss of some computing resources have been received, the cloud infrastructure node according to claim 9 or 10.

13. The non-transitory computer-readable medium causes the processor to further estimate a confidence level for each of the one or more hypothesized sources of the failure, at least partially based on some of the set of rules that correlate with the parameters associated with each hypothesized source of the failure; the failure notification message includes the confidence level, the cloud infrastructure node according to claim 9 or 10.

14. Generating the one or more hypothesized sources of the failure by processing the set of input data using the model includes generating an estimated level for each parameter using historical data related to each parameter included in the set of input data; identifying one or more abnormal parameters including actual levels having a threshold deviation from the corresponding estimated levels by comparing the estimated level of each parameter with the actual level of each parameter included in the set of input data; further including identifying the device or series of devices corresponding to each of the identified abnormal parameters. Each of the specified apparatus or series of apparatuses is the cloud infrastructure node according to claim 9 or 10, included as the one or more presumed sources of the failure. **Claim 15** A computer program for causing a processor to execute the method according to any one of claims 1 to 8.