Utilizing device signals to improve recovery of virtual machine host devices
The host failure recovery system uses a multi-layer power failure detection model to quickly and accurately detect and recover faulty host equipment, solving the problem of inefficient detection and recovery in existing cloud computing systems, and achieving efficient recovery of faulty host equipment.
Patent Information
- Application Number
- CN202380080034.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-12-08
- Publication Date
- 2025-07-04
AI Technical Summary
The existing cloud computing system is inefficient and inaccurate in detecting and recovering faulty host servers. It often takes more than 15 minutes to identify faulty host servers, with an accuracy rate of only about 50%.
The host failure recovery system is adopted, and a multi-layer power failure detection model is used, including a rule-based layer and a machine learning layer, based on the power consumption signals and usage characteristics of the host device, quickly and accurately determine the health status of the host device and initiate the recovery process.
It significantly reduces downtime for host servers and cloud computing systems, detecting unhealthy host devices two to three times faster, with an accuracy rate of up to 95%, without waiting for long periods of verification of failures, achieving fast and accurate failure recovery.
Smart Images

Figure CN120266100A_ABST
Abstract
Description
Background Art
[0001] In recent years, significant progress has been made in the hardware and software platforms for implementing cloud computing systems. Cloud computing systems typically utilize different types of virtual services (e.g., virtual machines, computing containers, software packages) that provide computing capabilities for various devices. These virtual services can be hosted by a group of host servers on a cloud computing system such as a server data center.
[0002] Despite the progress in the cloud computing field, current cloud computing systems face several technical deficiencies, especially in the area of recovering failed host servers. Specifically, a host server is a computing device in a cloud computing system that hosts one or more virtual machines. Sometimes, a host server fails due to one or more reasons. For example, one or more of the virtual machines on a host server stop operating, and the host server that is still operating also stops functioning properly. As another example, the host server has a hardware failure. Although in some cases, such as a hardware failure with a specific hardware error, current cloud computing systems can detect and resolve the failed host server. However, in most cases, current cloud computing systems cannot detect or determine in a timely manner when a host server has failed or has already failed.
[0003] For example, many current cloud computing systems typically take 15 minutes or longer to investigate a potentially failed host server, and even then, the accuracy rate of these systems correctly identifying the failed host server is only about 50%. This extended downtime is inefficient and has a negative impact on both the cloud computing system and the client.
[0004] These problems and other issues (covered below) result in existing computing systems being inefficient and inaccurate in host device failure detection and recovery. Brief Description of the Drawings
[0005] The detailed description provides one or more implementations with additional specificity and detail by using the drawings, as briefly described below.
[0006] Figure 1 Illustrates an example overview of a host failure recovery system for efficiently detecting and recovering failed host devices according to one or more implementations.
[0007] Figure 2A – Figure 2B Illustrates an example diagram of a computing system environment in which a host failure recovery system is implemented according to one or more implementations.
[0008] Figure 3A – Figure 3BIllustrated is an example block diagram of a host failure recovery system that determines the health of a host device according to one or more implementations.
[0009] Figure 4 Illustrated is an example block diagram for generating a power failure machine learning model to efficiently and accurately determine the health of a host device according to one or more implementations.
[0010] Figure 5 Illustrated is a series of example actions for initiating recovery of a host device based on processing a power consumption signal of the host device according to one or more implementations.
[0011] Figure 6 Illustrated are some components that may be included in a computer system. DETAILED DESCRIPTION
[0012] This document describes using a host failure recovery system to efficiently and accurately determine the health of a host device. For example, the host failure recovery system detects when a host server fails by using a power failure detection model that determines whether the host server is operating in a healthy power state (abbreviated as "healthy") or an unhealthy power state (abbreviated as "unhealthy"). Specifically, the host failure recovery system uses a multi-layer power failure detection model that determines a power consumption failure event on the host device. The failure detection model determines the health of the host device with high confidence based on the usage characteristics and / or power consumption signal of the host device. Additionally, the host failure recovery system can initiate a rapid recovery of the failed host device.
[0013] In fact, the implementations of the present disclosure solve one or more of the problems mentioned above and other problems in the art. Systems, computer-readable media, and methods that utilize a host failure recovery system can quickly and accurately detect failed host devices in a cloud computing system and implement rapid recovery, which significantly reduces the downtime of host servers and cloud computing systems.
[0014] For illustration, in one or more implementations, the host failure recovery system uses a power failure detection model to determine that a host device is unhealthy. For example, the host failure recovery system receives at a management entity a power consumption signal corresponding to the power being consumed by a host system on the host device. Based on the power consumption signal, and in some cases based on the usage characteristics of the host device, the failure detection model determines a power consumption failure of the host device. In some instances, the power failure detection model includes a first rule-based layer and a second machine learning layer with a power failure machine learning model. In response to determining that the host device is unhealthy, the host failure recovery system initiates recovery of the unhealthy host device from the management entity.
[0015] As described in this document, when compared to existing computing systems, the host failure recovery system provides several technical benefits with respect to image dataset generation. In fact, the host failure recovery system provides several practical applications that deliver benefits and / or solve problems by adding accuracy and efficiency to the power consumption failure and recovery process of host devices in a cloud computing system. Some of these technical benefits and practical applications will be discussed below and in this document.
[0016] As pointed out above, existing computer systems have the problem of extended downtime of host devices, which results in service interruptions and loss of processing opportunities and functionality. Specifically, some existing computer systems rely on binary signals that track "heartbeats" to determine whether a host device has potentially failed. For example, an agent on the host device provides regular heartbeats (e.g., at fixed intervals), and if the heartbeat is lost, the existing computer system begins diagnosing the host device to determine whether a failure mitigation action is needed. When a problem is potentially identified, the existing computer system will first wait a waiting period (e.g., a waiting period of 15 minutes typically) before implementing a mitigation action because these existing computer systems lack the ability to make up-to-date, highly confident determinations about the health of the host device.
[0017] In contrast, the host failure recovery system performs several actions to overcome and improve these problems. As explained below, the host failure recovery system significantly reduces the time required to detect an unhealthy host device and improves the accuracy of detecting an unhealthy host device. In fact, in cases where existing computer systems take up to 15 minutes to identify an unhealthy host device with a 50% accuracy or confidence rate, the host failure recovery system can identify an unhealthy host device two to three times faster (e.g., within 5 - 8 minutes) with a confidence rate of up to 95%.
[0018] To further illustrate, the power failure detection model provides the benefits of increased accuracy and efficiency when compared to existing computer systems. For example, based on real-time monitoring of the power consumption of the host device and in some instances based on the usage characteristics of the host device using the power failure detection model, the power failure detection model can determine the health of the host device faster. In this way, the host failure recovery system determines more quickly when a host device is unhealthy and repairs it, thus eliminating extended downtime.
[0019] To illustrate by way of example, by utilizing a two - layer power fault detection model, the host fault recovery system efficiently and accurately processes simple and complex signals corresponding to the power consumption of a host device to determine when the host device is unhealthy and needs to be repaired. For example, the host fault recovery system utilizes a first rule - based layer in the power fault detection model that uses heuristic methods to quickly detect more common fault signal patterns of the host device.
[0020] Additionally, the host fault recovery system utilizes a second machine - learning layer that typically has a power fault machine - learning model which accurately determines when the host server is unhealthy based on complex and / or uncommon fault signal patterns (such as combinations of usage characteristics of the host device). In various implementations, the usage characteristics of the host device include hardware information, software information, and / or workload information. Thus, compared to existing computer systems, the host fault recovery system determines unhealthy host devices with higher confidence and accuracy (e.g., typically over 90 - 95% confidence).
[0021] As another benefit, due to the higher accuracy in detecting faulty host devices, the host fault recovery system does not need to wait for a waiting period (e.g., 15 minutes or longer) to verify whether the host server has truly failed or is just under - loaded / over - loaded. In fact, in some instances, the host fault recovery system determines in real - time whether the host server is unhealthy and should be repaired.
[0022] Furthermore, in various implementations, the host fault recovery system trains the power fault machine - learning model that operates in the second layer of the power fault detection model. In these implementations, by training the power fault machine - learning model, the host fault recovery system ensures that the power fault machine - learning model accurately and robustly generates a dynamic power consumption range corresponding to the power usage range that the host device is expected to consume based on its usage characteristics.
[0023] As illustrated in the previous discussion, this document utilizes various terms to describe the features and advantages of one or more implementations described in this document. These terms are defined as follows and are used in different examples and contexts within this document.
[0024] For illustration, as an example of the terms used in this document, the term "cloud computing system" refers to a network of connected computing devices that provide various services to client devices. For example, a cloud computing system can be a distributed computing system that includes a collection of physical server devices (e.g., server nodes) organized in a hierarchical structure (including compute zones, clusters, virtual local area networks (VLANs), racks, load balancers, fault domains, etc.). Additionally, the features and functions described in connection with cloud computing systems can similarly pertain to a hierarchical structure of racks, fault domains, or other physical server devices. A cloud computing system can refer to a private cloud computing system or a public cloud computing system.
[0025] In some implementations, a cloud computing system includes a management entity or orchestrator for managing servers (e.g., server clusters, server racks, server nodes, and / or other groups of server devices of computing devices). In many instances, the management entity manages one or more host devices. For example, in this document, the term "host device" refers to a server device that includes one or more virtual machines. In some instances, a host device can include memory and other computing components and systems, such as a host system and a secondary system. Additionally, a host device also includes a host operating system and additional components, as provided below in connection with Figure 2B what is provided.
[0026] As another example, as used in this document, a "virtual machine" refers to an emulation of a computer system on a server node that provides the functionality of one or more applications on a cloud computing system. In various implementations, a host device allocates computing cores and / or memory to virtual machines running on the host device. A virtual machine can provide the functionality required to execute one or more operating systems.
[0027] For example, the term "power consumption" refers to the power or electrical energy consumption at a given time instance or over a given time period. In various instances, power consumption is measured over time in watts or joules. Typically, a power monitoring device measures the power consumption of a computing device. In some implementations, a power monitoring device and / or other components generate a power consumption signal to indicate the measured power consumption of a computing device at a given time instance or over a time period.
[0028] As an additional example, as used in this document, "usage characteristics" (with reference to a computing device) refers to information corresponding to the characteristics, attributes, settings, and / or features of a computing device. For example, as further described below, the usage characteristics of a host device include hardware information (e.g., the physical capabilities of the host device), software information (e.g., the programs, instructions, and configurations of the host device), and workload information (e.g., the operations assigned to the host device and / or the operations running on the host device). The usage characteristics may include additional information corresponding to the host device. In various implementations, the host failure recovery system utilizes the usage characteristics to determine a power consumption range customized for a particular host device based on the current settings of the particular host device.
[0029] In various implementations, the host failure recovery system utilizes a power failure detection model to determine the health status of a host device in a cloud computing system. For example, in this document, a "power failure detection model" includes a model for determining a power consumption failure of a host device based on a power consumption signal. In some cases, the power failure detection model includes multiple layers, such as a first rule-based layer and a second machine learning layer with a power failure machine learning model. Examples of the power failure detection model are provided below in conjunction with the figures.
[0030] As another example, as used in this document, the terms "healthy" and "unhealthy" refer to the power or functional state of a host device. For example, the term "healthy" indicates the proper usage and consumption of a host device, such as a healthy power state for the operations being performed by a given host device. A host device is "unhealthy" when some or all of the host device becomes unresponsive, consumes too much power, or consumes too little power. For example, a host device with one or more unresponsive or poorly performing virtual machines is unhealthy. The host failure recovery system can recover and / or repair an unhealthy host device through a recovery process.
[0031] For example, as used in this document, the term "recovery process" refers to improving the health status of a host server by repairing one or more components of the host device. For example, a management entity provides in-band or out-of-band instructions to the host device (e.g., via an auxiliary service controller) to perform one or more processes to improve the health status of the host device. In some implementations, the recovery process includes restarting one or more virtual machines or applications running on the host device. In some implementations, the recovery process includes moving one or more virtual machines to another host device. In some implementations, the recovery process includes rebooting or restarting the host device.
[0032] As an additional example, as used in this document, the term "machine learning model" refers to a computer model or computer representation that can be adjusted (e.g., trained) based on inputs to approximate an unknown function. For example, machine learning models can include, but are not limited to: transformer models, sequence-to-sequence models, neural networks (e.g., convolutional neural networks or deep learning models), decision trees (e.g., gradient boosting decision trees), quantile regression models, linear regression models, logistic regression models, random forest models, clustering models, support vector learning, Bayesian networks, regression-based models, principal component analysis, or combinations of the above models. As used herein, the term "machine learning model" includes deep learning models and / or shallow learning models.
[0033] Additional details regarding the host failure recovery system will now be provided. For illustration, Figure 1 An example overview for efficiently detecting and recovering a failed host device using a host failure recovery system according to one or more implementations is shown. As shown, Figure 1 A series of actions 100 is illustrated, where one or more of the actions may be performed by the host failure recovery system.
[0034] As Figure 1 shown, the series of actions 100 includes an action 110 of receiving a power consumption signal of the host device and usage characteristics of the host device. In various implementations, the host failure recovery system operates on a management entity or management device within a cloud computing system, where the management entity manages a group of host devices, and where the host failure recovery system monitors the health of the group of host devices.
[0035] As part of monitoring the health of the host device, the host failure recovery system receives a device signal from the host device. For example, the host failure recovery system receives a power consumption signal from the host device. Additionally, the host failure recovery system also receives usage characteristics corresponding to the host device, such as hardware information, software information, and / or workload information. Additional details regarding the host device signal are provided below in connection with Figures 2A - 2B and Figures 3A - 3B to provide additional details.
[0036] As shown, the series of actions 100 includes an action 120 of generating a power consumption range using a power failure detection model. In many implementations, the host failure recovery system utilizes a power failure detection model that includes two layers (correspondingly corresponding to different device health measurement tools). For example, in one or more implementations, the power failure detection model includes a first rule-based layer and a second machine learning model-based layer. In these implementations, the host failure recovery system utilizes the first layer when possible.
[0037] For illustration, action 120 includes sub-action 122 of the rule-based layer of the power fault detection model. For example, in various implementations, the rule-based layer determines whether the host device signal matches the signal pattern of one or more device fault rules. If so, the host fault recovery system utilizes these rules to obtain and / or identify the power consumption range of the host device. Otherwise, the host fault recovery system provides some or all of the host device signals to the second layer. Additional details regarding the rule-based layer of the power fault detection model are provided below in conjunction with Figures 3A - 3B to provide additional details regarding the rule-based layer of the power fault detection model.
[0038] As shown, action 120 includes sub-action 124 of the power fault detection machine learning model of the power fault detection model. In some implementations, the power fault detection model includes a power fault detection machine learning model that generates a dynamic processing range (e.g., power range) based on host device signals (such as the usage characteristics of the host device). In one or more implementations, the host fault recovery system utilizes the power fault detection model to generate a power consumption range customized according to the host device and its specific characteristics.
[0039] As noted above, the host fault recovery system utilizes the second layer of the power fault detection model to determine the power usage range (e.g., upper / lower power thresholds) of the host device given the current capabilities of the host device and the assigned operations. Additional details regarding the power fault detection machine learning model are provided below in conjunction with Figures 3A - 3B and Figure 4 to provide additional details regarding the power fault detection machine learning model.
[0040] As shown, a series of actions 100 includes action 130 of determining that the host device is unhealthy by comparing the power consumption signal with a power consumption threshold. For example, in various implementations, the host fault recovery system compares the current power consumption signal with a power consumption range or threshold (e.g., the upper or lower limit of the range). When the host device has a power consumption outside the power consumption range, the host fault recovery system determines that the host device is unhealthy. Additional details regarding determining that the host device is unhealthy are provided below in conjunction with Figures 3A - 3B to provide additional details regarding determining that the host device is unhealthy.
[0041] In addition, a series of actions 100 includes action 140 of recovering the host device based on determining that the host device is unhealthy. In various implementations, based on determining that the host device is unhealthy, the host fault recovery system on the management entity initiates the recovery process as described above to recover the host device. In this way, the host fault recovery system quickly and accurately discovers unhealthy host devices and repairs them before an extended timeout occurs. Additional details regarding repairing unhealthy host devices are provided below in conjunction with Figure 3B to provide additional details regarding repairing unhealthy host devices.
[0042] As mentioned above, Figures 2A - 2B additional details regarding the host failure recovery system are provided. For example, Figures 2A - 2B FIG. illustrates an example diagram of a computing system environment in which the host failure recovery system 206 is implemented according to one or more implementations. Specifically, Figure 2A example environmental components associated with the host failure recovery system are introduced. Additionally, Figure 2B various details regarding components and elements that contribute to explaining the functionality, operation, and actions of the host failure recovery system 206 are included.
[0043] As shown in the figure, Figure 2A a schematic diagram of an environment 200 (e.g., a digital media system environment) for implementing the host failure recovery system 206 is included. As shown in the figure, the environment 200 includes a cloud computing system 201 and client devices 205 connected via a network 280. The cloud computing system 201 includes various computing devices and components. Additional details regarding the computing devices are provided below in conjunction with Figure 6 to provide additional details regarding the computing devices.
[0044] In various implementations, the cloud computing system 201 is associated with a data center. For example, the cloud computing system 201 provides various services to the client devices 205. The cloud computing system 201 includes components that may be located in the same physical location or may be dispersed across multiple geographical locations.
[0045] As shown in the figure, the cloud computing system 201 includes a management entity 202 and multiple host computing devices (e.g., host device 240). For simplicity, only one host device with internal components is shown, but each of the host devices 240 may include the same or similar components. Additionally, the management entity 202 and the host devices may be implemented on one or more server devices or other computing devices.
[0046] The management entity 202 shown in the cloud computing system 201 includes a host management system 204 with the host failure recovery system 206. In various implementations, the host management system 204 manages one or more of the host devices 240. For example, the host management system 204 assigns operations, jobs, and / or tasks to one or more of the host devices 240 to be executed on one or more virtual machines. In various implementations, the host management system 204 tracks usage information of the host devices 240, such as the hardware capabilities, software versions and configurations, and / or workload assignments and progress of each device. The host management system 204 may also perform other management tasks corresponding to the host devices 240, as well as management tasks corresponding to managing the cloud computing system 201.
[0047] As mentioned above, the host management system 204 includes a host failure recovery system 206. As described in this document, the host failure recovery system 206 detects when a host device 240 fails and takes mitigation actions. As Figure 2A shown, the host failure recovery system 206 includes a power failure detection model 222. In various implementations, the host failure recovery system 206 utilizes the power failure detection model 222 to determine when a host device fails based on a power consumption signal (e.g., current power usage). Additional Figure 2B details regarding the power failure detection model 222 and the components of the host failure recovery system 206 are provided.
[0048] Note that this document describes using a power consumption signal to determine a failed host device. In some implementations, the host failure recovery system 206 determines a failed host device based on different metrics. For example, instead of tracking power consumption, the host failure recovery system 206 tracks network bandwidth usage and / or network packets to determine a failed host device (e.g., based on a range of expected packets that a host device should transmit). In fact, the host failure recovery system 206 can generate and / or use a packet failure detection model in a manner similar to the power failure detection model 222 described in this document to determine when a host device fails.
[0049] As shown, the host device 240 includes a host system 250 and an auxiliary system 262. In various implementations, the host system 250 and the auxiliary system 262 are separate systems on the same computing device. For example, the host system 250 is a first computer system having its own processor and memory (e.g., one or more host processors and host memory), network connections, and / or power supply, while the auxiliary system 262 is a second computing system on the same computing device (e.g., sharing the same motherboard) and having its own components (e.g., auxiliary processor, memory, network connections, power supply). In some instances, the host system 250 and the auxiliary system 262 communicate with each other and / or share components (e.g., one system includes a shared memory accessible by the other system).
[0050] As shown, the host system 250 includes virtual machines 252. For example, each host device in the host device 240 includes one or more virtual machines, computing containers, and / or other types of virtual services (or the ability to implement these functions). The host device 240 can include additional components, such as Figure 2B those shown and further described below.
[0051] Also as shown, the auxiliary system 262 includes an auxiliary service controller 264 and a host power monitoring device 266. In various implementations, the auxiliary service controller 264 is a dedicated microcontroller within the host device, separate from the host processor or host system 250. For example, in many implementations, the auxiliary service controller 264 includes its own processor and memory. In various implementations, the auxiliary service controller 264 is a baseband management controller (BMC).
[0052] In various implementations, the auxiliary system 262 monitors the power consumption (and / or other functions) of the host system 250. For example, the host power monitoring device 266 on (or at other locations within the host device) the auxiliary system 262 measures the host system power consumption 242 used by the host system 250. As shown, the host power monitoring device 266 reports the measured host system power consumption 244 to the auxiliary service controller 264, which passes the reported host system power consumption 246 to the host fault recovery system 206. Although the host system power consumption 242, the measured host system power consumption 244, and the reported host system power consumption 246 are described as separate signals, in various implementations, they are the same signal or versions of the same signal.
[0053] As Figure 2A further shown, the environment 200 includes client devices 205 that communicate with the cloud computing system 201 via the network 280. The client devices 205 can represent various types of computing devices, including mobile devices, desktop computers, or other types of computing devices. For example, the client devices access services provided by one or more virtual machines 252 in the host device of the cloud computing system 201.
[0054] In addition, the network 280 can include one or more networks that use one or more communication platforms or technologies to transmit data. For example, the network 280 can include the Internet or other data links that enable the transmission of electronic data between the corresponding client devices and the devices of the cloud computing system 201. Additional details regarding these computing devices and networks are provided below in connection with Figure 6 to provide additional details regarding these computing devices and networks.
[0055] As mentioned above, Figure 2B the host device 240 including the management entity 202 and the cloud computing system 201 includes additional components. Along with these additional components, Figure 2B additional details regarding the functions, operations, and actions of the host fault recovery system 206 are also provided.
[0056] For illustration, host device 240 shows host system 250 and auxiliary system 262, each with additional components. For example, host system 250 includes virtual machine 252 as introduced above, as well as host processor 254, host memory 256, host power supply 258, and host network card 260. Additionally, auxiliary system 262 includes auxiliary service controller 264 and host power monitoring device 266 as introduced previously, as well as auxiliary power supply 268 and auxiliary network card 270, which are separate from host power supply 258 and host network card 260.
[0057] In various implementations, management entity 202 performs a first set of interactions with host system 250 and a second set of interactions with host device 240. For example, management entity 202 communicates with host system 250 via host network card 260 to assign jobs, provide data for processing, receive status updates, and / or receive processed information. Additionally, management entity 202 interacts with auxiliary system 262 to monitor the health and status of host system 250 and host device 240. For example, management entity 202 communicates with auxiliary system 262 via auxiliary network card 270. In this way, even if communication with host system 250 fails or is interrupted, management entity 202 can monitor the status of host system 250. In some implementations, host system 250 and auxiliary system 262 share the same network connection.
[0058] As mentioned, host system 250 and auxiliary system 262 include separate power supplies. This protects auxiliary system 262 from power failures or interruptions in host system 250 (and vice versa). Additionally, in some instances, this allows host power monitoring device 266 on auxiliary system 262 to accurately monitor the power usage (e.g., power consumption) of host system 250 independently of the power used by auxiliary system 262 on the host device. However, in some instances, host system 250 and auxiliary system 262 share the same power supply. In these implementations, when reporting the host system power consumption 246 of host system 250 (or host processor 254) to management entity 202, auxiliary service controller 264 may ignore or disregard its own power usage when reporting the power consumption of host system 250, or only provide the total power usage of the host device.
[0059] Additionally, the host management system 204 within the management entity 202 includes various additional components and elements. For example, the host management system 204 includes a host device manager 210, a power usage manager 218, a host device recovery manager 228, and a storage manager 230. As shown, the host device manager 210 includes example elements corresponding to information of the host device 240, such as hardware information 212, software information 214, and workload information 216. The fault detection manager 220 is shown to include a power fault detection model 222, which includes a rule-based layer 224 and a machine learning layer 226. Additionally, the storage manager 230 includes examples of stored data, such as the power consumption signal 232 and usage characteristics 234 of the host device 240.
[0060] As just mentioned, the host management system 204 includes a host device manager 210. In one or more implementations, the host device manager 210 manages and / or monitors one or more aspects of the host device 240. For example, in some cases, the host device manager 210 communicates with the host system 250 (shown as a dashed line 238) to identify, provide, and / or receive usage characteristics 234 or other information associated with the operation of the host system 250. In one or more implementations, the usage characteristics 234 include hardware information 212, software information 214, and / or workload information 216. In some implementations, the host device manager 210 is located outside the host fault recovery system 206 and transmits usage characteristics 234 and / or other information about the host device 240 / host system 250 to the host fault recovery system 206.
[0061] As mentioned, in some instances, the host device manager 210 receives hardware information 212 about the host device 240. In some cases, the hardware information 212 corresponds to the host device's processing information and / or its ability to perform one or more operations associated with the cloud computing system 201. For example, the host device manager 210 identifies information about the physical capabilities of the host device 240 and / or the host system 250, such as information about the processor, random access memory (RAM), graphics card, storage device, power supply, network card, etc. of the host device 240. In some cases, the hardware information 212 is identified based on the stock keeping unit (SKU) of the host device 240 (e.g., the host fault recovery system 206 can identify the hardware information 212 of the host system 250 by using the SKU as an index value in a device inventory lookup table). In some instances, the hardware information 212 is related to the power consumption of the host system 250 because a given hardware configuration can correspond to the expected power usage or power capabilities of the host system 250.
[0062] As mentioned above, in various implementations, the host device manager 210 receives software information 214. For example, the software information 214 may include the software version and / or configuration settings of the host system 250. Generally speaking, the software information 214 corresponds to the authorization for the host device 240 to use the hardware for various purposes. In one or more implementations, the software information 214 is related to the hardware information 212 of the host device 240. In this way, the hardware information 212 and the software information 214 represent the capabilities of the host system 250 to perform various functions (and these functions may correspond to the expected or anticipated power consumption).
[0063] In some instances, the software on the host system 250 is configured to allow up to a given number of virtual machines to be deployed on the host system 250. In various instances, the software is configured to deploy a certain type of virtual machine 252 (e.g., a search engine index or a graphics processing) on the host system and / or corresponding to the type of operations that the virtual machine is authorized to perform. In various implementations, the software information 214 corresponds to the expected power consumption of the host system 250, which is based on the type of software applications running on the host system 250.
[0064] As mentioned above, in some implementations, the host device manager 210 receives workload information 216. In some cases, the workload information 216 includes information related to the virtual machines 252 associated with the host system 250. For example, the workload information 216 identifies the number and / or functions of the virtual machines 252 assigned to or operating on the host system 250. In some cases, the workload information 216 identifies the type of each virtual machine 252, which corresponds to the type of operations that the virtual machine is configured to perform. Some virtual machine operations (such as graphical user interface (GUI) functions) consume more power than other virtual machine operations (such as web search functions). In this way, the number and type of the virtual machines 252 included in the workload information 216 correspond to the expected power consumption of the host system 250.
[0065] In various implementations, the host device manager 210 (and / or the host management system 204) manages the workload and distributes it to the host system 250. For example, the host device manager 210 acts as an orchestrator that receives incoming requests and divides the request processing among one or more virtual machines and / or host systems.
[0066] In some implementations, the host device manager 210 collects the hardware information 212, the software information 214, and / or the workload information 216 as discussed in this document. Then, the host device manager 210 transmits this information to the storage manager 230 for storage as usage characteristics 234 on the host management system 204.
[0067] As mentioned above, the host management system 204 includes a power usage manager 218. As shown, the power usage manager 218 communicates with an auxiliary service controller 264. For example, the power usage manager 218 receives a power consumption signal that reports the constant or periodic power consumption 246 of the host system from the auxiliary system 262.
[0068] In various implementations, the auxiliary system 262 provides the reported host system power consumption 246 via a host device 240 that communicates with the management entity 202. In one or more implementations, the host device 240 sends the reported host system power consumption 246 to a host fault recovery system 206 on the management entity 202 via an auxiliary network card 270 and / or a host network card 260. In some implementations, the auxiliary service controller 264 has a direct connection to the power usage manager 218, such as via an out-of-band connection.
[0069] In some implementations, the reported host system power consumption 246 includes a measurement in watts or joules over time (e.g., joules per second), corresponding to the rate of electricity consumption at which the host system 250 performs its various functions. For example, the reported host system power consumption 246 indicates the power usage of the host system 250 at a single instant in time. As another example, the reported host system power consumption 246 indicates information about the power usage of the host system 250 over an elapsed period of time or time window.
[0070] In some implementations, the reported host system power consumption 246 includes raw information about the power usage of the host system 250, and the power usage manager 218 further processes and / or interprets the information (e.g., calculates the average power consumption over an elapsed period of time). In various implementations, the auxiliary system 262 may process some or all of the power consumption information before transmitting it to the power usage manager 218. In some instances, the power usage manager 218 transmits an original version and / or a processed version of the reported host system power consumption 246 to a storage manager 230 for storage as a power consumption signal 232 on the host management system 204. Accordingly, information associated with the current power usage of the host system 250 is transmitted to the host management system 204.
[0071] As mentioned above, the host management system 204 includes a fault detection manager 220. In various embodiments, the fault detection manager determines the fault state of the host system 250. For example, the fault detection manager 220 implements a power fault detection model 222 that considers the power consumption signal 232 and / or usage characteristics 234 to determine whether the host system 250 is healthy. In many implementations, the power fault detection model 222 is a two-layer model that includes a first rule-based layer 224. For example, the rule-based layer 224 considers the basic and / or common patterns of the power consumption signal 232 and / or usage characteristics 234 to parse whether the host system 250 is healthy.
[0072] In cases where the rule-based layer cannot parse the health state (e.g., cannot parse at all or cannot reach a given confidence value), the fault detection manager 220 implements a second machine learning layer that is trained to estimate or predict the expected power consumption range of the host system 250 based on the usage characteristics 234 and in some cases based on the power consumption signal 232. The fault detection manager 220 determines the health state of the host system 250 by comparing the power consumption range with the measured power consumption of the host system 250 (e.g., included in the power consumption signal 232). The power fault detection model 222 will be discussed in further detail below in conjunction with Figures 3A - 3B and Figure 4 In this way, the fault detection manager 220 can determine the health of the host system 250.
[0073] As mentioned above, the host management system 204 includes a host device recovery manager 228. In some implementations, the host device recovery manager 228 implements a recovery process based on the fault determination of the fault detection manager 220. For example, based on the determined fault of the host system 250, the host device recovery manager 228 generates a recovery signal to send to the host system 250. In some instances, the host device recovery manager 228 sends the recovery signal directly to the host system 250. In some instances, the host device recovery manager 228 sends the recovery signal to the host system 250 through the auxiliary system 262 (e.g., when the host system 250 is unresponsive, or when the auxiliary system 262 needs to implement a restart or reboot of some or all of the host system 250 and / or the host device 240).
[0074] In various implementations, the recovery signal includes instructions for recovering or repairing the host system 250. For example, the recovery signal includes instructions for restarting or rebooting the host system 250. As another example, the recovery signal indicates to restart or reboot a portion of the host system 250 (e.g., one or more virtual machines among the virtual machines 252 operating on the host system 250). Thus, in various implementations, once the host system 250 fails and the host system 250 follows the recovery instructions, the host device recovery manager 228 implements the recovery of the host system 250.
[0075] The host failure recovery system 206 can monitor, detect, and recover a failed host device, as discussed above. In some implementations, the host failure recovery system 206 performs the various features described in this document within an amount of time corresponding to the recovery time. For example, the recovery of the host device 240 (e.g., from detection to restart) can occur within the recovery time. In some instances, the recovery time can be real-time, 30 seconds, 1 minute, 2 minutes, 5 minutes, 8 minutes, 10 minutes, or any other duration. In this way, the host failure recovery system 206 can detect and recover a failed device in a shorter time than conventional systems, thereby providing a significant improvement.
[0076] Although Figures 2A - 2B illustrates a specific number, type, and arrangement of components within the environment 200 and / or the cloud computing system 201, various additional environmental configurations and arrangements are possible. For example, the environment 200 or the cloud computing system 201 can include additional client devices and / or computing devices. In some implementations, the cloud computing system 201 is a single computing device or a computer cluster.
[0077] As mentioned above, Figures 3A - 3B additional details are provided regarding host device signals and the power failure detection model (including the layers of the model). Specifically, Figures 3A - 3B illustrates an example block diagram of a host failure recovery system for determining the health of a host device according to one or more implementations.
[0078] For illustration, the host failure recovery system 206 analyzes the failure state of the host system 250 by implementing the power failure detection model 222. In some implementations, the host failure recovery system 206 inputs the information collected about the host device 240 into the power failure detection model 222. For example, in various instances, the inputs are the power consumption signal 332 and the usage characteristics 334, corresponding to those described above in connection with Figure 2A and Figure 2BThe described power consumption signal 232 and usage characteristics 234. Based on the input, the power failure detection model 222 generates an output. For example, the host failure recovery system considers the power consumption signal 332 and / or usage characteristics 334 of the host device 240, and outputs a determination of a healthy host device 322 or an unhealthy host device 324.
[0079] As discussed above, conventional systems only verify and / or monitor the health of host devices based on a lost heartbeat, and initiating a recovery operation can be delayed for up to 15 minutes. In some implementations, the host failure recovery system 206 continuously monitors the power consumption signal 332 and / or usage characteristics 334 of the host device, and uses this information to determine the fault state at any point in time. In various implementations, the host failure recovery system 206 periodically monitors the health of the host device every 30 seconds, 1 minute, 2 minutes, 5 minutes, 8 minutes, or 10 minutes. In certain implementations, the host failure recovery system 206 also waits for a lost heartbeat or other type of signal. However, the host failure recovery system 206 does not need to be further delayed in determining whether the host device has failed because the host failure recovery system 206 uses the power failure detection model 222 to make this determination quickly, as described herein. In fact, the host failure recovery system 206 can implement the various features described in this document at any desired interval to give a more real-time representation of the health of the host device 240, thus providing a significant improvement over conventional systems.
[0080] As discussed herein, the power failure detection model 222 is a two-layer model that includes a first rule-based layer 224 and a second machine learning layer 226. For example, when implementing the power failure detection model 222, the host failure recovery system 206 first implements the rule-based layer 224. As shown in the example implementation, the rule-based layer 224 includes a rule manager 302 and one or more rules. The rule manager 302 considers the host device input and determines whether one or more rules in the rule-based layer 224 apply to these inputs (e.g., whether one of the rules can resolve the health state of the host device 240 based on these inputs). If the rule applies, then the host failure recovery system 206 will continue by implementing one or more rules in the rule-based layer 224 and will not continue to implement the second machine learning layer 226 (e.g., skip the machine learning layer 226).
[0081] In some instances, the rules represent simple and / or common patterns of the power consumption signal 332 and / or usage characteristics 234. In these instances, the rules can resolve the fault state of the host device 240 with high confidence, thus eliminating the need to implement the more complex and procedurally cumbersome machine learning layer 226.
[0082] In various implementations, the rule - based layer 224 includes a power consumption signal rule 304. For example, in various instances, the power consumption signal rule 304 includes the rule: when the observed power consumption of the host device 240 is below 20W (or other power rate). When this rule is met or reached, the host device 240 is unhealthy. In fact, the power consumption signal rule 304 can include various conditional rules that can be easily applied to the power consumption signal 332 to determine the health status of the host device 240.
[0083] In some implementations, in addition to the power consumption signal rule 304, the rule - based layer 224 includes a usage characteristic rule 306, as Figure 3A shown. For example, certain devices can have a hardware configuration with known power consumption requirements (e.g., identifiable by SKU), and the effective power consumption above or below a threshold power can correspond to a rule determination that the host device 240 is unhealthy. For example, a host device with five active virtual machines should have a power consumption above 50W. Thus, the host failure recovery system 206 implements certain advanced rules to utilize the rule - based layer 224 to handle easily resolvable situations.
[0084] The host failure recovery system 206 is capable of creating and / or implementing any number of rules based on any number of host device inputs and their combinations to implement the techniques of the rule - based layer 224 as described in this document. In this way, in various instances, the host failure recovery system 206 uses the rule - based layer 224 of the power failure detection model 222 to determine whether the host device 240 is a healthy host device 322 or an unhealthy host device 324.
[0085] In some implementations, the rule manager 302 considers the host device inputs and determines that the rules of the rule - based layer 224 do not apply to these inputs (e.g., no rule can resolve the health status of the host device 240 based on the inputs). In these cases, the host failure recovery system 206 continues to implement the second machine - learning layer 226. For example, the power failure machine - learning model 310 considers the usage characteristics 334 of the host device (i.e., hardware information, software information, and / or workload information) and predicts, estimates, generates, determines, and / or plans the expected power consumption of the host device 240.
[0086] As shown, the power failure machine - learning model 310 generates a dynamic power consumption range 318. In many implementations, the dynamic power consumption range 318 is customized and / or specific to the host device. In fact, the host failure recovery system 206 utilizes the power failure machine - learning model 310 within the power failure detection model 222 to generate the dynamic power consumption range 318 based on the usage characteristics 334 of the host device 240, as described in more detail below.
[0087] In various implementations, the dynamic power consumption range 318 includes an upper limit (e.g., a maximum threshold) and a lower limit (e.g., a minimum threshold). In various implementations, the upper and lower limits can help predict the expected power consumption within a desired confidence interval or percentage. For example, in some instances, the dynamic power consumption range 318 predicts the expected power consumption with an accuracy of 95%.
[0088] In additional implementations, the power failure machine learning model 310 generates different dynamic power consumption ranges corresponding to different confidence levels (e.g., a first power consumption range with a lower confidence level has a smaller range than a second power consumption range with a higher confidence level). Thus, the machine learning layer 226 can be implemented to generate a dynamic power consumption range 318 with a given confidence interval corresponding to the expected power consumption of the host device 240.
[0089] As discussed in conjunction with Figure 4 the machine learning layer 226 includes a power failure machine learning model 310 that is trained to determine the dynamic power consumption range 318. Additionally, any number of usage characteristics can be used as a basis for training the power failure machine learning model 310 and / or generating the dynamic power consumption range 318, in addition to those usage characteristics discussed in this document.
[0090] As Figure 3B shown by action 320 in, the host failure recovery system 206 determines whether the host device 240 is a healthy host device 322 or an unhealthy host device 324 based on determining whether the power signal is below or above the power consumption range (or threshold). Specifically, the host failure recovery system 206 compares the dynamic power consumption range 318 with the power consumption signal 332 to determine whether the power signal is below or above the power consumption threshold (where the threshold corresponds to the dynamic power consumption range 318). When the power consumption signal 332 is above the maximum threshold or below the minimum threshold of the dynamic power consumption range 318 (shown as the "yes" path), the host failure recovery system 206 determines that the host device 240 is an unhealthy host device 324. Otherwise, when the power consumption signal 332 of the host device 240 is within the dynamic power consumption range 318 (shown as the "no" path), the host failure recovery system 206 determines that the host device 240 is a healthy host device 322.
[0091] In fact, the host failure recovery system 206 implements a second machine learning layer 226 of the power failure detection model 222, which generates a dynamic power consumption range 318 specific to the usage characteristics 334 of the host device 240. In this way, the host failure recovery system 206 can determine the health status of the host device 240 based on combinations of complex or uncommon usage characteristics 334. More details related to the power failure machine learning model 310 and related to determining the dynamic power consumption range 318 will be described below in conjunction with Figure 4 discussed.
[0092] As mentioned above, in some instances, the host failure recovery system 206 implements the power failure detection model 222 and determines that the host device 240 is unhealthy. In such an instance, the host failure recovery system 206 continues to implement a recovery process for the host device 240 (e.g., restart some or all of the components of the host device 240), such as those discussed above in conjunction with Figure 2B discussed. In fact, the host failure recovery system 206 initiates a recovery process 344, which is shown by linking two "B" annotations together.
[0093] In other instances, the host failure recovery system 206 does not make any changes to the host device 240 (e.g., delays the recovery process) and continues to monitor the power consumption signal 342 in preparation for future determination of the health status of the host device as described in this document, which is shown by linking two "A" annotations together.
[0094] As mentioned above, in many implementations, the host failure recovery system 206 considers the current state and recent events of the host device 240 when determining whether the host device 240 has failed. For example, the power failure detection model 222 detects a recent power fluctuation event, such as the host device 240 having been recently restarted. In such a case, the host device 240 may signal power usage below typical values (e.g., shutdown) and / or above typical values (e.g., startup). In fact, the host failure recovery system 206 can identify one or more routine processes associated with the power fluctuation event. In these implementations, the host failure recovery system 206 can delay for a period of time to determine the health or health status of the host device 240. In some implementations, the host failure recovery system 206 provides this information to the power failure machine learning model 310, which incorporates it as part of generating a dynamic power consumption range 318 customized for the host device 240.
[0095] Figure 4 A block diagram illustrating the training of a power failure machine learning model according to one or more implementations is shown. As shown, Figure 4Examples of power failure machine learning models generated by the host failure recovery system 206 are included. Specifically, Figure 4 A power failure detection model is shown that trains a power failure machine learning model to detect, classify, and quantify a dynamic power consumption range. In various implementations, the power failure machine learning model is implemented on a client device and / or a server device by the host failure recovery system 206.
[0096] As shown in the figure, Figure 4 It includes training data 402, a power failure machine learning model 310, and a loss model 420. The training data 402 includes training usage characteristics 404 and a set of true power consumption 406 corresponding to the training usage characteristics 404. For example, in one or more implementations, the training usage characteristics 404 include hardware information, software information, and workload information of the host device, while the true power consumption 406 includes the power consumption signal of the host device with a corresponding set of training usage characteristics 404. In some implementations, the training data 402 includes different types of usage characteristics and / or corresponding truths.
[0097] Also as shown in the figure, Figure 4 It includes a power failure machine learning model 310. In various implementations, the power failure machine learning model 310 is a quantile regression model that includes several machine learning network layers (e.g., neural network layers). For example, in one or more implementations, the power failure machine learning model 310 includes an input layer, a hidden layer, and an output layer. The power failure machine learning model 310 may include additional and / or different network layers not currently shown.
[0098] As just mentioned, the power failure machine learning model 310 includes an input layer. In some implementations, the input layer encodes the training usage characteristics 404 into vector information. For example, in various implementations, the input layer processes the input usage characteristics (e.g., the training usage characteristics 404) to encode the usage characteristic features corresponding to the hardware information, software information, and / or workload information into a numerical representation.
[0099] As mentioned above, the power failure machine learning model 310 includes one or more hidden layers. In one or more implementations, the hidden layers process feature vectors to find hidden encoded features. In various implementations, the hidden layers map or encode the input using features into feature vectors (i.e., potential object feature maps or potential object feature vectors). In various implementations, the hidden layers generate features based on the encoded usage feature data from the input layer. For example, in one or more implementations, the hidden layers process each usage feature through various neural network layers to transform, convert, and / or encode the data from the input usage features into a feature vector (e.g., a string of numbers representing the encoded usage feature data in a vector space).
[0100] As mentioned above, the power failure machine learning model 310 includes an output layer. In one or more implementations, the output layer processes feature vectors to decode the power consumption information from the input usage features. For example, the output layer can decode the feature vector to generate the upper quantile 410 and / or the lower quantile 412 of the power consumption corresponding to the input usage features. The upper quantile 410 and the lower quantile 412 define the dynamic power consumption range 418.
[0101] As Figure 4 shown, the host failure recovery system 206 can implement a loss model 420. In some instances, the loss model includes one or more loss functions, and the host failure recovery system 206 implements the loss functions of the loss model 420 to train the power failure machine learning model 310. For example, the host failure recovery system 206 uses the loss model 420 to determine the amount of error or loss corresponding to the dynamic power consumption range 418 output by the power failure machine learning model 310. For example, the host failure recovery system 206 uses the loss model 420 to compare a given true power consumption with a given dynamic power consumption range to generate an amount of error or loss (where the given dynamic power consumption range is generated by the power failure machine learning model 310 based on the training usage features 404 corresponding to the true power consumption 406 being compared).
[0102] In various implementations, the host failure recovery system 206 uses the amount of loss to train and optimize the neural network layers of the power failure machine learning model 310 via backpropagation and / or end-to-end learning. For example, the host failure recovery system 206 backpropagates the amount of loss via the feedback 422 to adjust the hidden layer (and / or other layers of the power failure machine learning model 310). In this way, the host failure recovery system 206 can iteratively adjust and train the power failure machine learning model 310 to learn a set of best-fit parameters that can accurately generate the upper quantile 410 and the lower quantile 412 that define the dynamic power consumption range 418.
[0103] Now turning toFigure 5 , which illustrates an example flowchart of a series of actions including initiating a recovery of a host device based on a power consumption signal of the host device according to one or more implementations. Although Figure 5 illustrates actions according to one or more implementations, alternative implementations may omit, add, reorder, and / or modify any of the actions shown. Additionally, Figure 5 each of the actions in Figure 5 may be performed as part of a method. Alternatively, a non-transitory computer-readable medium may include instructions that, when executed by at least one processor, cause a computing device to perform the actions of Figure 5 . In additional implementations, a system may perform the actions of
[0104] As shown, a series of actions 500 includes an action 510 of receiving a power consumption signal of a host device. For example, action 510 may involve receiving, at a management entity, a power consumption signal corresponding to the power consumption of a host system on the host device.
[0105] Further as shown, a series of actions 500 includes an action 520 of determining that the host device is unhealthy by using a power failure detection model that determines a power consumption fault event of the host device. For example, action 520 may involve using, at a management entity, a fault detection model to determine that the host device is unhealthy, the fault detection model determining a power consumption fault of the host device based on the power consumption signal.
[0106] Further as shown, a series of actions 500 includes an action 530 of recovering the host device based on determining that the host device is unhealthy. For example, action 530 may involve initiating, from a management entity, a recovery process for recovering the host device based on determining that the host device is unhealthy.
[0107] In various implementations, a series of actions 500 may include additional actions. For example, in some implementations, a series of actions 500 includes an action of generating a power failure detection model that includes a first rule-based layer and a second machine learning layer having a power failure machine learning model. Among these actions, the power failure detection model determines a power consumption fault of the host device based on usage characteristics of the host device and the power consumption signal. Additionally, a series of actions 500 may include an action of determining whether the host device is unhealthy by comparing the power consumption signal with a power consumption range or threshold determined for the host device by the power failure detection model based on one or more usage characteristics of the host device.
[0108] In some implementations, the power failure detection model includes a first rule-based layer and a second machine learning layer, and a series of actions 500 includes an action of determining whether to use the first rule-based layer or the second machine learning layer to determine that the host device is unhealthy based on the power consumption signal and usage characteristics received from the host device. Additionally, in one or more implementations, the action may include an action of determining to use the first rule-based layer based on identifying a rule corresponding to the power consumption signal from the host device and determining that the host device is unhealthy based on the power consumption signal satisfying the rule.
[0109] In various implementations, a series of actions 500 includes an action of determining to use the second machine learning layer with a power failure machine learning model based on a failure to identify a rule corresponding to the usage characteristics or power consumption signal of the host device. Additionally, the action may include using the power failure machine learning model to determine a power consumption range for the host device based on the usage characteristics of the host device. Further, the action may include determining that the host device is unhealthy based on the power consumption signal being outside the power consumption range of the host device.
[0110] In one or more implementations, the power consumption signal and usage characteristics received at the management entity correspond to operations being performed by a host processor of a host system of the host device. Additionally, in some instances, the usage characteristics of the host device include hardware information of the host device, software information of the host device, and / or workload information of virtual machines running on the host device.
[0111] In some implementations, the power consumption signal of the host system is received at the management entity from an auxiliary service controller on the host device, and the host device includes a processor separate from the host system and a power supply for the auxiliary service controller. In various cases, the auxiliary service controller uses a power measurement device on the host device to determine the power consumption signal.
[0112] In some instances, a series of actions 500 includes an action of providing a recovery instruction to the host system via an auxiliary controller on the host device (e.g., as shown by the dashed line 248). Additionally, in various implementations, a series of actions 500 includes an action of providing a recovery instruction for restarting the host system from the management entity to the host device.
[0113] In one or more implementations, a series of actions 500 includes receiving a second power consumption signal corresponding to a second host system on a second host device at the management entity. Additionally, the action may include determining that the second host device is healthy by using the power failure detection model based on the second power consumption signal. Further, in some instances, the action includes delaying a recovery process for the second host device based on determining that the second host device is healthy.
[0114] In various examples, a power failure detection model determines that a host device is unhealthy based on the power consumption signal of the host device being below a lower power consumption threshold determined by the power failure machine learning model for the host device. Additionally, in some implementations, the workload information of the usage characteristics of the host device includes the number of operations on the host device and the types of operations on the host device.
[0115] In one or more implementations, a series of actions 500 includes an action to delay initiating a recovery process based on detecting a mitigation event corresponding to the host device. In some examples, a management entity manages multiple host devices, and the host device is one of the multiple host devices. In some cases, the host device and the management entity are located in a data center.
[0116] In some examples, a series of actions 500 includes determining a first rule - based layer that does not utilize a power detection model, and an action to utilize a second machine - learning layer of the power detection model to determine that the host device is unhealthy in response to determining not to utilize the first rule - based layer of the power failure detection model. In some examples, the power failure machine learning model determines that the host device is unhealthy based on the power consumption signal of the host device being below the lower limit of the power consumption range determined by the power failure machine learning model for the host device. Additionally, in some examples, the power failure machine learning model determines that the host device is unhealthy based on the power consumption signal of the host device being above the upper limit of the power consumption range determined by the power failure machine learning model for the host device.
[0117] A "computer network" (hereinafter referred to as "network") is defined as one or more data links that can enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or other communication connection (wired, wireless, or a combination of wired and wireless), the computer treats the connection as a transmission medium. The transmission medium can include networks and / or data links, which can be used to carry the required program code in the form of computer - executable instructions or data structures, and the program code can be accessed by a general - purpose or special - purpose computer. The above combinations should also be included within the scope of computer - readable media.
[0118] In addition, the networks (i.e., computer networks) described in this document can represent a network or a collection of networks (such as the Internet, an enterprise intranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a cellular network, a wide area network (WAN), a metropolitan area network (MAN), or a combination of two or more such networks), and one or more computing devices can access the host failure recovery system 206 through these networks. In fact, the networks described in this document can include one or more networks that use one or more communication platforms or technologies to transmit data. For example, the network can include the Internet or other data links that can enable the transmission of electronic data between the corresponding client devices and components (e.g., server devices and / or virtual machines thereon) of a cloud computing system.
[0119] In addition, when arriving at various computer system components, program code means in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-transitory computer-readable storage medium (device) (and vice versa). For example, computer-executable instructions or data structures received through a network (i.e., a computer network) or a data link can be buffered in the RAM within a network interface module (NIC) and then ultimately transferred to the computer system RAM and / or a less volatile computer storage medium (device) at the computer system. Therefore, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize the transmission medium.
[0120] Computer-executable instructions include, for example, instructions and data that, when executed by at least one processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or a group of functions. In some implementations, the computer-executable instructions are executed by a general-purpose computer to transform the general-purpose computer into a special-purpose computer that implements the elements of the present disclosure. Computer-executable instructions can include, for example, binary code, intermediate format instructions (such as assembly language), or even source code. Although the subject matter has been described in language specific to structural features and / or method acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or acts described above. Instead, the described features and acts are disclosed as example forms for implementing the claims.
[0121] Figure 6 Certain components that can be included in a computer system 600 are illustrated. The computer system 600 can be used to implement the various computing devices, components, and systems described in this document.
[0122] In various implementations, computer system 600 can represent one or more of the client devices, server devices, or other computing devices described above. For example, computer system 600 can refer to various types of network devices capable of accessing data on a network (i.e., a computer network), a cloud computing system, or other systems. For example, a client device can refer to a mobile device, such as a mobile phone, smartphone, personal digital assistant (PDA), tablet computer, laptop computer, or wearable computing device (e.g., headphones or smartwatch). A client device can also refer to a non-mobile device, such as a desktop computer, a server node (e.g., from another cloud computing system), or other non-portable devices.
[0123] Computer system 600 includes a processor 601 (i.e., at least one processor). Processor 601 can be a general-purpose single-chip or multi-chip microprocessor (e.g., an advanced RISC (Reduced Instruction Set Computer) machine (ARM)), a special-purpose microprocessor (e.g., a digital signal processor (DSP)), a microcontroller, a programmable gate array, etc. Processor 601 can be referred to as a central processing unit (CPU). Although the illustrated processor 601 is only Figure 6 a single processor in the computer system 600 shown, in alternative configurations, a combination of processors (e.g., ARM and DSP) can be used.
[0124] Computer system 600 also includes a memory 603 that communicates electronically with processor 601. Memory 603 can be any electronic component capable of storing electronic information. For example, memory 603 can be embodied as random access memory (RAM), read-only memory (ROM), disk storage media, optical storage media, flash devices in RAM, on-board memory included in the processor, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, etc. (including combinations thereof).
[0125] Instructions 605 and data 607 can be stored in memory 603. Instructions 605 can be executed by processor 601 to implement some or all of the functions disclosed herein. Executing instructions 605 can involve using data 607 stored in memory 603. Any of the various examples of modules and components described in this document can be partially or fully implemented as instructions 605 stored in memory 603 and executed by processor 601. Any of the various examples of data described in this document can be in data 607 stored in memory 603 and used during the execution of instructions 605 by processor 601.
[0126] The computer system 600 may also include one or more communication interfaces 609 for communicating with other electronic devices. The one or more communication interfaces 609 may be based on wired communication technologies, wireless communication technologies, or both. Some examples of the one or more communication interfaces 609 include Universal Serial Bus (USB), Ethernet adapters, wireless adapters operating according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 wireless communication protocol, wireless communication adapters, and infrared (IR) communication ports.
[0127] The computer system 600 may also include one or more input devices 611 and one or more output devices 613. Some examples of the one or more input devices 611 include keyboards, mice, microphones, remote control devices, buttons, joysticks, trackballs, touchpads, and light pens. Some examples of the one or more output devices 613 include speakers and printers. A particular type of output device that is typically included in the computer system 600 is the display device 615. The display device 615 used in the implementations disclosed herein may utilize any suitable image projection technology, such as liquid crystal displays (LCDs), light emitting diodes (LEDs), gas plasma, electroluminescence, etc. A display controller 617 may also be provided for converting the data 607 stored in the memory 603 into text, graphics, and / or moving images (as appropriate) shown on the display device 615.
[0128] The various components of the computer system 600 may be coupled together by one or more buses, which may include a power bus, control signal buses, status signal buses, data buses, etc. For clarity, Figure 6 the various buses are illustrated as a bus system 619 in the figures.
[0129] Those skilled in the art will appreciate that the present disclosure may be practiced in a network computing environment having a variety of types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. The present disclosure may also be practiced in a distributed system environment where local and remote computer systems are linked (by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network (i.e., a computer network) and cooperate to perform tasks. In a distributed system environment, program modules may be located in local and remote memory storage devices.
[0130] Unless otherwise specifically described as implemented in a particular manner, the techniques described in this document can be implemented by hardware, software, firmware, or any combination thereof. Any features described as modules, components, etc. can also be implemented together in an integrated logic device or can be implemented separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part via a non-transitory processor-readable storage medium that includes instructions that, when executed by at least one processor, perform one or more of the methods described in this document. The instructions can be organized as routines, programs, objects, components, data structures, etc., which can perform specific tasks and / or implement specific data types and can be combined or distributed as needed in various implementations.
[0131] A computer-readable medium can be any available medium accessible by a general or special-purpose computer system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, implementations of the present disclosure can include at least two distinct types of computer-readable media: non-transitory computer-readable storage medium (device) and transmission medium.
[0132] As used in this document, a non-transitory computer-readable storage medium (device) can include RAM, ROM, EEPROM, CD-ROM, solid state drives (“SSD”) (e.g., RAM-based), flash memory, phase change memory (“PCM”), other types of memory, other optical disc storage devices, magnetic disk storage devices, or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of computer-executable instructions or data structures and that is accessible by a general or special-purpose computer.
[0133] Without departing from the scope of the claims, the steps and / or actions of the methods described in this document can be interchanged with each other. In other words, unless a particular order of steps or actions is required for the correct operation of the method described, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims.
[0134] The term “determine” encompasses a variety of actions and, accordingly, “determine” can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or other data structure), ascertaining, etc. Further, “determine” can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Still further, “determine” can include parsing, selecting, choosing, establishing, etc.
[0135] The terms "comprising", "including", and "having" are intended to be inclusive and mean that there may be additional elements in addition to the listed elements. Additionally, it should be understood that references to "one implementation" or "implementations" of the present disclosure are not intended to be construed as precluding the existence of additional implementations that also include the recited features. For example, any element or feature described herein with respect to an implementation may be combined with any element or feature of any other implementation described herein, where such elements or features are compatible.
[0136] The present disclosure may be embodied in other specific forms without departing from its spirit or characteristics. The described implementations are to be considered illustrative and not restrictive. Accordingly, the scope of the present disclosure is indicated by the appended claims rather than the foregoing description. Changes within the meaning and range of equivalents of the claims will be covered within their scope.
Claims
1. A computer-implemented method, comprising: at a management entity, receiving a power consumption signal corresponding to the power consumption of a host system on a host device; at the management entity, using a power fault detection model that determines a power consumption fault of the host device based on the power consumption signal to determine that the host device is unhealthy, wherein the power fault detection model includes a first rule-based layer and a second machine learning layer; determining whether to use the first rule-based layer or the second machine learning layer to determine that the host device is unhealthy based on the power consumption signal received from the host device and usage characteristics; and based on determining that the host device is unhealthy, initiating, from the management entity, a recovery process for restoring the host device.
2. The computer-implemented method according to claim 1, further comprising: determining to use the first rule-based layer based on identifying a rule corresponding to the power consumption signal from the host device; and determining that the host device is unhealthy based on the power consumption signal satisfying the rule.
3. The computer-implemented method according to claim 1 or claim 2, further comprising: determining to use the second machine learning layer having a power fault machine learning model based on a failure to identify a rule corresponding to the usage characteristics or the power consumption signal of the host device; using the power fault machine learning model to determine a power consumption range for the host device based on the usage characteristics of the host device; and determining that the host device is unhealthy based on the power consumption signal being outside the power consumption range for the host device.
4. The computer-implemented method according to claim 3, wherein the power consumption signal and the usage characteristics received at the management entity correspond to operations being performed by a host processor of the host system of the host device.
5. The computer-implemented method according to claim 1, wherein the usage characteristics of the host device include hardware information of the host device, software information of the host device, and workload information of virtual machines running on the host device.
6. The computer-implemented method according to claim 1, wherein: the power consumption signal of the host system is received at the management entity from an auxiliary service controller on the host device; and the host device includes a processor separate from the host system and a power supply for the auxiliary service controller.
7. The computer-implemented method according to claim 6, wherein the auxiliary service controller determines the power consumption signal using a power measurement device on the host device.
8. The computer-implemented method according to claim 1, further comprising providing a recovery instruction to the host system via an auxiliary service controller on the host device.
9. The computer-implemented method according to claim 1, wherein initiating the recovery process includes providing, from the management entity, a recovery instruction for restarting the host system to the host device.
10. The computer-implemented method according to claim 1, further comprising: At the management entity, receive a second power consumption signal corresponding to a second host system on a second host device; At the management entity, determine that the second host device is healthy based on the second power consumption signal by using the power failure detection model; And Based on determining that the second host device is healthy, delay the recovery process for recovering the second host device.
11. A system, comprising: At least one processor at a server device; And A computer memory including instructions that, when executed by the at least one processor at the server device, cause the system to perform operations, the operations including: At a management entity, receive a power consumption signal corresponding to the power consumption of a host system on a host device; At the management entity, use a power failure detection model that determines a power failure of the host device based on the usage characteristics or the power consumption signal of the host device to determine that the host device is unhealthy; and Based on determining that the host device is unhealthy, initiate a recovery process for recovering the host device from the management entity.
12. The system according to claim 11, wherein the power failure detection model determines that the host device is unhealthy based on the power consumption signal of the host device being lower than a lower power consumption threshold determined by a power failure machine learning model for the host device.
13. The system according to claim 11, wherein the workload information from the usage characteristics of the host device includes the number of operations on the host device and the type of operations on the host device.
14. The system according to claim 11, further comprising instructions that, when executed by the at least one processor, cause the system to perform additional operations, the additional operations including delaying the initiation of the recovery process based on detecting a mitigation event corresponding to the host device.
15. The system according to claim 11, wherein: The host device and the management entity are located in a data center; and The management entity manages a plurality of host devices, and the host device is one of the plurality of host devices.
16. The system according to claim 11, further comprising instructions that, when executed by the at least one processor, cause the system to perform additional operations, the additional operations including: Determine a first rule-based layer that does not utilize the power failure detection model; And In response to determining the first rule-based layer that does not utilize the power failure detection model, use a second machine learning layer of the power failure detection model to determine that the host device is unhealthy.
17. A computer-implemented method, comprising: At a management entity, receive a power consumption signal corresponding to the power consumption of a host processor and a host memory on a host device; Generate a power failure detection model, where the power failure detection model includes a first rule-based layer and a second machine learning layer with a power failure machine learning model, and the power failure detection model determines a power consumption failure of the host device based on usage characteristics and power consumption signals of the host device; At the management entity, determine that the host device is unhealthy by comparing the power consumption signal with a power consumption range determined for the host device by the power failure detection model based on one or more usage characteristics of the host device; And Based on determining that the host device is unhealthy, initiate a recovery process for restoring the host device from the management entity.
18. The computer-implemented method according to claim 17, wherein the power failure machine learning model determines that the host device is unhealthy based on the power consumption signal of the host device being lower than a lower threshold of the power consumption range determined for the host device by the power failure machine learning model.
19. The computer-implemented method according to claim 17, wherein the power failure machine learning model determines that the host device is unhealthy based on the power consumption signal of the host device being higher than an upper threshold of the power consumption range determined for the host device by the power failure machine learning model.
20. The computer-implemented method according to claim 17, wherein initiating the recovery process includes providing a recovery instruction for restarting the host device from the management entity to the host device.