Infrastructure management system for hardware fault repair
By defining a health model and minimum operational limits for hardware in the edge infrastructure, hardware can operate in a degraded state, which solves the problem of hardware failures not being repaired in a timely manner in the edge infrastructure, improves hardware utilization and repair efficiency, and ensures SLA compliance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2016-12-26
- Publication Date
- 2026-08-04
AI Technical Summary
In edge infrastructure, hardware failures cannot be repaired in a timely manner, resulting in low hardware utilization, affecting workload load balancing, and existing methods cannot effectively manage hardware failures, making it difficult to meet service level agreements (SLAs).
By defining the health model and minimum operating limits of the hardware through the infrastructure management system, the hardware can be allowed to operate in a degraded state. Combined with the monitoring, degraded state provisioning and return authorization (RMA) system, flexible repair of hardware failures can be achieved, ensuring that the hardware can still meet the minimum operating requirements when it fails.
It improves hardware utilization and repair efficiency, ensures that service level agreements (SLAs) are not affected, and enables flexible management and timely repair of hardware failures.
Smart Images

Figure CN108431836B_ABST
Abstract
Description
Background Technology
[0001] Large-scale networked systems are common platforms used in various settings for running applications and maintaining business and operational functions. For example, a data center (e.g., a physical cloud computing platform) can simultaneously provide various services (e.g., web applications, email services, search engine services, etc.) to multiple customers. These large-scale networked systems typically involve a large number of resources distributed throughout the data center, each of which is analogous to a physical machine or virtual machine (VM) running on a physical node or host. Data centers operate on hardware components that may occasionally fail. In some cases, the failed hardware component can be easily replaced. However, in other cases, the hardware component cannot be replaced immediately. Therefore, a comprehensive system for configuring and implementing data center hardware components, as well as failing data center hardware components, can improve overall data center hardware operation and distributed hardware management to meet defined objectives. Summary of the Invention
[0002] The embodiments described herein provide methods and systems for implementing an infrastructure management system that supports hardware failure remediation. The infrastructure management system can be implemented based on an infrastructure management system platform that includes operatively integrated components to reduce the impact of faulty hardware in the hardware infrastructure of a distributed computing system. The infrastructure management system supports configuration patterns that help define configuration profiles for hardware. Configuration patterns can be data structures used to represent or define configuration attributes of hardware in the computing infrastructure. Specifically, configuration patterns include a health model for the hardware. The health model is a technical representation of the hardware's computational conditions. The hardware configuration pattern and health model can be defined in a configuration profile. The health model further defines minimum operational limits for the hardware based on health metrics or optional and required components associated with the hardware. The minimum operational limit is used as a threshold to allow the hardware to operate in a degraded state rather than causing complete hardware failure. In this respect, the infrastructure management system improves hardware utilization by allowing hardware that would otherwise be designated as faulty to operate in a degraded state before repair or replacement.
[0003] During operation, it is determined that a hardware component failure has occurred. The hardware component is part of a hardware assembly. The hardware assembly's repair attributes are accessed. These repair attributes indicate the minimum operational limits for the hardware assembly. Minimum operational limits can be based on health metrics or the optional and required components of the hardware assembly. Minimum operational limits help determine whether the hardware assembly should operate in a degraded state.
[0004] The minimum operational requirements of the hardware assembly are met even when no faulty hardware components are present. Operation of the hardware assembly in a degraded state is initiated. A degraded state includes the hardware assembly operating without any hardware components. In embodiments, a hardware manager (e.g., operating system and return authorization) is associated with a degraded state configuration to facilitate the initiation of operations to operate and repair the hardware assembly in a degraded state. When a degraded state is anticipated, a degraded state configuration can be defined to support hardware assembly operations and infrastructure management operations for hardware assemblies operating in a degraded state.
[0005] The summary is provided to introduce, in a simplified form, some concepts that will be further described in the following detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation as an auxiliary means of determining the scope of the claimed subject matter. Attached Figure Description
[0006] The present invention will now be described in detail with reference to the accompanying drawings, in which:
[0007] Figure 1 This is a block diagram of an exemplary distributed computing infrastructure environment in which the embodiments described herein can be implemented;
[0008] Figure 2A and Figure 2B This is a block diagram of an exemplary implementation of an infrastructure management system for hardware fault repair according to the embodiments described herein;
[0009] Figure 3 This is a block diagram of an exemplary implementation of an infrastructure management system for hardware fault repair according to the embodiments described herein;
[0010] Figure 4 This is a flowchart illustrating an exemplary method for implementing an infrastructure management system for hardware fault repair according to embodiments described herein;
[0011] Figure 5 This is a flowchart illustrating an exemplary method for implementing an infrastructure management system for hardware fault repair according to embodiments described herein;
[0012] Figure 6 This is a block diagram of an exemplary computing environment suitable for implementing the embodiments described herein; and
[0013] Figure 7 This is a block diagram of an exemplary distributed computing system suitable for implementing the embodiments described herein. Detailed Implementation
[0014] Edge computing typically refers to pushing the boundaries of computing applications, data, and services from centralized nodes to the logical end of the network. In this way, a cloud computing network service provider's distributed computing system can include edge infrastructure supporting geographically dispersed customers. Edge infrastructure can be specifically deployed based on identified traffic and usage patterns within the distributed computing system. In this respect, client devices can access the distributed computing system from the central infrastructure of the edge infrastructure. Edge infrastructure can include hardware or hardware assemblies in data center racks, located as close as possible to the customers, and is not centralized.
[0015] As used interchangeably herein, the phrases and terms “hardware assembly,” “hardware manifest,” or “hardware” do not mean a component limited to any particular configuration, but rather broadly refer to any individual device, assembly of devices (e.g., network devices, computing devices, and power devices), and its components that can be integrated into a rack within a distributed computing infrastructure. A hardware assembly, hardware manifest, or hardware may include individual hardware components that can be independently defined or configured as hardware with reference to the functionality described herein. While the edge infrastructure and some specific challenges within the embodiments described herein are illustrated by way of example, it is conceivable that the described methods and systems can be implemented in other types of infrastructure with hardware. In one instance, the hardware may reside within a private enterprise network managed by a customer of a cloud computing network service provider. In another example, the hardware may reside within a data center managed by a cloud computing network service provider.
[0016] Edge infrastructure located in partner locations of cloud computing network service providers can present challenges in resolving hardware failures. Edge infrastructure in partner locations may have different policies regarding hardware access, control, and operational standards. Therefore, hardware failures may not be resolved immediately compared to infrastructure wholly owned and / or operated by the cloud computing network service provider. The repair schedule for failed edge infrastructure hardware may only be implemented interimly and / or may be delayed by several months. This limits the maximum number of hardware devices that can be flagged as unhealthy (i.e., failed) and placed offline before technicians can perform repairs. Edge infrastructure hardware also often has limited backup hardware, making the impact of hardware failures in the edge infrastructure very significant. For example, load balancing of workloads in the edge infrastructure becomes much more difficult when several machines fail and are offline.
[0017] Conventional methods for resolving hardware failures rely on the immediate removal, replacement, or repair of hardware or hardware assemblies. This hardware failure strategy primarily depends on the large volume of hardware in data centers wholly owned and controlled by cloud computing network service providers, or their immediate access to the data center. However, this solution may not always be feasible, and alternative solutions may be more efficient in certain situations. Furthermore, with the increasing implementation of edge infrastructure, the standard hardware failure strategy of immediately removing, replacing, or repairing hardware in bulk may become unacceptable, necessitating an alternative approach.
[0018] The embodiments described herein relate to simple and efficient methods, systems, and computer storage media for implementing an infrastructure management system that supports hardware failure remediation. At a high level, monitoring, degraded state provisioning, and return authorization (RMA) systems, processes, and components are configured to support hardware failure remediation. Hardware failure remediation allows a hardware assembly to operate in a degraded state, where healthy hardware components in the assembly operate alongside faulty hardware components. The infrastructure management system supports configuration patterns that help define hardware profiles. Specifically, the configuration patterns include a hardware health model. The health model is a technical representation of the hardware's computational conditions. In particular, the health model defines minimum operational limits for the hardware based on health metrics or optional and required components associated with the hardware. These minimum operational limits are used as a threshold to allow the hardware to operate in a degraded state rather than causing complete failure of the hardware assembly. This results in maximizing the utilization of the hardware assembly.
[0019] Infrastructure management systems can be implemented for distributed computing system infrastructures (e.g., cloud computing infrastructure). In particular, such systems can be implemented for edge infrastructures that are difficult to access for hardware failure remediation. Implementing hardware failure remediation can also advantageously improve RMA systems by allowing timely scheduling of repairs in distributed computing infrastructures to achieve better hardware utilization and efficiency. Timely scheduling of repairs can include planning hardware repairs so that service level agreements (SLAs) with customers are unaffected or minimally affected during the repair process. Timely scheduling of repairs can also be based on the availability of alternative hardware and technicians to perform the repair operations.
[0020] Improving hardware availability and utilization hinges on defining hardware resilience. Hardware resilience refers to relaxing mandatory health requirements for hardware. Hardware resilience can be based on health metrics or optional and required components of the hardware. For example, for a hardware assembly, a health model can be defined, which includes: health status, health metrics to track, optional and required components, minimum operational limits, and other attributes. A hardware assembly can include functional (healthy) components and faulty (unhealthy) components in the event of a failure. Functional and faulty components can be evaluated, and if the hardware assembly still meets the minimum operational limits, the hardware can be reconfigured and brought back online to operate in a degraded state while awaiting RMA action. Hardware resilience can be specifically defined as part of a configuration pattern that supports defining profiles for the hardware. Hardware resilience defined in minimum operational limits can be defined in the remediation attributes of the configuration pattern. Hardware resilience can be further defined or adjusted to conform to or be consistent with the Service Level Agreement (SLA) of the tenant using the hardware. An SLA is a contract between a cloud computing network service provider and a customer that defines the expected service. For example, optional components are defined for Stock Units (SKUs) to keep a particular SKU online without optional components while it still meets the agreed service level, rather than simply making the SKU operational. As an example, a machine may include multiple hard drives that are part of a standard deployment and typically go offline in the event of a machine failure. However, in the scenario described herein, the machine can be reconfigured to operate using fewer drives than all of the standard deployment if it still meets minimum operating limits. Furthermore, in some cases, minimum operating limits must also meet a tenant's SLA.
[0021] During operation, it is determined that a hardware component failure has occurred. The hardware component is part of a hardware assembly. The hardware assembly's repair attributes are accessed. These repair attributes indicate the minimum operational limits for the hardware assembly. Minimum operational limits can be based on health metrics or the optional and required components of the hardware assembly. Minimum operational limits help determine whether the hardware assembly should operate in a degraded state.
[0022] It is determined that the hardware assembly still meets minimum operational constraints even when operating without faulty hardware components. Operation of the hardware assembly in a degraded state is initiated. A degraded state includes the hardware assembly operating without any hardware components. In an embodiment, a hardware manager (e.g., an operating system) is associated with a degraded state configuration to facilitate the initiation of the operation and to enable the hardware assembly to operate in a degraded state. When a degraded state is anticipated, the degraded state configuration is defined to support hardware assembly operations and infrastructure management operations for the hardware assembly operating in a degraded state.
[0023] Accordingly, refer to Figure 1 The distributed computing infrastructure 100 supports an infrastructure management system platform that provides integrated functionality based on components of the platform described herein. The distributed computing infrastructure 100 includes an infrastructure management system 110, edge infrastructure 130, central infrastructure 140, administrator clients 150, vendor clients 160, and customer clients (170a and 170b). The components described herein communicate using a network (not shown), which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs). Such network environments are common in offices, enterprise-wide computer networks, intranets, and the Internet. Accordingly, the network is not further described herein.
[0024] For example, administrator client 150, supplier client 160, and customer clients (170a and 170b) may include those referenced herein. Figure 6 The described computing device 600 can be of any type. Administrator client 150, vendor client 160, and customer clients (170a and 170b) can provide access to the different components described herein. Specifically, as further described herein, administrator client 150 and vendor client 160 can access infrastructure management system 110 to perform one or more operations facilitated by infrastructure management system 110. Customer client 150a can access resources in distributed computing infrastructure 100 via central infrastructure 140, and customer client 150b can access resources in distributed computing infrastructure 100 via edge infrastructure 130.
[0025] As used herein, "platform" refers to any system, computing device, process, or service, or a combination thereof. A platform can be implemented as hardware, software, firmware, dedicated equipment, or any combination thereof. A platform can be integrated into a single device or distributed across multiple devices. The various components of a platform can be co-located or distributed. The platform can be formed from other platforms and their components.
[0026] In addition to those shown or as alternatives, other arrangements and elements (e.g., machines, interfaces, functions, sequences, and functional groupings, etc.) may be used, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or combined with other components and implemented in any suitable combination and location. The various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software. For example, various functions can be performed by a processor that executes instructions stored in memory.
[0027] Distributed computing infrastructure 100 can rely on infrastructure management system 110 for hardware fault repair. Infrastructure management system 110 is responsible for managing the infrastructure's hardware (e.g., edge infrastructure). Infrastructure management system 110 can be implemented via data center infrastructure management services to define and deploy hardware using configuration files with configuration attributes that express the requirements, health status, repair, and settings of a specific machine SKU. Administrator client 150 can facilitate the configuration and management of the infrastructure management system's operations using services, configuration patterns, configuration files, and SKUs. SKUs can be used to describe hardware, hardware assemblies, and hardware assembly components, whereby an SKU represents attributes associated with hardware and distinguishing it from other hardware (e.g., manufacturer, product description, BIOS, firmware, configuration, material, size, color, packaging, and warranty terms). It is conceivable that an SKU can also refer to a unique identifier or code relating to a specific inventory unit. Infrastructure management system 110 can, in particular, receive and store configuration patterns that include repair attributes indicating minimum operational limits. Minimum operational limits can be based on health metrics or on optional and required components of a hardware assembly for a specific machine SKU. In this regard, the minimum operational constraint refers to the basic health requirements of a subset of operable hardware components, compared to the requirement of having a complete set of operable hardware components. If the hardware assembly fails but some of its components meet the basic operational requirements, the hardware components will still be used; however, if the basic operational requirements are not met, the hardware components will not be used.
[0028] Health metrics defined in the health model of a hardware SKU can quantify the minimum operational limits used for hardware repair assessment. Minimum operational limits can be dynamic or static. Profiles can be updated to indicate different minimum operational limits based on multiple factors. For example, traffic patterns, accessibility to edge infrastructure, traced failure rates, and administrative actions may be factors that determine minimum operational limits and further support the dynamic allocation of minimum operational limits for hardware SKUs. Therefore, at least some hardware in the edge infrastructure can be associated with the health model to indicate minimum operational limits. In this respect, a hardware assembly failure does not render the entire hardware assembly inoperable unless the hardware failure causes the hardware assembly's operational limits to fall below the minimum operational limits indicated by the health metrics.
[0029] As an example, a machine associated with four physical disks can have a health model that instructs the infrastructure management system to monitor the number of healthy disks as an associated health metric. The health model based on this health metric can define a minimum operational limit as a machine operating with at least two disks. In this respect, a maximum of two disks on the machine may fail, and the machine may still be operational or be reproduced to operate with two disks. Reproducibility may be part of a comprehensive repair operation of the hardware assembly, allowing the hardware assembly to operate in a degraded state. It is conceivable that the SLA protocol can be a factor in defining the minimum operational limit. For example, if the SLA further requires at least three disks for the machine, then even if the machine can operate with two disks, it may be associated with a minimum operational limit of three disks for use by a tenant associated with the SLA because the machine cannot meet the SLA. Thus, it is possible to deactivate the hardware assembly for use with a first tenant having a first SLA, but not deactivate (or reproducible) the same hardware assembly for a second tenant having a second SLA. Other variations and combinations of defining and implementing minimum operational limits can be conceived using the embodiments described herein.
[0030] Degraded state configurations can include configurations defined within a distributed computing system infrastructure adapted for hardware failure recovery. As described herein, a degraded state configuration can be associated with a hardware manager (e.g., operating system, hypervisor, infrastructure controller, RMA portal) when a degraded state is anticipated. In particular, a degraded state configuration can include instructions on how the hardware manager should configure and operate the hardware assembly when it operates in a degraded state, which differs from a non-degraded state. For example, if a disk failure occurs, the physical disk is not statically mapped to a logical drive to account for operation in a degraded state. In this regard, a degraded state configuration can be pre-configured or defined within the hardware manager. In embodiments, a degraded state configuration can be pre-configured to change the hardware manager's traditional configuration when it is anticipated that hardware will operate in a degraded state.
[0031] Degraded status configurations can be associated with a data center manager. The data center manager can define new names or labels to capture the status of healthy hardware assemblies associated with unhealthy hardware components. The data center manager can mark hardware components as healthy or unhealthy, but also includes hardware component attribute fields (“attribute fields” or “machine attributes”) to indicate that a hardware component is unhealthy. Similarly, marking hardware assemblies with hardware components with attribute fields can help indicate faulty hardware components, allowing faulty hardware components to be ignored, and additionally, faulty hardware components can be replaced under RMA, as discussed below. As an example, for servers, additional attribute fields can indicate missing disks and bad disks. During monitoring, the monitor service operates to read attribute fields and avoids reporting errors of disks already marked as bad or missing. Infrastructure management systems based on degraded status configuration attributes can target edge infrastructure (e.g., edge SKUs and environments) for hardware repair functionality, while excluding centralized infrastructure. Configuration patterns can include hardware repair function trigger attributes. Configuration pattern-based profiles for hardware components and SKUs can specifically define trigger attributes to indicate when hardware repair functionality is applied to a specific hardware infrastructure. In this regard, degraded state configuration in distributed computing system components can support hardware assembly operations and infrastructure management operations for hardware assemblies running in degraded state.
[0032] Infrastructure Management System 100 can operate in conjunction with an existing RMA to support RMA, at least in part, based on the hardware fault repair capabilities described herein. As part of Infrastructure Management System 100, the RMA supports both timely and immediate RMA operations. The RMA can utilize attribute fields of a hardware assembly to identify specific health status information of its hardware components, which must be addressed under the RMA. Timely scheduled repairs can include planned repairs to the hardware assembly such that the Service Level Agreement (SLA) with the customer is unaffected or minimally affected during the repair. For example, in a conventional model, hardware identified as faulty would go offline, impacting and potentially causing the SLA to be unmet. Using the Infrastructure Management System, hardware is allowed to operate in a degraded state as long as the minimum health requirements are met. As a result, hardware can be scheduled and repaired while the SLA remains met. Immediate RMA operations can be performed to immediately repair the hardware to full health. In an embodiment, immediate RMA operations are based on compliance with the SLA requirements for the tenant associated with the hardware. The tenant can be removed from the hardware during repair so that the tenant does not encounter any unexpected failures.
[0033] Continue to refer to Figure 1 ,like Figure 1As shown, the infrastructure management system 110 may include multiple components supporting the provision of hardware fault repair capabilities as described herein. The infrastructure management system 110 includes a monitoring component 112 (WD 112), a data center manager component 114 (DM 114), a repair service component 116 (RS 116), a provisioning service component 118 (PS 118), and an RMA component 120 (RMA 120). The infrastructure management system 110 uses configuration modes and corresponding configuration files to monitor, configure, repair, provision, and provide status information for the RMA used by the hardware. In embodiments, the configuration modes are based on SKUs and configuration attributes as described herein. The configuration modes include repair attributes that indicate minimum operational limits for the hardware assembly. Minimum operational limits may be based on health metrics or optional and required components of the hardware assembly. Minimum operational limits support the determination of whether the hardware assembly should operate in a degraded state.
[0034] As an example, a disk description within a machine or node can be used to describe the failure lifecycle. For instance, at a higher level, WD 112 can access configuration modes and perform health checks on the disks based on these modes to ensure the minimum required number of disks are healthy. Configuration modes can specifically specify the required number of disks and the number of optional disks based on SKU definitions. Configuration modes can also be based on tenant SLAs, allowing degraded hardware to handle changes in optional disks. WD 112 can report that some disks have failed, but the minimum required number of disks are healthy. Based on the health information report that the hardware assembly meets minimum operational constraints, PS 118 re-provisions the hardware and installs the healthy disks. During degraded provisioning, PS 118 uses health status information and configuration modes from WD 112 to provision the hardware assembly in degraded state. Therefore, hardware assemblies operating in degraded state minimize the impact of these changes on tenants.
[0035] As discussed herein, degraded state configurations can be associated with hardware managers (e.g., operating systems, hypervisors, architecture controllers) when a degraded state is anticipated. In one exemplary implementation, degraded state configuration is supported via configuration modes and corresponding configuration files. For example, in specific cases involving disks and the operating system supporting those disks, configuration modes disrupt the static mapping between physical drives and logical drives (volumes). Depending on the machine's operating condition, different physical drive slots are used as system volumes. As long as enough physical disks are healthy enough to meet minimum requirements for logical drives and capacity metrics, the machine is utilized and marked as healthy. Volumes and logical drives in optional categories may not always be created, and tenant applications are aware of this configuration and anticipate that these volumes may not exist. Thus, it is conceivable that the operating system and other applications operating with full-capacity hardware are informed or notified (i.e., programmed and reconfigured) as needed to adapt to and tolerate hardware operating in a degraded state. For example, the operating system could be pre-configured to boot to any drive in anticipation of a boot drive failure, and thus be able to recover to a degraded state on any drive.
[0036] This example implementation is based on an exemplary hardware SKU of 4 JBODs (Just a Bunch of Disks). The hardware can operate and service traffic as long as 2 disks are online and healthy. This evaluation can be based on testing and / or meeting expected SLAs. The basic requirement for the SKU is 2 disks, and the remaining 2 disks are considered optional. In this embodiment, to accommodate this flexibility, physical drives are not statically mapped to logical drives. There may not be a fixed mapping from disk controller slots to logical disks, but the assignment is consistent. The lowest-functioning disk controller slot can be labeled as logical disk 0. If physical disk 0 exposed from the controller is unhealthy, then physical disk 1 exposed from the controller will become logical disk 0. No additional state is stored to calculate the mapping, as it is a consistent algorithm, provided that consistent hard drive validation checks are performed and the machine event audit log is accessible.
[0037] In another example, all four disks of a hardware SKU with four disks may fail. Minimum operational limits can instruct the hardware SKU to operate in a degraded state, with only two of the four failed disks operational. In practice, in some cases, it may be advantageous to repair at least a portion of the hardware. Specifically, hardware can be repaired to meet minimum operational limits. In this respect, the entire hardware SKU is not lost. Similarly, a portion of the hardware in a deployment rack can be repaired during a transition period before the rest of the hardware is repaired. For example, if there are 10 degraded blades in a rack, the repair operation could include repairing two blades to have enough blades to keep the rack operational, rather than repairing all 10 blades simultaneously. Allowing a portion of the hardware to be repaired and operational can also be specifically based on the satisfaction of the SLA of the tenant associated with the hardware.
[0038] The configuration file specifies the basic requirement as having healthy logical disks 0 and 1, while disks 2 and 3 are optional. Volume specifications remain unchanged in this regard. Volumes are still created based on logical disk assignments. If no matching logical disk is provided for a specified volume, the volume is not created. The physical-to-logical disk partial mode can be extended, and the volume information remains the same. As mentioned above, the local operating system can be configured to expect that optional volumes may not exist and to adjust its behavior based on which volumes have already been provisioned. Tenants can make assumptions, but the basic requirement for volumes always exists on healthy nodes.
[0039] Prior to the provisioning phase in degraded state, WD 112 can access machine 210 status information and configuration files in DM 114 to make decisions regarding health metrics. A key difference is that health requirements and WD 112 behavior will change based on the machine status stored in DM. If a hardware component is marked as faulty in DM 114, WD 112 may not monitor the unhealthy component. WD 112 only monitors healthy hardware and reports any faults to DM. In the specific case of disks, WD only monitors disks marked as healthy in DM and provisioned. Once a problem is detected, WD 112 reports the hardware problem to DM 114, as in the previous procedure. WD 112 can be configured to report all hardware problems. RS 116 can attempt to resolve issues after the upgrade disassembly mode. In the final mitigation scenario, RS 116 can request PS 118 to re-provision (i.e., degraded state provisioning) the machine.
[0040] refer to Figure 2A and Figure 2B , Figure 2A and Figure 2BA method for implementing an infrastructure management system is illustrated. DM 114 is responsible for managing the hardware in the distributed computing infrastructure. DM 114 is responsible for receiving and storing configuration patterns of the hardware. Configuration patterns can be developed for specific SKUs of hardware or hardware assemblies having individual hardware components (e.g., a physical machine having a disk, NIC (Network Interface Controller), memory, processor, chip, etc.). DM 114 also serves as a repository for health status information. The distributed computing infrastructure may include a machine 210 supported for hardware repair functions. Machine 210 represents exemplary hardware or hardware assemblies consistent with the functions described herein. DM 114 stores and provides access to health status information of the hardware infrastructure (e.g., edge infrastructure 130). The health status information is based on the configuration patterns and configuration files of the corresponding hardware in the hardware infrastructure. As described herein, configuration patterns can be defined based on health models and SKUs.
[0041] In step 212, WD 112 accesses and retrieves health status information from DM 114 to identify healthy hardware in the hardware infrastructure to be selectively monitored. WD 112 operates based on configuration information to monitor and report any hardware failures. It is conceivable that WD 112 may also report health SLA failures to DM that are optionally considered as a factor in the determination of whether a hardware component is marked as healthy or unhealthy, as described in more detail herein. In step 214, WD 112 utilizes the hardware profile and health status to determine how to monitor the edge infrastructure. For example, WD 112 determines which hardware and hardware components are healthy and need to be monitored, and which health metrics and / or health SLAs will be monitored, particularly those associated with hardware repair capabilities. In step 216, WD 112 monitors the edge infrastructure, machine 210, to identify faulty hardware components.
[0042] In step 218, based on monitoring, WD 112 can detect failures in hardware components within the hardware assembly (e.g., machine components in machine 210). The infrastructure management system 110 includes configuration modes that provide flexibility to define certain hardware components as needed and define other hardware components as optional and further health metrics for hardware components and corresponding threshold conditions. Hardware failures can be mapped to optional categories of hardware components or a certain health metric threshold condition (e.g., minimum operating limits). Optional hardware identified in the mapping indicates hardware components that may fail, and the hardware assembly can be restored without immediately repairing the optional hardware. Therefore, the machine continues to operate as long as the required components and health metrics are met. For example, a specific disk mapped to an optional requirement may fail while the computer is still marked as healthy, and a similar number of disks determined to be at or above the minimum operating limit may fail while the machine is still marked as healthy. Therefore, the hardware assembly can only be marked as unhealthy (or offline) if the minimum operating limit is not met. For example, WD 112 can monitor machine 210, and the infrastructure management system can ensure that the required minimum number of disks are healthy. Alternatively, when minimum operating limits are defined based on optional and required components, any optional component marked as unhealthy will not cause the machine to be marked as unhealthy. In step 220, WD 112 reports the detected hardware failure to DM 114. Machine 210 may be repaired before being processed for resupply. It is conceivable that the repair service step may be optional for some or all types of hardware assemblies.
[0043] RS 116 operates to perform repair operations (e.g., restarting system services, soft reboot, and hard reboot). In an exemplary implementation, a soft reboot may specifically refer to a software restart and a hard reboot may refer to a hardware reset. In step 222, RS 116 accesses hardware health status information from DM 114 (e.g., pulls hardware status) and in step 224 attempts a repair action to repair the faulty hardware component. As shown in step 226, if the repair action fails, RS 116 may transmit a request to perform a repair operation on the hardware component in step 228. In step 230, PS 118 initiates a repair operation for machine 210 (i.e., the hardware assembly including the faulty hardware component). The repair operation (e.g., resupply) may refer to: resupplying the degraded state of the functional hardware components of the hardware assembly while excluding the faulty hardware component. Resupply may be based on a configuration file in DM 114 and health status information updated via WD 112. The configuration file is used to verify configuration attributes. Functional hardware components are marked as healthy so that the infrastructure is not exposed to unhealthy hardware components for operation or monitoring. If an additional hardware problem is detected during redeployment, PS 118 may cause the redeployment operation to fail and mark machine 210 as unhealthy, causing machine 210 to lose rotation. In step 232, PS 118 initiates a redeployment operation to attempt to reconfigure machine 210. In one exemplary implementation, as part of the redeployment, PS 118 loads a pre-execution environment (PXE) onto machine 210. In step 234, the pre-execution environment accesses fault information stored in DM 114. Based on the fault type, the pre-execution environment alters its conventional provisioning behavior. As an example, the pre-execution environment verifies the health status of the disks on the machine in step 236. If a disk fails the health requirements but is not marked as unhealthy in DM 114, the pre-execution environment marks the disk as unhealthy.
[0044] The remaining healthy disks are compared to the basic health requirements (e.g., minimum operating limits) of the machine SKU. If the number of healthy disks matches the number required by the basic health requirements, the provisioning process continues. The pre-execution environment continues formatting and provisioning the disks. The pre-execution environment can use a boot tool to modify the BIOS settings and boot order in step 238 to enable the machine to operate on the newly selected system disk. In step 240, WD 112 stops monitoring the failed disks, and machine 210 sends an indication to PS 118 that the re-provisioning has been successfully completed. In step 242, a message indicating that the re-provisioning has been successful is sent to RS 116.
[0045] refer to Figure 2B An exemplary resupply implementation for the infrastructure management system 110 is shown. Figure 3The system includes machine 210, machine operating system 210 (MOS 250), PS 118, and DM 114. In one embodiment, the degraded state provisioning may include specific exemplary implementation details. When machine 210 boots, the machine transmits a PXE boot request in step 252. PS 118 receives the request and accesses (and / or updates) machine information from DM 114 in step 254 to determine a response. PS 118 chooses to load a pre-installed environment (PE) image onto the machine in step 256 and updates the state in DM 114 to reflect this action. After the PE image is loaded, machine 310 boots into PXE in step 258. In step 260, MOS 310 accesses and retrieves the machine's configuration files and DM 114 to obtain the machine's health status. MOS 310 performs diagnostics in step 262 to verify disk lifetime and health status. In step 264, updated machine information is transmitted to DM 114. For example, any disk detected as unhealthy but not marked as unhealthy in DM 114 is marked accordingly in DM 114. The repair operation involves verifying that the number of healthy physical disks matches the basic requirement for healthy disks specified in the configuration file. As long as the basic number of healthy disks exists, the provisioning process continues.
[0046] In step 266, the process establishes and supplies only healthy drives. For example, MOS 310 selects the first healthy physical disk as logical disk 0 to host the system volume. The supply process downloads the operating system image and installs the operating system on the system volume. The remaining healthy disks are supplied sequentially as remaining volumes, i.e., the next healthy physical disk corresponds to logical disk 1, and it is a matching volume. After setting up the drives, MOS 310 changes the boot settings to ensure that the first healthy physical disk is marked as the system boot disk. In step 268, MOS 310 updates the machine 210 information in DM 114, and in step 270, MOS 310 boots the machine to the operating system.
[0047] refer to Figure 3 , Figure 3 An implementation of an infrastructure management system for hardware fault repair is shown. Specifically, Figure 3 The RMA operation process for the infrastructure management system is shown. Figure 3 This includes supplier client 160, DM 114, PS 118, and machine 210. Figure 3It also includes an RMA component 120 with an RMA portal 302, an RMA status 304, and a synchronization agent 306. The RMA 120 provides an RMA portal that operates as a gateway or access point to view the hardware status in the distributed computing system. The RMA portal can provide access to view unhealthy hardware set to RMA. The RMA portal tracks and displays the hardware status. The hardware status is stored in the RMA status 304. The synchronization agent 306 facilitates the coordination of status changes between the RMA 120 and the DM 114. A vendor client 160 accesses the DM 114, which stores hardware health status information, via a publicly accessible portal.
[0048] As discussed in this article, hardware supported by an infrastructure management system can operate in a degraded state. The hardware can provide live traffic but has unhealthy hardware components. RMA 120 allows hardware to be marked with two status messages—“Degraded” and “PendingRMA”—to support hardware failure recovery via RMA components. A degraded state indicates that the machine is operating while a hardware component has failed, while PendingRMA indicates that the vendor has requested that the machine be moved to OFR (Out of Recovery). The vendor can also access the infrastructure management system and move the hardware from “PendingRMA” to immediate “RMA” based on the SLA requirements of the tenant associated with the hardware. During hardware recovery, the tenant can be removed from the hardware to prevent unexpected failures.
[0049] Continue to refer to Figure 3Initially, in step 310, the vendor can request to take the hardware (e.g., machine 210) offline via vendor client 160. When the vendor (i.e., service technician) requests to move the target machine 210 to OFR in DM 114, the RMA portal updates or submits the status to PendingRMA in the RMA portal at step 312. As shown in step 314, synchronization agent 306 is configured to periodically pull status information from RMA status 304. DM 114 is also configured to periodically pull status information from DM 114, as shown in step 316. In step 318, RMA 120 then determines the action to be taken for the hardware with the pending RMA status. When DM 114 is in a healthy state (because at least one hardware component is still operational), RMA 120 can update DM 114 in step 320 to request to move the hardware to OFR in DM 114. PS 118 is also configured to periodically retrieve status information from DM 114. Thus, PS 118 picks up the machine's state change, as shown in step 322, and in step 324, the machine's cancellation supply process begins. For example, in step 324, PS 118 may transmit a request to initiate wiping the machine, and in step 326, machine 210 is wiped. In step 328, the machine may optionally be shut down.
[0050] Continuing with the exemplary implementation of machine 210, after the machine has completed unsupply and shutdown, in step 330, PS 118 notifies DM 114 that machine 210 is in OFR within DM 114. When the DM 114 status is OFR, RMA 120 can update the RMA portal via steps 332, 334, 336, and 338 (these steps illustrate the periodic retrieval of status information) to mark the machine as OFR, which will then be displayed in the portal. In step 340, vendor client 160 can retrieve the status information from RMA 102, allowing the vendor to commence service after seeing the status update at step 342. It is conceivable that RMA portal 302 may not provide feedback to the user other than indicating that the machine is in a "waiting for RMA" status.
[0051] RMA 120 enables state synchronization between the RMA portal and DM 114. To support timely RMA, RMA can be configured to begin querying attribute fields (e.g., machine attributes) to identify machines in a "degraded" state. Additionally, timely RMA machine attributes can persist up to the RMA error description, as these are essentially hardware errors. The actions taken by the RMA service depend on the machine's RMA and DM states.
[0052] Furthermore, as mentioned above, as an example, if a machine has basic disk health requirements, the infrastructure management system can use PS 118 (e.g., PsAgent) to implement an agent service to determine whether the basic health requirements are met. If so, PS 118 will complete the provisioning with fewer hard drives and allow the machine to run in a degraded state. For each degraded machine, PsAgent sets machine attributes in the DM to indicate how many disks are missing and which disks are faulty. WD 112 monitors the number of disks and the required volumes. WD 112 can be updated to skip the verification of unused disks, i.e., retain the disks in the RMA machine attributes as appropriate.
[0053] Now go to Figure 4 This document provides a flowchart illustrating a method for implementing the functionality of an infrastructure management system for hardware fault repair. It begins, at box 410, by determining that a hardware component failure has occurred. The hardware component is part of a hardware assembly. At box 420, a repair operation is initiated to operate the hardware assembly in a degraded state. A degraded state includes the hardware assembly operating without the failed hardware component. At box 430, repair attributes of the hardware properties are accessed. These repair attributes indicate the minimum operational limits of the hardware assembly. The configuration mode includes multiple attributes for defining the configuration file for the corresponding hardware assembly. These attributes include repair attributes indicating the minimum operational limits of the hardware assembly. The health model is a representation of the calculated conditions of the hardware assembly. The minimum operational limits are defined based on health metrics or optional and required components associated with the hardware assembly.
[0054] In box 440, it is determined that the minimum operational requirements for the hardware assembly to operate without any faulty hardware components are met. In box 450, operation of the hardware assembly in a degraded state is initiated. A degraded state includes the hardware assembly operating without any faulty hardware components. The hardware manager associated with the hardware assembly is pre-configured with a degraded state configuration in anticipation of a degraded state for operating the hardware assembly. The degraded state configuration includes instructions for operating the hardware assembly in a degraded state.
[0055] Now go to Figure 5This document provides a flowchart illustrating a method for implementing the functionality of an infrastructure management system for hardware fault repair. Initially, at box 510, a degraded state configuration is configured for the hardware infrastructure in anticipation of a degraded state for operation. The degraded state configuration includes instructions for operating the hardware infrastructure in a degraded state. At box 520, it is determined that a hardware component failure has occurred. The hardware component is included in a hardware assembly of the hardware infrastructure. At box 530, repair attributes are accessed. The repair attributes indicate the minimum operational limits of the hardware assembly. At box 540, it is determined that the minimum operational limits of the hardware assembly are met for operation without the failed hardware component. At box 550, operation of the hardware assembly in a degraded state is initiated. A degraded state includes operation of the hardware assembly without the failed hardware component. At box 560, the operation is performed using the hardware assembly in the hardware infrastructure. The operation is performed based at least in part on the degraded state configuration.
[0056] Referring to an infrastructure management system, the embodiments described herein allow for hardware fault repair. An infrastructure management system service platform component refers to an integrated component used to provide hardware fault repair. An integrated component refers to the hardware architecture and software framework that support data access functions using the infrastructure management system service platform. Hardware architecture refers to physical components and their interrelationships, while software framework refers to software that provides functionality that can be implemented using hardware devices running that software. An end-to-end software-based infrastructure management system service platform can operate within the infrastructure management system service platform component to operate computer hardware to provide infrastructure management system service platform functionality. Therefore, the infrastructure management system service platform component can manage resources and provide services for infrastructure management system functionality. Any other variations and combinations of embodiments of the invention are contemplated.
[0057] As an example, an infrastructure management system service platform may include an API library containing specifications for routines, data structures, object classes, and variables that can support interaction between the device's hardware architecture and the infrastructure management system service platform's software framework. These APIs include configuration specifications for the infrastructure management system service platform, enabling driver components and other components within the system to communicate with each other within the infrastructure management system service platform, as described herein.
[0058] After briefly describing an overview of embodiments of the invention, exemplary operating environments in which embodiments of the invention may be implemented are described below to provide a general context for various aspects of the invention. In particular, reference is made first to… Figure 6This illustrates an exemplary operating environment for implementing embodiments of the invention, generally designated as computing device 600. Computing device 600 is merely one example of a suitable computing environment and is not intended to impose any limitation on the scope of the invention's use or functionality. Computing device 600 should also not be construed as having any dependency or requirement associated with any one or combination of the illustrated components.
[0059] This invention can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions (such as program modules) that are executed by a computer or other machine (such as a personal data assistant or other handheld device). Typically, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs a specific task or implements a specific abstract data type. This invention can be implemented in various system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This invention can also be implemented in distributed computing environments, where tasks are performed by remote processing devices linked via a communication network.
[0060] refer to Figure 6 The computing device 600 includes a bus 610 that directly or indirectly couples to the following devices: memory 612, one or more processors 614, one or more presentation components 616, input / output ports 618, input / output components 620, and an illustrative power supply 622. Bus 610 can represent one or more buses (such as an address bus, a data bus, or a combination thereof). Although Figure 6 The various block diagrams are shown with lines for clarity, but in reality, the depiction of the various components is not so clear, and in representation, the lines are more accurately gray and blurred. For example, presentation components such as display devices can be considered I / O components. Additionally, the processor has memory. It is recognized that this is the nature of the art, and it is reiterated... Figure 6 The figures are merely illustrative of exemplary computing devices that can be used in conjunction with one or more embodiments of the present invention. No distinction is made between categories such as “workstation,” “server,” “laptop,” and “handheld device,” as they all refer to the same thing. Figure 6 Within that range, and with reference to “computing device”, it can be expected.
[0061] Computing device 600 typically includes a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 600, and includes volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.
[0062] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by the computing device 700. Computer storage media does not include the signal itself.
[0063] Communication media typically implement computer-readable instructions, data structures, program modules, or other data using modulated data signals (such as carrier waves or other transmission mechanisms), and include any information transmission medium. The term "modulated data signal" refers to a signal in which one or more characteristics of the signal can be set or altered in a manner that allows information to be encoded within the signal. By way of example and not limitation, communication media include wired media (such as wired networks or direct wired connections) and wireless media (such as acoustic, RF, infrared, and other wireless media). Any combination of the above should also be included within the scope of computer-readable media.
[0064] Memory 612 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Computing device 600 includes one or more processors that read data from various entities such as memory 612 or I / O components 620. Presentation component 616 presents data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc.
[0065] I / O port 618 allows computing device 600 to be logically coupled to other devices, including I / O components 620, some of which may be built into it. Exemplary components include microphones, joysticks, game controllers, satellite antennas, scanners, printers, wireless devices, etc.
[0066] Now for reference Figure 7 , Figure 7 An exemplary distributed computing environment 700 in which an implementation of this disclosure can be employed is shown. Specifically, Figure 7A high-level architecture of the infrastructure management system (“System”) in the cloud computing platform 710 is illustrated, which supports seamless modification of software components. It should be understood that this and other arrangements described herein are illustrated by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, and functional groupings, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or combined with other components and implemented in any suitable combination and location. The various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory.
[0067] The data center can support a distributed computing environment 700 (e.g., centralized infrastructure and edge infrastructure) including a cloud computing platform 710, racks 720, and nodes 730 (e.g., computing devices, processing units, or blades) within the racks 720. The system can be implemented using a cloud computing platform 710 that runs cloud services across different data centers and geographical regions. The cloud computing platform 710 can implement a structure controller 740 component for provisioning and managing the allocation, deployment, upgrades, and management of cloud services. Typically, the cloud computing platform 710 is used to store data or run service applications in a distributed manner. The cloud computing infrastructure 710 in the data center can be configured to host and support the operation of endpoints for specific service applications. The cloud computing infrastructure 710 can be a public cloud, a private cloud, or a dedicated cloud.
[0068] A host 750 (e.g., an operating system or runtime environment) that runs a defined software stack on node 130 can be supplied to node 730. Node 730 can also be configured to perform specialized functions (e.g., compute nodes or storage nodes) within the cloud computing platform 710. Node 730 is allocated to run one or more portions of a tenant's service application. A tenant can refer to a customer that utilizes the resources of the cloud computing platform 710. The service application components of the cloud computing platform 710 that support a particular tenant can be referred to as tenant infrastructure or leases. The terms service application, application, or service are used interchangeably herein and broadly refer to any software or portions of software that run on top of or access storage and compute equipment locations within a data center.
[0069] When node 730 is supporting more than one individual service application, the node can be partitioned into virtual machines (e.g., virtual machine 752 and virtual machine 754). Physical machines can also run individual service applications simultaneously. Virtual machines or physical machines can be configured as personalized computing environments supported by resources 760 (e.g., hardware and software resources) in the cloud computing platform 710. It is conceivable that resources can be configured for specific service applications. Furthermore, each service application can be divided into functional parts, allowing each functional part to operate on a separate virtual machine. In the cloud computing platform 710, multiple servers can be used to run service applications and perform data storage operations in a cluster. In particular, servers can perform data operations independently but are exposed as a single device referred to as a cluster. Each server in the cluster can be implemented as a node.
[0070] Client device 180 can link to service applications in cloud computing platform 710. Client device 780 can be any type of computing device, which may correspond to, for example, the reference... Figure 7 The computing device 700 is described. A client device 780 can be configured to issue commands to a cloud computing platform 710. In embodiments, the client device 780 can communicate with the service application via Virtual Internet Protocol (IP) and a load balancer or other means that initiates communication requests to a designated endpoint in the cloud computing platform 710. Components of the cloud computing platform 710 can communicate with each other via a network (not shown), which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs).
[0071] Various aspects of the distributed computing environment 700 and the cloud computing platform 710 have been described. Note that the desired functionality within the scope of this disclosure can be achieved using any number of components. Although for clarity, Figure 7 The various components are shown with lines, but in reality, the depiction of the various components is not clear, and in terms of representation, the lines may be more accurately described as gray or blurred. Furthermore, although... Figure 7 Some components are described as single components, but the descriptions are exemplary in nature and number and are not to be construed as limiting all implementations of this disclosure.
[0072] The embodiments described in the preceding paragraphs can be combined with one or more of the specifically described alternatives. In particular, in the alternatives, the claimed embodiments may include references to more than one other embodiment. The claimed embodiments may specify further limitations on the claimed subject matter.
[0073] The subject matter of embodiments of the present invention has been described herein as specific to meet legal requirements. However, the description itself is not intended to limit the scope of this patent. Rather, the inventors have envisioned that the claimed subject matter may be implemented in other ways, incorporating other current or future techniques to include steps different from or similar to those described herein. Furthermore, although the terms “step” and / or “box” may be used herein to imply different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and if the order of the individual steps is explicitly described.
[0074] For the purposes of this disclosure, the word "comprising" has the same broad meaning as the word "including," and the word "access" includes "receiving," "quoting," or "retrieval." Additionally, unless otherwise indicated, words such as "a" and "one" include both plural and singular forms. Thus, for example, the constraint of "feature" is satisfied where one or more features are present. Furthermore, the term "or" includes conjunctions, adversative conjunctions, and both (therefore, a or b includes a or b, and a and b).
[0075] For the purposes of the detailed discussion above, embodiments of this disclosure are described with reference to a distributed computing infrastructure having an infrastructure management system; however, the infrastructure management system depicted herein is merely exemplary. Components may be configured to perform novel aspects of the embodiments, wherein being configured to include being programmed to perform a particular task or to implement a particular abstract data type using code. Furthermore, while embodiments of the invention may generally refer to the infrastructure management system and schematic diagrams described herein, it should be understood that the described techniques can be extended to the context of other implementations.
[0076] Embodiments of the invention have been described with respect to specific examples, but these specific examples are illustrative in all respects and not restrictive. Alternative embodiments will be apparent to those skilled in the art to which this invention pertains without departing from the scope of the invention.
[0077] As can be seen from the foregoing, the present invention is well suited to achieve all the objectives and purposes set forth above, and such advantages are clear and inherent to the structure.
[0078] It should be understood that certain features and sub-combinations are useful and can be used without reference to other features or sub-combinations. This is covered by and within the scope of the claims.
Claims
1. A system for implementing an infrastructure management system that supports hardware fault repair, the system comprising: The infrastructure management component is configured as follows: It is determined that a hardware component failure has occurred, and the hardware component is included in the hardware assembly; Initiate a repair operation for operating the hardware assembly in a degraded state, wherein the degraded state includes the hardware assembly operating without any faulty hardware components. Access the repair properties of the hardware assembly, wherein the repair properties indicate the minimum operating limits of the hardware assembly, wherein multiple different types of hardware assemblies are configured with corresponding minimum operating limits for the hardware components in the multiple different types of assemblies. Determine that the hardware assembly operates in the absence of any faulty hardware components, satisfying the minimum operational constraints of the hardware assembly. as well as Initiate operation of the hardware assembly in the degraded state, wherein the degraded state includes operation of the hardware assembly without any faulty hardware components. The minimum operating limit mentioned therein is dynamic; and The minimum operational constraint is quantified by a health metric defined in a health model, which also defines the minimum operational constraint based on the health metric.
2. The system of claim 1, wherein the configuration mode includes a plurality of attributes for defining a configuration file for a corresponding hardware assembly, the plurality of attributes including the repair attribute, the repair attribute indicating the minimum operational limit of the health model from the hardware assembly, wherein the health model is a representation of the computational conditions of the hardware assembly.
3. The system of claim 1, wherein the minimum operational limitation is defined based on a health metric associated with the hardware assembly or optional and required components associated with the hardware assembly.
4. The system of claim 1, wherein, in the event that a degraded state for operating the hardware assembly is anticipated, a hardware manager associated with the hardware assembly is pre-configured with a degraded state configuration, wherein the degraded state configuration includes instructions for operating the hardware assembly in the degraded state.
5. The system according to claim 1, further comprising: The Data Center Manager component is configured as follows: Provides access to health status information and configuration files of a hardware assembly, wherein the health status information includes health status information of each healthy and unhealthy hardware component of the hardware assembly. The monitor component is configured as follows: Access the health status information of the hardware assembly; Selectively monitoring hardware components of a hardware assembly, wherein the health status information indicates that the hardware components are healthy; and Report a fault in the hardware assembly, wherein at least one fault is based on a health SLA failure of the hardware assembly.
6. The system according to claim 5, further comprising: The service provisioning component is configured as follows: Based on the health status information and configuration file corresponding to the hardware assembly, a repair operation is performed on the hardware assembly in the degraded state. The health status information and configuration file are obtained from the data center manager component. The repair operation includes verifying the health status information of the hardware assembly.
7. The system of claim 6, wherein the supply service component is configured for: When the minimum operational limit is not met for the first tenant with the first SLA Disconnect the hardware assembly, where the first SLA is a factor in the minimum operational constraints; Identify a second tenant with a second SLA, wherein the minimum operating constraint is satisfied for the second tenant with the second SLA; as well as For the second tenant, a repair operation is performed on the hardware assembly.
8. The system according to claim 7, further comprising: The return authorization component is configured as follows: The timely RMA operation is performed at least in part based on the attribute fields of the hardware components in the hardware assembly, where the attribute fields indicate the corresponding hardware group of the hardware assembly. Health status information of the item.
9. A computer-implemented method for implementing an infrastructure management system, the method comprising: It is determined that a hardware component failure has occurred, and the hardware component is included in the hardware assembly; Access the repair properties of the hardware assembly, wherein the repair properties indicate the minimum operational limitations of the hardware assembly; Based on accessing the repair attributes, it is determined that the hardware assembly meets the minimum operational requirements for operation without any faulty hardware components. as well as Initiate operation of the hardware assembly in a degraded state, wherein the degraded state includes operation of the hardware assembly without the hardware component. The minimum operating limit mentioned therein is dynamic; and The minimum operational constraint is quantified by a health metric defined in a health model, which also defines the minimum operational constraint based on the health metric.
10. The method of claim 9, wherein the minimum operational constraints are defined based on: health metrics or optional and required components associated with a stock unit (SKU) of the hardware assembly and a service level agreement (SLA) associated with the hardware assembly.
11. The method of claim 9, wherein, in the event that a degraded state for operating the hardware assembly is anticipated, the hardware assembly is pre-configured with a degraded state configuration, wherein the degraded state configuration includes instructions for operating the hardware in the degraded state.
12. The method of claim 9, wherein when it is determined that a failure of the hardware component has occurred, a repair operation is initiated on the hardware assembly to operate in the degraded state based on health status information and a configuration file corresponding to the hardware assembly, wherein the repair operation includes verifying the health status information of the hardware assembly.
13. The method of claim 9, wherein initiating the operation of the hardware assembly in the degraded state further comprises: It has been determined that multiple hardware components of the hardware assembly have failed. A subset of the multiple hardware components to be repaired is determined based on the minimum operational constraints of the hardware assembly. Repair the subset of hardware components; as well as Perform a repair operation on the hardware assembly.
14. The method of claim 9, further comprising: When the minimum operational limit is not met for a first tenant with a first SLA, the hardware assembly is deactivated, where the first SLA is a factor in the minimum operational limit; Identify a second tenant with a second SLA, wherein the minimum operating constraint is satisfied for the second tenant with the second SLA; as well as For the second tenant, a repair operation is performed on the hardware assembly.
15. The method of claim 9, further comprising: Receive an instruction to initiate a Return Authorization (RMA) operation to repair the hardware assembly, wherein the instruction is partly based on the SLA requirements of the tenant associated with the hardware assembly.
16. A computer storage device having computer-executable instructions embodied thereon, which, when executed by one or more processors, cause the one or more processors to perform a method for implementing an infrastructure management system for hardware fault repair, the method comprising: In the event that a degraded state is anticipated for operating the hardware infrastructure, a degraded state configuration is configured for the hardware infrastructure, wherein the degraded state configuration includes instructions for operating the hardware infrastructure in the degraded state. It is determined that a hardware component failure has occurred, said hardware component being included in the hardware assembly of said hardware infrastructure; Access the repair properties of the hardware assembly, wherein the repair properties indicate the minimum operational limitations of the hardware assembly; Determine that the hardware assembly operates in the absence of any faulty hardware components, satisfying the minimum operational constraints of the hardware assembly. Initiating operation of the hardware assembly in the degraded state, wherein the degraded state includes the hardware assembly operating without the hardware component; and The operation is performed using the hardware components in the hardware infrastructure, wherein the operation is performed at least in part based on the degraded state configuration. The minimum operating limit mentioned therein is dynamic; and The minimum operational constraint is quantified by a health metric defined in a health model, which also defines the minimum operational constraint based on the health metric.
17. The device of claim 16, wherein, in the event of anticipating the degraded state for operating the hardware assembly, the hardware infrastructure is pre-configured with a degraded state configuration, wherein the degraded state configuration includes instructions for operating the hardware in the degraded state.
18. The device of claim 16, wherein the configuration mode-based configuration file includes a degraded state configuration accessed during a repair operation to configure the hardware assembly to operate in the degraded state.
19. The device of claim 16, wherein the degraded state configuration includes an attribute field for a corresponding hardware component of the hardware assembly, the attribute field indicating health status information for the corresponding hardware component of the hardware assembly, wherein the timely RMA operation is performed at least in part based on the attribute field of the hardware component in the hardware assembly.
20. The device of claim 16, wherein the minimum operational limitations are defined based on: health metrics, optional and required components, and service level agreements associated with the hardware assembly.