Verification and Prediction of Cloud Readiness
By simulating user workloads and failure injection in the cloud infrastructure, comparing sample performance before and after updates, it solves the problem that cloud service providers have difficulty verifying and predicting the impact of updates, and realizes effective prediction and update adaptability testing of user experience, improving the stability and user satisfaction of cloud services.
Patent Information
- Application Number
- CN202080089002.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-27
- Filing Date
- 2020-11-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-11-17
AI Technical Summary
Cloud service providers have difficulty effectively verifying and predicting the impact of hardware or software updates on user experience, especially in complex and evolving cloud infrastructure, where traditional methods have difficulty simulating diverse user workloads and scenarios, resulting in the possibility of introducing failures and disruptions in updates.
By simulating user workloads in cloud infrastructure, using statistical methods and fault injection techniques, comparing sample performance before and after updates, collecting telemetry data and analyzing, predicting the impact of updates on user experience.
It realizes sufficient testing of updates in actual user production workloads, predicts and verifys whether updates are suitable for deployment, reduces failures and interrupts introduced by updates, and improves the reliability and user experience of cloud services.
Smart Images

Figure CN114902192B_ABST
Abstract
Description
Background Art
[0001] Computer hardware such as a processing unit is fabricated on a silicon chip. Specifically, multiple transistors can be etched on a silicon wafer to implement a specific set of components. These components include logic gates, registers, arithmetic logic units, and so on. The specific configuration and interconnection of these components can be according to an instruction set architecture. Errors may be found in the hardware, such as defects, flaws, or vulnerabilities. Some of these errors may be hardware errors irreversibly etched on the silicon of the hardware. Other errors may originate from the software used to operate the hardware. When these errors are found, code can be developed and deployed to mitigate or eliminate these errors. Code can also be deployed to improve the functionality of the computer hardware. Summary of the Invention
[0002] To provide a basic understanding of some aspects described herein, a brief summary of the innovative subject matter is given below. This summary of the invention is not an extensive overview of the claimed subject matter. Its purpose is neither to identify key or important elements of the claimed subject matter nor to delineate the scope of the innovative subject matter. Its sole purpose is to present some concepts of the claimed subject matter in a brief form as a prelude to the more detailed description that is presented later.
[0003] An embodiment provides a method for validating and predicting cloud readiness. The method includes identifying a sample of components from a cloud infrastructure, where updates are applied to the sample to generate a processed sample, and the processed sample has a statistically sufficient size and relevant cloud-level diversity, and identifying a control sample of components from the cloud infrastructure, where the control sample is statistically comparable to the processed sample. The method also includes executing a workload set on the processed sample and the control sample. Additionally, the method includes predicting the impact of the updates on the user experience based on a comparison of telemetry captured during the execution of the workload set on the processed sample and the control sample.
[0004] Another embodiment provides a method. The method includes uploading an update and defining the scope of validation and prediction determined via test cases applied to the hardware update, where the scope includes determining the workload type, the allowed time, and the number of concurrent components during the execution of the test cases. The method also includes monitoring the execution of the test cases, where the test cases are executed at cloud-level scale and cloud-level diversity, and obtaining telemetry from one or more components under test.
[0005] In addition, another embodiment provides one or more computer-readable storage media for storing computer-readable instructions. The computer-readable instructions perform a method for validating and predicting cloud readiness. The method includes identifying a sample of components from a cloud infrastructure, where updates are applied to the sample to generate a processed sample, and the processed sample has a statistically sufficient size and relevant cloud-level diversity, and identifying a control sample of components from the cloud infrastructure, where the control sample is statistically comparable to the processed sample. The method further includes executing a workload set on the processed sample and the control sample. In addition, the method includes predicting the impact of the updates on the user experience based on a comparison of telemetry captured during the execution of the workload set on the processed sample and the control sample.
[0006] The following description and the accompanying drawings set forth in detail certain illustrative aspects of the claimed subject matter. However, these aspects merely represent several of the various ways in which the innovative principles may be employed, and the claimed subject matter is intended to include all such aspects and their equivalents. Other advantages and novel features of the claimed subject matter will become apparent from the following detailed description of this innovation when considered in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 is a block diagram of a cloud readiness standard and validation environment;
[0008] Figure 2 is a process flow diagram of a method for validating and predicting cloud readiness for hardware updates;
[0009] Figure 3 is a process flow diagram of a method for validating and predicting cloud readiness for firmware updates;
[0010] Figure 4 is a process flow diagram of a method for performing a statistical analysis of results and telemetry data captured during testing of a control sample and a processed sample; and
[0011] Figure 5 is a block diagram of an exemplary computing device configured to validate and predict cloud readiness in accordance with aspects of the disclosed subject matter. DETAILED DESCRIPTION
[0012] A cloud infrastructure can be defined as a combination of hardware and software used to implement cloud-based services. The hardware can include servers, racks, network switches, routers, quantum computers, storage devices, power supply units (PSUs), etc. The software can include firmware, operating systems, etc. Cloud computing services, or simply "the cloud", can refer to a network based on a cloud infrastructure that can deliver various computing services. The components used to support the cloud can be designed for multiple applications, such as storing and managing data, running applications, or delivering content or services, such as streaming video, webmail, office productivity software, or social media. Cloud service providers operate and maintain the network and associated services and applications, which can have components located all over the world that are communicatively coupled together and operate as a single ecosystem. Users can access cloud-enabled services according to a predefined protocol with the cloud service provider. Thus, users can be clients of the cloud service provider.
[0013] Cloud service providers are able to support a number of users and provide access to the cloud infrastructure. Each user's data can be isolated and remain invisible to other users. Cloud service providers can manage the cloud infrastructure by, for example, installing new hardware, replacing hardware, repairing hardware, installing new software, updating software, etc. Provisioning hardware and software for each user enables the cloud service provider to access each user's cloud usage data. Cloud usage data can be, for example, the type of workload, the amount of workload, expected response, resource consumption distribution across users and seasons, user interest in various stock keeping units (SKUs) of hardware, software, and services.
[0014] Therefore, the task of cloud service providers is to provide state-of-the-art services and maintain an ever-evolving cloud. In fact, the cloud is undergoing continuous changes at all levels. First, the workloads on the cloud, cloud management, cloud control structures, and software infrastructure are evolving. For example, hypervisors, OSs, and device drivers are constantly changing. In addition, cloud hardware and low-level infrastructure are changing and evolving. Although cloud service providers manage and maintain the evolving cloud, their hardware and related software can be provided to the cloud service provider by hardware vendors. Errors in the hardware and software underlying the cloud can cause the hardware or software to malfunction during operation, resulting in an interruption of the cloud services provided to users. Such interruptions can lead to a regression in the user experience.
[0015] Functional regression in cloud infrastructure can be the result of introducing bad hardware or software updates to the cloud infrastructure. For example, components of cloud infrastructure such as nodes, servers, hard disk drives, modems, switches, routers, racks, power units, rack managers including controllers, firmware, and sensors, multiple levels of network components, storage devices, cooling infrastructure, or high-voltage infrastructure may suffer from defects such as vulnerabilities or bugs. For example, the regression can be a problem introduced by a recently deployed update to the cloud hardware or software. In some cases, the regression indicates a problem or failure observed by a user during the execution of a cloud service for that user. Typically, the user has a protocol with the cloud service provider for a specific service level. Hardware errors can cause the cloud service provider to violate that protocol with the user. Generally, errors in the hardware and the resulting hardware failures can only be resolved by the hardware vendor. In response to discovering a hardware error or a hardware-based cloud failure, the hardware vendor can provide a software or hardware update to update, fix, or improve the hardware functionality.
[0016] Typically, an update is a modification to the cloud infrastructure. For example, code can be developed to update, fix, or improve hardware or software functionality. In an embodiment, the code can be referred to as an update. In an example, the update can be microcode, where microcode is a set of instructions that implement the configuration or reconfiguration of hardware. In an embodiment, microcode can be a programmable instruction layer that serves as an intermediary between the hardware and the instruction set architecture associated with the hardware. Similarly, an update can also be firmware that implements the basic functionality of the hardware. For example, the firmware can provide instructions that define how components such as video cards, keyboards, and mice communicate and perform certain functions. An update can also be newly installed hardware, replaced hardware, or repaired hardware.
[0017] When an error or defect is discovered in the cloud infrastructure, an update can be created to mitigate or eliminate the discovered error or defect. Even if an update that eliminates the error or defect is provided, the cloud service provider may still violate the agreement with the user due to the error or defect. In the case of a hardware error, the hardware vendor can provide an update to eliminate the error or defect. However, the hardware vendor may lack the knowledge and resources to fully test the hardware error or defect because the hardware vendor does not know the workloads, scenarios, and other aspects to be applied to the hardware. Additionally, the update itself may introduce a regression, and the update vendor generally cannot predict the impact of the update on the user experience. Moreover, it is generally challenging to verify whether the update is indeed suitable for deployment in the cloud infrastructure.
[0018] This technique enables the verification and prediction of cloud readiness. As used herein, updated cloud readiness indicates the suitability of the update for use in a cloud infrastructure. When evaluating the updated cloud readiness, this technique predicts the impact of the update on the user experience. In an embodiment, this prediction can occur by simulating or emulating user workloads. The prediction of the impact on the user experience can also occur using actual user workloads. When the predicted impact of the update on the user experience is unnoticeable to the user, the update can be verified. As used herein, when the update negatively changes the user's operation or access to the cloud infrastructure, the predicted impact of the update can be noticeable.
[0019] As described herein, an update can be a physical hardware update where the hardware is replaced, repaired, or placed in a specific configuration. An update can also be a software update where code is provided to eliminate or mitigate software errors. Examples of software updates include microcode (uCode) changes to a central processing unit (CPU), firmware and microcode changes to a graphics processing unit (GPU), firmware changes to a network chipset, basic input / output system (BIOS) changes, field-programmable gate array (FPGA) reprogramming, hybrid hard disk drive (HHD) firmware, solid-state drive (SSD) firmware, host operating system, network interface card (NIC), etc. Software updates can also include device-specific code such as device drivers.
[0020] For ease of description, this technique is described as evaluating and predicting the cloud readiness of updates to "components". The term "component" can refer to any electronic device, including sub-components that are commonly considered components. The term component can also refer to software. For example, this technique can verify and predict the readiness of a software update to a device driver as a component of a cloud infrastructure. This technique can also verify and predict the cloud readiness of a firmware update to a graphics processing unit (GPU) as a component of a component such as a graphics card installed in a Peripheral Component Interconnect Express (PCIe). In another example, this technique can verify and predict the readiness of a firmware update to a power supply unit (PSU) as a component that powers a component rack frequently used in a data center. Each rack can include components such as nodes, servers, hard disk drives, modems, switches, routers, and other electronic devices. In this example, samples can be identified for statistical analysis described below, but the specific update being tested is applied to components that indirectly affect the samples. Additional examples of components with applied updates for which their corresponding cloud readiness can be verified and predicted include updates to network infrastructure, cooling infrastructure, or storage services.
[0021] Ultimately, cloud readiness is defined by the user experience and the cost associated with providing a satisfactory user experience. For example, an update that results in a large number of virtual machine (VM) outages is not cloud ready. Similarly, a cloud service provider spending a lot on redundancy to hide such outages from the user may also mean that the underlying update is not cloud ready. Thus, in an embodiment, when the problems (outages, downtime, flickers, etc.) noticed by the user and their root cause are below a required threshold, and the overhead required to meet this threshold is included in terms of cost, effort, etc., the update is cloud ready. However, what the user notices can depend on the user and their corresponding workload. Also, what is noticed can change over time. Additionally, a specific threshold is not predefined. Instead, the threshold can be a function of user expectations, contracts, and actionable SLAs, which tend to evolve over time.
[0022] Accordingly, the cloud readiness criteria (CRC) described herein validate an update and predict whether the update is cloud ready. In an embodiment, the cloud readiness criteria can include testing, experimentation, statistical analysis, and prediction. In an embodiment, the cloud readiness criteria are defined by the measurement of problems noticed by the user caused by the update and the cost associated with keeping that measurement below a specific threshold. In an embodiment, the measurement of problems noticed by the user caused by the update and the cost associated with keeping that measurement below a specific threshold enable prediction of the user impact associated with a hardware update. In this way, the present technology enables validation and prediction of whether a hardware update meets the cloud readiness criteria as defined by the then-current set of tests, user expectations, contractual obligations, and the cloud management structure and infrastructure in fact.
[0023] As a preliminary matter, some of the figures depict concepts in the context of one or more structural components, which are variously referred to as functions, modules, features, elements, etc. The various components shown in the figures can be implemented in any manner, such as via software, hardware (e.g., discrete logic components), firmware, or any combination thereof. In some embodiments, the various components can reflect the use of the corresponding components in an actual implementation. In other embodiments, any single component shown in the figures can be implemented by multiple actual components. The description of any two or more separate components in the figures can reflect different functions performed by a single actual component. Details are provided below Figure 1 regarding a system that can be used to implement the functions shown in the figures.
[0024] Other drawings depict concepts in flowchart form. In this form, certain operations are described as constituting different boxes that are performed in a particular order. Such an implementation is exemplary and non-limiting. Certain boxes described herein can be combined together and performed in a single operation, certain boxes can be broken into multiple component boxes, and certain boxes can be performed in a different order than shown herein, including in parallel fashion of performing the boxes. The boxes shown in the flowchart can be implemented by software, hardware, firmware, manual processing, etc. As used herein, hardware can include a computer system, discrete logic components such as an application specific integrated circuit (ASIC), etc.
[0025] Regarding terms, the phrase "configured to" includes any manner in which any type of functionality can be constructed to perform the identified operation. The functionality can be configured to perform the operation using, for example, software, hardware, firmware, etc.
[0026] The term "logic" includes any functionality for performing a task. For example, each operation shown in the flowchart corresponds to the logic for performing that operation. The operation can be performed using, for example, software, hardware, firmware, etc.
[0027] As used herein, the terms "component", "system", "client", "server", etc. are intended to refer to a computer-related entity, whether hardware, software (e.g., in execution), or firmware or any combination thereof. For example, a component can be a process running on a processor, an object, an executable, a program, a function, a library, a subroutine, a computer, or a combination of software and hardware.
[0028] As used herein, the term "hardware vendor" can refer to any vendor of hardware components. In an example, the hardware components of a cloud infrastructure can be obtained from a cloud service provider. In other examples, the hardware components of a cloud infrastructure can be obtained from a third party other than the cloud service provider.
[0029] By way of illustration, both an application running on a server and the server can be components. One or more components can reside within a process, and a component can be located on one computer and / or distributed between two or more computers. The term "processor" is generally understood to refer to a hardware component, such as a processing unit of a computer system.
[0030] Furthermore, the claimed subject matter can be implemented as using standard programming and / or engineering techniques to produce software, firmware, hardware, or any combination thereof to control a computer to implement the disclosed subject matter's methods, apparatuses, or articles of manufacture. The term "article of manufacture" as used herein is intended to include a computer program accessible from any computer-readable storage device or medium.
[0031] Computer-readable storage media can include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, and magnetic strips), optical disks (e.g., compact disks (CDs) and digital versatile disks (DVDs)), smart cards, and flash memory devices (e.g., cards, sticks, and key drives). In contrast, computer-readable media (i.e., non-storage media) generally can additionally include communication media, such as transmission media for wireless signals, etc.
[0032] Given the frequent and complex updates applied to cloud infrastructures, the verification and prediction of cloud-readiness described herein enable verification of updates. For example, microcode updates, BIOS changes, and network card firmware may be updated more frequently than in the past. Additionally, the complexity of hardware updates is generally increasing, both with respect to the updates themselves and the interaction of the updates with other aspects of the networked hardware and infrastructure. Moreover, it has traditionally been difficult to verify updates for cloud-readiness because the user experience in the cloud is constantly evolving, including cloud infrastructure and management, the hardware used in the cloud, the user base, the workloads, and user expectations. In other words, it has traditionally been challenging to mimic the cloud because its actual specifications are constantly changing. Further, the complexity of the cloud is changing and generally increasing. For example, the integration between hardware and software is becoming more complex as concepts such as threat mitigation, optimization, and secure boot are applied to that integration. Cloud-hosted workloads are also diverse, which involves an increasing number of little-known areas in the underlying infrastructure (e.g., leading to the exposure of fairly complex CPU vulnerabilities). This diversity is also evolving regularly. In other words, the diversity of cloud-hosted workloads may expose hardware errors that would typically not be discovered during typical standard operations. Finally, due to the large number of workloads and test cases, testing cloud-readiness requires scaling even without the dynamic evolution of the cloud as described above.
[0033] Generally speaking, it is not feasible for hardware vendors to consider all possibilities, such as frequent and complex updates and a changing, diverse, and scalable cloud. Additionally, the statistical manifestations of problems and the quirky behavior of some errors are only exposed when the necessary scale and diversity of updates are applied during test cases. Moreover, errors can be elusive and require specialized tools and processes to detect. Ultimately, the cloud service provider is responsible for any errors observed by the user. The cloud service provider provides a buffer between the hardware vendor and the user. Thus, the originator of the update is excluded from issues attributed to hardware errors and is not responsible for the resulting damage.
[0034] Figure 1It is a block diagram of a cloud readiness standard environment 100. The environment 100 can be used to test updates to cloud infrastructure in a networked environment that supports or replicates cloud computing services. Computing services include, but are not limited to, devices for storing and managing data, executing applications, or delivering content, storage devices, networking, and software, or services such as streaming video, webmail, office productivity software, or social media. Once a failure occurs, cloud computing services may be interrupted, and the user experience may regress. The environment 100 enables sufficient testing of updates before deployment in actual user production workloads. Specifically, this technology enables prediction of the impact of updates on the user experience.
[0035] As Figure 1 shown, the partner team 102 can access the authoring system 104. The partner team 104 can represent a managed service provider with technical expertise in developing test cases. In an embodiment, the partner team can author test cases via the authoring system 104. The authoring system implements tasks such as providing virtual machines (VMs) for cloud services in test cases. Specifically, the authoring system 104 enables the partner team to develop test logic and check and adjust test parameters according to the partner team's expectations. As shown, the authoring system 104 includes a command-line input (CLI) and a portal. The authoring system 104 can be accessed via the command-line input and the portal. In an embodiment, scripts defining tasks associated with test cases can be input into the system 100 via the CLI and the portal.
[0036] Test cases can be executed on full-stack production blades to generate the most realistic cloud conditions. Additionally, relevant samples can be represented by deploying sufficient blades that represent relevant parts of the cloud infrastructure for statistically significant results. In an embodiment, test cases can be run in a sandbox or isolated test environment in production to protect user VMs. User business can be simulated, and synthetic workloads can be utilized to represent user usage scenarios. To simulate production failures, faults can be injected to evaluate the impact of the payload on resilience, enabling fast and consistent results. Injecting errors when executing a workload set on processed samples and control samples may expose errors in the updated or processed samples. In an embodiment, to reduce false positives, tests can be authored such that the only variable is the payload being tested.
[0037] The script obtained via the authoring system 104 enables the execution application programming interface (API) 106 to call a process for test cases that can be executed on multiple cloud infrastructure configurations for supporting cloud services. In an embodiment, the execution API 106 can take the script from the authoring system 104 as input. In an embodiment, the execution API 106 can also take continuous test data from the continuous integration continuous deployment (CICD) / integration system block 108 regarding a specific build of the hardware / software to be tested as input. The CICD / integration system block 108 implements continuous changes to updates during testing and the delivery of these updates. In an embodiment, the execution API 106 can be configured to deploy, configure, and manage test cases applied to the hardware configuration for supporting cloud networks.
[0038] The execution engine 110 can coordinate a specific workload to be applied to the test cases obtained from the execution API 106. Specifically, the execution engine 110 can arrange a sequence of specific tasks of the workflow and increase or slow down the tasks within the workflow as needed by the test cases. The execution engine 110 can take as input the scheduling script and updates for testing from the firmware repository 112.
[0039] The test cases and updates are sent to the test environment management system 114 and the cloud control plane (CCP) 116, and the CCP 116 can call specific stock keeping units (SKUs), virtual machines, etc. to execute the test cases in the isolation environment 118. In an embodiment, a specific configuration of a component can be referred to as an SKU. Each SKU can uniquely identify a product, service, or any combination thereof. In an embodiment, the SKU can indicate a specific arrangement of hardware that can be used by a cloud provider.
[0040] The isolation environment 118 implements protection for the workload, such as the test workload from the workload environment 120. Specifically, the isolation environment eliminates or limits the potential impact of possible spurious updates that are still under test. The isolation environment can vary according to its specific use. For example, a cloud service provider can choose to execute the payload 134 in a replicated environment that mimics the production cloud. However, the isolation environment can also be process isolation, where the VM executing the actual live production workload is used for testing by isolating the processes executed on the user VM. In this case, the tests for verification and prediction described herein include adding or injecting processes into the VMs that the user owns and uses in their live production workloads.
[0041] The isolated environment 118 can also be implemented by temporarily allocating nodes (components) of the cloud infrastructure for verification and prediction purposes. The inclusive CCP 116 controls specific components of the cloud infrastructure requests. In some isolated environments, only some virtual machines on the processed host nodes are allocated for verification and prediction related workloads. The isolated environment 118 can also be implemented using a more robust (and expensive) level of isolation by a dedicated entire rack (which has multiple host nodes), an entire subnet of compute nodes, etc.
[0042] As shown, faults from the fault injection library 122 are optionally injected into the isolated environment 118. In an embodiment, applying test cases to the A sample and the B sample as described below also includes fault injection. Fault injection generally represents edge cases and errors, i.e., those scenarios that are considered less likely or less credible. Deliberate fault injection accelerates the coverage of such edge cases and actually shortens the time required to expose related problems. The idea behind fault injection is to catalyze situations that rarely occur in natural circumstances. This achieves reducing the time required for evaluation updates and increasing the confidence in sufficient verification and prediction. Generally, a situation that causes an error may occur on one in a million node-days. Fault injection artificially increases the likelihood regarding coverage. Naturally, the same injection is applied qualitatively and quantitatively to both the A sample and the B sample, as described below.
[0043] As shown, the isolated environment 118 includes a host node 124 and a host node 126. In many cloud implementations, the host nodes 124 and 126 are used to host VMs provided to users. In other cases, the host nodes manage storage devices provided to user applications. Test cases can be applied to each of the host nodes 124 and the host node 126. The host node 124 includes virtual machines 128 and 130 used by the guest stack. Generally, the guest implements a complete compute stack, similar to a physical computer. The virtual machine 128 can include a guest agent (GA) 132. The virtual machine 130 can include a GA 133. The workload executed by the VM can be referred to as a guest. The present technology can include a guest agent having responsibilities analogous to those of a host agent. For example, the guest agent can collect telemetry from the guest operating system and processes executing on the guest agent. In an embodiment, the guest agent can be used to inject faults from the fault injection library 122.
[0044] Payload 134 can be applied to host node 124. A cloud-readiness standard host agent 136 resides on host node 124. Similarly, payload 138 can be applied to host node 126. A cloud-readiness standard host agent 140 resides on host node 126. The CRC host agents 136 and 140 collect various telemetry directly from their respective host nodes. This can be redundant when the telemetry collected by the control plane 116 is sufficient. However, when a host node is isolated, the CRC host agents 136 and 140 can have more access to detailed data (and thus, there is no concern about user data privacy and less concern about negatively impacting the overall performance of the host node due to intensive telemetry collection). In an embodiment, another purpose of the CRC host agents 136 and 140 is to control the host itself without constraints, such as inducing a reboot or injecting an artificial host-level fault.
[0045] The results from the test cases can be stored in a results repository 150. In an embodiment, the results store includes a telemetry repository 152, a scheduling repository 154, an execution repository 156, a host event repository 158, and a guest event repository 160. The results can include the test output applied to the sample, as described below. These results can include, for example, pass / fail indicators, the identification of the layer that failed in the case of a failure (such as hardware, host operating system, virtual machine, guest operating system, application, etc.), and the telemetry from such layers. As described herein, telemetry is defined as general measurements taken at various points in the sample under test. For example, telemetry can include performance indicators of hardware and software on various such layers, various measurements and / or time series data, such as voltage and temperature, and logs from various layers. Additionally, telemetry can include the output of test cases, logs, and the output of telemetry from the cloud control plane 116. Thus, the telemetry captured during the execution of a workload can include the output or results of the workload execution and the metrics obtained from the underlying cloud control plane, hardware configuration, or any combination thereof.
[0046] The test case scope can include any number of variables. For example, the workload type can define the test case. The workload type can be multi-variable. Different workload types can be described as compute-intensive, memory-focused, I / O or networking-limited, workload duration, workload intensity, or workload order. Additionally, various infrastructure variables can be defined for the test case. These infrastructure variables include virtual machine (VM) SKU, VM size, VM density, guest OS, geographical location, and others.
[0047] Results and telemetry from the result store can be transmitted to a scoring system 162. The scoring system 162 provides a score or rating for updates with test results stored in a result repository 150. For example, an update may not meet the readiness criteria required by external users of the service but may still have sufficient reliability to be used internally for selected workloads (which are considered less sensitive to host quality). An anomaly detection system 164 analyzes data patterns in the results and telemetry and identifies anomalies. For example, the anomaly detection system 164 can identify problem components in a large fleet through subtle perturbations in the received telemetry. A dashboard 166 is a visual representation of the cloud readiness criteria environment 100. The dashboard 166 can visually present quantitative aspects such as the number of A-sample nodes and B-sample nodes currently being executed.
[0048] In an embodiment, telemetry captured during the execution of a workload set is used to predict changes in the overhead costs associated with delivering a user experience of sufficient quality. A user experience of sufficient quality can be a user experience that is free from the negative impacts of inattention to the user experience. The overhead associated with ensuring a quality user experience can be evaluated, where the overhead and thresholds are defined in terms of: monetary cost, redundancy, energy consumed, hours of work generated due to increased redundancy, quelling quality issues, or any combination thereof.
[0049] Figure 1 The block diagram is exemplary and should not be considered limited to environment 100. Note that environment 100 can have more or fewer blocks than those shown in the Figure 1 example. Additionally, these blocks can be implemented via computing components of a computing device 500 such as Figure 5 shown.
[0050] Verification and prediction of the impact of an update on the user experience can be determined via statistical methods. The prediction can be derived from conclusions drawn from a comparison of the behavior of A-B samples over a defined timeline (e.g., several weeks) and extrapolation of data, as applied to the entire population to predict expected behavior over an extended period (or without a defined endpoint). Specifically, observations from a control sample can be compared with observations from a treated sample. As used herein, a control sample can be referred to as an "A sample," while a treated sample can be referred to as a "B sample." Each of the A sample and the B sample can be drawn from a given population. In an embodiment, the samples are "representatives" of the components in a test, either directly or indirectly. Thus, a representative sample as used herein refers to a subset of components from a population that are typical examples of the group, quality, or type associated with a particular update being tested. Thus, the update being tested can specify or notify the samples selected for comparison.
[0051] In an embodiment, the extrapolation used to derive a prediction of the impact of an update on user experience is an identity function applied to the entire population based on the comparison and conclusion of the A and B results. If the processed B sample performs better or worse than the control A sample within a finite time frame with statistical significance, then extrapolation can be performed for the behavior of the population. This behavior applied to the population can predict the update performance of the entire population represented by the sample over an infinite time span.
[0052] In another embodiment, the extrapolation is a practical function that is the opposite of the identity function. This practical function can be used for various reasons. For example, the A and B samples may not represent the broader corresponding populations as desired. In such a case, the extrapolation of the A - B comparison may need to consider a function of the differences between the samples and the broader populations. For example, consider a situation where the broader population has a slower CPU than the A - B samples. In this example, the extrapolation must "correct" the expected performance accordingly, which may be considered an insufficient update. Another reason for using a practical function is to introduce a safety margin. For example, if the power consumption limit of the samples is more lenient than that of the broader population and the update consumes a slightly higher amount of power, then this may render it insufficient (which is actually the extrapolation function affecting the prediction).
[0053] In the A - B comparison according to the present technology, the control sample and the processed sample are statistically comparable. As used herein, statistically comparable can refer to samples that are similar in quality, quantity, or type. For example, statistically comparable samples can have the same type and number of components. Thus, statistically comparable samples are those samples obtained from a population that meets the same sample requirements, such as the type and number of nodes, servers, hard disk drives, modems, switches, routers, virtual machines, or other components.
[0054] In the statistical test according to the present technology, two statistically comparable samples are selected. One sample is designated as A or the control sample, and the other sample is designated as B or the processed sample. The A or control sample is the hardware / software configuration to which the update has not been applied. The B or processed sample is the hardware / software configuration to which the update has been applied. It is important to compare the performance of the processed sample and the control sample because components typically experience background noise issues, where these issues are not immediately noticeable.
[0055] Statistical analysis can be performed on a number of A-B comparison techniques and then extrapolated to a broader population over an infinite time span to create a prediction of the impact on the user experience and to determine if an update is cloud-ready. For example, the A-B comparison can be based on overall performance. A processed B sample with poor performance can be a negative indicator and vice versa. Performance indicators include any of the following: time-based benchmarks (less elapsed time is generally considered better), degree of concurrency (more is better, but less is considered better in some cases), number of IO operations, total energy consumption (generally less is better), power consumption, number of computational cycles, memory footprint, and so on. These performance indicators may vary among different entities in the sample. Therefore, statistical methods are useful.
[0056] Another comparison can be to count the number of specific events. Some of these issues randomly manifest as failures and damages in aspects or components of the cloud infrastructure. Any isolated event may not be significant as it may seem like an occasional occurrence. When tracking such an event in a representative and large enough sample, extrapolation can be performed on such events. The number of such events can be counted and then extrapolated to the population based on that count. In an embodiment, events can be screened for events that occur for irrelevant or legitimate reasons.
[0057] Statistical analysis can be performed on leading indicators (LIs) prior to a failure or damage event in aspects of the cloud infrastructure and then extrapolated to a broader population over an infinite time span to create a prediction of the impact on the user experience and to determine if an update is cloud-ready. For example, the latency of a storage unit may increase significantly prior to a complete failure. The power consumption of a dual in-line memory module (DIMM) may deviate from the common pattern of that model prior to a failure. In an embodiment, several leading indicators are tracked. Whether the leading indicators deviate from and depart from the pattern (or other aspects) exhibited by healthy nodes can be automatically monitored. An alert can be issued regarding such an anomaly. In this way, the leading indicators can provide insights into update behavior and also reduce the sample size and the time required to obtain reliable experimental data.
[0058] In some cases, it is necessary to analyze first-time anomalies and collect this data as statistical data for A and B samples. An anomaly may be a component failure. In some cases, the anomaly may be more subtle. For example, the anomaly may be that a leading indicator deviates from the common pattern of a healthy sample (or population). In this example, if B exhibits fewer or more early failures than A, the update is an improvement or deterioration relative to A, respectively.
[0059] Cloud infrastructure components such as nodes, VMs, switches, storage, etc. typically experience some problems during execution. Therefore, this technique compares A / B samples to determine which problems can be attributed to the update when a certain type of problem is expected to occur in both samples. For example, the comparison can be between availability and downtime, uptime and downtime, interruptions, data loss, etc. In an embodiment, the comparison contrasts the details of such anomalies. For example, if the A and B samples fail unexpectedly due to different reasons, or even just due to different telemetry, these failures can indicate that the update introduced differences - which may be harmful. Additionally, general telemetry from each sample can also be compared. For example, general telemetry such as power consumption, thermal footprint, number of compute cycles, performance of storage devices, etc. can be compared for each sample.
[0060] The workloads used for testing include the complexity and diversity associated with typical cloud-based workloads. The scale and diversity inherent in the cloud can be used during testing to ensure there are sufficient resources to cover the test cases required to test cloud readiness criteria. Specifically, the A / B samples in the test are assigned to ensure there is sufficient support to execute the required workloads. In an embodiment, the samples in the test are scaled out to ensure sufficient components are used to run the required workloads. In an embodiment, scalability can refer to the ability to create or expand the compute / storage capacity of the samples in the test to accommodate typical cloud usage requirements. Additionally, a statistically sufficient scale can indicate a sample size that reduces the margin of error to below a predefined threshold. By scaling out the samples in the test, this technique ensures that users do not experience regressions due to a lack of sufficient components available for testing. In an embodiment, the size of the sample can determine the confidence level for the verification and prediction of cloud readiness.
[0061] In an embodiment, a P-value can be determined for the results of a statistical analysis. The P-value can indicate the data points for which the results are considered to have statistical significance. For example, a P-value threshold of 0.1 or 0.05 can be used, where values less than the threshold are statistically significant. In an embodiment, the P-value can be increased or decreased based on the nature of the component being updated and other considerations.
[0062] Diversity ensures that there are a sufficient variety of component instances in each sample for update testing. In an embodiment, the relevant cloud-level diversity implementation can be applied to various component instances updated in the processed samples and control samples. Generally, a cloud service provider can provide access to multiple SKUs. The samples used for testing are representative of these SKUs. To obtain samples for testing for a specific SKU, the applicability of hardware SKUs can be filtered first. For example, given a microcode change, only the blades with CPUs affected by this microcode change are considered to be included in predicting the correctness of the microcode change and validating the microcode change. In an embodiment, the samples of hardware SKUs should span equivalent categories, and the equivalent categories span the general population of hardware SKUs determined to be applicable to the current microcode change. For example, the general population of applicable hardware SKUs should include blades with CPUs having various step sizes, various memory types and sizes, various bus frequencies, various BIOSs, etc.
[0063] In addition, to verify a microcode change or update, multiple hardware instances are required for each equivalent category. The size of each equivalent category is sufficient for A / B analysis. In an example, each equivalent category can include dozens of nodes or more. In an embodiment, the equivalent category size can be skewed to amplify various signals. For example, if an update is suspected of causing quirky behavior in certain categories, the sample size can be increased accordingly. In addition, if a certain category has an error history, the corresponding sample size will increase. Finally, if a certain workload is known to cause errors, the number of nodes exposed to that workload will increase across equivalent categories.
[0064] By manipulating the scale and diversity of the associated A / B samples in the test, the test according to the present technology can repeatedly execute workloads at a statistically significant scale in different integration environments of the cloud operated and managed by a cloud service provider. The correctness of the update can be predicted based on the results of the workload execution of the applied update.
[0065] Figure 2 is a flowchart of a method 200 for implementing the verification and prediction of hardware updates. In an embodiment, method 200 enables a cloud service provider to predict the correctness of a hardware update, such as a newly installed component or a repaired component. The component can be, for example, a hardware device, such as a server, a rack, a network switch, a router, a quantum computer, a storage device, a power supply unit (PSU), etc. The hardware update according to the present technology can be a specific configuration of a component. This configuration of the component can be referred to as the target hardware configuration. Thus, at block 202, the target hardware configuration is deployed in the cloud infrastructure. Additionally, a baseline hardware configuration is identified. As described above, the target hardware configuration and the baseline hardware configuration can be compared. Specifically, the baseline hardware configuration can be regarded as a control sample or an A sample. The target hardware configuration can be regarded as a processed sample or a B sample.
[0066] At block 204, test variables for the experiment are defined. As used herein, an experiment refers to running a workload on components of a cloud infrastructure. In the case of a host node, the experiment can consist of the workload to be executed on the host. In an embodiment, the test variables include synthetic workloads and fault injection. The synthetic workload can be selected based on typical workloads of a baseline hardware configuration. Thus, the synthetic workload can include one or more tasks known to utilize or execute on at least a portion of the baseline hardware configuration. In an embodiment, the experiment can be a test case defined by a target hardware configuration, a synthetic workload, and fault injection. The test case is executed at cloud-native scale and with cloud-native diversity. In an embodiment, updates can be tested in a production environment by using a synthetic workload that is designed to generally cover hardware functionality, particularly aspects that have been changed. The use of the synthetic workload can be executed such that it does not affect the production workload. Additionally, fault injection can be selected to simulate test scenarios that rarely occur in typical situations.
[0067] At block 206, the test variables are applied to the target hardware configuration and the baseline hardware configuration. In other words, each A-sample hardware configuration and B-sample hardware configuration are tested under similar conditions. At block 208, telemetry data points are collected for analysis and comparison. Telemetry data points are collected for each target hardware configuration and baseline hardware configuration. At block 210, the cloud-readiness of the update is predicted based on the telemetry, reliability, and performance of the target hardware configuration relative to the baseline configuration.
[0068] Figure 3 is a flowchart of a method 300 for implementing verification and prediction of software updates. In an embodiment, method 200 enables a cloud service provider to predict the impact of a software update, such as microcode, on the user experience. Thus, the software update can be, for example, code, microcode, firmware, etc. In an embodiment, the software update is code for modifying or mediating hardware functionality. In some cases, the update can be code applied to update a physical component. For example, the update can be firmware for a physical component. In Figure 3 the example, the target firmware is described as a software update. However, any software update can be used according to the present technology.
[0069] At block 302, the target firmware is deployed to a sample of cloud infrastructure servers. An equivalent group of servers is identified as the baseline. As described above, the target firmware and the baseline firmware can be compared. Specifically, the baseline firmware can be a previous version of the firmware and is regarded as the control sample or A-sample. The target firmware can be regarded as the processed sample or B-sample.
[0070] At block 304, test variables are defined for the experiment. In an embodiment, the test variables include synthetic workloads and fault injection. The synthetic workload can be selected based on the typical workload of the baseline firmware. Thus, the synthetic workload may include one or more tasks known to be utilized or executed on at least a portion of the baseline firmware. In an embodiment, the experiment may be a test case defined by the target firmware, the synthetic workload, and the fault injection. The test case is executed at cloud-native scale and cloud-native diversity. In an embodiment, the update can be tested in a production environment by using a synthetic workload that is designed to generally cover the hardware functionality, particularly the aspects that are changed. The use of the synthetic workload can be executed such that it does not affect the production workload. Additionally, the fault injection can be selected to simulate test scenarios that rarely occur in typical situations.
[0071] At block 306, the test variables are applied to the target firmware and the baseline group. In other words, each of the Sample A and Sample B is tested under similar conditions. At block 308, telemetry data points are collected for analysis and comparison. Telemetry data points are collected for each target firmware and baseline firmware. At block 310, the cloud-readiness of the update is predicted based on the telemetry, reliability, and performance of the target firmware relative to the baseline configuration.
[0072] Figure 4 is a flowchart of a method 400 for implementing a statistical analysis according to the present technology. At block 402, telemetry data points are collected for analysis and comparison. For example, telemetry data points can be collected at Figure 2 block 208 of Figure 3 or block 308 of
[0073] At block 408, a pass / fail decision regarding the deployment of an update is made based on the predicted impact on the user experience. By determining the root cause of a user regression, the pass / fail decision can be achieved, and the regression can be prevented from leaking into production through a "fail" decision. In an embodiment, a regression can be an issue introduced by a recently deployed update to the cloud infrastructure. A regression can also be a latent vulnerability that remains dormant until a specific payload or a change in the cloud infrastructure causes the vulnerability to become prominent. As used herein, "become prominent" can refer to an issue that is noticed. In some cases, a user regression is indicated via the Annual Interruption Rate (AIR) or a peak in the User Deployment Performance / Reliability (TDP / R). The AIR and TDP / R metrics are two fundamental KPIs that can be used to understand, set baselines, and compare the impact of hardware, software, and configuration changes to the cloud data and control plane on the user experience. Predicting the impact of an update on the user experience can also include predicting the AIR impact of any update. In an embodiment, the AIR metric can measure the likelihood of a user experience interruption within a year. Thus, the AIR is a KPI focused on the user experience. In an embodiment, the AIR and TDP / R metrics are derived from the underlying cloud control plane, hardware configuration, or any combination thereof.
[0074] This technology can address the challenge of assessing cloud readiness in an evolving and complex dynamic environment, while existing methods are limited to statically configured laboratories. The latter are not only overwhelmed by the complexity and diversity of the cloud but also unable to keep up with the dynamic nature of cloud evolution. Additionally, this technology addresses the reality that large populations always exhibit edge cases and errors. Thus, this technology leverages the probabilistic behavior of large populations, while existing methods are limited to deterministic testing or best capture quirks.
[0075] In some cases, the verification and prediction of cloud readiness are a service. By enabling CICD integration, a third party (such as a hardware vendor) can analyze and evaluate cloud readiness criteria. The CICD integration can be achieved via Figure 1 the CICD / Integration System block 108. Continuous Integration (CI) and Continuous Delivery (CD) enable the transfer of test cases to the test environment with speed, security, and reliability. Specifically, Continuous Integration (CI) allows developers (such as partner teams) to integrate code into a shared repository multiple times while building test cases. Then, each check-in is verified by an automated build, allowing the team to detect issues early. By integrating regularly, errors can be detected faster and more easily located. Continuous Delivery (CD) is the implementation of updates into the verification and prediction process. In this way, updates and ultimately the hardware configuration and the cloud are always in a deployable state, even in the face of partner teams making changes daily.
[0076] To implement cloud readiness as a service, a cloud readiness standard system implements self-organizing requests that can concurrently verify cloud readiness standards. This requires dynamic response and provisioning of sufficient resources corresponding to the type of hardware updates, an automated control plane for managing concurrent verification, etc. Additionally, the cloud readiness standard system implements multiple users and multiple users across multiple accounts, tenants, or subscriptions. Finally, the cloud readiness standard system must be instrumented to allow programmable and scriptable integration. Programmable and scriptable integration enables the processes described below.
[0077] For example, an update can be deployed. In some cases, the update is deployed using metadata that notifies its applicability (filtering target hardware nodes) and deployment instructions or tools. For example, the deployment instructions or tools can include time delays, restarts, deployment guidance / scripts, and optionally tools related to the deployment guidance / scripts. The scope of verification and prediction is defined. As used herein, defining the scope means determining the type of workloads to be used during a test case, the allowed time, and the number of concurrent nodes. Next, test cases derived from the required verification / prediction scope are monitored. Specifically, APIs can be implemented to start, stop, obtain mid-progress status, mid-results, and final results. Telemetry is obtained from A / B samples in the test. Specifically, APIs can be implemented to obtain telemetry from systems in the test, which allows troubleshooting and debugging of those systems. The APIs can also reserve / release hardware capacity to ensure that the required hardware is available / released.
[0078] In this way, third parties can access the cloud readiness standards and can verify or predict the readiness of hardware updates. The access to the cloud readiness standards described herein can be provided as a service to third parties, resulting in a change in the economics of hardware update development. For example, CPU developers will be able to iterate changes to microcode in small steps, similar to software developers, and verify such incremental changes against the actual evolving cloud. Traditionally, CPU developers have been limited to statically configured labs. By extension, the verification pipeline for hardware updates (also known as the CICD pipeline) can be live integrated to rely on the cloud readiness standard service described herein.
[0079] Turning to Figure 5 , Figure 5 is a block diagram illustrating an exemplary computing device 500 configured to verify and predict cloud readiness in accordance with aspects of the disclosed subject matter. Figure 5is an example of a computing environment in which architecture 100 or portions thereof can be deployed, for example. Exemplary computing device 500 includes one or more processors (or processing units), such as processor 502 and memory 504. Processor 502 and memory 504, along with other components, are interconnected via system bus 510. System bus 510 can be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus, also known as a mezzanine bus. Refer to Figure 1 The memory and programs described can be deployed in Figure 5 the corresponding portions of.
[0080] Memory 504 generally (but not always) includes both volatile memory 506 and non-volatile memory 508. Volatile memory 506 holds or stores information as long as the memory is powered. In contrast, non-volatile memory 508 is capable of storing (or persisting) information even when power is unavailable. In general, RAM and CPU cache memory are examples of volatile memory 506, while ROM, solid state storage devices, memory storage devices, and / or memory cards are examples of non-volatile memory 508.
[0081] Computing device 500 can also include other removable / non-removable volatile / non-volatile computer storage media. By way of example only, computing device 500 can also include a hard disk drive that reads from or writes to non-removable, non-volatile magnetic media, a disk drive that reads from or writes to removable, non-volatile disks, and an optical disk drive that reads from or writes to removable, non-volatile optical disks such as a CD ROM or other optical media. Other removable / non-removable, volatile / non-volatile computer storage media that can be used in the exemplary operating environment include, but are not limited to, cassette tapes, flash memory cards, digital versatile disks, digital video tapes, solid state RAM, solid state ROM, and the like. The hard disk drive can be connected to system bus 510 via a non-removable memory interface. The disk drive and the optical disk drive can be connected to system bus 510 via a removable memory interface.
[0082] The computing device 500 may also include a Basic Input / Output System (BIOS) that contains basic routines such as those that help transfer information between elements within the computing device 500 during startup and is typically stored in the non-volatile memory 508. The volatile memory 506 typically contains data and / or program modules that are immediately accessible to and / or currently being operated on by the processor 502. By way of example, and not limitation, Figure 5 it may also include an operating system, application programs, other program modules, and program data.
[0083] The computing device 500 generally includes a variety of computer-readable media. Computer-readable media can be any available media that is accessible by the computing device 500 and includes both volatile and non-volatile media, removable and non-removable media. By way of example, and not limitation, computer-readable media may include computer storage media and communication media. Computer storage media is different from and does not include modulated data signals or carrier waves. It includes hardware storage media, including volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and is accessible by the computing device 500. Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a transmission mechanism and includes any information delivery media. The term "modulated data signal" refers to a signal whose one or more characteristics are set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.
[0084] The processor 502 executes instructions retrieved from the memory 504 (and / or from computer-readable media) when performing the various functions of the cloud-readiness criteria as described above. The processor 502 may be composed of any one of a number of available processors such as a single-processor, multi-processor, single-core unit, and multi-core unit.
[0085] In addition, the illustrated computing device 500 includes network communication components 512 for interconnecting the computing device with other devices and / or services via a computer network. In an embodiment, the computing device 500 may implement access to, such as Figure 1Access to the cloud readiness criteria and verification environment 100 shown. The network communication component 512, sometimes referred to as a network interface card or NIC, communicates over a network using one or more communication protocols via a physical / tangible (e.g., wired, optical, etc.) connection, a wireless connection, or both. As will be readily understood by those skilled in the art, a network communication component (such as network communication component 512) typically includes hardware and / or firmware components (and may also include or incorporate executable software components) that send and receive digital and / or analog signals over a transmission medium (i.e., the network).
[0086] The computing device 500 also includes an I / O subsystem 514. As will be understood, the I / O subsystem includes a collection of hardware, software, and / or firmware components that enable or facilitate communication between a user of the computing device 500 and the processing system of the computing device 500. In fact, via the I / O subsystem 514, a computer operator can provide input via one or more input channels, by way of illustration and not limitation, such as a touchscreen / haptic input device, buttons, a pointing device, audio input, optical input, an accelerometer, etc. Output or presentation of information can be made via one or more display screens (which may or may not be touch-sensitive), speakers, haptic feedback, etc. As will be readily understood, the interaction between the computer operator and the computing device 500 is enabled via the I / O subsystem 514 of the computing device.
[0087] The computing device 500 also includes a cloud readiness criteria manager 516 and an authoring system 518. The cloud readiness criteria manager 516 can be used to test updates to hardware in an isolated environment 520 that supports or replicates cloud computing services. Cloud computing services include, but are not limited to, devices for storing and managing data, executing applications, or delivering content, storage devices, networking, and software, or services such as streaming video, webmail, office productivity software, or social media. The authoring system 518 can implement the development of test cases, including test logic and test parameters. Alternatively or additionally, the functions described herein can be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0088] It should also be noted that the different embodiments described herein can be combined in different ways. That is, parts of one or more embodiments can be combined with parts of one or more other embodiments. All of this is contemplated herein.
[0089] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for verification and prediction of cloud readiness, comprising: Identifying a sample of components from a cloud infrastructure, wherein at least one of a hardware update or a software update is applied to the sample to generate a processed sample, and the processed sample has a statistically sufficient size and relevant cloud-level diversity, wherein the statistically sufficient size is a sample size that reduces a margin of error below a predefined threshold, and wherein the processed sample includes at least one of a hardware configuration or a software configuration of the components to which at least one of the hardware update or the software update has been applied; Identifying a control sample of components from the cloud infrastructure, wherein the control sample is statistically comparable to the processed sample, and wherein the control sample includes at least one of a hardware configuration or a software configuration of the components to which at least one of the hardware update or the software update has not been applied; Executing a workload set on the processed sample and the control sample; And Predicting an impact of at least one of the hardware update or the software update on a user experience based on a comparison of telemetry captured during execution of the workload set on the processed sample and the control sample.
2. The method according to claim 1 further comprises: Selecting a workload set representative of actual usage of a population represented by the control sample and the processed sample based on characteristics that accelerate the likelihood of discovering issues with at least one of the hardware update or the software update.
3. The method according to claim 1, wherein the comparison of the telemetry captured during the execution of the workload set of the processed sample and the control sample comprises: Comparing statistical analyses of metrics obtained from the telemetry.
4. The method according to claim 1, further comprising: Predicting a customer impact of at least one of the hardware update or the software update based on the determined reliability and performance of at least one of the hardware update or the software update.
5. The method according to claim 1, wherein the telemetry captured during the execution of the workload comprises: Outputs of the workload execution and metrics obtained from an underlying cloud control plane, a hardware configuration, or any combination thereof.
6. The method according to claim 1, wherein the relevant cloud-level diversity enables various component instances to which at least one of the hardware update or the software update is applicable in the processed sample and the control sample.
7. The method according to claim 1, further comprising: Injecting faults when executing the workload set on the processed sample and the control sample to discover errors associated with at least one of the hardware update or the software update.
8. The method according to claim 1, wherein a change in an overhead cost associated with a delivery-quality user experience is predicted using the telemetry captured during execution of the workload set.
9. The method according to claim 1, wherein an overhead associated with ensuring a quality user experience is evaluated, wherein the overhead is defined in terms of monetary cost, redundancy, energy consumed, labor hours due to increased redundancy, pacifying quality issues, or any combination of the foregoing.
10. The method according to claim 1, wherein at least one of the hardware update or the software update includes a microcode change to a central processing unit (CPU) that is to be applied to the cloud infrastructure and modifies aspects of the cloud infrastructure, a firmware or microcode change to a graphics processing unit (GPU), a firmware change to a network chipset, a basic input / output system (BIOS) code change, a field programmable gate array (FPGA) reprogramming, a hybrid hard disk drive (HHD) firmware, a solid state drive (SSD) firmware, a host operating system, a network interface card (NIC), or any combination of the foregoing.
11. The method according to claim 1, wherein at least one of the hardware update or the software update includes the installation of new or repaired servers, racks, network switches, routers, quantum computers, storage devices, power units, or any combination of the foregoing.
12. A computer-readable storage medium carrying computer-executable instructions that, when executed on a computing system including at least a processor, perform a method for verification and prediction of cloud readiness, the method comprising: identifying a sample of components from a cloud infrastructure, wherein at least one of a hardware update or a software update is applied to the sample to generate a processed sample, and the processed sample has a statistically sufficient size and relevant cloud-level diversity, wherein the statistically sufficient size is a sample size that reduces a margin of error below a predefined threshold, and wherein the processed sample includes at least one of a hardware configuration or a software configuration of the components to which at least one of the hardware update or the software update has been applied; identifying a control sample of components from the cloud infrastructure, wherein the control sample is statistically comparable to the processed sample, and wherein the control sample includes at least one of a hardware configuration or a software configuration of the components to which at least one of the hardware update or the software update has not been applied; performing a workload set on the processed sample and the control sample; and predicting an impact of at least one of the hardware update or the software update on a user experience based on a comparison of telemetry captured during the performance of the workload set on the processed sample and the control sample.
13. The computer-readable storage medium according to claim 12, wherein the comparison of the telemetry captured during the execution of the workload set of the processed sample and the control sample includes: Comparing statistical analyses of metrics derived from the telemetry.
14. The computer-readable storage medium according to claim 12, comprising: Selecting a workload set representative of actual usage of a population represented by the control sample and the processed sample based on characteristics that accelerate the likelihood of discovering issues with at least one of the hardware update or the software update.
15. The computer-readable storage medium according to claim 12, wherein cloud-level scale is defined by using statistical methods to compare results captured in response to performing the workload set on the processed sample and the control sample, as defined by a plurality of nodes.
Citation Information
Patent Citations
Using cloud-based data for industrial simulation
CN104144204A
Dynamic Cleaning for Malware Using Cloud Technology
US20130061325A1