Cloud Maintenance Duration and Risk Estimation via Telemetry Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The complexity of cloud computing environments makes it difficult for system administrators to accurately predict the duration and success of hardware/software upgrades and cloud reconfigurations, leading to reluctance in performing necessary maintenance operations, which can impact system efficiency and operability.

Innovation Solution

A system that collects telemetry data from multiple cloud computing systems and uses machine learning to analyze configuration information and past maintenance operations, providing accurate duration and risk estimates for future upgrades and reconfigurations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If cloud service providers perform hardware/software upgrades and cloud reconfigurations to improve system efficiency and capacity, then system operability and efficiency are improved, but system downtime occurs and maintenance scheduling becomes uncertain

Engineering Contradiction:
Improvesystem efficiencyVSAvoidsystem operability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary analysis of maintenance operations by collecting telemetry data from multiple cloud computing systems and using machine learning models to predict duration and risk estimates before maintenance is executed. This allows administrators to plan and schedule maintenance operations in advance with accurate timeframes, reducing unexpected downtime while maintaining system efficiency improvements.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If system administrators wait for accurate maintenance duration predictions before scheduling operations, then maintenance scheduling accuracy is improved, but maintenance operations are delayed

Engineering Contradiction:
Improvemaintenance scheduling accuracyVSAvoidmaintenance operation timing
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The system continuously collects telemetry data from cloud computing systems during and after maintenance operations, feeds this data back into machine learning models, and uses the learned patterns to improve future duration and risk predictions. This feedback loop enables accurate scheduling predictions without delaying necessary maintenance operations, as the system learns from historical data to provide increasingly accurate estimates.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If cloud computing environments are virtualized to improve resource utilization and scalability, then system capacity is improved, but system complexity increases making maintenance prediction difficult

Engineering Contradiction:
Improvecloud scalabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system creates a universal machine learning model that analyzes telemetry data across multiple different cloud computing systems with various virtualization configurations. By learning from diverse system architectures and maintenance operations, the model develops universal patterns for predicting maintenance duration and risk that work across different cloud environments, managing complexity through data-driven generalization rather than environment-specific rules.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11121919B2Methods and apparatus to determine a duration estimate and risk estimate of performing a maintenance operation in a networked computing environment
Publication Date: 2021.09.14 VMWARE INC
  • US11121919B2 patent drawing
  • US11121919B2 patent drawing
  • US11121919B2 patent drawing

AI summary

Methods, apparatus, and systems are disclosed for determining a duration and/or risk estimate of performing a maintenance operation in a networked computing environment. An example apparatus includes a client data datastore to store telemetry data, associated with the maintenance operation, system configuration data, and a first secret, all received from a first virtual computing component operating in a virtual computing environment. The apparatus further includes a configuration comparator to identify configuration changes between a previous and current configuration of the virtual computing environment. A client data analyzer selects, based on the changes, an analysis model, and applies the model to generate results including the duration estimate or a risk of failure. The client data analyzer stores the results with the first secret and the configuration data, and a result interface retrieves and transmits the stored analysis results associated with the first secret matching a second secret included in a request.