Siamese Network Self-Calibrating Cloud Hardware Health

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Determining the health state of hardware resources in cloud systems is complex due to the large variety and different types of servers and network equipment, and existing methods like stress tests are not always appropriate or effective, especially when there is a lack of sufficient training data for machine learning models.

Innovation Solution

A Siamese network is used to self-calibrate the health state of hardware resources by generating pairs of embeddings from current and reference data, allowing for the prediction of health states based on feature variables such as failure operation data, performance data, and power telemetry data, without the need for extensive training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional stress tests and statistical analyses are used to determine hardware health state, then measurement precision may be improved, but device complexity and loss of time increase significantly

Engineering Contradiction:
Improvehealth state determination accuracyVSAvoidmonitoring system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical stress testing with a neural network-based prediction system. The neural network model learns normal hardware behavior patterns from historical data and predicts health states by comparing current metrics against learned patterns, eliminating the need for complex stress tests and statistical analyses while maintaining determination accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a virtual model (neural network) that copies and simulates normal hardware behavior patterns. This digital twin approach allows the system to predict health states by comparing actual hardware metrics against the copied normal behavior model, avoiding the need for physical stress tests and complex monitoring infrastructure

Inventive Principle:
Principle #26Copying

2Measurement precision

If extensive training data is collected for machine learning models, then prediction accuracy improves, but loss of time and device complexity increase

Engineering Contradiction:
Improvehealth state prediction accuracyVSAvoidtraining data collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by training the neural network model offline using historical hardware metrics before deployment. The model learns normal behavior patterns in advance during a training phase, so that during operational monitoring, it can immediately predict health states without requiring additional data collection or processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The neural network model serves itself by learning from historical data and then autonomously predicting health states without requiring continuous human intervention or extensive real-time data collection. The model self-calibrates by comparing current metrics against its learned normal behavior patterns

Inventive Principle:
Principle #25Self-service

3Measurement precision

If multiple types of hardware resources are monitored with different methods, then measurement precision improves, but device complexity and difficulty of detecting increase

Engineering Contradiction:
Improvehealth state detection accuracyVSAvoidhealth state monitoring complexity
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent implements a universal neural network-based monitoring system that can detect and predict health states across multiple types of hardware resources (storage devices, processors, memory, etc.) using the same approach. The system collects metrics from various hardware components and applies the same predictive model to all, eliminating the need for different detection methods for different hardware types while maintaining accurate health state determination

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240168840A1Self-calibrating a health state of resources in the cloud
Publication Date: 2024.05.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240168840A1 patent drawing
  • US20240168840A1 patent drawing
  • US20240168840A1 patent drawing

AI summary

Systems and methods are provided for self-calibrating a health state of a hardware resource using a Siamese network based on a plurality of feature variables. The feature variables may include hardware failure data, performance degradation data, and power consumption data. The hardware failure data is based on machine operation records and warranty logs. The performance degradation data is based on hourly performance data and a number of client requests for performing functions. The power consumption data uses power telemetry and a processor (e.g., CPU) usage. The present disclosure uses a Siamese network with a plurality of trained neural networks in parallel to determine a correlation between incident data and reference data (e.g., representing a hardware resource in a healthy state). Use of the Siamese network enables self-calibrating a health status of servers in a cloud system without imposing stress tests or complex computations to classify the respective servers.