ML Anomaly Detection for Cloud Resource Usage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud computing environments face challenges in managing computing resource usage due to the difficulty in analyzing vast amounts of data across multiple accounts, leading to poor resource allocation, overuse, and mis-allocation, resulting in wasted resources and increased costs.
Innovation Solution
An anomaly detection platform utilizing machine learning models processes historical data from multiple cloud environments to identify trends and patterns, generating trained models that detect anomalous resource usage by processing current data to produce anomaly scores, enabling improved resource management and allocation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional data analysis methods are used to manage cloud computing resources, then manual monitoring and analysis can be performed, but the ability to analyze vast amounts of data across multiple accounts is insufficient, leading to poor resource allocation and overuse
Solution Approach 1:
The patent replaces manual data analysis methods with machine learning models that automatically process and analyze resource usage data. The ML models substitute human analysts and traditional monitoring tools, enabling scalable analysis of vast datasets across multiple cloud accounts without proportional increases in manual effort.
Solution Approach 2:
The patent introduces machine learning models as intermediaries between raw resource usage data and actionable insights. These models serve as mediators that transform complex multi-account data into anomaly scores and alerts, bridging the gap between data collection and resource management decisions.
2Measurement precision
If machine learning models are trained on historical data from multiple cloud environments, then anomaly detection accuracy is improved, but the complexity of model training and data processing increases
Solution Approach 1:
The patent segments the anomaly detection process into multiple independent machine learning models, each trained on specific aspects of resource usage data. This segmentation allows for specialized training on different data types (CPU, memory, storage, network) and simplifies the overall training complexity by breaking down the monolithic problem into manageable components.
Solution Approach 2:
The patent performs preliminary data processing and feature extraction during the training phase, preparing cleaned and structured historical data before model training. This preliminary action reduces the complexity of the actual training process by ensuring data is ready for consumption, and enables models to focus on learning patterns rather than handling raw data preprocessing during inference.
3Reliability
If multiple machine learning models are used to process resource usage data, then comprehensive anomaly detection is achieved, but the processing time and computational resources required increase
Solution Approach 1:
The patent divides the resource usage data into different feature sets (CPU utilization, memory consumption, storage usage, network traffic) and processes them through specialized ML models simultaneously. This segmentation enables parallel processing of different resource types, maintaining comprehensive detection while reducing overall processing time through concurrent execution.
Solution Approach 2:
The patent implements a multi-model approach where different ML models process different aspects of resource usage data in parallel. Rather than using a single comprehensive model that would require sequential processing, multiple specialized models work simultaneously on partial datasets, achieving comprehensive coverage with reduced total processing time through parallelization.
Data Source
AI summary
A device receives historical data associated with multiple cloud computing environments, trains one or more machine learning models, with the historical data, to generate trained machine learning models that generate outputs, and trains a model with the outputs to generate a trained model. The device receives particular data, associated with a cloud computing environment, that includes data identifying usage of resources associated with the cloud computing environment, and processes the particular data, with the trained machine learning models, to generate anomaly scores indicating anomalous usage of the resources associated with the cloud computing environment. The device processes the one or more anomaly scores, with the trained model, to generate a final anomaly score indicating anomalous usage of at least one of the resources associated with the cloud computing environment, and performs one or more actions based on the final anomaly score.


