Performance Model Training Across Pods for Resource-Stable Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Training and scoring of large numbers of performance models for anomaly detection in computing operations are computationally intensive and resource-heavy, making it prohibitive to manage and evaluate very large numbers of models effectively, especially in decentralized cloud computing environments.
Innovation Solution
The method involves training and scoring performance models by grouping them by type and selecting specific computing pods for training and scoring based on resource usage, using machine learning techniques and distributing the computational load across pods to optimize resource management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large numbers of performance models are trained and scored for anomaly detection, then detection accuracy is improved, but computational resource consumption increases prohibitively
Solution Approach 1:
The patent segments the large number of performance models into multiple groups, where each group is trained and scored independently. This division allows the system to manage computational resources more efficiently by processing smaller subsets of models simultaneously, reducing the peak resource consumption while maintaining comprehensive anomaly detection coverage across all models.
Solution Approach 2:
The patent implements partial action by training and scoring a subset of performance models in each batch rather than all models simultaneously. Multiple batches are processed sequentially or in parallel, allowing the system to achieve complete anomaly detection coverage over time while keeping resource consumption at manageable levels during each individual processing cycle.
2Stability of the object's composition
If computing pods are selected based on minimum resource usage change for training, then resource stability is improved, but selection complexity increases
Solution Approach 1:
The patent employs feedback mechanisms by monitoring resource usage changes in computing pods before and after training operations. This feedback information is used to dynamically select pods with minimum resource usage change for subsequent training tasks, creating a self-regulating system that adapts to actual resource conditions while maintaining stability.
Solution Approach 2:
The patent changes the selection criterion from static pod characteristics to dynamic parameter-based selection, specifically using resource usage change as the key parameter. This allows the system to identify and select pods that exhibit minimal resource fluctuation, thereby ensuring stable training operations while the selection logic adapts to changing system conditions.
3Productivity
If computing pods are selected based on maximum resource usage for scoring, then scoring efficiency is improved, but resource management complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-assessing the resource usage characteristics of computing pods before assigning scoring tasks. Pods are evaluated and selected based on their maximum resource usage capacity in advance, ensuring that the most capable pods are chosen for scoring operations, thereby improving efficiency while the selection criteria are established beforehand.
Solution Approach 2:
The patent introduces dynamic resource management by allowing pod selection for scoring to be based on real-time resource usage monitoring. The system dynamically identifies pods with maximum resource usage capacity and assigns scoring tasks accordingly, enabling flexible and efficient resource utilization while adapting to changing system conditions during operation.
Data Source
AI summary
A method is presented to facilitate the training of a very large number of machine-learning performance models used to detect anomalies in computing operations. The models are grouped together according to model type, and are allocated to different pods of a computing environment that is used to carry out the operations being monitored. Initial training of models in a group is carried out while monitoring resource usage, and a particular pod is selected for further training based on the resource usage. The pod selected for training preferably has a minimum change in resource usage before and after the initial training. A different pod can be selected for scoring the trained models. The pod selected for scoring preferably has a maximum resource usage during an initial scoring among all pods.


