Metric Agents for Real-Time GPU Metrics Collection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for collecting metrics related to computing workloads are inefficient and lack real-time capabilities, making it difficult to optimize resource utilization and troubleshoot performance issues effectively.

Innovation Solution

A framework that employs metric agents on each node to collect and store workload and system metrics, which are then transmitted to a remote analytics service for analysis, using a token-based authentication system and orchestration services like Kubernetes for efficient deployment and scaling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing techniques are used to collect metrics, then resource utilization can be monitored, but real-time capabilities are lacking and optimization is inefficient

Engineering Contradiction:
Improvemetrics collection efficiencyVSAvoidreal-time monitoring capability
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system segments the metrics collection function by deploying metric agents on individual nodes rather than using centralized collection. Each node independently collects and stores its own metrics locally, enabling parallel operations across the distributed system and achieving real-time capabilities without centralized bottlenecks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces metric agents as intermediary components that reside on each node and facilitate local metrics collection and storage. These agents act as mediators between the computing workloads and the monitoring system, enabling efficient real-time data capture without direct intervention from central monitoring infrastructure.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If centralized metrics collection is used, then data can be aggregated, but scalability and deployment efficiency are reduced

Engineering Contradiction:
Improvesystem scalabilityVSAvoiddeployment complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The architecture divides the monitoring system into independent node-level agents that operate autonomously. This segmentation eliminates the need for complex centralized deployment and allows the system to scale by simply adding more nodes with identical agent configurations, significantly improving adaptability and reducing deployment complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each node's metric agent independently collects, stores, and manages its own metrics without requiring external coordination or complex centralized management. This self-service approach simplifies the overall system architecture and enables easy scaling across distributed environments.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240070798A1Techniques to obtain metrics data
Publication Date: 2024.02.29 NVIDIA CORP
  • US20240070798A1 patent drawing
  • US20240070798A1 patent drawing
  • US20240070798A1 patent drawing

AI summary

Apparatuses, systems, and techniques to obtain metric data of a computing resource service provider. In at least one embodiment, metric data of one or more graphics processing unit (GPUs) is caused to be obtained from the one or more GPUs in an order from newest to oldest.