Distributed ML Metric Evaluation via Quantile Sketches

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for evaluating the performance of machine learning models require significant computational resources and long run times, making them inefficient for large datasets that exceed available memory.

Innovation Solution

The proposed solution involves partitioning a large dataset into multiple partitions, performing intermediate computations on each partition, and then merging the results to determine performance metrics of a machine learning model using quantile sketches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current techniques are used to evaluate machine learning model performance on large datasets, then accurate performance metrics can be obtained, but significant computational resources and long run times are required

Engineering Contradiction:
Improveperformance metric accuracyVSAvoidevaluation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent divides the large dataset into multiple partitions and processes each partition separately to compute intermediate quantile sketches. This segmentation allows parallel processing across multiple computing nodes, significantly improving evaluation efficiency while maintaining metric accuracy through proper merging of intermediate results.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If current techniques are used to evaluate machine learning model performance, then comprehensive performance metrics can be determined, but the evaluation requires more memory than is practically available

Engineering Contradiction:
Improveperformance metric accuracyVSAvoidmemory usage
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The dataset is partitioned into smaller subsets that can be processed independently with limited memory resources. Each partition generates intermediate quantile sketches that are much smaller than the original dataset, enabling evaluation of large datasets that would otherwise exceed available memory capacity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate quantile sketches as mediator structures that summarize each data partition. These sketches serve as compact representations that can be merged together to produce final performance metrics, avoiding the need to load the entire dataset into memory simultaneously.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If the dataset is partitioned into multiple partitions for distributed computation, then computational efficiency and memory usage are improved, but the complexity of merging results increases

Engineering Contradiction:
Improveevaluation speedVSAvoidmerging process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The quantile sketch data structure serves as an intermediary that simplifies the merging process. Instead of complex operations on raw data from multiple partitions, the patent merges quantile sketches using efficient sketch-combining algorithms that maintain accuracy while reducing computational complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250165856A1Distributed computation of machine learning model performance metrics
Publication Date: 2025.05.22 ORACLE INT CORP
  • US20250165856A1 patent drawing
  • US20250165856A1 patent drawing
  • US20250165856A1 patent drawing

AI summary

Techniques are described for determining at least one performance metric of a machine learning model. The techniques including, obtaining a dataset generated by at least using output of the machine learning model, partitioning the dataset into two or more partitions that include one or more elements from the dataset; and generating, for each respective partition, a respective first quantile sketch and a respective second quantile sketch based at least in part on each element in the respective partition. The techniques further including generating a first merged quantile sketch by merging each respective first quantile sketch, generating a second merged quantile sketch by merging each respective second quantile sketch, and determining the at least one performance metric of the machine learning model using the first merged quantile sketch and the second merged quantile sketch.