Data Selection for ML Models via Subset Metric Profiling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data profiling and selection techniques for machine learning models are limited by their focus on numeric and categorical variables, leading to suboptimal performance due to environmental and security constraints, requiring excessive computing resources and memory, especially in secure environments.

Innovation Solution

A computer-implemented method for data selection that generates models based on dataset metric values, identifies data subsets, and optimizes performance by calculating subset metric values, allowing for efficient data profiling and selection within secure environments without exhaustive searches, reducing resource requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current data profiling techniques are used that focus on numeric and categorical variables, then data selection can be performed, but model performance is suboptimal and excessive computing resources and memory are required

Engineering Contradiction:
Improvemodel performanceVSAvoidcomputing resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent changes the parameters used for data profiling from traditional numeric and categorical variables to include unstructured data characteristics and security environment constraints. This enables more effective data selection that improves model performance while reducing the computing resources needed by focusing on relevant data attributes.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the data selection process into multiple stages: initial data profiling, subset identification based on performance thresholds, and iterative refinement. This segmentation allows the system to avoid exhaustive search of all possible data subsets, thereby reducing memory and computing resource requirements while still achieving optimal model performance.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If exhaustive data profiling of multiple datasets is performed to optimize model performance, then model accuracy improves, but excessive computing resources and memory are required

Engineering Contradiction:
Improvemodel accuracyVSAvoidmemory
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies partial action by performing data profiling on a selective basis rather than exhaustively profiling all possible data subsets. The system identifies data subsets that meet performance thresholds and focuses profiling efforts on those candidates, achieving high model accuracy without requiring memory resources for exhaustive analysis of all potential subsets.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If data is transmitted within secure environments for profiling, then model optimization can occur, but security limitations and computing resource allocation are constrained

Engineering Contradiction:
ImprovesecurityVSAvoiddata profiling efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent introduces an intermediary data profiling system that operates within secure environments to evaluate data subsets before model training. This intermediary layer enables security-conscious data selection by assessing data characteristics and generating performance predictions without requiring transmission of confidential information, thereby maintaining security while improving profiling efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230385706A1Data selection for machine learning models based on data profiling
Publication Date: 2023.11.30 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20230385706A1 patent drawing
  • US20230385706A1 patent drawing
  • US20230385706A1 patent drawing

AI summary

A method, computer system, and a computer program for data selection is provided. The present invention may include generating a first model associated with a dataset. The present invention may further include determining a first model performance level associated with the first model based on a plurality of dataset metric values of the dataset. The present invention may further include a plurality of data subsets of a dataset based on the first model performance level failing to exceed a performance threshold and calculating a plurality of subset metric values associated with the plurality of data subsets. The present invention may further include generating a second model associated with at least one data subset based on the plurality of subset metric values and determining an optimization associated with the first model based on a second model performance level associated with the second model exceeding the performance threshold.