Intelligent Data Curation via Distributed Feature Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing size and complexity of data sets in 'big data' lead to inefficiencies in data preparation, as manual selection of necessary operations becomes time-consuming and overwhelming, causing bottlenecks in data availability for analysis and presentation.

Innovation Solution

A distributed processing system that divides data sets among processor cores, uses feature routines to detect structural and data features, generates metadata and context data, and employs suggestion models to recommend a subset of data preparation operations, which can be re-trained based on selected operations for improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual inspection is used to select data preparation operations, then personnel can identify needed operations, but the time required increases significantly with data set size

Engineering Contradiction:
Improveaccuracy of operation selectionVSAvoidtime for manual inspection
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically analyzing data sets and selecting appropriate preparation operations without human intervention. The automated system inspects data characteristics, determines needed operations, and executes them, eliminating the time-consuming manual inspection process while maintaining accurate operation selection.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical human inspection process with an automated computational system. Instead of personnel manually examining data sets to identify preparation needs, the system uses automated algorithms to analyze data characteristics and determine required operations, significantly reducing time while preserving selection accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If a battery of data preparation operations is performed on every data set, then all possible operations are covered, but unnecessary operations consume additional time and resources

Engineering Contradiction:
Improvecompleteness of data preparationVSAvoiddata preparation throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system extracts only the necessary data preparation operations from the full battery of possible operations. By analyzing data set characteristics and identifying which operations are actually needed, the system removes unnecessary operations from the processing pipeline, maintaining complete preparation where needed while eliminating waste, thus improving throughput without sacrificing reliability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by performing only the subset of data preparation operations that are actually needed for each data set, rather than executing the full battery of operations on every data set. This selective approach maintains sufficient preparation quality while significantly reducing redundant processing time and resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If the variety of data preparation operations increases to accommodate diverse uses, then more data needs can be met, but the complexity of selection becomes overwhelming

Engineering Contradiction:
Improvevariety of data uses supportedVSAvoidcomplexity of operation selection
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system handles the complexity of selecting from diverse operations through self-service automation. Instead of requiring personnel to navigate and select from an overwhelming variety of data preparation operations, the automated system independently analyzes data characteristics and intelligently selects appropriate operations from the available repertoire, maintaining high adaptability while eliminating selection complexity for users.

Inventive Principle:
Principle #25Self-service

4Quantity of substance

If data set size increases, then more data is available for analysis, but manual inspection becomes increasingly difficult and time-consuming

Engineering Contradiction:
Improvedata set sizeVSAvoidease of manual inspection
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent replaces the mechanical human inspection process with automated computational analysis. As data set sizes increase, the automated system efficiently handles the larger volumes without the diminishing returns that plague manual inspection, maintaining ease of operation by eliminating the need for human reviewers to navigate increasingly complex and voluminous data sets.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10909460B2Intelligent data curation
Publication Date: 2021.02.02 SAS INSTITUTE INC
  • US10909460B2 patent drawing
  • US10909460B2 patent drawing
  • US10909460B2 patent drawing

AI summary

An apparatus includes a processor to: provide a set of feature routines to a set of processor cores to detect features of a data set distributed thereamong; generate metadata indicative of the detected features; generate context data indicative of contextual aspects of the data set; provide the metadata and context data to each processor core, and distribute a set of suggestion models thereamong to enable derivation of a suggested subset of data preparation operations to be suggested to be performed on the data set; transmit indications of the suggested subset to a viewing device, and receive therefrom indications of a selected subset of data preparation operations selected to be performed; compare the selected and suggested subsets; and in response to differences therebetween, re-train at least one suggestion model of the set of suggestion models based at least on the combination of the metadata, context data and selected subset.