Intelligent Data Curation via Distributed Feature Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing size and complexity of data sets in 'big data' lead to inefficiencies in data preparation, as manual selection of necessary operations becomes time-consuming and overwhelming, causing bottlenecks in data availability for analysis and presentation.
Innovation Solution
A distributed processing system that divides data sets among processor cores, uses feature routines to detect structural and data features, generates metadata and context data, and employs suggestion models to recommend a subset of data preparation operations, which can be re-trained based on selected operations for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual inspection is used to select data preparation operations, then personnel can identify needed operations, but the time required increases significantly with data set size
Solution Approach 1:
The system performs self-service by automatically analyzing data sets and selecting appropriate preparation operations without human intervention. The automated system inspects data characteristics, determines needed operations, and executes them, eliminating the time-consuming manual inspection process while maintaining accurate operation selection.
Solution Approach 2:
The patent replaces the mechanical human inspection process with an automated computational system. Instead of personnel manually examining data sets to identify preparation needs, the system uses automated algorithms to analyze data characteristics and determine required operations, significantly reducing time while preserving selection accuracy.
2Reliability
If a battery of data preparation operations is performed on every data set, then all possible operations are covered, but unnecessary operations consume additional time and resources
Solution Approach 1:
The system extracts only the necessary data preparation operations from the full battery of possible operations. By analyzing data set characteristics and identifying which operations are actually needed, the system removes unnecessary operations from the processing pipeline, maintaining complete preparation where needed while eliminating waste, thus improving throughput without sacrificing reliability.
Solution Approach 2:
The patent applies partial action by performing only the subset of data preparation operations that are actually needed for each data set, rather than executing the full battery of operations on every data set. This selective approach maintains sufficient preparation quality while significantly reducing redundant processing time and resource consumption.
3Adaptability or versatility
If the variety of data preparation operations increases to accommodate diverse uses, then more data needs can be met, but the complexity of selection becomes overwhelming
Solution Approach 1:
The system handles the complexity of selecting from diverse operations through self-service automation. Instead of requiring personnel to navigate and select from an overwhelming variety of data preparation operations, the automated system independently analyzes data characteristics and intelligently selects appropriate operations from the available repertoire, maintaining high adaptability while eliminating selection complexity for users.
4Quantity of substance
If data set size increases, then more data is available for analysis, but manual inspection becomes increasingly difficult and time-consuming
Solution Approach 1:
The patent replaces the mechanical human inspection process with automated computational analysis. As data set sizes increase, the automated system efficiently handles the larger volumes without the diminishing returns that plague manual inspection, maintaining ease of operation by eliminating the need for human reviewers to navigate increasingly complex and voluminous data sets.
Data Source
AI summary
An apparatus includes a processor to: provide a set of feature routines to a set of processor cores to detect features of a data set distributed thereamong; generate metadata indicative of the detected features; generate context data indicative of contextual aspects of the data set; provide the metadata and context data to each processor core, and distribute a set of suggestion models thereamong to enable derivation of a suggested subset of data preparation operations to be suggested to be performed on the data set; transmit indications of the suggested subset to a viewing device, and receive therefrom indications of a selected subset of data preparation operations selected to be performed; compare the selected and suggested subsets; and in response to differences therebetween, re-train at least one suggestion model of the set of suggestion models based at least on the combination of the metadata, context data and selected subset.


