Intelligent Data Curation for Automated Cataloging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing size and complexity of data sets in 'big data' environments lead to inefficiencies in data preparation, as manual selection of necessary operations becomes time-consuming and overwhelming, causing bottlenecks that delay data availability for analysis and presentation.

Innovation Solution

A distributed processing system that analyzes metadata and context data to selectively include data sets in a catalog based on specified criteria, generating a score for each data set's likelihood of meeting those criteria, and suggesting relevant data preparation operations, thereby optimizing data preparation operations based on detected features and contextual aspects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual inspection is used to determine data preparation operations, then personnel can identify needed operations, but the time required increases significantly with data set size

Engineering Contradiction:
Improveaccuracy of identifying needed operationsVSAvoidtime for manual inspection
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual inspection with an automated system that uses machine learning models and algorithms to analyze data sets and determine required preparation operations. This substitution eliminates the need for personnel to manually inspect data, thereby resolving the contradiction between accurate identification and time consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables data sets to be automatically analyzed and characterized without human intervention. The automated characterization system performs the entire process of identifying data features, determining preparation needs, and generating operation recommendations independently, thus eliminating time loss while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

2Reliability

If a battery of data preparation operations is performed on every data set, then all possible operations are covered, but processing time and resources increase significantly

Engineering Contradiction:
Improvecompleteness of data preparationVSAvoiddata preparation throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements selective data preparation by applying only the subset of operations that are actually needed for each data set, rather than performing all possible operations universally. The automated characterization system identifies and applies only the necessary preparation steps, thus maintaining reliability while dramatically improving productivity by eliminating redundant operations.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system dynamically adjusts the data preparation process based on the specific characteristics of each data set. By changing the parameters of which operations are applied based on detected data features and requirements, the system ensures complete preparation when needed while avoiding unnecessary operations, thereby resolving the contradiction between reliability and productivity.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If the variety of data preparation operations increases to accommodate diverse data uses, then more uses can be supported, but the complexity of selection becomes overwhelming

Engineering Contradiction:
Improvevariety of supported data usesVSAvoidcomplexity of operation selection
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an automated characterization system as an intermediary between the diverse data sets and the variety of preparation operations. This intermediary automatically analyzes data features, determines requirements, and selects appropriate operations, thereby supporting diverse data uses while eliminating the complexity of manual selection. The system acts as a mediator that translates data characteristics into appropriate preparation actions.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the complex selection process into distinct automated steps: data characterization, feature detection, requirement determination, and operation selection. By dividing the overall process into manageable segments handled automatically by the system, the patent maintains high adaptability for diverse data uses while reducing the perceived complexity for users who no longer need to make these selections manually.

Inventive Principle:
Principle #1Segmentation

4Quantity of substance

If data set size increases, then more data is available for analysis, but the time for preparation operations increases correspondingly

Engineering Contradiction:
Improvedata set sizeVSAvoidpreparation time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent performs automated characterization and operation selection as preliminary actions before actual data preparation begins. By pre-analyzing data sets to identify features and determine required operations, the system prepares the groundwork efficiently, allowing subsequent preparation to proceed quickly regardless of data set size. This preliminary automated assessment eliminates the time penalty that would otherwise scale with data volume.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual data inspection and operation selection with automated computational systems that can process large data sets efficiently. This substitution enables the system to handle increasing data volumes without the linear increase in preparation time that would result from manual processing, as machines can analyze and characterize data much faster than human personnel.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11341414B2Intelligent data curation
Publication Date: 2022.05.24 SAS INSTITUTE INC
  • US11341414B2 patent drawing
  • US11341414B2 patent drawing
  • US11341414B2 patent drawing

AI summary

An apparatus includes processor(s) to: receive a request for a data catalog; in response to the request specifying a structural feature, analyze metadata of multiple data sets for an indication of including it, and to retrieve an indicated degree of certainty of detecting it for data sets including it; in response to the request specifying a contextual aspect, analyze context data of the multiple data sets for an indication of being subject to it, and to retrieve an indicated degree of certainty concerning it for data sets subject to it; selectively include each data set in the data catalog based on the request specifying a structural feature and/or a contextual aspect, and whether each data set meets what is specified; for each data set in the data catalog, generate a score indicative of the likelihood of meeting what is specified; and transmit the data catalog to the requesting device.