Intelligent Data Curation for Automated Cataloging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing size and complexity of data sets in 'big data' environments lead to inefficiencies in data preparation, as manual selection of necessary operations becomes time-consuming and overwhelming, causing bottlenecks that delay data availability for analysis and presentation.
Innovation Solution
A distributed processing system that analyzes metadata and context data to selectively include data sets in a catalog based on specified criteria, generating a score for each data set's likelihood of meeting those criteria, and suggesting relevant data preparation operations, thereby optimizing data preparation operations based on detected features and contextual aspects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual inspection is used to determine data preparation operations, then personnel can identify needed operations, but the time required increases significantly with data set size
Solution Approach 1:
The patent replaces manual inspection with an automated system that uses machine learning models and algorithms to analyze data sets and determine required preparation operations. This substitution eliminates the need for personnel to manually inspect data, thereby resolving the contradiction between accurate identification and time consumption.
Solution Approach 2:
The system enables data sets to be automatically analyzed and characterized without human intervention. The automated characterization system performs the entire process of identifying data features, determining preparation needs, and generating operation recommendations independently, thus eliminating time loss while maintaining accuracy.
2Reliability
If a battery of data preparation operations is performed on every data set, then all possible operations are covered, but processing time and resources increase significantly
Solution Approach 1:
The patent implements selective data preparation by applying only the subset of operations that are actually needed for each data set, rather than performing all possible operations universally. The automated characterization system identifies and applies only the necessary preparation steps, thus maintaining reliability while dramatically improving productivity by eliminating redundant operations.
Solution Approach 2:
The system dynamically adjusts the data preparation process based on the specific characteristics of each data set. By changing the parameters of which operations are applied based on detected data features and requirements, the system ensures complete preparation when needed while avoiding unnecessary operations, thereby resolving the contradiction between reliability and productivity.
3Adaptability or versatility
If the variety of data preparation operations increases to accommodate diverse data uses, then more uses can be supported, but the complexity of selection becomes overwhelming
Solution Approach 1:
The patent introduces an automated characterization system as an intermediary between the diverse data sets and the variety of preparation operations. This intermediary automatically analyzes data features, determines requirements, and selects appropriate operations, thereby supporting diverse data uses while eliminating the complexity of manual selection. The system acts as a mediator that translates data characteristics into appropriate preparation actions.
Solution Approach 2:
The patent segments the complex selection process into distinct automated steps: data characterization, feature detection, requirement determination, and operation selection. By dividing the overall process into manageable segments handled automatically by the system, the patent maintains high adaptability for diverse data uses while reducing the perceived complexity for users who no longer need to make these selections manually.
4Quantity of substance
If data set size increases, then more data is available for analysis, but the time for preparation operations increases correspondingly
Solution Approach 1:
The patent performs automated characterization and operation selection as preliminary actions before actual data preparation begins. By pre-analyzing data sets to identify features and determine required operations, the system prepares the groundwork efficiently, allowing subsequent preparation to proceed quickly regardless of data set size. This preliminary automated assessment eliminates the time penalty that would otherwise scale with data volume.
Solution Approach 2:
The patent replaces manual data inspection and operation selection with automated computational systems that can process large data sets efficiently. This substitution enables the system to handle increasing data volumes without the linear increase in preparation time that would result from manual processing, as machines can analyze and characterize data much faster than human personnel.
Data Source
AI summary
An apparatus includes processor(s) to: receive a request for a data catalog; in response to the request specifying a structural feature, analyze metadata of multiple data sets for an indication of including it, and to retrieve an indicated degree of certainty of detecting it for data sets including it; in response to the request specifying a contextual aspect, analyze context data of the multiple data sets for an indication of being subject to it, and to retrieve an indicated degree of certainty concerning it for data sets subject to it; selectively include each data set in the data catalog based on the request specifying a structural feature and/or a contextual aspect, and whether each data set meets what is specified; for each data set in the data catalog, generate a score indicative of the likelihood of meeting what is specified; and transmit the data catalog to the requesting device.


