Adaptive Data Ingestion Workflow Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems lack the ability to cognitively determine the optimal processing workflow for ingesting data into a corpus, leading to inefficient resource usage and reduced throughput due to either excessive or insufficient processing, which affects the quality and accuracy of cognitive systems like deep QA and machine learning models.
Innovation Solution
A method that identifies groups of fields with common metadata attributes, generates metrics, and assigns weight values to determine a natural language processing (NLP) measure and discreteness measure, selecting an appropriate processing workflow based on predefined thresholds to optimize data ingestion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If significant NLP processing is applied to all data before ingestion, then data quality and usability are improved, but computing resources and processing time are excessively consumed
Solution Approach 1:
The patent applies different processing workflows to different groups of fields based on their specific characteristics. Fields are evaluated individually using metrics such as NLP measure and discreteness measure, and only those fields requiring NLP processing receive it. This localized approach ensures high data quality where needed while avoiding unnecessary processing for fields that don't require it, thus resolving the contradiction between data quality and resource consumption.
2Use of energy by moving object
If no NLP processing is applied to data, then computing resources are conserved, but data usability and system performance deteriorate
Solution Approach 1:
The patent changes the parameter of processing intensity based on the characteristics of each field. By calculating metrics such as NLP measure (indicating need for natural language processing) and discreteness measure (indicating data structure), the system dynamically adjusts the processing applied to each field. This ensures that data quality and system performance are maintained at appropriate levels while avoiding unnecessary resource consumption.
3Ease of operation
If uniform processing workflow is applied to all fields, then processing simplicity is maintained, but resource efficiency and throughput are reduced
Solution Approach 1:
The patent segments the data fields into different groups based on their characteristics, evaluated through metrics such as NLP measure and discreteness measure. Each segment is then assigned an appropriate processing workflow from multiple available options. This segmentation allows the system to maintain simplicity in the overall process structure while optimizing resource efficiency and throughput by applying the right level of processing to each field group.
4Manufacturing precision
If extensive data curating and processing is performed, then corpus quality is improved, but processing time and resource consumption increase
Solution Approach 1:
The patent applies partial processing action by evaluating each field against metrics such as NLP measure and discreteness measure, and applying NLP processing only to the extent necessary for each field. This partial action approach ensures that corpus quality is improved sufficiently for system performance while avoiding excessive processing that would unnecessarily increase processing time and resource consumption.
Data Source
AI summary
Improved data ingestion techniques are provided. A data set comprising records is received, where each record contains one or more fields. A group of fields is identified, where each of the fields has a common metadata attribute. Metrics are determined for the group based on metadata associated with each field, and weight values are assigned to each of the metrics. A natural language processing (NLP) measure and a discreteness measure are generated for the group of fields based on the metrics and the weight values. A processing workflow is selected to use when ingesting data from the group of fields into a corpus, based on comparing the NLP measure and the discreteness measure to one or more predefined thresholds, and each of the fields in the group of fields are processed using the processing workflow.


