Intelligent Data Ingestion Filtering Irrelevant Items
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data analysis systems face inefficiencies and inaccuracies due to rigid ingestion methodologies, wasteful use of computing resources, and generation of inaccurate results from storing and analyzing irrelevant data items.
Innovation Solution
An intelligent data ingestion system that generates a merged data set during ingestion by filtering out irrelevant data items, combining relevant data items from multiple data sets based on specified criteria, and storing only the relevant data in non-temporary storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If conventional systems funnel entire data sets into permanent storage regardless of relevance, then data completeness is maintained, but computing resources are wasted storing and analyzing irrelevant data items
Solution Approach 1:
The system performs preliminary filtering during the data ingestion phase, evaluating and removing irrelevant data items before they are stored in permanent storage. This advance action prevents wasteful storage and subsequent analysis of irrelevant data, resolving the contradiction by maintaining only relevant data while conserving computing resources.
Solution Approach 2:
The system extracts and removes irrelevant data items from incoming data sets during ingestion, separating them from relevant data before storage. This extraction process ensures that only pertinent data items are stored and analyzed, eliminating resource waste while preserving data completeness for relevant items.
2Measurement precision
If conventional systems store and analyze all data items from requested data sets, then no data filtering is needed, but analytical accuracy deteriorates due to inclusion of irrelevant data items
Solution Approach 1:
The system performs data filtering and relevance evaluation during the ingestion phase, before analysis occurs. This preliminary action ensures that only relevant data items proceed to storage and analysis, guaranteeing analytical accuracy without requiring complex filtering operations during the analysis phase itself.
Solution Approach 2:
The system introduces an intermediary filtering layer during data ingestion that acts as a mediator between raw data and permanent storage. This intermediary process evaluates data relevance and selectively passes only appropriate items to storage, simplifying the overall system by handling filtering upfront rather than requiring complex ongoing filtering during analysis.
3Ease of operation
If conventional systems require multiple user interfaces and extensive user interactions to code and format data queries, then user control is enhanced, but operational efficiency decreases due to excessive user interactions
Solution Approach 1:
The system performs automated relevance evaluation and filtering of data items during ingestion without requiring user intervention. This self-service capability handles data preprocessing autonomously, eliminating the need for users to manually code and format complex queries across multiple interfaces, thereby enhancing operational efficiency while maintaining control through configurable parameters.
Solution Approach 2:
The system provides a unified data ingestion interface that handles multiple functions including data reception, relevance evaluation, filtering, and storage in a single integrated process. This multi-functional approach eliminates the need for multiple separate user interfaces and interactions, improving operational efficiency while preserving user control through a comprehensive single interface.
Data Source
AI summary
The present disclosure relates to systems, non-transitory computer-readable media, and methods for intelligently generating an efficient new data set for storage in non-temporary storage during ingestion of other inefficient data sets. In particular, in one or more embodiments, the disclosed systems ingests a subset of a second data set of operational data items along with a first data set of response data items based on correlations between the subset of operational data items and the response data items. Thus, the disclosed systems provide a robust solution to efficiently ingesting data sets while avoiding computing storage and processing waste.


