Automated Data Ingestion and Storage Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data analysis tools generate large amounts of additional data and metadata, leading to storage space challenges due to their inability to differentiate between high and low priority data, requiring tedious and costly manual intervention in data analysis environments.
Innovation Solution
A method using machine learning models and information governance to automatically detect data analysis requests, conduct shallow term assignments, match with ranked terms, flag irrelevant metadata, generate criticality rankings, and place low priority datasets into cold storage, thereby limiting data ingestion and storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data analysis tools scan large volumes of data, then data analysis capability is improved, but storage space consumption increases
Solution Approach 1:
The system performs preliminary actions by conducting shallow term assignments and generating criticality rankings before full data analysis. This preliminary classification allows the system to identify and prioritize high-value data early in the process, enabling selective analysis that reduces overall storage requirements while maintaining analytical effectiveness.
Solution Approach 2:
The system applies local quality by differentiating data based on criticality rankings and domain classifications. High-criticality data is retained and analyzed in detail, while low-criticality data is filtered out or stored differently. This localized differentiation optimizes storage space by not uniformly treating all data equally, instead focusing resources on the most valuable data segments.
2Measurement precision
If manual intervention is used to manage data storage, then storage control precision is improved, but operational complexity and cost increase
Solution Approach 1:
The system implements self-service by automatically performing data classification, criticality ranking, and storage optimization without requiring manual intervention. The machine learning models and governance frameworks enable the system to self-manage its own data storage needs, reducing operational complexity while maintaining precise control through automated decision-making processes.
Solution Approach 2:
The system employs feedback mechanisms where historical usage data continuously refines the criticality ranking models. This feedback loop allows the system to learn from past data access patterns and improve its automated storage decisions over time, achieving precise storage control through adaptive algorithms rather than manual management.
3Reliability
If all data is stored for future analysis, then data availability is improved, but storage efficiency deteriorates
Solution Approach 1:
The system performs preliminary criticality assessment and domain classification before storing data. By pre-determining which data is likely to be accessed again based on its characteristics and historical patterns, the system can selectively store only the necessary data, improving storage efficiency while maintaining availability for high-value data types.
Solution Approach 2:
The system changes storage parameters based on data criticality rankings and domain classifications. Different storage tiers, retention policies, and access controls are applied dynamically according to data characteristics. This parameter adaptation allows the system to optimize storage efficiency by assigning different storage treatments to different data segments rather than using uniform storage policies.
Data Source
AI summary
An embodiment for managing data using machine learning models and information governance. The embodiment may automatically detect a data analysis request made within a system and identify subject datasets. The embodiment may automatically conduct shallow term assignments on each row and column of data in the subject datasets and automatically match the shallow term assignments for each row and column with a stored set of ranked terms, and automatically flag rows or columns matching with ranked terms above a predetermined threshold ranking for further analysis. The embodiment may automatically and continuously monitor and detect irrelevant metadata types to prevent subsequent analysis and storage of data including the irrelevant metadata types. The embodiment may automatically generate a criticality ranking for stored analysis datasets. The embodiment may detect low priority analysis datasets having a criticality ranking below a criticality threshold, and automatically place the low priority analysis datasets into cold storage.

