Data Ingestion Platform for Heterogeneous Source Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data ingestion systems face difficulties in efficiently managing and processing data from diverse, dynamic, and heterogeneous sources, particularly in handling structured and unstructured data, and ensuring the secure management of sensitive information, which hinders effective data collection and usage.
Innovation Solution
A data ingestion platform that employs acquisition agents to collect raw data from various sources, categorization models to classify data, matching models to unify data into a schema, and profile models to generate subject and non-subject facts, while ensuring secure handling and reporting, with journaling for data traceability and privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional data ingestion systems are used to collect data from diverse sources, then data collection capability is provided, but efficiency and effectiveness are hindered due to difficulty in managing heterogeneous data formats and structures
Solution Approach 1:
The patent introduces an intermediary layer consisting of adapters and normalizers that mediate between diverse data sources and the core processing system. Adapters transform data from various sources into a standardized intermediate format, while normalizers further refine this data into a unified schema, thereby resolving the complexity of handling heterogeneous data formats without compromising ingestion efficiency
Solution Approach 2:
The data ingestion system is segmented into distinct modular components including source-specific adapters, data normalizers, schema validators, and processing engines. Each component handles specific aspects of data transformation independently, allowing the system to manage diverse data formats through specialized modules rather than a monolithic complex system
2Quantity of substance
If data from multiple heterogeneous sources is ingested, then data variety and volume are increased, but difficulty in processing and managing structured and unstructured data increases
Solution Approach 1:
The system dynamically adjusts processing parameters based on data characteristics. Different data types (structured, semi-structured, unstructured) are routed to appropriate processing pipelines with customized parameters. For example, JSON data triggers one set of parsing parameters while CSV data triggers another, enabling efficient processing of diverse data formats without manual configuration
Solution Approach 2:
Normalizer components act as intermediaries that automatically detect data structure types and apply appropriate transformation rules. These normalizers serve as mediators between raw heterogeneous data and the unified schema, automatically adapting processing approaches based on whether the input is structured, semi-structured, or unstructured data
3Quantity of substance
If sensitive or private data is processed, then comprehensive data collection is achieved, but special management requirements increase to meet commercial or regulatory requirements
Solution Approach 1:
The system performs preliminary classification and tagging of sensitive data during the ingestion phase. Data is automatically identified, categorized by sensitivity level, and marked with appropriate metadata before entering the processing pipeline. This preliminary action enables downstream systems to apply appropriate access controls and retention policies without adding complex management overhead later
Solution Approach 2:
The system implements feedback mechanisms that monitor data processing activities and automatically adjust handling of sensitive information. When sensitive data is detected, the system provides feedback to routing logic to direct this data through specialized processing channels with enhanced security measures, automatically adapting to compliance requirements without manual intervention
4Adaptability or versatility
If data is collected from unbounded variety of sources with different formats and interfaces, then data source diversity is increased, but difficulties in efficient collection, management, or use of disparate data increase
Solution Approach 1:
The patent implements a universal adapter framework that can interface with multiple different data sources through standardized mechanisms. Rather than requiring custom integration code for each source type, the system uses a multi-functional adapter layer that handles various protocols and formats through a unified interface, thereby maintaining data source diversity while preserving processing efficiency
Data Source
AI summary
Embodiments are directed to data ingestion over a network. Raw data and integrated data associated with a plurality of separate data sources may be provided such that the raw data includes content associated with a plurality of subjects. Categorization models may be employed to categorize the raw data based on various features, such as, format, structure, data source, variability, volume, or associated entities. Matching models may be determined based on the categorization of the of the raw data, the integrated data and the content associated with the plurality of subjects. Matching models may generate a plurality of unified facts based on the raw data and the integrated data such that each unified fact is associated with a score associated with a quality of its match with a unified schema.


