Data Ingestion Platform for Heterogeneous Source Normalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data ingestion systems face difficulties in efficiently managing and processing data from diverse, dynamic, and heterogeneous sources, particularly in handling structured and unstructured data, and ensuring the secure management of sensitive information, which hinders effective data collection and usage.

Innovation Solution

A data ingestion platform that employs acquisition agents to collect raw data from various sources, categorization models to classify data, matching models to unify data into a schema, and profile models to generate subject and non-subject facts, while ensuring secure handling and reporting, with journaling for data traceability and privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional data ingestion systems are used to collect data from diverse sources, then data collection capability is provided, but efficiency and effectiveness are hindered due to difficulty in managing heterogeneous data formats and structures

Engineering Contradiction:
Improvedata ingestion efficiencyVSAvoidsystem complexity for managing heterogeneous data
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary layer consisting of adapters and normalizers that mediate between diverse data sources and the core processing system. Adapters transform data from various sources into a standardized intermediate format, while normalizers further refine this data into a unified schema, thereby resolving the complexity of handling heterogeneous data formats without compromising ingestion efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The data ingestion system is segmented into distinct modular components including source-specific adapters, data normalizers, schema validators, and processing engines. Each component handles specific aspects of data transformation independently, allowing the system to manage diverse data formats through specialized modules rather than a monolithic complex system

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If data from multiple heterogeneous sources is ingested, then data variety and volume are increased, but difficulty in processing and managing structured and unstructured data increases

Engineering Contradiction:
Improvedata volume and varietyVSAvoiddifficulty in processing structured and unstructured data
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The system dynamically adjusts processing parameters based on data characteristics. Different data types (structured, semi-structured, unstructured) are routed to appropriate processing pipelines with customized parameters. For example, JSON data triggers one set of parsing parameters while CSV data triggers another, enabling efficient processing of diverse data formats without manual configuration

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Normalizer components act as intermediaries that automatically detect data structure types and apply appropriate transformation rules. These normalizers serve as mediators between raw heterogeneous data and the unified schema, automatically adapting processing approaches based on whether the input is structured, semi-structured, or unstructured data

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If sensitive or private data is processed, then comprehensive data collection is achieved, but special management requirements increase to meet commercial or regulatory requirements

Engineering Contradiction:
Improvecomprehensive data collectionVSAvoidspecial management overhead for sensitive data
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system performs preliminary classification and tagging of sensitive data during the ingestion phase. Data is automatically identified, categorized by sensitivity level, and marked with appropriate metadata before entering the processing pipeline. This preliminary action enables downstream systems to apply appropriate access controls and retention policies without adding complex management overhead later

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms that monitor data processing activities and automatically adjust handling of sensitive information. When sensitive data is detected, the system provides feedback to routing logic to direct this data through specialized processing channels with enhanced security measures, automatically adapting to compliance requirements without manual intervention

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If data is collected from unbounded variety of sources with different formats and interfaces, then data source diversity is increased, but difficulties in efficient collection, management, or use of disparate data increase

Engineering Contradiction:
Improvedata source diversityVSAvoidefficiency of collection, management, or use of disparate data
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements a universal adapter framework that can interface with multiple different data sources through standardized mechanisms. Rather than requiring custom integration code for each source type, the system uses a multi-functional adapter layer that handles various protocols and formats through a unified interface, thereby maintaining data source diversity while preserving processing efficiency

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11580323B2Data ingestion platform
Publication Date: 2023.02.14 ASTRUMU INC
  • US11580323B2 patent drawing
  • US11580323B2 patent drawing
  • US11580323B2 patent drawing

AI summary

Embodiments are directed to data ingestion over a network. Raw data and integrated data associated with a plurality of separate data sources may be provided such that the raw data includes content associated with a plurality of subjects. Categorization models may be employed to categorize the raw data based on various features, such as, format, structure, data source, variability, volume, or associated entities. Matching models may be determined based on the categorization of the of the raw data, the integrated data and the content associated with the plurality of subjects. Matching models may generate a plurality of unified facts based on the raw data and the integrated data such that each unified fact is associated with a score associated with a quality of its match with a unified schema.