Data Harmonization Platform with Metadata Mapping and ML

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Organizations face challenges in collecting, integrating, and analyzing large volumes of structured and unstructured data from multiple heterogeneous sources, as existing technologies struggle to efficiently process and harmonize data for business intelligence and big data systems, leading to difficulties in accessing and utilizing data for actionable insights.

Innovation Solution

A flexible and scalable data-to-decision platform that uses advanced software to aggregate, preprocess, and harmonize data from various sources through machine learning techniques, including pattern recognition, supervised learning, and metadata-based mapping, to transform and standardize data, enabling cost-effective data integration and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is collected from multiple heterogeneous sources, then data volume and variety increase, but data integration complexity and processing time increase

Engineering Contradiction:
Improvedata volumeVSAvoiddata integration complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments data processing into distinct modules: data collection from multiple sources, data preprocessing/cleansing, pattern recognition for feature extraction, supervised learning for classification, and metadata-based mapping. Each module handles specific aspects of data harmonization independently, reducing overall integration complexity while managing large data volumes from heterogeneous sources.

Inventive Principle:
Principle #1Segmentation

2Stability of the object's composition

If data from multiple sources is harmonized, then data consistency improves, but processing time and computational resources increase

Engineering Contradiction:
Improvedata consistencyVSAvoiddata processing time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The system performs preliminary data preprocessing and cleansing before main analysis, including removing duplicates, handling missing values, and standardizing formats. Metadata-based mapping is established in advance to define relationships between data elements from different sources, reducing processing time during actual data harmonization while ensuring consistency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces manual data harmonization processes with automated machine learning techniques, including pattern recognition algorithms and supervised learning models. These computational methods automatically identify and resolve inconsistencies across data sources, maintaining data consistency while significantly reducing processing time compared to manual approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If structured data is converted to standardized format, then data usability improves, but information loss and context loss occur

Engineering Contradiction:
Improvedata usabilityVSAvoidoriginal meaning and context
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system introduces metadata as an intermediary layer between raw data and standardized format. Metadata preserves original data characteristics, sources, and contextual information while enabling standardized processing. This intermediary structure allows data to be converted to usable standardized formats without losing original meaning or context, as metadata carries this information separately.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system nests multiple levels of data representation: raw data structures contain embedded metadata about original format and context, which is then nested within standardized data structures. This nested approach allows standardized processing at the outer level while preserving original information in inner layers, maintaining both usability and information integrity.

Inventive Principle:
Principle #7Nested doll (Nesting)

4Productivity

If automated data processing is implemented, then productivity increases, but system complexity and automation extent increase

Engineering Contradiction:
Improvedata processing productivityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements universal data processing components that handle multiple data types and sources through the same pipeline. The metadata-based mapping framework and machine learning models are designed to work across diverse data formats and sources, increasing productivity through automation while managing system complexity by reusing the same components for different tasks rather than creating separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11893008B1System and method for automated data harmonization
Publication Date: 2024.02.06 FRACTAL ANALYTICS PTE LTD
  • US11893008B1 patent drawing
  • US11893008B1 patent drawing
  • US11893008B1 patent drawing

AI summary

Systems and methods are provided to aggregate and analyze data from a plurality of data sources. The system may obtain data from a plurality of data sources. The system may also transform data from each of the plurality of data sources into a format that is compatible for combining the data from the plurality of data sources. The system uses a data harmonization module to organize, classify, analyze and thus relate previously unrelated data stored in multiple databases and/or associated with different organizations. The system can generate and publish master data of a plurality of business entities.