Schema Discovery for Unstructured Data Using Sidecar Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in discovering proper schemas for unstructured datasets, particularly in environments with large volumes of data and multiple tenants, where the diversity of unstructured data from various sources complicates the conversion from OLTP to OLAP, and existing methods struggle to efficiently analyze and utilize Big Data effectively.

Innovation Solution

A method that uses a sidecar to collect schema discovery rules during data conversion, generates multiple schemas for different tenants, exports unstructured data to SQL databases using ETL, monitors usage data, and optimizes schema discovery using machine learning to apply schemas with high usage to other tenants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple schemas are generated for different tenants using schema discovery rules, then the adaptability to handle diverse unstructured data from various sources is improved, but the device complexity increases due to the need for automated schema generation and management systems

Engineering Contradiction:
Improveadaptability to handle diverse unstructured dataVSAvoidcomplexity of automated schema generation system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically discovering schemas from unstructured data without requiring manual intervention. The schema generation process is automated through machine learning models that analyze data patterns and generate appropriate schemas autonomously, reducing the need for complex manual configuration while maintaining high adaptability to diverse data sources

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes parameters by dynamically adjusting schema structures based on the characteristics of incoming unstructured data. Different schemas are generated with varying parameters (data types, relationships, constraints) to accommodate the diversity of data sources, allowing the system to adapt its schema parameters automatically rather than using a fixed complex structure

Inventive Principle:
Principle #35Parameter changes

2Productivity

If schema discovery is optimized using usage data and machine learning, then the productivity of data processing is improved, but the device complexity increases due to monitoring and optimization mechanisms

Engineering Contradiction:
Improveproductivity of data processingVSAvoidcomplexity of monitoring and optimization system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements feedback by continuously monitoring usage data from the data processing operations and using this information to optimize schema discovery. The usage patterns are fed back into the machine learning models to improve schema generation accuracy and performance over time, creating a self-improving system that increases productivity without requiring proportionally complex monitoring infrastructure

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary actions by pre-generating schemas using machine learning models before actual data processing occurs. This advance preparation reduces processing time during production operations, improving productivity. The schemas are created in advance based on historical data patterns, eliminating the need for complex real-time schema generation during high-volume processing

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If unstructured data is converted to structured SQL databases using ETL, then the ease of operation for data analysis is improved, but the loss of information may occur during the transformation process

Engineering Contradiction:
Improveease of operation for data analysisVSAvoidinformation loss during transformation
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system performs preliminary schema discovery and validation before the ETL transformation process. By pre-defining the target schema structure based on the unstructured data characteristics, the system ensures that all relevant information is captured during transformation. This preliminary preparation reduces information loss by establishing appropriate data mappings and validation rules before the actual conversion to structured SQL databases

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters by adjusting transformation settings and data mapping parameters to preserve information during ETL. Different parameter configurations are used to handle various data types and relationships, ensuring that the structured representation in SQL databases maintains the essential characteristics and information from the original unstructured data while improving ease of analysis

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11947561B2Heterogeneous schema discovery for unstructured data
Publication Date: 2024.04.02 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11947561B2 patent drawing
  • US11947561B2 patent drawing
  • US11947561B2 patent drawing

AI summary

An embodiment for analyzing and tracking data flow to determine proper schemas for unstructured data. The embodiment may automatically use a sidecar to collect schema discovery rules during conversion of raw data to unstructured data. The embodiment may automatically generate multiple schemas for different tenants using the collected schema discovery rules. The embodiment may automatically use ETL to export unstructured data to SQL databases with the generated multiple schemas for the different tenants. The embodiment may automatically monitor usage data of the SQL databases and collect the usage data. The embodiment may automatically optimize schema discovery using the collected usage data. The embodiment may automatically discover schemas with hot usage and apply the discovered schemas with hot usage to other tenants for consumption and further monitoring.