Schema Discovery for Unstructured Data Using Sidecar Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in discovering proper schemas for unstructured datasets, particularly in environments with large volumes of data and multiple tenants, where the diversity of unstructured data from various sources complicates the conversion from OLTP to OLAP, and existing methods struggle to efficiently analyze and utilize Big Data effectively.
Innovation Solution
A method that uses a sidecar to collect schema discovery rules during data conversion, generates multiple schemas for different tenants, exports unstructured data to SQL databases using ETL, monitors usage data, and optimizes schema discovery using machine learning to apply schemas with high usage to other tenants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple schemas are generated for different tenants using schema discovery rules, then the adaptability to handle diverse unstructured data from various sources is improved, but the device complexity increases due to the need for automated schema generation and management systems
Solution Approach 1:
The system performs self-service by automatically discovering schemas from unstructured data without requiring manual intervention. The schema generation process is automated through machine learning models that analyze data patterns and generate appropriate schemas autonomously, reducing the need for complex manual configuration while maintaining high adaptability to diverse data sources
Solution Approach 2:
The system changes parameters by dynamically adjusting schema structures based on the characteristics of incoming unstructured data. Different schemas are generated with varying parameters (data types, relationships, constraints) to accommodate the diversity of data sources, allowing the system to adapt its schema parameters automatically rather than using a fixed complex structure
2Productivity
If schema discovery is optimized using usage data and machine learning, then the productivity of data processing is improved, but the device complexity increases due to monitoring and optimization mechanisms
Solution Approach 1:
The system implements feedback by continuously monitoring usage data from the data processing operations and using this information to optimize schema discovery. The usage patterns are fed back into the machine learning models to improve schema generation accuracy and performance over time, creating a self-improving system that increases productivity without requiring proportionally complex monitoring infrastructure
Solution Approach 2:
The system performs preliminary actions by pre-generating schemas using machine learning models before actual data processing occurs. This advance preparation reduces processing time during production operations, improving productivity. The schemas are created in advance based on historical data patterns, eliminating the need for complex real-time schema generation during high-volume processing
3Ease of operation
If unstructured data is converted to structured SQL databases using ETL, then the ease of operation for data analysis is improved, but the loss of information may occur during the transformation process
Solution Approach 1:
The system performs preliminary schema discovery and validation before the ETL transformation process. By pre-defining the target schema structure based on the unstructured data characteristics, the system ensures that all relevant information is captured during transformation. This preliminary preparation reduces information loss by establishing appropriate data mappings and validation rules before the actual conversion to structured SQL databases
Solution Approach 2:
The system changes parameters by adjusting transformation settings and data mapping parameters to preserve information during ETL. Different parameter configurations are used to handle various data types and relationships, ensuring that the structured representation in SQL databases maintains the essential characteristics and information from the original unstructured data while improving ease of analysis
Data Source
AI summary
An embodiment for analyzing and tracking data flow to determine proper schemas for unstructured data. The embodiment may automatically use a sidecar to collect schema discovery rules during conversion of raw data to unstructured data. The embodiment may automatically generate multiple schemas for different tenants using the collected schema discovery rules. The embodiment may automatically use ETL to export unstructured data to SQL databases with the generated multiple schemas for the different tenants. The embodiment may automatically monitor usage data of the SQL databases and collect the usage data. The embodiment may automatically optimize schema discovery using the collected usage data. The embodiment may automatically discover schemas with hot usage and apply the discovered schemas with hot usage to other tenants for consumption and further monitoring.


