Cross-Profiling Term Derivation for Data Lake Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data lake approaches face difficulties in efficiently defining business terms for both structured and unstructured data, leading to inefficient data search and analytics, as they typically define these terms separately without utilizing the definitions across data types.
Innovation Solution
A method and system that generate terms for structured and unstructured data files using respective profiling techniques, where the processor utilizes terms from structured data profiling to derive terms for unstructured data profiling, and vice versa, managing a term list to improve data search efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If business terms are defined separately for structured and unstructured data, then the definition process is simpler and more manageable, but the efficiency of data search and analytics deteriorates due to lack of cross-data-type term utilization
Solution Approach 1:
The patent merges the separate business term definition processes for structured and unstructured data into a unified system. The term derivation module combines terms from structured data profiling with unstructured data profiling results, allowing cross-utilization of term definitions across different data types while maintaining manageable complexity through modular architecture.
Solution Approach 2:
The patent creates a universal term list that serves multiple functions: it stores terms from structured data, terms from unstructured data, and derived terms that combine both. This universal term repository enables data search and analytics to benefit from cross-data-type term utilization, improving productivity without sacrificing ease of maintenance through the standardized term management interface.
2Productivity
If terms from structured data profiling are utilized in unstructured data profiling, then data search efficiency improves, but the complexity of the term derivation process increases
Solution Approach 1:
The patent introduces a term derivation module as an intermediary between structured data profiling and unstructured data profiling. This module receives terms from structured data, processes them according to derivation rules, and integrates them with unstructured data terms. The intermediary approach improves data search efficiency by enabling cross-term utilization while containing complexity within the modular term derivation component, making the system easier to maintain and update.
3Measurement precision
If a comprehensive term list integrating both structured and unstructured data terms is maintained, then the accuracy of business terms improves, but the time and resources required to manage the term list increase
Solution Approach 1:
The patent implements preliminary action by automatically deriving terms from structured data profiling results before unstructured data profiling is performed. The term derivation module pre-processes and organizes terms in advance, creating a foundation that improves business term accuracy. This preliminary organization reduces the time and resources needed for subsequent term list management, as the systematic structure enables more efficient updates and maintenance.
Data Source
AI summary
A method for performing data management that includes generating, using a processor of an agent server, terms for structured data files using structured data profiling and terms for unstructured data files using unstructured data profiling, wherein the structured data files and the unstructured data files are stored in a storage; and managing a term list, wherein the term list stores terms generated by the processor, wherein the processor utilizes terms generated through structured data profiling in deriving terms generated through unstructured data profiling.


