Data Lake Semantic Mapping Using Prolog Rules
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting semantic relationships between data fields in data lakes are manual, tedious, and prone to errors.
Innovation Solution
A computer system utilizing a descriptor generation unit and a rules engine based on deductive programming language (Prolog) to automatically generate descriptors and logical links between data fields, forming a knowledge base.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual extraction of semantic relationships is performed by experts, then accuracy can be maintained, but the process becomes tedious and time-consuming
Solution Approach 1:
The system enables self-service automated extraction of semantic relationships using Prolog-based inference engines that automatically analyze data fields and generate knowledge graphs without requiring manual expert intervention, thereby reducing time loss while maintaining accuracy through rule-based reasoning
Solution Approach 2:
The manual mechanical process of expert analysis is replaced by an automated computational system using Prolog programming language and inference engines that systematically process data fields, descriptors, and rules to extract semantic relationships automatically
2Reliability
If manual extraction by experts is used, then semantic relationships can be identified, but human errors are introduced
Solution Approach 1:
The system eliminates human error by implementing self-service automated extraction through Prolog-based inference engines that consistently apply predefined rules and logic without human intervention, ensuring reliable and reproducible results
Solution Approach 2:
The system incorporates feedback mechanisms where the inference engine continuously validates extracted relationships against the knowledge base and predefined rules, correcting inconsistencies and ensuring high reliability through iterative verification
3Productivity
If automated extraction is implemented using Prolog-based rules engine, then extraction speed and accuracy improve, but system complexity increases
Solution Approach 1:
The system segments the complex extraction task into distinct modular components: data field analysis, descriptor generation, rule application, and knowledge base construction, allowing each component to be independently managed and optimized while maintaining high productivity
Solution Approach 2:
The system introduces descriptors as intermediary elements that bridge raw data fields and semantic relationships, simplifying the overall process by creating intermediate representations that make the extraction more manageable and less complex
Data Source
Figure 1~2
AI summary
One of the aims of the present invention is to provide an objective and reproducible tool for automatically extracting the semantic relationships that may exist between the data fields in a data lake. To this end, the prior art proposes that said operation be manually carried out by experts in the field. However, such a solution is time-consuming and error-prone. Thus, the inventors propose using the capabilities of deductive programming languages such as Prolog to automatically create a knowledge base from descriptors of the data fields in the data lake. This solution is faster than in the prior art and less subject to errors.