Data Lake Semantic Mapping Using Prolog Rules

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting semantic relationships between data fields in data lakes are manual, tedious, and prone to errors.

Innovation Solution

A computer system utilizing a descriptor generation unit and a rules engine based on deductive programming language (Prolog) to automatically generate descriptors and logical links between data fields, forming a knowledge base.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual extraction of semantic relationships is performed by experts, then accuracy can be maintained, but the process becomes tedious and time-consuming

Engineering Contradiction:
Improveextraction accuracyVSAvoidextraction time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service automated extraction of semantic relationships using Prolog-based inference engines that automatically analyze data fields and generate knowledge graphs without requiring manual expert intervention, thereby reducing time loss while maintaining accuracy through rule-based reasoning

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical process of expert analysis is replaced by an automated computational system using Prolog programming language and inference engines that systematically process data fields, descriptors, and rules to extract semantic relationships automatically

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual extraction by experts is used, then semantic relationships can be identified, but human errors are introduced

Engineering Contradiction:
Improveextraction reliabilityVSAvoidhuman error
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The system eliminates human error by implementing self-service automated extraction through Prolog-based inference engines that consistently apply predefined rules and logic without human intervention, ensuring reliable and reproducible results

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms where the inference engine continuously validates extracted relationships against the knowledge base and predefined rules, correcting inconsistencies and ensuring high reliability through iterative verification

Inventive Principle:
Principle #23Feedback

3Productivity

If automated extraction is implemented using Prolog-based rules engine, then extraction speed and accuracy improve, but system complexity increases

Engineering Contradiction:
Improveextraction productivityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex extraction task into distinct modular components: data field analysis, descriptor generation, rule application, and knowledge base construction, allowing each component to be independently managed and optimized while maintaining high productivity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces descriptors as intermediary elements that bridge raw data fields and semantic relationships, simplifying the overall process by creating intermediate representations that make the extraction more manageable and less complex

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4457646B1Data lake management system and method
Publication Date: 2026.02.18 ALTEN
  • EP4457646B1 patent drawingFigure 1~2

AI summary

One of the aims of the present invention is to provide an objective and reproducible tool for automatically extracting the semantic relationships that may exist between the data fields in a data lake. To this end, the prior art proposes that said operation be manually carried out by experts in the field. However, such a solution is time-consuming and error-prone. Thus, the inventors propose using the capabilities of deductive programming languages such as Prolog to automatically create a knowledge base from descriptors of the data fields in the data lake. This solution is faster than in the prior art and less subject to errors.