Probabilistic Generative Model for Relational Schema Annotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The generation of domain knowledge graphs is challenging due to the scarcity of domain experts and the time-consuming process of building them from scratch, with conventional techniques being ineffective for complex schemas and failing to identify domain-specific terms and relationships.
Innovation Solution
A method and system that utilize a pre-defined probabilistic generative model to generate annotations and field-names for a relational schema, trained using either supervised Maximum Likelihood Estimation or unsupervised Expectation Maximization techniques, based on annotated or unannotated relational data, to create new field-names and annotations through stochastic generative and probabilistic inference processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If domain experts are used to build domain knowledge graph from scratch, then the quality and accuracy of domain concepts and relationships are improved, but the time consumption and resource requirements increase significantly
Solution Approach 1:
The system performs preliminary action by pre-defining a probabilistic generative model with domain knowledge structures before actual annotation tasks. The model is pre-trained with domain concepts and relationships, so when new schema elements need annotation, the system can quickly generate annotations using the pre-established knowledge framework rather than building from scratch each time.
Solution Approach 2:
The system uses copying by generating annotations for new schema elements based on existing annotated examples and domain knowledge patterns. Instead of requiring domain experts to manually create all annotations, the system copies proven annotation patterns and adapts them to new contexts through the probabilistic generative model, significantly reducing the need for expert intervention.
2Ease of manufacture
If conventional techniques annotate tables or fields with a single vertex/concept, then the annotation process is simple, but the effectiveness decreases for complex schemas with multiple domain terms
Solution Approach 1:
The system applies segmentation by breaking down complex schema elements into multiple domain concept vertices and relationships rather than using a single vertex. The probabilistic generative model generates structured annotations that can include multiple connected concepts, allowing complex fields to be represented by networks of related domain terms while maintaining systematic processing through the model framework.
Solution Approach 2:
The annotation system uses composite materials analogy by combining multiple domain concept vertices and relationship edges into composite annotation structures. Instead of single-concept annotations, the system creates composite annotations that integrate multiple domain terms and their relationships, providing more comprehensive and accurate representations of complex schema elements.
3Ease of manufacture
If existing techniques fail to identify domain-specific terms corresponding to concepts, then the implementation is straightforward, but the domain knowledge representation becomes inaccurate
Solution Approach 1:
The system introduces an intermediary probabilistic generative model that acts as a bridge between raw schema data and domain knowledge concepts. This intermediary model processes schema elements and maps them to appropriate domain terms by learning from training data, thereby improving term identification accuracy without requiring direct manual mapping by domain experts for each element.
Solution Approach 2:
The system replaces the mechanical/manual process of domain expert term identification with an automated probabilistic inference mechanism. The generative model uses statistical patterns and learned relationships to automatically identify and map domain-specific terms, substituting the manual mechanical process with an automated computational system that maintains or improves accuracy.
Data Source
AI summary
This disclosure relates generally to generating annotations and field-names for a relational schema. Typically, most domains have relational database (RDB) system built for them instead of domain ontologies and usually linguistic information of the schema is not used to recover the domain terms. The disclosed method and system facilitate generating annotations and field-names for a relational schema, while considering the linguistic information of a schema by using a trained model, trained through a proposed training technique. The trained model comprises of at least one knowledge graph and a set of associated parameters. The trained model is further used to perform a plurality of tasks, wherein the plurality of tasks include generating a plurality of new fieldnames for a relational schema through a stochastic generative process and for generating a new annotation for a fieldname of a relational schema through a probabilistic inference technique.


