Semantic Triple Filtering for Knowledge Base Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for creating and augmenting knowledge bases are inefficient due to the transmission of large amounts of poor quality and erroneous semantic triples, leading to increased bandwidth usage and processing requirements.
Innovation Solution
A computer-implemented method for generating and filtering semantic triples from unstructured text, involving taxonomic relation resolution, noun-centric relation extraction, and relevance scoring, to reduce the number and improve the quality of triples included in the knowledge base.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional methods transmit all extracted semantic triples to build a knowledge base, then the knowledge base can be populated with comprehensive data, but the bandwidth usage increases and transmission speed decreases due to the large volume of poor quality and erroneous triples
Solution Approach 1:
The system performs preliminary filtering and quality assessment of semantic triples before transmission to the knowledge base. A filtering module evaluates each triple's quality metrics (completeness, consistency, reliability) and selects only high-quality triples for transmission, thereby reducing the total volume of data transmitted while maintaining knowledge base population effectiveness
Solution Approach 2:
The system changes the parameter of triple quality by introducing quality thresholds and filtering criteria. By setting minimum quality standards for subject, predicate, and object elements, the system transforms the transmission dataset from containing all extracted triples to containing only those meeting quality parameters, thus reducing bandwidth usage without compromising knowledge base comprehensiveness
2Quantity of substance
If conventional methods include all extracted semantic triples in the knowledge base, then the knowledge base achieves high coverage, but processing requirements and storage needs increase due to erroneous and low quality triples
Solution Approach 1:
The system performs preliminary filtering and quality assessment of semantic triples before transmission to the knowledge base. A filtering module evaluates each triple's quality metrics (completeness, consistency, reliability) and selects only high-quality triples for transmission, thereby reducing the total volume of data transmitted while maintaining knowledge base population effectiveness
Solution Approach 2:
The system extracts and removes erroneous and low-quality triples from the dataset before knowledge base population. By identifying and excluding triples with missing elements, inconsistent relationships, or low confidence scores, the system separates useful data from harmful data, reducing processing requirements while maintaining high coverage of valid knowledge
3Loss of information
If the system transmits all semantic triples without filtering, then no information is lost, but bandwidth usage increases and transmission efficiency decreases
Solution Approach 1:
The system changes the parameter of triple quality by introducing quality thresholds and filtering criteria. By setting minimum quality standards for subject, predicate, and object elements, the system transforms the transmission dataset from containing all extracted triples to containing only those meeting quality parameters, thus reducing bandwidth usage without compromising knowledge base comprehensiveness
Data Source
AI summary
The present disclosure relates to a computer-implemented method of verifying a semantic triple generated for building a knowledge base including data patterns defining concepts associated with semantic triples derived from unstructured text. The method includes providing the semantic triple to a user interface, the semantic triple including a subject, an object, and a relation. The method also includes receiving, from the user interface, an acceptance or a rejection of the subject, the object, and the relation as relevant or not to the knowledge base. The method also includes transmitting the semantic triple for inclusion as a data pattern in the knowledge base in the event that all of the subject, the object, and the relation, have been accepted.


