Semantic Knowledge Base Using NLP and RDF Ontologies

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for extracting information from documents into a knowledge base are limited by their reliance on statistical techniques and lack of semantic understanding, failing to accurately infer knowledge-based facts and address contextual and identity references.

Innovation Solution

The use of Natural Language Processing (NLP) to transform documents into Resource Description Framework (RDF) triples, leveraging ontologies to define knowledge stores and employing a semantic matcher that understands context-specific vocabularies, enabling the extraction of information into a highly resolved knowledge base.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If statistical techniques are used for information extraction, then automation is improved, but measurement precision deteriorates

Engineering Contradiction:
Improveautomation of information extractionVSAvoidaccuracy of knowledge-based facts
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary layer of semantic understanding and contextual analysis between the statistical extraction process and the final knowledge base. This intermediary uses ontologies, entity recognition, and relationship modeling to bridge the gap between automated statistical methods and precise semantic meaning, allowing automation while improving accuracy through multiple processing stages.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The information extraction process is divided into multiple segmented stages: initial statistical extraction, entity recognition, relationship identification, contextual analysis, and verification. Each stage processes specific aspects of the data independently, allowing the system to maintain automation while improving precision through progressive refinement at each segment.

Inventive Principle:
Principle #1Segmentation

2Productivity

If keyword search and match techniques are used, then productivity is improved, but manufacturing precision deteriorates

Engineering Contradiction:
Improvespeed of information extractionVSAvoidaccuracy of inferred knowledge
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system performs preliminary actions by pre-defining ontologies, entity types, and relationship schemas before the actual information extraction process. This preliminary structuring enables rapid keyword-based retrieval to be followed by automated semantic validation, maintaining high productivity while improving precision through pre-established knowledge frameworks that guide the extraction process.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If Natural Language Processing is used to understand information semantics, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvesemantic understanding accuracyVSAvoidcomplexity of processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by implementing NLP and semantic analysis only where needed in the information extraction process, rather than applying complex processing to all data uniformly. Keyword-based extraction is used for straightforward cases, while NLP techniques are applied selectively to ambiguous or context-dependent segments, maintaining precision where required while managing overall system complexity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11409780B2Semantic knowledge base
Publication Date: 2022.08.09 SEMANTIC TECH
  • US11409780B2 patent drawing
  • US11409780B2 patent drawing
  • US11409780B2 patent drawing

AI summary

A system for categorising and referencing a document using an electronic processing device, wherein: the electronic o processing device reviews the content of the document to identify structures within the document; wherein the identified structures are referenced against a library of structures stored in a database; wherein the document is categorised according to the conformance of the identified structures with those of the stored library of structures; and wherein the categorised structure is added to the stored library.