Automated Knowledge Graph Generation from Mixed Text Tables

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for transforming text documents into knowledge graphs are manual and time-consuming, especially when dealing with mixed structured and unstructured data, which hampers the accuracy and efficiency of semantic analysis.

Innovation Solution

A machine learning text classifier infers natural language semantics from tabular data in text documents, using a pipeline of state-of-the-art NLP techniques to automatically generate knowledge graphs, integrating named entity recognition, relation extraction, and coreference resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual transformation of text documents to knowledge graphs is performed, then accuracy of semantic analysis can be maintained through human judgment, but productivity is severely reduced due to time-consuming manual construction

Engineering Contradiction:
Improveaccuracy of semantic analysisVSAvoidspeed of knowledge graph construction
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automated self-service transformation of text documents into knowledge graphs using machine learning models that automatically perform entity recognition, relation extraction, and graph construction without requiring manual human intervention for each document, thereby resolving the contradiction between maintaining accuracy and improving productivity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of reading and constructing knowledge graphs with automated machine learning-based NLP systems that use trained models to infer semantics and generate knowledge graphs automatically, substituting human cognitive work with computational processes that maintain accuracy while dramatically improving speed

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If mixed structured and unstructured data is processed manually, then contextual understanding can be achieved, but device complexity increases due to the need for multiple processing approaches

Engineering Contradiction:
Improvecontextual understandingVSAvoidcomplexity of processing pipeline
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal NLP processing pipeline that handles both structured tabular data and unstructured text data through the same machine learning-based architecture, using multi-functional models that can process different data types uniformly, thereby reducing device complexity while maintaining contextual understanding

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses a composite approach combining multiple NLP techniques (entity recognition, relation extraction, coreference resolution) into an integrated processing pipeline that handles mixed data types, creating a unified system that manages complexity through composition of specialized components working together

Inventive Principle:
Principle #40Composite materials

3Productivity

If automated NLP techniques are used, then productivity is improved through fast processing, but measurement precision may deteriorate due to difficulty in correlating discrepant document parts

Engineering Contradiction:
Improvespeed of processingVSAvoidaccuracy of dependency analysis
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent incorporates feedback mechanisms where the NLP system iteratively refines its analysis by using extracted entities and relations to inform subsequent processing steps, allowing the system to correct errors and improve accuracy in correlating discrepant document parts while maintaining high processing speed

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary processing steps such as entity recognition and coreference resolution before main relation extraction, preparing the data in advance to facilitate more accurate dependency analysis in later stages, thereby maintaining both speed and precision in the automated processing pipeline

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250036594A1Transforming tables in documents into knowledge graphs using natural language processing
Publication Date: 2025.01.30 ORACLE INT CORP
  • US20250036594A1 patent drawing
  • US20250036594A1 patent drawing
  • US20250036594A1 patent drawing

AI summary

Herein from tabular data in a text document, a machine learning text classification pipeline infers natural language syntax and semantics to prepare the tabular data for graph analytics. In an embodiment, a computer infers a respective classification of each column in a table in a text document that contains natural language. Based on the classifications of the columns in the table, the vertices of the knowledge graph are automatically generated. Based on those column classifications and automatic analysis of a particular document portion that does not contain the table, the edges of the knowledge graph are automatically generated. The knowledge graph may be generated and operated as a property graph. Pipeline subsystems herein include column classification, edge type identification, and subject/object detection that provide sufficient semantic enrichment and context sensitivity to faster generate a more accurate knowledge graph than the state of the art.