Automated Knowledge Graph Generation from Mixed Text Tables
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for transforming text documents into knowledge graphs are manual and time-consuming, especially when dealing with mixed structured and unstructured data, which hampers the accuracy and efficiency of semantic analysis.
Innovation Solution
A machine learning text classifier infers natural language semantics from tabular data in text documents, using a pipeline of state-of-the-art NLP techniques to automatically generate knowledge graphs, integrating named entity recognition, relation extraction, and coreference resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual transformation of text documents to knowledge graphs is performed, then accuracy of semantic analysis can be maintained through human judgment, but productivity is severely reduced due to time-consuming manual construction
Solution Approach 1:
The system enables automated self-service transformation of text documents into knowledge graphs using machine learning models that automatically perform entity recognition, relation extraction, and graph construction without requiring manual human intervention for each document, thereby resolving the contradiction between maintaining accuracy and improving productivity
Solution Approach 2:
The patent replaces the manual mechanical process of reading and constructing knowledge graphs with automated machine learning-based NLP systems that use trained models to infer semantics and generate knowledge graphs automatically, substituting human cognitive work with computational processes that maintain accuracy while dramatically improving speed
2Reliability
If mixed structured and unstructured data is processed manually, then contextual understanding can be achieved, but device complexity increases due to the need for multiple processing approaches
Solution Approach 1:
The patent implements a universal NLP processing pipeline that handles both structured tabular data and unstructured text data through the same machine learning-based architecture, using multi-functional models that can process different data types uniformly, thereby reducing device complexity while maintaining contextual understanding
Solution Approach 2:
The system uses a composite approach combining multiple NLP techniques (entity recognition, relation extraction, coreference resolution) into an integrated processing pipeline that handles mixed data types, creating a unified system that manages complexity through composition of specialized components working together
3Productivity
If automated NLP techniques are used, then productivity is improved through fast processing, but measurement precision may deteriorate due to difficulty in correlating discrepant document parts
Solution Approach 1:
The patent incorporates feedback mechanisms where the NLP system iteratively refines its analysis by using extracted entities and relations to inform subsequent processing steps, allowing the system to correct errors and improve accuracy in correlating discrepant document parts while maintaining high processing speed
Solution Approach 2:
The system performs preliminary processing steps such as entity recognition and coreference resolution before main relation extraction, preparing the data in advance to facilitate more accurate dependency analysis in later stages, thereby maintaining both speed and precision in the automated processing pipeline
Data Source
AI summary
Herein from tabular data in a text document, a machine learning text classification pipeline infers natural language syntax and semantics to prepare the tabular data for graph analytics. In an embodiment, a computer infers a respective classification of each column in a table in a text document that contains natural language. Based on the classifications of the columns in the table, the vertices of the knowledge graph are automatically generated. Based on those column classifications and automatic analysis of a particular document portion that does not contain the table, the edges of the knowledge graph are automatically generated. The knowledge graph may be generated and operated as a property graph. Pipeline subsystems herein include column classification, edge type identification, and subject/object detection that provide sufficient semantic enrichment and context sensitivity to faster generate a more accurate knowledge graph than the state of the art.


