Knowledge Expansion System for Document Structure Graphs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Inter-word relationship information lacks the knowledge desired by users, as existing techniques fail to expand knowledge effectively from written documents with complex structures like bulleted lists or tables, limiting the extraction of new knowledge.
Innovation Solution
A knowledge expansion system that extracts subgraphs from document structure graphs based on inter-word relationship information, creates rules for matching subgraph structures, and adds new knowledge to the inter-word relationship information, enabling expansion beyond traditional sentence-based knowledge extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional sentence-based knowledge extraction is used, then the extraction process is simple, but the knowledge expansion capability is limited and cannot handle complex document structures
Solution Approach 1:
The patent segments the document structure into hierarchical components (sections, subsections, paragraphs, sentences) and represents them as a document structure graph. This segmentation enables the system to handle complex document structures by breaking them down into manageable units that can be processed individually while maintaining their hierarchical relationships.
Solution Approach 2:
The patent transitions from traditional sentence-based extraction to a multi-dimensional approach by incorporating document structure hierarchy. The document structure graph adds structural dimensionality, allowing knowledge extraction to consider not only textual content but also the hierarchical organization and relationships between different document elements.
2Loss of information
If knowledge extraction is limited to sentences, then the processing is straightforward, but the ability to utilize document structure information is lost
Solution Approach 1:
The patent extracts document structure information by identifying and separating structural elements (headings, sections, paragraphs) from the textual content. The document structure graph extraction unit specifically extracts the hierarchical structure as a separate representation, preserving structural information that would be lost in traditional sentence-based processing.
Solution Approach 2:
The document structure graph serves multiple functions: it represents the hierarchical organization of the document, enables navigation between different structural levels, and provides the basis for structure-aware knowledge extraction. This multi-functional representation maximizes the utility of document structure information without requiring separate processing mechanisms.
3Quantity of substance
If the inter-word relationship information is expanded using general techniques, then some knowledge can be added, but the expansion is insufficient for documents with complex structures like bulleted lists or tables
Solution Approach 1:
The document structure graph acts as an intermediary between the raw document content and the knowledge extraction process. It mediates by providing a structured representation that captures hierarchical relationships, allowing the knowledge extraction to reliably identify and expand inter-word relationships based on the documented structure rather than relying solely on textual patterns.
Data Source
AI summary
A subgraph extraction means 71 extracts, from a document structure graph indicating a document structure, a subgraph as a part of the document structure graph on the basis of inter-word relationship information indicating a relationship between a word and a word. A rule creation means 72 creates a rule for extracting a subgraph having the same structure as the subgraph from the document structure graph. A knowledge addition means 73 extracts a subgraph from the document structure graph in accordance with the rule and adds the information indicated by the subgraph to the inter-word relationship information.


