Document Understanding Support Apparatus for Dynamic Word Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in efficiently extracting and understanding relevant information from electronic documents, as the importance of words varies by user and document type, and changes over time, leading to unsatisfactory support in information processing.

Innovation Solution

A document understanding support apparatus that uses machine-learning to extract and output relationships between words, employing a word extraction condition learning device, word extractor, word relationship extraction condition learning device, and output device to identify and present essential words and relationships, adapting to user needs and document types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If only words are extracted and displayed from electronic documents, then information extraction efficiency is improved, but understanding of document contents deteriorates

Engineering Contradiction:
Improveinformation extraction efficiencyVSAvoiddocument content understanding
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent segments the extracted information into two distinct categories: individual words and word relationships. The word extraction unit identifies important words, while the word relationship extraction unit identifies relationships between those words. This segmentation allows the system to present both isolated key terms and their contextual connections, thereby maintaining information extraction efficiency while preserving document content understanding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges the output of word extraction and word relationship extraction into a unified presentation. By combining both words and their relationships in the output, the system achieves both efficient information extraction and comprehensive content understanding, resolving the contradiction between these two objectives.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If extraction conditions are fixed for specific users and document types, then extraction precision is improved, but adaptability to different users and situations deteriorates

Engineering Contradiction:
Improveextraction precisionVSAvoidadaptability to different users and situations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic extraction conditions that can be adjusted based on user feedback and different usage scenarios. The system learns from user interactions and modifies its extraction criteria over time, allowing it to maintain high precision for specific users while simultaneously adapting to new users and situations. This dynamic adjustment mechanism resolves the contradiction between fixed precision and flexible adaptability.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10635897B2Document understanding support apparatus, document understanding support method, non-transitory storage medium
Publication Date: 2020.04.28 KK TOSHIBA
  • US10635897B2 patent drawing
  • US10635897B2 patent drawing
  • US10635897B2 patent drawing

AI summary

A document understanding support apparatus as one embodiment of the present invention includes a word extraction condition learning device, a word extractor, a word relationship extraction condition learning device, a word relationship extractor, and an output device. The word extraction condition learning device creates a word extraction condition for extracting words from a target electronic document by machine-learning based on feature values assigned to respective words. The word extractor extracts words satisfying the word extraction condition. The word relationship extraction condition learning device creates a word relationship extraction condition for extracting word relationships from the target electronic document by machine-learning based on feature values with respect to extraction target word relationships. The word relationship extractor extracts a word relationship satisfying the word relationship extraction condition. The output device outputs at least either the extracted words or the extracted word relationship.