Contextual Vocabulary Extraction System for Document Summarization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional vocabulary learning methods rely on memorization without context, leading to minimal memory retrieval cues and quick forgetting, and often fail to connect words to relevant subjects, while summarizing documents is a tedious and subjective process that lacks accuracy.

Innovation Solution

A system and method that analyzes documents to identify unfamiliar words and their contexts, providing a summary of document relevance and adherence to vocabulary norms by using statistical indicators and user-selectable views, such as sort order, word size, and color, to convey information about word frequency and subject matter association.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional memorization methods are used for vocabulary learning, then students can learn words and definitions, but memory retrieval cues are minimal and information is quickly forgotten

Engineering Contradiction:
Improvememory retentionVSAvoidvocabulary recall
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments vocabulary learning into contextualized units by extracting words along with their surrounding text contexts from documents. Instead of isolated word-definition pairs, the system divides content into meaningful contextual segments that preserve the natural usage environment, enabling students to learn vocabulary within authentic linguistic frameworks and improve retention through contextual associations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces textual context as an intermediary element between the word and the student's memory. By presenting words embedded in their original document contexts, the system creates a mediating layer that provides additional semantic cues and associations, serving as a retrieval scaffold that bridges the gap between isolated vocabulary and meaningful memory storage.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If vocabulary is compiled for the subject of vocabulary without subject associations, then words are organized alphabetically or by frequency, but the words are unrelated to relevant subjects and useless for learning other subjects

Engineering Contradiction:
Improvevocabulary organizationVSAvoidsubject relevance
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent makes the vocabulary extraction system multi-functional by enabling it to serve both vocabulary learning purposes and subject matter exploration. The same contextualized word extracts can be used for teaching vocabulary while simultaneously exposing students to domain-specific content and terminology, allowing the system to adapt to different educational objectives without requiring separate compilation processes for different subjects.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary subject classification and word extraction from relevant documents before the vocabulary is presented to students. By pre-processing documents to identify subject matter and extract pertinent vocabulary in context, the system prepares subject-relevant vocabulary lists in advance, ensuring that students encounter words that are both linguistically useful and thematically appropriate for their studies.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If manual compilation of document synopses is performed, then summaries can be created, but the process is tedious and the synopses are subjective and lack accuracy

Engineering Contradiction:
Improvedocument summarization accuracyVSAvoidsummarization time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent implements self-service document summarization by enabling the system to automatically extract and present relevant vocabulary and contexts from documents without human intervention. The computational system performs the summarization function autonomously by identifying key terms, extracting their contexts, and organizing them into concise representations, eliminating the need for manual synopsis compilation while maintaining objective accuracy based on actual document content.

Inventive Principle:
Principle #25Self-service

4Ease of operation

If authors do not adhere to vocabulary norms, then writing may be more spontaneous, but overuse of certain words ruins the flow of the document and distracts the reader

Engineering Contradiction:
Improvewriting spontaneityVSAvoiddocument quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements feedback mechanisms for vocabulary norm adherence by analyzing document text and identifying words that are overused relative to established norms. The system provides authors with quantitative feedback about their vocabulary usage patterns, highlighting specific words that appear excessively frequently and potentially disrupting document flow, enabling authors to revise and improve their writing while maintaining spontaneous composition.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8311808B2System and method for advancement of vocabulary skills and for identifying subject matter of a document
Publication Date: 2012.11.13 THINKMAP
  • US8311808B2 patent drawing
  • US8311808B2 patent drawing
  • US8311808B2 patent drawing

AI summary

A system and method for providing vocabulary information includes one or more computer processors that, for each of a plurality of words of a text, determine a relevance of the word to the text, and, for each of at least a subset of the plurality of words, output an indication of the respective determined relevance of the word to the text, where, for each of the plurality of words, the determination includes comparing a frequency of the word in the text to a frequency threshold.