Co-occurrence Knowledge Base for Document Disambiguation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Searching for information about entities in large document repositories can be ambiguous due to unstructured text, leading to imprecise data analysis, and there is a need for an intelligent system to detect and record co-occurring features across documents to improve information retrieval.

Innovation Solution

A system and method for building a knowledge base of feature co-occurrences by crawling documents, extracting features, and aggregating their co-occurrences across a corpus using entity extraction, topic modeling, and event detection modules, with a knowledge base aggregator storing co-occurrences exceeding a predetermined threshold for disambiguation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If unstructured text is used in large document repositories, then storage capacity and information volume are improved, but search precision and data analysis accuracy deteriorate

Engineering Contradiction:
Improveinformation volumeVSAvoidsearch precision
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments unstructured text into structured feature extracts (entities, topics, events, keywords) through automated extraction modules. This segmentation transforms the homogeneous unstructured text into heterogeneous structured data elements that can be precisely queried and analyzed, resolving the contradiction between storing large information volumes and maintaining search precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a knowledge base aggregator as an intermediary component that processes extracted features and creates co-occurrence relationships. This intermediary layer bridges the unstructured text and structured queries, enabling precise search by aggregating and indexing feature co-occurrences across documents without requiring direct manipulation of the original unstructured text.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual searching in large document repositories is performed, then search precision can be maintained, but time consumption and labor requirements increase

Engineering Contradiction:
Improvesearch precisionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements automated feature extraction and co-occurrence aggregation that performs itself without human intervention. The system automatically crawls documents, extracts features, aggregates co-occurrences, and builds the knowledge base autonomously, eliminating manual searching while maintaining precision through automated intelligent processing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical searching with automated computational processes. Instead of human readers manually scanning documents, the system uses automated feature extraction algorithms and knowledge base aggregation mechanisms to perform search and analysis tasks, dramatically reducing time consumption while maintaining or improving precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If feature extraction and co-occurrence aggregation are performed on entire document corpora, then disambiguation accuracy is improved, but processing complexity and computational resources increase

Engineering Contradiction:
Improvedisambiguation accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the document corpus into individual documents, extracts features from each document independently, and then aggregates co-occurrences at the feature level rather than processing entire corpora at once. This segmentation approach reduces processing complexity by handling smaller units independently while maintaining overall disambiguation accuracy through aggregated results.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by focusing extraction and aggregation on the most relevant features and their co-occurrences rather than processing every possible element in the entire corpus uniformly. The system identifies and processes only the necessary feature combinations that contribute to disambiguation, reducing computational complexity while maintaining accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9922032B2Featured co-occurrence knowledge base from a corpus of documents
Publication Date: 2018.03.20 FINCH COMPUTING LLC
  • US9922032B2 patent drawing
  • US9922032B2 patent drawing
  • US9922032B2 patent drawing

AI summary

A system for building a knowledge base of co-occurring features extracted from a document corpus is disclosed. The method includes a plurality of feature extraction software modules that may extract different features from each document in the corpus. The system may include a knowledge base aggregator module that may keep count of the co-occurrences of features in the different documents of a corpus and determine appropriate co-occurrences to store in a knowledge base.