Co-occurrence Knowledge Base for Document Disambiguation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Searching for information about entities in large document repositories can be ambiguous due to unstructured text, leading to imprecise data analysis, and there is a need for an intelligent system to detect and record co-occurring features across documents to improve information retrieval.
Innovation Solution
A system and method for building a knowledge base of feature co-occurrences by crawling documents, extracting features, and aggregating their co-occurrences across a corpus using entity extraction, topic modeling, and event detection modules, with a knowledge base aggregator storing co-occurrences exceeding a predetermined threshold for disambiguation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If unstructured text is used in large document repositories, then storage capacity and information volume are improved, but search precision and data analysis accuracy deteriorate
Solution Approach 1:
The patent segments unstructured text into structured feature extracts (entities, topics, events, keywords) through automated extraction modules. This segmentation transforms the homogeneous unstructured text into heterogeneous structured data elements that can be precisely queried and analyzed, resolving the contradiction between storing large information volumes and maintaining search precision.
Solution Approach 2:
The patent introduces a knowledge base aggregator as an intermediary component that processes extracted features and creates co-occurrence relationships. This intermediary layer bridges the unstructured text and structured queries, enabling precise search by aggregating and indexing feature co-occurrences across documents without requiring direct manipulation of the original unstructured text.
2Measurement precision
If manual searching in large document repositories is performed, then search precision can be maintained, but time consumption and labor requirements increase
Solution Approach 1:
The patent implements automated feature extraction and co-occurrence aggregation that performs itself without human intervention. The system automatically crawls documents, extracts features, aggregates co-occurrences, and builds the knowledge base autonomously, eliminating manual searching while maintaining precision through automated intelligent processing.
Solution Approach 2:
The patent replaces manual mechanical searching with automated computational processes. Instead of human readers manually scanning documents, the system uses automated feature extraction algorithms and knowledge base aggregation mechanisms to perform search and analysis tasks, dramatically reducing time consumption while maintaining or improving precision.
3Measurement precision
If feature extraction and co-occurrence aggregation are performed on entire document corpora, then disambiguation accuracy is improved, but processing complexity and computational resources increase
Solution Approach 1:
The patent segments the document corpus into individual documents, extracts features from each document independently, and then aggregates co-occurrences at the feature level rather than processing entire corpora at once. This segmentation approach reduces processing complexity by handling smaller units independently while maintaining overall disambiguation accuracy through aggregated results.
Solution Approach 2:
The patent applies partial action by focusing extraction and aggregation on the most relevant features and their co-occurrences rather than processing every possible element in the entire corpus uniformly. The system identifies and processes only the necessary feature combinations that contribute to disambiguation, reducing computational complexity while maintaining accuracy.
Data Source
AI summary
A system for building a knowledge base of co-occurring features extracted from a document corpus is disclosed. The method includes a plurality of feature extraction software modules that may extract different features from each document in the corpus. The system may include a knowledge base aggregator module that may keep count of the co-occurrences of features in the different documents of a corpus and determine appropriate co-occurrences to store in a knowledge base.


