Ontological Question Clustering from Document Assertions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in finding specific information or forming a cohesive understanding of topics due to dispersed assertions across multiple electronic documents, which are not effectively organized into meaningful clusters.
Innovation Solution
The method involves analyzing documents to identify entities and relationships, generating assertions, inverting these assertions to create questions, and clustering them around concepts and topics, with a combined graph facilitating traversal among topics, questions, assertions, and documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If information is stored across multiple electronic documents, then the quantity of information is increased, but the difficulty of finding specific information and forming cohesive understanding increases
Solution Approach 1:
The patent segments information into discrete assertions that can be independently extracted and organized from multiple documents. Each assertion represents a factual statement that can be separately identified, tagged, and clustered, making the overall information more searchable and manageable despite being distributed across many documents.
Solution Approach 2:
The patent introduces an intermediary assertion layer between raw documents and user queries. Assertions serve as mediators that connect dispersed document information to search queries, enabling cohesive understanding by organizing facts into structured units that can be systematically retrieved and synthesized.
2Loss of information
If assertions are extracted from multiple documents, then the completeness of knowledge is improved, but the complexity of organizing assertions into meaningful clusters increases
Solution Approach 1:
The patent segments the organization process into distinct stages: extraction of individual assertions, tagging with metadata, and clustering into topic groups. This segmented approach makes the overall organization task more manageable by breaking down the complex problem into smaller, systematic steps.
Solution Approach 2:
The patent changes the parameters of assertion organization by introducing metadata tags and clustering criteria. Assertions are organized based on configurable parameters such as topic categories, document sources, and semantic relationships, allowing flexible and systematic organization without requiring complex manual curation.
3Loss of information
If facts are identified and organized into clusters, then the level of understanding is improved, but the time and resources required for organization increase
Solution Approach 1:
The patent performs preliminary organization of assertions into clusters and topics before users need to access the information. By pre-processing documents to extract, tag, and cluster assertions in advance, the system reduces the time required for users to gain understanding, as the organizational work is completed beforehand.
Solution Approach 2:
The patent implements self-service organization through automated assertion extraction and clustering algorithms. The system automatically organizes facts into meaningful clusters without requiring manual intervention, reducing time and resource requirements while maintaining high levels of understanding and coherence.
Data Source
AI summary
Electronic documents are analyzed to identify assertions, which are inverted to generate questions that may be answered by the assertions. A document or a corpus of electronic documents may be analyzed to identify entities and relationships among entities within the text of the document(s). Assertions are identified based on the entities and relationships among the entities. Each assertion represents a fact about an entity, and a group of assertions represents a summary of the document or document corpus. The assertions are inverted to generate questions that may be answered by the assertions. The questions may be further analyzed to identify relevant concepts and topics and to cluster the questions around the concepts and topics. A combined graph may also be generated that facilitates traversal among topics, concepts, questions, assertions, document summaries, and documents.


