Ontological Question Clustering from Document Assertions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in finding specific information or forming a cohesive understanding of topics due to dispersed assertions across multiple electronic documents, which are not effectively organized into meaningful clusters.

Innovation Solution

The method involves analyzing documents to identify entities and relationships, generating assertions, inverting these assertions to create questions, and clustering them around concepts and topics, with a combined graph facilitating traversal among topics, questions, assertions, and documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If information is stored across multiple electronic documents, then the quantity of information is increased, but the difficulty of finding specific information and forming cohesive understanding increases

Engineering Contradiction:
Improvequantity of informationVSAvoidease of finding specific information
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments information into discrete assertions that can be independently extracted and organized from multiple documents. Each assertion represents a factual statement that can be separately identified, tagged, and clustered, making the overall information more searchable and manageable despite being distributed across many documents.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary assertion layer between raw documents and user queries. Assertions serve as mediators that connect dispersed document information to search queries, enabling cohesive understanding by organizing facts into structured units that can be systematically retrieved and synthesized.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If assertions are extracted from multiple documents, then the completeness of knowledge is improved, but the complexity of organizing assertions into meaningful clusters increases

Engineering Contradiction:
Improvecompleteness of knowledgeVSAvoidcomplexity of organizing assertions
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the organization process into distinct stages: extraction of individual assertions, tagging with metadata, and clustering into topic groups. This segmented approach makes the overall organization task more manageable by breaking down the complex problem into smaller, systematic steps.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters of assertion organization by introducing metadata tags and clustering criteria. Assertions are organized based on configurable parameters such as topic categories, document sources, and semantic relationships, allowing flexible and systematic organization without requiring complex manual curation.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If facts are identified and organized into clusters, then the level of understanding is improved, but the time and resources required for organization increase

Engineering Contradiction:
Improvelevel of understandingVSAvoidtime for organization
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary organization of assertions into clusters and topics before users need to access the information. By pre-processing documents to extract, tag, and cluster assertions in advance, the system reduces the time required for users to gain understanding, as the organizational work is completed beforehand.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service organization through automated assertion extraction and clustering algorithms. The system automatically organizes facts into meaningful clusters without requiring manual intervention, reducing time and resource requirements while maintaining high levels of understanding and coherence.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8370278B2Ontological categorization of question concepts from document summaries
Publication Date: 2013.02.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8370278B2 patent drawing
  • US8370278B2 patent drawing
  • US8370278B2 patent drawing

AI summary

Electronic documents are analyzed to identify assertions, which are inverted to generate questions that may be answered by the assertions. A document or a corpus of electronic documents may be analyzed to identify entities and relationships among entities within the text of the document(s). Assertions are identified based on the entities and relationships among the entities. Each assertion represents a fact about an entity, and a group of assertions represents a summary of the document or document corpus. The assertions are inverted to generate questions that may be answered by the assertions. The questions may be further analyzed to identify relevant concepts and topics and to cluster the questions around the concepts and topics. A combined graph may also be generated that facilitates traversal among topics, concepts, questions, assertions, document summaries, and documents.