Topical Document Search via Reference Network Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in efficiently searching large datasets of electronic data due to the need for well-formulated queries, often resulting in either irrelevant or excessive search results, as existing search tools struggle to limit the search to relevant subsets of data related by topic.

Innovation Solution

A method is provided to define a topical subset of data by following references between documents, allowing for a focused search within a defined search space, which includes originating data and iteratively adds data that contains references to or from the originating data, thereby narrowing the search to relevant documents related by common topics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a user searches the entire data set using traditional search tools, then the search covers all possible data, but the number of results becomes excessively large and includes irrelevant information

Engineering Contradiction:
Improverelevance of search resultsVSAvoidnumber of search results
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the entire data set into multiple topic-based subsets or clusters. Instead of searching all data at once, the system divides the search space into manageable topic groups, allowing users to search within relevant segments only, thereby reducing result quantity while maintaining reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces topic models and semantic intermediaries that act as mediators between user queries and the data set. These intermediaries interpret queries, identify relevant topics, and guide the search to appropriate data subsets, filtering out irrelevant results and improving the relevance-to-quantity ratio.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If a user formulates a precise query to reduce results, then the relevance of results improves, but the time spent formulating and adjusting queries increases

Engineering Contradiction:
Improverelevance of search resultsVSAvoidtime spent on query formulation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing the data set into topic-based clusters and pre-computing relevance relationships before actual searches. When a user submits a query, the system has already organized the data structure to enable rapid identification of relevant subsets, eliminating the need for users to spend time refining queries through trial and error.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where the system analyzes user queries, provides suggestions for optimization, and learns from user interactions to improve future query processing. This feedback loop reduces the time users need to spend formulating queries by guiding them toward more effective search strategies.

Inventive Principle:
Principle #23Feedback

3Productivity

If a user limits the search to a specific subset of data, then the search time decreases, but the risk of missing relevant information outside the subset increases

Engineering Contradiction:
Improvesearch speedVSAvoidcompleteness of search results
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent makes the search subset dynamic rather than static. The system automatically adjusts the search scope based on the query content, user preferences, and relevance analysis. If the initial subset yields insufficient results, the system dynamically expands the search to related subsets, ensuring completeness while maintaining efficiency.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent employs a nested structure where topic-based subsets are organized hierarchically, with broader topics containing narrower subtopics. The search can navigate through nested levels, starting from a focused subset and progressively expanding to parent topics if needed, thus balancing search speed with result completeness.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS9519707B2System and method for topical document searching
Publication Date: 2016.12.13 THE BUREAU OF NAT AFFAIRS
  • US9519707B2 patent drawing
  • US9519707B2 patent drawing
  • US9519707B2 patent drawing

AI summary

Systems and methods are providing for searching for documents within topically-defined clusters. A search space is defined, starting with one or more source documents, by examining references from one documents to another and following the networks of references to some level of indirection. Depending on the embodiment, references may be followed from a document containing a reference to a referred-to document, or from a referred-to document to a document containing a reference, or both. Once a search space has been defined, a query is executed, and documents within the search space that satisfy the query parameters are identified.In certain embodiments of the invention, the documents primarily relate to legal materials, and one or more source documents are associated with one or more topics within a topic directory. In such embodiments, a search query may be limited to one or more selected topics by executing the search query within a search space defined using the associated document or documents as the source.