Semantic Document Retrieval via Triple Extraction and Expansion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document search systems require significant skill and knowledge to find documents based on semantic similarity, as they often rely on word searches, Boolean operators, and vector space models, which are inefficient and prone to errors.

Innovation Solution

A system and method that ingest documents to extract and index library triples, allowing for the identification of similar triples based on semantic similarity, with the ability to expand triples using a semantic corpus and rank results by similarity scores, enabling quicker and more precise retrieval of similar documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If word search or Boolean operator methods are used, then documents can be found based on keyword matching, but the search requires significant skill and knowledge and is prone to errors

Engineering Contradiction:
Improvesearch accuracyVSAvoidsearch complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces manual keyword search mechanisms with an automated semantic similarity system that uses natural language processing and machine learning algorithms to automatically extract entities, relationships, and events from documents, eliminating the need for users to manually craft search queries with Boolean operators

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system transforms the search approach by changing from exact keyword matching to semantic similarity comparison, where documents are represented as structured data models containing entities, relationships, and events that can be compared based on their semantic content rather than literal text matches

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If vector space models are used for document analysis, then document importance can be determined, but the search process becomes time-consuming and less efficient

Engineering Contradiction:
Improvedocument relevanceVSAvoidsearch time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing by pre-extracting and structuring entities, relationships, and events from documents during an indexing phase, creating ready-to-query structured representations that enable rapid similarity comparison without requiring time-consuming vector space calculations during the actual search

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system segments documents into discrete structured components (entities, relationships, events) that can be independently processed and compared, allowing for more efficient search operations than analyzing entire documents as continuous vector representations

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If semantic similarity-based search is implemented, then document retrieval accuracy improves, but the system complexity increases

Engineering Contradiction:
Improvesemantic similarity accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces structured data models as intermediary representations between raw document text and similarity comparison operations. These models serve as a simplified interface that captures semantic meaning in an organized format, making the system more manageable despite the complexity of underlying NLP and machine learning components

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240111816A9System and method for finding similar documents based on semantic factual similarity
Publication Date: 2024.04.04 THOMSON REUTERS ENTERPRISE CENTRE GMBH
  • US20240111816A9 patent drawing
  • US20240111816A9 patent drawing
  • US20240111816A9 patent drawing

AI summary

The present disclosure is directed towards systems and methods for finding documents that are similar to a reference text. The inventive systems and methods examine a set of collected documents to determine the facts present in those documents by, for example, extracting triplets and expanding them. A user's input reference text is similarly examined to extract and expand triplets therein and the facts identified with respect to the reference text are used as a basis to find documents having similar facts. The present disclosure is also related to systems and methods for mining facts from documents relating to a primary source such as a piece of legislation and using the mined facts to improve the results of subsequent searches.