Semantic Document Retrieval via Triple Extraction and Expansion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document search systems require significant skill and knowledge to find documents based on semantic similarity, as they often rely on word searches, Boolean operators, and vector space models, which are inefficient and prone to errors.
Innovation Solution
A system and method that ingest documents to extract and index library triples, allowing for the identification of similar triples based on semantic similarity, with the ability to expand triples using a semantic corpus and rank results by similarity scores, enabling quicker and more precise retrieval of similar documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If word search or Boolean operator methods are used, then documents can be found based on keyword matching, but the search requires significant skill and knowledge and is prone to errors
Solution Approach 1:
The patent replaces manual keyword search mechanisms with an automated semantic similarity system that uses natural language processing and machine learning algorithms to automatically extract entities, relationships, and events from documents, eliminating the need for users to manually craft search queries with Boolean operators
Solution Approach 2:
The system transforms the search approach by changing from exact keyword matching to semantic similarity comparison, where documents are represented as structured data models containing entities, relationships, and events that can be compared based on their semantic content rather than literal text matches
2Measurement precision
If vector space models are used for document analysis, then document importance can be determined, but the search process becomes time-consuming and less efficient
Solution Approach 1:
The patent performs preliminary processing by pre-extracting and structuring entities, relationships, and events from documents during an indexing phase, creating ready-to-query structured representations that enable rapid similarity comparison without requiring time-consuming vector space calculations during the actual search
Solution Approach 2:
The system segments documents into discrete structured components (entities, relationships, events) that can be independently processed and compared, allowing for more efficient search operations than analyzing entire documents as continuous vector representations
3Measurement precision
If semantic similarity-based search is implemented, then document retrieval accuracy improves, but the system complexity increases
Solution Approach 1:
The patent introduces structured data models as intermediary representations between raw document text and similarity comparison operations. These models serve as a simplified interface that captures semantic meaning in an organized format, making the system more manageable despite the complexity of underlying NLP and machine learning components
Data Source
AI summary
The present disclosure is directed towards systems and methods for finding documents that are similar to a reference text. The inventive systems and methods examine a set of collected documents to determine the facts present in those documents by, for example, extracting triplets and expanding them. A user's input reference text is similarly examined to extract and expand triplets therein and the facts identified with respect to the reference text are used as a basis to find documents having similar facts. The present disclosure is also related to systems and methods for mining facts from documents relating to a primary source such as a piece of legislation and using the mined facts to improve the results of subsequent searches.


