Conceptual Inverted Index for Semantic Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional information retrieval technologies based on keyword indexing fail to effectively match documents when query keywords are not present, and semantic analysis techniques like LSA do not leverage large volumes of crowd-sourced data effectively.
Innovation Solution
Implementing a conceptual inverted index (CII) that populates entries with pointers to documents related to concepts in a concept graph, allowing for efficient querying and linking of text to concepts, and automatically updating knowledge bases with new data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If query expansion techniques are used to improve search quality by substituting search terms with synonyms and performing independent searches, then the quality of returned results is improved, but the computational cost and time required for searching increases significantly
Solution Approach 1:
The patent pre-computes and stores synonym relationships, concept hierarchies, and document-concept associations in inverted tables during an offline indexing phase. This preliminary action allows the search system to quickly retrieve pre-processed relationships without performing expensive computations during online query execution, thus improving response time while maintaining search quality
Solution Approach 2:
The inverted table structure enables the system to self-serve search queries by directly looking up document identifiers associated with queried concepts and their synonyms, eliminating the need for complex real-time joins and computations. The pre-organized data structure allows the system to serve queries efficiently without external computational assistance
2Speed
If traditional keyword indexing is used to enable fast document retrieval, then search speed is improved, but the ability to match documents when query keywords are not present deteriorates
Solution Approach 1:
The patent introduces concepts as intermediaries between keywords and documents. Instead of directly matching keywords to documents, the system maps keywords to concepts, then retrieves documents associated with those concepts. This intermediary layer enables semantic matching beyond exact keyword occurrence, improving document matching accuracy while maintaining fast retrieval through pre-computed concept-document relationships
Solution Approach 2:
The patent adds a conceptual dimension to traditional keyword indexing by organizing data in an inverted table structure that maps concepts to document identifiers. This dimensional transformation allows the system to operate in concept space rather than purely in keyword space, enabling semantic search capabilities while maintaining the efficiency of inverted index lookups
3Measurement precision
If LSA techniques are used to project document representations to latent semantic space, then semantic understanding is improved, but the ability to leverage large volumes of crowd-sourced data deteriorates
Solution Approach 1:
The patent segments the large-scale data processing into two distinct phases: offline concept extraction and indexing from crowd-sourced data, and online query processing. This segmentation allows the system to pre-process and incorporate knowledge from Wikipedia and other sources into structured concept graphs and inverted tables, enabling efficient utilization of large volumes of crowd-sourced data while maintaining fast response times during actual searches
Data Source
AI summary
According to an aspect, storing and querying conceptual indices (CIs) includes creating a conceptual inverted index (CII) from the CIs. The CII includes CII entries, each of which corresponds to a concept in a concept graph. Creating the CII includes populating each entry with pointers to documents selected from the CIs having likelihoods of being related to the concept that are greater than a threshold value, and the corresponding likelihoods. An aspect also includes receiving a query that includes a concept in the concept graph, and generating query results from a search that include the row at least a subset of the pointers to documents. Each of the CIs is associated with a corresponding document and includes a CI entry for each concept in the concept graph, and each of the CI entries specifies a value indicating a likelihood that the document is related to the concept in the concept graph.


