Syntactic Annotation Tokens for Distant Term Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Question answer systems face challenges in accurately ranking answers due to limitations in capturing syntactic relationships between distant terms in source documents, leading to suboptimal query search results.
Innovation Solution
A knowledge manager generates syntactic annotation tokens based on syntactic relationships within source documents, storing them in parallel fields to create a knowledge structure that enhances query searches by aligning terms and their annotations, allowing for improved ranking of answers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spannear queries are used to search for terms appearing close to each other, then query search capability is provided, but syntactic relationships between distant terms cannot be captured
Solution Approach 1:
The patent introduces a new dimension to the search system by adding syntactic annotation tokens that capture grammatical relationships between terms. Instead of only searching based on term proximity in the original text space, the system creates a syntactic dimension where terms are connected through their grammatical roles (subject, object, modifier, etc.), enabling detection of relationships between distant terms that share syntactic connections.
Solution Approach 2:
The patent introduces syntactic annotation tokens as intermediary elements between original terms. These tokens act as mediators that explicitly represent syntactic relationships (such as subject-verb-object connections) between terms that may be far apart in the original text, allowing the search system to traverse these intermediary connections to find semantically related terms regardless of their positional distance.
2Productivity
If traditional term proximity ranking is used, then simple query processing is maintained, but answer ranking accuracy deteriorates
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing syntactic annotation tokens for all terms in the corpus during an offline processing phase. This preliminary syntactic enrichment of the data allows the online query processing to efficiently leverage pre-established syntactic relationships without performing complex real-time syntactic analysis, thus maintaining query processing efficiency while improving answer ranking accuracy.
Solution Approach 2:
The patent replaces the mechanical proximity-based ranking system with a syntactic relationship-based ranking system. Instead of mechanically counting word distances and using simple proximity metrics, the system substitutes this with a semantic-aware ranking mechanism that queries pre-computed syntactic annotations to determine answer relevance based on grammatical relationships, thereby improving ranking accuracy while maintaining efficiency through the pre-processed syntax data.
Data Source
AI summary
An approach is provided in which a knowledge manager generates syntactic annotation tokens that correspond to syntactic relationships between terms included in a source document. The knowledge manager creates a knowledge structure that stores the syntactic annotation tokens in parallel fields and stores the source document terms in original text fields, which align to the parallel fields. In turn, the knowledge manager utilizes the knowledge structure to generate answers to questions based upon the syntactic annotation tokens.


