Syntactic Annotation Tokens for Distant Term Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Question answer systems face challenges in accurately ranking answers due to limitations in capturing syntactic relationships between distant terms in source documents, leading to suboptimal query search results.

Innovation Solution

A knowledge manager generates syntactic annotation tokens based on syntactic relationships within source documents, storing them in parallel fields to create a knowledge structure that enhances query searches by aligning terms and their annotations, allowing for improved ranking of answers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spannear queries are used to search for terms appearing close to each other, then query search capability is provided, but syntactic relationships between distant terms cannot be captured

Engineering Contradiction:
Improvequery search accuracyVSAvoidsyntactic relationship information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent introduces a new dimension to the search system by adding syntactic annotation tokens that capture grammatical relationships between terms. Instead of only searching based on term proximity in the original text space, the system creates a syntactic dimension where terms are connected through their grammatical roles (subject, object, modifier, etc.), enabling detection of relationships between distant terms that share syntactic connections.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces syntactic annotation tokens as intermediary elements between original terms. These tokens act as mediators that explicitly represent syntactic relationships (such as subject-verb-object connections) between terms that may be far apart in the original text, allowing the search system to traverse these intermediary connections to find semantically related terms regardless of their positional distance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional term proximity ranking is used, then simple query processing is maintained, but answer ranking accuracy deteriorates

Engineering Contradiction:
Improvequery processing efficiencyVSAvoidanswer ranking accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing syntactic annotation tokens for all terms in the corpus during an offline processing phase. This preliminary syntactic enrichment of the data allows the online query processing to efficiently leverage pre-established syntactic relationships without performing complex real-time syntactic analysis, thus maintaining query processing efficiency while improving answer ranking accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical proximity-based ranking system with a syntactic relationship-based ranking system. Instead of mechanically counting word distances and using simple proximity metrics, the system substitutes this with a semantic-aware ranking mechanism that queries pre-computed syntactic annotations to determine answer relevance based on grammatical relationships, thereby improving ranking accuracy while maintaining efficiency through the pre-processed syntax data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9904674B2Augmented text search with syntactic information
Publication Date: 2018.02.27 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US9904674B2 patent drawing
  • US9904674B2 patent drawing
  • US9904674B2 patent drawing

AI summary

An approach is provided in which a knowledge manager generates syntactic annotation tokens that correspond to syntactic relationships between terms included in a source document. The knowledge manager creates a knowledge structure that stores the syntactic annotation tokens in parallel fields and stores the source document terms in original text fields, which align to the parallel fields. In turn, the knowledge manager utilizes the knowledge structure to generate answers to questions based upon the syntactic annotation tokens.