Natural Language Query Parsing with Discriminator Scores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search query generation methods fail to identify discriminative terms effectively due to reliance on taxonomies like HUTT, leading to poor search results when HUTT terms are not good discriminators across all documents.
Innovation Solution
A computer-implemented method that parses natural language questions into parse trees, identifies argument positions, and calculates discriminator scores using TF-IDF to determine required terms, modifiers, and bigrams, adding them to the search query if they surpass predetermined threshold scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If search query generation relies on taxonomies like HUTT to identify required terms, then the process is systematic and structured, but the discriminative power of selected terms deteriorates when HUTT terms are not good discriminators across all documents
Solution Approach 1:
The patent changes the selection criterion from taxonomy-based categorization to discriminator score-based selection. Instead of relying on HUTT taxonomy levels, the system calculates a discriminator score for each term based on its ability to distinguish relevant documents from non-relevant ones, and selects terms that exceed a threshold score. This parameter change enables the system to adapt to any document corpus regardless of taxonomy coverage.
2Productivity
If only HUTT terms are included as required terms, then the query generation is simple and fast, but the search result quality deteriorates when HUTT terms fail to capture important discriminators
Solution Approach 1:
The system implements a feedback mechanism where discriminator scores are calculated based on the actual performance of terms in distinguishing relevant from non-relevant documents. Terms are selected based on this feedback signal (discriminator score threshold), creating a closed-loop system that continuously optimizes for search result quality while maintaining efficient query generation through automated scoring.
Solution Approach 2:
The discriminator score acts as an intermediary metric between the raw term frequency data and the final term selection decision. Instead of directly selecting terms based on taxonomy presence, the system uses the discriminator score as an intermediate evaluation layer that bridges the gap between simple term counting and quality-based term selection.
3Measurement precision
If the system considers modifiers and bigrams in addition to head terms, then the query captures more nuanced meaning, but the complexity of query generation increases
Solution Approach 1:
The patent segments the query generation process into distinct hierarchical levels: head terms, modifiers, and bigrams. Each level is processed separately with its own discriminator score calculation and threshold application. This segmentation allows the system to systematically evaluate and combine different types of terms while maintaining manageable process complexity through modular evaluation stages.
Data Source
AI summary
Embodiments can provide a computer implemented method, in a data processing system comprising a processor and a memory comprising instructions which are executed by the processor to cause the processor to implement an improved search query generation system, the method comprising inputting a natural language question; parsing the natural language question into a parse tree; identifying argument positions comprising one or more argument position terms; for each argument position: comparing a head term's discriminator score against a threshold discriminator score; and if the head term surpasses the threshold discriminator score, adding the head term as a required term to an improved search query; and outputting the improved search query.


