Search Query Normalization via Session Log Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current query normalization techniques, relying on text-matching, fail to recognize specific brand names, product names, and retail-specific jargon, leading to poor search experiences as they cannot generate all suitable items based on a search query, often returning irrelevant results due to mismatched user intent.
Innovation Solution
A system and method that analyzes session logs to generate sets of query reformulations, filters and categorizes normalization candidates, and compares categories to remove uncommon ones, thereby transforming search queries into accurate and relevant normalization candidates for web search engines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If query normalization based on text-matching is used, then the search engine can locate items matching the search query, but it fails to recognize specific brand names, product names, and retail-specific jargon, leading to poor search experiences
Solution Approach 1:
The patent introduces an intermediary component - a normalization candidate generation system that bridges the gap between simple text-matching and complex semantic understanding. This system uses session logs and query reformulations as intermediaries to infer user intent and generate normalized search queries that account for brand names, product names, and retail jargon without requiring the search engine itself to perform complex recognition tasks.
Solution Approach 2:
The patent applies preliminary action by generating normalization candidates and reformulations before the actual search execution. The system pre-processes search queries by analyzing session logs, generating potential normalization candidates, and selecting the most appropriate ones in advance, thereby improving both accuracy and adaptability without adding complexity during the actual search operation.
2Productivity
If text-matching query normalization is used, then the search engine returns results for the particular search query, but it cannot generate most suitable items based on user intent, leading to poor search experiences
Solution Approach 1:
The patent implements feedback mechanisms by utilizing session logs that capture user search behavior, query modifications, and interaction patterns. This feedback loop allows the system to learn from actual user intent and refine normalization candidate generation, thereby improving reliability while maintaining productivity through efficient log-based analysis rather than real-time complex processing.
Solution Approach 2:
The system applies self-service by enabling the search engine to automatically generate and select normalization candidates based on its own accumulated session logs and query reformulation data. This self-improving mechanism allows the system to enhance its understanding of user intent over time without requiring external intervention, thereby improving reliability while maintaining operational efficiency.
3Device complexity
If simple text-matching normalization is used, then the system operation is simple, but it cannot recognize retail-specific jargon and generates irrelevant results
Solution Approach 1:
The patent applies segmentation by breaking down the complex task of query normalization into distinct components: session log analysis, query reformulation generation, normalization candidate selection, and category-based filtering. This segmentation allows each component to be processed independently using relatively simple operations, maintaining overall system simplicity while achieving high precision through the coordinated execution of multiple specialized modules.
Data Source
AI summary
A system for generating normalization candidates for a search query includes a database for storing session logs with each session log including query data and a processor in communication with the database and configured to execute computer-readable instructions causing the processor to analyze session log data to generate sets of query reformulations for a plurality of search queries, select one of the sets containing a normalization candidate that matches the search query, filter the selected set of reformulations, tie the candidates in the selected set to a category, compare the categories of the candidates, remove at least one reformulation from the selected set when the category of one candidate is uncommon with the category of the other candidate, and store the remaining candidates in the database. A method and one or more non-transitory computer-readable storage media for generating stemming pairs for a search query are also disclosed.


