Key Phrase Discovery Using Search Query Logs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Advertisers face inefficiencies in manually composing lists of keyword variations for online keyword auctions, often missing important terms and variations like misspellings and semantically related keywords, which can lead to lost sales and revenue.

Innovation Solution

A system that extracts key phrases from search query logs and generates a Similarity Graph to determine similarity levels between keywords, enabling the automatic identification of frequent misspellings, keyword/acronym pairs, and semantically related keywords, thereby narrowing the search space for advertisers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If advertisers manually compose lists of keyword variations, then they can control the keyword selection process, but they often miss important terms and variations like misspellings and semantically related keywords

Engineering Contradiction:
Improvekeyword coverage completenessVSAvoidmissed keyword variations
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent introduces search engine query logs as an intermediary data source that captures actual user search behavior. These logs serve as a mediator between user intent and advertiser keyword selection, providing empirical evidence of what terms users actually use when searching for products or services. This intermediary data source reveals misspellings, variations, and semantically related keywords that manual composition would miss.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the manual mechanical process of composing keyword lists with an automated computational system. The system uses algorithms to process search query logs, automatically extract relevant keywords and variations, and generate comprehensive keyword lists. This substitution eliminates human limitations in identifying all possible keyword variations while maintaining systematic and reproducible keyword selection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If advertisers include more search terms and variations in advertising, then they can reach more potential customers, but the manual process is time-consuming and prone to omissions

Engineering Contradiction:
Improvekeyword list comprehensivenessVSAvoidkeyword research time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-processing and analyzing search query logs to identify patterns, frequent misspellings, and semantically related keywords before the advertising campaign begins. The system captures and stores this intelligence in advance, so when advertisers need keyword lists, the comprehensive work of identifying all variations has already been completed. This preliminary analysis eliminates the need for time-consuming manual research during campaign setup.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service by automatically generating comprehensive keyword lists without requiring advertiser intervention in the analysis process. The computational system serves itself by using its own resources to process query logs, extract keywords, and generate lists that advertisers can directly utilize. This self-service capability scales efficiently to handle large volumes of data and multiple advertisers simultaneously.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If the system processes large volumes of search query logs to find frequent misspellings and variations, then keyword accuracy improves, but computational complexity increases

Engineering Contradiction:
Improvekeyword identification accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential and most frequent keyword variations from the large volume of search query logs, rather than processing every single query in detail. The system identifies and extracts high-frequency misspellings, common variations, and semantically related terms that have the greatest impact on advertising effectiveness. This selective extraction reduces computational complexity while maintaining high accuracy for the most important keywords.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies parameter changes by using frequency thresholds and similarity metrics to filter and prioritize keyword candidates from the search query logs. The system transforms the raw data into structured keyword lists by applying parameters such as minimum frequency counts, spelling similarity thresholds, and semantic relevance scores. These parameter-based filtering mechanisms reduce the data volume requiring detailed processing while preserving the most valuable keywords.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7627559B2Context-based key phrase discovery and similarity measurement utilizing search engine query logs
Publication Date: 2009.12.01 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7627559B2 patent drawing
  • US7627559B2 patent drawing
  • US7627559B2 patent drawing

AI summary

Usage context obtained from search query logs is leveraged to facilitate in discovery and/or similarity determination of key search phrases. A key phrase extraction process extracts key phrases from raw search query logs and breaks individual queries into a vector of the key phrases. A Similarity Graph generation process then generates a Similarity Graph from the output of the key phrase extraction process. Information relating to the similarity levels between two key phrases can be employed to restrict a search space for tasks such as, for example, online keyword auctions and the like. Thus, instances can be employed to find frequent misspellings of a given keyword, keyword/acronym pairs, key phrases with similar intention, and/or keywords which are semantically related and the like.