Sibling Search Query Generation via Wildcard Substitution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current natural language processing systems require human intervention for acquiring training data and rely on explicit word or phrase types, making them inefficient and less robust for generating related search queries.

Innovation Solution

A system that processes search queries with wildcards to generate sibling search queries by comparing similarity measures from search query logs, independent of natural language processing, allowing for automatic generation of training data for machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If natural language processing systems are used to generate related search queries, then the system can understand semantic meaning, but the system requires human intervention for acquiring training data and relies on explicit word or phrase types making it inefficient

Engineering Contradiction:
Improvesemantic understanding accuracyVSAvoidtraining data acquisition efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system automatically generates sibling search queries by substituting wildcards with different word sequences from search query logs, eliminating the need for human intervention in training data acquisition. The system serves itself by using its own search query data to generate training examples for natural language processing models.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-processes search query logs to extract and organize word sequences that can be used as substitutions for wildcards. By preparing these word sequences in advance and storing them in a structured format, the system enables efficient generation of sibling search queries without requiring real-time human intervention.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If explicit word or phrase types are used in natural language processing systems, then the system can categorize search queries, but the system becomes less robust for generating related search queries

Engineering Contradiction:
Improvequery categorization capabilityVSAvoidrobustness in generating related queries
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system changes the approach from using explicit word types to using wildcard-based parameter substitution. By representing variable parts of search queries as wildcards and substituting them with word sequences from logs, the system maintains categorization capability while improving robustness through actual usage patterns rather than predefined types.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system copies actual search query patterns from the search query log and uses them as templates for generating sibling queries. Instead of relying on explicit type definitions, the system replicates real-world query structures and substitutes only the variable portions, thereby maintaining robustness and adaptability simultaneously.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If human intervention is used for acquiring training data, then the training data can be carefully curated, but the process becomes inefficient and time-consuming

Engineering Contradiction:
Improvetraining data qualityVSAvoidtraining data acquisition time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system automatically generates training data by substituting wildcards with word sequences from search query logs, eliminating the need for human curators. The system uses its own operational data to create training examples, significantly reducing acquisition time while maintaining quality through systematic processing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces the manual mechanical process of human data curation with an automated computational process. By using algorithmic substitution of wildcards with word sequences, the system achieves rapid generation of training data while maintaining consistency and quality through structured processing rules.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If natural language processing models are trained manually, then the models can be accurately trained, but the process requires significant human resources and time

Engineering Contradiction:
Improvemodel training accuracyVSAvoidtraining process automation level
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The system automatically generates training data and prepares it for model training without human intervention. By using wildcard substitution with word sequences from search query logs, the system creates training examples that maintain accuracy while enabling full automation of the training data preparation process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-generates training data by substituting wildcards with relevant word sequences before model training begins. This preliminary preparation of training data in an automated manner ensures accuracy is maintained while the data is ready for immediate use in training processes, reducing both time and human resource requirements.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11379527B2Sibling search queries
Publication Date: 2022.07.05 GOOGLE LLC
  • US11379527B2 patent drawing
  • US11379527B2 patent drawing
  • US11379527B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for determining a plurality of sibling search queries for an input search query. In one aspect, a method comprises: receiving an input search query that satisfies a context template comprising a sequence of one or more words and a wildcard, wherein a wildcard represents variable data, wherein the input search query satisfies the context template and comprises a target word sequence that corresponds to the wildcard in the context template; and determining a plurality of sibling search queries for the input search query, wherein each sibling search query satisfies the context template and comprises a sibling word sequence that corresponds to the wildcard in the context template.