Search Query Relation Determination Using Conditional Probability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search service systems face challenges in accurately determining the relation between search queries, leading to temporal and economic losses due to inefficient data analysis and classification methods that fail to consider meaningful relationships between terms, often registering unrelated queries and consuming excessive time and resources.

Innovation Solution

A method and system that collect and analyze search query data to determine relations using click rate information and session data, generating conditional probability and correlation information to identify relevant search queries, while excluding non-relevant queries and maintaining a systematic database structure for efficient processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If conventional methods classify and store related search queries manually, then the service operator can provide related search queries, but temporal losses and economic losses occur due to inefficient data analysis and classification

Engineering Contradiction:
Improvetemporal lossesVSAvoiddata analysis efficiency
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The system automatically classifies and stores related search queries using statistical analysis of search session data, eliminating the need for manual classification by service operators. The system self-learns relationships between queries through conditional probability calculations based on co-occurrence patterns in search logs.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual classification processes are replaced with automated statistical analysis mechanisms. The system uses conditional probability calculations and correlation analysis to automatically determine query relationships, substituting human operators with computational algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If statistical methods are used to extract related search queries, then time and costs are reduced, but the meaning relation between terms is not considered leading to inaccurate results

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidrelation determination accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system changes the parameters used for relation determination by incorporating multiple statistical measures (conditional probability, correlation coefficients) and filtering criteria (minimum occurrence thresholds, significance levels) to improve the accuracy of meaning relation detection while maintaining processing efficiency.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system uses feedback from statistical analysis results to refine query relations. By calculating conditional probabilities and correlation coefficients from search session data, the system continuously improves its ability to distinguish meaningful relationships from coincidental co-occurrences.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If all co-occurring queries are registered as related, then comprehensive coverage is achieved, but processing time and resources are excessively consumed

Engineering Contradiction:
Improvequery relation coverageVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system changes parameters by setting minimum occurrence thresholds and significance levels for query co-occurrence. Only query pairs that meet these statistical criteria are registered as related, filtering out spurious associations while maintaining comprehensive coverage of meaningful relationships.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system applies different quality standards to different query relations based on their statistical significance. High-confidence relations (meeting threshold criteria) are stored and used, while low-confidence relations are excluded, creating a quality-filtered set of related queries.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS7693904B2Method and system for determining relation between search terms in the internet search system
Publication Date: 2010.04.06 NHN CORP
  • US7693904B2 patent drawing
  • US7693904B2 patent drawing
  • US7693904B2 patent drawing

AI summary

A method of determining a relation between search queries and a system for executing the method are provided. A method of determining a relation between search queries, comprises: maintaining a database including a search session and search queries received from a user terminal during the search session; determining numbers of search sessions where first and second search queries are received during a predetermined time interval; calculating conditional probability based on the determined numbers of search sessions; calculating correlation by using a total number of search sessions and said numbers of search sessions; and determining a relation between said search queries based on said calculated conditional probability and said calculated correlation.