Query Segmentation via Probabilistic Scoring and Dynamic Identifier Insertion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

State-of-the-art search engines face challenges in processing complex search queries due to high computational resource demands and error-prone manual segmentation methods, which lead to low accuracy and inability to recognize new keywords.

Innovation Solution

A method that segments search queries by correlating semantic elements with predetermined search terms, modifying irrelevant elements with segmentation identifiers, and combining terms to form probabilistically weighted search queries, allowing for dynamic and accurate segmentation without manual intervention or complex databases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If statistics-based machine learning method is used for query segmentation, then segmentation can be automated, but computational resources and computational time are hugely consumed

Engineering Contradiction:
Improvequery segmentation automationVSAvoidcomputational resources consumption
Core Design Contradiction:
Extent of automationVSUse of energy by moving object

Solution Approach 1:

The patent segments the search query into multiple candidate segmentation schemes, where each scheme represents a different way of dividing the query into keywords. This allows the system to evaluate multiple possibilities efficiently without requiring exhaustive computational analysis of all potential segmentations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of segmentation accuracy by introducing a segmentation score calculation mechanism that evaluates each candidate segmentation scheme. This allows the system to automatically select the most appropriate segmentation without requiring manual intervention or excessive computational resources.

Inventive Principle:
Principle #35Parameter changes

2Extent of automation

If statistics-based machine learning method is used for query segmentation, then automation is achieved, but computational time is hugely consumed

Engineering Contradiction:
Improvequery segmentation automationVSAvoidcomputational time
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The patent generates a limited number of candidate segmentation schemes (e.g., top N schemes) rather than evaluating all possible segmentations. This partial action approach achieves sufficient automation while significantly reducing computational time compared to exhaustive methods.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent introduces a segmentation score parameter that enables rapid evaluation and comparison of candidate segmentation schemes. This parameter-driven approach allows the system to quickly identify the best segmentation without requiring extensive computational time.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manual text segmentation is used to obtain segmentation rules, then segmentation accuracy can be improved, but errors in manual segmentation propagate to segmentation rules and subsequent search query segmentation

Engineering Contradiction:
Improvequery segmentation accuracyVSAvoiderror propagation in segmentation rules
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent enables the system to automatically evaluate and select segmentation schemes without relying on manually created segmentation rules. The segmentation score calculation mechanism allows the system to self-correct and adapt to different query types, eliminating error propagation from manual segmentation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces a feedback mechanism through segmentation score evaluation, where the system assesses the quality of each candidate segmentation scheme and uses this information to select the best segmentation. This feedback loop prevents error propagation by continuously evaluating segmentation quality rather than relying on fixed manual rules.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If statistics-based machine-learning method is used, then existing keywords can be segmented, but new keywords that have not appeared in manual text segmentation cannot be recognized, increasing error rate

Engineering Contradiction:
Improvekeyword segmentation accuracyVSAvoidrecognition of new keywords
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent dynamically generates candidate segmentation schemes based on the actual search query rather than relying on pre-established segmentation rules. This dynamic approach allows the system to adapt to new keywords and emerging terminology, improving both accuracy and versatility simultaneously.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal segmentation evaluation mechanism that can handle both existing keywords and new keywords equally. The segmentation score calculation approach is applicable to any keyword combination, making the system versatile across different domains and query types without requiring domain-specific manual rules.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11003700B2Methods and systems for query segmentation in a search
Publication Date: 2021.05.11 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11003700B2 patent drawing
  • US11003700B2 patent drawing
  • US11003700B2 patent drawing

AI summary

The present application discloses a method for segmenting a search query. A server receives a search query including an ordered sequence of Chinese characters. For each Chinese character, one or more predetermined search terms are identified and then combined to form concatenated search queries, each concatenated search query including at least one segmentation identifier that separates the Chinese characters of the ordered sequence of Chinese characters. A specific concatenated search query is identified based on search probabilities of the concatenated search queries. The specific concatenated search query is further segmented into two or more search terms according to one or more locations of the at least one segmentation identifier in the specific concatenated search query. Finally, at least one new search term is identified from the two or more search terms such that one of the ordered sequence of Chinese characters occupies the first position of the new search term.