Query Segmentation via Probabilistic Scoring and Dynamic Identifier Insertion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
State-of-the-art search engines face challenges in processing complex search queries due to high computational resource demands and error-prone manual segmentation methods, which lead to low accuracy and inability to recognize new keywords.
Innovation Solution
A method that segments search queries by correlating semantic elements with predetermined search terms, modifying irrelevant elements with segmentation identifiers, and combining terms to form probabilistically weighted search queries, allowing for dynamic and accurate segmentation without manual intervention or complex databases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If statistics-based machine learning method is used for query segmentation, then segmentation can be automated, but computational resources and computational time are hugely consumed
Solution Approach 1:
The patent segments the search query into multiple candidate segmentation schemes, where each scheme represents a different way of dividing the query into keywords. This allows the system to evaluate multiple possibilities efficiently without requiring exhaustive computational analysis of all potential segmentations.
Solution Approach 2:
The patent changes the parameter of segmentation accuracy by introducing a segmentation score calculation mechanism that evaluates each candidate segmentation scheme. This allows the system to automatically select the most appropriate segmentation without requiring manual intervention or excessive computational resources.
2Extent of automation
If statistics-based machine learning method is used for query segmentation, then automation is achieved, but computational time is hugely consumed
Solution Approach 1:
The patent generates a limited number of candidate segmentation schemes (e.g., top N schemes) rather than evaluating all possible segmentations. This partial action approach achieves sufficient automation while significantly reducing computational time compared to exhaustive methods.
Solution Approach 2:
The patent introduces a segmentation score parameter that enables rapid evaluation and comparison of candidate segmentation schemes. This parameter-driven approach allows the system to quickly identify the best segmentation without requiring extensive computational time.
3Measurement precision
If manual text segmentation is used to obtain segmentation rules, then segmentation accuracy can be improved, but errors in manual segmentation propagate to segmentation rules and subsequent search query segmentation
Solution Approach 1:
The patent enables the system to automatically evaluate and select segmentation schemes without relying on manually created segmentation rules. The segmentation score calculation mechanism allows the system to self-correct and adapt to different query types, eliminating error propagation from manual segmentation.
Solution Approach 2:
The patent introduces a feedback mechanism through segmentation score evaluation, where the system assesses the quality of each candidate segmentation scheme and uses this information to select the best segmentation. This feedback loop prevents error propagation by continuously evaluating segmentation quality rather than relying on fixed manual rules.
4Measurement precision
If statistics-based machine-learning method is used, then existing keywords can be segmented, but new keywords that have not appeared in manual text segmentation cannot be recognized, increasing error rate
Solution Approach 1:
The patent dynamically generates candidate segmentation schemes based on the actual search query rather than relying on pre-established segmentation rules. This dynamic approach allows the system to adapt to new keywords and emerging terminology, improving both accuracy and versatility simultaneously.
Solution Approach 2:
The patent creates a universal segmentation evaluation mechanism that can handle both existing keywords and new keywords equally. The segmentation score calculation approach is applicable to any keyword combination, making the system versatile across different domains and query types without requiring domain-specific manual rules.
Data Source
AI summary
The present application discloses a method for segmenting a search query. A server receives a search query including an ordered sequence of Chinese characters. For each Chinese character, one or more predetermined search terms are identified and then combined to form concatenated search queries, each concatenated search query including at least one segmentation identifier that separates the Chinese characters of the ordered sequence of Chinese characters. A specific concatenated search query is identified based on search probabilities of the concatenated search queries. The specific concatenated search query is further segmented into two or more search terms according to one or more locations of the at least one segmentation identifier in the specific concatenated search query. Finally, at least one new search term is identified from the two or more search terms such that one of the ordered sequence of Chinese characters occupies the first position of the new search term.


