Voice Directory Search Using Hierarchical Language Model Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automated directory assistance systems face challenges in accurately recognizing business names due to the large number of similar names and imprecise user inputs, leading to reduced efficiency and accuracy in providing relevant information.
Innovation Solution
A voice-enabled business directory search system that uses a hierarchical tree structure of nodes associated with speech recognition language models, where each node is linked to specific business categories, allowing for more accurate recognition by narrowing down the range of potential matches based on user input and geographical location, and updates language models based on user interactions and search logs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If a conventional automated directory assistance system uses speech recognition to identify business names, then the system can automate the information retrieval process, but the recognition accuracy deteriorates due to the large number of similar business names and imprecise user inputs
Solution Approach 1:
The patent segments the speech recognition process into multiple stages: initial speech input recognition, category identification, and refined business name recognition. By dividing the recognition task into sequential steps with narrowing scope, the system maintains automation while improving accuracy at each stage.
Solution Approach 2:
The system performs preliminary actions by first identifying business categories and geographical locations before attempting to recognize the specific business name. This preliminary classification narrows the search space and provides context that improves subsequent name recognition accuracy.
2Adaptability or versatility
If the speech recognition system maintains a large database of all possible business names to improve recognition coverage, then the system can recognize a wider range of businesses, but the computational resources and processing time increase
Solution Approach 1:
The patent segments the business database into hierarchical categories and geographical regions. Instead of processing all business names simultaneously, the system only loads and processes the subset of businesses relevant to the identified category and location, reducing memory and computational requirements while maintaining comprehensive coverage.
Solution Approach 2:
The system applies local quality by tailoring the recognition database to each specific query context. Language models and business lists are dynamically selected based on the identified category and geographical location, ensuring high recognition coverage for relevant businesses while minimizing resource usage for irrelevant ones.
3Device complexity
If the system uses a single comprehensive language model for all business categories to maintain system simplicity, then the device complexity is reduced, but the speech recognition accuracy deteriorates for specific business types
Solution Approach 1:
The patent implements a dynamic language model selection mechanism that automatically chooses the appropriate specialized language model based on the identified business category. The system transitions from a static single-model approach to a dynamic multi-model approach, where model selection occurs in real-time based on contextual information.
Solution Approach 2:
The business category identification acts as an intermediary that bridges the user's speech input and the appropriate language model selection. This intermediary layer analyzes the speech input to determine the business category, then uses this information to select the most suitable language model for accurate name recognition.
Data Source
AI summary
A method of providing navigation directions includes receiving, at a user terminal, a query spoken by a user, wherein the query spoken by the user includes a speech utterance indicating (i) a category of business, (ii) a name of the business, and (iii) a location at which or near which the business is disposed; identifying, by processing hardware, the business based on the speech utterance; and providing navigation directions to the business via the user terminal.


