Voice Recognition Language Model Generation via Document-Type Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing language model generation techniques for voice recognition face challenges in accurately collecting a corpus similar to the voice recognition target from web pages, leading to reduced recognition accuracy due to retrieval of diverse web pages, inclusion of high-frequency words, and biased results from synonym inclusion and notation fluctuations.
Innovation Solution
A language model generating device and method that analyzes web page text, extracts appropriate words based on document type, generates a word set, and uses this set to retrieve relevant web pages, thereby generating a language model for voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If web pages are retrieved using a single retrieval word, then the retrieval process is simple, but the retrieved web pages include diverse topics and specialized fields, reducing recognition accuracy
Solution Approach 1:
The patent segments the retrieval process by dividing the corpus into multiple document types (e.g., news articles, academic papers, blogs) and extracting retrieval words specific to each document type. This segmentation allows the system to retrieve more targeted web pages for each document type, improving recognition accuracy while maintaining operational simplicity through automated classification.
Solution Approach 2:
The system dynamically adjusts retrieval strategies based on the analyzed corpus characteristics. It automatically determines the appropriate number of retrieval words and their weights based on the document type distribution and content analysis, making the retrieval process adaptive rather than static, thereby improving accuracy without requiring manual configuration.
2Measurement precision
If multiple retrieval words are used to improve retrieval accuracy, then recognition accuracy improves, but the retrieval process becomes more complex
Solution Approach 1:
The patent changes parameters such as the number of retrieval words, their weights, and selection criteria based on the analyzed corpus characteristics. The system automatically adjusts these parameters according to document type and content analysis results, achieving high retrieval accuracy through adaptive parameter optimization rather than fixed complex rules.
Solution Approach 2:
The system performs self-service by automatically analyzing the corpus, identifying document types, selecting appropriate retrieval words, and determining retrieval parameters without external intervention. This automation reduces the apparent complexity for users while maintaining high retrieval accuracy through sophisticated internal processing.
3Adaptability or versatility
If high-frequency words are included in the language model, then the language model covers more common terms, but words with high appearance frequency have unreasonably high language probability
Solution Approach 1:
The patent applies local quality by adjusting language probabilities based on local context and document type. Instead of uniformly high probabilities for all high-frequency words, the system modifies probabilities according to the specific document type and contextual relevance, ensuring that high-frequency words contribute appropriately to the language model without dominating it unreasonably.
4Adaptability or versatility
If synonyms and notation variations are included in the corpus, then the language model is more comprehensive, but the retrieval results become biased and corpus collection becomes insufficient
Solution Approach 1:
The patent performs preliminary action by pre-processing the corpus to normalize synonyms and notation variations before analysis. The system identifies and consolidates equivalent terms in advance, creating a standardized foundation that improves both comprehensiveness and reliability. This preliminary normalization prevents bias in retrieval results while maintaining the benefits of comprehensive language coverage.
Data Source
AI summary
A text in a corpus including a set of world wide web (web) pages is analyzed. At least one word appropriate for a document type set according to a voice recognition target is extracted based on an analysis result. A word set is generated from the extracted at least one word. A retrieval engine is caused to perform a retrieval process using the generated word set as a retrieval query of the retrieval engine on the Internet, and a link to a web page from the retrieval result is acquired. A language model for voice recognition is generated from the acquired web page.


