Speech Recognition Library Generation via Phonetic Transcription Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems face challenges in accurately recognizing phrases due to variations in pronunciation, as existing systems often rely on limited phonetic transcriptions that do not account for different pronunciations of names and locations based on nationality or geographic location.
Innovation Solution
A method and apparatus for generating a speech recognition library that identifies and incorporates multiple pronunciations of phrases by analyzing closed caption information from video segments, computing difference metrics, and adding new phonetic transcriptions to the library when they differ from existing ones, ensuring broader recognition capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a speech recognition system uses a limited phonetic transcription library, then the system complexity is reduced, but the recognition accuracy deteriorates due to inability to account for pronunciation variations based on nationality or geographic location
Solution Approach 1:
The patent segments the phonetic transcription library into multiple pronunciation variants organized by nationality or geographic location. Instead of maintaining a single monolithic library, the system divides phonetic transcriptions into distinct segments (e.g., American English, British English, other regional variants) that can be independently stored and retrieved, thereby improving recognition accuracy for diverse pronunciations while managing complexity through structured organization
Solution Approach 2:
The patent adds a new dimension to the phonetic transcription library by incorporating nationality or geographic location as an organizing category. This transforms the library from a simple flat structure into a multi-dimensional structure where phonetic transcriptions are indexed not only by phrase but also by regional variant, enabling the system to select appropriate pronunciations based on contextual information about the speaker's origin
2Adaptability or versatility
If the speech recognition library incorporates multiple pronunciations of phrases, then the adaptability to different nationalities and locations is improved, but the library size and processing complexity increase
Solution Approach 1:
The patent implements a universal phonetic transcription library structure that serves multiple functions: it stores phonetic transcriptions for multiple nationalities and geographic locations, supports different pronunciation variants for the same phrase, and enables the speech recognition system to adapt to diverse speakers. This multi-functional design allows the library to handle various pronunciation scenarios without requiring separate systems for each nationality or region
Solution Approach 2:
The patent performs preliminary organization of phonetic transcriptions by nationality and geographic location during library construction. By pre-segmenting and categorizing pronunciation variants before actual speech recognition operations, the system avoids the need for complex real-time analysis to determine which pronunciation variant to use, thereby reducing processing complexity during runtime while maintaining comprehensive coverage of multiple pronunciations
Data Source
AI summary
Methods and apparatus to generate a speech recognition library for use by a speech recognition system are disclosed. An example method comprises identifying a plurality of video segments having closed caption data corresponding to a phrase, the plurality of video segments associated with respective ones of a plurality of audio data segments, computing a plurality of difference metrics between a baseline audio data segment associated with the phrase and respective ones of the plurality of audio data segments, selecting a set of the plurality of audio data segments based on the plurality of difference metrics, identifying a first one of the audio data segments in the set as a representative audio data segment, determining a first phonetic transcription of the representative audio data segment, and adding the first phonetic transcription to a speech recognition library when the first phonetic transcription differs from a second phonetic transcription associated with the phrase in the speech recognition library.


