Automatic Speech Recognition Using Phonological Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic speech recognition systems face high computational loads and are inefficient in recognizing speech with strong regional or national accents, or from speakers with speech difficulties, as they rely on statistical and template matching methods that require extensive processing and training on specific user voices.
Innovation Solution
A method that identifies phonological features within acoustic signals by determining acoustic parameters such as root mean square amplitude, fundamental frequency, and formant frequencies, and separates these features into zones for comparison with a stored lexicon to identify words, using a sequence of phonological segments and penalty calculations to match with lexical entries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If statistical and template matching systems are used for speech recognition, then recognition accuracy can be maintained through training on user exemplars, but computational load increases significantly and training is required for each user
Solution Approach 1:
The speech signal is divided into time windows and further segmented into phonological segments based on acoustic parameters. Each segment is analyzed independently for phonological feature extraction, reducing the computational burden compared to analyzing the entire speech signal as a whole. This segmentation allows the system to process speech in manageable chunks without losing recognition accuracy.
Solution Approach 2:
The patent extracts only the essential phonological features from the acoustic signal rather than processing the complete spectral information. By identifying and extracting key phonological segments and their characteristics, the system reduces computational load while maintaining the ability to accurately recognize speech, eliminating the need for extensive template matching.
2Adaptability or versatility
If conventional speech recognition systems are used, then they can recognize standard speech patterns, but they fail to correctly recognise speech with strong regional or national accents or speech difficulties
Solution Approach 1:
The system changes the parameter space from traditional spectral features to phonological feature space. By representing speech in terms of phonological segments and their features rather than raw acoustic spectra, the system becomes more adaptable to different accents and speech patterns. This parameter transformation allows the same phonological representation to capture variations in pronunciation across different accents while maintaining recognition accuracy.
3Measurement precision
If detailed spectral analysis is performed on the speech signal, then recognition accuracy improves, but processing time and computational resources increase
Solution Approach 1:
The patent extracts only the essential phonological information from the speech signal by identifying phonological segments and their features, rather than performing complete spectral analysis. This extraction approach maintains measurement precision for the relevant phonological characteristics while significantly reducing processing time by avoiding unnecessary computational steps.
Solution Approach 2:
The speech signal is segmented into phonological units and processed in discrete time windows, allowing precise analysis of each segment's phonological features without the computational overhead of continuous detailed spectral analysis. This segmentation enables efficient processing while maintaining precision where it matters most.
Data Source
AI summary
The invention provides a method of automatic speech recognition. The method includes receiving a speech signal, dividing the speech signal into time windows, for each time window determining acoustic parameters of the speech signal within that window, and identifying phonological features from the acoustic parameters, such that a sequence of phonological features are generated for the speech signal, separating the sequence of phonological features into a sequence of zones, and comparing the sequences of zones to a lexical entry comprising a sequence of phonological segments to a stored lexicon to identify one or more words in the speech signal.


