Speech Recognition Modified Frame Using Discarded Feature Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies discard significant information from feature vectors, leading to reduced performance in recognizing utterances, as most data in the feature vector is projected out during frame generation, resulting in lost information that could aid the recognizer.
Innovation Solution
The proposed solution involves generating a modified frame by combining information from multiple frames using displacement matrices and radial basis functions, incorporating additional features from discarded dimensions to enhance speech recognition performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If feature vectors are projected to generate frames using traditional LDA technique, then computational complexity is reduced and processing speed is improved, but significant information is lost leading to reduced speech recognition accuracy
Solution Approach 1:
The patent recovers discarded features by generating multiple frames from different projections of the feature vector. Instead of permanently discarding features during LDA projection, the system creates multiple frames (y0, y1, y2) from different projection matrices, then combines them through displacement operations to reconstruct and utilize the previously lost information, thereby improving speech recognition accuracy
Solution Approach 2:
The patent transforms the problem from a single projection dimension to multiple dimensions by creating frames from different projection matrices. The displacement matrix M and radial basis function operate in this expanded dimensional space, allowing the system to capture speech characteristics that would be lost in any single projection, thus recovering information without increasing the final frame dimensionality
2Measurement precision
If multiple frames are generated and combined using displacement matrices and radial basis functions, then speech recognition accuracy is improved, but computational complexity and processing time increase
Solution Approach 1:
The patent performs preliminary actions by pre-computing multiple projection matrices and generating multiple frames (y0, y1, y2) before the actual speech recognition decision. The displacement matrix M and radial basis function transformations are calculated in advance, allowing the system to prepare enhanced feature representations ahead of time, thereby reducing the computational burden during real-time recognition
Solution Approach 2:
The patent merges multiple frames (y0, y1, y2) from different projections into a single enhanced frame x' through the displacement operation x' = y0 + M*φh(y1). This combining process integrates information from multiple sources while producing a single consolidated output, avoiding the need to process multiple separate frames independently and thus reducing overall computational complexity
Data Source
AI summary
Methods and apparatus related to speech recognition performed by a speech recognition device are disclosed. The speech recognition device can receive a plurality of samples corresponding to an utterance and generate a feature vector z from the plurality of samples. The speech recognition device can select a first frame y0 from the feature vector z, and can generate a second frame y1, where y0 and y1 differ. The speech recognition device can generate a modified frame x′ based on the first frame y0 and the second frame y1 and then recognize speech related to the utterance based on the modified frame x′. The recognized speech can be output by the speech recognition device.


