Audio Processing Device for Speech Recognition with Compressed Power Parameters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing devices for speech recognition face challenges in reducing hardware costs and circuit efficiency due to the high bit width requirements for speech feature extraction, particularly in implementing Mel-scale Frequency Cepstral Coefficients (MFCC) processing.
Innovation Solution
The proposed audio processing device incorporates a memory circuit, a power logarithmic circuit, and a Mel filter circuit that perform sequential operations to compress and process audio data, reducing the bit width of power spectrum parameters and sharing memory resources, thereby minimizing hardware costs and circuit area.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high bit width is used for speech feature extraction, then speech recognition accuracy is improved, but hardware cost and circuit area increase
Solution Approach 1:
The speech processing pipeline is divided into separate functional modules (FFT module, power spectrum module, Mel filterbank module, DCT module) that operate sequentially. Each module processes data with appropriate bit width for its specific function, rather than using high bit width throughout the entire system. This segmentation allows precision to be applied only where necessary while reducing overall hardware requirements.
Solution Approach 2:
The patent applies logarithmic compression to the power spectrum parameters, transforming the data representation from linear to logarithmic scale. This parameter change reduces the dynamic range and allows for lower bit width representation while maintaining speech recognition accuracy. The logarithmic transformation is particularly effective in reducing the bit width required for representing power spectrum values.
2Productivity
If multiple hardware circuit modules are used for speech feature extraction, then processing capability is improved, but manufacturing cost increases
Solution Approach 1:
The system is divided into distinct functional modules (FFT, power spectrum, Mel filterbank, DCT) that can be implemented using the same hardware resources at different times. This modular segmentation allows for efficient resource utilization while maintaining processing capability.
Solution Approach 2:
The patent implements a time-division multiplexing strategy where a single set of hardware resources (including memory circuits and processing units) serves multiple functions by processing different stages of speech feature extraction at different time periods. This multi-functionality reduces the total number of hardware components needed, thereby lowering manufacturing cost while maintaining full processing capability.
3Measurement precision
If memory circuit bit width is increased, then speech feature extraction precision is improved, but memory space requirement increases
Solution Approach 1:
Logarithmic compression is applied to the power spectrum parameters before storing them in memory, reducing the bit width required for memory storage while preserving the essential information for speech recognition. This parameter transformation allows high-precision feature extraction with reduced memory requirements.
Solution Approach 2:
The memory system is designed to handle different data types with different bit width requirements at different processing stages. By segmenting the memory usage according to processing stage and data type, the system optimizes memory space utilization while maintaining the precision needed for accurate speech feature extraction.
Data Source
AI summary
An audio processing device for speech recognition is provided, which includes a memory circuit, a power spectrum transfer circuit, and a feature extraction circuit. The power spectrum transfer circuit is coupled to the memory circuit, reads frequency spectrum coefficients of time-domain audio sample data from the memory circuit, generates compressed power parameters by performing a power spectrum transfer processing and a compressing processing according to the frequency spectrum coefficients, and writes the compressed power parameters into the memory circuit. The feature extraction circuit is coupled to the memory circuit, reads the compressed power parameters from the memory circuit, generates an audio feature vector by performing mel-filtering and frequency-to-time transfer processing according to the compressed power parameters. The bit width of the compressed power parameters is less than the bit width of the frequency spectrum coefficients.


