Audio Processing Device for Speech Recognition with Compressed Power Parameters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing devices for speech recognition face challenges in reducing hardware costs and circuit efficiency due to the high bit width requirements for speech feature extraction, particularly in implementing Mel-scale Frequency Cepstral Coefficients (MFCC) processing.

Innovation Solution

The proposed audio processing device incorporates a memory circuit, a power logarithmic circuit, and a Mel filter circuit that perform sequential operations to compress and process audio data, reducing the bit width of power spectrum parameters and sharing memory resources, thereby minimizing hardware costs and circuit area.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high bit width is used for speech feature extraction, then speech recognition accuracy is improved, but hardware cost and circuit area increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcircuit area
Core Design Contradiction:
Measurement precisionVSArea of stationary object

Solution Approach 1:

The speech processing pipeline is divided into separate functional modules (FFT module, power spectrum module, Mel filterbank module, DCT module) that operate sequentially. Each module processes data with appropriate bit width for its specific function, rather than using high bit width throughout the entire system. This segmentation allows precision to be applied only where necessary while reducing overall hardware requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies logarithmic compression to the power spectrum parameters, transforming the data representation from linear to logarithmic scale. This parameter change reduces the dynamic range and allows for lower bit width representation while maintaining speech recognition accuracy. The logarithmic transformation is particularly effective in reducing the bit width required for representing power spectrum values.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If multiple hardware circuit modules are used for speech feature extraction, then processing capability is improved, but manufacturing cost increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidmanufacturing cost
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The system is divided into distinct functional modules (FFT, power spectrum, Mel filterbank, DCT) that can be implemented using the same hardware resources at different times. This modular segmentation allows for efficient resource utilization while maintaining processing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a time-division multiplexing strategy where a single set of hardware resources (including memory circuits and processing units) serves multiple functions by processing different stages of speech feature extraction at different time periods. This multi-functionality reduces the total number of hardware components needed, thereby lowering manufacturing cost while maintaining full processing capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If memory circuit bit width is increased, then speech feature extraction precision is improved, but memory space requirement increases

Engineering Contradiction:
Improvefeature extraction precisionVSAvoidmemory space
Core Design Contradiction:
Measurement precisionVSVolume of stationary object

Solution Approach 1:

Logarithmic compression is applied to the power spectrum parameters before storing them in memory, reducing the bit width required for memory storage while preserving the essential information for speech recognition. This parameter transformation allows high-precision feature extraction with reduced memory requirements.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The memory system is designed to handle different data types with different bit width requirements at different processing stages. By segmenting the memory usage according to processing stage and data type, the system optimizes memory space utilization while maintaining the precision needed for accurate speech feature extraction.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11404046B2Audio processing device for speech recognition
Publication Date: 2022.08.02 XSAIL TECH CO LTD
  • US11404046B2 patent drawing
  • US11404046B2 patent drawing
  • US11404046B2 patent drawing

AI summary

An audio processing device for speech recognition is provided, which includes a memory circuit, a power spectrum transfer circuit, and a feature extraction circuit. The power spectrum transfer circuit is coupled to the memory circuit, reads frequency spectrum coefficients of time-domain audio sample data from the memory circuit, generates compressed power parameters by performing a power spectrum transfer processing and a compressing processing according to the frequency spectrum coefficients, and writes the compressed power parameters into the memory circuit. The feature extraction circuit is coupled to the memory circuit, reads the compressed power parameters from the memory circuit, generates an audio feature vector by performing mel-filtering and frequency-to-time transfer processing according to the compressed power parameters. The bit width of the compressed power parameters is less than the bit width of the frequency spectrum coefficients.