Audio Bit-Size Conversion Using Neural Accelerators for Low-Power ASR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech recognition (ASR) systems, particularly keyphrase detection, face inefficiencies due to the need for digital signal processors (DSPs) to convert high-definition audio signals from digital microphones, which increases power consumption and reduces accuracy by truncating bits, making them unsuitable for small devices.
Innovation Solution
An audio input sample bit-size conversion method using a neural network accelerator that scales 24-bit audio samples to 16-bit samples by dividing into high and low sample parts, applying gains, and recombining them to maintain signal integrity, allowing neural network accelerators to handle the conversion efficiently without DSPs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If DSP is used to convert 24-bit audio samples to 16-bit samples, then bit-size conversion is achieved, but power consumption increases and efficiency decreases
Solution Approach 1:
The patent replaces the DSP-based mechanical conversion system with a neural network accelerator that performs bit-size conversion through learned transformations. The neural network processes 24-bit audio samples and outputs 16-bit samples directly, eliminating the need for complex DSP algorithms and significantly reducing power consumption while maintaining conversion accuracy.
Solution Approach 2:
The patent changes the operational parameters of the audio processing system by transitioning from fixed-function DSP operations to adaptive neural network operations. The neural network dynamically adjusts its processing based on input characteristics, enabling efficient bit-size conversion with optimized power consumption compared to traditional DSP methods.
2Measurement precision
If DSP is used to perform complex conversion algorithms, then 24-bit to 16-bit conversion is achieved, but device efficiency lowers and power consumption rises
Solution Approach 1:
The patent substitutes the DSP-based conversion algorithm with a neural network accelerator that performs the same 24-bit to 16-bit conversion task. The neural network, trained on audio data, achieves accurate conversion while operating more efficiently on the hardware accelerator, thereby improving overall system productivity and reducing power consumption.
Solution Approach 2:
The patent applies preliminary training to the neural network using extensive audio datasets before deployment. This pre-training enables the network to perform accurate bit-size conversion directly during runtime without requiring complex DSP algorithms, thus improving processing efficiency and reducing the computational burden on the device.
3Reliability
If traditional ASR techniques are used, then speech recognition is achieved, but power consumption is too high for small stand-alone devices
Solution Approach 1:
The patent segments the ASR system into distinct functional components: a neural network accelerator for feature extraction and acoustic scoring, and a separate decoding component. This segmentation allows the power-intensive neural network operations to be performed efficiently on dedicated hardware, reducing overall power consumption while maintaining recognition accuracy on small devices.
Solution Approach 2:
The patent replaces traditional DSP-based feature extraction and acoustic scoring with a neural network accelerator. This substitution enables accurate speech recognition to be performed with significantly lower power consumption, making ASR viable for small stand-alone devices with limited power resources.
Data Source
AI summary
A method, system, and device are directed to audio input bit-size conversion for compatibility to audio processing systems with an expected input sample bit-size.


