Microphone Unit With Integrated Speech Coder For Low Power Transmission
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital microphone systems face high power consumption due to inefficient data bit rate on interfaces, particularly in Always-On Voice modes, which limits battery life in portable devices while requiring sufficient information for speech recognition and keyword detection.
Innovation Solution
Incorporating a microphone unit with on-chip or co-packaged integrated speech coder circuitry that transmits compressed speech data at a low bit rate, using speech coding techniques such as MDCT, CELP, or MFCC to reduce power consumption during data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If digital microphone data is transmitted at high bit rate to maintain speech recognition accuracy, then downstream processing performance is improved, but power consumption increases significantly
Solution Approach 1:
The patent extracts only the essential features from the audio signal using feature extraction circuitry (such as MFCC coefficients) rather than transmitting the complete audio waveform. This selective extraction of meaningful information reduces the data volume to be transmitted while preserving speech recognition accuracy, directly resolving the contradiction between transmission quality and power consumption.
Solution Approach 2:
The patent changes the representation parameters of audio data from raw waveform samples to compressed feature parameters (e.g., MFCC coefficients). This parameter transformation significantly reduces the bit rate required for transmission while maintaining the information content needed for accurate speech recognition, thereby reducing power consumption without sacrificing accuracy.
2Use of energy by moving object
If data bit rate is reduced to save power, then power consumption is minimized, but sufficient information for speech recognition is lost
Solution Approach 1:
The patent extracts only the essential features from the audio signal using feature extraction circuitry (such as MFCC coefficients) rather than transmitting the complete audio waveform. This selective extraction of meaningful information reduces the data volume to be transmitted while preserving speech recognition accuracy, directly resolving the contradiction between transmission quality and power consumption.
Solution Approach 2:
The patent introduces an intermediary feature extraction and compression stage between the microphone and the transmission interface. This intermediary processing layer converts the raw audio signal into a compact representation that retains essential speech information, acting as a mediator that enables low-bit-rate transmission without information loss for recognition purposes.
3Ease of operation
If microphone remains permanently enabled for voice commands, then voice functionality is always available, but battery life is significantly reduced
Solution Approach 1:
The patent performs preliminary feature extraction and compression of audio data at the microphone level before transmission to the main processing unit. This preliminary action reduces the data volume that needs to be transmitted and processed, enabling the microphone to remain permanently enabled for voice commands while significantly reducing the overall power consumption and extending battery life.
Solution Approach 2:
The patent extracts only the essential features from the audio signal using feature extraction circuitry (such as MFCC coefficients) rather than transmitting the complete audio waveform. This selective extraction of meaningful information reduces the data volume to be transmitted while preserving speech recognition accuracy, directly resolving the contradiction between transmission quality and power consumption.
Data Source
AI summary
A microphone unit has a transducer, for generating an electrical audio signal from a received acoustic signal; a speech coder, for obtaining compressed speech data from the audio signal; and a digital output, for supplying digital signals representing said compressed speech data. The speech coder may be a lossy speech coder, and may contain a bank of filters with center frequencies that are non-uniformly spaced, for example mel frequencies.


