Harmonic Distortion Augmentation for MEMS Microphone Arrays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech processing systems in automated clinical documentation face challenges due to imperfections in microelectro-mechanical system (MEMS) microphones, such as sensitivity variations, self-noise, frequency response mismatches, and harmonic distortions, which degrade speech processing performance and require costly calibration processes.
Innovation Solution
The method involves data augmentation techniques to generate harmonic distortion-based augmented signals by determining and accounting for microphone-specific imperfections, using harmonic distortion parameters and coefficients to adjust and enhance audio signals, thereby improving robustness to microphone system imperfections without the need for extensive calibration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional speech processing systems use MEMS microphones, then cost and device complexity are reduced, but microphone imperfections such as sensitivity variations, self-noise, frequency response mismatches, and harmonic distortions degrade speech processing performance
Solution Approach 1:
The patent creates virtual copies of microphone imperfections through data augmentation. Instead of physically calibrating each microphone to compensate for imperfections, the system generates synthetic audio signals that replicate the characteristic distortions (harmonic distortion, self-noise, frequency response) of real MEMS microphones. These augmented signals are then used to train speech processing models, allowing them to learn robust performance despite the inherent imperfections of cost-effective MEMS devices.
Solution Approach 2:
The patent modifies audio signal parameters to simulate microphone imperfections. By adjusting parameters such as adding harmonic distortion components, introducing self-noise, and applying frequency response filtering to training data, the system transforms perfect audio signals into representations that mimic real MEMS microphone behavior. This parameter transformation enables the training of robust speech processing systems without requiring expensive calibrated microphones.
2Measurement precision
If calibration processes are used to compensate for microphone imperfections, then speech processing accuracy is improved, but the process becomes costly and not feasible at large scale
Solution Approach 1:
The patent replaces the physical calibration process with a computational approach. Instead of performing time-consuming mechanical calibration procedures to measure and compensate for microphone imperfections, the system uses data augmentation techniques to virtually replicate these imperfections in training data. This substitution eliminates the need for complex calibration hardware and procedures while achieving the same goal of improving speech processing accuracy.
Solution Approach 2:
The patent performs preliminary data augmentation during the training phase to pre-compensate for microphone imperfections. By generating augmented training signals that include simulated microphone distortions before the actual speech processing task, the system prepares the model in advance to handle real-world imperfections. This preliminary action eliminates the need for post-deployment calibration and reduces operational complexity.
3Reliability
If data augmentation is applied to account for microphone imperfections, then robustness to microphone system imperfections is improved, but computational processing requirements increase
Solution Approach 1:
The patent applies partial data augmentation by selectively processing only the most critical aspects of audio signals. Instead of augmenting every possible parameter and condition exhaustively, the system focuses on the primary microphone imperfections (harmonic distortion, self-noise, frequency response) and applies targeted transformations. This partial action approach reduces computational burden while still achieving robustness to the most impactful microphone imperfections.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method, computer program product, and computing system for receiving a signal from each microphone of a plurality of microphones, thus defining a plurality of signals. Harmonic distortion associated with at least one microphone may be determined. One or more harmonic distortion-based augmentations may be performed on the plurality of signals based upon, at least in part, the harmonic distortion associated with the at least one microphone, thus defining one or more harmonic distortion-based augmented signals.