Dual-Channel ASR Spectrogram Balancing for Signal Variation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic speech recognition technologies face challenges in accurately transcribing dual-channel voice data due to variations in signal characteristics between channels, leading to reduced predictive accuracy and increased computational resources needed for training.
Innovation Solution
The method involves pre-processing dual-channel voice data by applying fast Fourier transformation, generating spectrograms, determining peak power, and creating balanced input spectrograms to equalize signal characteristics, which are then used by acoustic and language machine learning models to generate character probabilities and transcription outputs, improving predictive accuracy and training efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing automatic speech recognition technology is used to transcribe dual-channel voice data, then the system can process audio files, but the predictive accuracy is reduced due to variations in signal characteristics between channels
Solution Approach 1:
The patent applies parameter changes by transforming the voice data from time domain to frequency domain using Fast Fourier Transform (FFT), and then to spectrogram representation. This parameter transformation equalizes the signal characteristics between channels, allowing the ASR model to achieve higher predictive accuracy by processing data in a standardized frequency-based format rather than raw time-domain signals with varying characteristics.
Solution Approach 2:
The patent introduces spectrograms as an intermediary representation between the raw dual-channel voice data and the ASR model. The spectrogram serves as a mediator that converts variable signal characteristics into a standardized visual representation of frequency content over time, enabling the model to process both channels uniformly and improve predictive accuracy.
2Productivity
If existing automatic speech recognition technology processes dual-channel voice data without pre-processing, then the processing pipeline is simpler, but computational resources and training data requirements increase
Solution Approach 1:
The patent applies preliminary action by performing Fast Fourier Transform and spectrogram generation as pre-processing steps before the ASR model training. This preliminary transformation of the dual-channel voice data into balanced spectrograms prepares the data in advance, allowing the model to train faster and more efficiently without requiring excessive computational resources during the actual training process.
Data Source
AI summary
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for pre-processing dual-channel voice data for an automatic speech recognition mode. The method comprises creating one or more spectrograms for each channel of the dual-channel voice data by applying fast Fourier transform and generating power spectral density. The one or more balanced power spectrograms are created by merging the spectrograms of the channels, and are provided as input for acoustic and language processing by an automatic speech recognition machine learning model.


