Dual-Channel ASR Spectrogram Balancing for Signal Variation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic speech recognition technologies face challenges in accurately transcribing dual-channel voice data due to variations in signal characteristics between channels, leading to reduced predictive accuracy and increased computational resources needed for training.

Innovation Solution

The method involves pre-processing dual-channel voice data by applying fast Fourier transformation, generating spectrograms, determining peak power, and creating balanced input spectrograms to equalize signal characteristics, which are then used by acoustic and language machine learning models to generate character probabilities and transcription outputs, improving predictive accuracy and training efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing automatic speech recognition technology is used to transcribe dual-channel voice data, then the system can process audio files, but the predictive accuracy is reduced due to variations in signal characteristics between channels

Engineering Contradiction:
Improvepredictive accuracyVSAvoidsignal characteristic variation
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies parameter changes by transforming the voice data from time domain to frequency domain using Fast Fourier Transform (FFT), and then to spectrogram representation. This parameter transformation equalizes the signal characteristics between channels, allowing the ASR model to achieve higher predictive accuracy by processing data in a standardized frequency-based format rather than raw time-domain signals with varying characteristics.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces spectrograms as an intermediary representation between the raw dual-channel voice data and the ASR model. The spectrogram serves as a mediator that converts variable signal characteristics into a standardized visual representation of frequency content over time, enabling the model to process both channels uniformly and improve predictive accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If existing automatic speech recognition technology processes dual-channel voice data without pre-processing, then the processing pipeline is simpler, but computational resources and training data requirements increase

Engineering Contradiction:
Improvetraining speedVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent applies preliminary action by performing Fast Fourier Transform and spectrogram generation as pre-processing steps before the ASR model training. This preliminary transformation of the dual-channel voice data into balanced spectrograms prepares the data in advance, allowing the model to train faster and more efficiently without requiring excessive computational resources during the actual training process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12154589B2Systems and methods for processing bi-mode dual-channel sound data for automatic speech recognition models
Publication Date: 2024.11.26 OPTUM INC
  • US12154589B2 patent drawing
  • US12154589B2 patent drawing
  • US12154589B2 patent drawing

AI summary

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for pre-processing dual-channel voice data for an automatic speech recognition mode. The method comprises creating one or more spectrograms for each channel of the dual-channel voice data by applying fast Fourier transform and generating power spectral density. The one or more balanced power spectrograms are created by merging the spectrograms of the channels, and are provided as input for acoustic and language processing by an automatic speech recognition machine learning model.