Audio Source Separation Using Unified Acoustic and Language Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional systems for audio source separation in audio files, such as those used in Air Traffic Control communications, are inefficient and often fail to produce accurate results, especially when the time between utterances is short, and require non-speech breaks between speakers.

Innovation Solution

The use of machine learning and artificial intelligence to unify audio separation and automatic speech recognition techniques, training low-level classifiers and high-level clustering components to separate and identify audio sources without requiring non-speech breaks, using both acoustic and language information to handle complex patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional DSP techniques and simple rules are used for audio source separation, then the system is easier to implement, but the accuracy deteriorates when the time between utterances is short

Engineering Contradiction:
Improveease of implementationVSAvoidsource separation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent combines acoustic modeling with language modeling to create a unified speech recognition system. The acoustic model processes audio features while the language model provides contextual information, and their probabilities are combined to improve overall recognition accuracy, especially in challenging conditions with short utterance intervals.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system dynamically adjusts the duration of silence periods and the time windows used for processing based on the detected speech activity. By adapting these temporal parameters according to the actual speech patterns in the audio stream, the system maintains high accuracy even when utterances are closely spaced.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If traditional systems require non-speech breaks between speakers, then the processing is simpler, but the productivity deteriorates due to inability to handle continuous speech

Engineering Contradiction:
Improveprocessing complexityVSAvoidthroughput of continuous speech
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent implements continuous speech recognition by processing audio frames in a continuous stream without requiring silence breaks between speakers. The system maintains speech activity detection and recognition processes running continuously, allowing overlapping and back-to-back utterances to be processed seamlessly, thereby increasing throughput.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system performs preliminary speech activity detection on incoming audio frames before full recognition processing. This allows the system to prepare and buffer processing resources in advance, enabling continuous processing of speech without gaps while maintaining manageable complexity through staged processing.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If machine learning and AI techniques are used to unify audio separation and speech recognition, then the accuracy is improved, but the device complexity increases

Engineering Contradiction:
Improvesource separation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the complex audio processing task into distinct modular components: acoustic feature extraction, speech activity detection, acoustic modeling, language modeling, and probability combination. Each module performs a specific function and can be independently optimized and trained, reducing overall system complexity while maintaining high accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The unified system performs multiple functions simultaneously: it separates audio sources, performs speech recognition, and provides speaker identification using a single integrated machine learning framework. This multi-functionality reduces the need for separate specialized systems while achieving high accuracy across all tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240404528A1Systems and methods for separating and identifying audio in an audio file using machine learning
Publication Date: 2024.12.05 THE MITRE CORPORATION
  • US20240404528A1 patent drawing
  • US20240404528A1 patent drawing
  • US20240404528A1 patent drawing

AI summary

Disclosed herein are systems and methods for processing an audio file to perform audio Segmentation and Speaker Role Identification (SRID) by training low level classifier and high level clustering components to separate and identify audio from different sources in an audio file by unifying audio separation and automatic speech recognition (ASR) techniques in a single system. Segmentation and SRID can include separating audio in an audio file into one or more segments, based on a determination of the identity of the speaker, category of the speaker, or source of audio in the segment. In one or more examples, the disclosed systems and methods use machine learning and artificial intelligence technology to determine the source of segments of audio using a combination of acoustic and language information. In some examples, the acoustic and language information is used to classify audio in each frame and cluster the audio into segments.