Audio Stream Segmentation for Multi-Source Clarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in accurately recording and transcribing multi-party interactions, as they struggle to distinguish between multiple audio inputs from different sources, often resulting in confusion when multiple parties speak simultaneously and cross-talk occurs.

Innovation Solution

A computer system is configured to receive audio input from multiple sources, filter based on amplitude, and identify the primary source of audio, combining inputs into distinct speech blocks while discarding incidental captures, and displaying this information in real-time for user interpretation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple audio input devices are used to capture audio from multiple parties, then the ability to record multi-party interactions is improved, but the difficulty of determining which party is saying what increases

Engineering Contradiction:
Improvenumber of audio input devicesVSAvoidclarity of audio source attribution
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the multi-party audio stream by creating separate audio channels for each detected speaker. The system identifies distinct voice sources and routes their audio to separate channels, allowing each party's speech to be independently tracked and attributed. This segmentation resolves the confusion of mixed audio by dividing it into discrete, identifiable segments corresponding to individual speakers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that analyzes audio characteristics (such as voice patterns, pitch, and timbre) to identify and distinguish between different speakers. This intermediary system acts as a mediator between the raw multi-source audio input and the final output, adding metadata that attributes each audio segment to its source party without requiring physical separation of microphones.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If multiple audio input devices are used to capture audio from multiple parties, then the coverage of audio capture is improved, but the complexity of managing and processing audio streams increases

Engineering Contradiction:
Improvenumber of audio input devicesVSAvoidcomplexity of audio stream management
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent merges the management of multiple audio streams into a unified processing framework. Instead of handling each microphone's audio independently, the system combines all audio inputs into a single multi-channel stream that is processed together. This merging approach simplifies management by treating multiple sources as one integrated system, reducing the complexity of coordinating separate audio workflows.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements a universal audio processing system that can handle multiple audio sources, different speech patterns, and various output formats through a single platform. The system is designed to be multi-functional, capable of simultaneously performing audio capture, speaker identification, stream separation, and output generation across diverse scenarios without requiring separate specialized systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If audio from multiple sources is combined without filtering, then the completeness of audio recording is improved, but the accuracy of identifying primary speakers decreases

Engineering Contradiction:
Improvecompleteness of audio captureVSAvoidaccuracy of speaker identification
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies preliminary filtering and analysis to audio streams before they are fully processed and output. The system performs initial speaker identification and audio quality assessment on all incoming streams, pre-sorting and pre-tagging audio segments with source information. This preliminary action ensures that when audio is combined, the primary speakers are already identified and marked, maintaining accuracy while preserving completeness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent utilizes parameter changes in audio signals (such as variations in volume, frequency spectrum, and speech patterns) to distinguish primary speakers from incidental captures. By analyzing these parameters, the system can identify which audio sources represent intentional speech versus background or incidental captures, thereby maintaining speaker identification accuracy while retaining complete audio coverage.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8719032B1Methods for presenting speech blocks from a plurality of audio input data streams to a user in an interface
Publication Date: 2014.05.06 JEFFERSON AUDIO VIDEO SYSTEMS INC
  • US8719032B1 patent drawing
  • US8719032B1 patent drawing
  • US8719032B1 patent drawing

AI summary

A clear picture of who is speaking in a setting where there are multiple input sources (e.g., a conference room with multiple microphones) can be obtained by comparing input channels against each other. The data from each channel can not only be compared, but can also be organized into portions which logically correspond to statements by a user. These statements, along with information regarding who is speaking, can be presented in a user friendly format via an interactive timeline which can be updated in real time as new audio input data is received.