Multitrack Telephony Recording for Accurate Transcription

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current conference call recording technologies typically mix the audio of all participants, making it difficult to separate individual contributions for transcription, especially when using speech recognition software, and also pose challenges in isolating the contributions of a single participant.

Innovation Solution

A system and method for generating multitrack recordings of telephony communications, where each participant's audio contributions can be recorded in separate tracks, allowing for easy identification and transcription of individual participants' spoken words, using a media bridge and recording unit with an API that enables customers to specify recording formats and participant identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If mixed audio recording is used for conference calls, then recording simplicity is improved, but transcription accuracy of individual participants deteriorates

Engineering Contradiction:
Improverecording simplicityVSAvoidtranscription accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent divides the conference call audio into multiple separate tracks, with each track containing the audio of a specific participant. This segmentation allows transcription software to process individual participant audio separately, significantly improving transcription accuracy while maintaining recording simplicity through automated track assignment based on participant identifiers.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If speech recognition software is applied to mixed audio, then transcription process is simplified, but transcription accuracy of individual contributions deteriorates

Engineering Contradiction:
Improvetranscription process simplicityVSAvoidtranscription accuracy of individual contributions
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the audio into separate participant tracks, enabling speech recognition software to process each track independently. This approach maintains the simplicity of automated transcription while dramatically improving the accuracy of attributing spoken contributions to the correct participant, as the software no longer needs to separate overlapping voices in mixed audio.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If single participant isolation is attempted from mixed audio, then transcription of specific participant is improved, but system complexity increases

Engineering Contradiction:
Improvetranscription of single participantVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary separation of participant audio into distinct tracks during the recording phase, using participant identifiers (such as phone numbers or account IDs) to assign audio to specific tracks. This preliminary action eliminates the need for complex post-processing to isolate individual participants, as the separation is already accomplished before transcription begins.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11838442B2System and methods for creating multitrack recordings
Publication Date: 2023.12.05 VONAGE BUSINESS INC
  • US11838442B2 patent drawing
  • US11838442B2 patent drawing
  • US11838442B2 patent drawing

AI summary

Systems and methods for making a multitrack recording of a telephony communication, such as a conference call, record the contributions of each participant its own respective, separate recording track. In some instances, the contribution(s) of one or more participants is recorded in separate recording tracks, and the contributions of multiple other participants is mixed and recorded in a single recording track. An organizer or administrator of a telephony communication, such as a conference call, can instruct a multitrack recording system as to how to format a multitrack recording of the telephony communication via commands submitted through an application programming interface (API).