IP Telephony Audio Fidelity via Textual Backup

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

IP telephony systems face challenges in maintaining audio quality due to data packet loss and delay, particularly when high compression is used, leading to poor sound fidelity during transmissions over data networks.

Innovation Solution

The system transmits a textual representation of spoken audio input alongside the digital data created by a CODEC, allowing the receiving device to use this textual data as a backup to fill in missing or corrupted CODEC data, and employs error correction or redundancy to ensure the textual data's integrity, thereby improving audio quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If high compression is employed by CODECs to reduce data transmission size, then bandwidth efficiency is improved, but sound quality and audio fidelity deteriorate

Engineering Contradiction:
Improvebandwidth efficiencyVSAvoidsound quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The system performs speech-to-text conversion on the original spoken audio input before transmission, creating a textual representation that serves as preliminary data. This textual data is transmitted alongside or instead of traditional audio CODEC data, enabling the receiving device to reconstruct audio quality without requiring high-bitrate audio transmission, thus resolving the contradiction between bandwidth efficiency and sound quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a textual copy (transcription) of the spoken audio input that can be used to reconstruct the original audio at the receiving end. This textual copy serves as an alternative representation that preserves information while requiring less transmission bandwidth, allowing high compression ratios without sacrificing perceptual audio quality

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If data packets are transmitted over public data networks to enable IP telephony communications, then communication versatility is improved, but data packet loss and delay occur leading to poor audio fidelity

Engineering Contradiction:
Improvecommunication versatilityVSAvoidaudio fidelity
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system introduces textual representation data as an intermediary between the transmitted audio data and the final audio reconstruction. This textual intermediary can be used to fill in missing or corrupted audio data packets, ensuring reliable audio fidelity even when data packets are lost or delayed during transmission over public networks

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the data representation parameter from traditional audio CODEC formats to speech-to-text transcription formats. This parameter change enables the system to withstand packet loss and delay better, as textual data can be more easily reconstructed or regenerated from the transcription, thereby improving reliability while maintaining communication versatility

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9767802B2Methods and apparatus for conducting internet protocol telephony communications
Publication Date: 2017.09.19 VONAGE BUSINESS INC
  • US9767802B2 patent drawing
  • US9767802B2 patent drawing
  • US9767802B2 patent drawing

AI summary

IP telephony communications are conducted by sending both audio data produced by a CODEC that represents received spoken audio input, and a textual representation of the spoken audio input. A receiving device utilizes the textual representation of the spoken audio input to help recreate the spoken audio input when a portion of the CODEC data is missing. The textual representation can be generated by a speech-to-text function. Alternatively, the textual representation can be a notation of extracted phonemes.