Audio Stream Reconstruction for ASR Accuracy and Lower Bandwidth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing techniques modify digital audio streams irreversibly, affecting downstream applications like ASR and voice authentication, and require selecting between modified and unmodified streams, leading to suboptimal user experiences.

Innovation Solution

Generate audio streams from modified streams along with representation of modifications, enabling reconstruction of unmodified streams for downstream applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If audio streams are modified to improve user experience, then audio quality is improved, but downstream applications like ASR and voice authentication suffer from reduced accuracy

Engineering Contradiction:
Improveuser experienceVSAvoiddownstream application accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system segments the audio stream into multiple versions: modified audio streams for user consumption and unmodified audio streams for downstream applications. This segmentation allows each version to serve its specific purpose optimally without compromising the other.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different parts of the audio processing system use different versions of the audio stream. The modified stream is used where user experience is prioritized (audio output), while the unmodified stream is used where accuracy is critical (ASR, voice authentication).

Inventive Principle:
Principle #3Local quality

2Measurement precision

If unmodified audio streams are transmitted to downstream applications, then accuracy is improved, but network bandwidth requirements double

Engineering Contradiction:
Improvedownstream application accuracyVSAvoidnetwork bandwidth
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

Instead of transmitting both modified and unmodified audio streams, the system transmits the modified audio stream and generates a synthetic unmodified version at the receiving end. This copying approach reduces network bandwidth usage while maintaining the accuracy needed for downstream applications.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system adds a temporal dimension to the audio processing by generating the unmodified audio stream synthetically at a later stage rather than transmitting it directly. This allows the system to work with fewer simultaneous audio stream copies.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Quantity of substance

If modified audio streams are used for all applications, then network bandwidth is reduced, but downstream applications experience degraded performance

Engineering Contradiction:
Improvenetwork bandwidthVSAvoiddownstream application performance
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system introduces an intermediary component that synthesizes the unmodified audio stream from the modified audio stream. This intermediary generates the necessary unmodified audio data locally, eliminating the need to transmit separate unmodified streams while maintaining downstream application performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If both modified and unmodified audio streams are transmitted, then downstream application accuracy is maintained, but resource consumption increases

Engineering Contradiction:
Improvedownstream application accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The receiving system performs self-service by generating the unmodified audio stream synthetically from the received modified audio stream. This eliminates the need to receive and process separate unmodified audio streams, reducing resource consumption while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20260019507A1Generating audio streams from modified audio streams and information about the modifications to the audio streams
Publication Date: 2026.01.15 ZOOM VIDEO COMM INC
  • US20260019507A1 patent drawing
  • US20260019507A1 patent drawing
  • US20260019507A1 patent drawing

AI summary

Techniques for generating audio streams from modified audio streams and information about the modifications to the audio streams are provided. In an example method, a computing system joins a first client device to a video conference, to which a number of client devices are connected. The computing system receives, from the first client device, a modified first audio stream including information about the modification to the first audio stream. The computing system generates a second audio stream using the modified first audio stream and the information about the modification to the first audio stream. The computing system outputs the second audio stream.