Audio Stream Reconstruction for ASR Accuracy and Lower Bandwidth
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing techniques modify digital audio streams irreversibly, affecting downstream applications like ASR and voice authentication, and require selecting between modified and unmodified streams, leading to suboptimal user experiences.
Innovation Solution
Generate audio streams from modified streams along with representation of modifications, enabling reconstruction of unmodified streams for downstream applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If audio streams are modified to improve user experience, then audio quality is improved, but downstream applications like ASR and voice authentication suffer from reduced accuracy
Solution Approach 1:
The system segments the audio stream into multiple versions: modified audio streams for user consumption and unmodified audio streams for downstream applications. This segmentation allows each version to serve its specific purpose optimally without compromising the other.
Solution Approach 2:
Different parts of the audio processing system use different versions of the audio stream. The modified stream is used where user experience is prioritized (audio output), while the unmodified stream is used where accuracy is critical (ASR, voice authentication).
2Measurement precision
If unmodified audio streams are transmitted to downstream applications, then accuracy is improved, but network bandwidth requirements double
Solution Approach 1:
Instead of transmitting both modified and unmodified audio streams, the system transmits the modified audio stream and generates a synthetic unmodified version at the receiving end. This copying approach reduces network bandwidth usage while maintaining the accuracy needed for downstream applications.
Solution Approach 2:
The system adds a temporal dimension to the audio processing by generating the unmodified audio stream synthetically at a later stage rather than transmitting it directly. This allows the system to work with fewer simultaneous audio stream copies.
3Quantity of substance
If modified audio streams are used for all applications, then network bandwidth is reduced, but downstream applications experience degraded performance
Solution Approach 1:
The system introduces an intermediary component that synthesizes the unmodified audio stream from the modified audio stream. This intermediary generates the necessary unmodified audio data locally, eliminating the need to transmit separate unmodified streams while maintaining downstream application performance.
4Measurement precision
If both modified and unmodified audio streams are transmitted, then downstream application accuracy is maintained, but resource consumption increases
Solution Approach 1:
The receiving system performs self-service by generating the unmodified audio stream synthetically from the received modified audio stream. This eliminates the need to receive and process separate unmodified audio streams, reducing resource consumption while maintaining accuracy.
Data Source
AI summary
Techniques for generating audio streams from modified audio streams and information about the modifications to the audio streams are provided. In an example method, a computing system joins a first client device to a video conference, to which a number of client devices are connected. The computing system receives, from the first client device, a modified first audio stream including information about the modification to the first audio stream. The computing system generates a second audio stream using the modified first audio stream and the information about the modification to the first audio stream. The computing system outputs the second audio stream.


