Adaptive Speech Regeneration via Vector Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional teleconferencing systems face challenges with noise interference and data compression, leading to poor audio quality due to imperfect noise reduction techniques and bandwidth inefficiencies.
Innovation Solution
The use of an audio transformation model to extract and transmit speech and voice vector representations, which are then used to regenerate high-quality speech signals in the original voice, eliminating noise and reducing bandwidth requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional noise reduction techniques and CODECs are used, then audio signals can be transmitted, but audio quality deteriorates due to noise interference and compression artifacts
Solution Approach 1:
The patent extracts only the essential speech features (spectrogram representations) from the audio signal, separating them from the noisy components. This extraction approach transmits only the necessary information for speech regeneration, eliminating noise and compression artifacts while maintaining audio quality.
Solution Approach 2:
Instead of transmitting the original audio signal which contains noise, the patent creates a compressed representation (spectrogram) and uses an audio generation model to synthesize a clean copy of the speech signal at the receiving end. This copying process reconstructs the speech without preserving the original noise components.
2Loss of energy
If traditional CODECs compress audio signals, then bandwidth is reduced, but audio fidelity deteriorates due to lossy compression
Solution Approach 1:
The patent replaces traditional mechanical audio compression CODECs with a machine learning-based approach. Instead of using conventional compression algorithms that sacrifice fidelity, the system uses neural networks to extract essential features and regenerate high-fidelity speech, achieving both bandwidth efficiency and audio quality.
Solution Approach 2:
The patent transforms the audio signal into a different parameter space (spectrogram representations) that captures essential speech information more efficiently. This parameter transformation allows for compact representation and transmission while enabling high-fidelity regeneration through the audio generation model.
3Productivity
If audio signals are transmitted over bandwidth-limited channels, then communication is enabled, but noise interference increases and audio quality decreases
Solution Approach 1:
The system extracts only the essential speech information from the audio signal before transmission, removing noisy components. This extraction enables efficient transmission over bandwidth-limited channels while ensuring that only clean speech representations are regenerated at the receiving end.
Data Source
AI summary
Embodiments provide for employing an audio transformation model to resynthesize speech signals associated with a speaking entity. Examples can receive audio signals comprising speech signals that are captured by an audio capture device. Examples can divide the audio signals into audio segments and input the audio segments into an audio transformation model to generate a voice vector representation and a speech vector representation. The voice vector representation comprises characteristics related to a speaking voice associated with the speaking entity and the speech vector representation comprises one or more words spoken by the speaking entity. The one or more words comprised in the speech vector representation are associated with respective contextual attributes associated with the one or more words. The audio transformation model can utilize the voice vector representation and the speech vector representation to regenerate the speech signals associated with the speaking entity.


