Adaptive Speech Regeneration via Vector Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional teleconferencing systems face challenges with noise interference and data compression, leading to poor audio quality due to imperfect noise reduction techniques and bandwidth inefficiencies.

Innovation Solution

The use of an audio transformation model to extract and transmit speech and voice vector representations, which are then used to regenerate high-quality speech signals in the original voice, eliminating noise and reducing bandwidth requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional noise reduction techniques and CODECs are used, then audio signals can be transmitted, but audio quality deteriorates due to noise interference and compression artifacts

Engineering Contradiction:
Improveaudio qualityVSAvoidnoise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the essential speech features (spectrogram representations) from the audio signal, separating them from the noisy components. This extraction approach transmits only the necessary information for speech regeneration, eliminating noise and compression artifacts while maintaining audio quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of transmitting the original audio signal which contains noise, the patent creates a compressed representation (spectrogram) and uses an audio generation model to synthesize a clean copy of the speech signal at the receiving end. This copying process reconstructs the speech without preserving the original noise components.

Inventive Principle:
Principle #26Copying

2Loss of energy

If traditional CODECs compress audio signals, then bandwidth is reduced, but audio fidelity deteriorates due to lossy compression

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidaudio fidelity
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The patent replaces traditional mechanical audio compression CODECs with a machine learning-based approach. Instead of using conventional compression algorithms that sacrifice fidelity, the system uses neural networks to extract essential features and regenerate high-fidelity speech, achieving both bandwidth efficiency and audio quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the audio signal into a different parameter space (spectrogram representations) that captures essential speech information more efficiently. This parameter transformation allows for compact representation and transmission while enabling high-fidelity regeneration through the audio generation model.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If audio signals are transmitted over bandwidth-limited channels, then communication is enabled, but noise interference increases and audio quality decreases

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidnoise interference
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The system extracts only the essential speech information from the audio signal before transmission, removing noisy components. This extraction enables efficient transmission over bandwidth-limited channels while ensuring that only clean speech representations are regenerated at the receiving end.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240371357A1Adaptive speech regeneration
Publication Date: 2024.11.07 SHURE ACQUISITION HLDG INC
  • US20240371357A1 patent drawing
  • US20240371357A1 patent drawing
  • US20240371357A1 patent drawing

AI summary

Embodiments provide for employing an audio transformation model to resynthesize speech signals associated with a speaking entity. Examples can receive audio signals comprising speech signals that are captured by an audio capture device. Examples can divide the audio signals into audio segments and input the audio segments into an audio transformation model to generate a voice vector representation and a speech vector representation. The voice vector representation comprises characteristics related to a speaking voice associated with the speaking entity and the speech vector representation comprises one or more words spoken by the speaking entity. The one or more words comprised in the speech vector representation are associated with respective contextual attributes associated with the one or more words. The audio transformation model can utilize the voice vector representation and the speech vector representation to regenerate the speech signals associated with the speaking entity.