Speech Enhancement via Real-Time Embedding Vector Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech enhancement technologies require pre-registration of clear speech, making them unsuitable for real-time scenarios where the speaker cannot be determined in advance.

Innovation Solution

A method for speech enhancement that involves extracting embedding vectors from detected speech data, searching for a target embedding vector, generating a registration embedding vector, and calculating a masking value through correlation with audio feature vectors to enhance the target speech data in real-time without pre-registration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If pre-registration of clear speech is performed, then speech enhancement accuracy is improved, but real-time processing capability deteriorates

Engineering Contradiction:
Improvespeech enhancement accuracyVSAvoidreal-time processing capability
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts embedding vectors from audio data in real-time and performs speaker identification without requiring advance registration. The embedding vectors capture essential speaker characteristics dynamically, enabling the system to achieve accurate speech enhancement equivalent to pre-registration methods while maintaining real-time processing capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the traditional mechanical approach of pre-registering clear speech samples with an embedding vector-based representation system. By converting speaker characteristics into compact embedding vectors that can be extracted and compared in real-time, the system eliminates the time-consuming pre-registration step while preserving enhancement accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If embedding vector extraction and target search are performed, then speaker identification accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts only the essential speaker characteristics into compact embedding vectors, separating the critical identification features from the full audio signal. This extraction process reduces the data dimensionality while preserving speaker identity information, thereby improving identification accuracy without proportionally increasing computational complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the speaker identification problem from analyzing full audio signals to comparing embedding vectors. By changing the parameter representation from raw audio features to compressed embedding vectors, the system achieves higher identification accuracy while reducing computational burden through dimensionality reduction.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250210057A1Method, system, device, and storage medium for speech enhancement
Publication Date: 2025.06.26 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250210057A1 patent drawing
  • US20250210057A1 patent drawing
  • US20250210057A1 patent drawing

AI summary

The present disclosure relates to a field of computer technology and discloses a method, a system, a device, and storage medium for speech enhancement. The method for speech enhancement comprises acquiring audio data and, when speech data is detected in the audio data, extracting an embedding vector of the speech data; searching the embedding vector for a target embedding vector extracted from target speech data, and generating a registration embedding vector based on the target embedding vector; performing correlation calculation between the registration embedding vector and an audio feature vector of the audio data, to determine a masking value required for enhancing the target speech data; and enhancing, according to the masking value, the target speech data in the audio data.