Speech Enhancement Using Reference Feature Memory
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech enhancement technologies struggle to achieve satisfactory results when the quality of the speech signal is poor, as they often fail to extract sufficient features for effective enhancement.
Innovation Solution
The proposed method involves a feature memory system that extracts and records sound features of familiar individuals during speech processing. This information is used to improve the speech enhancement effect by utilizing known speech features, even in noisy environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If a single-pass speech enhancement algorithm is used to process speech signals in real-time, then the processing speed is improved, but the speech enhancement effect deteriorates when the speech signal quality is poor
Solution Approach 1:
The system performs preliminary feature extraction and storage during periods of good speech quality (quiet environment), creating a reference feature database before the actual enhancement is needed. This preliminary action enables the system to leverage pre-extracted features when processing noisy speech signals, improving enhancement effectiveness without compromising real-time processing speed.
Solution Approach 2:
The system creates a copy of the reference feature database that is updated and maintained separately from the real-time processing stream. This copied reference database can be retrieved and applied to enhance noisy speech signals, allowing the system to maintain fast processing while improving enhancement quality by leveraging stored reference features.
2Productivity
If speech enhancement is performed on sound with poor speech signal quality, then the real-time processing capability is maintained, but the feature extraction capability deteriorates
Solution Approach 1:
The reference feature database acts as an intermediary between the noisy input signal and the enhancement output. Instead of directly extracting features from poor-quality noisy signals, the system uses the pre-extracted reference features as an intermediary representation of the speaker's characteristics, which then guides the enhancement process and improves feature extraction accuracy.
Solution Approach 2:
The system performs preliminary feature extraction during periods of good signal quality to build the reference database before processing noisy signals. This preliminary action ensures that high-quality reference features are available when the actual enhancement is needed, maintaining both real-time processing capability and feature extraction precision.
3Reliability
If a feature memory system is added to extract and record sound features of familiar individuals, then the speech enhancement effect is improved, but the device complexity increases
Solution Approach 1:
The reference feature database serves multiple functions: it stores speaker-specific features for identification, provides reference data for enhancement, and enables the system to handle both known and unknown speakers. This multi-functionality justifies the added complexity by delivering benefits across multiple operational scenarios.
Solution Approach 2:
The system maintains a copied reference feature database that can be efficiently queried and applied during processing. By using a separate reference database rather than processing everything in real-time, the system manages complexity through data organization and retrieval mechanisms rather than computational complexity.
Data Source
AI summary
This application discloses a speech enhancement method (100, 200) and device. The speech enhancement method (100, 200) includes: receiving a current audio input signal having a speech portion and a non-speech portion (201); determining a voice feature of the speech portion in the current audio input signal (202); determining a speech quality of the current audio input signal (203); evaluating whether the speech quality meets a predetermined speech quality requirement (204); and creating or updating, in response to the speech quality meeting the predetermined speech quality requirement, a reference speech feature by using the voice feature, where the reference speech feature is used for enhancing the speech portion in an audio input signal (205).

