Whispered Speech Conversion via Acoustic Feature Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems struggle to accurately convert whispered speech to normal speech, leading to inaccuracies in understanding whispered inputs, particularly in private or meeting scenarios and for individuals with aphasia-like pronunciation.
Innovation Solution
A method and apparatus that utilize a whispered speech converting model trained with both whispered and normal speech data, incorporating acoustic features and preliminary recognition results, along with lip shape recognition, to convert whispered speech to normal speech, employing recurrent neural networks or attention-based mechanisms for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech recognition systems are used for whispered speech, then the system can process the input speech, but the recognition accuracy is low due to the irregular vibration characteristics of whispered speech
Solution Approach 1:
The patent transforms the acoustic features of whispered speech by mapping them to normal speech acoustic feature spaces through trained conversion models. This parameter transformation converts the irregular vibration characteristics of whispered speech into patterns resembling normal speech, enabling accurate recognition without requiring separate whispered speech recognition systems.
Solution Approach 2:
The patent introduces an acoustic feature conversion model as an intermediary between the whispered speech input and the speech recognition system. This model acts as a mediator that translates whispered speech characteristics into normal speech characteristics, allowing the existing recognition system to process whispered speech accurately without direct modification to the core recognition algorithms.
2Device complexity
If normal speech recognition models are directly applied to whispered speech, then the system structure remains simple, but the pronunciation differences cause inaccurate recognition
Solution Approach 1:
The patent segments the speech processing pipeline into distinct modules: an acoustic feature extraction module, an acoustic feature conversion module, and a speech recognition module. This segmentation allows the conversion model to be trained independently on whispered-to-normal speech pairs, while the recognition model can remain a standard system, thus balancing complexity and accuracy.
Solution Approach 2:
The patent performs preliminary conversion of acoustic features before the main recognition process. By pre-transforming whispered speech acoustic features into normal speech acoustic features using a trained conversion model, the system prepares the input data in advance, ensuring that the subsequent recognition stage receives optimized input without requiring complex real-time adjustments.
3Use of energy by moving object
If whispered speech is amplified to increase volume, then the speech can be heard better, but the pronunciation characteristics remain irregular and unrecognizable
Solution Approach 1:
The patent changes the acoustic feature parameters of whispered speech through a trained conversion model rather than simply amplifying the volume. This parameter transformation recovers the fundamental frequency and regular vibration patterns characteristic of normal speech, enabling accurate recognition while maintaining appropriate volume levels without relying on simple amplification.
Data Source
AI summary
A method, an apparatus and a device for converting a whispered speech, and a readable storage medium are provided. The method is implemented based on the whispered speech converting model. The whispered speech converting model is trained in advance by using recognition results and whispered speech training acoustic features of whispered speech training data as samples and using normal speech acoustic features of normal speech data parallel to the whispered speech training data as sample labels. A whispered speech acoustic feature and a preliminary recognition result of whispered speech data are acquired, then the whispered speech acoustic feature and the preliminary recognition result are inputted into a preset whispered speech converting model to acquire a normal speech acoustic feature outputted by the model. In this way, the whispered speech can be converted to a normal speech.


