Non-Audible Speech Signal Conversion Using Vocal Tract Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Converting non-audible speech signals into audible signals is challenging due to the lack of regular vocal cord vibrations, resulting in unnatural intonation and low speech recognition rates when using conventional methods that combine vocal tract and sound source feature value conversion models.
Innovation Solution
A speech processing method that calculates and converts non-audible speech signals into audible whispers using a vocal tract feature value conversion model, eliminating the need for sound source feature value conversion and reducing arithmetic load, allowing for high-speed real-time processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional methods combine vocal tract and sound source feature value conversion models to convert non-audible speech signals, then the conversion can be performed, but the speech recognition rate deteriorates and unnatural intonation occurs
Solution Approach 1:
The invention extracts and eliminates the sound source feature value conversion model from the conventional conversion system, retaining only the vocal tract feature value conversion model. This extraction resolves the technical contradiction by simplifying the conversion process to avoid the unnatural intonation and low speech recognition rates caused by including sound source conversion, while maintaining the essential vocal tract transformation needed for accurate speech conversion.
Solution Approach 2:
The invention segments the speech conversion process into two independent parts: non-audible speech signal collection through in-vivo conduction microphone and vocal tract feature value conversion. By separating the sound source conversion component, the system achieves higher speech recognition rates while reducing model complexity, as only the vocal tract transformation is applied to the collected non-audible speech features.
2Productivity
If conventional methods use both vocal tract and sound source feature value conversion models, then conversion can be achieved, but arithmetic load increases and processing speed decreases
Solution Approach 1:
The invention removes the sound source feature value conversion model from the processing pipeline, eliminating the redundant computational operations. This extraction directly reduces the arithmetic load by approximately half compared to conventional methods that process both vocal tract and sound source features, thereby enabling high-speed real-time conversion without sacrificing conversion accuracy.
Solution Approach 2:
The invention applies partial action by using only the necessary vocal tract feature value conversion while omitting the excessive sound source conversion operations. This partial processing approach maintains sufficient conversion quality for speech recognition while dramatically reducing the computational power required, thus achieving fast real-time processing.
3Measurement precision
If non-audible speech signals are converted using conventional models, then conversion is possible, but the output speech becomes unnatural and recognition accuracy decreases
Solution Approach 1:
The invention extracts and eliminates the problematic sound source feature value conversion component that causes unnatural intonation. By retaining only the vocal tract feature value conversion, the system produces natural-sounding output speech with high recognition accuracy, resolving the contradiction between conversion possibility and output quality.
Solution Approach 2:
The invention changes the conversion parameters by focusing solely on vocal tract feature transformations rather than attempting to reconstruct sound source characteristics. This parameter adjustment ensures that the conversion preserves the naturalness of speech while achieving high recognition accuracy, avoiding the artifacts introduced by conventional dual-model approaches.
Data Source
AI summary
[Problems]To convert a signal of non-audible murmur obtained through an in-vivo conduction microphone into a signal of a speech that is recognizable for (hardly misrecognized by) a receiving person with maximum accuracy.[Means for Solving Problems]A speech processing method comprising: a learning step (S7) for conducting a learning calculation of a model parameter of a vocal tract feature value conversion model indicating conversion characteristic of acoustic feature value of vocal tract, on the basis of a learning input signal of non-audible murmur recorded by an in-vivo conduction microphone and a learning output signal of audible whisper corresponding to the learning input signal recorded by a prescribed microphone, and then, storing a learned model parameter in a prescribed storing means; and a speech conversion step (S9) for converting a non-audible speech signal obtained through an in-vivo conduction microphone into a signal of audible whisper, based on a vocal tract feature value conversion model, with a learned model parameter obtained through the learning step set thereto.


