Non-Audible Speech Signal Conversion Using Vocal Tract Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Converting non-audible speech signals into audible signals is challenging due to the lack of regular vocal cord vibrations, resulting in unnatural intonation and low speech recognition rates when using conventional methods that combine vocal tract and sound source feature value conversion models.

Innovation Solution

A speech processing method that calculates and converts non-audible speech signals into audible whispers using a vocal tract feature value conversion model, eliminating the need for sound source feature value conversion and reducing arithmetic load, allowing for high-speed real-time processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods combine vocal tract and sound source feature value conversion models to convert non-audible speech signals, then the conversion can be performed, but the speech recognition rate deteriorates and unnatural intonation occurs

Engineering Contradiction:
Improvespeech recognition rateVSAvoidconversion model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The invention extracts and eliminates the sound source feature value conversion model from the conventional conversion system, retaining only the vocal tract feature value conversion model. This extraction resolves the technical contradiction by simplifying the conversion process to avoid the unnatural intonation and low speech recognition rates caused by including sound source conversion, while maintaining the essential vocal tract transformation needed for accurate speech conversion.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention segments the speech conversion process into two independent parts: non-audible speech signal collection through in-vivo conduction microphone and vocal tract feature value conversion. By separating the sound source conversion component, the system achieves higher speech recognition rates while reducing model complexity, as only the vocal tract transformation is applied to the collected non-audible speech features.

Inventive Principle:
Principle #1Segmentation

2Productivity

If conventional methods use both vocal tract and sound source feature value conversion models, then conversion can be achieved, but arithmetic load increases and processing speed decreases

Engineering Contradiction:
Improveprocessing speedVSAvoidarithmetic load
Core Design Contradiction:
ProductivityVSPower

Solution Approach 1:

The invention removes the sound source feature value conversion model from the processing pipeline, eliminating the redundant computational operations. This extraction directly reduces the arithmetic load by approximately half compared to conventional methods that process both vocal tract and sound source features, thereby enabling high-speed real-time conversion without sacrificing conversion accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention applies partial action by using only the necessary vocal tract feature value conversion while omitting the excessive sound source conversion operations. This partial processing approach maintains sufficient conversion quality for speech recognition while dramatically reducing the computational power required, thus achieving fast real-time processing.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If non-audible speech signals are converted using conventional models, then conversion is possible, but the output speech becomes unnatural and recognition accuracy decreases

Engineering Contradiction:
Improvespeech recognition rateVSAvoidconversion accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The invention extracts and eliminates the problematic sound source feature value conversion component that causes unnatural intonation. By retaining only the vocal tract feature value conversion, the system produces natural-sounding output speech with high recognition accuracy, resolving the contradiction between conversion possibility and output quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention changes the conversion parameters by focusing solely on vocal tract feature transformations rather than attempting to reconstruct sound source characteristics. This parameter adjustment ensures that the conversion preserves the naturalness of speech while achieving high recognition accuracy, avoiding the artifacts introduced by conventional dual-model approaches.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8155966B2Apparatus and method for producing an audible speech signal from a non-audible speech signal
Publication Date: 2012.04.10 NARA INSTITUTE OF SCIENCE AND TECHNOLOGY
  • US8155966B2 patent drawing
  • US8155966B2 patent drawing
  • US8155966B2 patent drawing

AI summary

[Problems]To convert a signal of non-audible murmur obtained through an in-vivo conduction microphone into a signal of a speech that is recognizable for (hardly misrecognized by) a receiving person with maximum accuracy.[Means for Solving Problems]A speech processing method comprising: a learning step (S7) for conducting a learning calculation of a model parameter of a vocal tract feature value conversion model indicating conversion characteristic of acoustic feature value of vocal tract, on the basis of a learning input signal of non-audible murmur recorded by an in-vivo conduction microphone and a learning output signal of audible whisper corresponding to the learning input signal recorded by a prescribed microphone, and then, storing a learned model parameter in a prescribed storing means; and a speech conversion step (S9) for converting a non-audible speech signal obtained through an in-vivo conduction microphone into a signal of audible whisper, based on a vocal tract feature value conversion model, with a learned model parameter obtained through the learning step set thereto.