Speech Extraction Using Vibration Signals for Speaker Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine-learning type speech extraction technologies struggle to accurately extract the speech signal of a specific speaker from a microphone input signal that includes speeches from multiple people.
Innovation Solution
An information processing apparatus that includes a first speech extraction processing section to generate a first speech extraction signal, a correction signal generation section to generate a correction signal from a vibration signal indicating the user's utterance, and a post-processing section to improve the accuracy of the utterance speech signal by post-processing the first speech extraction signal based on the correction signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine-learning type speech extraction technology is used to extract human speech from noisy signals, then speech extraction capability is improved, but the ability to extract speech from a specific speaker among multiple speakers deteriorates
Solution Approach 1:
The patent segments the speech extraction process into two distinct stages: first extracting speech components from the mixed signal using machine learning, then further processing these components to identify and separate specific speakers. This segmentation allows each stage to specialize in one aspect of the challenging task.
Solution Approach 2:
The patent introduces an intermediary processing stage that takes the output from the first speech extraction and prepares it for specific speaker identification. This intermediary step acts as a bridge between general speech extraction and speaker-specific extraction, enabling both capabilities to coexist.
2Device complexity
If only speech signals are processed for extraction, then processing simplicity is maintained, but extraction accuracy in multi-speaker environments deteriorates
Solution Approach 1:
The patent merges multiple signal types (speech signals and vibration signals) into a unified processing framework. By combining these different modalities, the system achieves better speaker-specific extraction accuracy while maintaining relatively simple processing through the integrated approach.
Solution Approach 2:
The patent creates a composite signal processing approach by combining speech signals with vibration signals from wearable devices. This composite methodology leverages the strengths of both signal types to improve extraction accuracy without significantly increasing overall system complexity.
3Measurement precision
If vibration signals from wearable devices are incorporated to improve speaker identification, then extraction accuracy is improved, but system complexity increases
Solution Approach 1:
The patent makes the speech extraction system multi-functional by enabling it to process both general speech extraction and speaker-specific extraction using a unified architecture. The same framework handles both speech signals from microphones and vibration signals from wearable devices, reducing the need for separate specialized systems.
Solution Approach 2:
The system uses the vibration signals from wearable devices to automatically identify and isolate the target speaker's speech components without requiring manual intervention or complex configuration. The system self-adjusts to focus on the correct speaker based on the vibration signal patterns.
Data Source
AI summary
To extract an utterance speech uttered by a specific user. An information processing apparatus includes a first speech extraction processing section that generates a first speech extraction signal by extracting an utterance speech component from a speech signal including an utterance speech uttered by a user, a correction signal generation section that generates a correction signal from a vibration signal indicating vibration of a part of the user that vibrates in conjunction with a user's utterance, and a post-processing section that generates an utterance speech signal indicating the utterance speech by post-processing the first speech extraction signal based on the correction signal.


