Speech Extraction Using Vibration Signals for Speaker Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine-learning type speech extraction technologies struggle to accurately extract the speech signal of a specific speaker from a microphone input signal that includes speeches from multiple people.

Innovation Solution

An information processing apparatus that includes a first speech extraction processing section to generate a first speech extraction signal, a correction signal generation section to generate a correction signal from a vibration signal indicating the user's utterance, and a post-processing section to improve the accuracy of the utterance speech signal by post-processing the first speech extraction signal based on the correction signal.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine-learning type speech extraction technology is used to extract human speech from noisy signals, then speech extraction capability is improved, but the ability to extract speech from a specific speaker among multiple speakers deteriorates

Engineering Contradiction:
Improvespeech extraction accuracyVSAvoidmulti-speaker discrimination capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the speech extraction process into two distinct stages: first extracting speech components from the mixed signal using machine learning, then further processing these components to identify and separate specific speakers. This segmentation allows each stage to specialize in one aspect of the challenging task.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing stage that takes the output from the first speech extraction and prepares it for specific speaker identification. This intermediary step acts as a bridge between general speech extraction and speaker-specific extraction, enabling both capabilities to coexist.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If only speech signals are processed for extraction, then processing simplicity is maintained, but extraction accuracy in multi-speaker environments deteriorates

Engineering Contradiction:
Improvesignal processing complexityVSAvoidspecific speaker extraction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges multiple signal types (speech signals and vibration signals) into a unified processing framework. By combining these different modalities, the system achieves better speaker-specific extraction accuracy while maintaining relatively simple processing through the integrated approach.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite signal processing approach by combining speech signals with vibration signals from wearable devices. This composite methodology leverages the strengths of both signal types to improve extraction accuracy without significantly increasing overall system complexity.

Inventive Principle:
Principle #40Composite materials

3Measurement precision

If vibration signals from wearable devices are incorporated to improve speaker identification, then extraction accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the speech extraction system multi-functional by enabling it to process both general speech extraction and speaker-specific extraction using a unified architecture. The same framework handles both speech signals from microphones and vibration signals from wearable devices, reducing the need for separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses the vibration signals from wearable devices to automatically identify and isolate the target speaker's speech components without requiring manual intervention or complex configuration. The system self-adjusts to focus on the correct speaker based on the vibration signal patterns.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250191607A1Information processing apparatus, information processing method, information processing program, and information processing system
Publication Date: 2025.06.12 SONY GROUP CORP
  • US20250191607A1 patent drawing
  • US20250191607A1 patent drawing
  • US20250191607A1 patent drawing

AI summary

To extract an utterance speech uttered by a specific user. An information processing apparatus includes a first speech extraction processing section that generates a first speech extraction signal by extracting an utterance speech component from a speech signal including an utterance speech uttered by a user, a correction signal generation section that generates a correction signal from a vibration signal indicating vibration of a part of the user that vibrates in conjunction with a user's utterance, and a post-processing section that generates an utterance speech signal indicating the utterance speech by post-processing the first speech extraction signal based on the correction signal.