Speech Processing Device Multi-Speaker Conversation Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques have low accuracy in extracting conversation groups of three or more speakers due to difficulties in detecting speech and establishing conversation between silent speakers and those who barely speak.

Innovation Solution

A speech processing device and method that detect speech from individual speakers, calculate degrees of established conversation in segments, and extract conversation groups based on long-time features, allowing for accurate identification of conversational partners in multi-speaker environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional technique based on sound-silent period alternation is used, then conversation groups of two speakers can be extracted, but accuracy drops significantly for conversation groups of three or more speakers

Engineering Contradiction:
Improveaccuracy in extracting conversation groupVSAvoidapplicability to multi-speaker conversations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The determination time period is divided into multiple segments, and the degree of established conversation is calculated separately for each segment. This segmentation allows the system to detect conversation patterns across different time windows, improving accuracy in identifying conversation groups of three or more speakers by analyzing local conversation dynamics within each segment rather than relying solely on global sound-silent alternation patterns.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of analysis by calculating the degree of established conversation not just based on sound-silent alternation, but by incorporating the temporal distribution of speech across multiple segments. This dimensional expansion from binary (sound/silent) to multi-dimensional (temporal distribution across segments) enables better detection of multi-speaker conversations where simple alternation patterns fail.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of manufacture

If degree of established conversation is calculated based on sound-silent alternation, then simple implementation is achieved, but low accuracy occurs when substantial listeners who barely speak are present

Engineering Contradiction:
Improvesimplicity of calculation methodVSAvoidaccuracy in detecting conversational partners
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

By segmenting the determination time period into multiple segments and calculating the degree of established conversation for each segment separately, the system can detect subtle conversation patterns that may be missed in a single global calculation. This allows substantial listeners who speak occasionally to be identified as conversational partners through their presence in specific segments, even though their overall speech duration is low.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter used for measuring established conversation from a simple sound-silent alternation metric to a segmented temporal distribution metric. This parameter change enables the system to capture the contribution of substantial listeners who speak intermittently, improving detection accuracy without requiring complex calculation methods.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9064501B2Speech processing device and speech processing method
Publication Date: 2015.06.23 PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
  • US9064501B2 patent drawing
  • US9064501B2 patent drawing
  • US9064501B2 patent drawing

AI summary

A speech processing device which can accurately extract a conversation group from among a plurality of speakers, even when a conversation group formed of three or more people is present. This device (400) comprises: a spontaneous speech detection unit (420) and a direction-specific speech detection unit (430) which separately detect, from a sound signal, uttered speech from the speakers; a conversation establishment level calculation unit (450) which calculates a conversation establishment level for each separated segment of the time being determined, for all of the pairings of two people, on the basis of the detected uttered speech; an extended-period characteristic amount calculation unit (460) which calculates an extended-period characteristic amount for the conversation establishment level of the time being determined, for each pairing; and a conversation-partner determination unit (470) which extracts a conversation group which forms a conversation on the basis of the calculated extended-period characteristic amount.