Speaker Identification via Operation Timing in Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems struggle to accurately recognize speech from conversation partners who are not familiar with the system, leading to inadequate recognition and the need for repeated speech from the partner.

Innovation Solution

A speech recognition device that includes an obtaining unit for capturing speech, a storage unit for storing the speech, an input unit for receiving operation inputs from the primary speaker, an utterance start detector to identify the start of speech, and a speaker identification unit to differentiate between the primary and secondary speakers based on timing inputs, allowing for reliable speech recognition of both speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the speech recognition system is operated by the first speaker (owner), then the first speaker can properly operate the system to make it recognize his speech, but the second speaker (conversation partner) cannot be recognized because he does not know how to use the system

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidoperation complexity for conversation partner
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically detects utterance timing and performs speech recognition without requiring the conversation partner to manually operate the device. The first speaker's operation input triggers automatic detection of the second speaker's utterance start position, enabling the system to serve itself in identifying and recognizing both speakers' speech.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If the system waits for manual operation input to detect speech start, then the first speaker's speech is recognized accurately, but the second speaker's speech is missed because he cannot provide operation input

Engineering Contradiction:
Improveutterance timing detection accuracyVSAvoidcapability to handle multiple speakers
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments the speech detection process by speaker role. The first speaker provides operation input that segments the conversation into distinct utterance sections. The utterance section detecting means then segments and identifies the second speaker's speech within these sections, allowing precise timing detection for both speakers independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The operation input from the first speaker serves as an intermediary trigger that enables the system to detect and identify the second speaker's utterance. This intermediary mechanism bridges the gap between manual operation and automatic multi-speaker recognition, allowing the system to adapt to conversations between familiar and unfamiliar users.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If the speech recognition system requires conversation partner to speak again when recognition fails, then recognition accuracy may improve, but communication efficiency deteriorates due to repeated speech

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoidcommunication efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system provides feedback by detecting utterance sections and identifying speakers automatically. When the second speaker utters speech, the system detects the utterance section based on the first speaker's operation input, identifies it as the second speaker's speech, and performs recognition without requiring the conversation partner to repeat himself, thus maintaining both reliability and efficiency.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11315572B2Speech recognition device, speech recognition method, and recording medium
Publication Date: 2022.04.26 PANASONIC HOLDINGS CORP
  • US11315572B2 patent drawing
  • US11315572B2 patent drawing
  • US11315572B2 patent drawing

AI summary

A speech recognition device includes: an obtaining unit which obtains a speech uttered in a conversation between a first speaker and a second speaker; a storage which stores the speech obtained; an input unit which receives operation input; an utterance start detector which, when the input unit receives the operation input, detects a start position of the speech; and a speaker identification unit which identifies a speaker of the speech as the first speaker who has performed the operation input or the second speaker who has not performed the operation input, based on (i) first timing at which the input unit has received the operation input and (ii) second timing indicating the detected start position of the speech. The first and second timing are set for each speech of the first and second speakers. A speech recognizer performs speech recognition on the speech whose speaker has been identified.