Speaker Identification via Operation Timing in Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems struggle to accurately recognize speech from conversation partners who are not familiar with the system, leading to inadequate recognition and the need for repeated speech from the partner.
Innovation Solution
A speech recognition device that includes an obtaining unit for capturing speech, a storage unit for storing the speech, an input unit for receiving operation inputs from the primary speaker, an utterance start detector to identify the start of speech, and a speaker identification unit to differentiate between the primary and secondary speakers based on timing inputs, allowing for reliable speech recognition of both speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the speech recognition system is operated by the first speaker (owner), then the first speaker can properly operate the system to make it recognize his speech, but the second speaker (conversation partner) cannot be recognized because he does not know how to use the system
Solution Approach 1:
The system automatically detects utterance timing and performs speech recognition without requiring the conversation partner to manually operate the device. The first speaker's operation input triggers automatic detection of the second speaker's utterance start position, enabling the system to serve itself in identifying and recognizing both speakers' speech.
2Measurement precision
If the system waits for manual operation input to detect speech start, then the first speaker's speech is recognized accurately, but the second speaker's speech is missed because he cannot provide operation input
Solution Approach 1:
The system segments the speech detection process by speaker role. The first speaker provides operation input that segments the conversation into distinct utterance sections. The utterance section detecting means then segments and identifies the second speaker's speech within these sections, allowing precise timing detection for both speakers independently.
Solution Approach 2:
The operation input from the first speaker serves as an intermediary trigger that enables the system to detect and identify the second speaker's utterance. This intermediary mechanism bridges the gap between manual operation and automatic multi-speaker recognition, allowing the system to adapt to conversations between familiar and unfamiliar users.
3Reliability
If the speech recognition system requires conversation partner to speak again when recognition fails, then recognition accuracy may improve, but communication efficiency deteriorates due to repeated speech
Solution Approach 1:
The system provides feedback by detecting utterance sections and identifying speakers automatically. When the second speaker utters speech, the system detects the utterance section based on the first speaker's operation input, identifies it as the second speaker's speech, and performs recognition without requiring the conversation partner to repeat himself, thus maintaining both reliability and efficiency.
Data Source
AI summary
A speech recognition device includes: an obtaining unit which obtains a speech uttered in a conversation between a first speaker and a second speaker; a storage which stores the speech obtained; an input unit which receives operation input; an utterance start detector which, when the input unit receives the operation input, detects a start position of the speech; and a speaker identification unit which identifies a speaker of the speech as the first speaker who has performed the operation input or the second speaker who has not performed the operation input, based on (i) first timing at which the input unit has received the operation input and (ii) second timing indicating the detected start position of the speech. The first and second timing are set for each speech of the first and second speakers. A speech recognizer performs speech recognition on the speech whose speaker has been identified.


