Barge-in Speech Control Unit for Dialogue Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Barge-in speech in spoken dialogue systems often includes repetitions and simple back-channels, leading to erroneous operations and reduced user convenience due to the engagement of these elements in dialogue control.

Innovation Solution

A spoken dialogue system that includes a barge-in speech control unit to determine whether to engage user speech elements based on their correspondence to predetermined morphemes in the previous system speech, using a barge-in speech determination model to differentiate between response candidates and non-engaged speech elements, thereby preventing erroneous operations and improving user convenience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all user speech elements are engaged in dialogue control, then the system responds to all user inputs, but erroneous operations occur due to repetitions and simple back-channels

Engineering Contradiction:
Improvedialogue control accuracyVSAvoiduser convenience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The user speech is segmented into individual speech elements, and each element is evaluated separately to determine whether it corresponds to a response candidate. This allows the system to distinguish between meaningful responses and meaningless repetitions at the element level, preventing erroneous engagement of non-responsive elements while maintaining engagement of valid responses.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A determination model is introduced as an intermediary between speech recognition and dialogue control. This model acts as a filter that predicts whether each speech element corresponds to a response candidate, allowing the system to selectively engage only those elements that are likely to be meaningful responses, thereby preventing erroneous operations while maintaining user convenience.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the system engages all barge-in speech, then it responds to all user inputs including repetitions, but this leads to erroneous operations

Engineering Contradiction:
Improvedialogue response efficiencyVSAvoiddialogue control accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The determination model performs preliminary evaluation of each speech element before it reaches the dialogue control unit. By predicting in advance whether a speech element corresponds to a response candidate, the system prevents erroneous operations from occurring in the first place, rather than having to correct them afterward. This preliminary filtering maintains high dialogue response efficiency while ensuring accuracy.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the system disengages repetitive speech elements, then erroneous operations are prevented, but legitimate responses may be missed

Engineering Contradiction:
Improvedialogue control accuracyVSAvoiddialogue response efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The determination model uses multiple parameters including acoustic features, text features, and contextual information to make its prediction. By changing and combining multiple parameters rather than relying on a single criterion, the system achieves high accuracy in distinguishing between legitimate responses and meaningless repetitions, preventing false disengagements while maintaining high dialogue response efficiency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11862167B2Voice dialogue system, model generation device, barge-in speech determination model, and voice dialogue program
Publication Date: 2024.01.02 NTT DOCOMO INC
  • US11862167B2 patent drawing
  • US11862167B2 patent drawing
  • US11862167B2 patent drawing

AI summary

A spoken dialogue device includes a recognition unit that recognizes an acquired user speech, a barge-in speech control unit that determines whether to engage a barge-in speech, a dialogue control unit that outputs a system response to a user based on a recognition result of the user speech other than the barge-in speech determined not to be engaged by the barge-in speech control unit, a response generation unit that generates a system speech based on the system response, and an output unit that outputs a system speech. When each user speech element included in the user speech corresponds to a predetermined morpheme included in the immediately previous system speech and does not correspond to a response candidate to the immediately previous system speech by a user, the barge-in speech control unit does not engage at least the user speech element.