Dialogue Speech Recognition Using Turn Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current dialogue speech recognition systems face challenges in accurately recognizing multiple speakers speaking simultaneously and lack a universal structure to improve recognition accuracy, as existing techniques are limited to specific contexts and disregard overlapping utterances.

Innovation Solution

A dialogue speech recognition system that utilizes turn information to differentiate linguistic likelihoods for speakers with and without the turn to speak, employing separate models for speech recognition based on the presence or absence of the turn to speak, allowing for improved accuracy in diverse dialogue scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single linguistic model is used for all speakers in a dialogue, then the device complexity is reduced, but the speech recognition accuracy deteriorates when multiple speakers speak simultaneously

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidlinguistic model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the linguistic model into multiple speaker-specific models (first linguistic model and second linguistic model) corresponding to different speakers. Each speaker has their own linguistic model that reflects their speech characteristics, vocabulary, and language patterns. This segmentation allows the system to accurately recognize speech from different speakers even when they speak simultaneously, resolving the contradiction between accuracy and complexity by organizing the complexity into manageable, speaker-specific components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by assigning different linguistic models to different speakers based on their individual characteristics. Instead of using a uniform linguistic model for all speakers, the system tailors the linguistic model to each speaker's local speech properties. This enables the system to handle overlapping speech effectively by applying the appropriate speaker-specific linguistic model to each speaker's utterances, thereby improving recognition accuracy without requiring an overly complex global model.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If overlapping utterances are disregarded in speech recognition, then the processing complexity is reduced, but the recognition accuracy deteriorates in dialogue scenarios with simultaneous speech

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by detecting overlapping utterances before conducting speech recognition. The system identifies which speakers are speaking simultaneously and prepares the appropriate speaker-specific linguistic models in advance. This preliminary detection and preparation allow the system to handle overlapping speech efficiently during the recognition phase, improving accuracy without requiring complex real-time processing of overlapping utterances.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces turn information as an intermediary element that mediates between the speech signals and the linguistic models. The turn information indicates whether a speaker has the turn to speak or the probability that they have the turn, serving as a bridge that helps the system appropriately apply speaker-specific linguistic models during overlapping utterances. This intermediary structure simplifies the processing by providing clear guidance on which speaker models to use, rather than requiring complex analysis of overlapping speech patterns.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If speaker-specific linguistic models are used, then the speech recognition accuracy for dialogue is improved, but the calculation time increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcalculation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by detecting overlapping utterances and determining turn information before conducting speech recognition. By identifying which speakers are speaking simultaneously and preparing the appropriate speaker-specific linguistic models in advance, the system avoids the need for complex real-time calculations during the recognition phase. This preliminary preparation reduces calculation time while maintaining the accuracy benefits of speaker-specific models.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by using speaker-specific linguistic models selectively based on turn information rather than applying all possible models to all speech. The system uses the first linguistic model when the first speaker has the turn and the second linguistic model when the second speaker has the turn, avoiding unnecessary calculations. This selective application of linguistic models reduces overall calculation time while maintaining high recognition accuracy for dialogue speech.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8818801B2Dialogue speech recognition system, dialogue speech recognition method, and recording medium for storing dialogue speech recognition program
Publication Date: 2014.08.26 NEC CORP
  • US8818801B2 patent drawing
  • US8818801B2 patent drawing
  • US8818801B2 patent drawing

AI summary

Disclosed is a dialogue speech recognition system that can expand the scope of applications by employing a universal dialogue structure as the condition for speech recognition of dialogue speech between persons. An acoustic likelihood computation means (701) provides a likelihood that a speech signal input from a given phoneme sequence will occur. A linguistic likelihood computation means (702) provides a likelihood that a given word sequence will occur. A maximum likelihood candidate search means (703) uses the likelihoods provided by the acoustic likelihood computation means and the linguistic likelihood computation means to provide a word sequence with the maximum likelihood of occurring from a speech signal. Further, the linguistic likelihood computation means (702) provides different linguistic likelihoods when the speaker who generated the acoustic signal input to the speech recognition means does and does not have the turn to speak.