Dialogue Speech Recognition Using Turn Information
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current dialogue speech recognition systems face challenges in accurately recognizing multiple speakers speaking simultaneously and lack a universal structure to improve recognition accuracy, as existing techniques are limited to specific contexts and disregard overlapping utterances.
Innovation Solution
A dialogue speech recognition system that utilizes turn information to differentiate linguistic likelihoods for speakers with and without the turn to speak, employing separate models for speech recognition based on the presence or absence of the turn to speak, allowing for improved accuracy in diverse dialogue scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single linguistic model is used for all speakers in a dialogue, then the device complexity is reduced, but the speech recognition accuracy deteriorates when multiple speakers speak simultaneously
Solution Approach 1:
The patent segments the linguistic model into multiple speaker-specific models (first linguistic model and second linguistic model) corresponding to different speakers. Each speaker has their own linguistic model that reflects their speech characteristics, vocabulary, and language patterns. This segmentation allows the system to accurately recognize speech from different speakers even when they speak simultaneously, resolving the contradiction between accuracy and complexity by organizing the complexity into manageable, speaker-specific components.
Solution Approach 2:
The patent applies local quality by assigning different linguistic models to different speakers based on their individual characteristics. Instead of using a uniform linguistic model for all speakers, the system tailors the linguistic model to each speaker's local speech properties. This enables the system to handle overlapping speech effectively by applying the appropriate speaker-specific linguistic model to each speaker's utterances, thereby improving recognition accuracy without requiring an overly complex global model.
2Measurement precision
If overlapping utterances are disregarded in speech recognition, then the processing complexity is reduced, but the recognition accuracy deteriorates in dialogue scenarios with simultaneous speech
Solution Approach 1:
The patent performs preliminary action by detecting overlapping utterances before conducting speech recognition. The system identifies which speakers are speaking simultaneously and prepares the appropriate speaker-specific linguistic models in advance. This preliminary detection and preparation allow the system to handle overlapping speech efficiently during the recognition phase, improving accuracy without requiring complex real-time processing of overlapping utterances.
Solution Approach 2:
The patent introduces turn information as an intermediary element that mediates between the speech signals and the linguistic models. The turn information indicates whether a speaker has the turn to speak or the probability that they have the turn, serving as a bridge that helps the system appropriately apply speaker-specific linguistic models during overlapping utterances. This intermediary structure simplifies the processing by providing clear guidance on which speaker models to use, rather than requiring complex analysis of overlapping speech patterns.
3Measurement precision
If speaker-specific linguistic models are used, then the speech recognition accuracy for dialogue is improved, but the calculation time increases
Solution Approach 1:
The patent performs preliminary action by detecting overlapping utterances and determining turn information before conducting speech recognition. By identifying which speakers are speaking simultaneously and preparing the appropriate speaker-specific linguistic models in advance, the system avoids the need for complex real-time calculations during the recognition phase. This preliminary preparation reduces calculation time while maintaining the accuracy benefits of speaker-specific models.
Solution Approach 2:
The patent applies partial action by using speaker-specific linguistic models selectively based on turn information rather than applying all possible models to all speech. The system uses the first linguistic model when the first speaker has the turn and the second linguistic model when the second speaker has the turn, avoiding unnecessary calculations. This selective application of linguistic models reduces overall calculation time while maintaining high recognition accuracy for dialogue speech.
Data Source
AI summary
Disclosed is a dialogue speech recognition system that can expand the scope of applications by employing a universal dialogue structure as the condition for speech recognition of dialogue speech between persons. An acoustic likelihood computation means (701) provides a likelihood that a speech signal input from a given phoneme sequence will occur. A linguistic likelihood computation means (702) provides a likelihood that a given word sequence will occur. A maximum likelihood candidate search means (703) uses the likelihoods provided by the acoustic likelihood computation means and the linguistic likelihood computation means to provide a word sequence with the maximum likelihood of occurring from a speech signal. Further, the linguistic likelihood computation means (702) provides different linguistic likelihoods when the speaker who generated the acoustic signal input to the speech recognition means does and does not have the turn to speak.


