Speech Recognition Using Conversation History for Short Utterance Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems face challenges in accurately recognizing words in large vocabulary continuous speech recognition, particularly when short speech or abbreviated sentences are inputted, leading to decreased recognition rates and incorrect selection of utterance candidates.

Innovation Solution

A speech recognition apparatus and method that utilizes a conversation history database to select appropriate candidates by comparing input speech signals with topic specifying information from past conversations, prioritizing candidates that match the conversation history and avoiding irrelevant outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional speech recognition systems use statistical language models to narrow down utterance candidates from large vocabulary continuous speech, then the system can handle large vocabulary recognition, but the recognition rate decreases when short speech or abbreviated sentences are inputted repeatedly

Engineering Contradiction:
Improvelarge vocabulary recognition capabilityVSAvoidrecognition rate
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary analysis of conversation history to identify topic specifying information before processing the current speech input. By pre-processing the conversation context and storing topic information in advance, the system can quickly retrieve relevant topics during speech recognition, thereby improving recognition accuracy for short or abbreviated speech without sacrificing large vocabulary capability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention introduces conversation history and topic specifying information as an intermediary between the speech signal and the final recognition result. This intermediary layer filters and prioritizes candidates based on contextual relevance, resolving the contradiction by enabling the system to leverage both large vocabulary coverage and context-aware precision

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the speech recognition system selects candidates based solely on highest recognition rate from conventional language models, then the selection process is simple and fast, but irrelevant candidates are selected when conversation context matters

Engineering Contradiction:
Improvecandidate selection speedVSAvoidcandidate relevance accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements feedback by continuously updating and utilizing conversation history to inform candidate selection. The topic specifying information extracted from conversation history feeds back into the candidate evaluation process, allowing the system to prioritize contextually relevant candidates while maintaining efficient selection through predefined topic matching mechanisms

Inventive Principle:
Principle #23Feedback

3Productivity

If word spotting method is used to extract recognition candidate words from continuous conversation speech, then extraction is efficient for small number of words, but accuracy falls as the number of words to be set increases

Engineering Contradiction:
Improveextraction efficiencyVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system applies local quality by treating different portions of the vocabulary differently based on conversation context. Instead of uniformly processing all candidates, the system identifies topic specifying information and applies enhanced processing or prioritization to contextually relevant candidates, thereby maintaining high efficiency while improving accuracy for important words

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS7415406B2Speech recognition apparatus, speech recognition method, conversation control apparatus, conversation control method, and programs for therefor
Publication Date: 2008.08.19 UBIQUITOUS AUTO SYNCHRONICITY LLC
  • US7415406B2 patent drawing
  • US7415406B2 patent drawing
  • US7415406B2 patent drawing

AI summary

An automatic conversation apparatus includes a speech recognizing unit receiving a speech signal and outputting characters/character string corresponding to the speech signal as a recognition result; a speech recognition dictionary storing unit storing a language model for determining candidates corresponding to the speech signal; a conversation database storing plural pieces of topic specifying information; a sentence analyzing unit analyzing the characters/character string outputted from the speech recognizing unit; and a conversation control unit storing a conversation history and acquiring an answer sentence based on an analysis result of the sentence analyzing unit. Speech recognizing unit includes a collating unit that outputs plural candidates based on the speech recognition dictionary storing unit; and a candidate determining unit comparing the plural candidates outputted from collating unit with topic specifying information corresponding to the conversation history with reference to the conversation database and outputs one candidate based on the comparison.