Speech Recognition Morphological Analysis Candidate Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in accurately processing input speech due to noise, variations in speech quality, and limited language support, leading to erroneous recognition and increased user burden, particularly in hands-free operations.

Innovation Solution

A speech processing apparatus and method that performs morphological analysis on speech recognition results, generating partial character string candidates by dividing the input into pre-defined units, allowing users to select and correct erroneous parts, thereby enabling continued processing without manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech recognition is performed using conventional methods, then the system can process input speech, but erroneous recognition occurs due to noise, speech quality variations, and limited language support

Engineering Contradiction:
Improverecognition accuracyVSAvoidnoise and speech quality variations
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the speech recognition process into multiple stages: initial recognition to generate candidate strings, morphological analysis to break down words into components, and selective correction where users can modify specific portions of recognized text. This segmentation allows the system to handle recognition errors more effectively by processing and correcting parts of the speech input rather than requiring complete re-recognition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback mechanisms by presenting multiple candidate recognition results to users, allowing them to review and correct errors. The morphological analysis provides structured feedback by breaking down words into morphemes, helping users identify and correct specific erroneous portions without having to re-input the entire speech.

Inventive Principle:
Principle #23Feedback

2Device complexity

If the system outputs only the most probable candidate as the correct recognition result, then the processing is simple, but the system cannot identify which part of the recognition is wrong

Engineering Contradiction:
Improveprocessing simplicityVSAvoiderror identification capability
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies segmentation by breaking down recognized words into morphological components (morphemes, stems, prefixes, suffixes). This segmentation enables the system to present structured candidate variations to users, making it possible to identify which specific part of the recognition is incorrect without significantly increasing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of requiring users to review and correct the entire recognition result, the system performs partial action by allowing correction of only specific erroneous portions identified through morphological analysis. This reduces the burden on users while maintaining the ability to identify and correct errors effectively.

Inventive Principle:
Principle #16Partial or excessive action

3Ease of operation

If keyboard manipulation is required to correct recognition results, then correction can be performed, but hands-free characteristics are lost and operational burden increases

Engineering Contradiction:
Improvecorrection capabilityVSAvoidhands-free operation
Core Design Contradiction:
Ease of operationVSExtent of automation

Solution Approach 1:

The system enables self-service by automatically performing morphological analysis on recognized speech and generating structured candidate corrections. Users can then select from these pre-processed options using simple input methods, allowing correction without full keyboard manipulation while maintaining hands-free characteristics for the majority of the correction process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The morphological analysis structure serves as an intermediary between full speech recognition and final text output. It provides a structured intermediate representation that enables easier correction through simplified user interaction, reducing the need for extensive keyboard manipulation while preserving hands-free operation benefits.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If multiple recognition candidates are generated and presented, then users can select correct parts, but the processing complexity and time increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the candidate generation process through morphological analysis, creating structured candidates based on word components rather than generating all possible permutations. This segmentation reduces the number of candidates that need to be processed and presented to users, thereby reducing processing time while maintaining recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial action by generating and presenting only the most relevant candidate corrections based on morphological analysis, rather than exhaustively listing all possible recognition variations. This approach maintains high recognition accuracy while minimizing the time required to process and present candidates to users.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8954333B2Apparatus, method, and computer program product for processing input speech
Publication Date: 2015.02.10 TOSHIBA DIGITAL SOLUTIONS CORP
  • US8954333B2 patent drawing
  • US8954333B2 patent drawing
  • US8954333B2 patent drawing

AI summary

An analyzing unit performs a morphological analysis of an input character string that is obtained by processing input speech. A generating unit divides the input character string in units of division previously decided, that is composed of one or plural morphemes, and generates partial character strings including part of components of the divided input character string. A candidate output unit outputs the generated partial character strings to a display unit. A selection receiving unit receives a partial character string selected from the outputted partial character strings as a target to be processed.