Speech Verification Model Using Merged Recognition and Language Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition technologies face challenges in accurately evaluating linguistic accuracy and confidence measures, often resulting in ungrammatical sentences due to limited feature consideration, and are complex to implement in practical language models.

Innovation Solution

A speech processing apparatus and method that extracts recognition feature information and language feature information to create a verification model using a learning process, incorporating conditional random fields, to improve the accuracy of speech recognition results by considering a broader range of linguistic features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If only recognition features from speech recognizing unit are used for confidence measure evaluation, then the evaluation process is simple, but the linguistic accuracy and confidence measure evaluation accuracy are insufficient

Engineering Contradiction:
Improvelinguistic accuracy evaluation accuracyVSAvoidfeature extraction and model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges recognition features (from speech recognizing unit) with language features (from language model) to create a comprehensive verification model. This combination allows the system to evaluate both recognition confidence and linguistic accuracy, resolving the contradiction by integrating multiple feature sources to improve evaluation accuracy without overwhelming complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The verification model serves multiple functions: it evaluates recognition confidence measures, verifies linguistic accuracy, and identifies ungrammatical sentences. By making the model multi-functional, the system achieves high evaluation accuracy across different aspects without requiring separate simple evaluation mechanisms for each function

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If comprehensive language model features are integrated into verification, then verification accuracy improves, but the complexity of implementing language model increases

Engineering Contradiction:
Improveverification accuracyVSAvoidlanguage model implementation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the verification process into distinct components: recognition feature extraction, language feature extraction, and verification model integration. This segmentation allows the complex verification task to be broken down into manageable parts, improving reliability while controlling implementation complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The verification model acts as an intermediary that bridges the speech recognizing unit and the language model. It integrates features from both sources and produces verification results, thereby improving verification accuracy while managing complexity through this intermediate verification layer rather than directly integrating all features into a single complex model

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8751226B2Learning a verification model for speech recognition based on extracted recognition and language feature information
Publication Date: 2014.06.10 NEC CORP
  • US8751226B2 patent drawing
  • US8751226B2 patent drawing
  • US8751226B2 patent drawing

AI summary

A speech processing apparatus 101 includes a recognition feature extracting unit 12 that extracts recognition feature information which is a characteristic of a speech recognition result 15 obtained by performing a speech recognition process on an inputted speech from the speech recognition result 15; a language feature extracting unit 11 that extracts language feature information which is a characteristic of a pre-registered language resource 14 from the language resource 14; and a model learning unit 13 that obtains a verification model 16 by a learning process based on the extracted recognition feature information and language feature information.