Filtering Model Training for Speech Recognition Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition technologies face challenges in accurately recognizing speech due to variations in environment, user articulation habits, and inconsistencies between training data and actual speech signals, leading to inaccurate recognition results.

Innovation Solution

A filtering model training method that determines syllable distances between original and recognized syllables, using these distances to train a filtering model that improves the matching of sound signals with the speech recognition engine, thereby enhancing recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech recognition is performed using a preset database and speech recognition engine, then speech recognition can be implemented, but recognition accuracy deteriorates due to environmental variations, user articulation habits, and inconsistencies between training data and actual speech signals

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidadaptability to environment and user articulation habits
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by collecting actual speech data from users in their real environments and training a filtering model in advance. This pre-trained filtering model is then used to process new speech inputs, adapting to individual user characteristics and environmental conditions before actual recognition occurs, thereby improving accuracy without requiring changes to the core speech recognition engine

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a filtering model as an intermediary component between the speech signal input and the speech recognition engine. This filtering model processes and adapts the speech signals based on user-specific characteristics and environmental conditions, serving as a mediator that bridges the gap between diverse real-world inputs and the fixed speech recognition engine, thereby improving recognition accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If a filtering model is trained using syllable distances between original and recognized syllables, then speech recognition accuracy is improved, but system complexity increases due to additional training processes and data processing

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidfiltering model training complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by using syllable distance as a specific training parameter to optimize the filtering model. Instead of training on raw speech signals directly, the system calculates syllable distances between original and recognized syllables and uses these distance metrics as training parameters. This transforms the complex speech recognition problem into a more manageable parameter optimization problem, improving accuracy while controlling complexity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system applies self-service by automatically collecting user speech data, calculating syllable distances, and training the filtering model without requiring manual intervention. The system serves itself by generating its own training data from actual usage and automatically optimizing its filtering model based on performance feedback, reducing the need for external configuration and maintenance

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11211052B2Filtering model training method and speech recognition method
Publication Date: 2021.12.28 YINWANG INTELLIGENT TECHNOLOGIES CO LTD
  • US11211052B2 patent drawing
  • US11211052B2 patent drawing
  • US11211052B2 patent drawing

AI summary

A filtering model training method includes obtaining N original syllables, obtaining N recognized syllables, and obtaining N syllable distances based on the N original syllables and the N recognized syllables, where the N syllable distances are in a one-to-one correspondence with N syllable pairs, the N original syllables and the N recognized syllables form the N syllable pairs, each syllable pair includes an original syllable and a recognized syllable that correspond to each other, and each syllable distance is used to indicate a similarity between an original syllable and a recognized syllable that are included in a corresponding syllable pair.