Regional Speech Recognition via Segmented Acoustic Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems struggle to accurately recognize regional dialects and vocabulary characteristics, as they are typically based on standard language models, leading to reduced recognition capabilities and loss of cultural heritage.

Innovation Solution

A regional-features-based speech recognition method and system that learns speech features by region using classified speech data, employing an acoustic model and language model trained on region-specific corpora, with vector labeling to prioritize region-specific words, and utilizing parallel processing by-region speech recognizers to predict and select the most accurate region category.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a standard language model is used for speech recognition, then the system structure remains simple and unified, but the recognition accuracy for regional dialects and vocabulary significantly deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech recognition system is segmented into multiple regional speech recognizers, each specialized for a specific region or dialect. The system divides the recognition task by creating separate acoustic models and language models for different regions (e.g., Seoul, Busan, Incheon), allowing each recognizer to optimize for its specific regional characteristics while maintaining overall system functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each regional speech recognizer is equipped with locally optimized acoustic models and language models trained on region-specific speech data. The system applies local quality by tailoring the recognition parameters, vocabulary, and linguistic patterns to match each region's unique speech characteristics, thereby improving recognition accuracy for local dialects without requiring a complete system redesign.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If speech data is classified and processed by region category, then recognition performance for regional dialects improves, but the processing time and computational resources increase

Engineering Contradiction:
Improvedialect recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary classification of input speech to identify the most likely region category before full recognition processing. By pre-classifying speech data into regional categories using acoustic features and statistical models, the system prepares the appropriate regional recognizer in advance, reducing the actual recognition time while maintaining high accuracy for regional dialects.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial processing by focusing computational resources only on the most relevant regional recognizers based on the input speech characteristics. Instead of running all regional recognizers simultaneously, the system selectively activates only those recognizers that match the detected regional features, thereby reducing overall processing time while maintaining high recognition accuracy for the identified dialect.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If region-specific acoustic models and language models are trained, then recognition accuracy for regional features improves, but the model training complexity and data requirements increase

Engineering Contradiction:
Improveregional speech recognition accuracyVSAvoidmodel training complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system employs universal training methodologies and shared model architectures that can be applied across multiple regions. By using a common framework for acoustic model and language model training that can accommodate different regional datasets, the system reduces training complexity while maintaining the ability to capture region-specific characteristics. The universal approach allows reuse of training procedures, feature extraction methods, and model structures across all regional recognizers.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system optimizes model training by adjusting key parameters such as data sampling rates, feature extraction parameters, and model architecture hyperparameters based on regional characteristics. By systematically varying these parameters to match regional speech patterns (e.g., accent intensity, vocabulary frequency, speech rate), the system achieves high recognition accuracy without requiring completely separate training pipelines for each region, thereby reducing overall training complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11488587B2Regional features based speech recognition method and system
Publication Date: 2022.11.01 LG ELECTRONICS INC
  • US11488587B2 patent drawing
  • US11488587B2 patent drawing
  • US11488587B2 patent drawing

AI summary

Disclosed is a regional-features-based speech recognition method, including learning speech features by region using speech data classified by region category, and recognizing input speech using an acoustic model and a language model generated through classification of a region category for the input speech and the learning. A user may use a dialect recognition service that is improved using learning based on artificial intelligence (AI) and enhanced mobile broadband (eMBB), ultra-reliable and low latency communications (URLLC), and massive machine-type communications (mMTC) techniques of 5G mobile communication.