Regional Dialect Speech Recognition via Acoustic Model Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems are inadequate for recognizing regional dialects, as they are based on standard dialects and often require manual transcription, leading to reduced recognition accuracy and increased time and expense.

Innovation Solution

A language modeling system that sorts and transcribes regional dialect-containing speech data, generates a regional dialect corpus, and trains acoustic and language models using specific features and deep learning to recognize regional dialects without conversion to standard dialects, employing a semi-automatic data processing method.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If standard dialect-based speech recognition systems are used, then recognition accuracy for standard dialect is maintained, but recognition accuracy for regional dialects deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoiddialect adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the speech recognition system into multiple dialect-specific components. It creates separate acoustic models and language models for different dialects (standard dialect, regional dialects) rather than using a single unified model. This segmentation allows the system to maintain high accuracy for each specific dialect while improving overall dialect adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of dialect classification and modeling. It adds dialect type as an additional dimension to the speech recognition system, creating a multi-dimensional framework that handles both standard and regional dialects. This is achieved by incorporating dialect-specific features and creating parallel processing paths for different dialect types.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If manual transcription is performed for speech data processing, then transcription accuracy is improved, but processing time and cost increase

Engineering Contradiction:
Improvetranscription accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing speech data to extract dialect-specific features before transcription. It performs preliminary dialect identification and feature extraction, which prepares the data in a way that facilitates more accurate and efficient transcription. This preliminary processing reduces the burden on manual transcription while maintaining or improving accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent enables the speech recognition system to perform self-service through automated dialect identification and feature extraction. The system automatically analyzes speech patterns, identifies dialect types, and extracts relevant features without requiring manual intervention. This self-service capability significantly reduces processing time and cost while maintaining high transcription accuracy through automated algorithms.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11056100B2Acoustic information based language modeling system and method
Publication Date: 2021.07.06 LG ELECTRONICS INC
  • US11056100B2 patent drawing
  • US11056100B2 patent drawing
  • US11056100B2 patent drawing

AI summary

Disclosed are a speech data based language modeling system and method. The speech data based language modeling method includes transcription of text data, and generation of a regional dialect corpus based on the text data and regional dialect-containing speech data and generation of an acoustic model and a language model using the regional dialect corpus. The generation of an acoustic model and a language model is performed by machine learning of an artificial intelligence (AI) algorithm using speech data and marking of word spacing of a regional dialect sentence using a speech data tag. A user is able to use a regional dialect speech recognition service which is improved using 5G mobile communication technologies of eMBB, URLLC, or mMTC.