Dialect Speech Recognition via Data Mining and Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in recognizing dialects, as they often convert dialect speech into standard dialect, leading to reduced recognition capabilities and requiring manual transcription, which is time-consuming and expensive, and are delayed by the presence of dialects.
Innovation Solution
A speech recognition method and system that selects and refines dialect speech data, extracts features, performs similar dialect clustering, and standardizes dialect corpora using a data mining device with modules for feature extraction, deep learning, core dialect extraction, and corpus standardization, enabling dialect recognition without conversion to standard dialect.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If dialect speech is converted to standard dialect through noise removal, then recognition accuracy is improved, but dialect recognition capability is reduced
Solution Approach 1:
The patent segments dialect speech processing into two distinct paths: standardization for recognition accuracy and preservation for dialect capability. The system separates noise removal operations from dialect feature preservation, allowing different processing streams to serve different objectives simultaneously.
Solution Approach 2:
The patent applies different quality standards to different aspects of speech processing. Standard dialect forms receive noise removal and standardization for accurate recognition, while dialect-specific features are preserved with minimal intervention to maintain dialect recognition capability. This local differentiation resolves the contradiction between accuracy and adaptability.
2Reliability
If manual transcription is performed for speech data processing, then processing completeness is improved, but time consumption and cost increase
Solution Approach 1:
The patent implements self-service through automated dialect clustering and standardization systems that process speech data without manual transcription. The system automatically identifies dialect patterns, clusters similar dialects, and standardizes them for recognition, eliminating the need for time-consuming manual transcription while maintaining processing completeness through algorithmic accuracy.
Solution Approach 2:
The patent replaces the mechanical process of manual transcription with automated computational systems. Machine learning algorithms and automatic speech processing systems substitute human transcription work, dramatically reducing time consumption and cost while maintaining or improving processing completeness through consistent, scalable automation.
3Adaptability or versatility
If dialect speech is recognized without conversion, then dialect recognition capability is improved, but recognition performance deteriorates
Solution Approach 1:
The patent introduces dialect clustering and standardization as intermediary processes between raw dialect speech and recognition systems. These intermediaries transform diverse dialects into standardized forms that maintain dialect identity while improving recognition performance. The standardization layer acts as a mediator that preserves dialect capability while enhancing recognition accuracy.
Solution Approach 2:
The patent changes key parameters of dialect speech through automated standardization processes. By systematically adjusting phonetic, phonological, and syntactic parameters while preserving dialect-specific characteristics, the system improves recognition performance without sacrificing dialect recognition capability. Parameter transformation enables both goals to coexist.
4Measurement precision
If speech data is refined to consistent form for language model learning, then model accuracy is improved, but applicability to non-standard dialects is reduced
Solution Approach 1:
The patent implements dynamic language model learning that adapts to different dialect types. Rather than using a fixed standardization approach, the system dynamically adjusts processing parameters based on detected dialect characteristics. This dynamic approach allows the model to maintain high accuracy for each dialect type while preserving broad dialect applicability across diverse speech varieties.
Solution Approach 2:
The patent creates a universal language model framework that handles both standard and non-standard dialects through a single system. The multi-functional processing pipeline applies appropriate standardization and clustering techniques based on input dialect type, enabling the model to achieve high accuracy across multiple dialect varieties without requiring separate specialized models for each dialect.
Data Source
AI summary
A data mining device, and a speech recognition method and system using the same are disclosed. The speech recognition method includes selecting speech data including a dialect from speech data, analyzing and refining the speech data including a dialect, and learning an acoustic model and a language model through an artificial intelligence (AI) algorithm using the refined speech data including a dialect. The user is able to use a dialect speech recognition service which is improved using services such as eMBB, URLLC, or mMTC of 5G mobile communications.


