一种基于融合语音学知识低资源语言语音转国际音标方法
By employing fine-grained phoneme segmentation and continuous fundamental frequency feature fusion, combined with phonological prior constraints and language model decoding, the problem of transcribing low-resource language speech into the International Phonetic Alphabet was solved. This approach enables efficient and low-cost phoneme and tone recognition, supporting the digitization of endangered languages and cross-linguistic information processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES
- Filing Date
- 2025-11-24
- Publication Date
- 2026-07-17
AI Technical Summary
Existing speech recognition systems struggle to accurately transcribe the International Phonetic Alphabet in low-resource language scenarios, particularly in tone and syllable recognition. Furthermore, traditional models are expensive to train and fail to meet the accuracy requirements of linguistic research.
By employing corpus preprocessing and label generation, pitch embedding and feature fusion, multi-task training and phonological consistency decoding, fine-grained phonological segmentation and a labeling system of initial consonants, final vowels and tone marks are adopted. Combined with continuous fundamental frequency features, a lightweight adapter is used to achieve transfer within a large model, and phonological prior constraints and language model decoding are added.
It significantly lowers the technical threshold for transcribing low-resource speech into the International Phonetic Alphabet, improves the accuracy of phoneme and tone recognition, reduces model output space and computational cost, provides interpretable results, and facilitates the digitization of endangered languages and cross-linguistic information processing.
Smart Images

Figure CN121415783B_ABST
Abstract
Citation Information
Patent Citations
Cross-language end-to-end speech recognition method for low resource Tujia language
CN109003601A
Modern voice collecting, recording, analyzing and displaying system
CN117612553A