一种基于融合语音学知识低资源语言语音转国际音标方法

By employing fine-grained phoneme segmentation and continuous fundamental frequency feature fusion, combined with phonological prior constraints and language model decoding, the problem of transcribing low-resource language speech into the International Phonetic Alphabet was solved. This approach enables efficient and low-cost phoneme and tone recognition, supporting the digitization of endangered languages ​​and cross-linguistic information processing.

CN121415783BActive Publication Date: 2026-07-17INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES
Filing Date
2025-11-24
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing speech recognition systems struggle to accurately transcribe the International Phonetic Alphabet in low-resource language scenarios, particularly in tone and syllable recognition. Furthermore, traditional models are expensive to train and fail to meet the accuracy requirements of linguistic research.

Method used

By employing corpus preprocessing and label generation, pitch embedding and feature fusion, multi-task training and phonological consistency decoding, fine-grained phonological segmentation and a labeling system of initial consonants, final vowels and tone marks are adopted. Combined with continuous fundamental frequency features, a lightweight adapter is used to achieve transfer within a large model, and phonological prior constraints and language model decoding are added.

Benefits of technology

It significantly lowers the technical threshold for transcribing low-resource speech into the International Phonetic Alphabet, improves the accuracy of phoneme and tone recognition, reduces model output space and computational cost, provides interpretable results, and facilitates the digitization of endangered languages ​​and cross-linguistic information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415783B_ABST
    Figure CN121415783B_ABST
Patent Text Reader

Abstract

本发明涉及一种基于融合语音学知识低资源语言语音转国际音标方法。本发明通过在大模型框架内注入标签、连续嵌入与音系软约束解码,降低了极低资源语音转写为国际音标的技术门槛。借助轻量化适配器,三十小时左右语料即可完成迁移,避免了重新训练整网的高昂成本;依托声母‑韵母‑调号一体化的标签体系以及连续基频特征的融合,系统对音素和声调的识别精度同步提升。直接输出符合规范且附带置信度和基频曲线的可解释结果,通过声母→韵母→调号转移概率与语言模型的联合解码,彻底解决了非法音节和调类混淆问题。为濒危语言的数字化存续、跨语言信息检索及文化传承提供具有可复制性的技术范式,具有重要的学术价值与社会意义。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Cross-language end-to-end speech recognition method for low resource Tujia language

    CN109003601A

  • Modern voice collecting, recording, analyzing and displaying system

    CN117612553A