一种数据动态平衡藏语多方言语音大模型训练方法及装置

By performing dynamic data balancing training on a large Tibetan multi-dialect speech model and utilizing temperature sampling strategies and complementary acoustic feature fusion, the data imbalance problem of the low-resource Tibetan multi-dialect speech model was solved, achieving efficient model performance improvement and recognition and translation effects.

CN122201263BActive Publication Date: 2026-07-17MINZU UNIVERSITY OF CHINA +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MINZU UNIVERSITY OF CHINA
Filing Date
2026-05-15
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing large-scale speech models suffer from extremely unbalanced data distribution in low-resource languages ​​such as Tibetan, leading to overfitting of high-resource dialects and the submergence of feature contributions from low-resource dialects, making it difficult to construct general and efficient large-scale speech models for multiple dialects.

Method used

By collecting multi-source heterogeneous speech data, combining and dividing it according to dialect attributes and task attributes, and using linear discriminant analysis and temperature sampling strategies to increase the sampling weight of scarce dialects, data clustering and mutual assistance strategies are implemented to achieve the alternation of different dialects in data batches with a relatively balanced proportion. Finally, acoustic feature complementarity fusion and model-induced training are carried out.

Benefits of technology

The model integrates low-resource dialect knowledge with high-resource dialect knowledge, which improves the recognition and translation performance of low-resource dialects, reduces the risk of overfitting, and ensures the model's ability to model the overall structure of the language. Its performance is close to or better than that of high-resource dialects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122201263B_ABST
    Figure CN122201263B_ABST
Patent Text Reader

Abstract

本发明公开一种数据动态平衡藏语多方言语音大模型训练方法及装置,涉及自然语言处理技术领域。方法包括:采集多源异构语音数据,按藏语方言属性与语音识别、语音翻译任务属性划分独立数据桶;再依有效性校验规则筛选样本,确定方言聚类中心,通过声学关联分析确立语言枢纽,结合方言化温度采样的均衡采样概率分布抽取含高低资源方言的混合训练批次,输入端到端语音大模型,最终得到融合高资源方言知识的低资源藏语方言模型。本发明能够平滑地放大稀缺数据的采样权重,同时保留高资源数据的结构化信息,从而诱导模型利用优势方言的声学特征辅助弱势方言的学习,确保了模型在多方言、多任务场景下的训练稳定性和综合性能。
Need to check novelty before this filing date? Find Prior Art