一种数据动态平衡藏语多方言语音大模型训练方法及装置
By performing dynamic data balancing training on a large Tibetan multi-dialect speech model and utilizing temperature sampling strategies and complementary acoustic feature fusion, the data imbalance problem of the low-resource Tibetan multi-dialect speech model was solved, achieving efficient model performance improvement and recognition and translation effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MINZU UNIVERSITY OF CHINA
- Filing Date
- 2026-05-15
- Publication Date
- 2026-07-17
AI Technical Summary
Existing large-scale speech models suffer from extremely unbalanced data distribution in low-resource languages such as Tibetan, leading to overfitting of high-resource dialects and the submergence of feature contributions from low-resource dialects, making it difficult to construct general and efficient large-scale speech models for multiple dialects.
By collecting multi-source heterogeneous speech data, combining and dividing it according to dialect attributes and task attributes, and using linear discriminant analysis and temperature sampling strategies to increase the sampling weight of scarce dialects, data clustering and mutual assistance strategies are implemented to achieve the alternation of different dialects in data batches with a relatively balanced proportion. Finally, acoustic feature complementarity fusion and model-induced training are carried out.
The model integrates low-resource dialect knowledge with high-resource dialect knowledge, which improves the recognition and translation performance of low-resource dialects, reduces the risk of overfitting, and ensures the model's ability to model the overall structure of the language. Its performance is close to or better than that of high-resource dialects.
Smart Images

Figure CN122201263B_ABST