一种基于鉴别器思想的文本蒸馏方法、系统和存储介质
By employing a text distillation method based on the discriminator concept, teacher and student models are trained using labeled and unlabeled text datasets. The student model parameters are optimized by combining mask training. This resolves the performance-scale contradiction in model compression and enables efficient application on low-resource devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU YIWISE INTELLIGENT TECH CO LTD
- Filing Date
- 2022-07-20
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies, knowledge distillation methods struggle to effectively compress model size without diminishing the learning capacity of student models, making them perform similarly to teacher models on resource-constrained devices.
We employ a text distillation method based on the discriminator concept. By acquiring labeled and unlabeled text datasets, we train the teacher and student models using knowledge distillation. We then combine this with mask training to test the learning performance of the student model and update the student model parameters using knowledge distillation loss and mask training loss.
While reducing the number of parameters in the student model, its performance is improved, enabling it to perform similarly to the teacher model on low-resource devices, and the student model can self-test and improve during the learning process.
Smart Images

Figure CN115271064B_ABST