Side-end speech recognition model compression method based on knowledge distillation

By combining knowledge distillation techniques from deep neural networks and Hamiltonian neural networks, the student model structure is dynamically adjusted, solving the adaptability problem of the side-end speech recognition model on resource-constrained devices and achieving efficient and stable speech recognition results.

CN122416992APending Publication Date: 2026-07-17

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Filing Date
2026-04-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing side-end speech recognition models are difficult to compress effectively on resource-constrained devices, and existing knowledge distillation methods fail to fully utilize the intermediate layer features of the teacher model, resulting in insufficient modeling of temporal evolution relationships in the student model and poor adaptability.

Method used

This paper adopts a method combining deep neural networks and Hamiltonian neural networks. Through knowledge distillation, intermediate layer features are extracted from the teacher model and passed to the student model. The structural parameters of the student model are dynamically adjusted, and the model is optimized in combination with edge device resource conditions. The dynamic modeling of Hamiltonian neural network is introduced to optimize temporal feature learning.

Benefits of technology

It improves speech recognition accuracy, reduces computational complexity and storage requirements, enhances the model's adaptability and operating efficiency on edge devices, and achieves stable, low-latency speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416992A_ABST
    Figure CN122416992A_ABST
Patent Text Reader

Abstract

本发明公开了基于知识蒸馏的侧端语音识别模型压缩方法,包括如下步骤:获取、预处理语音数据集,生成标准化的语音特征集;基于语音特征集利用深度神经网络构建、训练教师模型,包含卷积层、池化层和全连接层;将教师模型输出作为软标签,输入至简化的学生模型进行训练;提取教师模型中间层特征,传递至学生模型的对应层;根据边缘设备资源动态调整学生模型的参数,优化计算复杂度;利用哈密顿神经网络的动力学建模优化学生模型的时序特征学习;将训练后的学生模型部署至边缘设备,根据实时反馈进行微调。本发明实现了边缘设备上语音识别模型的高效压缩与部署,在保证准确性的同时也降低了计算复杂度和存储需求。
Need to check novelty before this filing date? Find Prior Art