一种基于CTC的非自回归端到端语音翻译方法
By introducing CTC technology and a pre-trained model into a non-autoregressive speech translation model, an acoustic and semantic encoder is constructed, which solves the problem of translation quality degradation in cross-modal modeling and achieves fast and accurate speech translation results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAONIU FANYI
- Filing Date
- 2023-01-06
- Publication Date
- 2026-07-17
AI Technical Summary
Non-autoregressive speech translation models suffer from a decline in translation quality in cross-modal modeling, and pre-training methods are difficult to effectively improve model performance.
We adopt a CTC-based non-autoregressive end-to-end speech translation method. We use a Transformer model structure with self-attention mechanism and CTC technology to construct an acoustic encoder, semantic encoder, adapter and decoder. We combine intermediate and top-level CTC loss, use pre-trained speech recognition and machine translation models to initialize parameters, and decode through a CTC greedy search strategy.
It achieves fast and accurate non-autoregressive speech translation with performance approaching that of autoregressive models, improving translation quality and accelerating the inference process.
Smart Images

Figure CN116227503B_ABST