一种基于CTC的非自回归端到端语音翻译方法

By introducing CTC technology and a pre-trained model into a non-autoregressive speech translation model, an acoustic and semantic encoder is constructed, which solves the problem of translation quality degradation in cross-modal modeling and achieves fast and accurate speech translation results.

CN116227503BActive Publication Date: 2026-07-17XIAONIU FANYI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAONIU FANYI
Filing Date
2023-01-06
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Non-autoregressive speech translation models suffer from a decline in translation quality in cross-modal modeling, and pre-training methods are difficult to effectively improve model performance.

Method used

We adopt a CTC-based non-autoregressive end-to-end speech translation method. We use a Transformer model structure with self-attention mechanism and CTC technology to construct an acoustic encoder, semantic encoder, adapter and decoder. We combine intermediate and top-level CTC loss, use pre-trained speech recognition and machine translation models to initialize parameters, and decode through a CTC greedy search strategy.

Benefits of technology

It achieves fast and accurate non-autoregressive speech translation with performance approaching that of autoregressive models, improving translation quality and accelerating the inference process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116227503B_ABST
    Figure CN116227503B_ABST
Patent Text Reader

Abstract

本发明公开一种基于CTC的非自回归端到端语音翻译方法,步骤为:构建编码器‑解码器结构的非自回归端到端语音翻译模型;构建语音识别模型和机器翻译模型作为预训练模型,使用音频文件和源语言文本数据训练语音识别模型;使用源语言文本和目标语言文本数据训练机器翻译模型;以语音识别模型和机器翻译模型作为预训练模型,使用两个模型的参数来初始化非自回归语音翻译模型的参数;在语音翻译数据集上对参数进行微调,完成训练过程;语音翻译模型的编码器进行编码,解码器根据编码器的输出结果进行解码,生成最终的目标语言文本。本发明增强了语音翻译模型的编码能力,在获得与自回归模型取得相似性能的情况下,能够获得较大的速度提升。
Need to check novelty before this filing date? Find Prior Art