The invention discloses a cross-
modal attention, global memory and dynamic
convolution-based
speech translation model training method and device, and a
speech translation method and device, and relates to the technical field of
speech processing and
machine translation. A
speech translation model is designed to comprise a speech
encoder, a text embedding layer, a cross-
modal attention adapter, a large
language model decoder, a global memory network, a dynamic
convolution decoder and an output layer. The cross-
modal attention adapter projects audio features and performs multi-head cross attention fusion with text embedding; the global memory network updates and enhances historical memory on the basis of a gating mechanism and a Transform
Encoder; and the dynamic
convolution decoder performs multi-scale convolution extraction on the decoded hidden representation and fuses with the memory, so that the translation quality is improved. According to the method, deep fusion of voice and text, continuous memory with contextual coherence and high-quality translation generation can be realized, the end-to-end voice translation performance is remarkably improved, and the actual requirements of real-time and high-quality end-to-end voice translation in a complex scene are met.