The invention relates to the technical field of
subtitle translation, in particular to a multi-
modal time sequence alignment AI video translation method and
system, and the method comprises the steps: 1, carrying out the multi-
modal analysis of a to-be-translated video, and obtaining audio separation data, voiceprint
feature data and visual
time sequence data; 2, performing cross-
language translation and context optimization on the basis of the voice of the audio separation data to generate a target language text, and synthesizing target language voice retaining the original voice color in combination with the voiceprint
feature data and the target language text; generating a
mouth shape animation matched with the target language voice based on the lip key
point data and the limb action
time sequence data; and step 3, performing four-dimensional alignment on the target language voice, the translated text, the
mouth shape animation and the limb action sequence through a cross-
modal time sequence
encoder, and dynamically adjusting the
layout of the bilingual subtitles to adapt to a video picture. According to the method and the device, multi-mode synchronization can be taken into consideration during video translation, so that the body actions such as voice, subtitles and mouth shapes are kept aligned.