The invention relates to the technical field of streaming media detection, in particular to a complex element new communication streaming media detection method based on a multi-
modal deep learning recognition technology, which comprises the following steps: S1, multi-
modal time-space synchronization preprocessing: mapping video key frames, audio clips and bullet screen texts to a unified time axis through a combined time-space calibration technology, establishing spatial
semantic association; s2, hierarchical multi-
modal feature
distillation is carried out, and discriminative multi-
granularity features including local details, global
semantics and cross-modal association
modes are extracted from all modals; and S3, establishing a dynamic graph modal
interaction network, constructing a learnable multi-modal
relation graph, and dynamically modeling cross-modal semantic interaction. According to the complex element new communication streaming media detection method based on the multi-modal
deep learning recognition technology, the problem that cross-modal complex semantic
collaboration cannot be captured through single-
modal analysis or shallow fusion, so that the detection missed judgment rate is high is solved.