The invention discloses a multi-
modal sentiment analysis model based on hierarchical adaptive cross-
modal fusion, belongs to the field of
natural language processing, and is used for solving the problems that in the prior art, multi-
granularity sentiment
feature extraction is insufficient, a cross-
modal fusion mechanism is rigid, and the distribution difference between different-source modals is large. The method comprises the following steps: firstly, extracting features of texts, audios and visual modalities from original video data, and coding the features into advanced semantic features; secondly, multi-
granularity information is fused through a hierarchical feature extractor to generate enhanced single-mode features; then, a self-adaptive cross-modal fusion network with a text as a core is adopted to realize bidirectional interaction and dynamic weighted fusion between modals; further, a
dynamic contrast learning mechanism is introduced to align modal distribution in a unified
potential space; and finally, inputting the optimized multi-modal features into a classifier and outputting an emotion analysis result. According to the model, through collaborative optimization of multi-
granularity feature extraction, adaptive fusion and comparative learning, the accuracy of
sentiment analysis and the robustness of the model are remarkably improved.