The application discloses a kind of multimodal
sentiment analysis method, device, equipment and storage medium.The method includes: constructing the
sentiment analysis dataset containing text, audio and visual sequence and inputting pre-training model, extracting each modality initial feature using
feature coding module, comparing learning and dynamic fusion are carried out through
modal common
information processing module to obtain cross-
modal common feature, with the aid of
modal specificity
information processing module, feature enhancement and
pooling are implemented to generate global enhanced features, finally, each kind of feature is spliced and regression prediction is executed to output sentiment intensity prediction value by fusion prediction module;Since the application cooperates modeling modal common and specificity information, the modal heterogeneity is considered and the average
processing bottleneck is broken, different dimension features can be dynamically distinguished, the modal conflict can be effectively relieved, and the prediction accuracy and robustness of
sentiment analysis can be significantly improved.