The invention belongs to the technical field of
animal behavior recognition and emotion calculation, and particularly relates to a text and
mask multi-
modal guided dynamic video
emotion recognition method, which comprises the following steps of: constructing a dynamic golden snub monkey emotion
data set, and constructing multi-
modal input containing videos, key points, masks and text description; obtaining a golden snub monkey
video sequence, and extracting face key points and
mask information; inputting the
video sequence, the key points, the face
mask and the text description into a double-flow guide module to generate a double-flow guide prompt; inputting a double-flow guidance prompt into a space attention priori module, and performing feature enhancement through a channel attention mechanism and a multi-head space-time
perceptron; extracting features from the text description through a
text enhancement network; and finally, adaptively adjusting the contribution weight of the text
semantics to the visual features through a dynamic gating fusion module to realize accurate recognition of the emotion of the golden snub monkey. According to the method, the problem of cross-species dynamic
emotion recognition is effectively solved through multi-
modal information collaborative fusion and dynamic modeling.