The invention is applied to the technical field of image
subtitle generation, and particularly discloses an image
subtitle method based on an attention and
state space model, which comprises the following steps: S1, constructing an image
subtitle dynamic
hybrid network model based on an attention and
state space; according to the image subtitle method based on the attention and the
state space model, an
encoder and a decoder serve as a framework, a
hybrid encoder is constructed, multi-
modal features of an image serve as input of the
encoder, an attention mechanism is adopted to capture relevance in the
modal features,
serialization features are extracted in combination with the state
space model, and the image subtitle method based on the attention and the state
space model is obtained. Fusion is carried out through a self-adaptive gating mechanism, and rich feature information is provided for the decoder; according to the method,
word embedding and multi-
modal feature interaction are carried out in a decoder, a dependency relationship between
word embedding and multi-modal features is obtained, a dynamic fusion module is designed to realize multi-modal dynamic fusion, multi-modal heterogeneous features of an image can be fully utilized to provide richer feature information, and
sentence description better conforming to human
cognition is generated.