The invention provides a uniform video segmentation method and
system based on sparse space-time propagation. The method comprises the following steps: generating an initial segmentation
mask for a video frame; extracting spatio-
temporal context features of a current frame and a historical frame by using a 3D spatio-temporal
convolution module, enhancing the spatio-
temporal context features, and fusing and coding the spatio-
temporal context features and an initial segmentation
mask of the current frame into a memory containing spatio-temporal contexts; aggregating the global multi-frame spatio-
temporal information into each memory frame during memory construction; the space-time aggregation reading module executes space-time aggregation reading based on a double-attention mechanism, and retrieves most
relevant information from a
memory bank; and performing sparse matching and fusion by adopting a non-local mechanism, and outputting a final video segmentation result. According to the method, multi-frame spatio-
temporal information is efficiently coded into a
memory bank, accurate and efficient sparse connection is established between the query frame and the
memory bank by using a dynamic 3D spatio-temporal
convolution module and a spatio-temporal aggregation reading strategy, video-level spatio-temporal relation modeling is realized, and a relatively high
processing speed is kept while the segmentation precision is remarkably improved.