The invention provides a
video prediction method and
system,
computer equipment and a storage medium, and belongs to the field of
computer vision, and the method comprises the steps: obtaining a continuous
video sequence, and extracting a time feature, a bottom-layer texture, a middle-layer contour and a high-layer semantic multi-dimensional feature from the continuous
video sequence; calculating attention based on the front theta-layer
spatial memory to obtain the
current time step
spatial memory; based on the previous
time step hidden state, the previous tau step time memory and the
current time feature, generating the time memory of the
current time step, fusing the time memory, the
spatial memory and the previous
time step hidden state, and outputting a top layer hidden feature; and carrying out weighted fusion on the layered spatial features by using weights, adding the layered spatial features with top hidden features, then carrying out up-sampling, and finally generating a prediction frame. Dynamic changes and detail structures are considered through multi-dimensional
feature extraction; the spatial-temporal characteristic coherence is enhanced by using an attention mechanism; through
feature fusion and iteration, historical and current information is integrated, video frame detail pictures are greatly restored, and the video precision is improved.