A
video decoding method employing a UNet neural network and a Transform neural network, comprising:
parsing a first
syntax element from a
bitstream, the first
syntax element indicating whether to enable neural network inter-frame prediction; in response to the first
syntax element indicating that neural network inter prediction is enabled, a second syntax element is parsed from the
bitstream, the second syntax element indicating N frames after an intra-coded frame (
Intra frame), where: inter prediction is performed on the N frames using the UNet DM (
Diffusion Model), performing inter-frame prediction on a frame after the N frames by using a combination of UNet DM (
Data Management) and Transform DM (
Data Management); and in response to the first syntax element indicating that neural network inter prediction is enabled,
parsing a third syntax element from the
bitstream, the third syntax element indicating one of: (a) inter prediction is performed on frames following the N frames using UNet DM and Transform DM, respectively, to generate two frames, respectively, and selecting an optimal prediction
frame based on a quality comparison between the two frames; or (b) respectively carrying out inter-frame prediction on the frames after the N frames by using UNet DM and Transform DM so as to respectively generate two frames, and carrying out frame fusion on the two frames so as to obtain a final prediction frame.