Innovations in systems, methods, and
software for features of a neural image or video codec are described herein. For example, a neural video
encoder can receive a current video frame,
encode the current video frame to produce encoded data, and output the encoded data as part of a
bitstream. As part of the encoding, the
encoder can determine a current latent representation for the current video frame, and
encode the current latent representation using an
entropy model network that includes one or more convolutional
layers. As part of the encoding the current latent representation, the
encoder can estimate statistical characteristics of a quantized version of the current latent representation based at least in part on a previous latent representation for a previous video frame, and entropy code the quantized version of the current latent representation based at least in part on the estimated statistical characteristics.