A
system processing hardware executes a
machine learning (ML) model-based video compression
encoder to receive
uncompressed video content and corresponding motion compensated video content, compare the uncompressed and motion compensated video content to identify an image space residual, transform the image space residual to a latent space representation of the
uncompressed video content, and transform, using a trained
image compression ML model, the motion compensated video content to a latent space representation of the motion compensated video content. The ML model-based video compression
encoder further encodes the latent space representation of the image space residual to produce an encoded latent residual, encodes, using the trained
image compression ML model, the latent space representation of the motion compensated video content to produce an encoded latent video content, and generates, using the encoded latent residual and the encoded latent video content, a compressed video content corresponding to the
uncompressed video content.