A method for performing content-based video compression using
reinforcement learning (RL) is provided. The method includes obtaining frame information associated with a frame from a video. The frame information comprises quantization parameter (QP) information associated with the frame, and the QP information indicates an initial compression level for encoding aspects of the frame. The frame information and additional information are processed by an RL agent to generate a generated QP map indicating a plurality of updated values associated with a plurality of
macro-blocks (MBs) of the frame. A
bitstream is generated comprising a plurality of bits for the
frame based on the generated QP map. Specifically, the plurality of updated values from the generated QP map indicates an amount of allocated bits from the
bitstream to allocate for each of the plurality of MBs. The
bitstream is provided to a downstream model.