The invention discloses an initial QP selection optimization
algorithm based on hierarchical
reinforcement learning. According to the
algorithm, action QP value selection is carried out from a rough
target range to a fine
target range by adopting a hierarchical
reinforcement learning strategy; a two-layer decision selection architecture is adopted, and an enhanced
Q learning algorithm is adopted; a
reward value, namely a Q value, is obtained through CTU-level
code rate control and an actual coding process in sequence; observing the next state and the last reward, performing TD iteration on the Q value, and updating the Q value to a corresponding Q table; performing an initial QP coding test based on the Q table obtained by training; the current state information is extracted, the Q table is inquired according to the current state, and the Q table gives a preferred action, namely, a preferred QP value; and then carrying out subsequent operation according to a coding process. According to the initial QP selection optimization algorithm based on hierarchical
reinforcement learning, a long-term
rate distortion target and a current coding target of
video sequence coding are considered at the same time, and the method has a good capability of guiding
code rate control optimization coding.