The invention discloses a double-
branch aesthetic
image processing method based on active
reinforcement learning. Firstly, a unified Markov
decision process framework is constructed, two subtasks of
color enhancement and composition optimization are processed in parallel, and a
reinforcement learning training
data set is constructed based on images. Thirdly, initializing a double-
branch strategy network, a
value network and a pixel-level agent
system, providing pixel-level and channel-level instant feedback rewards through a pre-trained aesthetic model, and updating network parameters in combination with human subjective preference constraints; and after training is completed, a high-quality
image sequence after progressive optimization is output. According to the invention, the pixels are defined as intelligent agents, so that the
motion space and
fineness of adjustment are expanded; and meanwhile, an aesthetic evaluation model is taken as a reward core, so that the result is consistent with expert modification in objective indexes and subjective scores, and personalized and progressive aesthetic modification requirements are met.