There is provided a computer-implemented method (200) for training a
reinforcement learning agent to adjust or maintain a parameter of a first
cell of a communication network. The method (200) comprises obtaining (202) a dataset comprising a first state of the first
cell at a first time instance, a second state of the first
cell at a second time instance, a first action, wherein the first action, when co-occurring with the first state at any time instance, causes the first cell to transition to the second state, and a first
reward value for performing the first action, determining (204), based on the dataset, a second action, wherein the second action, when co-occurring with the second state at any time instance, causes the first cell to transition to the first state, and a second
reward value for performing the second action, obtaining (206) a first information comprising the first state, the second state, the first action, and the first
reward value, obtaining (208) second information comprising the second state, the first state, the second action, and the second reward value, and training (210) the
reinforcement learning agent based on the first information and the second information.