The application provides a multi-agent
reinforcement learning active power distribution network collaborative optimization method and
system, belongs to the field of active power distribution network operation optimization and
intelligent control, and constructs a global linearized
voltage sensitivity model; each distributed resource is regarded as an agent, a local observation vector of each agent is input into a policy network to obtain an original action, and a real node
voltage is calculated; whether the real node
voltage satisfies a voltage safety constraint is checked, if yes, the original action is taken as an execution instruction, and if not, step 4 is entered; in step 4, online correction is performed, active injection adjustment, reactive injection adjustment and node voltage relaxation variable are solved jointly, and a candidate safety action is obtained; in step 5,
power flow checking is performed, and a
linearization error compensation term is updated; if the real node voltage does not satisfy the voltage safety constraint, the updated
linearization error compensation term is substituted into step 4, an action correction amount is solved, and steps 4 and 5 are iteratively executed; and the safety projection layer prediction accuracy and the action correction reliability are improved.