The application discloses a general-purpose learning method and
system for a physical
robot, which comprises a special talent model
library, a monitor set, a general talent model and a multi-
robot distributed exploration cluster. The special talent model
library contains L pre-trained special talent models. The monitor set contains L monitors, each of which is arranged with a special talent model and used for evaluating whether a current input state belongs to the use range of the special talent model. The multi-
robot distributed exploration cluster contains multiple physical robots, each of which is arranged in a different task scene. In the application, multiple robots are introduced to perform different tasks in parallel, the monitor is used to dynamically determine when the special talent model needs to be involved, the appropriate special talent model is selected for guidance, the DAGGER type iterative data aggregation mechanism is adopted, the general talent model continuously learns on the
state distribution generated by self-exploration in multiple scenes, and finally, the general talent model which can adapt to multiple scenes and surpass a single special talent is obtained.