The application discloses an implementation method and
system for self-evolution learning of an
intelligent agent, and belongs to the technical field of
intelligent agent reinforcement learning, large language models, memory and cognitive intelligence, and transfer learning; the method comprises the following steps: acquiring an environment state; inputting the environment state into a pre-trained strategy to output an optimal action and executing the action; after executing the action, a feedback
signal is acquired; the feedback
signal comprises an environment response and a
task completion degree; and a new strategy is obtained by optimizing the strategy according to the feedback
signal. The task performance of the application is continuously upgraded: the
intelligent agent can be continuously optimized when facing long-range or repetitive tasks, and the intelligent agent will become more and more skilled in
processing cross-platform complex tasks. The application reduces the research and application cost: the intelligent agent can reduce the dependence on manual work, can autonomously discover
reinforcement learning rules, does not need to continuously manually annotate data, and relies on environment feedback for iterative optimization.