Training method of action decision-making model of intelligent agent, action decision-making method and device

By constructing a state topology diagram and training action feedback model, combining the action value model, using human preference data to optimize the agent's action decision model, the problem of insufficient action decision-making of the agent is solved, and more accurate action execution and task completion are achieved.

CN120297323APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202410039699.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the performance of the action decision model of the agent is insufficient, resulting in the agent being unable to perform appropriate actions and unable to complete the task.

Method used

By constructing a state topology diagram, training the action feedback model and action value model, combining the action decision model, collaboratively optimizing the action decision process of the agent, using human preference data for supervised learning, and improving the accuracy of the action decision model.

Benefits of technology

It improves the accuracy and performance of the action decision model, allowing the agent to perform more accurate actions in a given state, meets human intentions and preferences, and improves task completion ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297323A_ABST
    Figure CN120297323A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of an action decision model of an intelligent agent and an action decision method and device, and belongs to the technical field of computers. According to the method, the state topological graph is constructed on the basis of the historical trajectory, experience distribution of actions of the intelligent agent can be fully reflected, the information utilization rate of the historical trajectory is higher, more information is brought, the action feedback model is guided and trained on the basis of the state topological graph, the accuracy of the action feedback model is improved, and the accuracy of the action feedback model is improved. A state topological graph, an action feedback model and a training process of a constraint action value model are combined to obtain an action value model with better accuracy and better performance, and the action value model is utilized to assist in training to obtain an action decision model with better accuracy, so that an accurate decision on which action is executed by an intelligent agent in a given state is facilitated; and the action decision model can be combined with a large model to mutually promote training, so that the performance of the two parties is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method for training an action decision model of an agent, an action decision method, and a device thereof. Background Art

[0002] With the development of computer and robot technologies, agents represented by robots, robotic arms, and large models have attracted wide attention. Currently, the application scenarios of agents have gradually expanded to cover many industrial application scenarios such as robot control, video, and games, and also involve the interaction and non-interaction scenarios between agents and humans.

[0003] Which action an agent executes is decided by an action decision model. When the performance of the action decision model is insufficient, the agent will be unable to execute appropriate actions and ultimately unable to complete the task. Therefore, how to train a more accurate action decision model to improve the performance of the action decision model is an urgent problem to be solved. Summary of the Invention

[0004] Embodiments of this application provide a method for training an action decision model of an agent, an action decision method, and a device thereof, which can train a more accurate action decision model, improve the performance of the action decision model, and thus decide more accurate actions for the agent. The technical solution is as follows:

[0005] On the one hand, a method for training an action decision model of an agent is provided, and the method includes:

[0006] Based on multiple historical trajectories of an agent, construct a state topology graph, where each historical trajectory includes multiple actions, and each action is used to control the transition between different states. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action;

[0007] Based on the state topology graph, train the action feedback model of the agent, where the action feedback model is used to provide a feedback signal of the environment where the agent resides on the action executed by the agent;

[0008] Based on the state topology graph and the action feedback model, train the action value model of the agent, where the action value model is used to evaluate the value of the action executed by the agent on the environment;

[0009] Based on the action value model, train the action decision model of the agent, where the action decision model is used to decide the action that the agent should execute in a given state.

[0010] On the one hand, an action decision method of an agent is provided, and the method includes:

[0011] When the state at the current moment is observed in the environment, the state is input into the action decision model of the agent, and the action decision model is used to decide the action that the agent should execute in a given state;

[0012] Through the action decision model, determine the execution probability of each of the multiple candidate actions for the agent, and the execution probability represents the possibility that the agent executes the candidate action in the state;

[0013] Based on the execution probabilities of the multiple candidate actions respectively, determine the target action that the agent executes at the current moment from the multiple candidate actions;

[0014] Wherein, the action decision model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment on the action executed by the agent, and the action value model is used to evaluate the value of the action executed by the agent on the environment.

[0015] On the one hand, a training device for an action decision model of an agent is provided, and the device includes:

[0016] A topology graph construction module, configured to construct a state topology graph based on multiple historical trajectories of the agent. Each historical trajectory includes multiple actions, and each action is used to control the transition between different states. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action;

[0017] A feedback training module, configured to train the action feedback model of the agent based on the state topology graph, and the action feedback model is used to provide a feedback signal of the environment where the agent resides on the action executed by the agent;

[0018] An action value training module, configured to train the action value model of the agent based on the state topology graph and the action feedback model, and the action value model is used to evaluate the value of the action executed by the agent on the environment;

[0019] A decision training module, configured to train the action decision model of the agent based on the action value model, and the action decision model is used to decide the action that the agent should execute in a given state.

[0020] In some embodiments, the topology graph construction module is configured to:

[0021] Initialize the state topology graph;

[0022] For any action in any of the historical trajectories, if a directed edge indicating the action is queried in the state topology graph, update the access count associated with the directed edge, where the access count indicates the query frequency of the directed edge;

[0023] If a directed edge indicating the action is not queried in the state topology graph, determine the starting state and the reaching state associated with the action, and based on the starting node indicating the starting state, add a reaching node indicating the reaching state and a directed edge pointing from the starting node to the reaching node.

[0024] In some embodiments, the feedback training module includes:

[0025] A trajectory sampling sub-module, configured to perform trajectory sampling based on the state topology graph to obtain multiple pairs of sampled trajectories, where each pair of the sampled trajectories includes a pair of sampled trajectories with equal lengths;

[0026] A labeling acquisition sub-module, configured to acquire the labeling results of the multiple pairs of sampled trajectories, where the labeling results indicate the satisfaction degrees of different sampled trajectories in each pair of the sampled trajectories with respect to the task performed by the intelligent agent;

[0027] A feedback training sub-module, configured to train the action feedback model based on the state topology graph and the labeling results when the feedback model update condition is satisfied.

[0028] In some embodiments, the trajectory sampling sub-module includes:

[0029] A random sampling unit, configured to randomly sample from the node set of the state topology graph to obtain multiple sampling points;

[0030] A trajectory sampling unit, configured to start trajectory sampling along the directed edge starting from any one of the multiple sampling points, and stop sampling when the trajectory length reaches the sampling length to obtain a sampled trajectory;

[0031] A trajectory pairing unit, configured to pair multiple sampled trajectories according to the trajectory lengths to obtain multiple pairs of sampled trajectories.

[0032] In some embodiments, the trajectory sampling unit is configured to:

[0033] If there is only one directed edge starting from the sampling point, use the reaching node pointed to by the directed edge as the next sampling point;

[0034] If there are at least two directed edges starting from the sampling point, randomly select the directed edge with the highest or lowest empirical action value from the at least two directed edges, and use the arrival node pointed to by the randomly selected directed edge as the next sampling point. The empirical action value indicates the value of the impact that the agent is expected to have on the environment when performing actions according to historical experience.

[0035] In some embodiments, the apparatus further includes:

[0036] A feedback update module, configured to update the estimated feedback value of each directed edge in the state topology graph after the action feedback model is updated. The estimated feedback value indicates the feedback signal that the environment is expected to generate when the agent performs the action indicated by the directed edge.

[0037] In some embodiments, the action value training module includes:

[0038] A function acquisition sub-module, configured to obtain an empirical action value function based on the state topology graph. The empirical action value function is used to provide an empirical action value for evaluating the actions performed by the agent based on the state topology graph;

[0039] An action value training sub-module, configured to train the action value model based on the empirical action value function and the action feedback model.

[0040] In some embodiments, the action value training sub-module includes:

[0041] An action value estimation unit, configured to obtain, in any iteration, the estimated action value of the action indicated by each directed edge in the state topology graph through the action value model. The estimated action value indicates the value of the impact that the action value model predicts the agent will have on the environment when performing the action;

[0042] A constraint loss acquisition unit, configured to obtain a constraint loss term based on the empirical action value function and the estimated action value. The constraint loss term characterizes the distribution difference between the empirical distribution and the model distribution of the action value;

[0043] An action value loss acquisition unit, configured to obtain an action value loss term based on the action feedback model and the estimated action value. The action value loss term characterizes the difference between the estimated action value of the action by the model and the target action value. The target action value characterizes the optimization target of the action value based on the action distribution;

[0044] An action value training unit, configured to iteratively train the action value model based on the constraint loss term and the action value loss term.

[0045] In some embodiments, the constraint loss acquisition unit includes:

[0046] A node determination subunit, configured to determine a plurality of concerned nodes from the state topology graph, where the concerned nodes indicate the states that need to be concerned when the agent implements a task;

[0047] An experience action value determination subunit, configured to, for any one of the concerned nodes, based on the experience action value function, determine the experience action values of each action in the support action set of the concerned node, where the support action set includes the action sets indicated by each directed edge starting from the concerned node, and the experience action value indicates the value of the impact expected to be generated on the environment when the agent executes the action according to historical experience;

[0048] An action value error determination subunit, configured to determine the action value error of the concerned node based on the experience action values and the estimated action values of each action in the support action set, where the action value error characterizes the difference between the experience action value and the estimated action value of each action in the support action set;

[0049] A constraint loss acquisition subunit, configured to obtain the constraint loss term based on the action value errors of each concerned node.

[0050] In some embodiments, the action value error determination subunit is configured to:

[0051] Determine the experience action value vector of the concerned node based on the experience action values of each action in the support action set;

[0052] Determine the estimated action value vector of the concerned node based on the estimated action values of each action in the support action set;

[0053] Determine the action value error based on the experience action value vector and the estimated action value vector.

[0054] In some embodiments, the action value loss acquisition unit is configured to:

[0055] Randomly sample the directed edge set of the state topology graph to obtain a plurality of sampled edges, and for any one of the sampled edges, determine the sampled state indicated by the starting node of the sampled edge and the sampled action indicated by the sampled edge;

[0056] Based on the action feedback model, determine the estimated feedback value of the agent executing the sampled action in the sampled state;

[0057] Based on the action decision model, determine the execution probability of the agent for the sampled action in the sampled state;

[0058] Determine the target action value of the sampling action based on the estimated feedback value, execution probability, and estimated action value of the agent executing the sampling action in the sampling state;

[0059] Obtain the action value loss term based on the target action value and the estimated action value.

[0060] In some embodiments, the decision training module is configured to:

[0061] In any iteration, determine, through the action decision model, a decision vector of the agent in the current state, where the decision vector indicates the possibility of the agent executing various actions at the current moment;

[0062] Based on the action value model, determine a scoring vector of the action value of the agent in the current state, where the scoring vector indicates the value that the agent is expected to bring when executing each action at the current moment;

[0063] Based on the decision vector and the scoring vector, determine the decision loss term of the agent at the current moment, where the decision loss term characterizes the error between the action executed by the agent's decision and the task expectation;

[0064] Iteratively train the action decision model based on the decision loss term.

[0065] In some embodiments, the apparatus further includes an action value update module, and the action value update module includes:

[0066] A node sampling sub-module, configured to sample from the node set of the state topology graph to obtain a plurality of nodes to be updated when the action value update condition is satisfied;

[0067] An action set determination sub-module, configured to determine a support action set for each of the nodes to be updated, where the support action set includes the action sets indicated by each directed edge starting from the node to be updated;

[0068] An action value update sub-module, configured to update the empirical action value of each action in the support action set, where the empirical action value indicates the value of the impact that the agent is expected to have on the environment when executing the action according to historical experience.

[0069] In some embodiments, the action value update sub-module is configured to:

[0070] For any action in the support action set, determine a plurality of target nodes that can be reached by executing the action starting from the node to be updated;

[0071] Determine the empirical transition probability of the node to be updated based on the access times of each directed edge connecting the node to be updated and each target node;

[0072] Determine the estimated feedback value of performing the action starting from the node to be updated through the action feedback model, where the estimated feedback value indicates the feedback signal expected to be generated by the environment when the agent performs the action;

[0073] Obtain the candidate action value of the action based on the empirical transition probability, the estimated feedback value, and the empirical action value of performing the action starting from the node to be updated;

[0074] Assign the maximum value among the candidate action values of each action in the support action set to the empirical action value of performing the action starting from the node to be updated.

[0075] On the one hand, an action decision device for an agent is provided, and the device includes:

[0076] An input module, configured to input the state into the action decision model of the agent when the state at the current moment is observed in the environment, where the action decision model is used to decide the action that the agent should perform in a given state;

[0077] A probability determination module, configured to determine, through the action decision model, the execution probability of each of multiple candidate actions of the agent, where the execution probability represents the possibility that the agent performs the candidate action in the state;

[0078] An action determination module, configured to determine the target action that the agent performs at the current moment from the multiple candidate actions based on the execution probabilities of the multiple candidate actions;

[0079] Wherein, the action decision model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, each directed edge connecting a pair of nodes indicates an action, the action feedback model is used to provide the feedback signal of the environment to the action performed by the agent, and the action value model is used to evaluate the value of the action performed by the agent on the environment.

[0080] On the one hand, a computer device is provided, and the computer device includes one or more processors and one or more memories. At least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the training method of the action decision model of the agent or the action decision method of the agent in any of the above possible implementation manners.

[0081] On the one hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the training method of the action decision model of the agent or the action decision method of the agent in any of the above possible implementation manners.

[0082] On the one hand, a computer program product is provided. The computer program product includes one or more computer programs, and the one or more computer programs are stored in a computer-readable storage medium. One or more processors of a computer device can read the one or more computer programs from the computer-readable storage medium, and the one or more processors execute the one or more computer programs, so that the computer device can execute the training method of the action decision model of the agent or the action decision method of the agent in any of the above possible implementation manners.

[0083] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0084] By constructing a non-parametric state topology graph based on historical trajectories, this state topology graph can fully reflect the empirical distribution of the agent's actions, has a higher information utilization rate for historical trajectory information, and brings more information. Furthermore, based on the state topology graph, the training of the action feedback model is guided, so that the action feedback model can calculate a more accurate estimated feedback value, improving the accuracy of the action feedback model. Then, by combining the state topology graph and the action feedback model, the training process of the action value model can be constrained to obtain an action value model with better accuracy and performance, making the calculation of the estimated action value by the action value model more accurate and able to conform to human intentions or preferences to a certain extent. Finally, an action decision model with better accuracy is trained with the assistance of the action value model with better accuracy, which helps to make accurate decisions on what actions the agent should execute in a given state. Description of the Drawings

[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.

[0086] Figure 1 It is a principle flowchart of a preference-based reinforcement learning algorithm provided by the embodiments of the present application;

[0087] Figure 2It is a schematic diagram of the implementation environment of a training method for an action decision-making model of an agent provided by an embodiment of the present application;

[0088] Figure 3 It is a flowchart of a training method for an action decision-making model of an agent provided by an embodiment of the present application;

[0089] Figure 4 It is a flowchart of a training method for an action decision-making model of an agent provided by an embodiment of the present application;

[0090] Figure 5 It is a training framework diagram of an action decision-making model provided by an embodiment of the present application;

[0091] Figure 6 It is a flowchart of an action decision-making method of an agent provided by an embodiment of the present application;

[0092] Figure 7 It is a performance comparison diagram of an agent in the box-pushing task provided by an embodiment of the present application;

[0093] Figure 8 It is a performance comparison diagram of a robot in the construction task provided by an embodiment of the present application;

[0094] Figure 9 It is a schematic structural diagram of a training device for an action decision-making model of an agent provided by an embodiment of the present application;

[0095] Figure 10 It is a schematic structural diagram of an action decision-making device of an agent provided by an embodiment of the present application;

[0096] Figure 11 It is a schematic structural diagram of a control terminal provided by an embodiment of the present application;

[0097] Figure 12 It is a schematic structural diagram of a training server provided by an embodiment of the present application. Detailed implementation manners

[0098] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0099] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor are the quantity and execution order limited.

[0100] As used in this application, the term "at least one" means one or more, and "a plurality" means two or more. For example, a plurality of historical trajectories means two or more historical trajectories.

[0101] The term "including at least one of A or B" in this application covers the following cases: including only A, including only B, and including both A and B.

[0102] The user-related information (including but not limited to the user's device information, personal information, behavioral information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application, when applied to specific products or technologies using the methods of the embodiments of this application, are all obtained with the user's permission, consent, authorization, or full authorization from all parties. Moreover, the collection, use, and processing of relevant information, data, and signals need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the annotation results for trajectory pairs or segment pairs involved in this application are all obtained under full authorization.

[0103] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0104] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0105] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0106] A pre-trained model (Pre-Training Model, PTM), also known as a foundation model or large model, refers to a deep neural network (Deep Neural Network, DNN) with a large number of parameters. It is trained on a vast amount of unlabeled data, and by leveraging the function approximation ability of the large-parameter DNN, the PTM extracts common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (Parameter-Efficient Fine-Tuning, PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified according to the data modalities they process into: language models, such as the Embeddings from Language Models (ELMo), the Bidirectional Encoder Representations from Transformers (BERT), the Generative Pre-Trained Transformer (GPT), etc.; visual models, such as the Swin-Transformer, the Vision Transformer (ViT), the Vision Mixture of Experts (V-MoE), etc.; speech models, such as the VALL-E model, etc.; multi-modal models, which refer to models that establish feature representations of two or more data modalities, such as the Vision-BERT (ViBERT), the Contrastive Language-Image Pre-Training (CLIP), the visual language model Flamingo, the Generalist-Agent (Gato), etc. Pre-trained models are important tools for outputting Artificial Intelligence Generated Content (AIGC) and can also serve as a general interface connecting multiple specific task models.

[0107] Autonomous driving technology refers to the ability of a vehicle to drive itself without driver operation. It generally includes technologies such as high-precision maps, environmental perception, computer vision, behavior decision-making, path planning, and motion control. Autonomous driving encompasses multiple development paths, including single-vehicle intelligence, vehicle-road cooperation, and networked cloud control. Autonomous driving technology has broad application prospects. Currently, in addition to the fields of logistics, public transportation, taxis, and intelligent transportation, it will further develop in the future.

[0108] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common applications include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content, conversational interactions, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0109] The solution provided in the embodiments of this application relates to the machine learning technology of artificial intelligence, specifically to reinforcement learning (RL), also known as re-inforcement learning, evaluation learning, or enhanced learning. It is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of an intelligent agent achieving maximum reward or specific goals through learning strategies during the interaction with the environment.

[0110] The classic model of reinforcement learning is the standard Markov Decision Process (MDP). Given certain conditions, reinforcement learning can be divided into model-based reinforcement learning (Model-Based RL) and model-free reinforcement learning (Model-Free RL), as well as active reinforcement learning (Active RL) and passive reinforcement learning (Passive RL). Variants of reinforcement learning include inverse reinforcement learning, hierarchical reinforcement learning, and reinforcement learning for partially observable systems. The algorithms used to solve reinforcement learning problems can be divided into two categories: policy search algorithms and value function algorithms.

[0111] Reinforcement learning theory is inspired by behaviorist psychology, focuses on online learning, and attempts to maintain a balance between exploration and exploitation. Different from supervised learning and unsupervised learning, reinforcement learning does not require any pre-given data. Instead, it obtains learning information and updates model parameters by receiving rewards (feedback) from the environment for actions. Reinforcement learning problems have been discussed in fields such as information theory, game theory, and automatic control, and are used to explain equilibrium states under bounded rationality, design recommendation systems, and robot interaction systems. Some complex reinforcement learning algorithms have a certain degree of general intelligence to solve complex problems and can reach the human level in Go and video games.

[0112] In some task scenarios, deep learning models can be used in reinforcement learning to form Deep Reinforcement Learning (DRL). Deep reinforcement learning combines the perception ability of deep learning and the decision-making ability of reinforcement learning to achieve end-to-end learning from perception to action, and can directly control according to the input signal. It is an artificial intelligence method closer to the human thinking mode. Deep reinforcement learning has the potential to enable robots to truly and fully autonomously learn one or more skills.

[0113] Next, the terms or concepts involved in deep reinforcement learning technology will be explained.

[0114] Agent refers to an entity with intelligence, that is, a software or hardware entity that can act autonomously. It is a very important concept in the field of artificial intelligence. Any independent entity that can think and interact with the environment can be abstracted as an agent. In the field of artificial intelligence, an agent has also been translated as proxy, agent, intelligent agent, agent, etc. In other words, an agent refers to a computational entity that resides in an environment, can continuously and autonomously play a role, and has characteristics such as residency, reactivity, sociality, and initiative. It can be either hardware (such as a robot) or software. An agent can interpret data obtained from the environment that reflects events occurring in the environment and perform actions that affect the environment.

[0115] Environment refers to the space or scenario where the agent resides. The environment can be a part of the real world or a part of the virtual world.

[0116] State refers to the state presented by the agent when interacting with the environment at a certain moment. State is a concept related to the time axis. The agent can observe the same or different states at different moments. For example, at time t, the agent observes state s, and at time t + 1, the agent observes another state s'.

[0117] Action refers to the behavior or action that an agent executes and can affect the environment. The agent affects the environment through actions, and an agent can usually execute one or more actions. For example, taking the agent as a robot, the agent can execute various actions such as walking, running, and jumping in the environment.

[0118] Trajectory, that is, the action trajectory of the agent, refers to a series of actions executed by the agent in a continuous time period and arranged in chronological order. Usually, when the agent executes a certain task, it starts interacting with the environment from the action at the initial moment until the task is completed or the number of actions reaches the set length and then stops. The complete action sequence formed is called a trajectory. For example, when multiple robot carts cooperate to transport goods, starting from the starting point until the goods are delivered or the number of walking steps reaches the preset number of steps and then stops. During this time period from starting to stopping, the continuous actions executed by each robot cart constitute an action sequence, which is called a trajectory of this robot cart.

[0119] Segment refers to a part of the continuous actions segmented or intercepted from the agent's trajectory. A segment is a subset of the trajectory. Since the length of the trajectory is usually long, for the convenience of annotation, each trajectory is segmented or intercepted according to the sampling length, so as to obtain several segments with equal lengths (all equal to the sampling length).

[0120] Policy is defined by a policy function and is usually presented as an action decision model, such as a policy neural network or other parametric models, etc., used to make decisions to control the movement of the agent according to the observed state. For example, the policy function π is presented as a probability density function. Given any state s, through the policy function π, the probability that the agent executes any action under the condition of the given state s can be determined. This probability represents the possibility that the agent makes this action.

[0121] Reward: After the agent executes a certain action, the environment can give a reward to the agent. Usually, the above reward is defined by a reward function, and the reward function is usually presented as an action feedback model, such as a reward neural network or other parametric models, etc. The goal of reinforcement learning is to maximize the total reward obtained by the agent for a series of actions executed in a time period.

[0122] State Transition. An agent can transition between different states by performing actions in the environment. The process of transferring from the old state at the previous moment to the new state at the current moment is called state transition. For example, at time t, after the agent observes state s, it performs action a in the environment, causing a transition from state s to another state s' at time t + 1. The above state transition from state s to state s' can be abstracted as a conditional probability density function P. Given the state s and action a at the current moment, the conditional probability density function P can predict the probability of transitioning to another state s' at the next moment.

[0123] Agent-Environment Interaction. After observing a certain state at the current moment, the agent will perform corresponding actions. After the agent takes an action, the environment will be affected by the action and update the state at the next moment, completing the state transition from the current moment to the next moment. At the same time, the environment will also return a feedback signal (or called a reward signal) to the agent.

[0124] Action-Value Function. A function used to evaluate the value brought by an agent performing a certain action at a certain moment. The action-value function Q is related to the policy function π. For the same agent, different policy functions π will result in different action-value functions Q. For example, when the policy function π remains unchanged, the action-value function Q can reflect the quality of the agent performing action a in the current state s at the current moment. The action-value function is usually presented as an action-value model, such as an action-value neural network or other parametric models.

[0125] Robot. It includes all machines that simulate human behavior or thoughts and other organisms (such as robotic dogs, robotic cats, robotic vehicles, etc.). Some computer programs are even called robots (such as chatbots, dialogue robots, etc.). In the embodiments of this application, the robot refers to an artificial machine device that can automatically perform tasks to replace or assist humans. The artificial machine device can be in anthropomorphic form or zoomorphic form and is generally an electromechanical device controlled by a computer program or an electronic circuit. Usually, a robot consists of a vision sensor, a robotic arm, and a main control computer.

[0126] Robotic Arm. It is a complex system with high precision, multiple inputs and outputs, high nonlinearity, and strong coupling widely used in the field of robotics. Due to its unique operational flexibility, in addition to being coupled to the robot body, the robotic arm can also be coupled to any other artificial machine device.

[0127] In recent years, deep reinforcement learning technology has made rapid progress. Using deep reinforcement learning technology, agents can master various complex tasks and skills, covering robot control, video, games, and numerous industrial applications. However, the key to the success of deep reinforcement learning is the need for a reward function carefully designed by a human engineer. In many practical application scenarios of reinforcement learning, formulating an appropriate reward function has always been a challenging topic. The quality of the reward function largely depends on the designer's in-depth understanding of the core logic of the problem and relevant background knowledge. For example, formulating a reward function for text generation tasks is particularly difficult, and the key lies in how to measure the quality of text generation with a single scalar value. Despite the great efforts made by human engineers in reward design, some studies still point out that there are many problems in current algorithms and application scenarios, such as the "Reward Hacking" phenomenon, which means that in order to obtain more rewards, agents may take some behaviors that are beyond the expectations of human engineers and even harmful. In this context, the agent will focus on exploiting the defects of the reward function to obtain the maximum reward and ignore whether its behavior meets the expectations, which may lead to unexpected and potentially risky behaviors.

[0128] In view of this, preference-based reinforcement learning (PbRL) technology has received extensive attention and spawned a series of algorithms. Compared with algorithms that rely on reward functions designed by human engineers, preference-based reinforcement learning technology uses human preferences to learn action feedback models (such as reward neural networks). Specifically, humans can provide preferences for a pair of action trajectories of the agent. For example, present a pair of historical trajectories made by the agent to the technician, and the technician annotates which historical trajectory in this pair conforms to human preferences, thereby implicitly indicating the goal of the behavior or task that the agent needs to learn. Among them, the historical trajectory refers to the action trajectory made by the agent in a past time period. At each historical moment in this past time period, there is a unique and definite state and action. By learning from the human feedback signal (that is, whether the trajectory conforms to human preferences), the agent can complete a specific task or master a certain behavior required by humans.

[0129] In the embodiments of this application, an effective preference-based reinforcement learning algorithm is involved. Without the need for a human engineer to design a reward function, it can learn an action feedback model from human preferences to promote the training of the agent's action decision model. Tests show that this algorithm can train the agent to exhibit novel behaviors and, to a certain extent, mitigate the challenges from the reward hacking phenomenon.

[0130] The following will be combined with Figure 1 , to illustrate the basic framework of the preference-based reinforcement learning algorithm.Figure 1 This is a principle flowchart of a preference-based reinforcement learning algorithm provided by an embodiment of the present application. As Figure 1 shown, π θ represents the action decision model of the agent, and θ is the parameter set of the action decision model. In the case of observing the state s, the action decision model π θ decides the action a that the agent needs to execute. After the agent executes the action a and interacts with the environment, the environment is updated to another state s'; represents the action feedback model, and ψ is the parameter set of the action feedback model. The action feedback model is used to estimate the reward estimate value that the environment should feedback to the agent in the new state s'.

[0131] Define the quadruple as the transfer data, which can indicate the state transition from state s to state s' and its related information. The transfer data is stored in the experience replay buffer (Replay Buffer). In other words, the experience replay buffer is used to store the historical trajectory of the agent.

[0132] Using the historical trajectories in the experience replay buffer, different trajectory pairs can be constructed in ways such as random sampling or non-random sampling. By presenting each pair of trajectories to the technician and asking the technician to select which trajectory in this pair of trajectories is more in line with the preference, and recording the results of the technician's annotation of conforming or non-conforming, one positive sample trajectory and one negative sample trajectory can be generated from each pair of trajectories, thus realizing the agent's query of human preferences for these trajectory pairs.

[0133] Alternatively, in the case where the lengths of each trajectory are generally long, in order to improve the query efficiency, different fragment pairs can also be constructed from the historical trajectories in various ways such as truncation, cutting, or sampling. By presenting each pair of fragments to the technician and asking the technician to select which fragment in this pair of fragments is more in line with the preference, and recording the results of the technician's annotation of conforming or non-conforming, one positive sample fragment and one negative sample fragment can be generated from each pair of fragments, thus realizing the agent's query of human preferences for these fragment pairs. For example, after setting the sampling length, each trajectory is cut into a series of fragments with a length not exceeding the sampling length, and then a pair of fragments with equal length is randomly selected and paired from all the cut fragments, so as to sample a number of fragment pairs. The construction method of the fragment pairs is not limited here.

[0134] To a certain extent, the preference data obtained by the above method (i.e., the annotation results of each pair of trajectories or segments) can reflect human expectations or desires for the behavior of the agent. By using this preference data, it is possible to learn and recover the underlying reward function through supervised learning techniques. This reward function can specify the reward value for choosing a certain action in a given state, thereby providing feedback to the action-value function to obtain a more accurate action-value function. The action-value function can then assist in optimizing the agent's policy function, enabling the policy function to make better and more appropriate decisions regarding the agent's actions. By repeating the above process, the agent can use human preference data to complete the training of the policy function.

[0135] In other words, based on the preference data, a supervised learning technique can be used to train an action-feedback model that estimates the reward value more accurately. The more accurate the reward estimate given by the action-feedback model, the better the performance of the associated action-value model. This enables the action-value model to evaluate the action value more accurately, which in turn guides the agent's action decision model, ultimately optimizing to obtain an action decision model with better performance and more accurate decisions, thus realizing the training of the action decision model based on the preference data.

[0136] Since preference-based reinforcement learning techniques rely on preference data, that is, the annotation results of pairs of trajectories or segments by technical personnel, a large amount of preference data needs to be manually annotated by humans. Therefore, preference-based reinforcement learning techniques require high labor costs in many application scenarios and have low efficiency in using preference data. In addition, when constructing pairs of trajectories or segments, some historical trajectories are usually randomly selected, and then different sampling methods are used to screen and pair them to form a pair of random trajectories to query human preferences. Therefore, only existing historical trajectories can be used for preference queries. Due to the randomness of the construction of trajectory pairs, it is very likely that both trajectories in a pair of trajectories do not conform to human preferences, resulting in low query efficiency.

[0137] In view of this, the embodiments of the present application relate to a method for training an action decision model of an agent, and propose an efficient preference-based reinforcement learning framework that can make full use of the preference data of humans or human experts to accurately estimate the experience of the action value, thereby assisting in realizing the learning of the action-value function, that is, the action-value model. This can act on the action-feedback model and the action decision model, significantly improving the performance of the entire agent's policy training process.

[0138] Specifically, by utilizing the historical trajectories stored in the experience replay buffer, a non-parametric statistical model, i.e., the state topology graph, is constructed. Using the state topology graph, an empirical action-value function can be learned. The empirical action-value function can provide at least two advantages: on the one hand, based on trajectory sampling on the state topology graph, trajectory pairs or segment pairs with more information can be constructed to assist in querying human preferences, improving the construction efficiency of trajectory pairs or segment pairs and also improving the query efficiency of preference data; on the other hand, it can constrain the learning of the action-value model to optimize and obtain a better-performing action-value model, that is, regularize the neural network-based action-value function to make the action-value model more accurate in estimating action values, thereby further accelerating the policy learning process, alleviating the overestimation error and extrapolation error in the action-value function learning process, and also improving the training efficiency of the action decision model and the learning efficiency of the entire policy learning process.

[0139] The embodiments of the present application are applicable to tasks including robot collaboration, robotic arm control, and any tasks involving human scenarios, such as, for example, the scenario of collaborative carts, the scenario of intelligent question answering, the scenario of robot dancing, etc. Taking the scenario of collaborative carts as an example, with the development of modern industry and technology, the demand for multiple robotic carts to collaborate in transporting goods is increasing day by day. To ensure that these robotic carts can cooperate effectively and efficiently, applying the preference-based reinforcement learning framework of the embodiments of the present application can train an action decision model that conforms to human intentions and preferences for the robotic carts.

[0140] The preference-based reinforcement learning framework of the embodiments of the present application can achieve the efficient utilization of human preference data. Taking the scenario of collaborative carts as an example, the robotic carts can more accurately understand human preferences, reducing the number of times of intervention and adjustment by technicians, reducing the development cost, and improving the utilization rate of preference data; in addition, it can achieve the efficient query of human preference data. By using historical data to construct a non-parametric state topology graph, more informative trajectory pairs or segment pairs can be constructed, thereby improving the construction efficiency of trajectory pairs or segment pairs, improving the query efficiency of preference data, and improving the learning efficiency of each model; in addition, the action-value model optimized by the empirical action-value function can more accurately estimate the action value of each action in a specific environment, that is, improving the accuracy of the action-value model, thereby optimizing the accuracy of the action feedback model and the action decision model. Taking the scenario of collaborative carts as an example, when the action decision model is used for multiple robotic carts to work together, it can provide better policy suggestions for each robotic cart and control each robotic cart to make more expected actions.

[0141] Still taking the scenario of multi-robot cooperation as an example, in the task of multi-robot cooperative transportation, through the preference-based reinforcement learning framework of the embodiments of the present application, the robotic carts can more accurately meet the needs and preferences of humans. For example, when multiple robotic carts are required to cooperate to move a cargo with a specific shape and weight, the action decision-making model can provide the best movement, cooperation, and path planning strategies for the robotic carts through the previously collected preference data, ensuring the safe, fast, and efficient movement of the cargo.

[0142] Furthermore, the preference-based reinforcement learning framework of the embodiments of the present application can also be combined with large models to promote each other. In other words, using a large amount of data and refined models can also promote the training of the action decision-making model. For example, the preference-based reinforcement learning framework of the embodiments of the present application can promote each other with large language models (LLMs). The preference-based reinforcement learning technology helps the fine-tuning process in large language models, and using large language models as preference models to annotate trajectories can improve the ability of the agent to solve complex control tasks.

[0143] The following describes the system architecture of the embodiments of the present application.

[0144] Figure 2 It is a schematic diagram of the implementation environment of a training method for an action decision-making model of an agent provided by the embodiments of the present application. Refer to Figure 2 In this implementation environment, it includes: an agent 201, an agent control system 202, an environment 203, and a training server 204.

[0145] The agent 201 refers to any entity with intelligence, that is, a software or hardware entity that can act autonomously. The agent 201 resides in the environment 203, can think independently and continuously play a role autonomously, and can execute one or more actions to interact with the environment 203. The actions executed by the agent 201 may affect the environment 203, and the environment 203 may change its state, that is, complete a state transition, under the influence of the actions of the agent. The agent 201 can also be regarded as a computing entity with characteristics such as residence, reactivity, sociality, and initiative, and can be either hardware (such as robots, robotic arms, robotic carts) or software (such as question-and-answer robots, game AIs, chess-playing AIs).

[0146] The agent control system 202 refers to a system or algorithm for controlling the actions of the agent 201. The agent control system 202 at least includes an action decision-making model, which is used to decide the action that the agent 201 should execute in a given state. Since the agent 201 can execute multiple actions in a given state, through the action decision-making model, it is possible to decide which action can obtain the optimal feedback signal at the current moment and in the given state. Therefore, the accuracy of the action decision-making model determines the accuracy and intelligence of the actions of the agent 201, and even determines whether the agent 201 can complete the given task. In addition to the action decision-making model, the agent control system 202 may also be loaded with other functional modules such as an operating system, a voice interaction module, a graphical interaction module, a path planning module, and an agent navigation module to control the agent 201 to achieve more diverse interaction or display functions. The embodiments of the present application do not specifically limit this.

[0147] In some embodiments, the agent control system 202 is a control module built into the agent 201, that is, the agent 201 and the agent control system 202 are coupled in the same entity. Or, the agent control system 202 and the agent 201 are two independent entities. For example, the agent control system 202 is an independent control terminal, and the control terminal can encapsulate the decided action into a control signal and send it to the agent 201, so as to remotely control the agent 201 to perform this action based on the control signal. At this time, at least a signal receiver needs to be installed on the agent 201. The embodiments of the present application do not specifically limit whether the agent 201 and the agent control system 202 are integrated in the same physical entity.

[0148] The environment 203 refers to the space or scene where the agent 201 resides. The environment 203 can be a part of the real world or a part of the virtual world. For example, if the agent 201 is a robot, the environment 203 can be the activity area of the robot in three-dimensional space when performing tasks. Another example is that if the agent 201 is a game AI, the environment 203 is the virtual game scene where the game AI conducts activities during the game.

[0149] Under the control of the agent control system 202, the agent 201 can interact with the environment 203. For example, after the agent 201 observes a certain state at the current moment, the agent control system 202 calls the action decision-making model to decide the target action to be executed from multiple candidate actions, and controls the agent 201 to execute the target action. After the agent 201 executes the target action, the environment 203 will be affected by the target action and update to the state at the next moment, completing the state transition from the current moment to the next moment. At the same time, the environment 203 will also return a feedback signal (or called a reward signal) to the agent 201.

[0150] The training server 204 refers to a computer device used to train the action decision-making model of an agent. Using the training method of the action decision-making model of the agent according to the embodiments of the present application, under the preference-based reinforcement learning framework, the training server 204 can fully learn and understand human preference data based on the constructed state topology graph, and finally train an action decision-making model with better accuracy and performance under the supervision of the preference data, so that the action decision-making model can make actions that more conform to human intentions or expectations for the agent, improving the satisfaction of humans with the actions of the agent. The training process of the action decision-making model can be completed locally, such as local offline training on the training server 204, or can be completed in the cloud, such as distributed training jointly by multiple servers to improve the training efficiency. The embodiments of the present application do not specifically limit this.

[0151] The agent 201, the agent control system 202, the environment 203, and the training server 204 can be directly or indirectly connected through wired or wireless communication methods. The present application does not limit this here.

[0152] The agent 201 can be an artificial machine device such as a robot, a robotic arm, a robot car, a drone, an autonomous vehicle, etc., or can be an intelligent terminal, such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.

[0153] The agent control system 202 can be a control module integrated inside the agent 201, or can be a control terminal independent of the agent 201. The control terminal includes but is not limited to smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, etc.

[0154] The environment 203 can be a part of the real world, such as the activity area of an artificial machine device, or can be a virtual world, a virtual scene, a virtual environment, a simulation environment, etc. provided by a computer device (such as a game server, a simulation device, an electronic device, etc.). The embodiments of the present application do not specifically limit this.

[0155] The training server 204 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0156] For ease of understanding, the following uses the scenario of multiple robotic carts collaborating in transportation as an example to illustrate the interaction between the agent and the environment. In the scenario of cart collaboration, the agent refers to multiple robotic carts, and the agent control system refers to the action decision-making model of the robotic carts, which can be integrated on the control chip of the robotic cart or can be an independent control terminal. The environment refers to the activity space or area when the robotic cart transports goods.

[0157] Using the training method of the action decision-making model of the agent in the embodiments of the present application, under the preference-based reinforcement learning framework, the training server can train an action decision-making model with better accuracy and performance. The training process can be completed locally, such as local offline training on the training server, or can be completed in the cloud, such as distributed training jointly by multiple servers to improve the training efficiency. The embodiments of the present application do not make specific limitations on this.

[0158] In some embodiments, after training is completed, the training server sends the parameter set of the action decision-making model to each robotic cart. Thus, each robotic cart makes decisions and executes actions under the control of its own built-in action decision-making model. Multiple robotic carts collaborate to move the goods from the starting point to the ending point to complete the goods transportation task.

[0159] In other embodiments, after training is completed, the training server sends the parameter set of the action decision-making model to the master control terminal of these robotic carts. The master control terminal is responsible for scheduling the actions of each robotic cart. Therefore, under the control of the action decision-making model, the master control terminal needs to make decisions on the actions of each robotic cart and send the control signals of each decision-making action to the corresponding robotic cart respectively, to achieve the macro scheduling of multiple robotic carts by the same master control terminal, and control multiple robotic carts to collaborate to move the goods from the starting point to the ending point to complete the goods transportation task. The embodiments of the present application do not make specific limitations on this.

[0160] Next, the basic process of the training method of the action decision-making model of the agent in the embodiments of the present application will be described.

[0161] Figure 3 is a flowchart of a training method of an action decision-making model of an agent provided by an embodiment of the present application. Refer to Figure 3 , this embodiment is executed by a computer device, which can be the training server 204 in the above-mentioned implementation environment, or can be other devices used to train the action decision-making model. This embodiment includes the following steps:

[0162] 301. The computer device constructs a state topology graph based on multiple historical trajectories of the agent. Each of these historical trajectories contains multiple actions, and each action is used to control the transition between different states. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action.

[0163] The agent involved in the embodiments of the present application refers to any entity with intelligence, that is, a software or hardware entity that can act autonomously. The agent resides in the environment, can think independently and continuously play a role autonomously, and can execute one or more actions to interact with the environment. The actions executed by the agent may affect the environment, and the environment may change its state, that is, complete a state transition, due to the actions of the agent. The agent can also be regarded as a computing entity with characteristics such as residency, reactivity, sociality, and initiative. It can be either hardware (such as a robot, a robotic arm, a robotic vehicle) or software (such as a question-and-answer robot, a game AI, a chess-playing AI). The embodiments of the present application do not specifically limit the implementation manner of the agent.

[0164] The historical trajectory involved in the embodiments of the present application refers to a series of actions executed by the agent in a past continuous time period and arranged in chronological order. When the agent executes a certain task in a past continuous time period, it starts interacting with the environment from the action at the initial moment and stops until the task is completed or the number of actions reaches the set length. The formed complete action sequence is called a historical trajectory.

[0165] For example, when multiple robotic vehicles cooperate to transport goods, starting from the starting point until the goods are delivered or the number of walking steps reaches the preset number of steps and then stops. During the time period from the start to the stop, the continuous action sequence formed by the actions executed by each robotic vehicle is called a historical trajectory of this robotic vehicle.

[0166] For another example, when the game AI executes a game level-breaking task, the game AI has been interacting with the game program in the game from the start moment until the number of action steps reaches the preset number of steps and fails to break through the level, or successfully breaks through the level within the preset number of steps. During the time period from the start to the successful or failed level break, the continuous action sequence executed by the game AI is called a historical trajectory of this game AI.

[0167] For any historical trajectory, it contains multiple actions that the agent has executed in the past time period. Each action is used to control the transition from the state at a certain moment to the state at the next moment. Therefore, the historical trajectory can also be called the historical action sequence of the agent.

[0168] The state topology graph involved in the embodiments of the present application refers to a graph (Graph) structure data with states as nodes, which can represent the topology relationship between states and the transitions between different states. The graph structure data is defined by a set of nodes and a set of directed edges. The set of nodes refers to the set composed of all nodes in the state topology graph, and each node indicates a state. The set of directed edges refers to the set composed of all directed edges in the state topology graph, and each directed edge indicates an action.

[0169] Since state transitions have a chronological order on the time axis, the edges connecting two nodes are directional (i.e., they are directed edges), and each directed edge will connect a pair of nodes. Due to the directionality (or say, the pointing property), the pair of nodes connected by a directed edge can be divided into a starting node and an arriving node. The starting node refers to the node from which the directed edge starts, and the arriving node refers to the node that the directed edge finally reaches or points to. From the perspective of states, a directed edge indicates an action, and this action will cause a state transition process between two different states. Distinguishing from the time axis, the state with an earlier timestamp is the starting state, and the state with a later timestamp is the arriving state. Therefore, the starting node of a directed edge indicates the starting state, and the arriving node of a directed edge indicates the arriving state.

[0170] The state topology graph can be stored in a computer device using one or more data structures, such as hash tables, arrays, dictionaries, Key-Value key-value pairs, queues, etc. Here, the data structure of the state topology graph is not specifically limited.

[0171] In some embodiments, for a given type of agent, the computer device collects multiple historical trajectories of the agent. These historical trajectories can belong to the same or different past time periods. For example, for a robotic vehicle, the historical trajectories of different robotic vehicles transporting goods within the past 1 hour can be collected. Another example is that for a game AI, the historical trajectories of the game AI of the same character in multiple past historical matches can be collected.

[0172] When collecting historical trajectories, historical trajectories can be randomly selected from a large number of historical trajectories to ensure the randomness of the historical trajectories; a trajectory length interval can also be preset, and multiple historical trajectories are randomly selected from the set of historical trajectories that meet the trajectory length interval, so that the historical trajectories are neither too long nor too short, thereby eliminating some low-quality trajectory samples through the trajectory length interval; multiple historical trajectories can also be simulated through simulation software, which saves the collection cost and improves the collection efficiency. The embodiments of the present application do not specifically limit the collection method of historical trajectories.

[0173] In some embodiments, after collecting multiple historical trajectories, based on these multiple historical trajectories, all the states observed during the construction phase and all the actions that have been executed can be obtained. As a result, nodes of the state topology graph can be constructed according to the observed states, and directed edges in the state topology graph can be constructed according to the executed actions, ultimately constructing a state topology graph. That is, the node set of the state topology graph reflects all the observed states, and the directed edge set reflects all the executed actions.

[0174] In step 301 above, by constructing a non-parametric state topology graph based on historical trajectories, this state topology graph can fully reflect the empirical distribution of the agent's actions, bringing more information, having a higher information utilization rate for historical trajectories, and also having a high construction efficiency for the state topology graph. In addition, the state topology graph also supports convenient dynamic updates. For example, once a new state is observed, only a new node needs to be added to the node set of the state topology graph. Another example is that once a new action is observed that leads to a state transition that has never occurred, only a new directed edge needs to be added to the directed edge set of the state topology graph.

[0175] 302. The computer device trains the action feedback model of the agent based on this state topology graph, and this action feedback model is used to provide a feedback signal of the environment where the agent resides for the action executed by the agent.

[0176] Since the agent resides in the environment and interacts with the environment through actions, the action feedback model involved in the embodiments of the present application is used to calculate or estimate the feedback signal of the environment for the action executed by the agent. Generally, after the agent executes a certain action in a certain state, a feedback signal will be generated through the action feedback model. This feedback signal can be implemented as an estimated feedback value. That is, the state and the action are input into the action feedback model, and an estimated feedback value is output. The magnitude of the estimated feedback value represents the degree of reward of the environment for the action. The larger the estimated feedback value, the higher the degree of reward. The smaller the estimated feedback value, the lower the degree of reward (or even there may be no response or a penalty). Thus, the action feedback model is also called the reward model, and the feedback signal is also called the reward signal. The action feedback model reflects the reward function of reinforcement learning, and the action feedback model can be a reward neural network or other parametric models.

[0177] In some embodiments, based on the state topology graph constructed in step 301, multiple pairs of sampled trajectories can be obtained by random sampling or non-random sampling from the state topology graph. Each pair of sampled trajectories contains a pair of sampled trajectories with equal lengths. It should be noted that for the convenience of data storage, it can be required that all pairs of sampled trajectories have equal lengths. For example, each sampled trajectory in each pair of sampled trajectories is controlled to have an equal length. For example, the length of all sampled trajectories is 10 to improve the memory access efficiency of the sampled trajectories. Alternatively, it can also be required that only the two sampled trajectories in each pair of sampled trajectories have equal lengths, but it is not required that all pairs of sampled trajectories have equal lengths. For example, the lengths of the two sampled trajectories in a certain pair of sampled trajectories are both 10, but the lengths of the two sampled trajectories in another pair of sampled trajectories are both 15. The embodiments of the present application do not make specific limitations on this. Next, multiple pairs of sampled trajectories can be directly presented to the technician for annotation, and the annotation results of the multiple pairs of sampled trajectories can be collected. These annotation results are used to indicate the satisfaction degree of different sampled trajectories in each pair of sampled trajectories with respect to the task performed by the agent. In other words, the sampled trajectories are presented to the technician in pairs for annotation. The technician needs to annotate which sampled trajectory in each pair of sampled trajectories conforms to the preference (or expectation). In this way, each pair of sampled trajectories will be divided into a positive sample trajectory and a negative sample trajectory by the annotation result. The positive sample trajectory refers to the sampled trajectory with the annotation result conforming to the preference, and the negative sample trajectory refers to the sampled trajectory with the annotation result not conforming to the preference. Therefore, the annotation result reflects the preference degree of humans for the two sampled trajectories in each pair of sampled trajectories, which can also be called the preference data of humans for the sampled trajectories. This method of collecting preference data based on trajectory pairs has a simple process for obtaining annotation results, few annotation times, and high collection efficiency.

[0178] In some other embodiments, based on the state topology graph constructed in step 301, multiple sampling trajectories can be obtained by random sampling or non-random sampling from the state topology graph. At this time, it is not necessary to pair the trajectories according to whether their lengths are equal. Instead, each sampling trajectory is first cut into multiple sampling segments by means of truncation, cutting, etc. Then, from all the sampling segments formed after cutting each sampling trajectory, the segments are paired according to whether their lengths are equal, so as to obtain multiple pairs of sampling segments. Each pair of sampling segments contains a pair of sampling segments with equal lengths. It should be noted that, for the convenience of data storage, it can be required that all paired sampling segments have equal lengths. For example, each sampling trajectory in each pair of sampling segments has an equal length. For example, the length of all sampling segments is 5 to improve the memory access efficiency of the sampling segments; or, it can also be only required that the two sampling segments in each pair of sampling segments have equal lengths, but it is not required that all paired sampling segments have equal lengths. For example, the lengths of the two sampling segments in a certain pair of sampling segments are both 3, while the lengths of the two sampling segments in another pair of sampling segments are both 5. The embodiments of the present application do not make specific limitations on this. Then, the multiple pairs of sampling segments can be directly presented to the technician for annotation, and the annotation results of the multiple pairs of sampling segments are collected. These annotation results are used to indicate the satisfaction degree of different sampling segments in each pair of sampling segments with respect to the task performed by the intelligent agent. In other words, the sampling segments are presented to the technician in pairs for annotation. The technician needs to annotate in each pair of sampling segments which sampling segment meets the preference (or expectation). In this way, each pair of sampling segments will be divided into a positive sample segment and a negative sample segment by the annotation result. The positive sample segment refers to the sampling segment with the annotation result meeting the preference, and the negative sample segment refers to the sampling segment with the annotation result not meeting the preference. Therefore, the annotation result reflects the preference degree of humans for the two sampling segments in each pair of sampling segments, and can also be called the preference data of humans for the sampling segments. This preference data collection method based on segment pairs has a higher data utilization rate for sampling trajectories, can generate richer sample data, and can obtain more informative preference data.

[0179] Further, after collecting the annotation results, that is, the preference data, the preference data can be used to guide the training process of the action feedback model. Since the preference data can reflect the expectation or expectation of humans for the behavior of the intelligent agent, using the preference data as a supervision signal to perform supervised learning on the action feedback model can enable the action feedback model to learn and restore the potential reward function, so that the action feedback model can calculate a more accurate estimated feedback value, thereby improving the accuracy of the action feedback model.

[0180] 303. The computer device trains the action value model of the agent based on the state topology graph and the action feedback model, and the action value model is used to evaluate the value of the impact of the actions executed by the agent on the environment.

[0181] The action value model involved in the embodiments of the present application is used to evaluate the value of the impact of the actions executed by the agent on the environment, and can measure the pros and cons of the impact brought by the actions for task completion. The measurement index of the action value is called the action value, and the action value calculated by using the action value model is actually an estimate or approximation of the true action value by the action value model, so it is called the estimated action value. For example, for a given state and an executed action, the state and the action are input into the action value model, and an estimated action value is output. The magnitude of the estimated action value represents whether executing the corresponding action in the given state helps the agent complete the task. The larger the estimated action value, the higher the value of executing the corresponding action, and the more helpful it is to complete the task. The smaller the estimated action value, the lower the value of executing the corresponding action, and the less helpful it is to complete the task. The action value model reflects the action value function of reinforcement learning, and the action value model can be an action value neural network or other parametric models.

[0182] In some embodiments, based on the state topology graph constructed in step 301 and the action feedback model trained in step 302, since the state topology graph can reflect the empirical distribution of each action in the historical trajectory, and thus can guide the training of the action value model from the perspective of statistical experience, while the action feedback model guides the training of the action value from the perspective of the feedback signal of the environment (if the reward given by the environment is greater, then the corresponding action value should also be higher). Therefore, by combining the state topology graph and the action feedback model, the training process of the action value model can be constrained, so as to obtain an action value model with better accuracy and performance, make the calculation of the estimated action value by the action value model more accurate, and can conform to human intentions or preferences to a certain extent.

[0183] 304. The computer device trains the action decision model of the agent based on the action value model, and the action decision model is used to decide the action that the agent should execute in a given state.

[0184] The action decision-making model involved in the embodiments of the present application is used to decide which action an intelligent agent should execute in a given state, so as to assist the intelligent agent in completing action decision-making, and further control the execution actions of the intelligent agent under the condition of reasonable path planning. For example, given an observed state, the state is input into the action decision-making model, and the execution probabilities of multiple candidate actions are output, so that the target action to be finally executed can be selected from multiple candidate actions. Among them, the value of each execution probability represents the possibility of the intelligent agent executing a candidate action. The larger the execution probability, the greater the possibility that the intelligent agent executes the corresponding candidate action, and the smaller the execution probability, the smaller the possibility that the intelligent agent executes the corresponding candidate action. Optionally, the candidate action with the largest execution probability can be directly selected as the target action, or alternatively, one can be randomly selected from the top N candidate actions with the largest execution probabilities as the target action, or alternatively, sampling can be performed according to the execution probabilities on the probability distribution of the candidate actions, and the target action finally determined to be executed by the intelligent agent can be randomly decided, such that the sampling process follows the probability distribution. The embodiments of the present application do not specifically limit this. The action decision-making model reflects the policy function of reinforcement learning, and the action decision-making model can be a policy neural network or other parametric models.

[0185] In some embodiments, based on the action value model trained in step 303, since the action value model can feedback the estimated action values of each action in a given state, and when the parameters of the action decision-making model are different, different actions may be decided in the same state. Therefore, the estimated action value given by the action value model to the target action decided can reflect the accuracy of the parameters of the action decision-making model itself. Therefore, the action value model can assist in completing the training and optimization of the action decision-making model. And since a more optimal action value model has been obtained in step 303, the finally optimized action decision-making model will also have higher accuracy and better performance.

[0186] In some other embodiments, the action value model and the action decision model are co-trained and optimized, that is, in each iteration of the iterative training, the action value model and the action decision model are updated once according to the state topology graph and the action feedback model. There is no order of priority between the optimization of the action value model and the action decision model. The action value model can be updated first and then the action decision model, or the action decision model can be updated first and then the action value model, or the action value model and the action decision model can be updated simultaneously. Specific limitations are not provided here. The above optimization process is iteratively executed until the action decision model meets the decision optimization stop condition. At this time, the trained action decision model is obtained. The decision optimization stop condition can be that the loss function value tends to converge, or the number of iteration steps reaches the set number of steps, etc. Specific limitations are not provided here for the decision optimization stop condition. In the co-training and optimization process, the action value model and the action decision model are updated once in each iteration, which can enable the action value model and the action decision model to guide each other, thereby further improving the accuracy of the finally optimized action decision model.

[0187] All the above optional technical solutions can be combined arbitrarily to form the optional embodiments of the present disclosure, which will not be elaborated here one by one.

[0188] The method provided in the embodiment of the present application constructs a non-parametric state topology graph based on the historical trajectory. This state topology graph can fully reflect the empirical distribution of the agent's actions, has a higher information utilization rate for the historical trajectory information, and brings more information. Furthermore, based on the state topology graph, the training of the action feedback model is guided, enabling the action feedback model to calculate a more accurate estimated feedback value, improving the accuracy of the action feedback model. Then, by combining the state topology graph and the action feedback model, the training process of the action value model can be constrained to obtain an action value model with better accuracy and performance, making the calculation of the estimated action value by the action value model more accurate, and being able to conform to human intentions or preferences to a certain extent. Finally, the action value model with better accuracy is used to assist in training an action decision model with better accuracy, which helps to make a precise decision on which action the agent should execute in a given state.

[0189] In the previous embodiment, the basic process of the training method of the agent's action decision model was introduced, that is, by constructing a state topology graph with more information, a more accurate action feedback model is trained, and then a more accurate action value model is trained, and finally a more accurate action decision model is trained. In the embodiment of the present application, the detailed process of the training method of the agent's action decision model will be described.

[0190] Figure 4 is a flowchart of a training method for an action decision model of an agent provided in the embodiment of the present application, asFigure 4 As shown, this embodiment is executed by a computer device, which can be the training server 204 in the above-mentioned implementation environment, or other devices for training an action decision-making model. Taking the computer device as the training server as an example, this embodiment includes the following steps:

[0191] 401. The training server constructs a state topology graph based on multiple historical trajectories of the agent. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action.

[0192] In some embodiments, for a given type of agent, the training server first initializes its action decision-making model, and then uses the initialized action decision-making model to control the interaction between the agent and the environment, collects multiple candidate historical trajectories of the agent, and then screens from the multiple candidate historical trajectories to obtain multiple historical trajectories. Each historical trajectory contains multiple actions, and each action is used to control the transition between different states. These historical trajectories can belong to the same or different past time periods. For example, for a robotic vehicle, the historical trajectories of different robotic vehicles transporting goods within the past 1 hour can be collected. Another example is that for a game AI, the historical trajectories of the game AI of the same character in multiple past historical games can be collected.

[0193] When screening historical trajectories from candidate historical trajectories, historical trajectories can be randomly selected from a large number of candidate historical trajectories to ensure the randomness of historical trajectories; or a trajectory length interval can be preset, and multiple historical trajectories are randomly selected from the set of candidate historical trajectories that meet the trajectory length interval, so that the historical trajectories are not too long or too short, thereby eliminating some low-quality candidate historical trajectories through the trajectory length interval.

[0194] In other embodiments, using the initialized action decision-making model, the interaction between the agent and the environment is simulated and calculated in simulation software, so that multiple historical trajectories can be directly simulated and calculated, which saves the acquisition cost of historical trajectories and improves the acquisition efficiency of historical trajectories. The embodiments of the present application do not specifically limit the acquisition method of historical trajectories.

[0195] In an exemplary scenario, in combination with Figure 5 for example, Figure 5 is a training framework diagram of an action decision-making model provided by an embodiment of the present application. As Figure 5 shown, taking the action decision-making model as the policy neural network π φ and the action feedback model as the reward neural network and the action value model as the action value neural network Q θ as an example. At the beginning of the algorithm, the policy neural network π is initializedφ , Reward Neural Network and Action-Value Neural Network Q θ . Next, based on the reward neural network and the action-value neural network Q θ , the policy neural network π φ is used to control the interaction between the agent and the environment, and multiple historical trajectories are obtained through any of the above acquisition methods. In some embodiments, when obtaining historical trajectories, not only a series of actions executed by the agent are obtained, but also a series of states brought about by these dynamics are recorded, and the reward neural network is used to calculate the estimated feedback value for each state transition. All the collected data will be stored in the constructed state topology graph when constructing the state topology graph, which can further enhance the information content contained in the state topology graph.

[0196] In some embodiments, after obtaining multiple historical trajectories, based on these multiple historical trajectories, all the states observed during the construction phase and all the actions that have been executed can be obtained. Thus, the nodes of the state topology graph can be constructed according to the observed states, and the directed edges in the state topology graph can be constructed according to the executed actions, and finally a state topology graph is constructed. That is, the node set of the state topology graph reflects all the observed states, and the directed edge set reflects all the executed actions. Schematically, the above multiple historical trajectories are all stored in the experience replay buffer of the training server, and a dynamic and directed state topology graph is constructed using each historical trajectory in the experience replay buffer.

[0197] Hereinafter, an example of a possible construction method of the state topology graph will be described. This construction method includes steps A1 to A3:

[0198] A1. The training server initializes the state topology graph.

[0199] In some embodiments, during the construction phase of the state topology graph, a state topology graph is first initialized, that is, the node set and the directed edge set of the state topology graph are created. Optionally, both the node set and the directed edge set are initialized to an empty set. For example, the state topology graph G is represented as G=(V, E), and both the node set V and the directed edge set E are initialized to an empty set.

[0200] In an exemplary scenario, the node set V is defined as The directed edge set E is defined as

[0201] On the one hand, since the nodes in the node set V indicate the state s, relevant information about the state s indicated by each node needs to be recorded in the node set. Optionally, the node set V is implemented as an array object, and each node in the array object records at least the node number, the state s indicated by the node, and its empirical action value The empirical action value will be described in detail in step B4 below and will not be elaborated here.

[0202] On the other hand, since the directed edges in the directed edge set E indicate the state transition process from state s to another state s' when the agent executes action a, that is, each directed edge represents the transition from state s to another state s' through action a. Therefore, relevant information about each action a and the state transition process it indicates needs to be recorded in the directed edge set. Optionally, the directed edge set is implemented as a dictionary object, and each directed edge in the dictionary object records at least the action a associated with this directed edge, the estimated feedback value The visit count N(s, a, s') and the empirical action value Both the estimated feedback value and the empirical action value will be described in detail in step B4 below and will not be elaborated here, while the visit count indicates the query frequency of this directed edge during the training phase. It should be noted that the estimated feedback value of each directed edge is provided by the action feedback model. Therefore, whenever the parameter set of the action feedback model is trained and updated once, the estimated feedback values of all the directed edges in the directed edge set will also be updated accordingly. For the detailed process, see step 403 below. In addition, the visit count of each directed edge will be updated dynamically in real time as the query frequency increases during the training phase. For the detailed process, see step A2 below. In addition, the empirical action value of each directed edge will also be refreshed when the action value update condition is met. For the detailed process, see step B4 below.

[0203] In some embodiments, each node in the node set V generates a unique state hash value for the state s indicated by the node through a hash function. Similarly, each directed edge in the directed edge set E also generates a unique action hash value for the action a indicated by the directed edge through a hash function. In this way, a key-value data structure that is convenient for access and query can be generated, where the key name Key can be the action hash value, and the key value Value can be the state hash value of the state reached by the action, which can ensure that the time complexity of the query process is Thus, the query efficiency for the state topology graph is improved.

[0204] It should also be noted that for each node in the node set V, a support action set can also be maintained The supported action set represents all the actions taken by the agent at state s, and its function is to assist in updating the information of the state topology graph itself, which will be described in detail in step B4 below.

[0205] A2. For any action in any one of the historical trajectories, if a directed edge indicating this action is queried in the state topology graph, the access count associated with this directed edge is updated, and this access count indicates the query frequency of this directed edge.

[0206] In some embodiments, for multiple historical trajectories obtained through the above acquisition method, since the construction methods of the nodes and directed edges in the state topology graph for each historical trajectory are similar, only one of the multiple historical trajectories is taken as an example for illustration. Since this historical trajectory contains multiple actions that guide state transitions in chronological order, for any action in this historical trajectory, in the set of directed edges of the state topology graph, it is queried whether there is a directed edge indicating this action. If a directed edge indicating this action is queried, then the operation of not adding new nodes, not adding new directed edges, and only updating the access count associated with the queried directed edge in step A2 will be executed; if no directed edge indicating this action is queried, then the operations of adding new nodes and adding new directed edges involved in the following step A3 will be executed.

[0207] In this step A2, for any action a in the historical trajectory, when a directed edge indicating this action is queried in the set of directed edges, it means that this directed edge itself already exists. Therefore, there is no need to add new nodes or new directed edges, and only the access count N(s, a, s') recorded on the currently queried directed edge needs to be updated. Optionally, the access count N(s, a, s') is incremented by 1. In other words, the value obtained by adding 1 to the original value of the access count N(s, a, s') is assigned to the access count N(s, a, s'), that is: N(s, a, s') ← N(s, a, s') + 1.

[0208] A3. If the training server does not query a directed edge indicating this action in the state topology graph, the starting state and the reaching state associated with this action are determined, and based on the starting node indicating this starting state, a reaching node indicating this reaching state and a directed edge pointing from this starting node to this reaching node are added.

[0209] In this step A3, for any action a in the historical trajectory, when no directed edge indicating this action is found in the set of directed edges, it means that there is no directed edge indicating this action in the state topology graph, indicating that the situation of executing this action from the starting state has not been observed in the experience replay buffer. Therefore, after determining the starting state and the reaching state of this action, the starting node indicating the starting state is found in the node set of the state topology graph. Then, a reaching node indicating the reaching state is newly added to the node set, and moreover, a directed edge pointing from the starting node to the reaching node is newly added to the set of directed edges.

[0210] In an exemplary scenario, a quadruple is defined as the transfer data, which can indicate the state transition from the starting state s to the reaching state s' and its related information. The transfer data is stored in the experience replay buffer (Replay Buffer). For a newly observed transfer data a reaching node indicating the reaching state s' is newly added to the node set, and a directed edge pointing from the starting state s to the reaching state s' is newly added to the set of directed edges. The directed edge is associated with the action a. In addition, for the newly added directed edge, its empirical action value can be initialized and its visit count N(s, a, s') = 1 is initialized.

[0211] In the above steps A1 - A3, by iteratively traversing each action in each historical trajectory, the node set and the set of directed edges of the state topology graph can be continuously filled. After traversing all the actions in all the historical trajectories, a constructed state topology graph will be obtained. In some other embodiments, all the nodes can also be constructed at once according to the complete set of states in the historical trajectory, and then all the directed edges can be constructed according to the complete set of actions in the historical trajectory. The embodiments of the present application do not specifically limit the construction method of the state topology graph.

[0212] In this step 401, based on the historical trajectory, a non - parametric state topology graph is constructed. This state topology graph can fully reflect the empirical distribution of the agent's actions, bringing more information, having a higher information utilization rate for the historical trajectory, and also having a high construction efficiency. In addition, the state topology graph also supports convenient dynamic updates. For example, once a new state is observed, only a new node needs to be added to the node set of the state topology graph. Another example is that once a new action results in an unoccurred state transition, only a new directed edge needs to be added to the set of directed edges of the state topology graph. Or when a recorded state transition appears again, only the visit count recorded on its directed edge needs to be updated. In this way, as the historical trajectory accumulates continuously, the state topology graph will become more and more perfect, realizing the adaptive update of the state topology graph.

[0213] 402. The training server trains the action feedback model of the agent based on the state topology graph, and the action feedback model is used to provide a feedback signal of the environment where the agent resides for the actions executed by the agent.

[0214] In some embodiments, based on the state topology graph constructed in step 401, multiple pairs of sampled trajectories can be derived from the state topology graph, presented to technicians for preference annotation, and the annotation results of each pair of sampled trajectories are obtained. The annotation results of each pair of sampled trajectories are collectively referred to as preference data. In other embodiments, since the annotation process of each pair of sampled trajectories by technicians is relatively time-consuming and laborious, each pair of sampled trajectories can be input into a large model, and the large model outputs the annotation results of each pair of sampled trajectories. The annotation results provided by the large model can also reflect a certain preference tendency to a certain extent. Therefore, the annotation results of each pair of sampled trajectories are also called preference data. The embodiments of the present application do not specifically limit whether the annotation results are annotated by technicians or by the large model.

[0215] Furthermore, after collecting the annotation results, that is, preference data, the preference data can be used to guide the training process of the action feedback model. Since the preference data can reflect the expectations or desires of humans (or the large model) for the behavior of the agent, therefore, using the preference data as a supervision signal, the action feedback model can be iteratively trained in the manner of supervised learning to update the parameter set of the action feedback model and obtain the trained action feedback model. This process is also called the update process of the action feedback model. Since the preference data is introduced as a supervision signal, the action feedback model can learn and recover the potential reward function, enabling the action feedback model to calculate a more accurate estimated feedback value, thereby improving the accuracy of the action feedback model.

[0216] Hereinafter, an example of a possible training method for the action feedback model will be described. The training method includes steps B1 to B4:

[0217] B1. The training server performs trajectory sampling based on the state topology graph to obtain multiple pairs of sampled trajectories, and each pair of the sampled trajectories includes a pair of sampled trajectories with equal lengths.

[0218] In some embodiments, the training server may randomly sample or non - randomly sample from the state topology graph to obtain multiple pairs of sampled trajectories, where each pair of sampled trajectories contains a pair of sampled trajectories with equal lengths. It should be noted that, for the convenience of data storage, it can be required that all pairs of sampled trajectories have equal lengths. For example, each sampled trajectory in each pair of sampled trajectories is controlled to have an equal length. For example, the length of all sampled trajectories is 10 to improve the memory access efficiency of the sampled trajectories; or, it can also be only required that the two sampled trajectories in each pair of sampled trajectories have equal lengths, but it is not required that all pairs of sampled trajectories have equal lengths. For example, the lengths of the two sampled trajectories in one pair of sampled trajectories are both 10, while the lengths of the two sampled trajectories in another pair of sampled trajectories are both 15. The embodiments of the present application do not specifically limit this.

[0219] In some embodiments, the training server sets the sampling length and ensures that the trajectory lengths of all sampled trajectories are equal to the sampling length. Then, it randomly pairs from all the sampled trajectories to form multiple pairs of sampled trajectories. This can increase the randomness of the pairing process, and the trajectory lengths of each pair of sampled trajectories are equal to the sampling length.

[0220] In other embodiments, the training server sets the sampling length and ensures that the trajectory lengths of all sampled trajectories do not exceed the sampling length. Then, it pairs according to the trajectory lengths from all the sampled trajectories, so that the two sampled trajectories included in each pair of successfully paired sampled trajectories have equal trajectory lengths (and must not exceed the sampling length), forming multiple pairs of sampled trajectories. In this way, the trajectory lengths of each pair of sampled trajectories are not necessarily equal, which can assist in collecting preference data under different lengths of sampled trajectories and improve the diversity of each pair of sampled trajectories. The embodiments of the present application do not specifically limit the sampling method and pairing method of the multiple pairs of sampled trajectories.

[0221] Next, an example of a possible trajectory sampling method will be described. This trajectory sampling method includes the following steps B11 - B13:

[0222] B11. The training server randomly samples from the node set of the state topology graph to obtain multiple sampling points.

[0223] In some embodiments, the training server first randomly samples from the node set to obtain multiple sampling points. Each sampling point is a node in the node set and can indicate a state. Each sampling point randomly sampled in this step B11 can be used as the starting point of a sampled trajectory. In other words, a group of sampling points is randomly selected from the node set of the state topology graph as the starting points of a group of sampled trajectories.

[0224] B12. The training server starts trajectory sampling from any one of the multiple sampling points along the directed edge starting from this sampling point, and stops sampling when the trajectory length reaches the sampling length, obtaining a sampling trajectory.

[0225] In some embodiments, for each of the multiple sampling points in step B11, this sampling point can be used as the starting point of a sampling trajectory, and trajectory sampling is gradually performed along the directed edge starting from the starting point. That is, when a directed edge pointing to an arrival node is found along a certain directed edge starting from the starting point, this arrival node is the next sampling point and also the second node of the sampling trajectory. Repeat the above sampling process until no extended directed edge can be found at a certain arrival node, or the trajectory length reaches the sampling length. At this time, stop sampling to obtain a sampling trajectory.

[0226] In some embodiments, if a Key-Value data structure is constructed based on the state topology graph, then for any sampling point in a sampling trajectory (which may be the starting point or a node found after continuing sampling from the starting point), it is possible to determine how many action possibilities there are in total in the state indicated by this sampling point. Then, randomly select an action from all possible actions, and use the action hash value of the selected action as an index to query whether the constructed Key-Value data structure is hit. If the Key-Value data structure can be hit, the state hash value stored in the Value is taken out, so that the node number of the next sampling point in the node set can be retrieved based on the state hash value. The Key-Value data structure can ensure that the time complexity of the query process is Thus, the query efficiency for the state topology graph is improved. Here, only the Key-Value data structure is used as an example for illustration. The training server can also improve the query efficiency through other data structures such as linked lists, dynamic arrays, and adjacency matrices based on the state topology graph. The embodiments of this application do not specifically limit this.

[0227] In some embodiments, for any sampling point in a sampling trajectory, since there can be more than one directed edge starting from this sampling point, it can be classified and discussed in the following two cases:

[0228] Case 1: If there is only one directed edge starting from this sampling point, the training server uses the arrival node pointed to by this directed edge as the next sampling point.

[0229] For any sampling point, if there is only one directed edge starting from the sampling point, there is no need to select the directed edge. The arrival node pointed to by this directed edge is directly used as the next sampling point. Then continue to determine whether there is only one directed edge starting from the next sampling point. Repeat the above process until no extended directed edge is found on a certain arrival node, or the trajectory length reaches the sampling length, the trajectory sampling is completed, and a sampling trajectory is obtained.

[0230] Case 2: If there are at least two directed edges starting from the sampling point, the training server randomly selects the directed edge with the highest or lowest experience action value from the at least two directed edges, and uses the arrival node pointed to by the randomly selected directed edge as the next sampling point. The experience action value indicates the value of the impact that the agent is expected to have on the environment when performing an action according to historical experience.

[0231] For any sampling point, if there are at least two directed edges from the sampling point, this involves the decision process of which directed edge to follow to determine the next sampling point. And because in the state topology graph, an empirical action value is also recorded for each directed edge Therefore, the empirical action value can be determined from the at least two directed edges. The first directed edge with the highest value and the second directed edge with the lowest empirical action value, a directed edge is randomly selected from the first directed edge and the second directed edge. After the directed edge for this trajectory sampling is selected from the at least two directed edges, the arrival node pointed to by the randomly selected directed edge is used as the next sampling point, and then it is determined whether there is only one directed edge starting from the next sampling point. The above process is repeated until no extended directed edge is found on a certain arrival node, or the trajectory length reaches the sampling length, the trajectory sampling is completed, and a sampled trajectory is obtained.

[0232] Here, the second case is based only on the empirical action value Taking the highest or lowest directed edge as an example, the method of implementing trajectory sampling is explained. This is to sample both actions with higher empirical action values and actions with lower empirical action values, thereby enriching the amount of information in the sampled trajectory and ensuring the diversity of the sampled trajectory. This only makes it possible for different sampled trajectories to have greater differences or distinctions due to the randomness of sampling, which facilitates the acquisition of labeling results that are easy to distinguish and learn, and avoids occasionally sampling two sampling trajectories with poor sample quality. The labeling results of such a paired pair of sampling trajectories are not very meaningful, because it is not helpful to learn which type of sampling trajectory is more satisfactory (both sampling trajectories are unsatisfactory).

[0233] In some other embodiments, the empirical action value may be selected from the at least two directed edges each time. The directed edge with the highest value, or alternatively, the empirical action value can be selected from the at least two directed edges each time. The directed edge with the lowest value, or alternatively, it is not filtered according to the empirical action value. Instead, a directed edge is randomly selected directly from the at least two directed edges. In this way, trajectory sampling can also be completed based on the state topology graph, and sampling randomness can be ensured. The embodiments of the present application do not specifically limit this.

[0234] In still other embodiments, forking can also start from this sampling point. One sampling trajectory continues to perform trajectory sampling along the directed edge with the highest empirical action value. Another sampling trajectory continues to perform trajectory sampling along the directed edge with the lowest empirical action value. In this way, a pair of sampling trajectories that can be successfully paired can be quickly obtained. A part of this pair of sampling trajectories overlaps, but two different sampling segments are generated from this sampling point onwards, thereby providing a pair of sampling trajectories with very high contrast and greater information volume for the subsequent supervised learning of the action feedback model, which helps to improve the accuracy of the learned action feedback model from the perspective of improving the sampling effect of the sampling trajectory.

[0235] B13. The training server pairs multiple sampling trajectories according to the trajectory length to obtain multiple pairs of sampling trajectories.

[0236] In some embodiments, for each sampling point in step B11, step B12 is executed to output a sampling trajectory starting from the sampling point. By traversing all the sampling points in step B11, multiple sampling trajectories can be obtained, but the lengths of these sampling trajectories may be different. On the one hand, different sampling lengths can be adopted during each trajectory sampling, resulting in different trajectory lengths even if the trajectory length is equal to the sampling length because of the different sampling lengths. On the other hand, if the sampling length has not been reached during sampling but there is no directed edge for extension at a certain reach node, sampling will stop, and at this time the trajectory length is less than the sampling length. Therefore, even if the sampling lengths are the same, the trajectory lengths are different. In view of this, the multiple sampling trajectories finally obtained may have the same or different trajectory lengths, but for paired sampling trajectories, it is necessary to ensure that the lengths of the two sampling trajectories are equal. Therefore, pairing is performed according to the trajectory lengths of each sampling trajectory, so that a pair of sampling trajectories with equal lengths is successfully paired to form a pair of sampling trajectories. In this way, the lengths of each pair of sampling trajectories are not necessarily equal, which can assist in collecting preference data under different lengths of sampling trajectories and improve the diversity of each pair of sampling trajectories. The embodiments of the present application do not specifically limit the sampling method and pairing method for multiple pairs of sampling trajectories. Or, for the sake of simplicity, the training server can also preset a trajectory length in advance, discard the sampling trajectories that do not meet the trajectory length, and then randomly pair all the remaining sampling trajectories with equal lengths, which improves the construction efficiency of the sampling trajectories.

[0237] In the above steps B11 to B13, a sampling method for trajectory pairs is provided, so that preference data is collected and annotation results are obtained in units of trajectory pairs, which has high data collection efficiency.

[0238] In some other embodiments, after traversing all the sampling points in step B11 and obtaining multiple sampling trajectories through step B12, instead of pairing at the trajectory level, each sampling trajectory is cut into multiple sampling segments by means of interception, cutting, etc. Then, from all the sampling segments formed after cutting each sampling trajectory, segment pairing is performed according to whether the lengths are equal, so as to obtain multiple pairs of sampling segments. Each pair of sampling segments contains a pair of sampling segments with equal lengths. It should be noted that for the convenience of data storage, it can be required that all paired sampling segments have equal lengths. For example, each sampling trajectory in each pair of sampling segments is controlled to have an equal length. For example, the length of all sampling segments is 5 to improve the memory access efficiency of the sampling segments; or, it can also be only required that the lengths of the two sampling segments in each pair of sampling segments are equal, but it is not required that all paired sampling segments have equal lengths. For example, the lengths of the two sampling segments in a certain pair of sampling segments are both 3, but the lengths of the two sampling segments in another pair of sampling segments are both 5. The embodiments of the present application do not specifically limit this.

[0239] In the above optional manner, a sampling method for fragment pairs is provided, so that preference data is collected and annotation results are obtained in units of fragment pairs. The data utilization rate of the sampling trajectory is higher, and more abundant sample data can be generated.

[0240] B2. The training server collects the annotation results of the multiple pairs of sampling trajectories, and the annotation results indicate the satisfaction degree of different sampling trajectories in each pair of sampling trajectories with respect to the task performed by the intelligent agent.

[0241] In some embodiments, after obtaining multiple pairs of sampling trajectories in step B1, the training server directly presents the multiple pairs of sampling trajectories to the technician for annotation, and collects the annotation results of the multiple pairs of sampling trajectories. These annotation results are used to indicate the satisfaction degree of different sampling trajectories in each pair of sampling trajectories with respect to the task performed by the intelligent agent. In other words, the sampling trajectories will be presented to the technician in pairs for annotation, and the technician needs to annotate whether each sampling trajectory in each pair of sampling trajectories meets the preference (or expectation). In this way, each pair of sampling trajectories will be divided into a positive sample trajectory and a negative sample trajectory by the annotation results. The positive sample trajectory refers to the sampling trajectory with the annotation result meeting the preference, and the negative sample trajectory refers to the sampling trajectory with the annotation result not meeting the preference. Therefore, the annotation results reflect the preference degree of humans for the two sampling trajectories in each pair of sampling trajectories. Optionally, each pair of sampling trajectories can also be directly input into the large model, and the large model is used for annotation to output the annotation results of each pair of sampling trajectories. The embodiments of the present application do not specifically limit this. This preference data collection method based on trajectory pairs has a simple acquisition process for annotation results, few annotation times, and high acquisition efficiency.

[0242] In some other embodiments, if multiple pairs of sampled segments are obtained through the optional method introduced in step B13, the multiple pairs of sampled segments can be directly presented to the technician for annotation, and the annotation results of the multiple pairs of sampled segments are collected. These annotation results are used to indicate the satisfaction degree of different sampled segments in each pair of sampled segments with respect to the task performed by the agent. In other words, the sampled segments are presented to the technician in pairs for annotation, and the technician needs to annotate in each pair of sampled segments which sampled segment meets the preference (or expectation). In this way, each pair of sampled segments will be divided into a positive sample segment and a negative sample segment by the annotation result. The positive sample segment refers to the sampled segment with the annotation result meeting the preference, and the negative sample segment refers to the sampled segment with the annotation result not meeting the preference. Therefore, the annotation result reflects the preference degree of humans for the two sampled segments in each pair of sampled segments. Optionally, each pair of sampled segments can also be directly input into the large model, and the large model is used to perform the annotation and output the annotation result of each pair of sampled segments. The embodiments of the present application do not specifically limit this. This method of collecting preference data based on segment pairs has a higher data utilization rate for the sampled trajectory, can generate richer sample data, and can obtain more informative preference data.

[0243] B3. When the feedback model update condition is satisfied, the training server trains the action feedback model based on the state topology graph and the annotation result.

[0244] In some embodiments, since the annotation of the sampled trajectory by the technician may be a relatively continuous and long-term process, a feedback model update condition can be preset. In this way, when the feedback model update condition is satisfied, the action feedback model is updated based on the preference data collected during the period from the start after the last update to the current moment. For example, the feedback model update condition can be regular updates, such as updating every two days, every week, or updating once every 500 pairs of annotation results of the sampled trajectory, etc. The specific content of the feedback model update condition is not specifically limited here.

[0245] When the feedback model update condition is satisfied, the training server starts the update process of the action feedback model. Using the preference data as the supervision signal, the action feedback model can be iteratively trained in the way of supervised learning to update the parameter set of the action feedback model and obtain the trained action feedback model.

[0246] In some embodiments, in the preference data collection method based on trajectory pairs, when using preference data as a supervision signal to implement supervised learning, for each pair of sampled trajectories, the annotation result indicates that this pair of sampled trajectories contains one positive sample trajectory and one negative sample trajectory. Using the action feedback model, for the positive sample trajectory in the annotation result, calculate the estimated feedback value of each action and sum them up to obtain the positive sample feedback sum value. Similarly, for the negative sample trajectory in the annotation result, calculate the estimated feedback value of each action and sum them up to obtain the negative sample feedback sum value. Furthermore, introduce a preference loss term into the loss function of the action feedback model to control the positive sample feedback sum value of each positive sample trajectory to be as high as possible, and control the negative sample feedback sum value of each negative sample trajectory to be as low as possible. Thus, on the premise of using preference data as a supervision signal, the parameter set of the action feedback model can be optimized through the preference loss term, and the parameter set of the action feedback model is iteratively trained until the loss function value converges, or when the number of iteration steps reaches the set number of steps, stop training and complete an update of the parameter set of the action feedback model.

[0247] In other embodiments, in the preference data collection method based on segment pairs, when using preference data as a supervision signal to implement supervised learning, for each pair of sampled segments, the annotation result indicates that this pair of sampled segments contains one positive sample segment and one negative sample segment. Using the action feedback model, for the positive sample segment in the annotation result, calculate the estimated feedback value of each action and sum them up to obtain the positive sample feedback sum value. Similarly, for the negative sample segment in the annotation result, calculate the estimated feedback value of each action and sum them up to obtain the negative sample feedback sum value. Furthermore, introduce a preference loss term into the loss function of the action feedback model to control the positive sample feedback sum value of each positive sample segment to be as high as possible, and control the negative sample feedback sum value of each negative sample segment to be as low as possible. Thus, on the premise of using preference data as a supervision signal, the parameter set of the action feedback model can be optimized through the preference loss term, and the parameter set of the action feedback model is iteratively trained until the loss function value converges, or when the number of iteration steps reaches the set number of steps, stop training and complete an update of the parameter set of the action feedback model.

[0248] In steps B1 - B3, after collecting the annotation result, i.e., the preference data, by using the preference data to guide the training process of the action feedback model, since the preference data can reflect the expectations or desires of humans (or large models) for the behavior of the intelligent agent, this process is also called the update process of the action feedback model. Due to the introduction of preference data as a supervision signal, the action feedback model can learn and recover the potential reward function, enabling the action feedback model to calculate more accurate estimated feedback values, thereby improving the accuracy of the action feedback model.

[0249] Furthermore, since the sampling trajectory is collected by means of trajectory sampling, there may be samples in the sampling trajectory that have not appeared in the historical trajectory. This is because the state topology graph can connect the state transitions of each historical trajectory in the past and connect the same states involved in different historical trajectories. Therefore, new sampling trajectories can be spliced during trajectory sampling. These sampling trajectories are not directly obtained from the historical trajectory, but are the splicing of segments in one or more historical trajectories. Therefore, the segments are real, but the whole trajectory is a new spliced trajectory, which can greatly enrich the data diversity of the sampling trajectory.

[0250] In some other embodiments, the idea of positive and negative sample contrast learning can also be introduced to combine supervised learning and contrast learning to implement iterative training of the action feedback model. The embodiments of the present application do not specifically limit the training method of the action feedback model.

[0251] Steps B1 to B3 introduce how to update the action feedback model using the state topology graph. Under the preference-based reinforcement learning framework of the embodiments of the present application, when the action value update condition is met, the empirical action value of the action indicated by some or all of the directed edges in the state topology graph can also be updated. This process is called the graph update process. Since the computational cost of calculating the empirical action value for the entire graph update each time is relatively high, in step B4, the example of updating the empirical action value of the action indicated by some directed edges each time will be used for illustration.

[0252] B4. When the action value update condition is met, the training server updates the empirical action value of the action indicated by some directed edges in the state topology graph. The empirical action value indicates the value of the impact expected on the environment when the agent executes the action according to historical experience.

[0253] In some embodiments, since even if only the empirical action value of the action indicated by some directed edges is updated each time, there is still a certain computational overhead in the update process of the empirical action value. Therefore, an action value update condition can be preset. In this way, when the action value update condition is met, a part of the nodes are first selected as the nodes to be updated. Then, since it was introduced in step A1 that each node maintains a support action set Then, according to the preset action value update rule, the empirical action values of all the actions included in the support action sets of all the nodes to be updated can be updated, which can save the computational overhead of each graph update process and improve the update efficiency of the state topology graph. Optionally, the action value update condition can be periodic update, such as updating once every two days, once a week, or updating once every 200 transfer data are collected in the experience replay buffer, etc. The specific content of the action value update condition is not specifically limited here.

[0254] It should be noted that both the action feedback model and the empirical action value can be updated regularly or triggered by conditions. However, if both are updated regularly, it is not required that their update cycles be the same. They can be updated independently according to their own cycles. For example, the action feedback model is updated every two days, while the empirical action value is updated every day. The embodiments of the present application do not make specific limitations on this.

[0255] When the action value update condition is met, the training server starts the graph update process of the state topology graph, that is, for the support action set of each node to be updated Update the support action set Update the empirical action value of each action in the support action set. Usually, according to the preset action value update rule, it is recalculated and reassigned on the basis of the original value to correct the possible deviation of the original value.

[0256] Next, an example of a possible action value update rule will be described. The action value update rule includes the following steps B41 to B43:

[0257] B41. When the action value update condition is met, the training server samples from the node set of the state topology graph to obtain multiple nodes to be updated.

[0258] In some embodiments, when the action value update condition is met, samples are taken from the node set of the state topology graph to obtain multiple nodes to be updated. Optionally, multiple nodes to be updated are directly randomly sampled from the node set, or probability sampling is performed on the node set according to the latest access timestamp of the nodes to obtain multiple nodes to be updated. Here, probability sampling means that nodes with the largest latest access timestamp (i.e., the closest to the current moment) have a greater probability of being sampled as nodes to be updated. Therefore, these nodes may be nodes that have just been accessed recently and are also relatively important nodes at the current stage of the training process. Therefore, it is necessary to update the empirical action values of each action in their support action sets. In this way, even if only a subset of the node set is sampled each time, a good graph update effect can be ensured.

[0259] In an exemplary scenario, probability sampling is performed on the node set according to the latest access timestamp of the nodes to obtain multiple nodes to be updated, and the multiple nodes to be updated can form a subset of the node set Then, for the subset The multiple nodes to be updated therein are sampled in reverse order. Sampling in reverse order means that starting from the node to be updated with the largest latest access timestamp, step B42 is executed until the node to be updated with the smallest latest access timestamp is sampled in reverse order, and finally the graph update process for all nodes to be updated is completed. Sampling in reverse order can ensure that the nodes to be updated that are accessed more recently are updated with the empirical action values earlier, which can improve the calculation efficiency and update efficiency of the empirical action values. Optionally, for a subset the multiple nodes to be updated therein can be sampled in order, or for a subset the multiple nodes to be updated therein can be sampled in a random order, which is not specifically limited in the embodiments of the present application.

[0260] It should be noted that since the action value update rules for different nodes to be updated are the same, an example of the action value update process for one node to be updated will be given below in combination with steps B42 to B43.

[0261] B42. The training server determines the support action set for each of the nodes to be updated, and the support action set includes the action sets indicated by the directed edges starting from the node to be updated.

[0262] For the state topology graph, each node to be updated will have a support action set In the action value update rule for each node to be updated, it can be set that the maximum operator for graph update is to operate on the support action set i.e., update all the actions in the support action set instead of updating all the actions of each directed edge in the entire directed edge set, which can further reduce the computational overhead in the update process of each node to be updated and improve the action value update efficiency of a single node.

[0263] In some embodiments, for each node to be updated, all the directed edges starting from the node to be updated are determined, and the union of all the actions indicated by these directed edges is the support action set of the node to be updated Since the update of the empirical action value is completed on the support action set no action values of actions that have not been seen will be considered during the update process, because there are no corresponding directed edges for actions that have not been seen and they will not be collected into the state topology graph, which can avoid overestimating the empirical action value, that is, try to ensure that the empirical action value is not overestimated or underestimated too much, making the estimation of the empirical action value more accurate.

[0264] B43. The training server updates the empirical action value of each action in the support action set.

[0265] In some embodiments, for each node to be updated, the support action set of the node to be updated is determined through step B42 After that, for the support action set For each action a in the set, based on the original value of the empirical action value of the action a, according to a preset action value update rule, recalculation and assignment are performed on the basis of the original value to correct the possible deviation of the original value

[0266] In some embodiments, the action value update rule can be set for the iterative update process of all actions in the support action set So as to quickly realize the traversal update of the empirical action values of each action. Schematically, any iteration in the above iterative update process is described. Each iteration updates the empirical action value of one action in the support action set Combined with the following steps B43a to B42e for description

[0267] B43a. For any action in the support action set, the training server determines a plurality of target nodes that can be reached by executing the action starting from the node to be updated

[0268] In some embodiments, for each node s to be updated, the support action set of the node s to be updated is determined through step B42 After that, for the support action set For each action in the set Since starting from the node s to be updated, executing the action a may reach multiple different target nodes s', therefore, the training server determines a plurality of target nodes s'∈S that can be reached by executing the action a starting from the node s to be updated. S refers to the state set composed of all states indicated by all nodes included in the node set V

[0269] B43b. The training server determines the empirical transition probability of the node to be updated based on the access times of each directed edge connecting the node to be updated and each target node

[0270] In some embodiments, for each target node s' found in step B43a, a directed edge starting from the node s to be updated and reaching the target node s' can be uniquely determined from the directed edge set, and then the access times N(s, a, s') of this directed edge are queried. Repeating the above operation of obtaining access times for all target nodes s'∈S can query the access times N(s, a, s') of all directed edges starting from the node to be updated. Summing up the queried all access times N(s, a, s'), the entire support action set is obtained The total number of visits, and then, the value obtained by dividing the number of visits of the current directed edge by the total number of visits is used as the empirical transition probability from the node s to be updated to the target node s'.

[0271] For example, using to represent the empirical transition probability that the node s to be updated reaches the target node s' through the action a, then the empirical transition probability is expressed by the following formula:

[0272]

[0273] where N(s, a, s') represents the number of visits of the directed edge from the node s to be updated to the target node s' through the action a, represents the support action set the total number of visits.

[0274] B43c. The training server determines, through the action feedback model, the estimated feedback value for executing the action starting from the node to be updated, and the estimated feedback value indicates the feedback signal expected to be generated by the environment when the agent executes the action.

[0275] In some embodiments, for the node s to be updated and each action a in the support action set the node s to be updated and the action a are input into the action feedback model, and the estimated feedback value for executing the action a starting from the node s to be updated is calculated through the action feedback model For example, the action feedback model refers to a reward neural network The node s to be updated and the action a are input into the reward neural network and the estimated feedback value of the node s to be updated relative to the action a is output For each action a in the action set an estimated feedback value can be calculated through this step B43c.

[0276] B43d. The training server obtains the candidate action value of the action based on the empirical transition probability, the estimated feedback value, and the empirical action value for executing the action starting from the node to be updated.

[0277] In some embodiments, based on the empirical transition probability calculated for each target node s' in step B43b, and based on the estimated feedback value calculated for each action a in step B43c and the original value of the empirical action value recorded for the node s to be updated and the action a itself a candidate action value can be calculated for the node s to be updated and the action a.

[0278] In some embodiments, for the node s to be updated and the action a, the calculation method of the candidate action value of the action a is as shown in the following formula:

[0279]

[0280] Wherein, represents the estimated feedback value calculated for the node s to be updated and the action a, and γ is a hyperparameter representing the weighting intensity. represents the empirical transition probability that the node s to be updated reaches the target node s' through the action a. represents the original value of the empirical action value recorded for the node s to be updated and the action a in the state topology graph itself.

[0281] B43e. The training server assigns the maximum value among the candidate action values of each action in the support action set to the empirical action value of executing this action starting from the node s to be updated.

[0282] In some embodiments, when the node s to be updated remains unchanged, for the support action set of the node s to be updated, for each action a, the candidate action value of executing the action a on the node s to be updated can be calculated through steps B43a - B43d. By traversing all the actions in the support action set all the candidate action values of all the actions can be obtained. Then, the maximum value among the candidate action values is used as the new value Q(s, a) and assigned to the empirical action value recorded in the state topology graph for the node s to be updated and the action a. In other words, the empirical action value recorded on the directed edge starting from the node s to be updated and indicating the action a is replaced from the original value with the new value Q(s, a), thus realizing the update of the action value. By iteratively executing steps B43a - B43e for all the nodes to be updated, an action value iterative update can be completed for all the sampled nodes to be updated.

[0283] In some embodiments, for the node s to be updated and the action a, the assignment method of the new value Q(s, a) of the empirical action value is as shown in the following formula:

[0284]

[0285] In the above steps B43a - B43e, for each action in the support action set by introducing the empirical transition probability based on the access times, the original value of the empirical action value can be weighted using the empirical transition probability, and then combined with the information influence of the reward signal dimension introduced by the estimated feedback value provided by the action feedback model to calculate the candidate action value (i.e., the selectable new value) of each action. Finally, for the support action set Assign the maximum value among all newly calculated candidate action values to the empirical action value of this action, thereby updating the empirical action value of this action. For the empirical action values of all actions in the support action set of each node to be updated iteratively execute the above update process, which can quickly re-estimate the empirical action value for each node to be updated, greatly improving the efficiency of action value update.

[0286] In some other embodiments, the state topology graph can also be configured as a graph neural network, such that the graph neural network learns appropriate values of its empirical action values under the guidance of the estimated feedback value and the access times. The embodiments of the present application do not specifically limit the action value update method.

[0287] In steps B41 - B43, a possible action value update rule for the empirical action value of each action in the support action set of the node to be updated is provided, which can quickly refresh and estimate the empirical action value with less computational overhead.

[0288] In some other embodiments, a full graph update can also be performed each time the action value is updated. In this way, the empirical action values of each action in the state topology graph can be updated in a timely and sufficient manner, and the empirical distribution of the latest action space can be restored, rather than being limited to the support action sets of some nodes. The embodiments of the present application do not specifically limit whether to perform a full graph update on the empirical action values.

[0289] It should be noted that the action value update process involved in step B4 is an optional step, and the training server may not update the empirical action value. The embodiments of the present application do not specifically limit this.

[0290] 403. After the action feedback model is updated, the training server updates the estimated feedback value of each directed edge in the state topology graph. The estimated feedback value indicates the feedback signal that the environment is expected to generate when the intelligent agent executes the action indicated by the directed edge.

[0291] In some embodiments, as introduced in step 402, the action feedback model may be updated each time the feedback model update condition is met. After each update of the action feedback model, the training server can use the updated action feedback model to, conversely, re-label the estimated feedback values recorded on each directed edge in the state topology graph. In other words, using the updated action feedback model, the estimated feedback values of each directed edge in the state topology graph are recalculated and overwritten and refreshed. Therefore, the performance of the action feedback model is improved each time it is updated. Through the updated action feedback model, the estimated feedback value of each directed edge can be recalculated to achieve re-labeling of the estimated feedback value.

[0292] For example, assume that the action feedback model is a reward neural network Every time the reward neural network After the update is completed, for any node s and action a in the state topology graph, a directed edge can be determined, and the node s and action a are input into the updated reward neural network to output the estimated feedback value of node s relative to action a The new value is then used to overwrite and store the old value in the directed edge, replacing the old value

[0293] In the above process, after each update of the action feedback model, the estimated feedback values maintained on each directed edge in the state topology graph are re-labeled using the action feedback model. This maximizes the utilization of the latest action feedback model, ensuring that the estimated feedback values are the best-performing and most accurate new values at the current moment, and can mitigate the impact of the non-stationary reward function. This is because during the regular update of the action feedback model, the underlying reward function it learns is constantly fine-tuning and optimizing. Therefore, the reward function is not always stationary. By promptly overwriting the old values with new values, it can be ensured that the estimated feedback values stored in the state topology graph can reflect the reward signals evaluated by the latest action feedback model, thereby enhancing the accuracy of the information contained in the state topology graph

[0294] 404. The training server obtains an empirical action value function based on the state topology graph. This empirical action value function is used to provide empirical action values for evaluating the actions performed by the agent based on the state topology graph

[0295] In some embodiments, since both the action values and estimated feedback values in the state topology graph may be updated, when obtaining the empirical action value function, the latest state topology graph needs to be used for guiding the calculation. Among them, the empirical action value function has been introduced in step B43e. Only the old values of the empirical action values for each action need to be given during the initialization phase, and subsequently, the empirical action values for each node relative to each action can be continuously updated through the methods introduced in steps B43a - B43e, so that the empirical action values can reflect the empirical distribution of the action values on the historical trajectory after multiple updates. Here, the empirical action value function will not be elaborated further. Please refer to step B43e

[0296] 405. The training server trains the action value model of the agent based on the empirical action value function and the action feedback model. This action value model is used to evaluate the value of the actions performed by the agent on the environment

[0297] In some embodiments, based on the empirical action-value function in step 404 and the action feedback model trained in step 402, the action-value model of the agent can be guided for training. This action-value model is used to evaluate the value of the impact of the actions executed by the agent on the environment. Since the empirical action-value function can reflect the empirical distribution of each action in the historical trajectory, it can guide the training of the action-value model from the perspective of statistical experience. The action feedback model, on the other hand, guides the training of the action value from the perspective of the feedback signal of the environment (if the reward given by the environment is greater, then the corresponding action value should also be higher). Therefore, by combining the empirical action-value function and the action feedback model, the training process of the action-value model can be constrained, so as to obtain an action-value model with better accuracy and performance, making the calculation of the estimated action value by the action-value model more accurate and able to conform to human intentions or preferences to a certain extent.

[0298] Next, a possible training method for the action-value model will be described by way of example in combination with steps C1 to C4. In this training method, for the loss function of the action-value model, a constraint loss term is constructed using the empirical action-value function to regularize the action-value model, thereby alleviating the overestimation error and extrapolation error in the learning process of the action-value model for the action-value function. At the same time, an action-value loss term is also considered to measure the difference between the estimated action value provided by the action-value model and the target action value in the optimization learning. Under the combined action of the constraint loss term and the action-value loss term, an action-value model with better accuracy and performance can be assisted in training, and the generalization of the action-value model can be improved. The training method of this action-value model includes the following steps C1 to C4:

[0299] C1. In any iteration, the training server obtains the estimated action value of the action indicated by each directed edge in the state topology diagram through this action-value model. This estimated action value indicates the value of the impact of the action-value model predicting that the agent executes this action on the environment.

[0300] In some embodiments, based on the empirical action-value function in step 404 and the action feedback model trained in step 402, iterative training is performed on the action-value model. In any iteration during the iterative training process of the action-value model, an estimated action value can be calculated for each directed edge in the state topology diagram through the initial action-value model. This estimated action value represents the estimation of the action value of the action indicated by the directed edge by the action-value model. The estimated action value is different from the empirical action value. The empirical action value is an empirical value derived from the empirical distribution of the state topology diagram, while the estimated action value is a predicted value calculated from the action distribution learned by the action-value model.

[0301] In some embodiments, for any directed edge in the state topology graph, the starting state s indicated by the starting node of the directed edge t and the action a indicated by the directed edge t are input together into the action value model Q θ to obtain the estimated action value Q t of the action a θ (s t , a t ). Repeating the above operations can obtain the estimated action values of executing each action at any node.

[0302] C2. The training server obtains a constraint loss term based on the empirical action value function and the estimated action value, and the constraint loss term characterizes the distribution difference between the empirical distribution and the model distribution of the action value.

[0303] In some embodiments, for any directed edge in the state topology graph, a set of data associated with the directed edge can be derived from the state topology graph: the starting state s indicated by the starting node of the directed edge t , the action a indicated by the directed edge t , the arrival state s indicated by the arrival node of the directed edge t+1 , the estimated feedback value calculated by the action feedback model according to the starting state s t and the action a t the empirical action value calculated by the empirical action value function according to the starting state s t and the action a t where t refers to the timestamp of the starting state s t , and t + 1 refers to the timestamp of the arrival state s t+1 . Therefore, for any directed edge, a set of data for this directed edge can be obtained from the state topology graph For simplicity, this set of data is referred to as the description data of this directed edge.

[0304] Furthermore, for any node s in the state topology graph, the set of actions indicated by all the directed edges starting from the node s constitutes the support action set of the node Support action set The empirical action values of all the actions in reflect the empirical distribution of the actions at the node s, while the estimated action values Q of all the actions in the support action set θ (s t , a t ) reflect the action distribution learned by the model at the node s. Therefore, for each node, in the support action set ​​Construct a constraint loss term for the action value model based on the difference between the empirical distribution and the action distribution. For a single node, this constraint loss term only considers the empirical action values of each action in its supported action set for each action on it and the predicted action value Q θ (s t , a t ). This helps to use the empirical action value function to provide constraints for the training process of the action value model Q θ , thus normalizing the action value model Q θ and accelerating the learning of the action value model Q θ for the empirical distribution.

[0305] In some embodiments, in each training iteration of the action value model Q θ , for each node in the state topology graph, according to the empirical action values of each action in its supported action set and the predicted action value Q for each action on it θ (s t , a t ), calculate the action value error of this node, and take the mathematical expectation of the action value errors of all nodes to obtain a constraint loss term. In this way, the information content of the constraint loss term is rich, and the expression ability of the constraint loss term is also stronger, which can better provide certain constraints for the action value model Q θ .

[0306] In some other embodiments, in each training iteration of the action value model Q θ , if the state topology graph contains a large number of nodes and directed edges, the training cost brought by full-graph calculation is relatively large. Then, according to a pre-set sampling rule, a part of the concerned nodes can be sampled from the state topology graph, and only for each concerned node, use the empirical action values of each action in its supported action set for each action on it and the predicted action value Q θ (s t , a t ) to calculate the action value error of this concerned node, and take the mathematical expectation of the action value errors of all concerned nodes to obtain a constraint loss term. In this way, the calculation cost of the constraint loss term is small, and the training efficiency of the action value model is improved. The embodiments of the present application do not make specific limitations on this.

[0307] Next, taking the method of constructing the constraint loss term using the action value errors of the concerned nodes as an example, the calculation process of the constraint loss term will be described in combination with steps C21 - C24:[[]]END

[0308] C21. The training server determines multiple focus nodes from the state topology graph, and the focus nodes indicate the states that need to be focused on when the agent implements the task.

[0309] In some embodiments, multiple focus nodes are obtained by randomly sampling from the node set of the state topology graph, or the nodes in the node set are arranged in order according to the latest access order, and the latest accessed multiple nodes are sampled as multiple focus nodes. Alternatively, sampling is performed from the node set in descending order of the number of accesses, and the multiple nodes with the highest number of accesses are sampled as multiple focus nodes. Or, in the same sampling manner as the nodes to be updated introduced in step B41, probability sampling is performed from the node set according to the latest access timestamp of the nodes to obtain multiple focus nodes. The embodiments of the present application do not specifically limit the sampling method of the focus nodes.

[0310] C22. For any one of the focus nodes, the training server determines the empirical action values of each action in the support action set of the focus node based on the empirical action value function, and the support action set includes the action sets indicated by each directed edge starting from the focus node.

[0311] In some embodiments, for any one of the focus nodes determined in step C21, the action sets indicated by all the directed edges starting from the focus node constitute the support action set of the focus node, and the empirical action value function can be used to calculate the empirical action values of each action in the support action set of the focus node. For example, for a certain focus node s, the support action set of the focus node s is determined and each action in the support action set is input into the empirical action value function to obtain the empirical action values of each action.

[0312] C23. The training server determines the action value error of the focus node based on the empirical action values and the estimated action values of each action in the support action set, and the action value error characterizes the difference between the empirical action value and the estimated action value of each action in the support action set.

[0313] In some embodiments, for any one of the focus nodes s determined in step C21, considering the support action set of the focus node s the empirical action values of each action in the support action set can be obtained in step C22, while the estimated action values of each action in the support action set can be obtained in step C1. Since the empirical action values of all the actions in the support action set reflect the empirical distribution of the actions on the focus node s, and the support action set while the estimated action values of each action in the support action set can be obtained in step C1. Since the empirical action values of all the actions in the support action set reflect the empirical distribution of the actions on the focus node s, and the support action set The predicted action values of all actions in [support action set] reflect the action distribution learned by the model at the attention node s. Therefore, for each attention node, according to the support action set the difference between the empirical distribution and the action distribution in [support action set], the action value error of the current attention node can be obtained.

[0314] In some embodiments, for each action in the support action set of the attention node s the difference between the empirical action value and the predicted action value of this action can be calculated, and the average value of the absolute values of the differences of all actions can be used as the final action value error of the attention node s. Alternatively, the mean square error of the absolute values of the differences of all actions can also be used as the final action value error of the joint node s. The embodiments of the present application do not specifically limit this.

[0315] In some other embodiments, in addition to calculating the average value and the mean square error, the action value error of the attention node s can also be calculated from the level of the action value vector. This will be described below in combination with steps C23a to C23c:

[0316] C23a. The training server determines the empirical action value vector of the attention node based on the empirical action values of the respective actions in the support action set.

[0317] In some embodiments, for the support action set of the attention node s the empirical action values of all actions in the support action set are concatenated into a row vector or a column vector to obtain the empirical action value vector of the attention node The empirical action value vector refers to the feature vector of the empirical action values of all actions in the support action set

[0318] C23b. The training server determines the predicted action value vector of the attention node based on the predicted action values of the respective actions in the support action set.

[0319] In some embodiments, for the support action set of the attention node s the predicted action values of all actions in the support action set are concatenated into a row vector or a column vector to obtain the predicted action value vector Q θ (s, ·) of the attention node. The predicted action value vector Q θ (s, ·) refers to the feature vector of the predicted action values of all actions in the support action set

[0320] C23c. The training server determines the action value error based on the empirical action value vector and the predicted action value vector.​​

[0321] In some embodiments, the exponential normalization function softmax is used to perform exponential normalization on the empirical action value vector calculated in step C23a to obtain a normalized empirical value vector, which represents the empirical policy of the action value and is denoted as Similarly, the exponential normalization function softmax is used to perform exponential normalization on the predicted action value vector Q θ (s, ·) as well, to obtain a normalized predicted value vector, which represents the predicted policy of the action value model and is also referred to as the model soft policy, denoted as π soft(θ) (s) = SoftmaxQ θ (s, ·).

[0322] In some embodiments, the KL distance (Kullback-Leibler Divergence, also known as KL divergence, relative entropy) between the normalized empirical value vector and the normalized predicted value vector π soft(θ) (s) is determined as the action value error of the attention node s, denoted as This can introduce the relative entropy information between the empirical policy and the model soft policy into the action value error of the attention node s, thereby improving the accuracy of the action value error of a single attention node.

[0323] In other embodiments, the Euclidean distance or cosine distance between the normalized empirical value vector and the normalized predicted value vector π soft(θ) (s) can also be determined as the action value error of the attention node s. This can consider the distance between the empirical policy and the model soft policy in the vector space in the action value error of the attention node s, and can also improve the accuracy of the action value error of a single attention node. The embodiments of the present application do not make specific limitations on this.

[0324] In steps C23a to C23c, a possible implementation manner of calculating the action value error of the attention node s from the level of the action value vector is provided. This can make the action value error measure the difference degree between the empirical policy and the model soft policy in the vector space as comprehensively and accurately as possible, and improve the accuracy of the action value error of a single attention node. Optionally, the action value error may not be calculated from the vector level, such as directly calculating the average value or the equalization error. The embodiments of the present application do not make specific limitations on this.

[0325] C24. The training server obtains the constraint loss term based on the action value errors of each attention node.

[0326] In some embodiments, for any attention node s determined in step C21, the action value error of the attention node s can be calculated according to steps C22 to C23. Then, based on the action value errors of each attention node, the constraint loss term for this training iteration can be constructed.

[0327] In some embodiments, the mathematical expectation of the action value errors of each attention node is used as the constraint loss term. At this time, the constraint loss term is expressed by the following formula:

[0328]

[0329] Wherein, represents the constraint loss term of the action value model, θ represents the parameter set of the action value model, represents the mathematical expectation, s represents the attention node, G represents the state topology graph, KL refers to the KL distance between two vectors, represents the empirical policy, π soft(θ) (s) represents the model soft policy.

[0330] In some other embodiments, in addition to using the mathematical expectation of the action value errors of each attention node as the constraint loss term, the average value of the action value errors of each attention node or the equalization error, etc. can also be used as the constraint loss term. The embodiments of the present application do not make specific limitations on this.

[0331] In steps C21 to C24, a possible implementation manner of constructing the constraint loss term based on the action value error of the attention node is provided. The constraint loss term constructed in such a manner can provide a constraint during the training process of the action value model Q θ so that the empirical action value function derived from the state topology graph can be used as the lower bound of the action value model Q θ , and its optimization goal is to try to control the estimated action value given by the action value model Q θ to be not less than the empirical action value provided by the empirical action value function .

[0332] Further, since only the actions in the support action set of a part of the attention nodes that are concerned are considered in the constraint loss term, there is no need to calculate the action value error for non-attention nodes, which greatly saves the calculation cost of the constraint loss term and improves the training efficiency of the action value model Q θ , thereby accelerating the training process of the subsequent action decision model.

[0333] Further, by θConsidering the constraint loss term in the loss function can pay more attention to the empirical distribution of actions that have occurred in historical experience. Of course, it will not completely abandon the potential distribution of actions that have not been executed. It just bets a certain constraint on the empirical distribution, making the empirical action value the lower bound of the estimated action value, effectively reducing the overestimation and extrapolation error of the action value model Q θ for the overestimation and extrapolation error of the action value.

[0334] C3. The training server obtains the action value loss term based on the action feedback model and the estimated action value. The action value loss term represents the difference between the estimated action value of the model for an action and the target action value, and the target action value represents the optimization target of the action value based on the action distribution.

[0335] In some embodiments, during the training iteration of the action value model, in addition to the constraint loss term constructed through step C2 in its loss function, there is also an action value loss term. The action value loss term is used to measure the degree of difference between the estimated action value provided by the model and the target action value, and the target action value refers to the optimization target of the action value based on the action distribution. The target action value needs to use the complete action distribution in the entire state topology graph to calculate during the calculation process, and it needs to involve the action decided by the action decision model and the estimated feedback value given by the action feedback model. Therefore, when constructing the action value loss term, it is necessary to utilize the action feedback model and the action decision model. Here, the action feedback model uses the parameter set obtained after training and optimization in step 402, and the action decision model uses the parameter set obtained after the latest training iteration at the current moment.

[0336] In some embodiments, for time step t, find the transition data at time step t from the state topology graph Use the various transition data recorded in the state topology graph to construct the target action value, and then use the estimated action value and the target action value to construct the action value loss term.

[0337] Next, in combination with steps C31 - C35, a possible construction method of the action value loss term will be illustrated by way of example. In this construction method, the action value loss term is implemented as a soft Bellman residual, as follows:

[0338] C31. The training server randomly samples the set of directed edges of the state topology graph to obtain multiple sampled edges. For any sampled edge, determine the sampled state indicated by the starting node of the sampled edge and the sampled action indicated by the sampled edge.

[0339] In some embodiments, in each training iteration of the action value model, steps C31 - C34 are used to recalculate the target action value for this iteration, so that the target action value will also calculate the latest fitting result as the action value model is optimized.

[0340] In some embodiments, although all directed edges in the state topology graph can reflect the complete action distribution, the computational overhead of calculating the target action value on the directed edges of the entire graph is large. Therefore, when calculating the target action value, random sampling can be performed from the directed edge set to obtain multiple sampled edges. Since each randomly sampled edge can also reflect the same action distribution as all directed edges, as long as the target action value is calculated based on the sampled edges, the optimization goal of the action value based on the action distribution can be better reflected, and the computational overhead of the target action value is greatly reduced, thereby improving the training efficiency of the action value model.

[0341] In other embodiments, the target action value may also be calculated based on all directed edges in the state topology graph, so that the target action value used in each training iteration has higher accuracy, which is not specifically limited in the embodiments of the present application.

[0342] In this step C31, the target action value is calculated based on the sampled edge as an example. After randomly sampling from the directed edge set to obtain multiple sampled edges, for any sampled edge, the sampling state s indicated by the starting node of the sampled edge can be determined. t and the sampling action a indicated by the sampling edge t .

[0343] C32. The training server determines, based on the action feedback model, an estimated feedback value of the agent performing the sampled action under the sampling state.

[0344] In some embodiments, for each sampled edge, the sampling state s of the sampled edge is t and sampling action a t Input to the action feedback model Feedback model through action Output for the sampling state s t and sampling action a t Calculated estimated feedback value Optionally, estimate the feedback value It can be the action feedback model after the latest update. The real-time calculation may also be the value of the latest version cached on the directed edge in the state topology graph, and the embodiments of the present application do not specifically limit this.

[0345] C33. The training server determines the probability of the agent executing the sampled action under the sampling state based on the action decision model.

[0346] In some embodiments, the sampling state s indicated by the starting node of each sampling edge t , the sampled state st Input to the action decision model π φ , through the action decision model π φ Output the execution probability π t of the sampled action a t calculated for the sampled state s φ (a t | s t ). This execution probability represents the likelihood that the agent will execute the sampled action a t in the sampled state s t . Optionally, since the action decision model and the action value model can be co-trained in the same pace or separately trained according to different iteration cycles, the execution probability π φ (a t | s t ) refers to the real-time calculation using the action decision model π φ after the latest update is completed.

[0347] C34. The training server determines the target action value of the sampled action based on the estimated feedback value, execution probability, and estimated action value of the agent executing the sampled action in this sampled state.

[0348] In some embodiments, for each sampled edge, given the sampled state s t and the sampled action a t of the sampled edge, using the estimated feedback value obtained in step C32 the execution probability π φ (a t | s t ) obtained in step C33, and the estimated action value Q θ (s t , a t ) obtained in step C1, the target action value Q target can be determined.

[0349] Next, a method of defining the target action value Q target using information entropy will be described. The calculation formula of the target action value Q target is as follows:

[0350]

[0351] where Q target represents the target action value, represents the estimated feedback value, s t represents the sampled state, a t represents the sampled action, γ is a hyperparameter, and π φ (a t | s tCharacterize the execution probability (for the same sampling state, if there are multiple possible sampling actions, the execution probabilities of its respective sampling actions can form a decision vector), Q θ (s t ,a t ) characterizes the estimated action value, α is a learnable temperature parameter (the variable α can update its value as the parameter set of the action value model is iteratively tuned), logπ φ (a t |s t ) characterizes the information entropy of the decision vector, and the temperature parameter α is used to control the weight provided by the information entropy.

[0352] C35. The training server obtains the action value loss term based on the target action value and the estimated action value.

[0353] In some embodiments, based on the target action value Q target obtained in step C34 and the estimated action value Q θ (s t ,a t ) obtained in step C1, an action value loss term can be constructed. For example, directly averaging the differences between the target action value Q target and the estimated action value Q θ (s t ,a t ) for each sampling action, or equalizing the error, etc., can quickly calculate its action value loss term.

[0354] In some other embodiments, for each sampling action, the square value of the difference between its target action value Q target and the estimated action value Q θ (s t ,a t ) can also be obtained, and then the mathematical expectation of the square values calculated for each sampling action is obtained to get the action value loss term. In this case, the action value loss term is as follows:

[0355]

[0356] Among them, Q θ (s t ,a t ) characterizes the estimated action value, s t characterizes the sampling state, a t characterizes the sampling action, Q target characterizes the target action value, G characterizes the state topology graph, τ t characterizes the transition data at time step t in the state topology graph

[0357] In steps C31 - C35, a method for constructing an action value loss term based on the soft Bellman residual is provided. In this way, the information entropy about the decision vector can be introduced into the action value loss term, making the information content of the action value loss term richer, and the action value loss term can more accurately measure the difference degree between the estimated action value and the target action value.

[0358] C4. The training server iteratively trains the action value model based on the constraint loss term and the action value loss term.

[0359] In some embodiments, based on the constraint loss term obtained in step C2 and the action value loss term obtained in step C3, the two can be summed or weighted and summed to obtain the loss function value of the current training iteration.

[0360] In one example, taking the weighted sum of the constraint loss term and the action value loss term as an example, the loss function expression of the action value function is as follows:

[0361]

[0362] Among them, J Q (θ) represents the loss function value of the action value model, represents the action value loss term, λ is a hyperparameter (referring to the weighting factor of the constraint loss term), represents the constraint loss term.

[0363] When the technical personnel set the action value optimization stop condition, through the methods provided in the above steps C1 - C4, the loss function value of the action value model in each training iteration can be calculated, and then it can be judged whether the action value optimization stop condition is satisfied. If the action value optimization stop condition is satisfied, then the training is stopped, and the trained action value model is obtained. If the action value optimization stop condition is not satisfied, the next training iteration is continued. The action value optimization stop condition can be that the loss function value tends to converge, or the number of iteration steps reaches the set number of steps, etc. Here, the action value optimization stop condition is not specifically limited.

[0364] In the training method of the action value model provided in steps C1 - C4, for the loss function of the action value model, a constraint loss term is constructed by using the empirical action value function to regularize the action value model, so as to alleviate the over - estimation error and extrapolation error in the learning process of the action value model for the action value function. At the same time, an action value loss term is also considered to measure the difference between the estimated action value provided by the action value model and the target action value in the optimization learning. Under the joint action of the constraint loss term and the action value loss term, it is possible to assist in training an action value model with better accuracy and performance, and improve the generalization of the action value model.

[0365] In the above steps 404 - 405, a possible implementation manner is provided for the training server to train the action value model of the agent based on the state topology graph and the action feedback model. Since the empirical action value function can be derived from the state topology graph, the empirical action value function is used to construct a constraint loss term, and the action feedback model and the action decision model are used to construct an action value loss term. The loss function constructed by combining the two can fully reflect the difference degrees between the action distribution predicted by the model and the empirical distribution and the optimization target respectively, making the finally trained action value model have better generalization ability.

[0366] In some other embodiments, in the loss function of the action value model, only the action value loss term can be considered and the constraint loss term can be not considered. This can save the computational cost of the loss function value and improve the training efficiency of the action value model. The embodiments of the present application do not make specific limitations on this.

[0367] 406. The training server trains the action decision model of the agent based on the action value model, and the action decision model is used to decide the action that the agent should execute in a given state.

[0368] In some embodiments, based on the action value model trained in step 405, since the action value model can feedback the estimated action value of each action in a given state, and when the parameters of the action decision model are different, different actions may be decided in the same state. Therefore, the estimated action value given by the action value model for the decided target action can reflect the accuracy of the parameters of the action decision model itself, where the target action refers to the action finally selected to be executed according to the decision vector. Therefore, the action value model can assist in completing the training and optimization of the action decision model. Since a more optimal action value model has been obtained in step 405, the finally optimized action decision model will also have higher accuracy and better performance.

[0369] In some other embodiments, the action value model and the action decision model are co-trained and optimized, that is, in each iteration of the iterative training, the action value model and the action decision model are updated once according to the state topology graph and the action feedback model. There is no order preference for the optimization of the action value model and the action decision model. The action value model can be updated first and then the action decision model, or the action decision model can be updated first and then the action value model, or the action value model and the action decision model can be updated simultaneously. There is no specific limitation here. Iteratively execute the above optimization process until the action decision model meets the decision optimization stop condition. At this time, the trained action decision model is obtained. The decision optimization stop condition can be that the loss function value tends to converge, or the number of iteration steps reaches the set number of steps, etc. There is no specific limitation on the decision optimization stop condition here. In the co-training and optimization process, the action value model and the action decision model are updated once in each iteration, which can make the action value model and the action decision model guide each other, thereby further improving the accuracy of the finally optimized action decision model.

[0370] Hereinafter, in combination with steps D1 to D4, a possible training method of the action decision model will be described. In this training method, a decision loss term is introduced into the loss function of the action decision model, and the estimated action value of the action output by the action value model is considered in the decision loss term, so that the training process of the action decision model can be completed under the guidance of the action value model. The description is as follows:

[0371] D1. In any iteration, the training server determines, through the action decision model, the decision vector of the agent in the current state, and the decision vector indicates the possibility of the agent performing various actions at the current moment.

[0372] In some embodiments, for the given state s at the current moment t t , the state s t is input into the action decision model π φ . Through the action decision model π φ , the decision vector π t (a θ | s t ) is calculated for the state s t . The decision vector includes the execution probability of the agent performing each possible action in the state s t , and each execution probability represents the possibility of the agent performing a possible action in the state s t .

[0373] D2. The training server determines, based on the action value model, a scoring vector of the action values of the agent in the current state, where the scoring vector indicates the value expected to be brought by the agent performing each action at the current moment.

[0374] In some embodiments, for a given state s at the current moment t t , the action decision model π φ outputs a decision vector π t calculated for the state s φ (a t |s t ). The decision vector π φ (a t |s t ) includes the execution probabilities of the agent performing each possible action in the state s t . Then, the state s t and each possible action a t can form a state-action pair. For each pair of the state s t and the action a t , they are input into the action value model Q θ . The action value model Q θ outputs an estimated action value Q t for each pair of the state s t and the action a θ (s t , a t ). Using the estimated action values Q t of the state s t with respect to all possible actions a θ (s t , a t ), a scoring vector Q t (s θ ) for the state s t can be constructed. Optionally, since the action decision model π φ and the action value model Q θ can be co-trained in the same pace or separately trained according to different iteration cycles, the scoring vector Q θ (s t ) refers to being calculated in real time using the action value model Q θ after the latest update is completed.

[0375] Taking the co-training framework as an example, in each training iteration, the action decision model π φ calculates the decision vector π t for the agent in the current state s φ (a t |s t),the agent then determines which action to execute according to the decision vector π φ (a t |s t ) and then executes the action a t and interacts with the environment. The environment gives an estimated feedback value through the action feedback model and refreshes the next state s t+1 . Then, the access times of the state topology graph are updated, or new nodes and directed edges are added. Next, the action value model Q θ gives the estimated action value Q θ (s t t , a θ ). Thus, in this training iteration, for the action value model Q θ and the action decision model π φ , their respective parameter sets are updated once (the updates of the two have no order distinction and can also be updated simultaneously), and then enter the next training iteration. In the previous step 405, the loss function of the action value model Q θ and its training method were introduced, while in this step 406, the loss function of the action decision model π φ and its training method are introduced.

[0376] D3. The training server determines the decision loss term of the agent at the current moment based on the decision vector and the scoring vector, and the decision loss term represents the error between the action executed by the agent's decision and the task expectation.

[0377] In some embodiments, based on the decision vector π φ (a t |s t ) obtained in step D1 and the scoring vector Q θ (s t ) obtained in step D2, the decision loss term of the agent at the current moment t is calculated.

[0378] In an example, the difference degree between the information entropy of the decision vector π φ (a t |s t ) and the scoring vector Q θ (s t ) is used to construct the decision loss term, and the expression of the decision loss term is as follows:

[0379]

[0380] where J π (φ) represents the decision loss term, φ represents the parameter set of the action decision model π, represents for each state s in the state topology graph G tFind the mathematical expectation of the value in the square brackets, π φ (a t |s t ) represents the decision vector, π φ (a t |s t ) T represents the transposed vector of the decision vector, α is a learnable temperature parameter, logπ φ (a t |s t ) represents the information entropy of the decision vector, and the temperature parameter α is used to control the weight provided by the information entropy, Q θ (s t ) represents the scoring vector. It should be noted that the temperature parameter α in the above formula can reuse the α used when obtaining the target action value Q target in step C34, but if the target action value Q target is not defined using information entropy, then the value of the variable α can be updated along with the iterative tuning of the parameter set of the action decision model, and the embodiments of the present application do not specifically limit this.

[0381] D4. The training server iteratively trains the action decision model based on this decision loss term.

[0382] In some embodiments, the loss function value of the action decision model is equal to the decision loss term obtained in step D3. Then, at this time, when the technical personnel set the decision optimization stop condition, through the methods provided in the above steps D1 to D4, the loss function value of the action decision model in each training iteration can be calculated, and then it can be judged whether the decision optimization stop condition is satisfied. If the decision optimization stop condition is satisfied, then the training is stopped, and the trained action decision model is obtained. If the decision optimization stop condition is not satisfied, the next training iteration is continued. The decision optimization stop condition can be that the loss function value tends to converge, or the number of iterative steps reaches the set number of steps, etc. Here, the decision optimization stop condition is not specifically limited.

[0383] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated here one by one.

[0384] The method provided by the embodiments of the present application constructs a non-parametric state topology graph based on historical trajectories. This state topology graph can fully reflect the empirical distribution of the actions of the intelligent agent, has a higher information utilization rate for historical trajectories, and brings more information. Furthermore, based on the state topology graph, it guides the training of the action feedback model, enabling the action feedback model to calculate a more accurate estimated feedback value, improving the accuracy of the action feedback model. Then, by combining the state topology graph and the action feedback model, it can constrain the training process of the action value model, obtain an action value model with better accuracy and performance, make the calculation of the estimated action value by the action value model more accurate, and can, to a certain extent, conform to human intentions or preferences. Finally, using the action value model with better accuracy to assist in training an action decision model with better accuracy, which helps to make precise decisions on which action the intelligent agent should execute in a given state.

[0385] In the above embodiments, the training process of the action decision model of the intelligent agent designed in the present application is introduced in detail. Next, in combination with Figure 5 , taking the action decision model as the policy neural network π φ , the action feedback model as the reward neural network , and the action value model as the action value neural network Q θ as an example, a possible implementation manner of this training process will be illustrated by way of example.

[0386] As Figure 5 shown, at the beginning of the algorithm, the policy neural network π φ , the reward neural network , and the action value neural network Q θ are initialized. Then, the policy neural network π φ is used to control the interaction between the intelligent agent and the environment, and multiple historical trajectories are collected. Specifically, the state s observed at any moment is input into the policy neural network π φ , and the decision vector π φ (a|s) in the state s is predicted through the policy neural network π φ . According to the decision vector, the action a to be executed is determined, and the intelligent agent is controlled to execute the action a. After the intelligent agent executes the action a, it interacts with the environment, causing the environment to update the new state s' at the next moment. Then, the environment can output a triple (s, a, s') to the reward neural network and output an estimated feedback value . Then, the quadruple (s, a, The transition data at a time step is stored in the experience replay buffer as (‘s’). The above operations are looped until the agent's task is completed, the task fails, or the set length is reached, obtaining a historical trajectory, and this historical trajectory will be stored in the experience replay buffer in the form of transition data at multiple time steps.

[0387] Further, based on a series of transition data stored in the experience replay buffer, a state topology graph G is constructed. Specifically, the nodes in the state topology graph G are constructed based on the states observed in the historical trajectory, and the directed edges in the state topology graph G are constructed based on the actions executed in the historical trajectory. It should be noted that since there may be unobserved states, there may be unexecuted actions in the entire action distribution. Therefore, the state topology graph G can be regarded as a subset of the complete action distribution in the environmental space.

[0388] Further, trajectory sampling is performed on the state topology graph G to obtain multiple pairs of sampled trajectories. Since the efficiency of trajectory sampling is high, it can efficiently construct effective and more informative trajectory pairs to query human preferences. Moreover, trajectory sampling allows the use of the topological structure of the state topology graph G to splice new sampled trajectories. The sampled trajectories are not necessarily directly obtained historical trajectories, but are formed by connecting different transition segments in multiple historical trajectories. In this way, the sampled trajectories do not actually occur in history, but the transition segments in the sampled trajectories are all real historical occurrences. These sampled trajectories can be presented to technicians or large models in pairs, and the technicians or large models can make annotations to obtain the annotation results of each pair of sampled trajectories. These annotation results reflect the preference data of humans (or learned by the large model). Then, using the preference data as a supervision signal, the state topology graph G is used to assist the reward neural network to perform supervised learning to obtain an optimized reward neural network Optionally, at regular intervals, the reward neural network is updated according to the preference data collected during the period for the parameter set ψ. After each update of the parameter set ψ, it is necessary to recalculate and relabel the estimated feedback value for each directed edge in the state topology graph G to achieve a full-graph update of the estimated feedback value, thus maximizing the utilization of historical transitions and alleviating the adverse effects brought by non-stationary reward functions.

[0389] Further, an empirical action value function can be derived from the state topology graph G such that the empirical action value function can calculate an empirical action value for each state s indicated by a node using the quadruple (s, a, ) The action-value neural network Q can be trained in reverse θ and the policy neural network π φ . Specifically, using the empirical action value to provide constraints and regularization for the action-value neural network Q θ such that the empirical action value serves as a lower bound for the estimated action value Q θ given by the action-value neural network Q θ (s, ·), thereby fully considering the information provided by the empirical distribution of actions and preventing the action-value neural network Q θ from overestimating the action value, reducing the overestimation error and extrapolation error of the action-value neural network Q θ so that the action-value neural network Q θ can more accurately estimate the value generated by each action in a given environment. Then, using the estimated action value Q θ (s, ·) to introduce evaluation information for the action value to the policy neural network π φ , ultimately improving the generalization and policy performance of the policy neural network π φ and enhancing the training efficiency and learning efficiency of the policy neural network π φ . In addition, the policy neural network π φ can try to provide better policy suggestions for the agent, thereby making decisions on high-quality actions that conform to human preferences or intentions, which means that the agent can more accurately understand human preferences, reducing the number of times of human intervention required when the agent executes tasks and reducing the labor cost.

[0390] It should be noted that the action-value neural network Q θ and the policy neural network π φ can be iteratively trained based on a co-training framework, that is, in each iteration, the parameters θ and φ of the action-value neural network Q θ and the policy neural network π φ are each updated once, and finally a trained action-value neural network Q θ and a policy neural network π φ can be obtained.

[0391] It should also be noted that during the entire training iteration process, as the historical trajectories continue to increase, if a directed edge already existing in the state topology graph G is discovered again, the access count of this directed edge is updated. If an action indicated by a non-existent directed edge is discovered, then a new directed edge and its corresponding reachable node are added to the state topology graph G. In this way, the state topology graph G will be continuously improved during the adaptive update. Further, the empirical action value function can also update the empirical action value of each directed edge regularly according to the access count, realizing the adaptive update and optimization of the empirical action value.

[0392] Taking the scenario of multiple robotic carts collaborating for transportation as an example, through the preference-based reinforcement learning framework of the embodiments of the present application, the robotic carts can more precisely meet the needs and preferences of humans. For example, when multiple robotic carts are required to cooperate to move a cargo with a specific shape and weight, the trained policy neural network π φ can provide the optimal movement, cooperation, and path planning strategies for each robotic cart based on the previous preference data, ensuring the safe, fast, and efficient movement of the cargo.

[0393] In the above various embodiments, the construction and update processes of the state topology graph are introduced in detail, and the training processes of the action feedback model, action value model, and action decision model are also introduced in detail. In the embodiments of the present application, the process of using the trained action decision model to control the agent to execute actions will be described in detail.

[0394] Figure 6 is a flowchart of a method for an agent to make action decisions provided by an embodiment of the present application. As Figure 6 shown, this embodiment is executed by a computer device, which can be the agent 201 or the agent control system 202 in the above-mentioned implementation environment, or other control terminals installed with the action decision model. Taking the computer device as the control terminal as an example for illustration, this embodiment includes the following steps:

[0395] 601. When the control terminal observes the state at the current moment in the environment, it inputs this state into the action decision model of the agent, and this action decision model is used to decide the action that the agent should execute in the given state.

[0396] Among them, the action decision-making model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal on the action executed by the agent in the environment, and the action value model is used to evaluate the value of the impact of the action executed by the agent on the environment. For the training processes of the action feedback model, the action value model, and the action decision-making model, please refer to the above respective embodiments in detail, and will not be elaborated here.

[0397] In some embodiments, the training server can train a more accurate and better-performing action decision-making model under a preference-based reinforcement learning framework by using the training method of the action decision-making model of the agent provided in the above embodiments. Then, the training server sends the parameter set of the action decision-making model to the control terminal of the agent, so that the control terminal obtains the trained action decision-making model.

[0398] Optionally, the training process of the action decision-making model can be completed locally on the training server, such as local offline training on the training server, or can be completed in the cloud, such as distributed training jointly performed by multiple servers to improve the training efficiency. The embodiments of the present application do not make specific limitations on this.

[0399] Optionally, the control terminal of the agent can be a control module built into the agent or a control device independent of the agent, and multiple agents can be macroscopically scheduled by the same control device, which is also called the master control terminal at this time. The embodiments of the present application do not make specific limitations on this.

[0400] In some embodiments, after the control terminal obtains the trained action decision-making model π φ for the environment where the agent resides, if the state s at the current moment is observed in this environment, the state s can be input into the action decision-making model π of the agent φ and enter the following step 602.

[0401] 602. The control terminal determines the execution probability of the agent for each of multiple candidate actions through the action decision-making model, and the execution probability represents the possibility of the agent executing the candidate action in this state.

[0402] In some embodiments, after the control terminal inputs the state s into the action decision-making model π φ of the agent, through the action decision-making model π φ a decision vector π φ (a|s) is calculated for the state s, and the decision vector π φ(a|s) includes the execution probability of the agent performing each possible action a in state s, and each execution probability characterizes the likelihood of the agent performing an action a in state s. For example, the agent can perform a total of 7 candidate actions: moving forward, moving backward, moving left, moving right, jumping, squatting, and staying still. However, in a certain state s, the agent moves to a corner, making it impossible for the agent to perform the actions of moving left and moving backward. Then, the action decision model π φ will still calculate 7 execution probabilities for the 7 candidate actions respectively. At this time, the execution probabilities of moving left and moving backward both approach 0.

[0403] 603. The control terminal determines the target action to be performed by the agent at the current moment from the multiple candidate actions based on the execution probability of each of the multiple candidate actions.

[0404] In some embodiments, the control terminal is based on the decision vector π in step 602 φ (a|s), and can decide a target action from multiple candidate actions. The target action refers to the optimal and most human-preferred candidate action planned by the action decision model π for the agent in the current state s. φ

[0405] In some embodiments, the candidate action with the highest execution probability can be directly selected as the target action. Or, a target action can also be randomly selected from the top N candidate actions with the highest execution probabilities. Or, sampling can also be performed according to the execution probability on the probability distribution of the candidate actions, and the target action finally determined to be executed by the agent can be randomly decided, so that the sampling process follows the probability distribution. The embodiments of the present application do not specifically limit this.

[0406] Taking the scenario of multiple robot carts collaborating in transportation as an example, after the action decision model π φ is trained, if the robot cart itself is its own control terminal, the training server will send the parameter set of the action decision model π φ to each robot cart, so that each robot cart makes decisions and executes actions under the control of the action decision model π φ built in itself. Multiple robot carts cooperate to move the goods from the starting point to the ending point to complete the goods transportation task; or, if there is a master control terminal for multiple robot carts, the master control terminal is responsible for scheduling the actions of each robot cart. Then, the training server will send the parameter set of the action decision model π φ to the master control terminal of these robot carts. The master control terminal makes decisions in the action decision model π φUnder the control of [the controller], the actions of each robotic vehicle are determined, and the control signals for each determined action are respectively sent to the corresponding robotic vehicle, so as to achieve the macro scheduling of multiple robotic vehicles by the same master control terminal. To complete the cargo transportation task, multiple robotic vehicles need to cooperate to move the cargo from the starting point to the ending point. The embodiments of the present application do not specifically limit this.

[0407] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated one by one here.

[0408] The method provided by the embodiments of the present application conducts collaborative training through a state topology graph, an action feedback model, and an action value model, and obtains an action decision model with better performance and generalization ability. Given the state of an agent, the action decision model will try its best to determine high-quality actions that conform to human preferences or intentions for the agent, which means that the agent can more accurately understand human preferences, reduce the number of times of human intervention required when the agent executes tasks, and reduce the labor cost.

[0409] In the above embodiments, it is introduced in detail how the action decision model trained by the preference-based reinforcement learning algorithm makes decisions on the actions of the agent in the application scenario. This preference-based reinforcement learning algorithm uses the empirical action value function to assist in trajectory sampling from the state topology graph, and normalizes the neural network-based action value function (i.e., the action value model) to achieve efficient learning. Secondly, experiments and tests show that the algorithm provided by the embodiments of the present application outperforms other preference-based reinforcement learning algorithms in various complex tasks and greatly improves the human feedback efficiency. Herein, the human feedback efficiency, simply referred to as feedback efficiency, refers to how to maximize the learning effect with limited human feedback (since the preference data labeled by human experts is usually relatively expensive, traditional algorithms often require more human feedback for learning, so the human feedback efficiency is relatively low). Finally, through the use of aligned empirical estimates, the algorithm provided by the embodiments of the present application shows obvious advantages in the performance comparison with other traditional algorithms, especially in scenarios where only limited human feedback labels are available. At the same time, the algorithm provided by the embodiments of the present application can train an accurate action value function (i.e., the action value model) and a better policy (i.e., the action decision model).

[0410] Next, the test performance of some tasks in the test scenario under different solutions will be described.

[0411] Figure 7 is a performance comparison diagram of an agent in the box-pushing task provided by the embodiments of the present application, as Figure 7As shown in the figure, the task of a robot pushing a box is selected as the experimental task in the test scenario. For the robot, the box-pushing task is a relatively complex task. The action decision-making model of the robot car is trained using this solution, traditional solutions 1-3, and the baseline solution 4 respectively. The action decision-making model trained with various solutions is used to control the robot to complete the box-pushing task, and the average return learning curves of various solutions are plotted. Among them, (a)-(c) are the test results of the average return learning curve when the human-labeled feedback is 300, and (d)-(f) are the test results of the average return learning curve when the human-labeled feedback is 1000. The environmental sizes of (a) and (d) are both 5×5, the environmental sizes of (b) and (e) are both 6×6, and the environmental sizes of (c) and (f) are both 7×7. The unit of the environmental size is the size specified by the map scale.

[0412] Figure 7 In (a)-(f), the comparison of 6 groups of average return learning curves is shown. The horizontal axis is the time axis (Timesteps, time steps), and the vertical axis is the average learning return rate (Episode Return). The higher the average learning return rate, the better the performance of the action decision-making model. For each comparison of the average return learning curves, in each experimental task, the baseline solution 4 refers to the highest performance obtained using the true reward function. Therefore, the baseline solution 4 is plotted as a horizontal line, representing the upper bound of the performance of the experimental task.

[0413] From Figure 7 it can be clearly seen that this solution not only surpasses the current traditional solutions 1-3 in terms of performance, but also is particularly outstanding in terms of sample efficiency. Moreover, in the initial stage of training, this solution can quickly achieve high performance in most tasks, and the speed is significantly faster than other traditional solutions 1-3. It should be noted that this solution only requires a small number of human preference labels (that is, the annotation results of human beings for trajectory pairs or segment pairs, also referred to as preference data) to be close to the performance of the baseline solution, which reflects extremely high feedback efficiency.

[0414] At the same time, from Figure 7 it can also be observed that some baseline solutions are significantly affected by random factors in some tasks, resulting in unstable performance in some experiments and large fluctuations in the learning curve. On the contrary, there are also some traditional solutions whose learning curves show a downward trend in more difficult tasks, which implies that these solutions are more sensitive to the quality of trajectory samples. Specifically, when a large number of sampled trajectory pairs are of low quality, it may interfere with the training of the action value function (that is, the action value model), and thus affect the performance of the overall strategy (that is, the action decision-making model).

[0415] Figure 8It is a performance comparison chart of a robot in a construction task provided by an embodiment of the present application. As Figure 8 shown, for the field of robot construction, the robot construction task is selected as the experimental task under the test scenario. For robots, the construction task is also a relatively complex task. In the test, for the CraftEnv environment in the field of robot construction, three construction tasks are selected: Strip-Shaped, Block-Shaped, and Two-Story tasks, to conduct test research. The action decision models of the robot are trained respectively using the present solution, traditional solutions 1-3, and the baseline solution 4, and the action decision models trained with various solutions are used to control the robot to complete the above three construction tasks, and the average return learning curves of various solutions are drawn.

[0416] Figure 8 (a)-(c) in it show the comparison of 3 groups of average return learning curves. (a) is the construction task of the strip-shaped building, (b) is the construction task of the block-shaped building, and (c) is the construction task of the two-story building. The horizontal axis is the time axis, and the vertical axis is the average learning return rate. In addition, the human-annotated feedback for the three construction tasks is 1000. These tasks fully demonstrate the complexity and diversity of construction tasks in the real world, requiring the agent robot (i.e., the intelligent agent) to be able to accurately operate the building components to meet specific building design requirements.

[0417] From Figure 8 it can be clearly seen that the present solution not only outperforms traditional solutions 1-3 in terms of algorithm performance, but also has a great improvement in feedback efficiency. More notably, although the number of samples used in the present solution is much less than that of traditional solution 2, its performance is already comparable to that of traditional solution 2, and even exceeds traditional solution 2. Specifically, in the strip-shaped building task, the present solution only uses 30% of the sample quantity and has already surpassed the average performance of traditional solution 2. This result can be observed through the comparison of the curves representing the present solution and traditional solution 2 in Figure 8 and the performance of the present solution is particularly outstanding. These research results verify that when dealing with complex tasks, the present solution can not only achieve efficient learning, but also greatly reduce the required number of feedbacks, thus greatly improving the learning efficiency.

[0418] It should also be noted that the state topology graph used in the embodiments of the present application adopts a simple non-parametric model, i.e., a graph model, but is not limited to a specific non-parametric model or its topological structure, and can be replaced with various novel and effective model structures according to requirements. In addition, there is no fixed choice for the underlying reinforcement learning algorithm in this solution. In practical applications, an appropriate reinforcement learning algorithm can be selected according to specific scenarios and requirements. At the same time, the learning framework of this solution can also be combined with other technologies, such as exploration of unsupervised learning, expansion of time-series data, and data augmentation using pseudo-labels, etc.

[0419] Figure 9 It is a schematic structural diagram of a training device for an action decision model of an agent provided by an embodiment of the present application, as Figure 9 shown. The device includes:

[0420] A topology graph construction module 901, configured to construct a state topology graph based on multiple historical trajectories of the agent. Each of the historical trajectories includes multiple actions, and each of the actions is used to control the transition between different states. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action;

[0421] A feedback training module 902, configured to train an action feedback model of the agent based on the state topology graph. The action feedback model is used to provide a feedback signal of the environment where the agent resides to the action executed by the agent;

[0422] An action value training module 903, configured to train an action value model of the agent based on the state topology graph and the action feedback model. The action value model is used to evaluate the value of the action executed by the agent on the environment;

[0423] A decision training module 904, configured to train an action decision model of the agent based on the action value model. The action decision model is used to decide the action that the agent should execute in a given state.

[0424] The device provided by the embodiment of the present application constructs a non-parametric state topology graph based on the historical trajectory. This state topology graph can fully reflect the empirical distribution of the actions of the intelligent agent, has a higher utilization rate of the information of the historical trajectory, and brings more information. Furthermore, based on the state topology graph, it guides the training of the action feedback model, enabling the action feedback model to calculate a more accurate estimated feedback value, improving the accuracy of the action feedback model. Then, by combining the state topology graph and the action feedback model, it can constrain the training process of the action value model, obtain an action value model with better accuracy and performance, make the calculation of the estimated action value by the action value model more accurate, and can, to a certain extent, conform to human intentions or preferences. Finally, using the action value model with better accuracy to assist in training an action decision model with better accuracy, thus contributing to making a precise decision on which action the intelligent agent should execute in a given state.

[0425] In some embodiments, the topology graph construction module 901 is configured to:

[0426] Initialize the state topology graph;

[0427] For any action in any one of the historical trajectories, if a directed edge indicating the action is queried in the state topology graph, update the access count associated with the directed edge, where the access count indicates the query frequency of the directed edge;

[0428] If a directed edge indicating the action is not queried in the state topology graph, determine the starting state and the reaching state associated with the action, and based on the starting node indicating the starting state, add a reaching node indicating the reaching state and a directed edge pointing from the starting node to the reaching node.

[0429] In some embodiments, based on Figure 9 the device composition, the feedback training module 902 includes:

[0430] A trajectory sampling sub-module, configured to perform trajectory sampling based on the state topology graph to obtain multiple pairs of sampled trajectories, where each pair of the sampled trajectories includes a pair of sampled trajectories with equal lengths;

[0431] A label acquisition sub-module, configured to acquire the label results of the multiple pairs of sampled trajectories, where the label results indicate the satisfaction degree of each pair of the sampled trajectories in distinguishing different sampled trajectories with respect to the task executed by the intelligent agent;

[0432] A feedback training sub-module, configured to train the action feedback model based on the state topology graph and the label results when the feedback model update condition is satisfied.

[0433] In some embodiments, based on Figure 9 the device composition, the trajectory sampling sub-module includes:

[0434] A random sampling unit, configured to randomly sample from the node set of the state topology graph to obtain a plurality of sampling points;

[0435] A trajectory sampling unit, configured to start trajectory sampling along a directed edge starting from any one of the plurality of sampling points, and stop sampling when the trajectory length reaches the sampling length, to obtain a sampling trajectory;

[0436] A trajectory pairing unit, configured to pair according to the trajectory length based on a plurality of sampling trajectories to obtain multiple pairs of sampling trajectories.

[0437] In some embodiments, the trajectory sampling unit is configured to:

[0438] If there is only one directed edge starting from the sampling point, use the arrival node pointed to by the directed edge as the next sampling point;

[0439] If there are at least two directed edges starting from the sampling point, randomly select the directed edge with the highest or lowest empirical action value from the at least two directed edges, and use the arrival node pointed to by the randomly selected directed edge as the next sampling point, where the empirical action value indicates the value of the impact expected to be generated on the environment when the intelligent agent executes an action according to historical experience.

[0440] In some embodiments, based on Figure 9 the composition of the device, the device further includes:

[0441] A feedback update module, configured to update the estimated feedback value of each directed edge in the state topology graph after the action feedback model is updated, where the estimated feedback value indicates the feedback signal expected to be generated by the environment when the intelligent agent executes the action indicated by the directed edge.

[0442] In some embodiments, based on Figure 9 the composition of the device, the action value training module 903 includes:

[0443] A function acquisition sub-module, configured to acquire an empirical action value function based on the state topology graph, where the empirical action value function is used to provide an empirical action value for evaluating the actions executed by the intelligent agent based on the state topology graph;

[0444] An action value training sub-module, configured to train the action value model based on the empirical action value function and the action feedback model.

[0445] In some embodiments, based on Figure 9 the composition of the device, the action value training sub-module includes:

[0446] An action value prediction unit, configured to obtain, in any iteration, the predicted action values of the actions indicated by each directed edge in the state topology graph through the action value model, where the predicted action values indicate the values of the impact on the environment when the action value model predicts that the agent executes the actions;

[0447] A constraint loss acquisition unit, configured to obtain a constraint loss term based on the empirical action value function and the predicted action values, where the constraint loss term represents the distribution difference between the empirical distribution and the model distribution of the action values;

[0448] An action value loss acquisition unit, configured to obtain an action value loss term based on the action feedback model and the predicted action values, where the action value loss term represents the difference between the predicted action value of the action by the model and the target action value, and the target action value represents the optimization target of the action value based on the action distribution;

[0449] An action value training unit, configured to iteratively train the action value model based on the constraint loss term and the action value loss term.

[0450] In some embodiments, based on Figure 9 the composition of the device, the constraint loss acquisition unit includes:

[0451] A node determination subunit, configured to determine a plurality of attention nodes from the state topology graph, where the attention nodes indicate the states that the agent needs to pay attention to when implementing the task;

[0452] An empirical action value determination subunit, configured to, for any one of the attention nodes, determine the empirical action values of each action in the support action set of the attention node based on the empirical action value function, where the support action set includes the set of actions indicated by each directed edge starting from the attention node, and the empirical action value indicates the value of the expected impact on the environment when the agent executes the action according to historical experience;

[0453] An action value error determination subunit, configured to determine the action value error of the attention node based on the empirical action values and the predicted action values of each action in the support action set, where the action value error represents the difference between the empirical action value and the predicted action value of each action in the support action set;

[0454] A constraint loss acquisition subunit, configured to obtain the constraint loss term based on the action value errors of each attention node.

[0455] In some embodiments, the action value error determination subunit is configured to:

[0456] Determine the empirical action value vector of the attention node based on the empirical action values of each action in the support action set;

[0457] Determine the estimated action value vector of the concerned node based on the estimated action values of each action in the support action set;

[0458] Determine the action value error based on the empirical action value vector and the estimated action value vector.

[0459] In some embodiments, the action value loss acquisition unit is configured to:

[0460] Randomly sample the directed edge set of the state topology graph to obtain multiple sampled edges. For any sampled edge, determine the sampled state indicated by the starting node of the sampled edge and the sampled action indicated by the sampled edge;

[0461] Based on the action feedback model, determine the estimated feedback value of the agent executing the sampled action in the sampled state;

[0462] Based on the action decision model, determine the execution probability of the agent for the sampled action in the sampled state;

[0463] Based on the estimated feedback value, execution probability, and estimated action value of the agent executing the sampled action in the sampled state, determine the target action value of the sampled action;

[0464] Based on the target action value and the estimated action value, obtain the action value loss term.

[0465] In some embodiments, the decision training module 904 is configured to:

[0466] In any iteration, through the action decision model, determine the decision vector of the agent in the current state, where the decision vector indicates the possibility of the agent executing various actions at the current moment;

[0467] Based on the action value model, determine the scoring vector of the action value of the agent in the current state, where the scoring vector indicates the value expected to be brought by the agent executing each action at the current moment;

[0468] Based on the decision vector and the scoring vector, determine the decision loss term of the agent at the current moment, where the decision loss term represents the error between the action decided by the agent to execute and the task expectation;

[0469] Based on the decision loss term, iteratively train the action decision model.

[0470] In some embodiments, based on Figure 9 the device composition, the device further includes an action value update module, and the action value update module includes:

[0471] The node sampling sub-module is used to sample from the node set of the state topology diagram when the action value update condition is met, and obtain multiple nodes to be updated;

[0472] The action set determination sub-module is used to determine the support action set of each node to be updated, and the support action set includes the action sets indicated by each directed edge starting from the node to be updated;

[0473] The action value update sub-module is used to update the empirical action value of each action in the support action set, and the empirical action value indicates the value of the impact that the agent is expected to have on the environment when executing the action according to historical experience.

[0474] In some embodiments, the action value update sub-module is used for:

[0475] For any action in the support action set, determine multiple target nodes that can be reached by executing the action starting from the node to be updated;

[0476] Based on the access times of each directed edge connecting the node to be updated and each target node, determine the empirical transition probability of the node to be updated;

[0477] Through the action feedback model, determine the estimated feedback value of executing the action starting from the node to be updated, and the estimated feedback value indicates the feedback signal that the environment is expected to generate when the agent executes the action;

[0478] Based on the empirical transition probability, the estimated feedback value, and the empirical action value of executing the action starting from the node to be updated, obtain the candidate action value of the action;

[0479] Assign the maximum value among the candidate action values of each action in the support action set to the empirical action value of executing the action starting from the node to be updated.

[0480] It should be noted that: when training the action decision model of the agent provided in the above embodiments, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device (such as a training server) is divided into different functional modules to complete all or part of the functions described above. In addition, the training device of the action decision model of the agent provided in the above embodiments and the embodiment of the training method of the action decision model of the agent belong to the same concept, and the specific implementation process can be seen in the embodiment of the training method of the action decision model of the agent, which will not be repeated here.

[0481] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated here one by one.

[0482] Figure 10 is a schematic structural diagram of an action decision-making device for an agent provided by an embodiment of the present application. As Figure 10 shown, the device includes:

[0483] An input module 1001, configured to input the state at the current moment into the action decision-making model of the agent when the state at the current moment is observed in the environment, where the action decision-making model is used to decide the action that the agent should execute in a given state;

[0484] A probability determination module 1002, configured to determine, through the action decision-making model, the execution probability of each of multiple candidate actions for the agent, where the execution probability represents the possibility that the agent executes the candidate action in the state;

[0485] An action determination module 1003, configured to determine, based on the execution probability of each of the multiple candidate actions, a target action that the agent executes at the current moment from the multiple candidate actions;

[0486] Among them, the action decision-making model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment on the action executed by the agent, and the action value model is used to evaluate the value of the action executed by the agent on the environment.

[0487] The device provided by the embodiment of the present application obtains an action decision-making model with better performance and better generalization ability through collaborative training of a state topology graph, an action feedback model, and an action value model, so that in a given state of the agent, the action decision-making model will try its best to decide a high-quality action that conforms to human preferences or intentions for the agent, that is, it means that the agent can more accurately understand human preferences, reduces the number of times of manual intervention required when the agent executes tasks, and reduces the labor cost.

[0488] It should be noted that when the action decision-making device of the agent provided in the above embodiment decides the action of the agent, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of a computer device (such as an agent or a control terminal of the agent) is divided into different functional modules to complete all or part of the functions described above. In addition, the action decision-making device of the agent provided in the above embodiment and the embodiment of the action decision-making method of the agent belong to the same concept, and the specific implementation process is detailed in the embodiment of the action decision-making method of the agent, which will not be repeated here.

[0489] Any combination of the above optional technical solutions can form an optional embodiment of the present disclosure, which will not be elaborated herein one by one.

[0490] Figure 11 It is a schematic structural diagram of a control terminal provided by an embodiment of the present application. As Figure 11 shown, the control terminal is an exemplary illustration of a computer device. The control terminal is used to control the actions of the intelligent agent, and can be a control system built into the intelligent agent or a control device independent of the intelligent agent. Optionally, the device type of the control terminal 1100 includes: smart phone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) player, notebook computer or desktop computer. The control terminal 1100 may also be referred to by other names such as user equipment, portable control terminal, laptop control terminal, desktop control terminal, etc.

[0491] Generally, the control terminal 1100 includes a processor 1101 and a memory 1102.

[0492] Optionally, the processor 1101 includes one or more processing cores, such as a 4-core processor, an 8-core processor, etc. Optionally, the processor 1101 is implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). In some embodiments, the processor 1101 includes a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 integrates a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 further includes an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0493] In some embodiments, the memory 1102 includes one or more computer-readable storage media, optionally non-transitory. Optionally, the memory 1102 further includes high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1102 is used to store at least one program code, and the at least one program code is used to be executed by the processor 1101 to implement the action decision method of the agent provided in various embodiments of the present application.

[0494] In some embodiments, the control terminal 1100 may further optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, the memory 1102, and the peripheral device interface 1103 can be connected through a bus or signal lines. Each peripheral device can be connected to the peripheral device interface 1103 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1108.

[0495] The peripheral device interface 1103 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 are implemented on separate chips or circuit boards, and this embodiment does not limit this.

[0496] The radio frequency circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1104 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts the received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. Optionally, the radio frequency circuit 1104 communicates with other control terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1104 further includes a circuit related to NFC (Near Field Communication), which is not limited in this application.

[0497] The display screen 1105 is used to display the UI (User Interface). Optionally, the UI includes graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch display screen, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signals can be input as control signals to the processor 1101 for processing. Optionally, the display screen 1105 is also used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there is one display screen 1105, which is set on the front panel of the control terminal 1100; in other embodiments, there are at least two display screens 1105, which are respectively set on different surfaces of the control terminal 1100 or are in a foldable design; in still other embodiments, the display screen 1105 is a flexible display screen, which is set on the curved surface or the folding surface of the control terminal 1100. Even more optionally, the display screen 1105 is set to an irregular non-rectangular shape, that is, a special-shaped screen. Optionally, the display screen 1105 is prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0498] The camera component 1106 is used to collect images or videos. Optionally, the camera component 1106 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the control terminal, and the rear camera is disposed on the back of the control terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the background blurring function by fusing the main camera and the depth camera, the panoramic shooting and VR (Virtual Reality) shooting functions or other fusion shooting functions by fusing the main camera and the wide-angle camera. In some embodiments, the camera component 1106 further includes a flash. Optionally, the flash is a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which is used for light compensation under different color temperatures.

[0499] In some embodiments, the audio circuit 1107 includes a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 1101 for processing, or input them to the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo collection or noise reduction, there are multiple microphones, which are respectively disposed at different parts of the control terminal 1100. Optionally, the microphone is an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signal from the processor 1101 or the radio frequency circuit 1104 into sound waves. Optionally, the speaker is a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1107 further includes a headphone jack.

[0500] The power supply 1108 is used to supply power to each component in the control terminal 1100. Optionally, the power supply 1108 is alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 1108 includes a rechargeable battery, the rechargeable battery supports wired charging or wireless charging. The rechargeable battery is also used to support fast charging technology.

[0501] In some embodiments, the control terminal 1100 further includes one or more sensors 1110. The one or more sensors 1110 include but are not limited to: an acceleration sensor 1111, a gyroscope sensor 1112, a pressure sensor 1113, an optical sensor 1114, and a proximity sensor 1115.

[0502] In some embodiments, the acceleration sensor 1111 detects the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the control terminal 1100. For example, the acceleration sensor 1111 is used to detect the components of the gravitational acceleration on the three coordinate axes. Optionally, the processor 1101 controls the display screen 1105 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1111. The acceleration sensor 1111 is also used to collect game or user movement data.

[0503] In some embodiments, the gyroscope sensor 1112 detects the body direction and rotation angle of the control terminal 1100, and the gyroscope sensor 1112 cooperates with the acceleration sensor 1111 to collect the 3D actions of the user on the control terminal 1100. The processor 1101 implements the following functions according to the data collected by the gyroscope sensor 1112: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0504] Optionally, the pressure sensor 1113 is disposed on the side frame of the control terminal 1100 and / or the lower layer of the display screen 1105. When the pressure sensor 1113 is disposed on the side frame of the control terminal 1100, it can detect the holding signal of the user on the control terminal 1100, and the processor 1101 performs left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1113. When the pressure sensor 1113 is disposed on the lower layer of the display screen 1105, the processor 1101 controls the operable controls on the UI interface according to the pressure operation of the user on the display screen 1105. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0505] The optical sensor 1114 is used to collect the ambient light intensity. In one embodiment, the processor 1101 controls the display brightness of the display screen 1105 according to the ambient light intensity collected by the optical sensor 1114. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 1105 is decreased. In another embodiment, the processor 1101 also dynamically adjusts the shooting parameters of the camera module 1106 according to the ambient light intensity collected by the optical sensor 1114.

[0506] The proximity sensor 1115, also known as the distance sensor, is usually disposed on the front panel of the control terminal 1100. The proximity sensor 1115 is used to collect the distance between the user and the front of the control terminal 1100. In one embodiment, when the proximity sensor 1115 detects that the distance between the user and the front of the control terminal 1100 is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from the lit state to the off state; when the proximity sensor 1115 detects that the distance between the user and the front of the control terminal 1100 is gradually increasing, the processor 1101 controls the display screen 1105 to switch from the off state to the lit state.

[0507] Those skilled in the art can understand that Figure 11 the structure shown in does not constitute a limitation on the control terminal 1100, and can include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0508] Figure 12 is a schematic structural diagram of a training server provided by an embodiment of the present application. As Figure 12 shown, the training server is an exemplary illustration of a computer device. The training server 1200 may vary greatly due to different configurations or performances. The training server 1200 includes one or more processors (Central Processing Units, CPUs) 1201 and one or more memories 1202. Among them, at least one computer program is stored in the memory 1202, and the at least one computer program is loaded and executed by the one or more processors 1201 to implement the training method of the action decision model of the agent provided in the above various embodiments. Optionally, the training server 1200 also has components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The training server 1200 also includes other components for implementing the functions of the device, which will not be elaborated here.

[0509] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including at least one computer program. The at least one computer program can be executed by a processor in the terminal to complete the training method of the action decision model of the agent or the action decision of the agent in the above various embodiments. For example, the computer-readable storage medium includes ROM (Read-Only Memory), RAM (Random-Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tapes, floppy disks, and optical data storage devices, etc.

[0510] In an exemplary embodiment, a computer program product is further provided, including one or more computer programs, and the one or more computer programs are stored in a computer-readable storage medium. One or more processors of a computer device can read the one or more computer programs from the computer-readable storage medium, and the one or more processors execute the one or more computer programs, so that the computer device can execute to complete the training method of the action decision model of the agent or the action decision of the agent in the above embodiment.

[0511] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. Optionally, the program is stored in a computer-readable storage medium. Optionally, the above-mentioned storage medium is a read-only memory, a magnetic disk, an optical disc, or the like.

[0512] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for an action decision-making model of an intelligent agent, characterized in that The method includes: Based on multiple historical trajectories of the agent, construct a state topology graph, where each historical trajectory contains multiple actions, and each action is used to control the transition between different states. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action; Based on the state topology graph, train the action feedback model of the agent, where the action feedback model is used to provide a feedback signal of the environment where the agent resides on the action executed by the agent; Based on the state topology graph and the action feedback model, train the action value model of the agent, where the action value model is used to evaluate the value of the impact of the action executed by the agent on the environment; Based on the action value model, train the action decision model of the agent, where the action decision model is used to decide the action that the agent should execute in a given state.

2. The method according to claim 1, characterized in that The constructing a state topology graph based on multiple historical trajectories of the agent includes: Initialize the state topology graph; For any action in any of the historical trajectories, if a directed edge indicating the action is queried in the state topology graph, update the access count associated with the directed edge, where the access count indicates the query frequency of the directed edge; If a directed edge indicating the action is not queried in the state topology graph, determine the start state and the reach state associated with the action, and based on the start node indicating the start state, add a reach node indicating the reach state and a directed edge pointing from the start node to the reach node.

3. The method according to claim 1, characterized in that, The training the action feedback model of the agent based on the state topology graph includes: Perform trajectory sampling based on the state topology graph to obtain multiple pairs of sampled trajectories, where each pair of sampled trajectories contains a pair of sampled trajectories with equal lengths; Collect the annotation results of the multiple pairs of sampled trajectories, where the annotation results indicate the satisfaction degree of each pair of sampled trajectories in distinguishing different sampled trajectories with respect to the task executed by the agent; When the feedback model update condition is satisfied, train the action feedback model based on the state topology graph and the annotation results.

4. The method according to claim 3, characterized in that, The performing trajectory sampling based on the state topology graph to obtain multiple pairs of sampled trajectories includes: Randomly sample from the node set of the state topology graph to obtain multiple sampling points; Starting from any one of the multiple sampling points, start trajectory sampling along the directed edge starting from the sampling point, and stop sampling when the trajectory length reaches the sampling length to obtain a sampled trajectory; Based on multiple sampled trajectories, pair them according to the trajectory length to obtain multiple pairs of sampled trajectories.

5. The method according to claim 4, wherein The starting from the sampling point, start trajectory sampling along the directed edge starting from the sampling point, and stop sampling when the trajectory length reaches the sampling length to obtain a sampled trajectory includes: If there is only one directed edge starting from the sampling point, use the reach node pointed to by the directed edge as the next sampling point; If there are at least two directed edges starting from the sampling point, randomly select the directed edge with the highest or lowest empirical action value from the at least two directed edges, and use the arrival node pointed to by the randomly selected directed edge as the next sampling point. The empirical action value indicates the value of the impact expected on the environment when the agent executes an action according to historical experience.

6. The method according to claim 3, characterized in that, The method further includes: After the action feedback model is updated, update the estimated feedback value of each directed edge in the state topology graph. The estimated feedback value indicates the feedback signal expected to be generated by the environment when the agent executes the action indicated by the directed edge.

7. The method according to claim 1, wherein Training the action value model of the agent based on the state topology graph and the action feedback model includes: Based on the state topology graph, obtain an empirical action value function, which is used to provide an empirical action value for evaluating the actions executed by the agent based on the state topology graph. Based on the empirical action value function and the action feedback model, train the action value model.

8. The method according to claim 7, wherein Training the action value model based on the empirical action value function and the action feedback model includes: In any iteration, through the action value model, obtain the estimated action value of the action indicated by each directed edge in the state topology graph. The estimated action value indicates the value of the impact expected on the environment when the action value model predicts that the agent executes the action. Based on the empirical action value function and the estimated action value, obtain a constraint loss term, which represents the distribution difference between the empirical distribution and the model distribution of the action value. Based on the action feedback model and the estimated action value, obtain an action value loss term, which represents the difference between the estimated action value of the action by the model and the target action value. The target action value represents the optimization target of the action value based on the action distribution. Based on the constraint loss term and the action value loss term, iteratively train the action value model.

9. The method according to claim 8, wherein Obtaining the constraint loss term based on the empirical action value function and the estimated action value includes: Determine a plurality of attention nodes from the state topology graph. The attention nodes indicate the states that the agent needs to pay attention to when achieving the task. For any one of the attention nodes, based on the empirical action value function, determine the empirical action value of each action in the support action set of the attention node. The support action set contains the set of actions indicated by each directed edge starting from the attention node. The empirical action value indicates the value of the impact expected on the environment when the agent executes the action according to historical experience. Based on the empirical action value and the estimated action value of each action in the support action set, determine the action value error of the attention node, which represents the difference between the empirical action value and the estimated action value of each action in the support action set. Based on the action value errors of each attention node, obtain the constraint loss term.

10. The method according to claim 9, wherein Determining the action value error of the attention node based on the empirical action value and the estimated action value of each action in the support action set includes: Determine the empirical action value vector of the concerned node based on the empirical action values of each action in the support action set; Determine the predicted action value vector of the concerned node based on the predicted action values of each action in the support action set; Determine the action value error based on the empirical action value vector and the predicted action value vector; 11. The method according to claim 8, wherein The obtaining of the action value loss term based on the action feedback model and the predicted action value includes: Randomly sample the set of directed edges of the state topology graph to obtain multiple sampled edges. For any sampled edge, determine the sampled state indicated by the starting node of the sampled edge and the sampled action indicated by the sampled edge; Based on the action feedback model, determine the predicted feedback value when the agent executes the sampled action in the sampled state; Based on the action decision model, determine the execution probability of the agent for the sampled action in the sampled state; Based on the predicted feedback value, execution probability, and predicted action value when the agent executes the sampled action in the sampled state, determine the target action value of the sampled action; Obtain the action value loss term based on the target action value and the predicted action value; 12. The method according to claim 1, wherein The training of the action decision model of the agent based on the action value model includes: In any iteration, through the action decision model, determine the decision vector of the agent in the current state, where the decision vector indicates the possibility of the agent executing various actions at the current moment; Based on the action value model, determine the scoring vector of the action value of the agent in the current state, where the scoring vector indicates the value expected to be brought by the agent executing each action at the current moment; Based on the decision vector and the scoring vector, determine the decision loss term of the agent at the current moment, where the decision loss term represents the error between the action decided by the agent to execute and the task expectation; Iteratively train the action decision model based on the decision loss term; 13. The method according to claim 1, characterized in that The method further includes: When the action value update condition is satisfied, sample from the set of nodes of the state topology graph to obtain multiple nodes to be updated; Determine the support action set of each node to be updated, where the support action set contains the set of actions indicated by each directed edge starting from the node to be updated; Update the empirical action value of each action in the support action set, where the empirical action value indicates the value expected to affect the environment when the agent executes the action according to historical experience; 14. The method according to claim 13, wherein The updating of the empirical action value of each action indicated by each directed edge in the support action set includes: For any action in the support action set, determine the multiple target nodes that can be reached by executing the action starting from the node to be updated; Based on the number of visits of each directed edge connecting the node to be updated and each target node, determine the empirical transition probability of the node to be updated; Through the action feedback model, determine the predicted feedback value when executing the action starting from the node to be updated, where the predicted feedback value indicates the feedback signal expected to be generated by the environment when the agent executes the action; Obtain a candidate action value for the action based on the empirical transition probability, the estimated feedback value, and the empirical action value of performing the action starting from the node to be updated; Assign the maximum value among the candidate action values of each action in the support action set to the empirical action value of performing the action starting from the node to be updated.

15. A method for action decision-making of an intelligent agent, characterized in that, The method includes: When observing the state at the current moment in the environment, input the state into the action decision model of the intelligent agent, where the action decision model is used to decide the action that the intelligent agent should perform in a given state; Determine, through the action decision model, the execution probability of the intelligent agent for each of multiple candidate actions, where the execution probability represents the likelihood of the intelligent agent performing the candidate action in the state; Based on the execution probabilities of the multiple candidate actions, determine the target action that the intelligent agent performs at the current moment from the multiple candidate actions; Among them, the action decision model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment on the action performed by the intelligent agent, and the action value model is used to evaluate the value of the action performed by the intelligent agent on the environment.

16. A training device for an action decision-making model of an intelligent agent, characterized in that, The device includes: A topology graph construction module, configured to construct a state topology graph based on multiple historical trajectories of the intelligent agent. Each historical trajectory includes multiple actions, and each action is used to control the transition between different states. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action; A feedback training module, configured to train the action feedback model of the intelligent agent based on the state topology graph, where the action feedback model is used to provide a feedback signal of the environment where the intelligent agent resides on the action performed by the intelligent agent; An action value training module, configured to train the action value model of the intelligent agent based on the state topology graph and the action feedback model, where the action value model is used to evaluate the value of the action performed by the intelligent agent on the environment; A decision training module, configured to train the action decision model of the intelligent agent based on the action value model, where the action decision model is used to decide the action that the intelligent agent should perform in a given state.

17. An action decision-making device for an intelligent agent, characterized in that, The device includes: An input module, configured to input the state into the action decision model of the intelligent agent when observing the state at the current moment in the environment, where the action decision model is used to decide the action that the intelligent agent should perform in a given state; A probability determination module, configured to determine, through the action decision model, the execution probability of the intelligent agent for each of multiple candidate actions, where the execution probability represents the likelihood of the intelligent agent performing the candidate action in the state; An action determination module, configured to determine the target action that the intelligent agent performs at the current moment from the multiple candidate actions based on the execution probabilities of the multiple candidate actions; Among them, the action decision-making model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment on the action executed by the intelligent agent, and the action value model is used to evaluate the value of the impact of the action executed by the intelligent agent on the environment.

18. A computer device, characterized in that, The computer device includes one or more processors and one or more memories. At least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the training method of the action decision-making model of the intelligent agent according to any one of claims 1 to 14, or the action decision-making method of the intelligent agent according to claim 15.

19. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the training method of the action decision-making model of the intelligent agent according to any one of claims 1 to 14, or the action decision-making method of the intelligent agent according to claim 15.

20. A computer program product, characterized in that, The computer program product includes at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the training method of the action decision-making model of the intelligent agent according to any one of claims 1 to 14, or the action decision-making method of the intelligent agent according to claim 15.

Citation Information

Cited By

  • WiFi communication quality optimization method of ESP32 microcontroller

    CN120475421A

  • Mirror image-based agent training system

    CN121303182A

  • Intelligent agent real-time decision-making method and system based on Riemannian manifold and medium

    CN122311474A

  • Training method for action decision-making model of intelligent agent, and action decision-making method and apparatus

    WO2025148822A1