Training method for action decision-making model of intelligent agent, and action decision-making method and apparatus

By constructing a state topology diagram and training action feedback model, using human preference data to optimize the agent's action decision model, the problem of insufficient performance of the agent's action decision model is solved, and more efficient and accurate action decisions are achieved, which are suitable for scenarios such as robot collaboration and robot arm control.

WO2025148822A1PCT designated stage expired Publication Date: 2025-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2025/070705
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2025-01-06
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the prior art, the performance of the action decision model of the agent is insufficient, resulting in the agent being unable to perform appropriate actions and failing to complete tasks. Especially in industrial application scenarios such as robot control, video and games, there are challenges in formulating appropriate reward functions, and preferential reinforcement learning technology requires high labor costs and inefficient preference data query.

Method used

By constructing a state topology diagram, training the action feedback model and action value model, using human preference data for supervised learning, optimizing the action decision model, improving the accuracy and efficiency of the model, reducing manual intervention, and improving the utilization rate and query efficiency of the preferred data.

Benefits of technology

It realizes more accurate action decisions for the agent in scenarios such as robot collaboration and robot arm control, reduces development costs, improves the learning efficiency of the model and the accuracy of action decisions, meets human intentions and preferences, and reduces the risk of reward hacking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070705_17072025_PF_FP_ABST
    Figure CN2025070705_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a training method for an action decision-making model of an intelligent agent, and an action decision-making method and apparatus. The training method comprises: constructing a state topology graph on the basis of historical trajectories; on the basis of the state topology graph, training an action feedback model of an intelligent agent, wherein the action feedback model is used for providing a feedback signal of an environment where the intelligent agent resides in respect of an action executed by the intelligent agent; on the basis of the state topology graph and the action feedback model, training an action value model of the intelligent agent, wherein the action value model is used for providing an estimated action value of the action executed by the intelligent agent, and the estimated action value indicates a metric value used for measuring the impact of the action executed by the intelligent agent on the environment; and on the basis of the action value model, training an action decision-making model of the intelligent agent, wherein the action decision-making model is used for deciding an action to be executed by the intelligent agent in a given state.
Need to check novelty before this filing date? Find Prior Art

Description

Training method, action decision method and device for action decision model of intelligent agent

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 9, 2024, with application number 202410039699.6, and invention name “Training method, action decision method and device for action decision model of intelligent body”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a training method, an action decision method, and an apparatus for an action decision model of an intelligent agent. Background Art

[0003] With the development of computer and robotics technology, intelligent agents, represented by robots, robotic arms, and large models, have attracted widespread attention. Currently, the application scenarios of intelligent agents have gradually expanded to include a wide range of industrial applications, including robotic control, video, and gaming, as well as interactive and non-interactive scenarios between intelligent agents and humans.

[0004] Technical content

[0005] The embodiments of the present application provide a training method, an action decision method and an apparatus for an action decision model of an intelligent agent, which can train a more accurate action decision model, improve the performance of the action decision model, and thus make more accurate action decisions for the intelligent agent.

[0006] The present invention provides a method for training an action decision model of an intelligent agent, which is executed by a computer device. The method includes:

[0007] Based on multiple historical trajectories of the intelligent agent, a state topology graph is constructed, where each historical trajectory contains multiple actions, each action is used to control the transition between different states, each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action;

[0008] Based on the state topology graph, training an action feedback model of the agent, wherein the action feedback model is used to provide a feedback signal of the environment in which the agent resides on the action performed by the agent;

[0009] Training an action value model of the agent based on the state topology graph and the action feedback model, wherein the action value model is used to provide an estimated action value for an action performed by the agent, wherein the estimated action value indicates a metric value used to measure the impact of the action performed by the agent on the environment;

[0010] Based on the action value model, an action decision model of the intelligent agent is trained, and the action decision model is used to decide the action that the intelligent agent should perform in a given state.

[0011] The present application also provides an action decision method for an intelligent agent, which is executed by a computer device. The method includes:

[0012] When the current state of the environment is observed, the state is input into the action decision model of the agent, and the action decision model is used to decide the action that the agent should perform in the given state;

[0013] Determining, by means of the action decision model, an execution probability of the agent for each of a plurality of candidate actions, wherein the execution probability represents the likelihood of the agent executing the candidate action in the state;

[0014] Determining a target action for the agent to perform at the current moment from the multiple candidate actions based on the execution probabilities of the multiple candidate actions;

[0015] Among them, the action decision model is obtained by collaborative training based on a state topology graph, an action feedback model and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment to the action performed by the agent, and the action value model is used to provide an estimated action value for the action performed by the agent. The estimated action value indicates a measurement value for measuring the impact of the action performed by the agent on the environment.

[0016] The present application also provides a device for training an action decision model of an intelligent agent, the device comprising:

[0017] A topology graph construction module is used to construct a state topology graph based on multiple historical trajectories of the intelligent agent, each of which contains multiple actions, each of which is used to control the transition between different states, each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action;

[0018] A feedback training module, configured to train an action feedback model of the agent based on the state topology graph, wherein the action feedback model is configured to provide a feedback signal from the environment in which the agent resides to an action performed by the agent;

[0019] an action value training module, configured to train an action value model of the agent based on the state topology graph and the action feedback model, wherein the action value model is configured to provide an estimated action value for an action performed by the agent, wherein the estimated action value indicates a metric value for measuring the impact of the action performed by the agent on the environment;

[0020] A decision training module is used to train the action decision model of the intelligent agent based on the action value model, and the action decision model is used to decide the action that the intelligent agent should perform in a given state.

[0021] The present application also provides an action decision-making device for an intelligent agent, the device comprising:

[0022] An input module is used to input the current state observed in the environment into the action decision model of the intelligent agent, and the action decision model is used to determine the action that the intelligent agent should perform in the given state;

[0023] a probability determination module, configured to determine, by means of the action decision model, an execution probability of the agent for each of a plurality of candidate actions, wherein the execution probability represents a likelihood that the agent will execute the candidate action in the state;

[0024] an action determination module, configured to determine a target action to be performed by the agent at the current moment from among the multiple candidate actions based on the execution probabilities of the multiple candidate actions;

[0025] Among them, the action decision model is obtained by collaborative training based on a state topology graph, an action feedback model and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment on the action performed by the agent, and the action value model is used to provide an estimated action value for the action performed by the agent. The estimated action value indicates a measurement value for measuring the impact of the action performed by the agent on the environment.

[0026] An embodiment of the present application also provides a computer device, which includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement a training method for an action decision model of an intelligent agent or an action decision method for an intelligent agent as described in any possible implementation method described above.

[0027] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement a training method for an action decision model of an intelligent agent or an action decision method of an intelligent agent as described in any possible implementation method described above.

[0028] The present application also provides a computer program product, comprising one or more computer programs stored in a computer-readable storage medium. One or more processors of a computer device can read the one or more computer programs from the computer-readable storage medium, and the one or more processors execute the one or more computer programs, so that the computer device can perform the method for training an action decision model of an intelligent agent or the method for action decision of an intelligent agent according to any of the possible implementations described above.

[0029] BRIEF DESCRIPTION OF THE DRAWINGS

[0030] FIG1 is a schematic flow chart of a preference-based reinforcement learning algorithm provided in an embodiment of the present application;

[0031] FIG2 is a schematic diagram of an implementation environment of a method for training an action decision model of an intelligent agent provided in an embodiment of the present application;

[0032] FIG3 is a flow chart of a method for training an action decision model of an intelligent agent provided in an embodiment of the present application;

[0033] FIG4 is a flow chart of a method for training an action decision model of an intelligent agent provided in an embodiment of the present application;

[0034] FIG5 is a diagram of a training framework of an action decision model provided in an embodiment of the present application;

[0035] FIG6 is a flow chart of an action decision-making method for an intelligent agent provided in an embodiment of the present application;

[0036] FIG7 is a performance comparison diagram of an agent on a box-pushing task provided by an embodiment of the present application;

[0037] FIG8 is a performance comparison diagram of a robot provided in an embodiment of the present application in a construction task;

[0038] FIG9 is a schematic structural diagram of a training device for an action decision model of an intelligent agent provided in an embodiment of the present application;

[0039] FIG10 is a schematic diagram of the structure of an action decision-making device of an intelligent agent provided in an embodiment of the present application;

[0040] FIG11 is a schematic structural diagram of a control terminal provided in an embodiment of the present application;

[0041] FIG12 is a schematic diagram of the structure of a training server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0043] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.

[0044] In this application, the term "at least one" means one or more, and the term "plurality" means two or more. For example, a plurality of historical traces means two or more historical traces.

[0045] In this application, the term "including at least one of A or B" refers to the following situations: including only A, including only B, and including both A and B.

[0046] The user-related information (including but not limited to the user's device information, personal information, behavioral information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application, when applied to specific products or technologies in the manner of the embodiments of this application, are all permitted, agreed, authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant information, data and signals must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the annotation results for track pairs or fragment pairs involved in this application are all obtained with full authorization.

[0047] With the advancement of computer and robotics technology, intelligent agents, represented by robots, robotic arms, and large models, have attracted widespread attention. The application scenarios of intelligent agents have gradually expanded to include numerous industrial applications, including robotic control, video, and gaming. The action decision model determines the action that the agent performs. If the performance of the action decision model is insufficient, the agent will be unable to execute the appropriate action and ultimately fail to complete the task. Therefore, how to train a more accurate action decision model to improve its performance is an urgent problem.

[0048] The solution provided in the embodiments of the present application involves machine learning technology of artificial intelligence, specifically reinforcement learning (RL), also known as reinforced learning, evaluation learning or enhanced learning, which is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of how intelligent agents maximize rewards or achieve specific goals through learning strategies during their interaction with the environment.

[0049] The classic model for reinforcement learning is the standard Markov decision process (MDP). Depending on the given conditions, reinforcement learning can be divided into model-based reinforcement learning (Model-Based RL) and model-free reinforcement learning (Model-Free RL), as well as active reinforcement learning (Active RL) and passive reinforcement learning (Passive RL). Variants of reinforcement learning include inverse reinforcement learning, hierarchical reinforcement learning, and reinforcement learning for partially observable systems. Algorithms used to solve reinforcement learning problems can be divided into two categories: policy search algorithms and value function algorithms.

[0050] Reinforcement learning theory, inspired by behaviorist psychology, focuses on online learning and attempts to maintain a balance between exploration and exploitation. Unlike supervised and unsupervised learning, reinforcement learning does not require any pre-defined data. Instead, it obtains learning information and updates model parameters by receiving rewards (feedback) from the environment in response to actions. Reinforcement learning has been discussed in fields such as information theory, game theory, and automatic control, and has been used to explain equilibrium states under bounded rationality, design recommendation systems, and robotic interaction systems. Some sophisticated reinforcement learning algorithms possess a degree of general intelligence for solving complex problems, reaching human-level performance in Go and video games.

[0051] In some task scenarios, deep learning models can be used in reinforcement learning, forming deep reinforcement learning (DRL). Deep reinforcement learning combines the perception capabilities of deep learning with the decision-making capabilities of reinforcement learning to achieve end-to-end learning from perception to action. It can directly control based on input signals and is an AI approach that is closer to human thinking. Deep reinforcement learning has the potential to enable robots to truly and fully autonomously learn one or even multiple skills.

[0052] Below, we will explain the terms or concepts involved in deep reinforcement learning technology.

[0053] An agent is an intelligent entity, that is, a software or hardware entity that can act autonomously. It is a very important concept in the field of artificial intelligence. Any independent entity that can think and interact with its environment can be abstracted as an agent. In the field of artificial intelligence, the term "agent" has also been translated as "agent," "agent," "intelligent subject," "agent," etc. To put it another way, an agent is a computational entity that resides in a certain environment, can function continuously and autonomously, and possesses characteristics such as residency, responsiveness, sociality, and initiative. It can be either hardware (such as a robot) or software. An agent can interpret data obtained from the environment that reflects events occurring in the environment and perform actions that affect the environment.

[0054] Environment refers to the space or scene where the agent resides. The environment can be a part of the real world or a part of the virtual world.

[0055] State refers to the state of an agent's interaction with its environment at a given moment. State is a concept related to time. An agent can observe the same or different states at different moments. For example, at time t, the agent observes state s, and at time t+1, it observes another state s'.

[0056] Actions are behaviors or actions performed by an agent that affect the environment. Agents influence the environment through actions, and agents can typically perform one or more actions. For example, a robot agent can perform various actions within the environment, such as walking, running, and jumping.

[0057] Trajectory, or the trajectory of an agent's actions, refers to a series of actions performed by the agent over a continuous period of time, arranged in chronological order. Typically, when an agent performs a task, it continuously interacts with the environment from the initial action until the task is completed or the number of actions reaches a set length, forming a complete action sequence called a trajectory. For example, when multiple robotic carts collaborate to transport goods, they start from the starting point and stop when the goods are delivered or the number of steps reaches a preset number. The action sequence formed by the continuous actions performed by each robotic cart during the time period from departure to stop is called a trajectory for that robotic cart.

[0058] A segment is a segment of a continuous action segmented or cut out from an agent's trajectory. A segment is a subset of a trajectory. Since trajectories are typically long, to facilitate annotation, each trajectory is segmented or cut out according to the sampling length, resulting in multiple segments of equal length (equal to the sampling length).

[0059] A policy, defined by a policy function, is typically presented as an action decision model, such as a policy neural network or other parametric model. It is used to make decisions based on observed states to control the movement of an agent. For example, a policy function π is presented as a probability density function. Given any state s, the policy function π can determine the probability that the agent will perform any action under the given state s. This probability represents the likelihood that the agent will perform that action.

[0060] Rewards: After an agent performs an action, the environment can reward the agent. These rewards are typically defined by a reward function, often expressed as an action-feedback model, such as a reward neural network or other parameterized model. The goal of reinforcement learning is to maximize the total reward earned by an agent for a series of actions performed within a timeframe.

[0061] State transition refers to the process by which an agent moves between different states in an environment by executing actions. The process of transitioning from the old state at the previous moment to the new state at the current moment is called a state transition. For example, after observing state s at time t, the agent executes action a in the environment, causing it to transition from state s to another state s' at time t+1. The state transition from state s to state s' can be abstracted as a conditional probability density function P. Given the current state s and action a, the conditional probability density function P can predict the probability of transitioning to another state s' at the next moment.

[0062] The interaction between the agent and the environment (Agent Environment Interaction): After the agent observes a certain state at the current moment, it will perform the corresponding action; after the agent takes action, the environment will be affected by the action and updated to the state of the next moment, completing the state transfer from the current moment to the next moment. At the same time, the environment will also return a feedback signal (or reward signal) to the agent.

[0063] An action-value function (A-VF) is a function used to evaluate the value of an agent's action at a given moment. The action-value function Q is related to the policy function π. The same agent will have different action-value functions Q when using different policy functions π. For example, if the policy function π remains unchanged, the action-value function Q can reflect the impact of the agent's action a on the environment under the current state s. The action-value function is often presented as an action-value model, such as an action-value neural network or other parametric models.

[0064] Robots include all machines that simulate human behavior or thoughts and other living things (such as robot dogs, robot cats, robot cars, etc.). Some computer programs are even called robots (such as chat robots, dialogue robots, etc.). The robots involved in the embodiments of this application refer to artificial machine devices that can automatically perform tasks to replace or assist human work. The artificial machine devices can be anthropomorphic or anthropomorphic, and are generally electromechanical devices controlled by computer programs or electronic circuits. Typically, a robot consists of a visual sensor, a robotic arm, and a main control computer.

[0065] A mechanical arm is a complex system widely used in robotics, characterized by high precision, multiple inputs and outputs, high nonlinearity, and strong coupling. Due to its unique operational flexibility, the mechanical arm can be coupled not only to the robot body but also to any other man-made machine device.

[0066] Deep reinforcement learning (DRL) technology has made rapid progress in recent years. Using DRL, intelligent agents can master a wide range of complex tasks and skills, encompassing robotics control, video games, and numerous industrial applications. However, the key to the success of DRL is a carefully designed reward function. Designing an appropriate reward function remains a challenging task in many practical reinforcement learning scenarios. The quality of a reward function depends heavily on the designer's deep understanding of the core logic of the problem and relevant background knowledge. For example, developing a reward function for text generation tasks is particularly challenging, as the key lies in measuring the quality of generated text using a single scalar value. Despite significant efforts in reward design, research has highlighted numerous issues in current algorithms and applications, such as "reward hacking," where an agent, in order to maximize rewards, engages in behaviors that are unexpected or even harmful to the human engineer. In this scenario, the agent focuses on exploiting flaws in the reward function to maximize rewards, while ignoring whether its own behavior meets expectations. This can lead to unintended and potentially risky behavior.

[0067] In light of this, preference-based reinforcement learning (PbRL) technology has received widespread attention and has spawned a series of algorithms. Compared to algorithms that rely on reward functions designed by human engineers, preference-based reinforcement learning techniques leverage human preferences to learn action-feedback models (such as reward neural networks). Specifically, humans can provide preferences for a pair of action trajectories of an intelligent agent. For example, a technician can be presented with a pair of historical trajectories performed by the intelligent agent, who can then label which of these historical trajectories meets the human preference, thereby implicitly indicating the behavior or task goal that the intelligent agent needs to learn. A historical trajectory refers to the action trajectory of the intelligent agent over a past period of time, with each historical moment in that past period having a unique and deterministic state and action. By learning from human feedback (i.e., whether the trajectory meets human preferences), the intelligent agent can complete a specific task or master a certain behavior required by humans.

[0068] This application embodiment involves an effective preference-based reinforcement learning algorithm. Without requiring human engineers to design a reward function, it can learn an action-reward model from human preferences to facilitate the training of an intelligent agent's action decision model. Testing has shown that this algorithm can train intelligent agents to exhibit novel behaviors and, to a certain extent, mitigate the challenge of reward hacking.

[0069] The basic framework of the preference-based reinforcement learning algorithm will be described below in conjunction with Figure 1. Figure 1 is a principle flow chart of a preference-based reinforcement learning algorithm provided by an embodiment of the present application. As shown in Figure 1, Characterize the action decision model of the intelligent agent, is the parameter set of the action decision model. When the state s is observed, the action decision model The decision agent needs to perform action a. After performing action a, the agent interacts with the environment, causing the environment to update to another state s'; Representing action feedback models, It is the parameter set of the action feedback model, which is used to estimate the reward value that the environment should feedback to the agent when performing action a in state s to reach the new state s'.

[0070] Defining a quad The transfer data can indicate the state transition from state s to state s' and its related information. The transfer data is stored in the experience replay buffer (Replay Buffer). In other words, the experience replay buffer is used to store the historical trajectory of the agent.

[0071] By utilizing the historical trajectories in the experience replay buffer, different trajectory pairs can be constructed through random sampling or non-random sampling. By presenting each pair of trajectories to a technician and allowing him to choose which trajectory in the pair is more in line with his preferences, and recording the results of compliance or non-compliance marked by the technician, a positive sample trajectory and a negative sample trajectory can be generated for each pair of trajectories, thereby enabling the intelligent agent to query humans' preferences for these trajectory pairs.

[0072] Alternatively, in cases where the lengths of individual trajectories are generally long, to improve query efficiency, different pairs of segments can be constructed from historical trajectories through various methods such as truncation, cutting, or sampling. Each pair of segments is presented to a technician, who is then asked to select which of the pair is more in line with their preferences. The technician's annotations of compliance or noncompliance are recorded, and a positive and a negative sample segment are generated for each pair of segments, thereby enabling the agent to query humans' preferences for these segments. For example, after setting a sampling length, each trajectory is cut into a series of segments whose length does not exceed the sampling length. Then, a pair of segments of equal length is randomly selected from all the cut segments, thereby sampling a number of segment pairs. The method for constructing the segment pairs is not limited here.

[0073] To a certain extent, the preference data obtained through the above query method (i.e., the labeling results for each pair of trajectories or segments) can reflect human expectations or desires for the agent's behavior. Using this preference data, supervised learning techniques can be used to learn and recover the underlying reward function. This reward function specifies the reward value for selecting an action in a given state, thereby providing feedback to the action-value function, resulting in a more accurate action-value function. The action-value function can then help optimize the agent's policy function, enabling it to make better and more appropriate decisions about the agent's actions. By repeating the above process, the agent can complete the training of its policy function using human preference data.

[0074] To put it another way, based on preference data, supervised learning technology can train an action feedback model that estimates reward values ​​more accurately. The more accurate the reward values ​​given by the action feedback model, the better the performance of the related action value model will be, making the action value model more accurate in its assessment of action values, which in turn guides the action decision model of the intelligent agent, and ultimately optimizes to obtain an action decision model with better performance and more accurate decisions, thus realizing the training of the action decision model based on preference data.

[0075] Because preference-based reinforcement learning relies on preference data—that is, technicians' annotation results for trajectory or segment pairs require humans to manually annotate a large amount of preference data—preference-based reinforcement learning techniques incur high labor costs in many application scenarios and utilize preference data inefficiently. Furthermore, when constructing trajectory or segment pairs, historical trajectories are typically randomly sampled, then screened and paired using various sampling methods to form a random pair of trajectories for querying human preferences. Therefore, preference queries can only be performed using existing historical trajectories. Due to the random nature of trajectory pair construction, it's very likely that neither trajectory in a pair will meet human preferences, making such queries inefficient.

[0076] In view of this, an embodiment of the present application relates to a training method for an action decision model of an intelligent agent, and proposes an efficient preference-based reinforcement learning framework that can make full use of the preference data of humans or human experts to make accurate empirical estimates of action values, thereby assisting in the learning of action value functions, i.e., action value models. This can act on action feedback models and action decision models, so that the strategy training process of the entire intelligent agent can achieve significant performance improvements.

[0077] Specifically, by utilizing the historical trajectories stored in the experience replay buffer, a non-parametric statistical model, namely the state topology graph, is constructed. An empirical action-value function can be learned using the state topology graph. The empirical action-value function can provide at least two advantages: first, based on trajectory sampling on the state topology graph, more informative trajectory pairs or fragment pairs can be constructed to assist in querying human preferences, thereby improving the efficiency of constructing trajectory pairs or fragment pairs and the efficiency of querying preference data; second, it can constrain the learning of the action-value model to optimize and obtain an action-value model with better performance, that is, regularizing the action-value function based on the neural network so that the action-value model can estimate the action value more accurately, thereby further accelerating the policy learning process, alleviating the over-estimation error and extrapolation error in the action-value function learning process, and also improving the training efficiency of the action decision model and the learning efficiency of the entire policy learning process.

[0078] The embodiments of this application are applicable to tasks involving robot collaboration, robotic arm control, and any human-related scenarios, such as vehicle collaboration, intelligent question-answering, and robot dancing. Taking the vehicle collaboration scenario as an example, with the development of modern industry and technology, the demand for multiple robotic vehicles to collaborate in transporting goods is growing. To ensure that these robotic vehicles can cooperate effectively and efficiently, the preference-based reinforcement learning framework of the embodiments of this application can be applied to train an action decision model for the robotic vehicles that conforms to human intentions and preferences.

[0079] The preference-based reinforcement learning framework of the embodiment of the present application can achieve efficient utilization of human preference data. Taking the car collaboration scenario as an example, the robot car can understand human preferences more accurately, reduce the number of times technical personnel need to intervene and adjust, reduce development costs, and improve the utilization rate of preference data; in addition, it can achieve efficient query of human preference data. By using historical data to construct a non-parametric state topology graph, it can construct more informative trajectory pairs or fragment pairs, thereby improving the construction efficiency of trajectory pairs or fragment pairs, improving the query efficiency of preference data, and improving the learning efficiency of each model; in addition, the action value model optimized by the empirical action value function can more accurately estimate the action value of each action in a specific environment, that is, improve the accuracy of the action value model, thereby optimizing the accuracy of the action feedback model and the action decision model. Taking the car collaboration scenario as an example, when multiple robot cars work together, the action decision model can provide better strategy recommendations for each robot car, and control each robot car to make more expected actions.

[0080] Taking the collaborative vehicle scenario as an example, in the case of multi-robot collaborative transport tasks, the preference-based reinforcement learning framework of the present application embodiment can help the robotic vehicles more accurately meet human needs and preferences. For example, when multiple robotic vehicles need to collaborate to move a cargo of a specific shape and weight, the action decision model can use previously collected preference data to provide the robotic vehicles with optimal movement, cooperation, and path planning strategies, ensuring the safe, fast, and efficient movement of the cargo.

[0081] Furthermore, the preference-based reinforcement learning framework of the embodiment of the present application can also be combined with large models to promote each other. In other words, the use of large amounts of data and sophisticated models can also promote the training of action decision models. For example, the preference-based reinforcement learning framework of the embodiment of the present application can promote each other with large-scale language models (LLM). The preference-based reinforcement learning technology helps the fine-tuning process in large-scale language models, and using large-scale language models as preference models to label trajectories can improve the ability of intelligent agents to solve complex control tasks.

[0082] The following describes the system architecture of the embodiment of the present application.

[0083] FIG2 is a schematic diagram of an implementation environment of a method for training an action decision model of an intelligent agent provided by an embodiment of the present application. Referring to FIG2 , the implementation environment includes: an intelligent agent 201 , an intelligent agent control system 202 , an environment 203 and a training server 204 .

[0084] Agent 201 refers to any intelligent entity, that is, a software or hardware entity capable of autonomous activity. Agent 201 resides in environment 203, capable of independent thought and continuous autonomous functioning, and capable of performing one or more actions to interact with environment 203. The actions performed by agent 201 may affect environment 203, which, influenced by the agent's actions, may change its state, i.e., complete a state transition. Agent 201 can also be considered a computational entity possessing characteristics such as residency, responsiveness, sociality, and initiative. It can be either hardware (e.g., a robot, robotic arm, or robotic car) or software (e.g., a question-answering robot, game AI, or chess AI).

[0085] The intelligent agent control system 202 refers to a system or algorithm for controlling the actions of the intelligent agent 201. The intelligent agent control system 202 contains at least an action decision model, which is used to decide the action that the intelligent agent 201 should perform in a given state. Since the intelligent agent 201 can perform multiple actions in a given state, the action decision model can determine which action to perform at the current moment and in a given state to obtain the optimal feedback signal. Therefore, the accuracy of the action decision model determines the accuracy and intelligence of the action of the intelligent agent 201, and even determines whether the intelligent agent 201 can complete the given task. In addition to the action decision model, the intelligent agent control system 202 can also be equipped with other functional modules such as an operating system, a voice interaction module, a graphic interaction module, a path planning module, and an intelligent agent navigation module to control the intelligent agent 201 to achieve richer and more diverse interaction or display functions. The embodiments of the present application do not specifically limit this.

[0086] In some embodiments, the intelligent body control system 202 is a control module built into the intelligent body 201, that is, the intelligent body 201 and the intelligent body control system 202 are coupled in the same entity, or, the intelligent body control system 202 and the intelligent body 201 are two independent entities. For example, the intelligent body control system 202 is an independent control terminal, and the control terminal can encapsulate the decided action as a control signal and send it to the intelligent body 201, thereby remotely controlling the intelligent body 201 to perform this action based on the control signal. At this time, at least a signal receiver needs to be installed on the intelligent body 201. The embodiment of the present application does not specifically limit whether the intelligent body 201 and the intelligent body control system 202 are inherited in the same physical entity.

[0087] Environment 203 refers to the space or scene in which agent 201 resides. Environment 203 can be a portion of the real world or a portion of a virtual world. For example, if agent 201 is a robot, environment 203 can be the three-dimensional space where the robot moves while performing a task. For another example, if agent 201 is a game AI, environment 203 can be the virtual game scene in which the game AI operates during a game.

[0088] Under the control of the agent control system 202, the agent 201 can interact with the environment 203. For example, after the agent 201 observes a certain state at the current moment, the agent control system 202 calls the action decision model to determine the target action to be executed from multiple candidate actions, and controls the agent 201 to execute the target action. After the agent 201 executes the target action, the environment 203 will be affected by the target action and update the state at the next moment, completing the state transition from the current moment to the next moment. At the same time, the environment 203 will also return a feedback signal (or reward signal) to the agent 201.

[0089] The training server 204 refers to a computer device used to train the action decision model of the intelligent agent. Using the training method of the action decision model of the intelligent agent in the embodiment of the present application, under the preference-based reinforcement learning framework, the training server 204 can fully learn and understand the human preference data on the basis of the constructed state topology map, and finally train an action decision model with better accuracy and better performance under the supervision of the preference data, so that the action decision model can make decisions for the intelligent agent that are more in line with human intentions or expectations, and improve human satisfaction with the action of the intelligent agent. The training process of the action decision model can be completed locally, such as local offline training of the training server 204, or it can be completed in the cloud, such as distributed training by multiple servers to speed up training efficiency. The embodiment of the present application does not specifically limit this.

[0090] The agent 201, the agent control system 202, the environment 203 and the training server 204 can be directly or indirectly connected through wired or wireless communication, which is not limited in this application.

[0091] The intelligent agent 201 can be a man-made machine device such as a robot, a robotic arm, a robotic car, a drone, an unmanned vehicle, or an intelligent terminal such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to these.

[0092] The intelligent body control system 202 can be a control module integrated into the intelligent body 201, or it can be a control terminal independent of the intelligent body 201. The control terminal includes but is not limited to smart phones, tablet computers, laptops, desktop computers, smart speakers, smart watches, etc.

[0093] Environment 203 can be a part of the real world, such as the activity area of ​​a man-made machine device, or it can be a virtual world, virtual scene, virtual environment, simulation environment, etc. provided by a computer device (such as a game server, simulation device, electronic device, etc.). The embodiments of the present application do not specifically limit this.

[0094] The training server 204 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0095] To facilitate understanding, the following example illustrates the interaction between the agent and the environment using a collaborative transport scenario involving multiple robotic carts. In this collaborative transport scenario, the agents refer to the multiple robotic carts. The agent control system refers to the robotic cart's action decision-making model, which can be integrated into the robotic cart's control chip or a standalone control terminal. The environment refers to the space or area in which the robotic carts operate when transporting goods.

[0096] By using the training method of the action decision model of the intelligent agent in the embodiment of the present application, under the preference-based reinforcement learning framework, the training server can train an action decision model with better accuracy and better performance. The training process can be completed locally, such as local offline training of the training server, or it can be completed in the cloud, such as distributed training by multiple servers to speed up training efficiency. The embodiment of the present application does not specifically limit this.

[0097] In some embodiments, after training is completed, the training server sends the parameter set of the action decision model to each robot car, so that each robot car decides and executes the action under the control of its own built-in action decision model. Multiple robot cars collaborate to move the goods from the starting point to the end point to complete the cargo transportation task.

[0098] In other embodiments, after training is completed, the training server sends the parameter set of the action decision model to the master control terminal of these robot carts. The master control terminal is responsible for scheduling the actions of each robot cart. Therefore, the master control terminal needs to decide the actions of each robot cart under the control of the action decision model, and send the control signal of each action decided to the corresponding robot cart respectively, so as to realize the macro scheduling of multiple robot carts by the same master control terminal, and control multiple robot carts to collaborate to move the goods from the starting point to the end point to complete the cargo transportation task. The embodiments of the present application do not specifically limit this.

[0099] The following describes the basic process of the training method of the action decision model of the intelligent agent in the embodiment of the present application.

[0100] FIG3 is a flow chart of a method for training an action decision model of an intelligent agent provided in an embodiment of the present application. Referring to FIG3 , this embodiment is executed by a computer device, which may be the training server 204 in the aforementioned implementation environment, or another device for training an action decision model. This embodiment includes the following steps:

[0101] 301. The computer device constructs a state topology diagram based on multiple historical trajectories of the intelligent agent. Each historical trajectory contains multiple actions, and each action is used to control the transition between different states. Each node in the state topology diagram indicates a state, and each directed edge connecting a pair of nodes indicates an action.

[0102] The historical trajectory referred to in the embodiments of this application refers to a series of actions performed by an agent in a continuous period of time in the past, arranged in chronological order. When the agent performs a task in a continuous period of time in the past, it continuously interacts with the environment from the initial action until the task is completed or the number of actions reaches a set length, forming a complete action sequence called a historical trajectory.

[0103] For example, when multiple robot carts work together to transport goods, they start from the starting point and stop when the goods are delivered or the number of walking steps reaches a preset number. The action sequence composed of the continuous actions performed by each robot cart during the time period from starting to stopping is called a historical trajectory of this robot cart.

[0104] For example, when a game AI performs a game level-breaking task, it interacts with the game program from the start of the game until the number of action steps reaches a preset number and the game fails to break the level, or the game succeeds within the preset number of steps. The action sequence composed of the continuous actions performed by the game AI from the start of the game to the success or failure of the level-breaking is called a historical trajectory of this game AI.

[0105] For any historical trajectory, it contains multiple actions that the agent has performed in the past period of time. Each action is used to control the transition from the state at a certain moment to the state at the next moment. Therefore, the historical trajectory can also be called the historical action sequence of the agent.

[0106] The state topology graph involved in the embodiments of this application refers to a graph structured data with states as nodes, which can represent the topological relationships between states and transitions between different states. The graph structured data is defined by a node set and a directed edge set. The node set is the set of all nodes in the state topology graph, each node indicating a state, and the directed edge set is the set of all directed edges in the state topology graph, each directed edge indicating an action.

[0107] Since state transitions have a timeline order, the edge connecting two nodes is directional (i.e., a directed edge), and each directed edge connects a pair of nodes. Due to its directionality (or directivity), the pair of nodes connected by a directed edge can be divided into a starting node and an arrival node. The starting node refers to the node starting from the directed edge, and the arrival node refers to the node that the directed edge ultimately reaches or points to. From a state perspective, a directed edge indicates an action that will result in a state transition between two different states. From the timeline, the state with an earlier timestamp is the starting state, and the state with a later timestamp is the arrival state. Therefore, the starting node of a directed edge indicates the starting state, and the arrival node of a directed edge indicates the arrival state.

[0108] The state topology graph may be stored in a computer device using one or more data structures, such as a hash table, an array, a dictionary, a key-value pair, a queue, etc. The data structure of the state topology graph is not specifically limited here.

[0109] In some embodiments, the computer device collects multiple historical trajectories of a given type of intelligent agent. These historical trajectories may belong to the same or different past time periods. For example, for a robot car, the historical trajectories of different robot cars transporting goods in the past hour may be collected. For another example, for a game AI, the historical trajectories of the same game AI in multiple past historical games may be collected.

[0110] When collecting historical trajectories, historical trajectories can be randomly extracted from a large number of historical trajectories to ensure randomness. Alternatively, a trajectory length interval can be preset, and multiple historical trajectories can be randomly extracted from a trajectory set consisting of historical trajectories that meet the trajectory length interval, so that the historical trajectories are not too long or too short, thereby eliminating some low-quality trajectory samples based on the trajectory length interval. Alternatively, multiple historical trajectories can be simulated using simulation software, which saves collection costs and improves collection efficiency. The embodiments of this application do not specifically limit the method for collecting historical trajectories.

[0111] In some embodiments, after collecting multiple historical traces, all states observed during the construction phase and all actions executed can be obtained based on these multiple historical traces. This allows the nodes of a state topology graph to be constructed based on the observed states, and the directed edges in the state topology graph to be constructed based on the executed actions, ultimately constructing a state topology graph. That is, the node set of the state topology graph reflects all observed states, and the directed edge set reflects all executed actions.

[0112] In step 301 above, a non-parametric state topology graph is constructed based on the historical trajectory. This state topology graph can fully reflect the empirical distribution of the agent's actions, providing more information, higher information utilization of the historical trajectory, and high efficiency in constructing the state topology graph. In addition, the state topology graph also supports convenient dynamic updates. For example, once a new state is observed, it is only necessary to add a new node to the node set of the state topology graph. For example, once a new action is observed that leads to a state transition that has never occurred before, it is only necessary to add a new directed edge to the directed edge set of the state topology graph.

[0113] 302. The computer device trains an action feedback model of the agent based on the state topology graph, where the action feedback model is used to provide feedback signals of the environment in which the agent resides regarding the actions performed by the agent.

[0114] Because the agent resides in the environment and interacts with the environment through action, the action feedback model involved in the embodiment of the present application is used to calculate or estimate the feedback signal of the action performed by the agent by the environment. Usually, after the agent performs a certain action in a certain state, a feedback signal can be generated by the action feedback model. The feedback signal can be implemented as an estimated feedback value, that is, the state and action are input into the action feedback model, and an estimated feedback value is output. The value of the estimated feedback value characterizes the reward degree of the environment for the action. The larger the estimated feedback value is, the higher the reward degree is, and the smaller the estimated feedback value is, the lower the reward degree is (may even have no reaction or bring punishment). Thus, the action feedback model is also referred to as a reward model, and the feedback signal is also referred to as a reward signal. What the action feedback model embodies is the reward function of reinforcement learning, and the action feedback model can be a reward neural network or other parameter models.

[0115] In some embodiments, based on the state topology map constructed in step 301, random sampling or non-random sampling can be performed from the state topology map to obtain multiple pairs of sampling trajectories, each pair of sampling trajectories includes a pair of sampling trajectories of equal length. It should be noted that, in order to facilitate data storage, all pairs of sampling trajectories can be required to be of equal length, for example, each sampling trajectory in each pair of sampling trajectories is controlled to be of equal length, for example, the length of all sampling trajectories is 10, so as to improve the memory access efficiency of the sampling trajectories; or, only the two sampling trajectories in each pair of sampling trajectories are required to be of equal length, but not all pairs of sampling trajectories are required to be of equal length, for example, the length of the two sampling trajectories in a certain pair of sampling trajectories is 10, but the length of the two sampling trajectories in another pair of sampling trajectories is 15. This embodiment of the present application does not specifically limit this. Then, the multiple pairs of sampling trajectories can be directly presented to the technician for annotation, and the annotation results of the multiple pairs of sampling trajectories are collected. These annotation results are used to indicate the preference degree of different sampling trajectories in each pair of sampling trajectories relative to the task performed by the intelligent agent. To put it another way, sampling trajectories are presented to technicians in pairs for annotation. In each pair, technicians are asked to note which sampling trajectory best matches their preference (or expectation). This way, each pair is divided into a positive sample trajectory and a negative sample trajectory. The positive sample trajectory's indicator annotation result is the sampling trajectory that matches the preference, while the negative sample trajectory's indicator annotation result is the sampling trajectory that does not match the preference. Therefore, the annotation results reflect the degree of preference for each of the two sampling trajectories in each pair and can also be called human preference data for the sampling trajectories. This method of collecting preference data based on trajectory pairs is simple, requires minimal annotation, and is highly efficient.

[0116] In other embodiments, based on the state topology map constructed in step 301, random sampling or non-random sampling can be performed from the state topology map to obtain multiple sampling trajectories. However, it is not necessary to pair the trajectories according to whether the lengths are equal. Instead, each sampling trajectory is first cut into multiple sampling segments by interception, cutting, etc., and then, from all the sampling segments obtained after cutting each sampling trajectory, the segments are paired according to whether the lengths are equal, thereby obtaining multiple pairs of sampling segments. Each pair of sampling segments includes a pair of sampling segments of equal length. It should be noted that, in order to facilitate data storage, it is possible to require that all paired sampling segments have equal lengths. For example, each sampling trajectory in each pair of sampling segments is controlled to have equal lengths, for example, the lengths of all sampling segments are 5, so as to improve the memory access efficiency of the sampling segments. Alternatively, it is also possible to require that only the two sampling segments in each pair of sampling segments have equal lengths, but not all paired sampling segments have equal lengths. For example, the lengths of the two sampling segments in a pair of sampling segments are both 3, but the lengths of the two sampling segments in another pair of sampling segments are both 5. This embodiment of the present application is not specifically limited to this. Next, multiple pairs of sampling segments can be directly presented to technicians for labeling, and the labeling results of multiple pairs of sampling segments can be collected. These labeling results are used to indicate the degree of preference of different sampling segments in each pair of sampling segments relative to the task performed by the intelligent agent. In other words, the sampling segments will be presented to technicians in pairs for labeling. The technicians need to label which sampling segment is more in line with the preference (or expectation) in each pair of sampling segments. In this way, each pair of sampling segments will be divided into a positive sample segment and a negative sample segment. The positive sample segment index annotation result is the sampling segment that meets the preference, and the negative sample segment index annotation result is the sampling segment that does not meet the preference. Therefore, the labeling result reflects the degree of human preference for each of the two sampling segments in each pair of sampling segments, and can also be called human preference data for the sampling segments. This method of collecting preference data based on segment pairs has a higher data utilization rate for the sampling trajectory, can produce richer sample data, and can obtain preference data with more information.

[0117] Furthermore, after collecting the labeled results, i.e., preference data, the preference data can be used to guide the training process of the action feedback model. Since the preference data can reflect human expectations or desires for the behavior of the intelligent agent, using the preference data as a supervisory signal to supervise the learning of the action feedback model can enable the action feedback model to learn and restore the underlying reward function, so that the action feedback model can calculate a more accurate estimated feedback value, thereby improving the accuracy of the action feedback model.

[0118] 303. The computer device trains an action value model of the agent based on the state topology diagram and the action feedback model. The action value model is used to provide an estimated action value for the action performed by the agent. The estimated action value indicates a metric value for measuring the impact of the action performed by the agent on the environment.

[0119] The action value model involved in the embodiment of the present application is used to evaluate the value of the impact of the action performed by the intelligent agent on the environment. The evaluation result can measure the degree of pros and cons of the impact of the action on the completion of the task. The measurement index of the evaluation result is called the action value, and the action value calculated by the action value model is actually an estimate or estimation of the real action value by the action value model, so it is called the estimated action value. For example, for a given state and the action performed, the state and action are input into the action value model, and an estimated action value is output. The value of the estimated action value represents whether the execution of the corresponding action in the given state helps the intelligent agent to complete the task. The larger the estimated action value, the higher the value of executing the corresponding action, and the more helpful it is to complete the task. The smaller the estimated action value, the smaller the value of executing the corresponding action, and the less helpful it is to complete the task. The action value model reflects the action value function of reinforcement learning. The action value model can be an action value neural network or other parameter model.

[0120] In some embodiments, the state topology map constructed in step 301 can reflect the empirical distribution of each action in the historical trajectory, thereby guiding the training of the action value model from the perspective of statistical experience, while the action feedback model trained in step 302 guides the training of the action value model from the perspective of the feedback signal of the environment to the action (the greater the reward of the environment to the action, the higher the corresponding action value should be). Therefore, combining the state topology map and the action feedback model can constrain the training process of the action value model, thereby obtaining an action value model with better accuracy and better performance, making the action value model more accurate in calculating the estimated action value and in line with human intentions or preferences to a certain extent.

[0121] 304. The computer device trains an action decision model of the intelligent agent based on the action value model. The action decision model is used to decide the action that the intelligent agent should perform in a given state.

[0122] The action decision model involved in the embodiment of the present application is used to decide what action the intelligent agent should perform in a given state, thereby assisting the intelligent agent in completing the action decision, and then controlling the execution action of the intelligent agent under the condition of reasonable path planning. For example, given an observed state, the state is input into the action decision model, and the execution probability of each of multiple candidate actions is output, so that the target action to be finally executed can be selected from the multiple candidate actions. Among them, the value of the execution probability of each candidate action represents the possibility of the intelligent agent executing the corresponding candidate action. The greater the execution probability, the greater the possibility of the intelligent agent executing the corresponding candidate action, and the smaller the execution probability, the smaller the possibility of the intelligent agent executing the corresponding candidate action. In some embodiments, the candidate action with the highest execution probability can be selected as the target action, or a random selection can be made from the top N candidate actions with the highest execution probability as the target action, or sampling can be performed on the probability distribution of the candidate actions according to the execution probability, and the target action is determined by sampling, so that the sampling process obeys the probability distribution. The embodiment of the present application does not specifically limit this. The action decision model reflects the policy function of reinforcement learning. The action decision model can be a policy neural network or other parameter model.

[0123] In some embodiments, since the action value model trained in step 303 can provide an estimated action value for each action under a given state, and different target actions may be decided under the same state when the parameters of the action decision model are different, the estimated action value of the target action decided by the action decision model is determined by the action value model, which can reflect the accuracy of the parameters of the action decision model itself. Therefore, the action value model can assist in completing the training optimization of the action decision model. Since a better-performing action value model has been obtained in step 303, the action decision model finally optimized will also have higher accuracy and better performance.

[0124] In other embodiments, the action value model and the action decision model are trained and optimized in a collaborative manner, that is, during each iteration of the iterative training, the action value model and the action decision model are updated once according to the state topology diagram and the action feedback model. There is no order of priority in the optimization of the action value model and the action decision model. The action value model can be updated first and then the action decision model, or the action decision model can be updated first and then the action value model, or the action value model and the action decision model can be updated simultaneously. The above optimization process is iteratively performed until the action decision model meets the decision optimization stopping condition, at which time the trained action decision model is obtained. The decision optimization stopping condition can be that the loss function value tends to converge, or the number of iteration steps reaches the set number of steps, etc. The decision optimization stopping condition is not specifically limited here. In the collaborative training optimization process, the action value model and the action decision model are updated once in each iteration, so that the action value model and the action decision model can guide each other, thereby further improving the accuracy of the action decision model finally optimized.

[0125] All of the above technical solutions can be combined in any way to form the embodiments of the present disclosure, and will not be described in detail here.

[0126] In the previous embodiment, the basic process of the training method for the action decision model of an intelligent agent was introduced. That is, by constructing a state topology graph with a larger amount of information, a more accurate action feedback model was trained, which in turn led to a more accurate action value model, and ultimately a more accurate action decision model. In the embodiment of this application, the detailed process of the training method for the action decision model of an intelligent agent will be described.

[0127] FIG4 is a flow chart of a method for training an action decision model of an intelligent agent provided in an embodiment of the present application. As shown in FIG4 , this embodiment is executed by a computer device, which may be the training server 204 in the aforementioned implementation environment, or another device for training action decision models. Taking the computer device as the training server as an example, this embodiment includes the following steps:

[0128] 401. The training server constructs a state topology graph based on multiple historical trajectories of the intelligent agent, where each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action.

[0129] In some embodiments, the training server first initializes the action decision model of a given type of intelligent agent, then uses the initialized action decision model to control the interaction between the intelligent agent and the environment, collects multiple candidate historical trajectories of the intelligent agent, and then screens the multiple candidate historical trajectories to obtain multiple historical trajectories. Each historical trajectory contains multiple actions, and each action is used to control the transition between different states. These historical trajectories can belong to the same or different past time periods. For example, for a robot car, the historical trajectories of different robot cars transporting goods in the past 1 hour can be collected. For example, for a game AI, the historical trajectories of the same character of the game AI in multiple historical games in the past can be collected.

[0130] When screening historical trajectories from candidate historical trajectories, historical trajectories can be randomly selected from a large number of candidate historical trajectories to ensure the randomness of the historical trajectories. Alternatively, a trajectory length interval can be preset and multiple historical trajectories can be randomly selected from the trajectory set consisting of candidate historical trajectories that meet the trajectory length interval, so that the historical trajectories are not too long or too short, thereby eliminating some low-quality candidate historical trajectories based on the trajectory length interval.

[0131] In other embodiments, the initialized action decision model is used to simulate the interaction between the intelligent agent and the environment in the simulation software, thereby simulating and calculating multiple historical trajectories. This saves the cost of collecting historical trajectories and improves the efficiency of collecting historical trajectories. The embodiments of the present application do not specifically limit the method of collecting historical trajectories.

[0132] In an exemplary scenario, an example is given in conjunction with FIG5 , which is a training framework diagram of an action decision model provided in an embodiment of the present application. As shown in FIG5 , the action decision model is used as the strategy neural network. The action feedback model is a reward neural network The action value model is the action value neural network Q θ For example. At the beginning of the algorithm, initialize the policy neural network Reward Neural Network and action-value neural network Q θ Then, based on the reward neural network and action-value neural network Q θ , using policy neural network To control the interaction between the agent and the environment, multiple historical trajectories are obtained through any of the above collection methods. In some embodiments, when obtaining historical trajectories, not only a series of actions performed by the agent are obtained, but also a series of states brought about by these actions are recorded, and a reward neural network is used to calculate the state of the agent. The estimated feedback value is calculated for each state transition. These collected data will be stored in the constructed state topology map when the state topology map is constructed, which can further increase the amount of information contained in the state topology map.

[0133] In some embodiments, after obtaining multiple historical trajectories, all states observed during the construction phase and all actions performed can be obtained based on the multiple historical trajectories, so that the nodes of the state topology graph can be constructed based on the observed states, and the directed edges in the state topology graph can be constructed based on the actions performed, and finally a state topology graph can be constructed. That is, the node set of the state topology graph reflects all observed states, and the directed edge set reflects all actions performed. Exemplarily, the above-mentioned multiple historical trajectories are stored in the experience replay buffer of the training server, and a dynamic, directed state topology graph is constructed using the various historical trajectories in the experience replay buffer.

[0134] The following is an example of a possible method for constructing a state topology graph, which includes steps A1-A3:

[0135] A1. The training server obtains the initialized state topology map.

[0136] In some embodiments, during the state topology graph construction phase, an initialized state topology graph is first obtained. In some examples, both the node set and the directed edge set of the initialized state topology graph are empty sets. For example, the state topology graph G is represented as G = (V, E), where both the node set V and the directed edge set E are empty sets.

[0137] In an exemplary scenario, the node set V is defined as Define the directed edge set E as

[0138] Since the nodes in the node set V indicate state s, the node set needs to record the relevant information of the state s indicated by each node. In some examples, the node set V is implemented as an array object, and each node in the array object records at least the node number, the state s indicated by the node, and the experience action value Experience Action Value This will be described in detail in step B4 below and will not be repeated here.

[0139] Since the directed edges in the directed edge set E indicate action a, the agent's execution of action a will result in a state transition process from state s to another state s', that is, each directed edge represents a transition from state s to another state s' through action a. Therefore, the directed edge set needs to record the relevant information of each action a and the state transition process it indicates. In some examples, the directed edge set is implemented as a dictionary object, and each directed edge in the dictionary object records at least the action a associated with this directed edge, the estimated feedback value, and the action a associated with this directed edge. Number of visits N(s,a,s′), and experience action value Estimated feedback value and the experience action value All of this will be explained in detail in step B4 below and will not be repeated here. The number of visits N(s,a,s′) indicates the query frequency for this directed edge during the training phase. It should be noted that the estimated feedback value of each directed edge is provided by the action feedback model. Therefore, every time the parameter set of the action feedback model is updated once, the estimated feedback values ​​of all directed edges in the directed edge set will also be updated. For detailed process, please refer to step 403 below. In addition, the number of visits to each directed edge will be dynamically updated in real time during the training phase as the query frequency increases. For detailed process, please refer to step A2 below. In addition, the empirical action value of each directed edge will also be updated when the action value update condition is met. For detailed process, please refer to step B4 below.

[0140] In some embodiments, each node in the node set V generates a unique state hash value for the state s indicated by the node through a hash function. Similarly, each directed edge in the directed edge set E also generates a unique action hash value for the action a indicated by the directed edge through a hash function. In this way, a key-value data structure that is easy to access and query can be generated. Taking the action indicated by the directed edge as an example, the corresponding key name Key can be the action hash value, and the corresponding key value Value can be the state hash value of the arrival state of the action. This ensures that the time complexity of the query process is This can improve the query efficiency of the state topology graph.

[0141] It should also be noted that for each node in the node set V, a supported action set can also be maintained. The supported action set represents all actions taken by the agent in state s. Its role is to assist in updating the information of the state topology graph itself, which will be explained in detail in step B4 below.

[0142] A2. For any action in any historical trajectory, the training server queries the state topology graph for the directed edge indicating the action.

[0143] In some embodiments, when constructing a state topology graph using multiple historical traces, each historical trace constructs nodes and directed edges in the state topology graph in a similar manner. Therefore, only one of the multiple historical traces is used as an example for description. Because the historical trace contains multiple actions that guide state transitions in chronological order, for each action in the historical trace, a query is performed within the directed edge set of the state topology graph to determine whether a directed edge indicating the action exists. If a directed edge indicating the action is found, the operation in step A3 is performed. If no directed edge indicating the action is found, the operation in step A4 below is performed.

[0144] A3: In response to the training server querying the directed edge indicating the action, updating the number of visits associated with the directed edge, where the number of visits indicates the query frequency of the directed edge.

[0145] In step A3, for any action a in the historical trajectory, if a directed edge indicating that action is found in the directed edge set, this indicates that the directed edge already exists. Therefore, there is no need to add a new node or directed edge. Instead, only the access count N(s,a,s′) recorded on the directed edge needs to be updated. In some examples, the access count N(s,a,s′) is incremented by 1. In other words, the value obtained by adding 1 to the original value of the access count N(s,a,s′) is assigned to the access count N(s,a,s′), i.e., N(s,a,s′)←N(s,a,s′)+1.

[0146] A4. In response to the training server failing to find a directed edge indicating the action, determining the starting state and arrival state associated with the action, querying the starting node indicating the starting state and the arrival node indicating the arrival state in the state topology graph; if the starting node indicating the starting state is not found, adding a starting node indicating the starting state; if the arrival node indicating the arrival state is not found, adding an arrival node indicating the arrival state; and adding a directed edge from the starting node to the arrival node.

[0147] In this step A4, for any action a in the historical trajectory, if no directed edge indicating the action is found in the directed edge set, it means that there is no directed edge indicating this action in the state topology graph, which means that the experience replay buffer has not observed the situation of executing this action from the starting state. Therefore, after determining the starting state and the arrival state of this action, the starting node indicating the starting state and the arrival node indicating the arrival state are further queried in the node set of the state topology graph. If the starting node indicating the starting state is not found, a starting node indicating the starting state can be added to the node set. If the arrival node indicating the arrival state is not found, an arrival node indicating the arrival state can be added to the node set, and a directed edge from the starting node to the arrival node is added to the directed edge set.

[0148] In an exemplary scenario, the four-tuple is defined For transfer data, it can indicate the state transition from the starting state s to the arrival state s' and its related information. The transfer data is stored in the experience replay buffer (Replay Buffer). When a new transfer data is observed When , for the corresponding newly added node, its experience action value can be initialized For the corresponding newly added directed edges, their empirical action values ​​can be initialized And initialize its access times N(s,a,s′)=1.

[0149] In steps A1-A4 above, by iteratively traversing each action in each historical trajectory, the node set and directed edge set of the state topology graph can be continuously populated. After traversing all actions in all historical trajectories, a completed state topology graph will be obtained. In other embodiments, all nodes can be constructed at once according to the full set of states in the historical trajectory, and all directed edges can be constructed according to the full set of actions in the historical trajectory. The embodiments of this application do not specifically limit the method for constructing the state topology graph.

[0150] In this step 401, a non-parametric state topology diagram is constructed based on the historical trajectory. This state topology diagram can fully reflect the empirical distribution of the action of the intelligent agent, brings more information, has a higher information utilization rate for the historical trajectory, and the construction efficiency of the state topology diagram is also high. In addition, the state topology diagram also supports dynamic updates very conveniently. For example, once a new state is observed, it is only necessary to add a new node to the node set of the state topology diagram. For example, once a new action is observed that leads to a state transition that has never occurred, it is only necessary to add a directed edge to the directed edge set of the state topology diagram. Or, if a certain recorded state transition occurs again, it is only necessary to update the number of visits recorded on its directed edge. In this way, as the historical trajectory continues to accumulate, the state topology diagram will become more and more perfect, realizing the adaptive update of the state topology diagram.

[0151] 402. The training server trains an action feedback model of the agent based on the state topology graph, where the action feedback model is used to provide feedback signals of the environment in which the agent resides on the actions performed by the agent.

[0152] In some embodiments, based on the state topology map constructed in step 401, multiple pairs of sampling trajectories can be derived from the state topology map and presented to a technician for preference labeling to obtain the labeling results of each pair of sampling trajectories. The labeling results of each pair of sampling trajectories are collectively referred to as preference data. In other embodiments, since the labeling process for each pair of sampling trajectories is relatively time-consuming and laborious for the technician, each pair of sampling trajectories can be input into a large model, and the large model outputs the labeling results of each pair of sampling trajectories. The labeling results provided by the large model can also reflect certain preference tendencies to a certain extent. Therefore, the labeling results of each pair of sampling trajectories are also called preference data. The embodiments of the present application do not specifically limit whether the labeling results are labeled by the technician or by the large model.

[0153] Furthermore, after collecting the labeled results, i.e., preference data, the preference data can be used to guide the training process of the action feedback model. Since the preference data can reflect the expectations or anticipations of humans (or large models) for the behavior of the intelligent agent, the preference data can be used as a supervisory signal to iteratively train the action feedback model in a supervised learning manner to update the parameter set of the action feedback model and obtain a trained action feedback model. This process is also called the update process of the action feedback model. Since the preference data is introduced as a supervisory signal, the action feedback model can learn and restore the potential reward function, so that the action feedback model can calculate a more accurate estimated feedback value, thereby improving the accuracy of the action feedback model.

[0154] The following is an example of a possible training method for the motion feedback model, which includes steps B1-B4:

[0155] B1. The training server performs trajectory sampling based on the state topology graph to obtain multiple pairs of sampling trajectories, each pair of sampling trajectories containing a pair of sampling trajectories of equal length.

[0156] In some embodiments, the training server may randomly sample or non-randomly sample from the state topology diagram to obtain multiple pairs of sampling trajectories, wherein each pair of sampling trajectories includes a pair of sampling trajectories of equal length. It should be noted that, in order to facilitate data storage, it is possible to require that all pairs of sampling trajectories are of equal length, for example, to control each sampling trajectory in each pair of sampling trajectories to be of equal length, for example, the length of all sampling trajectories is 10, so as to improve the memory access efficiency of the sampling trajectories; or, it is also possible to only require that the two sampling trajectories in each pair of sampling trajectories are of equal length, but it is not required that all pairs of sampling trajectories are of equal length, for example, the length of the two sampling trajectories in a pair of sampling trajectories is 10, but the length of the two sampling trajectories in another pair of sampling trajectories is 15. This embodiment of the present application does not specifically limit this.

[0157] In some embodiments, the training server sets the sampling length and ensures that the trajectory length of all sampling trajectories is equal to the sampling length, and then randomly pairs all the sampling trajectories to form multiple pairs of sampling trajectories. This can increase the randomness of the pairing process, and the trajectory length of each pair of sampling trajectories is equal to the sampling length.

[0158] In other embodiments, the training server sets the sampling length and ensures that the trajectory length of all sampling trajectories does not exceed the sampling length. Then, all sampling trajectories are paired according to the trajectory length, so that the trajectory lengths of the two sampling trajectories contained in each pair of successfully paired sampling trajectories are equal (and must not exceed the sampling length), forming multiple pairs of sampling trajectories. In this way, preference data under sampling trajectories of different lengths can be collected, thereby improving the diversity of multiple pairs of sampling trajectories. The embodiments of the present application do not specifically limit the sampling method and pairing method of multiple pairs of sampling trajectories.

[0159] Below, a possible trajectory sampling method is described by way of example. The trajectory sampling method includes the following steps B11-B13:

[0160] B11. The training server randomly samples from the node set of the state topology graph to obtain multiple sampling points.

[0161] In some embodiments, the training server first randomly samples from the node set to obtain multiple sampling points, each of which can be used as the starting point of a sampling trajectory.

[0162] B12. The training server starts sampling a trajectory along the directed edge starting from any sampling point among the multiple sampling points, and stops sampling when the trajectory length reaches the sampling length, thereby obtaining a sampling trajectory.

[0163] In some embodiments, for each sampling point, this sampling point can be used as the starting point of a sampling trajectory, and trajectory sampling can be performed step by step along the directed edge starting from the starting point. That is, along a certain directed edge starting from the starting point, the arrival node pointed to by this directed edge is found as the second node of the sampling trajectory. Then, along a certain directed edge starting from the second node, the arrival node pointed to by this directed edge is found as the third node of the sampling trajectory. The above sampling process is repeated until no extended directed edge is found on a certain arrival node, or the trajectory length reaches the sampling length. At this time, sampling is stopped to obtain a sampling trajectory.

[0164] In some embodiments, if a Key-Value data structure is constructed based on the state topology graph, then for any node in a sampling trajectory, it is possible to determine the total number of possible actions in the state indicated by the node, and then randomly select an action from all possible actions, and use the action hash value of the selected action as an index to query whether it hits the constructed Key-Value data structure. If it can hit the Key-Value data structure, the state hash value stored in the Value is taken out, so that the node number of the next node can be reversely checked based on the state hash value. Here, only the Key-Value data structure is used as an example for explanation. The training server can also improve the query efficiency based on the state topology graph through other data structures such as linked lists, dynamic arrays, adjacency matrices, etc., and the embodiments of the present application do not specifically limit this.

[0165] In some embodiments, during the trajectory sampling process, after a node in the trajectory is determined, since there may be more than one directed edge starting from the node, it can be divided into the following two cases for classification and discussion:

[0166] Case 1: If there is only one directed edge starting from the node, the training server will use the arrival node pointed by the directed edge as the next node of the node in the sampling trajectory.

[0167] If there is only one directed edge starting from the node, there is no need to select a directed edge. The arrival node pointed to by this directed edge is directly taken as the next node. Then continue to determine whether there is only one directed edge starting from the next node. Repeat the above process until no extended directed edge is found on a certain arrival node, or the trajectory length reaches the sampling length, the trajectory sampling is completed, and a sampling trajectory is obtained.

[0168] Case 2: If there are at least two directed edges starting from the node, the training server selects the directed edge with the highest or lowest empirical action value from the at least two directed edges, and uses the arrival node pointed to by the selected directed edge as the next node. The empirical action value of the selected directed edge is determined by the empirical action value function, and the empirical action value function is used to provide the empirical action value of the action performed on the intelligent agent based on the state topology graph.

[0169] If there are at least two directed edges from the node, this involves the decision process of which directed edge to follow to determine the next node in the sampling trajectory. Since an empirical action value is also recorded for each directed edge in the state topology graph, Therefore, the first directed edge with the highest empirical action value and the second directed edge with the lowest empirical action value can be determined from the at least two directed edges, and a directed edge can be randomly selected from the first and second directed edges. After selecting a directed edge from the at least two directed edges, the arrival node pointed to by the selected directed edge is used as the next node. The determination is then continued to determine whether there is only one directed edge starting from the next node. The above process is repeated until no extended directed edge is found at a certain arrival node, or the trajectory length reaches the sampling length. Trajectory sampling is completed, and a sampled trajectory is obtained.

[0170] Case 2 selects directed edges from those with the highest or lowest empirical action values, so that both actions with higher and lower empirical action values ​​can be sampled. This enriches the information content of the sampled trajectories and ensures the diversity of the sampled trajectories. This allows for greater differentiation between different sampled trajectories due to the randomness of sampling, making it easier to obtain annotation results that are easy to distinguish and learn, and reducing the probability of pairing two sample trajectories with poor sample quality, because the annotation results of such a pair of sample trajectories are of little significance and are not conducive to the training of the action feedback model.

[0171] In other embodiments, only the directed edge with the highest empirical action value may be selected from the at least two directed edges, or only the directed edge with the lowest empirical action value may be selected from the at least two directed edges, or, instead of screening according to the empirical action value, a directed edge may be randomly selected directly from the at least two directed edges. In this way, trajectory sampling can be completed based on the state topology graph and sampling randomness can be guaranteed. The embodiments of the present application do not specifically limit this.

[0172] In some other embodiments, bifurcation can also be performed starting from this node, with one sampling trajectory continuing to sample along the directed edge with the highest empirical action value, and the other sampling trajectory continuing to sample along the directed edge with the lowest empirical action value. In this way, a pair of sampling trajectories that can be successfully paired can be quickly obtained. Part of this pair of sampling trajectories overlaps, but two different sampling segments are generated from this node onwards. Such a pair of sampling trajectories has higher contrast and greater information content, which helps to improve the accuracy of the learned motion feedback model from the perspective of improving the sampling effect of the sampling trajectory.

[0173] B13. The training server pairs the multiple sampling trajectories according to the trajectory length to obtain multiple pairs of sampling trajectories.

[0174] In some embodiments, for each sampling point in step B11, a sampling trajectory starting from the sampling point is obtained through step B12. By traversing all the sampling points in step B11, multiple sampling trajectories can be obtained, but the trajectory lengths of these sampling trajectories may be different. A different sampling length can be used each time a trajectory is sampled, resulting in a trajectory length that varies due to different sampling lengths. In addition, when performing trajectory sampling, if sampling is stopped because no extended directed edge can be found at a certain arrival node, the trajectory length of the sampling trajectory obtained at this time is less than the sampling length. Therefore, even if the sampling length is the same, the trajectory length is different. In view of this, the multiple sampling trajectories finally obtained may have the same or different trajectory lengths. To ensure that the trajectory lengths of the two paired sampling trajectories are equal, the multiple sampling trajectories are paired according to the trajectory length, so that preference data under sampling trajectories of different lengths can be collected, thereby improving the diversity of each pair of sampling trajectories. The embodiment of the present application does not specifically limit the pairing method of multiple pairs of sampling trajectories. Alternatively, for simplicity, the training server can also pre-set a trajectory length, discard sampling trajectories that do not meet the trajectory length, and then randomly pair them among all remaining sampling trajectories of equal length. This can improve the efficiency of constructing multiple pairs of sampling trajectories.

[0175] In the above steps B11-B13, a sampling method of trajectory pairs is provided, which collects preference data in units of trajectory pairs, and has high data collection efficiency.

[0176] In other embodiments, after traversing all sampling points in step B11 and obtaining multiple sampling trajectories through step B12, pairing is not performed at the trajectory level. Instead, each sampling trajectory is cut into multiple sampling segments by interception, cutting, etc., and then, from all the sampling segments formed after cutting each sampling trajectory, the segments are paired according to whether their lengths are equal, thereby obtaining multiple pairs of sampling segments. Each pair of sampling segments includes a pair of sampling segments of equal length. It should be noted that, in order to facilitate data storage, it is possible to require that all paired sampling segments have equal lengths, for example, each sampling trajectory in each pair of sampling segments has equal lengths, for example, the lengths of all sampling segments are 5, so as to improve the memory access efficiency of the sampling segments; alternatively, it is possible to require only the two sampling segments in each pair of sampling segments to be equal in length, but not all paired sampling segments are required to be equal in length. For example, the lengths of the two sampling segments in a pair of sampling segments are both 3, but the lengths of the two sampling segments in another pair of sampling segments are both 5. This embodiment of the present application does not specifically limit this.

[0177] In the above embodiment, a segment pair sampling method is provided, which collects preference data in units of segment pairs, has a higher data utilization rate for the sampling trajectory, and can generate richer sample data.

[0178] B2. The training server collects the labeling results of the multiple pairs of sampling trajectories, and the labeling results indicate the preference degree of each of the two sampling trajectories in each pair of sampling trajectories relative to the task performed by the intelligent agent.

[0179] In some embodiments, after obtaining multiple pairs of sampling trajectories in step B1, the training server directly presents the multiple pairs of sampling trajectories to the technician for labeling, and collects the labeling results of the multiple pairs of sampling trajectories. These labeling results are used to indicate the preference degree of each of the two sampling trajectories in each pair of sampling trajectories relative to the task performed by the agent. In other words, the sampling trajectories will be presented to the technician in pairs for labeling. The technician needs to label which sampling trajectory is more in line with the preference (or expectation) in each pair of sampling trajectories. In this way, each pair of sampling trajectories will be divided into a positive sample trajectory and a negative sample trajectory. The positive sample trajectory is the sampling trajectory that is labeled as being in line with the preference, and the negative sample trajectory is the sampling trajectory that is labeled as not being in line with the preference. Therefore, the labeling results reflect the human preference degree for each of the two sampling trajectories in each pair of sampling trajectories. In some examples, each pair of sampling trajectories can also be input into a large model, and the large model will label each pair of trajectories. This embodiment of the present application does not specifically limit this. This preference data collection method based on trajectory pairs has a simple process, a small number of labeling times, and high collection efficiency.

[0180] In other embodiments, after obtaining multiple pairs of sampling segments as described in step B13, the multiple pairs of sampling segments can be presented to technicians for annotation, and the annotation results of the multiple pairs of sampling segments are collected. These annotation results are used to indicate the satisfaction of different sampling segments in each pair of sampling segments relative to the task performed by the agent. In other words, the sampling segments will be presented to technicians in pairs for annotation. The technicians need to mark which sampling segment is more in line with the preference (or expectation) in each pair of sampling segments. In this way, each pair of sampling segments will be divided into a positive sample segment and a negative sample segment. The positive sample segment refers to the sampling segment that is marked as meeting the preference, and the negative sample segment refers to the sampling segment that is marked as not meeting the preference. Therefore, the annotation results reflect the degree of human preference for each of the two sampling segments in each pair of sampling segments. In some examples, each pair of sampling segments can also be input into a large model, and the large model will annotate each pair of sampling segments. This embodiment of the present application does not specifically limit this. This method of collecting preference data based on segment pairs has a higher data utilization rate for the sampling trajectory, can generate richer sample data, and obtain preference data with more information.

[0181] B3. When the feedback model update condition is met, the training server trains the action feedback model based on the state topology graph and the annotation result.

[0182] In some embodiments, since labeling sampling trajectories may be a continuous, long-term process, a feedback model update condition can be pre-defined. When the feedback model update condition is met, the action feedback model is updated based on the preference data collected between the last update and the current moment. For example, the feedback model update condition can be a periodic update, such as every two days, once a week, or once every 500 pairs of sampling trajectories have been labeled. The specific content of the feedback model update condition is not specifically limited here.

[0183] When the feedback model update conditions are met, the training server uses the preference data as a supervision signal and iteratively trains the action feedback model in a supervised learning manner to update the parameter set of the action feedback model and obtain a trained action feedback model.

[0184] In some embodiments, under a preference data collection method based on trajectory pairs, when the preference data is used as a supervisory signal for supervised learning, for each pair of sampling trajectories, the labeling result indicates that the pair of sampling trajectories contains a positive sample trajectory and a negative sample trajectory. The action feedback model is used to determine the estimated feedback value of each action in the positive sample trajectory and then sum them to obtain the positive sample feedback sum value. Similarly, the estimated feedback value of each action in the negative sample trajectory is determined and then summed to obtain the negative sample feedback sum value. Then, a preference loss term is introduced into the loss function of the action feedback model to control the positive sample feedback sum value of each positive sample trajectory to be as high as possible and the negative sample feedback sum value of each negative sample trajectory to be as low as possible, thereby optimizing the parameter set of the action feedback model through the preference loss term, and iteratively training the parameter set of the action feedback model until the loss function value converges or the number of iteration steps reaches the set number of steps, then the training is stopped to complete an update of the parameter set of the action feedback model.

[0185] In other embodiments, under a preference data collection method based on segment pairs, when the preference data is used as a supervisory signal for supervised learning, for each pair of sampling segments, the labeling result indicates that the pair of sampling segments contains a positive sample segment and a negative sample segment, and the action feedback model is used to determine the estimated feedback value of each action in the positive sample segment and then sum them to obtain the positive sample feedback sum value. Similarly, the estimated feedback value of each action in the negative sample segment is determined and then summed to obtain the negative sample feedback sum value. Then, a preference loss term is introduced into the loss function of the action feedback model to control the positive sample feedback sum value of each positive sample segment to be as high as possible and the negative sample feedback sum value of each negative sample segment to be as low as possible, thereby optimizing the parameter set of the action feedback model through the preference loss term, and iteratively training the parameter set of the action feedback model until the loss function value converges or the number of iteration steps reaches the set number of steps, then the training is stopped, and an update of the parameter set of the action feedback model is completed.

[0186] In steps B1-B3, after collecting the labeling results, i.e., preference data, the preference data is used to guide the training process of the action feedback model. Since the preference data can reflect the expectations or expectations of humans (or large models) for the behavior of the intelligent agent, the action feedback model can learn and restore the potential reward function, so that the action feedback model can calculate a more accurate estimated feedback value, thereby improving the accuracy of the action feedback model.

[0187] Furthermore, sampling trajectories are collected through trajectory sampling. Samples that have not appeared in historical trajectories may appear in the sampling trajectories. This is because the state topology graph can connect the state transitions of past historical trajectories and connect the same states involved in different historical trajectories, thereby splicing new sampling trajectories during trajectory sampling. These new sampling trajectories are not obtained from a single historical trajectory, but are obtained by splicing fragments from multiple historical trajectories. They do not belong to any single historical trajectory, thus enriching the data diversity of the sampling trajectories.

[0188] In other embodiments, the idea of ​​positive and negative sample contrast learning can also be introduced to combine supervised learning and contrastive learning to achieve iterative training of the action feedback model. The embodiments of the present application do not specifically limit the training method of the action feedback model.

[0189] Steps B1 and B3 describe how to use the state topology graph to update the action feedback model. However, within the preference-based reinforcement learning framework of the present embodiment, the empirical action values ​​for the actions indicated by some or all directed edges in the state topology graph can also be updated when the action value update conditions are met. This process is called a graph update. Because updating the empirical action values ​​for the entire graph requires a high computational overhead, it's possible to update only the empirical action values ​​for actions indicated by some directed edges at a time to reduce computational overhead and improve update efficiency. This partial update is used as an example in step B4 below.

[0190] B4. When the action value update condition is met, the training server updates the empirical action value of the action indicated by some directed edges in the state topology graph according to the empirical action value function, where the empirical action value function is used to calculate the empirical action value of the action based on the state topology graph.

[0191] In some embodiments, even if only the empirical action values ​​for actions indicated by a portion of directed edges are updated each time, the update process still incurs a certain amount of computational overhead. Therefore, a preset action value update condition can be set. When the action value update condition is met, the graph update process is executed, thereby effectively reducing computational overhead. In some examples, the action value update condition can be updated periodically, such as every two days, once a week, or once every 200 transfer data collected in the empirical replay buffer. The specific content of the action value update condition is not specifically limited here.

[0192] It should be noted that both the action feedback model and the empirical action value can be updated periodically or conditionally triggered. However, if both are updated periodically, it is not required that the update cycles of the two be consistent. The two can be updated separately according to independent cycles. For example, the action feedback model is updated every two days, while the empirical action value is updated every day. The embodiments of the present application do not specifically limit this.

[0193] When the action value update conditions are met, the training server starts the graph update process of the state topology graph, that is, for each node to be updated, the supported action set Update the supported action set The empirical action value of each action in is usually recalculated and assigned based on the original value according to a preset empirical action value function to correct possible deviations from the original value.

[0194] In some embodiments, a map update process includes the following steps B41 to B43:

[0195] B41. When the action value update condition is met, the training server samples the node set in the state topology graph to obtain multiple nodes to be updated.

[0196] In some embodiments, the training server randomly samples multiple nodes to be updated from the node set, or performs probability sampling in the node set according to the latest access timestamp of the node to obtain multiple nodes to be updated. The probability sampling here means that the node with the largest latest access timestamp (that is, the closer to the current moment) has a greater probability of being sampled as the node to be updated. These nodes are nodes that have just been visited recently and may be more important nodes at the current stage of the training process. Updating the empirical action values ​​of each action in the supporting action set of these nodes can ensure a better graph update effect.

[0197] In an exemplary scenario, a probability sampling is performed in the node set according to the latest access timestamp of the node to obtain multiple nodes to be updated. The multiple nodes to be updated can constitute a subset of the node set. Then, for the subset The multiple nodes to be updated in the graph are executed in the order of the latest access timestamp from large to small, thereby completing the graph update process for all nodes to be updated. The order of the latest access timestamp from large to small can ensure that the more recently accessed nodes to be updated, the earlier the experience action value is updated, which can improve the update efficiency of the experience action value. In some embodiments, it is also possible to update the graph for the subset. The multiple nodes to be updated in are processed in the sampling order, or for a subset The multiple nodes to be updated are processed in a random order, and the embodiments of the present application do not specifically limit this.

[0198] It should be noted that, since the action value update process of different nodes to be updated is the same, the action value update process of one node to be updated is taken as an example below, and steps B42 to B44 are explained.

[0199] B42. The training server determines a supported action set for each node to be updated, where the supported action set includes a set of actions indicated by each directed edge starting from the node to be updated.

[0200] For the state topology graph, each node to be updated will have a set of supported actions In the action value update rule of each node to be updated, the maximum operator of the graph update is set to support the action set Up operation, that is, update the supported action set Instead of updating all the actions of all directed edges in the entire directed edge set, this can further reduce the computational overhead of the update process of each node to be updated and improve the efficiency of updating the action value of a single node.

[0201] In some embodiments, for each node to be updated, all directed edges starting from the node to be updated are determined, and the set of all actions indicated by these directed edges is the supported action set of the node to be updated. When updating the empirical action value on the supported action set, actions that are not included in the supported action set are not considered. This can avoid over-estimation of the empirical action value, that is, try to ensure that the empirical action value is not estimated to be too large or too small, making the estimation of the empirical action value more accurate.

[0202] B43. The training server updates the empirical action value of each action in the supported action set.

[0203] In some embodiments, for each node to be updated, the supported action set of the node to be updated is determined in step B42. Afterwards, for each action a in the supported action set, the empirical action value is recalculated and assigned according to the preset empirical action value function to correct the possible deviation of the original value.

[0204] In some embodiments, the empirical action value function can be set to The iterative update process of all actions in the supported action set is performed to quickly update the empirical action value of each action. Schematically, any iteration in the above iterative update process is described, and each iteration updates the empirical action value of an action in the supported action set, in conjunction with the following steps B43a-B42e:

[0205] B43a. For any action in the supported action set, the training server determines multiple target nodes that can be reached by executing the action starting from the node to be updated.

[0206] In some embodiments, for the supported action set For each action a in , since it is possible to reach multiple different target nodes by executing action a starting from the node to be updated s, the training server determines multiple target nodes s'∈S that can be reached by executing action a starting from the node to be updated s, where S refers to all nodes contained in the node set V. For example, assuming that executing action a (such as a straight-line action) starting from the node to be updated s may reach the target node s1 (for example, s1 indicates a falling state) or the target node s2 (for example, s2 indicates a backward state), then the multiple target nodes that may be reached by executing action a starting from the node to be updated s include s1 and s2. It should be noted that in this case, there will be a directed edge from s to s1 and a directed edge from s to s2 in the state topology graph, and these two directed edges both indicate action a.

[0207] B43b. The training server determines the empirical transition probability of the node to be updated reaching each target node through the action based on the number of visits to each directed edge from the node to be updated to each target node through the action.

[0208] In some embodiments, for each target node s' found in step B43a, a directed edge starting from the node to be updated s and reaching the target node s' can be uniquely determined from the directed edge set, and then the number of visits N(s,a,s′) of this directed edge is queried. The above operation of obtaining the number of visits is repeated for all target nodes s'∈S, and the number of visits N(s,a,s′) of all directed edges indicating action a starting from the node to be updated can be queried. The sum of all the queried number of visits N(s,a,s′) is calculated to obtain the total number of visits of all directed edges indicating action a starting from the node to be updated s. Then, the value obtained by dividing the number of visits of the current directed edge by the total number of visits is used as the empirical transition probability of the node to be updated s to this target node s'.

[0209] For example, using Characterize the empirical transition probability of the node to be updated s reaching the target node s' through action a, then the empirical transition probability is expressed as the following formula:

[0210] Among them, N(s,a,s′) represents the number of visits to the directed edge from the node to be updated s to the target node s′ through action a, ∑ s′∈S N(s,a,s′) represents the total number of visits to all directed edges indicating action a starting from the node to be updated s.

[0211] B43c. The training server determines, through the action feedback model, an estimated feedback value of executing the action starting from the node to be updated, where the estimated feedback value indicates a feedback signal that is expected to be generated when the environment executes the action on the agent.

[0212] In some embodiments, for the node to be updated s and the supported action set For each action a, the node to be updated s and action a are input into the action feedback model, and the action feedback model is used to calculate the estimated feedback value of executing the action a starting from the node to be updated s. For example, the action feedback model can be a reward neural network Input the node to be updated s and action a into the reward neural network Output estimated feedback value For action sets For each action a, an estimated feedback value can be calculated through this step B43c.

[0213] B43d. The training server updates the empirical action value of the action based on the empirical transition probability of the node to be updated reaching each target node through the action, the estimated feedback value, and the empirical action value of each target node.

[0214] In some embodiments, based on the empirical transition probability calculated for each target node s' in step B43b, and based on the estimated feedback value calculated for the action a in step B43c And the original value of the experience action value recorded for the target node s' An empirical action value can be calculated for the node s to be updated and the action a.

[0215] In some embodiments, for the node to be updated s and action a, the empirical action value of action a is The calculation method of is shown in the following formula (i.e., the empirical action value function):

[0216] in, Represents the estimated feedback value calculated for the node s to be updated and the action a, γ is a hyperparameter representing the weighted strength, Characterizes the empirical transition probability of the node to be updated s reaching the target node s' through action a.

[0217] B43e. The training server assigns the maximum value among the experience action values ​​of each action in the supported action set to the experience action value of the node to be updated.

[0218] In some embodiments, when the node to be updated s remains unchanged, the supported action set for the node to be updated s is For each action a in the above example, the empirical action value of executing action a on the node to be updated s can be calculated by steps B43aˉB43d, and the traversal of the supported action set All actions in the game can get the experience action value of each action, and the maximum value of each experience action value can be used as the new value. Assign the value to the experience action value recorded for the node to be updated s in the state topology graph. In other words, replace the experience action value recorded on the node to be updated s from the original value to the new value Iteratively execute steps B43a-B43e for all nodes to be updated to complete an iterative update of the action value.

[0219] In some embodiments, for the node to be updated s, the new value of the empirical action value is The assignment method is shown in the following formula:

[0220] In the above steps B43a-B43e, for the supported action set For each action in , by introducing the experience transition probability based on the number of visits, the experience transition probability can be used to weight the original value of the experience action value of each target node that may be reached by executing action a from the node to be updated s, and then combined with the estimated feedback value provided by the action feedback model to calculate the experience action value of the action a. Ultimately, the action set will be supported The maximum value of the newly calculated empirical action values ​​of all actions is assigned to the empirical action value of the node to be updated, thereby updating the empirical action value of the node to be updated.

[0221] In other embodiments, the state topology graph can also be configured as a graph neural network, so that the graph neural network learns the appropriate value of its empirical action value under the guidance of the estimated feedback value and the number of visits. The embodiments of the present application do not specifically limit the action value update method.

[0222] In steps B41-B43, a possible action value updating method for the empirical action value of each action in the supported action set of the node to be updated is provided, which can quickly implement the updating of the empirical action value with less computational overhead.

[0223] In other embodiments, a full graph update may be performed each time an action value is updated, so that the empirical action value of each action in the state topology graph can be updated, and the empirical distribution of each action can be restored more accurately. The embodiments of the present application do not specifically limit whether a full graph update is used for the empirical action value.

[0224] It should be noted that step B4 involves an action value update process, and the training server may not update the experience action value, which is not specifically limited in the embodiment of the present application.

[0225] 403. After the action feedback model is updated, the training server updates the estimated feedback value of each directed edge in the state topology graph, where the estimated feedback value indicates the feedback signal that is expected to be generated when the environment performs the action indicated by the directed edge on the agent.

[0226] In some embodiments, as described in step 402, the action feedback model may be updated each time a feedback model update condition is met. After each update of the action feedback model, the training server can use the updated action feedback model to re-mark the estimated feedback value recorded on each directed edge in the state topology graph. In other words, using the updated action feedback model, the estimated feedback value for each directed edge in the state topology graph is recalculated and updated. Therefore, each update of the action feedback model results in improved performance.

[0227] For example, the action feedback model is a reward neural network Each time the reward neural network After the update is completed, for any node s and action a in the state topology graph, the node s and action a are input into the updated reward neural network In the output, the estimated feedback value of executing action a from node s is output And overwrite the old value with the new value output.

[0228] In the above process, after each update of the action feedback model, the action feedback model is used to re-label the estimated feedback value maintained on each directed edge in the state topology graph, ensuring that the estimated feedback value is the most accurate new value at the current moment. This can improve and mitigate the impact of non-stationary reward functions. This is because the action feedback model is regularly updated, and the potential reward function it learns is constantly fine-tuned and optimized. Therefore, the reward function does not always tend to be stable. By promptly overwriting the old value with the new value, it can ensure that the estimated feedback value stored in the state topology graph can reflect the reward signal evaluated by the latest action feedback model, which in turn improves the accuracy of the information contained in the state topology graph.

[0229] 404. Based on the state topology graph and the action feedback model, train an action value model of the agent, wherein the action value model is used to provide an estimated action value for the action performed by the agent, and the estimated action value indicates a metric value for measuring the impact of the action performed by the agent on the environment.

[0230] In some embodiments, both the empirical action value and the estimated feedback value in the state topology graph may be updated. For the empirical action value, it is necessary to use the empirical action value function and the latest state topology graph to calculate and update it. The empirical action value function has been introduced in step B43e. It is only necessary to give the old value of the empirical action value of each node during the initialization phase. Subsequently, the empirical action value of each node relative to each action can be continuously updated using the method described in steps B43a-B43e, so that after multiple updates, the empirical action value can reflect the empirical distribution of the action value on the historical trajectory.

[0231] In some embodiments, based on the empirical action value function and the action feedback model, the action value model of the intelligent agent can be guided to be trained. This action value model is used to evaluate the impact of the actions performed by the intelligent agent on the environment. Since the empirical action value function can reflect the empirical distribution of each action in the historical trajectory, it can guide the training of the action value model from the perspective of statistical experience, while the action feedback model guides the training of the action value from the perspective of the reward signal of the environment (the greater the reward given by the environment, the higher the corresponding action value should be). Therefore, combining the empirical action value function and the action feedback model can constrain the training process of the action value model, thereby obtaining an action value model with better accuracy and better performance, making the action value model more accurate in calculating the estimated action value and in line with human intentions or preferences to a certain extent.

[0232] Below, we will combine steps C1-C4 to illustrate a possible training method for the action value model. In this training method, a constraint loss term is constructed using the empirical action value function for the loss function of the action value model to regularize the action value model, thereby alleviating the over-estimation error and extrapolation error of the action value model in the learning process of the action value function. At the same time, an action value loss term is also considered to measure the difference between the estimated action value provided by the action value model and the target action value in the optimization learning. Under the joint action of the constraint loss term and the action value loss term, it is possible to assist in training an action value model with better accuracy and performance, and to improve the generalization of the action value model. The training method of this action value model includes the following steps C1-C4:

[0233] C1. In any iteration, the training server obtains the estimated action value of each directed edge in the state topology graph through the action value model.

[0234] In some embodiments, the action value model is iteratively trained based on the empirical action value function and the action feedback model. In any iteration of the iterative training process of the action value model, an estimated action value can be calculated for each directed edge in the state topology graph using the initial action value model. This estimated action value represents the action value of the action indicated by the directed edge as estimated by the action value model. The estimated action value is different from the empirical action value. The empirical action value is an empirical value derived from the empirical distribution of the state topology graph, while the estimated action value is a predicted value calculated from the action distribution learned by the action value model.

[0235] In some embodiments, for any directed edge in the state topology graph, the starting state s indicated by the starting node of the directed edge is t and the action a indicated by the directed edge t , together with the input to the action value model Q θ In the example, we get action a t The estimated action value Q θ (s t ,a t ). Repeat the above steps to get the estimated action value of any action indicated by a directed edge.

[0236] C2. The training server obtains a constraint loss term based on the empirical action value function and the estimated action value, where the constraint loss term represents the distribution difference between the empirical distribution of the action value and the model distribution.

[0237] In some embodiments, for any directed edge in the state topology graph, a set of data associated with the directed edge can be derived from the state topology graph: the starting state s indicated by the starting node of the directed edge t , the action a indicated by the directed edget , the arrival state s indicated by the arrival node of the directed edge t+1 , the action feedback model is based on the starting state s t and action a t Calculated estimated feedback value The empirical action value function is based on the starting state s t and action a t Calculated experience action value Where t refers to the starting state s t The timestamp of t+1 is the time when the state s is reached. t+1 Therefore, for any directed edge, a set of data for this directed edge can be obtained from the state topology graph. For the sake of simplicity, this set of data is called the description data of this directed edge.

[0238] Furthermore, for any node s in the state topology graph, the set of actions indicated by all directed edges starting from the node s constitutes the supported action set of the node. Supported action sets Experience action value of all actions in Reflects the empirical distribution of actions on node s, and supports the estimated action value Q of all actions in the action set θ (s t ,a t ) reflects the action distribution learned by the model on node s. Therefore, we can support the action set for each node The difference between the experience distribution and the action distribution is used to construct the constraint loss term of the action value model. This constraint loss term only considers the supported action set for a single node. The experience action value of each action and the estimated action value Q θ (s t ,a t ), which helps to utilize the empirical action value function Let Q be the action value model θ The training process provides constraints to normalize the action value model Q θ , accelerated action value model Q θ For learning about empirical distributions.

[0239] In some embodiments, in the action-value model Q θ In each training iteration, each node in the state topology graph can be trained according to its supported action set The experience action value of each action and the estimated action value Q θ (s t ,at ) to calculate the action value error of the node, and find the mathematical expectation of the action value errors of all nodes to obtain a constraint loss term. In this way, the constraint loss term contains rich information, the stronger the ability of the constraint loss term to express, the better it can be for the action value model Q θ Provide certain constraints.

[0240] In some other embodiments, in the action-value model Q θ In each training iteration, if the state topology graph contains many nodes and directed edges, the training overhead caused by the full graph calculation is large. In this case, a part of the focus nodes can be sampled in the state topology graph according to the pre-set sampling rules, and only for each focus node, the supported action set is used. The experience action value of each action and the estimated action value Q θ (s t ,a t ) is used to calculate the action value error of the focus node, and the mathematical expectation of the action value errors of all focus nodes is obtained to obtain a constrained loss term. In this way, the computational overhead of the constrained loss term is small, and the training efficiency of the action value model is improved. The embodiment of the present application does not make specific limitations on this.

[0241] Next, we will use the method of constructing the constraint loss term by using the action value error of the focus node as an example, and combine steps C21-C24 to illustrate the calculation process of the constraint loss term:

[0242] C21. The training server determines a plurality of focus nodes from the state topology diagram, where the focus nodes indicate the states that the agent needs to focus on when implementing the task.

[0243] In some embodiments, random sampling is performed in the node set of the state topology graph to obtain multiple focus nodes, or the nodes in the node set are arranged in order according to the order of the latest access, and the multiple nodes that have been recently accessed are sampled as multiple focus nodes, or sampling is performed from the node set in order of the number of visits from high to low, and the multiple nodes with the highest number of visits are sampled as multiple focus nodes, or, the same sampling method of the node to be updated introduced in step B41 can be used to perform probabilistic sampling from the node set according to the latest access timestamp of the node to obtain multiple focus nodes. The embodiment of the present application does not specifically limit the sampling method of the focus nodes.

[0244] C22. For any of the focus nodes, the training server determines the empirical action value of each action in the supported action set of the focus node based on the empirical action value function, where the supported action set includes the set of actions indicated by each directed edge starting from the focus node.

[0245] In some embodiments, for any focus node determined in step C21, the set of actions indicated by all directed edges starting from the focus node constitutes the supported action set of the focus node, and the empirical action value function is used to calculate the supported action set of the focus node. The empirical action value of each action in the supported action set of the focus node can be calculated. For example, for a certain focus node s, determine the supported action set of the focus node s Using empirical action-value function Determine the supported action set The experienced action value of each action in .

[0246] C23. The training server determines the action value error of the focus node based on the empirical action value and the estimated action value of each action in the supported action set, where the action value error represents the difference between the empirical action value and the estimated action value of each action in the supported action set.

[0247] In some embodiments, for any focus node s determined in step C21, the supported action set of the focus node s is considered. In step C22, the supported action set can be obtained The experience action value of each action in step C1 can be obtained from the supported action set. The estimated action value of each action in . Since the support action set The empirical action values ​​of all actions in reflect the empirical distribution of actions on the focus node s, while the support action set The estimated action values ​​of all actions in reflect the action distribution learned by the model on the focus node s. Therefore, for each focus node, the supported action set The difference between the experience distribution and the action distribution is used to obtain the action value error of the current focus node.

[0248] In some embodiments, the supported action set for the focus node s is For each action in , the difference between the empirical action value and the estimated action value of this action can be calculated, and the absolute values ​​of the differences of all actions can be averaged to serve as the final action value error of the focus node s. Alternatively, the mean square error of the absolute values ​​of the differences of all actions can be calculated to serve as the final action value error of the focus node s, which is not specifically limited in this embodiment of the application.

[0249] In other embodiments, in addition to calculating the average value and the mean square error, the action value error of the focus node s may also be calculated from the level of the action value vector. This will be explained below in conjunction with steps C23a-C23c:

[0250] C23a. The training server determines the empirical action value vector of the focus node based on the empirical action value of each action in the supported action set.

[0251] In some embodiments, the supported action set for the focus node s is Action sets will be supported The empirical action values ​​of all actions in are spliced ​​into a row vector or column vector to obtain the empirical action value vector of the focus node Experience Action Value Vector Refers to the supported action set The feature vector of the empirical action values ​​of all actions in .

[0252] C23b. The training server determines the estimated action value vector of the focus node based on the estimated action value of each action in the supported action set.

[0253] In some embodiments, the supported action set for the focus node s is Action sets will be supported The estimated action values ​​of all actions in are spliced ​​into a row vector or column vector to obtain the estimated action value vector Q of the focus node θ (s,·), estimated action value vector Q θ (s,·) refers to the set of supported actions The feature vector of the estimated action values ​​of all actions in .

[0254] C23c. The training server determines the action value error based on the empirical action value vector and the estimated action value vector.

[0255] In some embodiments, the exponential normalization function softmax is used to normalize the empirical action value vector calculated in step C23a. Perform exponential normalization to obtain the normalized experience value vector, which represents the experience strategy of the action value, denoted as Similarly, the estimated action value vector Q calculated in step C23b is normalized using the exponential normalization function softmax. θ (s,·) is also exponentially normalized to obtain a normalized estimated value vector, which represents the estimated strategy of the action value model, also known as the model soft strategy, denoted by π soft(θ) (s)=SoftmaxQ θ (s,·).

[0256] In some embodiments, the normalized experience value vector and the normalized estimate vector π soft(θ)The KL distance (Kul lback-Leibler Divergence, also known as KL divergence, relative entropy) between (s) is determined as the action value error of the focus node s, denoted as In this way, the relative entropy information between the empirical strategy and the model soft strategy can be introduced into the action value error of the focus node s, thereby improving the accuracy of the action value error of a single focus node.

[0257] In other embodiments, the normalized experience value vector and the normalized estimate vector π soft(θ) (s) is determined as the action value error of the focus node s. In this way, the distance between the empirical strategy and the model soft strategy in the vector space can be considered in the action value error of the focus node s, and the accuracy of the action value error of a single focus node can be improved. The embodiment of the present application does not make specific limitations on this.

[0258] In steps C23a and C23c, a possible implementation method for calculating the action value error of the focus node s at the action value vector level is provided. This allows the action value error to measure the difference between the empirical strategy and the model soft strategy in vector space as comprehensively and accurately as possible, thereby improving the accuracy of the action value error of a single focus node. In some embodiments, the action value error can be calculated without taking the vector level into account, such as by directly calculating the average or averaged error. This is not specifically limited in the present embodiments.

[0259] C24. The training server obtains the constraint loss term based on the action value error of each focus node.

[0260] In some embodiments, for any focus node s determined in step C21, the action value error of the focus node s can be calculated according to steps C22ˉC23, and then, based on the action value errors of each focus node, the constraint loss term of this training iteration can be constructed.

[0261] In some embodiments, the mathematical expectation of the action value error of each focus node is used as the constraint loss term. In this case, the constraint loss term is expressed as the following formula:

[0262] in, represents the constraint loss term of the action-value model, θ represents the parameter set of the action-value model, represents mathematical expectation, s represents the focus node, G represents the state topology graph, and KL refers to the KL distance between two vectors. represents the experience strategy, π soft(θ) (s) represents the model soft policy.

[0263] In other embodiments, in addition to calculating the mathematical expectation of the action value error of each focus node as the constraint loss term, the average value or average error of the action value error of each focus node can also be used as the constraint loss term. The embodiments of the present application do not specifically limit this.

[0264] In steps C21-C24, a possible implementation method for constructing a constraint loss term based on the action value error of the focus node is provided. The constraint loss term constructed in this way can be used in the action value model Q θ Provide a constraint during the training process so that the empirical action value function guided by the state topology graph Can be used as an action value model Q θ The lower bound of the action value model Q is to control the action value model Q as much as possible. θ The estimated action value given is not less than the empirical action value function The experience action value provided.

[0265] Furthermore, since only a part of the supported action sets of the concerned nodes are considered in the constraint loss term In this way, there is no need to calculate the action value error for non-focus nodes, which greatly saves the computational overhead of the constraint loss term and improves the action value model Q θ The training efficiency can be improved, thereby accelerating the training process of the subsequent action decision model.

[0266] Furthermore, by using the action value model Q θ Considering the constraint loss term in the loss function can pay more attention to the empirical distribution of actions that have appeared in historical experience. Of course, it will not completely abandon the potential distribution of actions that have not been executed. It just has certain constraints on the empirical distribution, making the empirical action value the lower bound of the estimated action value, which effectively reduces the action value model Q θ Overestimation and extrapolation errors of action values.

[0267] C3. The training server obtains an action value loss term based on the action feedback model and the estimated action value. The action value loss term represents the difference between the model's estimated action value for the action and the target action value. The target action value represents the optimization target for the action value based on the action distribution.

[0268] In some embodiments, in the training iteration of the action value model, in addition to the constraint loss term constructed by step C2, its loss function also includes an action value loss term. The action value loss term is used to measure the degree of difference between the estimated action value provided by the model and the target action value, and the target action value refers to the optimization target of the action value based on the action distribution. During the calculation process, the target action value needs to use the complete action distribution in the entire state topology graph to realize the calculation, and needs to involve the action decided by the action decision model and the estimated feedback value given by the action feedback model. Therefore, when constructing the action value loss term, it is necessary to use the action feedback model and the action decision model. The action feedback model here uses the parameter set obtained after the training optimization in step 402, and the action decision model uses the parameter set obtained after the latest training iteration at the current moment.

[0269] In some embodiments, for time step t, the transition data at time step t is found from the state topology graph. The target action value is constructed using the various transfer data recorded in the state topology graph, and then the action value loss term is constructed using the estimated action value and the target action value.

[0270] Below, in combination with steps C31-C35, an example of a possible construction method of the action value loss term is given. In this construction method, the action value loss term is implemented as a soft Bellman residual as follows:

[0271] C31. The training server randomly samples the directed edge set of the state topology graph to obtain multiple sampling edges. For any sampling edge, the training server determines the sampling state indicated by the starting node of the sampling edge and the sampling action indicated by the sampling edge.

[0272] In some embodiments, in each training iteration of the action value model, the target action value of this iteration is recalculated through steps C31-C34, so that the target action value will also calculate the latest fitting result as the action value model is optimized.

[0273] In some embodiments, although all directed edges in the state topology graph can reflect the complete action distribution, the computational overhead of calculating the target action value on the directed edges of the entire graph is large. Therefore, when calculating the target action value, random sampling can be performed from the directed edge set to obtain multiple sampling edges. Since each randomly sampled sampling edge can also reflect the same action distribution as all directed edges, as long as the target action value is calculated based on the sampled edges, it can also better reflect the optimization goal of the action value based on the action distribution, and greatly reduce the computational overhead of the target action value, thereby improving the training efficiency of the action value model.

[0274] In other embodiments, the target action value may also be calculated based on all directed edges in the state topology graph, so that the target action value used in each training iteration has higher accuracy. This embodiment of the present application does not specifically limit this.

[0275] In this step C31, the target action value is calculated based on the sampled edge as an example. After randomly sampling from the directed edge set to obtain multiple sampled edges, the sampling state s indicated by the starting node of any sampled edge can be determined. t and the sampling action a indicated by the sampling edge t .

[0276] C32. The training server determines, based on the action feedback model, an estimated feedback value of the agent performing the sampled action under the sampling state.

[0277] In some embodiments, for each sampled edge, the sampling state s of the sampled edge is t and sampling action a t Input to the action feedback model Feedback model through action Output for the sampling state s t and sampling action a t Calculated estimated feedback value In some embodiments, the estimated feedback value It can be the action feedback model after the latest update. The real-time calculation may also be the value of the latest version cached on the directed edge in the state topology graph, and the embodiments of the present application do not specifically limit this.

[0278] C33. The training server determines the probability of the agent executing the sampled action under the sampling state based on the action decision model.

[0279] In some embodiments, the sampling state s indicated by the starting node of each sampling edge t , the sample state s t Input to the action decision model Action decision model Output for the sampling state s t Calculate the sampling action a t The probability of execution The execution probability represents the agent's state in the sampled state s t Execute sampling action a t In some embodiments, since the action decision model and the action value model can be trained in tandem with the same synchronization or trained separately in different iteration cycles, the execution probability Refers to the action decision model after the latest update Calculated in real time.

[0280] C34. The training server determines the target action value of the sampling action based on the estimated feedback value, execution probability and estimated action value of the agent performing the sampling action in the sampling state.

[0281] In some embodiments, for each sampled edge, the sampling state s at the sampled edge t and sampling action a t Given the situation, the estimated feedback value obtained in step C32 is used The execution probability obtained in step C33 And the estimated action value Q obtained in step C1 θ (s t ,a t ), the target action Q can be determined target .

[0282] Next, we will use information entropy to define the target action value Q target The target action value Q target The calculation formula is as follows:

[0283] Among them, Q target Represents the target action value, Represents the estimated feedback value, s t Characterize the sampling state, a t Characterizes the sampling action, γ is a hyperparameter, Represents the execution probability (for the same sampling state, if there are multiple possible sampling actions, then the execution probability of each sampling action can form a decision vector), Q θ (s t ,a t ) represents the estimated action value, α is a learnable temperature parameter (the variable α can update its value as the parameter set of the action value model is iteratively adjusted). Characterizes the information entropy of the decision vector, and the temperature parameter α is used to control the weight provided by the information entropy.

[0284] C35. The training server obtains the action value loss item based on the target action value and the estimated action value.

[0285] In some embodiments, based on the target action value Q obtained in step C34 target And the estimated action value Q obtained in step C1 θ (s t ,a t), an action value loss term can be constructed. For example, directly calculate the target action value Q of each sampled action target and the estimated action value Q θ (s t ,a t ) to quickly calculate the action value loss term.

[0286] In other embodiments, for each sampled action, its target action value Q can also be obtained. target and the estimated action value Q θ (s t ,a t ), and then calculate the mathematical expectation of the square values ​​calculated for each sampled action to obtain the action value loss term. The action value loss term in this case is as follows:

[0287] Among them, Q θ (s t ,a t ) represents the estimated action value, s t Characterize the sampling state, a t Characterize the sampling action, Q target represents the target action value, G represents the state topology, τ t Represents the transfer data at time step t in the state topology graph

[0288] In steps C31-C35, a method for constructing an action value loss term based on the soft Bellman residual is provided, which can introduce the information entropy about the decision vector into the action value loss term, making the action value loss term more informative and more accurately measuring the degree of difference between the estimated action value and the target action value.

[0289] C4. The training server iteratively trains the action value model based on the constraint loss term and the action value loss term.

[0290] In some embodiments, based on the constraint loss term obtained in step C2 and the action value loss term obtained in step C3, the two can be summed or weighted summed to obtain the loss function value of this training iteration.

[0291] In an example, taking the weighted sum of the constraint loss term and the action value loss term as an example, the loss function expression of the action value model is as follows:

[0292] Among them, J Q (θ) represents the loss function value of the action value model, Characterizes the action value loss term, λ is a hyperparameter (referring to the weighting factor of the constraint loss term), Characterize the constraint loss term.

[0293] When the technician sets the action value optimization stopping condition, the method provided by the above steps C1-C4 can be used to calculate the loss function value of the action value model in each training iteration, and then determine whether the action value optimization stopping condition is met. If the action value optimization stopping condition is met, then training is stopped and the trained action value model is obtained. If the action value optimization stopping condition is not met, the next training iteration is started. The action value optimization stopping condition can be that the loss function value tends to converge, or the number of iteration steps reaches the set number of steps. The action value optimization stopping condition is not specifically limited here.

[0294] In the training method of the action value model provided in steps C1ˉC4, a constraint loss term is constructed using the empirical action value function for the loss function of the action value model to regularize the action value model, thereby alleviating the over-estimation error and extrapolation error of the action value model in the learning process of the empirical action value function. At the same time, an action value loss term is also considered to measure the difference between the estimated action value provided by the action value model and the target action value in the optimization learning. Under the joint action of the constraint loss term and the action value loss term, an action value model with better accuracy and better performance can be assisted in training, and the generalization of the action value model can be improved.

[0295] In the above step 404, a possible implementation method is provided for the training server to train the action value model of the intelligent agent based on the state topology graph and the action feedback model. Since the empirical action value function can be derived from the state topology graph, the empirical action value function is used to construct the constraint loss term, and the action feedback model and the action decision model are used to construct the action value loss term. The loss function constructed by combining the two can fully reflect the degree of difference between the action distribution predicted by the model relative to the empirical distribution and the optimization target, so that the final trained action value model has better generalization.

[0296] In other embodiments, in the loss function of the action value model, only the action value loss term can be considered without considering the constraint loss term. This can save the computational overhead of the loss function value and improve the training efficiency of the action value model. The embodiments of the present application do not make specific limitations on this.

[0297] 405. The training server trains an action decision model of the intelligent agent based on the action value model. The action decision model is used to decide the action that the intelligent agent should perform in a given state.

[0298] In some embodiments, based on the action value model trained in step 404, since the action value model can feedback the estimated action value of each action in a given state, and different parameters of the action decision model may result in different actions being decided in the same state, the action value model is used to evaluate the target action determined by the action value model to obtain the estimated action value, which can reflect the accuracy of the parameters of the action decision model itself, where the target action refers to the action selected to be executed based on the decision vector. Therefore, the action value model can assist in completing the training and optimization of the action decision model. Since a more optimal action value model has been obtained in step 404, the ultimately optimized action decision model will also have higher accuracy and better performance.

[0299] In other embodiments, the action value model and the action decision model are trained and optimized in a collaborative manner, that is, during each iteration of the iterative training, the action value model and the action decision model are updated once according to the state topology diagram and the action feedback model. There is no order of priority in the optimization of the action value model and the action decision model. The action value model can be updated first and then the action decision model, or the action decision model can be updated first and then the action value model, or the action value model and the action decision model can be updated simultaneously. The above optimization process is iteratively performed until the action decision model meets the decision optimization stopping condition, at which time the trained action decision model is obtained. The decision optimization stopping condition can be that the loss function value tends to converge, or the number of iteration steps reaches the set number of steps, etc. The decision optimization stopping condition is not specifically limited here. In the collaborative training optimization process, the action value model and the action decision model are updated once in each iteration, so that the action value model and the action decision model can guide each other, thereby further improving the accuracy of the action decision model finally optimized.

[0300] The following describes a possible training method for the action decision model in conjunction with steps D1-D4. In this training method, a decision loss term is introduced into the loss function of the action decision model, and the action value estimated by the action value model for the action output is taken into account in the decision loss term. In this way, the action decision model can be trained under the guidance of the action value model. The following is an explanation:

[0301] D1. In any iteration, the training server determines the decision vector of the agent at the current moment through the action decision model. The decision vector indicates the possibility of the agent performing various actions at the current moment.

[0302] In some embodiments, for a given state s at the current time t t , change the state s t Input to the action decision model Action decision model Output for state s t Calculate the decision vector The decision vector includes the agent in state s t The probability of executing each possible action under state s. Each execution probability represents the agent’s state s. t The probability of performing a possible action.

[0303] D2. The training server determines a scoring vector of the action value of the agent at the current moment based on the action value model. The scoring vector indicates the estimated action value of each of the multiple actions performed by the agent at the current moment.

[0304] In some embodiments, for a given state s at the current time t t , action decision model Output for state s t Calculate the decision vector Decision Vector Including the agent in state s t The probability of executing each possible action is, then, the state s t and every possible action a t Can form a state-action pair, each pair of state s t and action a t Input to the action-value model Q θ In the action value model Q θ For each pair of states s t and action a t Both output an estimated action value Q θ (s t ,a t ), using state s t With respect to all possible actions a t The estimated action value Q θ (s t ,a t ), we can construct a state s t The scoring vector Q θ (s t ). In some embodiments, due to the action decision model and the action-value model Q θ It can be trained in the same synchronization or trained separately in different iteration cycles, so the scoring vector Q θ (s t ) refers to the action value model Q after the latest update θ Calculated in real time.

[0305] Taking the collaborative training framework as an example, in each training iteration, the action decision model The current state s of the agent will be calculated t The decision vector under The agent then uses the decision vector Decide which action to perform, then perform action a t And interact with the environment, the environment through action feedback model Give estimated feedback value and reaches the next state s t+1 , then update the number of visits to the state topology or add new nodes and directed edges, and then the action value model Q θ Given the estimated action value Q θ (s t ,a t ), and then in this training iteration, for the action value model Q θ and action decision model Each of them will update its own parameter set once (there is no order in which the two updates are made, and they can also be updated simultaneously), and then enter the next training iteration. In the previous step 404, the action value model Q θ The loss function and its training method, and this step 405 introduces the action decision model The loss function and its training method.

[0306] D3. The training server determines the decision loss term of the agent at the current moment based on the decision vector and the scoring vector. The decision loss term represents the error between the action executed by the agent and the task expectation.

[0307] In some embodiments, based on the decision vector obtained in step D1 and the scoring vector Q obtained in step D2 θ (s t ), calculate the decision loss term of the agent at the current time t.

[0308] In one example, using a decision vector The information entropy and scoring vector Q θ (s t ) to construct the decision loss term, and the expression of the decision loss term is as follows:

[0309] in, Represents the decision loss term, Representing action decision models The parameter set, Representation of each state s in the state topology graph G t Find the mathematical expectation of the values ​​in the brackets. Represents the decision vector, represents the transposed vector of the decision vector, α is a learnable temperature parameter, Characterize the information entropy of the decision vector, the temperature parameter α is used to control the weight provided by the information entropy, Q θ (s t ) represents the scoring vector. It should be noted that the temperature parameter α in the above formula can be reused to obtain the target action value Q in step C34 target α is used when, but if the target action value Q target If the definition is not based on information entropy, the variable α can update its value along with the iterative adjustment of the parameter set of the action decision model. The embodiment of the present application does not impose any specific limitation on this.

[0310] D4. The training server iteratively trains the action decision model based on the decision loss term.

[0311] In some embodiments, the loss function value of the action decision model is equal to the decision loss term obtained in step D3. At this time, if the technician sets the decision optimization stopping condition, the loss function value of the action decision model in each training iteration can be calculated through the method provided by the above steps D1-D4, and then it can be determined whether the decision optimization stopping condition is met. If the decision optimization stopping condition is met, then training is stopped and the trained action decision model is obtained. If the decision optimization stopping condition is not met, the next training iteration is started. The decision optimization stopping condition can be that the loss function value tends to converge, or the number of iteration steps reaches the set number of steps, etc. The decision optimization stopping condition is not specifically limited here.

[0312] All of the above technical solutions can be combined in any way to form the embodiments of the present disclosure, and will not be described in detail here.

[0313] The method provided in the embodiment of the present application constructs a non-parametric state topology map based on the historical trajectory. This state topology map can fully reflect the empirical distribution of the intelligent agent's actions, has a higher information utilization rate for the historical trajectory, and also brings more information. Then, based on the state topology map, the action feedback model is guided to train, so that the action feedback model can calculate a more accurate estimated feedback value, thereby improving the accuracy of the action feedback model. Then, by combining the state topology map and the action feedback model, the training process of the action value model can be constrained to obtain an action value model with better accuracy and better performance, so that the action value model can calculate the estimated action value more accurately and can conform to human intentions or preferences to a certain extent. Finally, the more accurate action value model is used to assist in training a more accurate action decision model, thereby helping to make accurate decisions on what action the intelligent agent should perform in a given state.

[0314] In the above embodiment, the training process of the action decision model of the intelligent agent provided by the present application is described in detail. The following will be combined with Figure 5 to take the action decision model as the strategy neural network. The action feedback model is a reward neural network The action value model is the action value neural network Q θ As an example, a possible implementation method of the training process is illustrated.

[0315] As shown in Figure 5, at the beginning of the algorithm, the policy neural network is initialized Reward Neural Network and action-value neural network Q θ Next, using the policy neural network To control the interaction between the agent and the environment, multiple historical trajectories are collected. Specifically, the state s observed at any moment is input into the policy neural network Through the policy neural network To predict the decision vector under state s According to the decision vector, the action a to be executed is determined, and the agent is controlled to execute the action a. After the agent executes the action a, it interacts with the environment so that the environment updates the new state s' at the next moment. Then, the environment can reward the neural network. Input a triple (s, a, s') and output an estimated feedback value Next, the quadruple The transfer data at one time step is stored in the experience replay buffer. The above operations are performed cyclically until the agent's task is completed, failed, or the set length is reached. A historical trajectory is obtained, and this historical trajectory is stored in the experience replay buffer in the form of transfer data at multiple time steps.

[0316] Furthermore, based on the series of transition data stored in the experience replay buffer, a state topology graph G is constructed. Specifically, the nodes in the state topology graph G are constructed based on the states observed in the historical trajectory, and the directed edges in the state topology graph G are constructed based on the actions executed in the historical trajectory. It should be noted that since there may be unobserved states, there may also be unexecuted actions in the entire action distribution. Therefore, the state topology graph G can be considered a subset of the complete action distribution in the environment space.

[0317] Furthermore, trajectory sampling is performed on the state topology graph G to obtain multiple pairs of sampled trajectories. Since trajectory sampling is very efficient, it is possible to efficiently construct effective and more informative trajectory pairs. In addition, trajectory sampling allows the topological structure of the state topology graph G to be used to splice out new sampled trajectories. The sampled trajectories are not necessarily directly obtained historical trajectories, but are formed by connecting different trajectory fragments from multiple historical trajectories. In this way, the sampled trajectories are not historical events, but the trajectory fragments in the sampled trajectories are all historical events. These sampled trajectories can be presented in pairs to technicians or large models, who will annotate them to obtain the annotation results for each pair of sampled trajectories. These annotation results reflect the preference data of humans (or those learned by the large model). Then, the preference data is used as a supervisory signal, and the state topology graph G is used to assist the reward neural network. Perform supervised learning to obtain an optimized reward neural network In some embodiments, the reward neural network is updated regularly based on the preference data collected during a certain period. Parameter set Parameter Set After each update, it is necessary to calculate the reward neural network based on the latest reward neural network. The estimated feedback value for each directed edge in the state topology graph G Recalculation and relabeling are performed to update the entire graph of estimated feedback values, which maximizes the utilization of historical trajectories and alleviates the adverse effects of non-stationary reward functions.

[0318] Furthermore, an empirical action value function can be derived from the state topology graph G Make the experience action value function For each state s indicated by a node, an empirical action value can be calculated Reusing the quad In turn, train the action value neural network Q θ and policy neural networks Specifically, using the experience action value To the action value neural network Q θ Provide constraints and regularize so that the empirical action value As an action-value neural network Q θ The estimated action value given The lower bound of , thus fully considering the information provided by the empirical distribution of actions, and preventing the action value neural network Q θ For overestimation of action values, reduce the action value neural network Q θ The overestimation error and extrapolation error make the action value neural network Q θIt can more accurately estimate the action value of each action in a given environment. Then, the estimated action value Q θ (s,·) to policy neural network Introducing evaluation information for action values ​​to ultimately improve the strategy neural network The generalization and policy performance of the policy neural network The training efficiency and learning efficiency of the policy neural network It can provide better strategic recommendations for intelligent agents, so that they can make decisions on high-quality actions that are in line with human preferences or intentions, reducing the number of times human intervention is required when the intelligent agent performs tasks and reducing labor costs.

[0319] It should be noted that the action value neural network Q θ and policy neural networks It can be iteratively trained based on a collaborative training framework, that is, the action value neural network Q θ and policy neural networks Update the parameter sets θ and Finally, we can get a trained action value neural network Q θ and policy neural networks

[0320] It should also be noted that during the entire training iteration process, as historical trajectories continue to increase, if a directed edge already in the state topology graph G is discovered again, the visit count of this directed edge is updated. If a directed edge indicating an action is discovered that does not exist, then if the state topology graph G does not have the corresponding starting node or arrival node, the corresponding starting node or arrival node is added, and the corresponding directed edge is added to the state topology graph G. In this way, the state topology graph G will be continuously improved during the adaptive update. Furthermore, the empirical action value function can also periodically update the empirical action value of each directed edge and each node based on the visit count, achieving adaptive update and optimization of the empirical action value.

[0321] Taking the scenario of multiple robot carts cooperating in transportation as an example, through the preference-based reinforcement learning framework of the embodiment of the present application, the robot carts can more accurately meet human needs and preferences. For example, when multiple robot carts are required to cooperate in moving a cargo of a specific shape and weight, the trained policy neural network It can provide the best movement, cooperation and path planning strategies for each robotic vehicle to ensure the safe, fast and efficient movement of goods.

[0322] In the above embodiments, the construction and update process of the state topology map is described in detail, and the training process of the action feedback model, action value model and action decision model is described in detail. In the embodiments of the present application, the process of using the trained action decision model to control the intelligent agent to perform actions will be described in detail.

[0323] FIG6 is a flow chart of an agent action decision method provided in an embodiment of the present application. As shown in FIG6 , this embodiment is executed by a computer device, which can be the agent 201 or the agent control system 202 in the above implementation environment, or other control terminals equipped with an action decision model. Taking the computer device as an example of the control terminal, this embodiment includes the following steps:

[0324] 601. When the current state is observed in the environment, the control terminal inputs the state into the action decision model of the intelligent agent, and the action decision model is used to decide the action that the intelligent agent should perform in a given state.

[0325] The action decision model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide feedback signals from the environment to the actions performed by the agent, and the action value model is used to evaluate the impact of the actions performed by the agent on the environment. The training processes for the action feedback model, action value model, and action decision model are detailed in the above embodiments and will not be repeated here.

[0326] In some embodiments, the training server uses the training method of the action decision model of the intelligent agent provided in the above embodiments to train an action decision model with better accuracy and better performance under the preference-based reinforcement learning framework. Then, the training server sends the parameter set of the action decision model to the control terminal of the intelligent agent, so that the control terminal obtains the trained action decision model.

[0327] In some embodiments, the training process of the action decision model can be completed locally on the training server, such as local offline training on the training server, or it can be completed in the cloud, such as distributed training by multiple servers to speed up training efficiency. The embodiments of the present application do not specifically limit this.

[0328] In some embodiments, the control terminal of the intelligent agent can be a control module built into the intelligent agent, or it can be a control device independent of the intelligent agent, and multiple intelligent agents can be macro-scheduled by the same control device, which is also called a master control terminal. The embodiments of the present application do not specifically limit this.

[0329] In some embodiments, the control terminal obtains the trained action decision model Later, for the environment where the agent resides, if the current state s is observed in the environment, the state s can be input into the action decision model of the agent , proceed to step 602 below.

[0330] 602. The control terminal determines the execution probability of the agent for each of the multiple candidate actions through the action decision model. The execution probability of each candidate action represents the possibility of the agent executing the corresponding candidate action in the current state.

[0331] In some embodiments, the control terminal inputs the state s into the action decision model of the agent Later, through the action decision model Calculate a decision vector for state s Decision Vector It includes the execution probability of the agent performing each possible action a in state s. Each execution probability represents the possibility of the agent performing an action a in state s. For example, the agent can perform a total of 7 candidate actions: moving forward, moving backward, moving left, moving right, jumping, squatting, and standing still. However, in a certain state s, the agent moves to a corner, making it impossible for the agent to move left or backward. Then the action decision model Seven execution probabilities will still be calculated for the seven candidate actions, but the execution probabilities of moving left and moving backward are close to 0.

[0332] 603. The control terminal determines a target action to be performed by the agent at the current moment from the multiple candidate actions based on the execution probabilities of the multiple candidate actions.

[0333] In some embodiments, the control terminal is based on the decision vector in step 602 The execution probability of each candidate action given in can be used to decide a target action from multiple candidate actions. The target action refers to the optimal candidate action that the action decision model plans for the agent in the current state s and is most in line with human preferences.

[0334] In some embodiments, the candidate action with the highest execution probability can be directly selected as the target action, or one can be randomly selected from the top N candidate actions with the highest execution probability as the target action, or sampling can be performed on the probability distribution of the candidate actions according to the execution probability, and the target action to be performed by the intelligent agent can be randomly decided so that the sampling process obeys the probability distribution. The embodiments of the present application do not specifically limit this.

[0335] Taking the scenario of multiple robot cars cooperating in transportation as an example, in the action decision model After training, if the robot car itself is its own control terminal, the training server will The parameter set is sent to each robot car, so that each robot car has its own built-in action decision model Under the control of the decision-making model, multiple robot cars collaborate to transport the goods from the starting point to the end point to complete the cargo transportation task; or, if multiple robot cars have a master control terminal, the master control terminal is responsible for scheduling the actions of each robot car, then the training server will The parameter set is sent to the master control terminal of these robot cars, and the master control terminal uses the action decision model The action of each robot car is decided under the control of , and the control signal of each action is sent to the corresponding robot car respectively, so as to realize the macro-scheduling of multiple robot cars by the same master control terminal, and control multiple robot cars to cooperate in transporting the goods from the starting point to the end point to complete the cargo transportation task. The embodiments of the present application do not make specific limitations on this.

[0336] All of the above technical solutions can be combined in any way to form the embodiments of the present disclosure, and will not be described in detail here.

[0337] The method provided in the embodiment of the present application obtains an action decision model with better performance and better generalization ability through collaborative training of a state topology graph, an action feedback model and an action value model, so that under the state of a given intelligent agent, the action decision model can make high-quality actions for the intelligent agent that are in line with human preferences or intentions, which means that the intelligent agent can understand human preferences more accurately, reduce the number of times human intervention is required when the intelligent agent performs tasks, and reduce labor costs.

[0338] In the above embodiments, the action decision model trained by the preference-based reinforcement learning algorithm is described in detail, and how to realize the decision-making of the action of the intelligent agent in the application scenario. This preference-based reinforcement learning algorithm uses the empirical action value function to assist in trajectory sampling in the state topology diagram, and normalizes the action value model based on the neural network to achieve efficient learning. Experiments and tests show that the algorithm provided by the embodiment of the present application surpasses other preference-based reinforcement learning algorithms in various complex tasks, and greatly improves the human feedback efficiency, wherein the human feedback efficiency is referred to as feedback efficiency, which refers to how to maximize the learning effect under limited human feedback (the preference data annotated by human experts is usually expensive, and traditional algorithms often require more human feedback to learn, so the human feedback efficiency is low). By utilizing the aligned empirical estimation, the algorithm provided by the embodiment of the present application shows obvious advantages in the performance comparison with other traditional algorithms, especially in the scenario where only limited human feedback labels are available. At the same time, the algorithm provided by the embodiment of the present application can train an accurate action value model and a better action decision model.

[0339] The following describes the test performance of some tasks in different scenarios.

[0340] Figure 7 is a performance comparison chart of an intelligent agent on a box-pushing task provided by an embodiment of the present application. As shown in Figure 7, the task of pushing a box by a robot is selected as the experimental task in the test scenario. For a robot, the box-pushing task is a relatively complex task. This scheme, traditional schemes 1-3, and baseline scheme 4 are used to train the action decision model of the robot car respectively, and the action decision model trained by various schemes is used to control the robot to complete the box-pushing task, and the average reward learning curves of various schemes are plotted. Among them, (a)ˉ(c) is the test result of the average reward learning curve when the human-labeled feedback (Feedback) is 300, and (d)ˉ(f) is the test result of the average reward learning curve when the human-labeled feedback is 1000. The environment sizes of (a) and (d) are both 5×5, the environment sizes of (b) and (e) are both 6×6, and the environment sizes of (c) and (f) are both 7×7. The unit of the environment size is the size specified by the map scale.

[0341] Figure 7 (a) and (f) show a comparison of six sets of average reward learning curves. The horizontal axis represents time steps, and the vertical axis represents the average learning return rate (Episode Return). A higher average learning return rate indicates better performance of the action decision model. For each set of average reward learning curve comparisons, in each experimental task, Baseline 4 represents the highest performance achieved using the true reward function. Therefore, Baseline 4 is plotted as a horizontal line, representing the upper bound of performance for the experimental task.

[0342] As Figure 7 clearly shows, this solution not only surpasses existing traditional solutions 1-3 in performance, but also excels in sample efficiency. Furthermore, in the early stages of training, this solution rapidly achieves high performance on most tasks, significantly faster than other traditional solutions 1-3. Notably, this solution only requires a small amount of human preference labels (i.e., human annotations of track pairs or segment pairs, also referred to as preference data) to approach the performance of the baseline solution, demonstrating extremely high feedback efficiency.

[0343] Figure 7 also shows that some baseline solutions are significantly affected by random factors in some tasks, resulting in unstable performance in some experiments and large fluctuations in the learning curve. Conversely, some traditional solutions exhibit a downward trend in the learning curves of more difficult tasks, indicating that these solutions are sensitive to trajectory sample quality. Specifically, when a large number of sampled trajectory pairs are of low quality, this may interfere with the training of the action-value model, thereby affecting the performance of the overall strategy (i.e., the action decision model).

[0344] FIG8 is a performance comparison diagram of a robot provided in an embodiment of the present application in a construction task. As shown in FIG8 , for the field of robot construction, a robot construction task was selected as the experimental task in the test scenario. For the robot, the construction task is also a relatively complex task. In the test, for the CraftEnv environment in the field of robot construction, three construction tasks were selected: strip-shaped building (Strip-Shaped), block-shaped building (Block-Shaped) and simple two-story building (Two-Story) tasks for testing and research. This solution, traditional solutions 1-3 and baseline solution 4 were used to train the robot's action decision model respectively, and the action decision models trained by various solutions were used to control the robot to complete the above three construction tasks, and the average reward learning curves of various solutions were drawn.

[0345] Figure 8 (a) and (c) show a comparison of three sets of average reward learning curves: (a) for the strip building task, (b) for the block building task, and (c) for the two-story building task. The horizontal axis represents time, and the vertical axis represents the average learning reward. Furthermore, the human-labeled feedback for all three construction tasks is 1000. These tasks fully demonstrate the complexity and diversity of real-world construction tasks, requiring the agent (i.e., intelligent agent) to precisely manipulate building components to meet specific architectural design requirements.

[0346] As can be clearly seen in Figure 8, this solution not only outperforms traditional solutions 1-3 in algorithmic performance, but also significantly improves feedback efficiency. More notably, despite using far fewer samples than traditional solution 2, this solution's performance is comparable to or even exceeds that of traditional solution 2. Specifically, in the linear building task, this solution surpassed the average performance of traditional solution 2 using only 30% of the samples. This result can be observed by comparing the curves representing this solution and traditional solution 2 in Figure 8, where the performance of this solution is particularly outstanding. These results demonstrate that this solution not only achieves efficient learning when handling complex tasks, but also significantly reduces the amount of feedback required, thereby greatly improving learning efficiency.

[0347] It should also be noted that the state topology graph used in the embodiment of the present application adopts a simple non-parametric model, namely a graph model, but is not limited to a specific non-parametric model or its topological structure, and can be replaced with various novel and effective model structures as needed. In addition, this solution does not have a fixed choice for the underlying reinforcement learning algorithm. In practical applications, an appropriate reinforcement learning algorithm can be selected according to specific scenarios and needs. At the same time, the learning framework of this solution can also be combined with other technologies, such as unsupervised learning exploration, time series data expansion, and data enhancement using pseudo-labels.

[0348] FIG9 is a schematic diagram of the structure of a training device for an action decision model of an intelligent agent provided in an embodiment of the present application. As shown in FIG9 , the device includes:

[0349] A topology map construction module 901 is configured to construct a state topology map based on multiple historical trajectories of the agent. Each historical trajectory includes multiple actions, each of which is used to control the transition between different states. Each node in the state topology map indicates a state, and each directed edge connecting a pair of nodes indicates an action.

[0350] A feedback training module 902 is configured to train an action feedback model of the agent based on the state topology graph, wherein the action feedback model is configured to provide feedback signals from the environment in which the agent resides regarding actions performed by the agent;

[0351] An action value training module 903 is configured to train an action value model of the agent based on the state topology graph and the action feedback model, wherein the action value model is configured to provide an estimated action value for an action performed by the agent, wherein the estimated action value indicates a metric value for measuring the impact of the action performed by the agent on the environment;

[0352] The decision training module 904 is used to train the action decision model of the intelligent agent based on the action value model. The action decision model is used to decide the action that the intelligent agent should perform in a given state.

[0353] It should be noted that the training device for the action decision model of an intelligent agent provided in the above embodiment only uses the division of the above functional modules as an example to illustrate when training the action decision model of an intelligent agent. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device (such as a training server) can be divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the action decision model of an intelligent agent provided in the above embodiment and the embodiment of the training method for the action decision model of an intelligent agent are based on the same concept. The specific implementation process is detailed in the embodiment of the training method for the action decision model of an intelligent agent, and will not be repeated here.

[0354] All of the above technical solutions can be combined in any way to form the embodiments of the present disclosure, and will not be described in detail here.

[0355] FIG10 is a schematic diagram of the structure of an action decision device of an intelligent agent provided in an embodiment of the present application. As shown in FIG10 , the device includes:

[0356] Input module 1001 is used to input the current state of the environment into the action decision model of the agent, and the action decision model is used to determine the action that the agent should perform in the given state;

[0357] A probability determination module 1002 is configured to determine, using the action decision model, the execution probability of the agent for each of a plurality of candidate actions, where the execution probability of each candidate action represents the likelihood that the agent will execute the candidate action in the given state;

[0358] An action determination module 1003 is configured to determine a target action to be performed by the agent at the current moment from among the multiple candidate actions based on the execution probabilities of the multiple candidate actions.

[0359] Among them, the action decision model is obtained through collaborative training based on a state topology graph, an action feedback model and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment on the action performed by the agent, and the action value model is used to provide an estimated action value for the action performed by the agent. The estimated action value indicates a measurement value for measuring the impact of the action performed by the agent on the environment.

[0360] It should be noted that the action decision device of the intelligent agent provided in the above embodiment only uses the division of the above functional modules as an example to illustrate when deciding the action of the intelligent agent. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device (such as the intelligent agent or the intelligent agent's control terminal) is divided into different functional modules to complete all or part of the functions described above. In addition, the action decision device of the intelligent agent provided in the above embodiment and the embodiment of the action decision method of the intelligent agent belong to the same concept. The specific implementation process is detailed in the embodiment of the action decision method of the intelligent agent, which will not be repeated here.

[0361] All of the above technical solutions can be combined in any way to form the embodiments of the present disclosure, and will not be described in detail here.

[0362] FIG11 is a schematic structural diagram of a control terminal provided in an embodiment of the present application. As shown in FIG11 , the control terminal is an exemplary illustration of a computer device. The control terminal is used to control the actions of an intelligent agent. It may be a control system built into the intelligent agent or a control device independent of the intelligent agent. In some embodiments, the device types of the control terminal 1100 include: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV), a laptop computer, or a desktop computer. The control terminal 1100 may also be referred to as a user device, a portable control terminal, a laptop control terminal, a desktop control terminal, or other names.

[0363] Typically, the control terminal 1100 includes a processor 1101 and a memory 1102 .

[0364] In some embodiments, processor 1101 includes one or more processing cores, such as a quad-core processor or an octa-core processor. In some embodiments, processor 1101 is implemented in hardware using at least one of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). In some embodiments, processor 1101 includes a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, processor 1101 is integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the display screen. In some embodiments, processor 1101 also includes an AI (Artificial Intelligence) processor, which is used to handle computing operations related to machine learning.

[0365] In some embodiments, the memory 1102 includes one or more computer-readable storage media, which in some embodiments are non-transitory. In some embodiments, the memory 1102 also includes high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1102 is used to store at least one program code, which is executed by the processor 1101 to implement the action decision method of the intelligent agent provided in various embodiments of the present application.

[0366] In some embodiments, the control terminal 1100 may further include a peripheral device interface 1103 and at least one peripheral device. The processor 1101, memory 1102, and peripheral device interface 1103 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1108.

[0367] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 are implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0368] The RF circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. In some embodiments, the RF circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. In some embodiments, the RF circuit 1104 communicates with other control terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1104 also includes circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0369] The display screen 1105 is used to display a UI (User Interface). In some embodiments, the UI includes graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch screen display, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signals can be input as control signals to the processor 1101 for processing. In some embodiments, the display screen 1105 is also used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there is one display screen 1105, which is provided as a front panel of the control terminal 1100; in other embodiments, there are at least two display screens 1105, which are provided on different surfaces of the control terminal 1100 or are foldable; in still other embodiments, the display screen 1105 is a flexible display screen, which is provided on a curved surface or a foldable surface of the control terminal 1100. In some embodiments, the display screen 1105 is provided as a non-rectangular irregular shape, i.e., a special-shaped screen. In some embodiments, the display screen 1105 is made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0370] The camera assembly 1106 is used to capture images or videos. In some embodiments, the camera assembly 1106 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the control terminal, and the rear camera is set on the back of the control terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1106 also includes a flash. In some embodiments, the flash is a monochrome temperature flash, or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which is used for light compensation at different color temperatures.

[0371] In some embodiments, the audio circuit 1107 includes a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 1101 for processing, or input them into the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there are multiple microphones, which are respectively arranged at different parts of the control terminal 1100. In some embodiments, the microphone is an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 1101 or the radio frequency circuit 1104 into sound waves. In some embodiments, the speaker is a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1107 also includes a headphone jack.

[0372] Power supply 1108 is used to power various components in control terminal 1100. In some embodiments, power supply 1108 is AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1108 includes a rechargeable battery, the rechargeable battery supports wired charging or wireless charging. The rechargeable battery is also configured to support fast charging technology.

[0373] In some embodiments, the control terminal 1100 further includes one or more sensors 1110 , including but not limited to: an acceleration sensor 1111 , a gyroscope sensor 1112 , a pressure sensor 1113 , an optical sensor 1114 , and a proximity sensor 1115 .

[0374] In some embodiments, the accelerometer 1111 detects the magnitude of acceleration along the three coordinate axes of the coordinate system established by the control terminal 1100. For example, the accelerometer 1111 is used to detect the components of gravity acceleration along the three coordinate axes. In some embodiments, the processor 1101 controls the display screen 1105 to display the user interface in either a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 1111. The accelerometer 1111 is also used to collect game or user motion data.

[0375] In some embodiments, the gyroscope sensor 1112 detects the orientation and rotation angle of the control terminal 1100. The gyroscope sensor 1112 and the acceleration sensor 1111 collaborate to collect the user's 3D movements on the control terminal 1100. The processor 1101 implements the following functions based on the data collected by the gyroscope sensor 1112: motion sensing (for example, changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0376] In some embodiments, the pressure sensor 1113 is disposed on the side frame of the control terminal 1100 and / or the lower layer of the display screen 1105. When the pressure sensor 1113 is disposed on the side frame of the control terminal 1100, it can detect the user's grip signal of the control terminal 1100, and the processor 1101 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 1113. When the pressure sensor 1113 is disposed on the lower layer of the display screen 1105, the processor 1101 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 1105. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0377] Optical sensor 1114 is used to detect ambient light intensity. In one embodiment, processor 1101 controls the display brightness of display screen 1105 based on the ambient light intensity detected by optical sensor 1114. Specifically, when the ambient light intensity is high, the display brightness of display screen 1105 is increased; when the ambient light intensity is low, the display brightness of display screen 1105 is decreased. In another embodiment, processor 1101 also dynamically adjusts the capture parameters of camera assembly 1106 based on the ambient light intensity detected by optical sensor 1114.

[0378] Proximity sensor 1115, also known as a distance sensor, is typically located on the front panel of control terminal 1100. Proximity sensor 1115 is used to detect the distance between the user and the front of control terminal 1100. In one embodiment, when proximity sensor 1115 detects that the distance between the user and the front of control terminal 1100 is gradually decreasing, processor 1101 controls display screen 1105 to switch from a screen-on state to a screen-off state. When proximity sensor 1115 detects that the distance between the user and the front of control terminal 1100 is gradually increasing, processor 1101 controls display screen 1105 to switch from a screen-off state to a screen-on state.

[0379] Those skilled in the art will appreciate that the structure shown in FIG11 does not limit the control terminal 1100 and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.

[0380] FIG12 is a schematic diagram of the structure of a training server provided in an embodiment of the present application. As shown in FIG12 , a training server is an exemplary illustration of a computer device. The training server 1200 may vary significantly due to different configurations or performance. The training server 1200 includes one or more processors (CPUs) 1201 and one or more memories 1202. The memories 1202 store at least one computer program, which is loaded and executed by the one or more processors 1201 to implement the training methods for the action decision models of intelligent agents provided in the various embodiments described above. In some embodiments, the training server 1200 also has components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The training server 1200 also includes other components for implementing device functions, which are not described in detail here.

[0381] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including at least one computer program. The at least one computer program can be executed by a processor in a terminal to implement the method for training the action decision model of the intelligent agent or the action decision of the intelligent agent in each of the above-mentioned embodiments. For example, the computer-readable storage medium includes ROM (Read-Only Memory), RAM (Random-Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device.

[0382] In an exemplary embodiment, a computer program product is also provided, comprising one or more computer programs stored in a computer-readable storage medium. One or more processors of a computer device can read the one or more computer programs from the computer-readable storage medium, and the one or more processors execute the one or more computer programs, so that the computer device can perform the training method of the action decision model of the intelligent agent or the action decision of the intelligent agent in the above-mentioned embodiment.

[0383] Those skilled in the art will understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing related hardware through a program. In some embodiments, the program is stored in a computer-readable storage medium. In some embodiments, the above-mentioned storage medium is a read-only memory, a disk, or an optical disk, etc.

[0384] The above description is merely an embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A training method for an action decision model of an agent, executed by a computer device, the method comprising: Based on multiple historical trajectories of the agent, construct a state topology graph, each of the historical trajectories comprising a plurality of actions, each action being used to control the transition between different states, each node in the state topology graph indicating a state, and each directed edge connecting a pair of nodes indicating an action; Based on the state topology graph, train the action feedback model of the agent, the action feedback model being used to provide a feedback signal of the environment where the agent resides on the action executed by the agent; Based on the state topology graph and the action feedback model, train the action value model of the agent, the action value model being used to provide an estimated action value for the action executed by the agent, the estimated action value indicating a metric value for measuring the impact of the action executed by the agent on the environment; Based on the action value model, train the action decision model of the agent, the action decision model being used to decide the action that the agent should execute in a given state.

2. The method according to claim 1, wherein, The constructing a state topology graph based on multiple historical trajectories of the agent comprises: Obtain an initialized state topology graph; For any action in any of the historical trajectories, query for a directed edge indicating the action in the state topology graph; In response to querying a directed edge indicating the action, update the access count associated with the directed edge, the access count indicating the query frequency of the directed edge; In response to not querying a directed edge indicating the action, determine the starting state and the reaching state associated with the action, query for a starting node indicating the starting state and a reaching node indicating the reaching state in the state topology graph, if the starting node indicating the reaching state is not queried, then add a reaching node indicating the reaching state, if the starting node indicating the starting state is not queried, then add a starting node indicating the starting state, and add a directed edge from the starting node to the reaching node.

3. The method according to any one of claims 1-2, wherein, The training the action feedback model of the agent based on the state topology graph comprises: Perform trajectory sampling based on the state topology graph to obtain multiple pairs of sampled trajectories, each pair of the sampled trajectories comprising a pair of sampled trajectories of equal length; Collect the annotation results of the multiple pairs of sampled trajectories, the annotation results indicating the preference degrees of different sampled trajectories in each pair of the sampled trajectories with respect to the task executed by the agent; Based on the state topology graph and the annotation results, train the action feedback model.

4. The method according to claim 3, wherein, The performing trajectory sampling based on the state topology graph to obtain multiple pairs of sampled trajectories comprises: Randomly sample from the node set of the state topology graph to obtain a plurality of sampling points; Starting from any one of the plurality of sampling points, start trajectory sampling along the directed edge starting from the sampling point, and stop sampling when the trajectory length reaches the sampling length to obtain a sampled trajectory; Pair the multiple sampled trajectories according to the trajectory length to obtain multiple pairs of sampled trajectories.

5. The method according to claim 4, wherein Among them, the trajectory sampling starts along the directed edge starting from the sampling point and stops sampling when the trajectory length reaches the sampling length. The obtained sampling trajectory includes: If there is only one directed edge starting from the sampling point, the arrival node pointed to by the directed edge is used as the next node in the sampling trajectory, and the trajectory sampling continues along the directed edge starting from this next node until the trajectory length reaches the sampling length and then stops sampling; If there are at least two directed edges starting from the sampling point, select the directed edge with the highest or lowest empirical action value from the at least two directed edges, and use the arrival node pointed to by the selected directed edge as the next node in the sampling trajectory, and continue the trajectory sampling along the directed edge starting from this next node until the trajectory length reaches the sampling length and then stops sampling; wherein, the empirical action value is determined by the empirical action value function, and the empirical action value function is used to provide the empirical action value of the action executed by the agent based on the state topology graph.

6. The method according to any one of claims 3-4, further comprising: After the action feedback model is trained, update the estimated feedback value of each directed edge in the state topology graph, and the estimated feedback value indicates the feedback signal that the environment is expected to generate when the agent executes the action indicated by the directed edge.

7. The method according to any one of claims 1 to 6, wherein Training the action value model of the agent based on the state topology graph and the action feedback model includes: In any iteration, through the action value model, obtain the estimated action value of the action indicated by each directed edge in the state topology graph; Based on the empirical action value function and the estimated action value, obtain a constraint loss term, and the constraint loss term characterizes the distribution difference between the empirical distribution and the model distribution of the action value; Based on the action feedback model and the estimated action value, obtain an action value loss term, and the action value loss term characterizes the difference between the estimated action value of the action by the model and the target action value, and the target action value characterizes the optimization target of the action value based on the action distribution; Based on the constraint loss term and the action value loss term, iteratively train the action value model.

8. The method according to claim 7, wherein The obtaining of the constraint loss term based on the empirical action value function and the estimated action value includes: Determine a plurality of attention nodes from the state topology graph, and the attention nodes indicate the states that the agent needs to pay attention to when realizing the task; For any one of the attention nodes, based on the empirical action value function, determine the empirical action value of each action in the support action set of the attention node, and the support action set includes the set of actions indicated by each directed edge starting from the attention node; Based on the empirical action value and the estimated action value of each action in the support action set, determine the action value error of the attention node, and the action value error characterizes the difference between the empirical action value and the estimated action value of each action in the support action set; Based on the action value errors of each attention node, obtain the constraint loss term.

9. The method according to claim 7, wherein The determining of the action value error of the attention node based on the empirical action value and the estimated action value of each action in the support action set includes: Determine the empirical action value vector of the concerned node based on the empirical action values of each action in the support action set; Determine the predicted action value vector of the concerned node based on the predicted action values of each action in the support action set; Determine the action value error based on the empirical action value vector and the predicted action value vector; 10. The method according to any one of claims 6-9, characterized in that, The obtaining of the action value loss term based on the action feedback model and the predicted action value includes: Randomly sample the set of directed edges of the state topology graph to obtain multiple sampled edges. For any sampled edge, determine the sampled state indicated by the starting node of the sampled edge and the sampled action indicated by the sampled edge; Based on the action feedback model, determine the predicted feedback value when the agent executes the sampled action in the sampled state; Based on the action decision model, determine the execution probability of the agent for the sampled action in the sampled state; Based on the predicted feedback value, execution probability, and predicted action value when the agent executes the sampled action in the sampled state, determine the target action value of the sampled action; Obtain the action value loss term based on the target action value and the predicted action value; 11. The method according to any one of claims 1-10, characterized in that, The training of the action decision model of the agent based on the action value model includes: In any iteration, through the action decision model, determine the decision vector of the agent in the current state. The decision vector indicates the possibility of the agent executing various actions at the current moment; Based on the action value model, determine the scoring vector of the action values of the agent in the current state. The scoring vector indicates the predicted action values of the agent executing various actions at the current moment; Based on the decision vector and the scoring vector, determine the decision loss term of the agent at the current moment. The decision loss term represents the error between the action decided and executed by the agent and the task expectation; Iteratively train the action decision model based on the decision loss term; 12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: When the action value update condition is satisfied, sample in the set of nodes of the state topology graph to obtain multiple nodes to be updated; Determine the support action set of each node to be updated. The support action set contains the set of actions indicated by each directed edge starting from the node to be updated; Update the empirical action of each action in the support action set Assign the maximum value among the empirical action values of each action in the support action set to the empirical action value of the node to be updated; 13. The method according to claim 12, wherein The updating of the empirical action value of each action in the support action set includes: For any action in the support action set, determine the multiple target nodes that can be reached by executing the action starting from the node to be updated; Based on the access times of each directed edge connecting the node to be updated and each target node, determine the empirical transition probability of the node to be updated reaching each target node through the action; Through the action feedback model, determine the predicted feedback value of executing the action starting from the node to be updated; Update the empirical action value of the action based on the experience transition probability of each target node reached by the action through the to-be-updated node, the estimated feedback value, and the empirical action value of each target node. Assign the maximum value among the candidate action values of each action in the support action set to the empirical action value of the action executed starting from the to-be-updated node.

14. A method for an agent to make an action decision, executed by a computer device, the method comprising: When the state at the current moment is observed in the environment, input the state into the action decision model of the agent. Through the action decision model, determine the execution probability of each of multiple candidate actions for the agent, where the execution probability represents the possibility of the agent executing the candidate action in the state. Based on the execution probabilities of the multiple candidate actions, determine the target action executed by the agent at the current moment from the multiple candidate actions. Wherein, the action decision model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment on the action executed by the agent, and the action value model is used to provide an estimated action value for the action executed by the agent. The estimated action value indicates a metric value used to measure the impact of the action executed by the agent on the environment.

15. A training device for an action decision model of an agent, executed by a computer device, the device comprising: A topology graph construction module, configured to construct a state topology graph based on multiple historical trajectories of the agent. Each historical trajectory contains multiple actions, and each action is used to control the transition between different states. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. A feedback training module, configured to train the action feedback model of the agent based on the state topology graph. The action feedback model is used to provide a feedback signal of the environment where the agent resides on the action executed by the agent. An action value training module, configured to train the action value model of the agent based on the state topology graph and the action feedback model. The action value model is used to provide an estimated action value for the action executed by the agent. The estimated action value indicates a metric value used to measure the impact of the action executed by the agent on the environment. A decision training module, configured to train the action decision model of the agent based on the action value model. The action decision model is used to decide the action that the agent should execute in a given state.

16. An action decision device for an agent, executed by a computer device, the device comprising: An input module, configured to input the state into the action decision model of the agent when the state at the current moment is observed in the environment. The action decision model is used to decide the action that the agent should execute in a given state. A probability determination module, configured to determine, by using the action decision model, the execution probability of each of multiple candidate actions for the agent, where the execution probability represents the possibility of the agent executing the candidate action in the state; An action determination module, configured to determine, based on the execution probability of each of the multiple candidate actions, a target action that the agent executes at the current moment from the multiple candidate actions; Wherein, the action decision model is obtained through collaborative training based on a state topology graph, an action feedback model, and an action value model. Each node in the state topology graph indicates a state, and each directed edge connecting a pair of nodes indicates an action. The action feedback model is used to provide a feedback signal of the environment on the action executed by the agent, and the action value model is used to provide an estimated action value for the action executed by the agent. The estimated action value indicates a metric value used to measure the influence of the action executed by the agent on the environment.

17. A computer device, characterized in that, The computer device includes one or more processors and one or more memories. At least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the training method of the action decision model of the agent according to any one of claims 1 to 13, or the action decision method of the agent according to claim 14.

18. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the training method of the action decision model of the agent according to any one of claims 1 to 13, or the action decision method of the agent according to claim 14.

19. A computer program product, characterized in that, The computer program product includes at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the training method of the action decision model of the agent according to any one of claims 1 to 13, or the action decision method of the agent according to claim 14.

Citation Information

Patent Citations

  • Training method of action decision-making model of intelligent agent, action decision-making method and device

    CN120297323A

  • Intelligent decision model training method and device, equipment and storage medium

    CN115648204A

  • Multi-agent deep reinforcement learning training method based on PPO

    CN116306979A

  • Multi-agent collaborative strategy training method and system based on plot memory

    CN116360435A

  • Multi-agent collaborative decision-making method based on deep reinforcement learning under limited communication resources

    CN116456480A

Cited By

  • Data sharing method and system for multi-agent platform

    CN120832328A

  • Multi-agent optimization training method and system based on perceptual analysis

    CN121052315A

  • Intelligent agent decision behavior abnormity monitoring method and device, medium and product

    CN121093236A

  • Complex task collaborative decision-making method and device based on AI intelligent agent and medium

    CN121598988A

  • Multi-dimensional professional equipment simulation training resource scheduling method

    CN121858305A