An intelligent agent trajectory optimization method based on human feedback and task goals

By designing mathematical expressions of task goals and building a reinforced learning environment, and combining with human feedback optimization reward model, the problems of complex design and high debugging cost in the existing technology are solved, and efficient and precise optimization of the trajectory of the agent is achieved.

CN119962565BActive Publication Date: 2025-06-13FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510449875.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing technology relies on domain expert knowledge when designing reward functions. It is difficult to design reasonable reward functions for complex tasks, the debugging cost is high, and the feedback signal is not intuitive, which makes the agent unable to obtain the optimal operating trajectory and it is difficult to complete the task efficiently and accurately.

Method used

The trajectory optimization method of the agent based on human feedback and task objectives is adopted, and the mathematical expression of the task objective is designed to build a reinforcement learning environment. The sampling algorithm is used to select trajectory fragments and update the reward model according to human preferences, to determine whether the trajectory meets the task needs, and to extract the trajectory fragments from the advantageous and non-advantage containers to optimize the reward model, and finally conduct reinforcement learning and training the agent.

Benefits of technology

Reduce the burden of human feedback, improve the efficiency of reinforcement learning and training, improve the accuracy of the trajectory of the agent, reduce the cost of human participation, and enhance the accurate prediction ability of the reward model for task goals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962565B_ABST
    Figure CN119962565B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for optimizing the running trajectory of an agent based on human feedback and task objectives, including: designing a mathematical expression of the task objective according to the task objective; building a reinforcement learning environment according to the task; randomly sampling two different trajectory segments starting from the same state, and annotating the trajectory segments according to human preferences to update the reward model; starting from the same state, judging whether the current trajectory of the agent meets the task requirements according to the mathematical expression of the task objective, storing the trajectories that meet the task requirements in the advantage container, and storing the trajectories that do not meet the requirements in the non-advantage container; randomly extracting trajectory segments from the advantage container and the non-advantage container for optimizing the reward model; training the agent through reinforcement learning according to the optimized reward model, and outputting the optimal running trajectory of the agent. Compared with the prior art, the present invention can reduce the human feedback burden, improve the reinforcement learning training efficiency, and enhance the accuracy of the agent's running trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agent motion control, and more particularly to an optimization method for the running trajectory of an agent based on human feedback and task objectives. Background Art

[0002] In recent years, reinforcement learning, as an advanced method, has surpassed traditional models with more flexible and efficient end-to-end control capabilities. Especially in complex systems, it reduces the dependence on highly specialized mathematical models. Its application scope is extensive, covering multiple fields such as robotics, games, autonomous driving vehicles, and large language models. In the practical application of reinforcement learning, reward functions are crucial as they fundamentally guide the behavior of agents.

[0003] In complex tasks, especially in complex robot applications, designing reward functions is a highly challenging task that usually requires a large amount of domain expertise to effectively guide the behavior of agents. Under the same environment and task, different reward functions may have a significant impact on the behavior learning of agents. The current mainstream solutions are divided into manually designing reward functions, generating reward functions with large language models, and optimizing reward models based on preference-based reinforcement learning. Among them, the process of human-designed reward functions is as follows: 1) Manually design the reward function according to the task objective and prior knowledge; 2) Use the reinforcement learning environment to test the effectiveness of the reward function, and iteratively optimize the reward function by adjusting parameters and rules; 3) Verify the performance of the reward function in a specific task until the agent's behavior meets the expected goal;

[0004] The process of generating reward functions with large language models is as follows: 1) Use large language models to generate complex reward functions; 2) Iteratively improve the quality of the reward functions generated by large language models through the reward reflection process;

[0005] The process of generating reward models with preference-based reinforcement learning methods is as follows: 1) The preferences of experts for agent trajectory segments; 2) Optimize the reward prediction model.

[0006] However, the above several types of methods all have their inherent drawbacks. For example, the drawbacks of human-designed reward functions are: 1) Dependence on domain expert knowledge: Designing reward functions requires a deep understanding of task objectives and environmental characteristics, with high requirements for domain expertise; 2) Difficulty of complex tasks: For complex tasks, it is difficult to design comprehensive, reasonable reward functions that do not cause unexpected behaviors; 3) High debugging cost: Designing reward functions usually requires repeated trials and adjustments, which is time-consuming and costly; 4) Non-intuitiveness: In some tasks, the feedback signals of reward functions are not easily directly related to the behavior effects of agents, which may lead to difficulties in policy learning or deviation from the goal.

[0007] The problems with generating reward functions using large language models are as follows: 1) High cost, as calling a suitable large language model requires a large amount of money; 2) The ability to generate reward functions is limited by the capabilities of the large language model itself; 3) There is no guarantee that each generated reward function is better than its previous version.

[0008] The problems with generating reward models using preference-based reinforcement learning methods are as follows: 1) Heavy human burden, as a large amount of preference data for trajectory segment pairs is required; 2) The non-transitive preferences and attention limitations of humans lead to errors in trajectory segment preferences, resulting in a decline in the prediction ability of the reward model.

[0009] All of the above deficiencies can lead to the agent not being able to obtain the optimal operating trajectory in practical applications, making it difficult to ensure the efficient and accurate completion of related tasks. Summary of the Invention

[0010] The purpose of the present invention is to overcome the above-mentioned deficiencies of the existing technologies and provide an intelligent agent operating trajectory optimization method based on human feedback and task objectives, which can reduce the human feedback burden, improve the efficiency of reinforcement learning training, and enhance the accuracy of the intelligent agent's operating trajectory.

[0011] The purpose of the present invention can be achieved through the following technical solutions: An intelligent agent operating trajectory optimization method based on human feedback and task objectives, comprising the following steps:

[0012] S1. According to the task objective, design a mathematical expression of the task objective;

[0013] S2. Build a reinforcement learning environment according to the task;

[0014] S3. Starting from the same state, use a sampling algorithm to select two different trajectory segments, and label the trajectory segments according to human preferences to update the reward model;

[0015] S4. Starting from the same state, according to the mathematical expression of the task objective, judge whether the current trajectory of the agent meets the task requirements, and store the trajectory that meets the task requirements in the advantage container, and the trajectory that does not meet the requirements in the non-advantage container;

[0016] S5. Use a sampling algorithm to extract trajectory segments from the advantage container and the non-advantage container respectively for optimizing the reward model;

[0017] S6. Perform reinforcement learning training on the agent according to the optimized reward model, and output the optimal operating trajectory of the agent.

[0018] Furthermore, step S1 specifically converts the task requirements of the task objective into a mathematical form description for quantitatively evaluating whether a trajectory segment meets the task objective.

[0019] Further, the reinforcement learning environment established in step S2 includes a state space, an action space, and environmental dynamics, and is used to simulate the dynamic interaction between the agent and the environment.

[0020] Further, the specific process of step S3 is as follows:

[0021] S31. Use a sampling algorithm to sample trajectory segment pairs from the reinforcement learning environment;

[0022] S32. Show the trajectory segment pairs to human evaluators and select the trajectory that better meets the task objective based on preferences;

[0023] S33. Convert the human feedback into preference data for updating the parameters of the reward model;

[0024] S34. Repeat the above steps S31 - S33 until the performance of the reward model reaches the preset standard.

[0025] Further, the trajectory segment is specifically a sequence of states, actions, and rewards generated after the agent executes a policy in the reinforcement learning environment, where the policy is specifically a rule or mapping function for the agent to select actions under specific states.

[0026] Further, the preference data is specifically the priority information expressed by human evaluators for the selection of the trajectory segment that better meets the task objective between two trajectory segments.

[0027] Further, the reward model is specifically a model established by learning human preference data and is used to predict the relative superiority or inferiority of trajectory segments in terms of the task objective.

[0028] Further, the specific process of step S4 is as follows:

[0029] S41. Calculate the quantization index of the trajectory segment using the mathematical expression of the task objective;

[0030] S42. Store the trajectories with quantization indices meeting the preset threshold in the advantage container; store the trajectories with quantization indices not reaching the preset threshold in the non - advantage container.

[0031] Further, the specific process of step S5 is as follows:

[0032] S51. Use a sampling algorithm to randomly sample high - quality trajectory segments from the advantage container and contrastive trajectory segments from the non - advantage container to form training data;

[0033] S52. Use the training data to optimize the reward model through supervised learning.

[0034] Further, the specific process of step S6 is as follows:

[0035] S61. Update the agent's policy based on the optimized reward model;

[0036] S62. Verify the new policy in the reinforcement learning environment and collect new trajectory segments;

[0037] S63. Optimize the reward model using the feedback information of the new trajectory segments;

[0038] S64. Repeat the above steps S61 - S63 until the policy performance of the agent converges, and output the running trajectory of the current agent, which is the optimal running trajectory.

[0039] Compared with the prior art, the present invention has the following advantages:

[0040] The present invention first designs a mathematical expression of the task objective according to the task objective; then builds a reinforcement learning environment according to the task; starting from the same state, selects two different trajectory segments using a sampling algorithm and labels the trajectory segments according to human preferences to update the reward model; then starting from the same state, according to the mathematical expression of the task objective, determines whether the current trajectory of the agent meets the task requirements, and stores the trajectories that meet the task requirements in the advantage container, and the trajectories that do not meet the requirements are stored in the non - advantage container; then uses the sampling algorithm to extract trajectory segments from the advantage container and the non - advantage container respectively to optimize the reward model; finally, performs reinforcement learning training on the agent according to the optimized reward model, and outputs the optimal running trajectory of the agent. Thereby, it can reduce the human feedback burden, improve the reinforcement learning training efficiency, and enhance the accuracy of the agent's running trajectory.

[0041] The present invention designs a mathematical expression of the task objective according to the task objective, that is, transforms the task requirements into a mathematical form description, which is used for subsequent quantitative evaluation of whether the trajectory segment meets the task objective, can reduce the heavy dependence on expert preference data, thereby reducing the human preference feedback requirements and lowering the human participation cost. And by introducing additional task - related information through the mathematical expression of the task objective, it can effectively alleviate the non - transitivity and error effects in preference feedback, making the training process more robust.

[0042] The present invention shows random - sampled trajectory segment pairs to human evaluators, selects the trajectories that are more in line with the task objective based on preferences, and then converts human feedback into preference data to update the parameters of the reward model. Thus, it combines the task objective with human feedback, and further optimizes the reward model using the mathematical expression of the task objective, making it more accurately guide the agent's behavior, thereby improving the quality of the reward model.

[0043] The present invention first annotates trajectory segments according to human preferences and updates the reward model, then determines whether the trajectory of the agent meets the task requirements, stores the trajectories that meet the task requirements in the advantage container, and stores the trajectories that do not meet the requirements in the non-advantage container; then extracts trajectory segments from the advantage container and the non-advantage container to optimize the reward model, realizing a phased learning process that combines human feedback and task objectives, making the training more efficient and conducive to the gradual optimization of the reward model and the policy. This reinforcement learning framework based on human feedback and task objectives is applicable to diverse and highly complex task environments, especially in the field of robotics, and can effectively improve the task completion efficiency and accuracy of the agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic flow chart of the method for optimizing the running trajectory of an agent based on human feedback and task objectives according to the present invention;

[0045] Figure 2 It is a schematic diagram of the process of training an agent by reinforcement learning in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] Embodiment

[0048] As Figure 1 shown, a method for optimizing the running trajectory of an agent based on human feedback and task objectives includes the following steps:

[0049] S1. According to the task objective, design a mathematical expression of the task objective, specifically converting the task requirements into a mathematical description for quantitatively evaluating whether the trajectory segment meets the task objective in step S4 later;

[0050] S2. Build a reinforcement learning environment according to the task. The reinforcement learning environment refers to a dynamic system that simulates the interaction between the agent and the environment, including components such as the state space, action space, and environment dynamics;

[0051] S3. Starting from the same state, use a sampling algorithm to select two different trajectory segments, and annotate the trajectory segments according to human preferences to update the reward model;

[0052] S4. Starting from the same state, according to the mathematical expression of the task objective, determine whether the current trajectory of the agent meets the task requirements, and store the trajectories that meet the task requirements in the advantage container, and store the trajectories that do not meet the requirements in the non-advantage container;

[0053] Among them, the dominant container is specifically a data structure for storing trajectory segments evaluated as meeting the task objectives, and the non-dominant container is specifically a data structure for storing trajectory segments evaluated as not meeting the task objectives;

[0054] S5. Use a sampling algorithm to extract trajectory segments from the dominant container and the non-dominant container respectively for optimizing the reward model;

[0055] S6. Perform reinforcement learning training on the agent according to the optimized reward model, and output the optimal running trajectory of the agent.

[0056] This embodiment applies the above solution, pre-designs the mathematical expression of the task objective and builds a reinforcement learning environment: analyzes the task objective and its requirements, converts the key indicators or characteristics of task completion into a mathematical expression in mathematical form. Among them, the mathematical expression should have sparsity and can evaluate the task completion situation through high-level features of the trajectory; in addition, defines the environment model of reinforcement learning, including the state space, action space and dynamic rules of the environment; and determines the initial conditions and termination conditions according to the task objective, and sets the interaction mode between the agent and the environment.

[0057] This embodiment completes the reinforcement learning training process of the agent, as Figure 2 shown. First, initialize the policy and the reward model. Among them, the policy refers to the rule or mapping function for the agent to select actions in a specific state, and the reward model refers to a model established by learning human preference data and is used to predict the relative superiority or inferiority of trajectory segments in terms of task objectives.

[0058] After that, randomly sample trajectory segment pairs (that is, including two trajectory segments) from the reinforcement learning environment, show the trajectory segment pairs to human evaluators, and select the trajectory that better meets the task objective based on preference; convert the human feedback into preference data to update the parameters of the reward model; repeat the above process until the performance of the reward model reaches the expected standard. Among them, the trajectory segment refers to a sequence of states, actions and rewards generated by the agent after executing the policy in the reinforcement learning environment; the preference data refers to the priority information expressed by the human evaluator for the choice of the trajectory that better meets the task objective between two trajectory segments.

[0059] This embodiment randomly selects two trajectory segments and invites experts to label the preference of the superiority or inferiority of the trajectory according to the understanding of the task objective; then inputs the labeled preference data into the reward prediction model, and uses a preference learning algorithm (such as a contrast loss function) to optimize the model parameters, so that the reward model gradually learns the experts' understanding of the task.

[0060] Starting from the same state again, calculate the quantization metrics of the trajectory segments using the mathematical expression of the task objective; store the trajectories with quantization metrics meeting the preset threshold in the advantage container and those not reaching the preset threshold in the non-advantage container; then randomly sample high-quality trajectory segments from the advantage container and contrastive trajectory segments from the non-advantage container to form training data; use the training data to optimize the reward model through supervised learning; in this embodiment, starting from the same state, multiple trajectories are generated, and the task completion degree of each trajectory is evaluated using the pre-designed mathematical expression of the task objective; according to the task completion degree, the trajectories meeting the task requirements are stored in the advantage container, and those not meeting the task requirements are stored in the non-advantage container; this process does not rely on expert feedback and can autonomously classify trajectories using task-related structured information. In addition, in this embodiment, trajectory segments are extracted from the advantage container and the non-advantage container to form contrast data of high-task-completion-degree trajectories and low-task-completion-degree trajectories; the trajectory pairs are input into the reward prediction model, and the reward model is further optimized through contrastive learning to make up for the deficiency of preference feedback data. This stage can further enhance the accurate prediction ability of the reward model for the task objective.

[0061] Finally, based on the optimized reward model, update the agent's policy; and verify the new policy in the reinforcement learning environment and collect new trajectory segments; use the feedback information of the new trajectory segments to further optimize the reward model; repeat the above process until the policy performance of the agent converges, that is, the state where the agent's policy is stable in performance on the task objective and no longer changes significantly after multiple optimizations. In this embodiment, the reward signal generated by the optimized reward model is used to guide the training of the agent's policy by the reinforcement learning algorithm; the agent continuously adjusts its policy through interaction with the environment and finally learns the optimal policy that can complete the task objective; by verifying the performance of the agent's policy, it is ensured that its behavior meets the task requirements and is applicable to the actual application scenario. Thus, effectively combining human feedback and task objectives can reduce the dependence on expert preferences, improve the prediction accuracy of the reward model, and make the agent perform more efficiently and robustly in diverse tasks, ensuring that the agent can complete relevant tasks efficiently and accurately according to the optimal operation trajectory.

Claims

1. A method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives, characterized in that: The following steps are involved: S1. Design a mathematical expression of the task objective according to the task objective; S2, build a reinforcement learning environment based on the task; S3, starting from the same state, a sampling algorithm is used to select two different trajectory segments, and the trajectory segments are labeled according to human preferences to update the reward model; S4, starting from the same state, according to the mathematical expression of the task goal, determine whether the current trajectory of the agent meets the task requirements, and store the trajectory that meets the task requirements into the advantage container, and the trajectory that does not meet the requirements is stored in the non-advantage container; S5, using a sampling algorithm to extract trajectory segments from the dominant container and the non-dominant container respectively for optimizing the reward model; S6. Perform reinforcement learning training on the intelligent agent according to the optimized reward model, and output the optimal running trajectory of the intelligent agent; The specific process of step S3 is: S31, using a sampling algorithm to sample trajectory segments from a reinforcement learning environment; S32, present trajectory segment pairs to human evaluators and select the trajectory that better meets the task goal based on preference; S33, converting human feedback into preference data for updating the parameters of the reward model; S34, repeating the above steps S31 to S33 until the performance of the reward model reaches the preset standard; The specific process of step S5 is: S51, using a sampling algorithm to randomly sample high-quality trajectory segments from the dominant container and to sample comparative trajectory segments from the non-dominant container to form training data; S52. Use the training data to perform supervised learning optimization on the reward model.

2. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The step S1 specifically converts the task requirements of the task objectives into a mathematical description, which is used to quantitatively evaluate whether the trajectory segment meets the task objectives.

3. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The reinforcement learning environment constructed in step S2 includes a state space, an action space and environmental dynamics, which are used to simulate the dynamic interaction between the agent and the environment.

4. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The trajectory segment is specifically a sequence of states, actions, and rewards generated after the agent executes a strategy in a reinforcement learning environment, wherein the strategy is specifically a rule or mapping function for the agent to select an action in a specific state.

5. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The preference data is specifically priority information expressed by a human evaluator regarding a selection between two trajectory segments that better meets a task objective.

6. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The reward model is specifically a model established by learning human preference data, and is used to predict the relative merits of trajectory segments in terms of task objectives.

7. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 2, characterized in that: The specific process of step S4 is as follows: S41, calculating the quantitative index of the trajectory segment using the mathematical expression of the task goal; S42, storing the trajectory whose quantitative index meets the preset threshold into the advantage container; Trajectories whose quantitative indicators do not reach the preset threshold are stored in the non-dominant container.

8. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 4, characterized in that: The specific process of step S6 is as follows: S61. Update the agent's strategy based on the optimized reward model. S62, verify the new strategy in the reinforcement learning environment and collect new trajectory segments; S63, optimizing the reward model using feedback information of the new trajectory segment; S64, repeat the above steps S61 to S63 until the strategy performance of the agent reaches convergence, and output the running trajectory of the current agent, which is the optimal running trajectory.

Citation Information

Patent Citations

  • Crawler automatic driving method fusing human feedback information and deep reinforcement learning

    CN117032208A

  • Electric propeller trajectory tracking control method and system based on artificial intelligence

    CN118466227A