Intelligent agent moving trajectory optimization method based on human feedback and task target

By designing mathematical expressions of task goals and building a reinforced learning environment, and updating the reward model in combination with human feedback, the operation trajectory of the agent is optimized, and the problems of complex design of reward functions and high debugging costs in the existing technology are solved, and efficient and precise optimization of the operation trajectory of the agent is achieved.

CN119962565AActive Publication Date: 2025-05-09FUDAN UNIVERSITY

Patent Information

Application Number
CN202510449875.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-09
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing technology relies on domain expert knowledge when designing reward functions. It is difficult to design reasonable reward functions for complex tasks, the debugging cost is high, and the feedback signal is not intuitive, which makes the agent unable to obtain the optimal operating trajectory.

Method used

The agent's operation trajectory optimization method based on human feedback and task goals is adopted, and the agent is trained through the mathematical expression of the task goals is designed, and the reward model is updated using sampling algorithms and human preference annotations, and the agent is trained through the optimized reward model.

Benefits of technology

Reduce the burden of human feedback, improve the efficiency of reinforcement learning and training, improve the accuracy of the trajectory of the agent, reduce the cost of human participation, and enhance the prediction accuracy of the reward model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962565A_ABST
    Figure CN119962565A_ABST
Patent Text Reader

Abstract

The invention relates to an agent moving trajectory optimization method based on human feedback and a task target, and the method comprises the steps: designing a task target mathematical expression according to the task target; building a reinforcement learning environment according to the task; starting from the same state, randomly sampling two different track fragments, and labeling the track fragments according to human preferences to update the reward model; starting from the same state, according to a task target mathematical expression, judging whether the current trajectory of the intelligent agent meets task requirements or not, storing the trajectory meeting the task requirements into a dominant container, and storing the trajectory not meeting the task requirements into a non-dominant container; randomly extracting track fragments from the dominant container and the non-dominant container for optimizing the reward model; and performing reinforcement learning to train the intelligent agent according to the optimized reward model, and outputting to obtain the optimal moving trajectory of the intelligent agent. Compared with the prior art, the method can reduce the feedback burden of human beings, improve the reinforcement learning training efficiency, and improve the accuracy of the moving trajectory of the intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent body motion control, and in particular to an intelligent body running trajectory optimization method based on human feedback and task objectives. Background Art

[0002] In recent years, reinforcement learning, as an advanced method, has surpassed traditional models with more flexible and efficient end-to-end control capabilities, especially reducing the reliance on highly specialized mathematical models in complex systems. Its application range is wide, covering many fields such as robots, games, self-driving vehicles, and large language models. In the practical application of reinforcement learning, reward functions are crucial, and they fundamentally guide the behavior of intelligent agents.

[0003] In complex tasks, especially in complex robotics applications, designing reward functions is an extremely challenging task that usually requires a lot of domain expertise to effectively guide the behavior of the agent. In the same environment and task, different reward functions may have a significant impact on the behavioral learning of the agent. The current mainstream solutions are divided into manually designed reward functions, large language models to generate reward functions, and preference-based reinforcement learning to optimize reward models. Among them, the process of human reward function design is: 1) manually design reward functions based on task objectives and prior knowledge; 2) use reinforcement learning environments to test the effectiveness of reward functions, and iteratively optimize reward functions by adjusting parameters and rules; 3) verify the performance of the reward function in specific tasks until the agent behavior meets the expected goals; The process of generating a reward function from a large language model is as follows: 1) Generate a complex reward function using a large language model; 2) Iterate and improve the quality of the reward function generated by the large language model through a reward reflection process; The process of generating reward models using preference-based reinforcement learning methods is as follows: 1) expert preferences for agent trajectory segments; 2) optimization of the reward prediction model.

[0004] However, the above methods all have their inherent shortcomings. For example, the shortcomings of human-designed reward functions are: 1) Dependence on domain expert knowledge: Reward function design requires a deep understanding of task objectives and environmental characteristics, and has high requirements for domain expertise; 2) Difficulty of complex tasks: For complex tasks, it is difficult to design a comprehensive, reasonable and non-unexpected reward function; 3) High debugging cost: Reward function design usually requires repeated trials and adjustments, which is time-consuming and costly; 4) Non-intuitive: In some tasks, the feedback signal of the reward function is not easy to be directly associated with the behavior effect of the intelligent agent, which may lead to difficulty in strategy learning or deviation from the goal.

[0005] The problems with generating reward functions from large language models are: 1) It is costly, and calling a suitable large language model requires a lot of money; 2) The ability to generate reward functions is limited by the capabilities of the large language model itself; 3) There is no guarantee that each generated reward function is better than its previous version.

[0006] The problems with generating reward models using preference-based reinforcement learning methods are: 1) the human burden is heavy and requires a large amount of preference data for trajectory segments; 2) humans’ own intransitive preferences and attention limitations lead to errors in trajectory segment preferences, which reduces the predictive ability of the reward model.

[0007] The above defects will lead to the inability of the intelligent agent to obtain the optimal operation trajectory in practical applications, making it difficult to ensure the efficient and accurate completion of related tasks. Summary of the invention

[0008] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide an intelligent agent operation trajectory optimization method based on human feedback and task objectives, which can reduce the burden of human feedback, improve the efficiency of reinforcement learning training, and enhance the accuracy of the intelligent agent operation trajectory.

[0009] The object of the present invention can be achieved by the following technical solution: A method for optimizing the running trajectory of an intelligent agent based on human feedback and task objectives, comprising the following steps: S1. Design a mathematical expression of the task objective according to the task objective; S2, build a reinforcement learning environment based on the task; S3, starting from the same state, a sampling algorithm is used to select two different trajectory segments, and the trajectory segments are labeled according to human preferences to update the reward model; S4, starting from the same state, according to the mathematical expression of the task goal, determine whether the current trajectory of the agent meets the task requirements, and store the trajectory that meets the task requirements into the advantage container, and the trajectory that does not meet the requirements is stored in the non-advantage container; S5, using a sampling algorithm to extract trajectory segments from the dominant container and the non-dominant container respectively for optimizing the reward model; S6. Perform reinforcement learning training on the intelligent agent based on the optimized reward model, and output the optimal running trajectory of the intelligent agent.

[0010] Furthermore, the step S1 specifically converts the task requirements of the task objectives into a mathematical description, which is used to quantitatively evaluate whether the trajectory segment meets the task objectives.

[0011] Furthermore, the reinforcement learning environment constructed in step S2 includes a state space, an action space and environmental dynamics, which are used to simulate the dynamic interaction between the agent and the environment.

[0012] Furthermore, the specific process of step S3 is as follows: S31, using a sampling algorithm to sample trajectory segments from a reinforcement learning environment; S32, present trajectory segment pairs to human evaluators and select the trajectory that better meets the task goal based on preference; S33, converting human feedback into preference data for updating the parameters of the reward model; S34. Repeat the above steps S31 to S33 until the performance of the reward model reaches the preset standard.

[0013] Furthermore, the trajectory segment is specifically a sequence of states, actions, and rewards generated after the agent executes a strategy in a reinforcement learning environment, wherein the strategy is specifically a rule or mapping function for the agent to select an action under a specific state.

[0014] Furthermore, the preference data is specifically priority information expressed by a human evaluator on the selection of the two trajectory segments that better meets the task goal.

[0015] Furthermore, the reward model is specifically a model established by learning human preference data, and is used to predict the relative merits of trajectory segments in terms of task objectives.

[0016] Furthermore, the specific process of step S4 is as follows: S41, calculating the quantitative index of the trajectory segment using the mathematical expression of the task goal; S42, storing the trajectories whose quantitative indicators meet the preset threshold into a dominant container; storing the trajectories whose quantitative indicators do not meet the preset threshold into a non-dominant container.

[0017] Furthermore, the specific process of step S5 is as follows: S51, using a sampling algorithm to randomly sample high-quality trajectory segments from the dominant container and to sample comparative trajectory segments from the non-dominant container to form training data; S52. Use the training data to perform supervised learning optimization on the reward model.

[0018] Furthermore, the specific process of step S6 is as follows: S61. Update the agent's strategy based on the optimized reward model. S62, verify the new strategy in the reinforcement learning environment and collect new trajectory segments; S63, optimizing the reward model using feedback information of the new trajectory segment; S64. Repeat the above steps S61 to S63 until the strategy performance of the agent reaches convergence, and output the running trajectory of the current agent, which is the optimal running trajectory.

[0019] Compared with the prior art, the present invention has the following advantages: The present invention first designs a mathematical expression of the task objective according to the task objective; then builds a reinforcement learning environment according to the task; starting from the same state, a sampling algorithm is used to select two different trajectory segments, and the trajectory segments are labeled according to human preferences to update the reward model; starting from the same state, according to the mathematical expression of the task objective, it is judged whether the current trajectory of the intelligent agent meets the task requirements, and the trajectory that meets the task requirements is stored in the advantage container, and the trajectory that does not meet the requirements is stored in the non-advantage container; then a sampling algorithm is used to extract trajectory segments from the advantage container and the non-advantage container respectively, which are used to optimize the reward model; finally, reinforcement learning training is performed on the intelligent agent according to the optimized reward model, and the optimal running trajectory of the intelligent agent is output. This can reduce the burden of human feedback, improve the efficiency of reinforcement learning training, and improve the accuracy of the running trajectory of the intelligent agent.

[0020] The present invention designs the mathematical expression of the task goal according to the task goal, that is, converts the task requirement into a mathematical description, which is used for the subsequent quantitative evaluation of whether the trajectory segment meets the task goal, which can reduce the heavy reliance on expert preference data, thereby reducing the demand for human preference feedback and reducing the cost of human participation. In addition, by introducing additional task-related information through the mathematical expression of the task goal, the non-transitivity and error effects in the preference feedback can be effectively alleviated, making the training process more robust.

[0021] The present invention displays randomly sampled pairs of trajectory segments to human evaluators, and selects trajectories that are more in line with task objectives based on preferences. The human feedback is then converted into preference data to update the parameters of the reward model, thereby combining the task objectives with human feedback. The mathematical expression of the task objectives is used to further optimize the reward model, so that it can more accurately guide the behavior of the intelligent agent, thereby improving the quality of the reward model.

[0022] The present invention firstly labels the trajectory segments according to human preferences and updates the reward model, then determines whether the trajectory of the intelligent agent meets the task requirements, and stores the trajectories that meet the task requirements into the advantage container, and stores the trajectories that do not meet the requirements into the non-advantage container; then extracts the trajectory segments from the advantage container and the non-advantage container to optimize the reward model, realizing a staged learning process combining human feedback and task goals, making the training more efficient and conducive to the gradual optimization of the reward model and strategy. The reinforcement learning framework based on human feedback and task goals is suitable for diverse and highly complex task environments, especially in the field of robotics, and can effectively improve the efficiency and accuracy of the task completion of the intelligent agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a flow chart of the intelligent agent running trajectory optimization method based on human feedback and task objectives of the present invention; Figure 2 Schematic diagram of the process of training an intelligent agent through reinforcement learning in an embodiment. DETAILED DESCRIPTION

[0024] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] Example

[0026] like Figure 1 As shown, a method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives includes the following steps: S1. According to the task objective, a mathematical expression of the task objective is designed, specifically, the task requirement is converted into a mathematical description, which is used for quantitatively evaluating whether the trajectory segment meets the task objective in the subsequent step S4; S2. Build a reinforcement learning environment based on the task. The reinforcement learning environment refers to a dynamic system that simulates the interaction between the agent and the environment, including components such as state space, action space, and environmental dynamics; S3, starting from the same state, a sampling algorithm is used to select two different trajectory segments, and the trajectory segments are labeled according to human preferences to update the reward model; S4, starting from the same state, according to the mathematical expression of the task goal, determine whether the current trajectory of the agent meets the task requirements, and store the trajectory that meets the task requirements into the advantage container, and the trajectory that does not meet the requirements is stored in the non-advantage container; The dominant container is specifically a data structure for storing trajectory segments that are evaluated to meet the task objectives, and the non-dominant container is specifically a data structure for storing trajectory segments that are evaluated to not meet the task objectives; S5, using a sampling algorithm to extract trajectory segments from the dominant container and the non-dominant container respectively for optimizing the reward model; S6. Perform reinforcement learning training on the intelligent agent based on the optimized reward model, and output the optimal running trajectory of the intelligent agent.

[0027] This embodiment applies the above scheme, pre-designs the mathematical expression of the task goal and builds a reinforcement learning environment: analyzes the task goal and its requirements, converts the key indicators or characteristics of the task completion into a mathematical expression in mathematical form, wherein the mathematical expression should be sparse and be able to evaluate the task completion through the high-level features of the trajectory; in addition, defines the environmental model of reinforcement learning, including the state space, action space and dynamic rules of the environment; and determines the initial conditions and termination conditions according to the task goal, and sets the interaction mode between the agent and the environment.

[0028] This embodiment completes the reinforcement learning training process of the intelligent agent, such as Figure 2As shown, first initialize the strategy and reward model, where the strategy refers to the rule or mapping function for the agent to select actions in a specific state, and the reward model refers to the model established by learning human preference data, which is used to predict the relative merits of trajectory segments in terms of task objectives.

[0029] Then, trajectory segments pairs (i.e., including two trajectory segments) are randomly sampled from the reinforcement learning environment, the trajectory segment pairs are shown to the human evaluator, and the trajectory that better meets the task goal is selected based on preference; the human feedback is converted into preference data to update the parameters of the reward model; the above process is repeated until the performance of the reward model reaches the expected standard, where a trajectory segment refers to a sequence of states, actions, and rewards generated after the agent executes a strategy in the reinforcement learning environment; preference data refers to the priority information expressed by the human evaluator on the choice between the two trajectory segments that better meets the task goal.

[0030] This embodiment randomly selects two trajectory segments and invites experts to label the pros and cons of the trajectories according to their understanding of the task objectives. The labeled preference data is then input into the reward prediction model, and the model parameters are optimized using a preference learning algorithm (such as a contrastive loss function), so that the reward model gradually learns the expert's understanding of the task.

[0031] Starting from the same state again, the quantitative index of the trajectory segment is calculated using the mathematical expression of the task target; the trajectory whose quantitative index meets the preset threshold is stored in the dominant container, and the trajectory whose quantitative index does not meet the preset threshold is stored in the non-dominant container; then high-quality trajectory segments are randomly sampled from the dominant container, and comparative trajectory segments are sampled from the non-dominant container to form training data; the reward model is optimized for supervised learning using the training data; this embodiment starts from the same state, generates multiple trajectories, and uses the pre-designed mathematical expression of the task target to evaluate the task completion of each trajectory; according to the task completion, the trajectory that meets the task requirements is stored in the dominant container, and the trajectory that does not meet the task requirements is stored in the non-dominant container; this process does not rely on expert feedback, and the task-related structured information can be used to autonomously classify the trajectory. In addition, this embodiment extracts trajectory segments from the dominant container and the non-dominant container to form comparative data of high task completion trajectories and low task completion trajectories; the trajectory pairs are input into the reward prediction model, and the reward model is further optimized through comparative learning, thereby making up for the lack of preference feedback data. This stage can further enhance the ability of the reward model to accurately predict the task target.

[0032] Finally, based on the optimized reward model, the agent's strategy is updated; and the new strategy is verified in the reinforcement learning environment, and new trajectory segments are collected; the feedback information of the new trajectory segments is used to further optimize the reward model; the above process is repeated until the agent's strategy performance reaches convergence, that is, the agent's strategy is stable in performance on the task goal after multiple optimizations and no longer changes significantly. This embodiment uses the reward signal generated by the optimized reward model to guide the reinforcement learning algorithm to train the agent's strategy; the agent continuously adjusts the strategy by interacting with the environment, and finally learns the optimal strategy that can complete the task goal; by verifying the performance of the agent's strategy, it is ensured that its behavior meets the task requirements and is suitable for actual application scenarios. In this way, the effective combination of human feedback and task goals can reduce the dependence on expert preferences, improve the prediction accuracy of the reward model, and make the agent more efficient and robust in diversified tasks, ensuring that the agent can efficiently and accurately complete related tasks according to the optimal running trajectory.

Claims

1. A method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives, characterized in that: The following steps are involved: S1. Design a mathematical expression of the task objective according to the task objective; S2, build a reinforcement learning environment based on the task; S3, starting from the same state, a sampling algorithm is used to select two different trajectory segments, and the trajectory segments are labeled according to human preferences to update the reward model; S4, starting from the same state, according to the mathematical expression of the task goal, determine whether the current trajectory of the agent meets the task requirements, and store the trajectory that meets the task requirements into the advantage container, and the trajectory that does not meet the requirements is stored in the non-advantage container; S5, using a sampling algorithm to extract trajectory segments from the dominant container and the non-dominant container respectively for optimizing the reward model; S6. Perform reinforcement learning training on the intelligent agent based on the optimized reward model, and output the optimal running trajectory of the intelligent agent.

2. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The step S1 specifically converts the task requirements of the task objectives into a mathematical description, which is used to quantitatively evaluate whether the trajectory segment meets the task objectives.

3. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The reinforcement learning environment constructed in step S2 includes a state space, an action space and environmental dynamics, which are used to simulate the dynamic interaction between the agent and the environment.

4. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The specific process of step S3 is: S31, using a sampling algorithm to sample trajectory segments from a reinforcement learning environment; S32, present trajectory segment pairs to human evaluators and select the trajectory that better meets the task goal based on preference; S33, converting human feedback into preference data for updating the parameters of the reward model; S34. Repeat the above steps S31 to S33 until the performance of the reward model reaches the preset standard.

5. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 4, characterized in that: The trajectory segment is specifically a sequence of states, actions, and rewards generated after the agent executes a strategy in a reinforcement learning environment, wherein the strategy is specifically a rule or mapping function for the agent to select an action in a specific state.

6. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 4, characterized in that: The preference data is specifically priority information expressed by a human evaluator regarding a selection between two trajectory segments that better meets a task objective.

7. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 4, characterized in that: The reward model is specifically a model established by learning human preference data, and is used to predict the relative merits of trajectory segments in terms of task objectives.

8. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 2, characterized in that: The specific process of step S4 is as follows: S41, calculating the quantitative index of the trajectory segment using the mathematical expression of the task goal; S42, storing the trajectory whose quantitative index meets the preset threshold into the advantage container; Trajectories whose quantitative indicators do not reach the preset threshold are stored in the non-dominant container.

9. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 1, characterized in that: The specific process of step S5 is as follows: S51, using a sampling algorithm to randomly sample high-quality trajectory segments from the dominant container and to sample comparative trajectory segments from the non-dominant container to form training data; S52. Use the training data to perform supervised learning optimization on the reward model.

10. The method for optimizing the trajectory of an intelligent agent based on human feedback and task objectives according to claim 5, characterized in that: The specific process of step S6 is as follows: S61. Update the agent's strategy based on the optimized reward model. S62, verify the new strategy in the reinforcement learning environment and collect new trajectory segments; S63, optimizing the reward model using feedback information of the new trajectory segment; S64. Repeat the above steps S61 to S63 until the strategy performance of the agent reaches convergence, and output the running trajectory of the current agent, which is the optimal running trajectory.

Citation Information

Patent Citations

  • Crawler automatic driving method fusing human feedback information and deep reinforcement learning

    CN117032208A

  • Electric propeller trajectory tracking control method and system based on artificial intelligence

    CN118466227A

  • Federal reinforcement learning system, method and equipment for multi-agent trusted interactive decision control

    CN118982061A

Cited By

  • Unmanned ship cluster joint search and rescue method, device and equipment and storage medium

    CN121115794A

  • An unmanned ship cluster joint search and rescue method, device, equipment and storage medium

    CN121115794B