Action Space Alignment Method and System for Embodied Agents Based on Large Language Models
The dual alignment method using PEFT and retrieval-based generation effectively addresses the inefficiencies and biases in existing action space alignment for embodied agents, ensuring safe and effective action generation with reduced computational costs and improved interpretability.
Patent Information
- Application Number
- CN202411116365.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-14
AI Technical Summary
The existing embodied agents based on large language models have format illusions and state illusions in action space alignment, and traditional reinforcement learning methods require customized reward functions and high cost and sparse reward problems, resulting in inconsistent behavior and inefficiency of agents.
The method of combining internal alignment and external alignment is adopted to ensure that the actions of the agent output are aligned with a safe and effective action set through efficient parameter fine-tuning and search generation based methods, including efficient fine-tuning of Q-Lora parameters and ROUGE-L score calculation, as well as In-Context Learning to enhance the adaptability and output accuracy of the model.
It improves the flexibility and resource efficiency of the model, reduces the possibility of hallucination phenomena, enhances the interpretability and controllability of the model, and ensures the effectiveness and safety of the actions generated by the agent.
Smart Images

Figure CN119005359B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for aligning the action space of an embodied agent based on a large language model, and belongs to the technical field of reasoning and planning. Background Art
[0002] In recent years, representative large language models (LLMs) such as GPT-4, Claude-3, Gemini, and Llama have demonstrated excellent performance in a wide range of natural language processing and generation tasks. They not only excel in natural language processing but also have rapidly developed in fields such as reasoning and planning, bringing new vitality to embodied intelligence technology. Many recent studies have explored the capabilities of LLMs-based embodied agents in aspects such as intention understanding, logical reasoning, and task planning, demonstrating the potential of large language models as a core planner in embodied agents.
[0003] However, aligning the capabilities and behaviors of embodied agents with a safe and effective action space remains a major challenge. Research has shown that these model-based agents may inadvertently learn biases, discriminatory, or harmful content in the training data, resulting in outputs that deviate from human expectations. In particular, these embodied agents may produce misleading "hallucination" behaviors, generating false or unfounded content. Specifically, the action "hallucinations" of agents mainly include format hallucinations, that is, an incorrect control action; and state hallucinations, that is, actions that cannot be executed based on incorrect state judgments. For example Figure 1 as shown, an unaligned agent may generate a malformed action instruction "gotocoffeetable 1", while the correct instruction format should include a space between "go" and "to". This formatting error causes the agent to fail to reach "coffeetable 1" successfully and further triggers a state hallucination, generating an unexecutable action instruction "put vase 3in / on coffeetable 1". Since the agent is not actually at the location of "coffeetable 1", it cannot complete this action. Therefore, it is particularly important to adopt effective alignment techniques to ensure that the behaviors of LLMs-based agents are consistent with a feasible and safe operation space.
[0004] The currently widely adopted model alignment method is mainly based on the following alignment paradigm: first is supervised fine-tuning (SFT), followed by reinforcement learning (RL). This method first fine-tunes the model on domain-specific training data, aiming to improve its practicality by enhancing the model's ability to follow instructions and alleviate hallucinations to a certain extent. Nevertheless, when faced with unfamiliar inputs, a model relying solely on supervised fine-tuning may still give incorrect answers.
[0005] In an embodied environment, implementing action space alignment for large language model-based embodied agents using reinforcement learning (RL) faces several major challenges:
[0006] (1) Customized reward functions are required: RL usually needs to customize reward functions for different environments, which may require collecting additional datasets or performing complex manual annotations. This step not only increases the burden of preparatory work but also may introduce human biases.
[0007] (2) High computational resource consumption: RL training usually requires a large amount of computational resources. Since the scale of a general reward model may be comparable to that of the language model itself, this makes reinforcement learning training particularly expensive in terms of resource consumption, and the training process often lacks stability, which further increases the training difficulty and cost.
[0008] (3) The problem of sparse rewards exists: In an embodied environment, an agent may encounter the problem of sparse rewards, that is, it is difficult to obtain rewards for correct action or behavior sequences through direct feedback. This situation is not conducive to RL training because it will lead to low learning efficiency and it is difficult for the agent to find an effective learning path in a complex environment;
[0009] To more effectively ensure that the model output is consistent with human preferences, many studies have added a reinforcement learning stage based on human feedback after supervised fine-tuning. For example, InstructGPT, RAFT, and Constitutional AI, etc., use human feedback to perform reinforcement learning to train the model to learn and understand human preferences. This method trains a reward model to guide the learning process based on human evaluations of the model output, so as to improve the performance and output quality of the model. Summary of the Invention
[0010] To solve the above problems, the present invention provides a method and system for action space alignment of large language model-based embodied agents, which uses internal alignment and external alignment in cooperation, not only ensuring the effectiveness of the actions generated by the agent, but also enhancing the overall performance of the model. The internal and external alignment methods are used to alleviate the action hallucination problem of LLM-based embodied agents.
[0011] The technical solution of the present invention is: In the first aspect, the present invention provides a method for action space alignment of large language model-based embodied agents, including internal alignment and external alignment;
[0012] In the internal alignment, a parameter-efficient fine-tuning method is used to promote the efficient model adaptation of the agent;
[0013] In the external alignment, a retrieval-based generation method is used to ensure that the actions output by the agent are aligned with a safe and effective action set.
[0014] Furthermore, in the internal alignment, the method for efficiently fine-tuning parameters to promote the efficient model adaptation of the agent specifically includes:
[0015] The Q-Lora parameter-efficient fine-tuning method PEFT is utilized to promote the efficient model adaptation of the agent;
[0016] The goal of internal alignment is to train an llm-based embodied agent π on the embodied dataset θ , which learns the ability of expert demonstrations in solving the embodied inference task x task ; The expert demonstration trajectory segments are presented in the form of action-environment state pairs, where each pair of data represents the corresponding environment state s t after the execution of the action a t ; Suppose let τ t = [(a0, s0), (a1, s1), (a2, s2),..., (a t-1 , s t-1 )] represent the trajectory segment at time t; The goal of supervised fine-tuning is to learn to generate an action sequence {a t ~ π θ (x task , τ t )} to efficiently complete the task; In short, this is to seek to find a set of parameters θ * , such that the loss function θ under the guidance of the embodied agent π is minimized;
[0017]
[0018] The cross-entropy loss is selected as the loss function which measures the difference between π θ and the actions generated by the expert demonstration, represents the expected value or expected loss under the policy of the embodied agent π θ , Θ represents the parameter space, containing all possible parameter sets θ.
[0019] Furthermore, in the external alignment, the method based on retrieval generation is used to ensure that the actions output by the agent are aligned with the set of safe and effective actions, specifically including:
[0020] In the method based on retrieval generation, the goal of retrieval-augmented generation is to select the optimal action from a set of safe and effective actions as the final output of the model; When the agent π θ after internal alignment generates a candidate action a c , the first step of external action space alignment is to calculate the candidate action a cThe ROUGE-L score between each valid action v in the valid action space V, where the ROUGE-L score is used to evaluate the similarity between two sentences and reflects a c The similarity with all valid actions v; the calculation formula of the ROUGE-L score is as follows:
[0021]
[0022] Here, LCS(a c , v) represents the length of the longest common subsequence between the candidate action a c and the valid action v, and len(a c ) is the length of the candidate action a c ; by calculating the ROUGE-L score between a c and the valid action v, the action with the highest score is selected as the final output of the model; if multiple valid actions have the highest and equal similarity values with a c , the selected action set C is formed.
[0023] Furthermore, after obtaining the selected action set C, a policy model π p is introduced to select the best action; the In-Context Learning (ICL) of the large language model is used for action space alignment; the introduction of ICL aims to enhance the model's performance on new tasks by modifying the input query of the model, rather than directly updating the model weights; the currently retrieved valid action set C is used as an external input, allowing the LLM to obtain information about the physical world outside the execution context, thereby anchoring the LLM's response to the retrieved actions and reducing the possibility of hallucination; the final action is represented as:
[0024] a f = π p (x task , τ t , C)
[0025] where x task represents the embodied reasoning task, and τ t represents the trajectory segment at time t.
[0026] In a second aspect, the present invention provides an action space alignment system for an embodied intelligent agent based on a large language model, including a module for executing the method described in the first aspect above.
[0027] In a third aspect, the present invention provides a processor for running a program, wherein when the program runs, it executes the action space alignment method for an embodied intelligent agent based on a large language model described in any item of the first aspect.
[0028] Fourthly, the present invention provides a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the action space alignment method of the embodied intelligent agent based on the large language model described in the first aspect.
[0029] The beneficial effects of the present invention are as follows:
[0030] 1. The newly proposed internal and external alignment strategy of the present invention shows higher flexibility and resource efficiency compared with the traditional reinforcement learning (RL)-based alignment method, and contributes to improving the model interpretability. Specifically, the core advantages are as follows:
[0031] First, it eliminates the need to design a reward function and perform reinforcement learning (RL) training for specific embodied tasks.
[0032] Second, the selection of the policy model within the framework is flexible. It can be either a fine-tuned open-source small model (with about 7B model parameters) or a powerful commercial large language model.
[0033] Third, the effective action space retrieval can generate actions by retrieving evidence from the physical environment (such as a safe and effective action set), naturally reducing the risk of generating hallucinated content. Finally, the retrieval-based generation policy can determine the source of the answers generated by the LLM, enhancing the interpretability and transparency of the model decision-making process.
[0034] 2. The present invention not only ensures the effectiveness of the actions generated by the intelligent agent, but also enhances the overall performance of the model, proving its effectiveness. In addition, the method of the present invention requires lower computing resources: during internal alignment, an efficient parameter fine-tuning method is utilized; while in external alignment, compared with the traditional RL-based model alignment technology, the retrieval-based external alignment shows higher efficiency and stability. Moreover, it improves the controllability and interpretability of the model output. The internal and external alignment methods provide an effective solution to alleviate the action hallucination of the LLM-based embodied body. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is the internal and external alignment framework diagram in the present invention;
[0036] Figure 2 It is an example of the action space alignment of the embodied intelligent agent based on the large language model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] Embodiment 1: As Figure 1-2 shown, in the first aspect, the action space alignment method of the embodied intelligent agent based on the large language model in the present invention includes internal alignment and external alignment;
[0038] In the internal alignment, the parameter-efficient fine-tuning method is used to promote the efficient model adaptation of the agent;
[0039] The parameter-efficient fine-tuning method for promoting the efficient model adaptation of the agent specifically includes: as Figure 2 shown. Specifically, in the internal alignment, in order to effectively balance the computational cost and prevent the catastrophic forgetting problem in the base model; the Q-Lora parameter-efficient fine-tuning method PEFT is utilized to promote the efficient model adaptation of the agent;
[0040] The goal of internal alignment is to train an LLM-based embodied agent π on the embodied dataset θ , which learns the ability of expert demonstrations to solve the embodied reasoning task x task ; The expert demonstration trajectory segments are presented in the form of action-environment state pairs, where each pair of data represents the corresponding environment state s t after executing the action z t ; Suppose let τ t = [(a0, s0), (a1, s1), (a2, s2),...,(a t-1 , s t-1 )] represent the trajectory segment at time t; The goal of supervised fine-tuning is to learn to generate an action sequence {a t ∼ π θ (x task , τ t )} to efficiently complete the task; In short, this is to seek to find a set of parameters θ * such that the loss function θ under the guidance of the embodied agent π is minimized;
[0041]
[0042] The cross-entropy loss is selected as the loss function which measures the difference between π θ and the actions generated by the expert demonstration, denotes the expected value or expected loss under the policy of the embodied agent π θ , Θ represents the parameter space, containing all possible parameter sets θ. This method of internal alignment not only reduces the action hallucination of the agent, but also allows the embodied agent to make reasonable actions in the new environment.
[0043] In the external alignment, this is different from the traditional method of updating model parameters using reinforcement learning. At this stage, the method based on retrieval generation is used to ensure that the actions output by the agent are aligned with the set of safe and effective actions. Specifically, it includes:
[0044] In the retrieval-based generation method, the goal of retrieval-augmented generation is to select the optimal action from a set of safe and effective actions as the final output of the model; when the agent π after internal alignment θ generates a candidate action z c When, the first step of external action space alignment is to calculate the ROUGE-L score between the candidate action z c and each valid action v in the valid action space V, where the ROUGE-L score is used to evaluate the similarity between two sentences, reflecting z c The similarity with all valid actions v; the calculation formula of the ROUGE-L score is as follows:
[0045]
[0046] Here, LCS(z c , v) represents the length of the longest common subsequence between the candidate action z c and the valid action v, len(a c ) is the length of the candidate action a c ; by calculating the ROUGE-L score between a c and the valid action v, select the action with the highest score as the final output of the model; if multiple valid actions have the highest and equal similarity values with a c , then form the set of candidate actions C.
[0047] Furthermore, after obtaining the set of candidate actions C, introduce the policy model π p to select the best action; utilize the In-Context Learning of the large language model, that is, the ICL ability for action space alignment; the introduction of ICL aims to enhance the performance of the model on new tasks by modifying the input query of the model, rather than directly updating the model weights; compared with simple ICL prompt engineering, the results generated by retrieval-based ICL are more accurate and the possibility of model hallucinations is lower. This is because taking the currently retrieved set of valid actions C as the external input allows the LLM to obtain information about the physical world outside the execution context, thus anchoring the response of the LLM to the retrieved actions and reducing the possibility of hallucination phenomena; the final action is represented as:
[0048] a f = π p (x task , τ t , C)
[0049] where, x task represents the embodied reasoning task, τ tIt represents a trajectory segment at time t. In this way, the model can achieve precise and practical action alignment in the field of embodied intelligence while maintaining high-quality output to ensure the safety and effectiveness of actions.
[0050] In a second aspect, the present invention provides an action space alignment system for an embodied intelligent agent based on a large language model, including a module for executing the method described in the first aspect above.
[0051] In a third aspect, the present invention provides a processor for running a program, wherein when the program runs, it executes the action space alignment method for an embodied intelligent agent based on a large language model described in any item of the first aspect.
[0052] In a fourth aspect, the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the action space alignment method for an embodied intelligent agent based on a large language model described in the first aspect.
[0053] Figure 2 This is an example of action space alignment for an embodied intelligent agent based on a large language model in the present invention. In the figure, the unaligned intelligent agent is on the left and the aligned intelligent agent is on the right. There are two typical action hallucinations: (a) Format hallucination, the intelligent agent generates an incorrectly formatted action, such as missing a space symbol in "goto coffeetable 1"; (b) State hallucination, when the intelligent agent misjudges its state and attempts an unexecutable action, such as placing an object in a place it has not reached yet. The aligned intelligent agent can alleviate the above two hallucinations. First, for format hallucination, the aligned intelligent agent will find the action with the highest similarity in the set of valid actions for replacement. All actions in the set of valid actions must be correctly formatted / valid. Second, for state hallucination, if the intelligent agent has not successfully reached in front of the table, there is no possibility of any table-related operations in the set of valid actions. The aligned intelligent agent (external alignment based on retrieval-augmented generation) ensures that the final action output will only come from the set of valid actions, so the present invention fundamentally solves the state environment problem.
[0054] To verify the effectiveness of this method, the present invention has conducted extensive experiments on the #Unseen and #Seen test tasks in the ALFWorld simulation environment. The tasks in #Seen are similar to the rooms where the agent is trained, but have different object positions, quantities, etc., while the #Unseen tasks involve performing new tasks in completely unseen rooms and items. In this environment, the agent needs to generate a series of actions based on instructions and text / visual image observations at each step to effectively complete the task. The agent has a maximum of 30 executable steps to complete the task. The present invention mainly evaluated two types of agents. First, the first type is the base model, that is, the model only fine-tuned with efficient parameters (first-stage alignment), and the second type is the alignment model, that is, the model undergoes two-stage alignment. Table 1 presents the completed experimental results.)
[0055] Table 1
[0056]
[0057] As shown in Item 1 of Table 1, the model after internal alignment (i.e., the SFT model) is endowed with the basic ability to complete embodied tasks. For example, the SR performance of OPT and Bloomz in unknown scenarios reaches 79.10% and 48.51% respectively. However, relying solely on efficient parameter fine-tuning, the model may generate ineffective "hallucination" actions, and the degree of mitigation is reflected by the LC score. In particular, Llama-2 performs poorly in LC and does not exceed 50% in both visible and unseen tasks. This highlights the limitations of relying solely on internal model alignment to ensure safe and effective action generation. Therefore, it is necessary to address these challenges through external alignment.)
[0058] After external alignment based on RAG, as shown in Item 2 of Table 1, the ability of all models to output effective actions has been significantly improved, and most metrics have experienced varying degrees of enhancement. By adopting RAG-based external alignment, we ensure that the actions generated by the LLM-based embodied agent are safe and effective, as evidenced by the LC of all models reaching 100%. In addition, with the improvement of the ability to output effective operations, the SR, GCS, and RDI metrics of all models have been significantly improved. For example, in the unseen scenario, as the LC increases, the SR of Llama increases sharply from 38.81% to 75.37%, the GCS increases from 47.51% to 81.34%, and the RDI decreases from 9.70% to 2.20%.
[0059] In summary, through extensive experiments conducted on multiple models, it has been verified that our method not only ensures the effectiveness of the actions generated by the agent but also enhances the overall performance of the model, demonstrating its effectiveness. In addition, our method requires lower computational resources: during internal alignment, we utilize an efficient parameter fine-tuning method. In external alignment, compared with traditional reinforcement learning-based model alignment techniques, retrieval-based external alignment shows higher efficiency and stability. Moreover, it improves the controllability and interpretability of the model output. The internal and external alignment methods provide an effective solution for alleviating the action hallucination of LLM-based embodied agents.
[0060] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. Action space alignment method for embodied agents based on large language models, characterized in that: Including internal alignment and external alignment; In the internal alignment, an efficient parameter fine-tuning method is used to promote the efficient model adaptation of the agent; In the external alignment, a retrieval-based generation method is used to ensure that the actions output by the agent are aligned with a safe and effective action set; In the internal alignment, the use of an efficient parameter fine-tuning method to promote the efficient model adaptation of the agent specifically includes: The Q-Lora efficient parameter fine-tuning method PEFT is utilized to promote the efficient model adaptation of the agent; The goal of internal alignment is to train an LLM-based embodied agent π on an embodied dataset θ , which learns the ability of expert demonstrations to solve the embodied inference task x task ; Expert demonstration trajectory segments are presented in the form of action-environment state pairs, where each pair of data represents the corresponding environment state s t after executing the action a t ; Suppose let τ t = [(a0, s0), (a1, s1), (a2, s2),...,(a t-1 , s t-1 )] represent the trajectory segment at time t; The goal of supervised fine-tuning is to learn to generate an action sequence {a t ~ π θ (c task , τ t )} to efficiently complete the task; In short, this is to seek a set of parameters θ * such that the loss function θ under the guidance of the embodied agent π is minimized; Select cross-entropy loss as the loss function It measures π θ and the difference between the actions generated by the expert demonstrations, representing the expected value or expected loss under the embodied agent π θ policy, Θ represents the parameter space containing all possible parameter sets θ; In the external alignment, the use of a retrieval-based generation method to ensure that the actions output by the agent are aligned with a safe and effective action set specifically includes: In the retrieval-based generation method, the goal of retrieval-augmented generation is to select the optimal action from a set of safe and effective actions as the final output of the model; when the agent π θ generates a candidate action a c At this time, the first step of external action space alignment is to calculate the ROUGE-L score between the candidate action a c and each valid action v in the valid action space V, where the ROUGE-L score is used to evaluate the similarity between two sentences, reflecting the similarity between a c and all valid actions v; the calculation formula of the ROUGE-L score is as follows: Here, LCS(a c , v) represents the length of the longest common subsequence between the candidate action a c and the valid action v, and len(a c ) is the length of the candidate action a c ; by calculating the ROUGE-L score between a c and the valid action v, the action with the highest score is selected as the final output of the model; if multiple valid actions have the highest and equal similarity values with a c , the set of candidate actions C is formed.
2. The action space alignment method of the embodied intelligent agent based on the large language model according to claim 1, characterized in that: After obtaining the selected action set C, introduce the policy model π p to select the best action; utilize the In-Context Learning (ICL) of the large language model to align the action space; the introduction of ICL aims to enhance the model's performance on new tasks by modifying the input query of the model, rather than directly updating the model weights; take the currently retrieved valid action set C as an external input, allowing the LLM to obtain information about the physical world outside the execution context, thereby anchoring the LLM's response to the retrieved actions and reducing the likelihood of hallucination phenomena; the final action is represented as: a f = π p (x task , τ t , C) where x task represents an embodied inference task, and τ t represents a trajectory segment at time t.
3. An action space alignment system for an embodied intelligent agent based on a large language model, characterized in that Including a module for performing the method according to any one of claims 1-2.
4. A processor, characterized in that, The processor is used to run a program, wherein, when the program runs, it executes the action space alignment method of the embodied agent based on the large language model according to any one of claims 1-2.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the action space alignment method of the embodied agent based on the large language model according to any one of claims 1-2.