Multi-agent reinforcement learning exploration method and device based on large language model

CN118333183BActive Publication Date: 2026-09-25TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410433959.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2026-09-25
Estimated Expiration
2044-04-11

AI Technical Summary

Technical Problem

[0006]本申请提供一种基于大语言模型的多智能体强化学习探索方法及装置,以解决相关技术中的探索方法要么需要人类专业知识来设计子目标,要么在识别有用的子目标方面存在困难,因此将任务相关信息整合到子目标中时较为困难,且由于难以引入任务相关的先验信息等原因,探索效率较低,如何将大语言模型引入决策问题中,降低成本与时间,满足各种实际情境的应用,实现高效探索等问题

Benefits of technology

[0021]本申请第五方面实施例提供一种计算机程序产品,所述计算机程序被执行时,以用于实现如上的基于大语言模型的多智能体强化学习探索方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118333183B_ABST
    Figure CN118333183B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of large language models, in particular to a multi-agent reinforcement learning exploration method and device based on a large language model, wherein the method comprises the following steps: generating a key state discrimination function by using a large language model based on at least one preset prompt template; finding a task-related key state with explicit semantics and expression in a sampled trajectory based on the key state discrimination function; and obtaining a multi-agent reinforcement learning exploration result by taking the key state as prior information. The application can generate a key state discrimination function in a round of dialogue by using a large language model to perform subsequent key state recognition, introduce language-form knowledge of the large language model into a decision-making task, greatly reduce the cost caused by frequent calling of the large language model, and effectively promote efficient exploration of multi-agents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large language model technology, and in particular to a multi-agent reinforcement learning exploration method and apparatus based on large language models. Background Technology

[0002] Reinforcement learning is a cutting-edge research area in machine learning, aiming to find optimal policies for sequential decision-making problems. Multi-agent reinforcement learning further considers decision-making problems involving multiple agents, which has many real-world applications and is currently a popular research direction. Achieving efficient exploration is a relentless pursuit for researchers in reinforcement learning and multi-agent reinforcement learning. Large language models have shown great potential in reinforcement learning and multi-agent reinforcement learning, demonstrating remarkable capabilities in various downstream tasks. Increasingly, research is beginning to utilize the rich intrinsic knowledge and capabilities of large language models to solve decision-making problems.

[0003] In related technologies, common exploration methods include introducing random exploration or encouraging maximum diversity or novelty. While these methods have achieved good results in some scenarios, they exhibit significant redundant exploration due to a lack of effective task-related guidance. Other methods emphasize encouraging influential behavior during agent interactions, which may lead to unexpected alliances or require additional human prior knowledge. Furthermore, some research focuses on sub-goal-based methods, using sub-goals to guide effective exploration.

[0004] Introducing large language models into decision-making problems remains challenging. A major challenge lies in integrating the linguistic knowledge of large language models into specific tasks, which are typically represented in symbolic form. Creating language twins is one approach, but it requires significant manual effort and incurs substantial costs. Some research utilizes large language models as high-level planners, assuming the existence of predefined models that perform low-level control or can translate symbolic states. Despite significant progress, these approaches rely too heavily on potentially unavailable high-quality low-level controllers or translation models, limiting their applicability, particularly in diverse real-world contexts. Methods like reward design based on large language models or fine-tuning large language models are either only applicable to simple tasks or are time-consuming and require substantial training data and resources.

[0005] However, the exploration methods in related technologies either require human expertise to design sub-goals or have difficulties in identifying useful sub-goals. Therefore, it is difficult to integrate task-related information into sub-goals, and the exploration efficiency is low due to the difficulty in introducing task-related prior information. How to introduce large language models into decision-making problems, reduce costs and time, meet the application needs of various practical situations, and achieve efficient exploration is an urgent problem to be solved. Summary of the Invention

[0006] This application provides a multi-agent reinforcement learning exploration method and apparatus based on a large language model to address the problems in related technologies, such as the need for human expertise to design sub-goals, difficulties in identifying useful sub-goals, challenges in integrating task-related information into sub-goals, and low exploration efficiency due to the difficulty in incorporating task-related prior information. The application aims to introduce a large language model into decision-making problems to reduce costs and time, meet the application requirements of various practical scenarios, and achieve efficient exploration.

[0007] The first aspect of this application provides a multi-agent reinforcement learning exploration method based on a large language model, comprising the following steps: generating a key state discriminant function using a large language model based on at least one preset prompt template; searching for task-related key states with display semantics and expression in the sampled trajectory based on the key state discriminant function; and obtaining the multi-agent reinforcement learning exploration result by using the key states as prior information.

[0008] Optionally, in one embodiment of this application, the step of finding task-related key states with display semantics and expression in the sampled trajectory based on the key state discrimination function includes: receiving the state at any time step as an input variable, and outputting a Boolean value indicating whether the current input state belongs to its corresponding key state, so as to identify and label each state in the trajectory and obtain identification and labeling results; and determining the key state based on the identification and labeling results.

[0009] Optionally, in one embodiment of this application, obtaining the multi-agent reinforcement learning exploration results by treating the key states as prior information includes: introducing intrinsic rewards at each time step of the sub-trajectories of the trajectory to guide the training of the multi-agent based on the intrinsic rewards, and combining extrinsic rewards to obtain subspace-based hindsight intrinsic rewards; and using a tree structure to record the transition relationships between the key states based on the subspace-based hindsight intrinsic rewards for memory-based exploration.

[0010] Optionally, in one embodiment of this application, the expression for the intrinsic reward is: , in, Indicates that time step t is based on critical states. intrinsic reward. I Indicates an intrinsic reward indicator. represent The state at any given moment, It is the label of a certain sub-track. This is the key state corresponding to this sub-trajectory. It represents a distance metric, such as Manhattan distance; A subspace mapping function maps the complete state space to a subspace consisting of a subset of its elements. Among them This represents an element in the state space. This represents a subspace spanned by some elements in the state space. For example, if the state space has 5 dimensions, and only the 2nd and 3rd dimensions are related to the reward, then only these two dimensions are selected to calculate the distance. This represents the subspace mapping function.

[0011] Optionally, in one embodiment of this application, the hindsight intrinsic reward is: , in, This represents the intrinsic reward after time step t. Indicates that time step t is based on critical states. intrinsic reward. It refers to extrinsic reward. , These represent scaling factors for extrinsic and intrinsic rewards, respectively.

[0012] Optionally, in one embodiment of this application, the subspace-based hindsight intrinsic reward, which uses a tree structure to record the transition relationships between the key states for memory-based exploration, includes: based on the key state chain of the trajectory, searching for the corresponding branch in the key state memory tree to sample the next key state that satisfies the preset most likely condition; and using the next key state as the target of the last sub-trajectory in the trajectory to apply the hindsight intrinsic reward to the entire trajectory.

[0013] A second aspect of this application provides a multi-agent reinforcement learning exploration device based on a large language model, comprising: a generation module for generating a key state discrimination function using a large language model based on at least one preset prompt template; a search module for searching for task-related key states with display semantics and expression in a sampled trajectory based on the key state discrimination function; and an exploration module for obtaining multi-agent reinforcement learning exploration results by using the key states as prior information.

[0014] Optionally, in one embodiment of this application, the searching module includes: a receiving unit, configured to receive the state at any time step as an input variable, and output a Boolean value indicating whether the current input state belongs to its corresponding key state, so as to identify and label each state in the trajectory and obtain identification and labeling results; and a determining unit, configured to determine the key state based on the identification and labeling results.

[0015] Optionally, in one embodiment of this application, the exploration module includes: a training unit, configured to introduce intrinsic rewards at each time step of a sub-trajectory of the trajectory, so as to guide the training of the multi-agent based on the intrinsic rewards, and to obtain a subspace-based hindsight intrinsic reward in combination with extrinsic rewards; and a recording unit, configured to record the transition relationships between the key states using a tree structure based on the subspace-based hindsight intrinsic reward for memory-based exploration.

[0016] Optionally, in one embodiment of this application, the expression for the intrinsic reward is: , in, Indicates that time step t is based on critical states. intrinsic reward. I Indicates an intrinsic reward marker. represent The state at any given moment, It is a label for a certain sub-track. This is the key state corresponding to this sub-trajectory. It represents a distance metric, such as Manhattan distance; A subspace mapping function maps the complete state space to a subspace consisting of a subset of its elements. Among them This represents an element in the state space. This represents a subspace spanned by some elements in the state space. For example, if the state space has 5 dimensions, and only the 2nd and 3rd dimensions are related to the reward, then only these two dimensions are selected to calculate the distance. This represents the subspace mapping function.

[0017] Optionally, in one embodiment of this application, the hindsight intrinsic reward is: , in, This represents the intrinsic reward after time step t. Indicates that time step t is based on critical states. intrinsic reward. It refers to extrinsic reward. , These represent scaling factors for extrinsic and intrinsic rewards, respectively.

[0018] Optionally, in one embodiment of this application, the recording unit includes: a search subunit, used to search for a corresponding branch in the key state memory tree based on the key state chain of the trajectory, so as to sample the next key state that satisfies the preset most likely condition; and a determination subunit, used to take the next key state as the target of the last sub-trajectory in the trajectory, so as to apply the posterior intrinsic reward to the entire trajectory.

[0019] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-agent reinforcement learning exploration method based on a large language model as described in the above embodiments.

[0020] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multi-agent reinforcement learning exploration method based on a large language model.

[0021] A fifth aspect of this application provides a computer program product, which, when executed, is used to implement the above-described multi-agent reinforcement learning exploration method based on a large language model.

[0022] This application's embodiments can design prompt templates and utilize a large language model to generate key state discriminant functions to identify key states in the trajectory. Combined with a subspace-based hindsight intrinsic reward method, a key state memory tree, and corresponding exploration and planning methods, multi-agent exploration is completed, yielding accurate exploration results. This achieves the generation of key state discriminant functions in a single round of dialogue using a large language model for subsequent key state identification. It allows the introduction of linguistic knowledge from the large language model into decision-making tasks, significantly reducing the cost associated with frequent calls to the large language model. The subspace-based hindsight intrinsic reward and key state memory tree effectively promote efficient multi-agent exploration by enhancing reward density and enabling more organized exploration. This addresses the problems in related technologies where exploration methods either require human expertise to design sub-goals or face difficulties in identifying useful sub-goals, making it difficult to integrate task-related information into sub-goals. Furthermore, the difficulty in incorporating task-related prior information leads to low exploration efficiency. This addresses the challenges of introducing large language models into decision-making problems to meet the application needs of various practical scenarios, reducing costs and time while achieving efficient exploration.

[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0024] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating a multi-agent reinforcement learning exploration method based on a large language model, according to an embodiment of this application. Figure 2 This is a schematic diagram of a large language model structure according to an embodiment of this application; Figure 3 This is a schematic diagram illustrating an exploration based on key state guidance according to one embodiment of this application; Figure 4 This is a schematic diagram of the structure of a multi-agent reinforcement learning exploration device based on a large language model according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Attached image description: 10-Multi-agent reinforcement learning exploration device based on large language model; 100-Generation module, 200-Search module and 300-Exploration module; 501-Memory, 502-Processor and 503-Communication interface. Detailed Implementation

[0026] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0027] The following describes, with reference to the accompanying drawings, a multi-agent reinforcement learning exploration method and apparatus based on a large language model according to embodiments of this application. Addressing the challenges of exploration methods in the related technologies mentioned above, which either require human expertise to design sub-goals or struggle to identify useful sub-goals, thus making it difficult to integrate task-related information into sub-goals, and resulting in low exploration efficiency due to the difficulty in incorporating task-related prior information, this application provides a multi-agent reinforcement learning exploration method based on a large language model. In this method, prompt templates can be designed, and a key state discriminant function generated using a large language model can be used to identify key states in the trajectory. This is combined with a subspace-based hindsight intrinsic reward method, a key state memory tree, and corresponding exploration and planning methods to complete the multi-agent exploration and obtain accurate exploration results. This enables the generation of key state discriminant functions in a single round of dialogue using a large language model for subsequent key state identification. This allows the introduction of linguistic knowledge from the large language model into decision-making tasks, significantly reducing the cost associated with frequent calls to the large language model. Subspace-based hindsight intrinsic rewards and key state memory trees effectively promote efficient multi-agent exploration by enhancing reward density and promoting more organized exploration. This addresses the challenges of exploration methods in related technologies, which either require human expertise to design sub-goals or struggle to identify useful sub-goals, making it difficult to integrate task-related information into sub-goals and resulting in low exploration efficiency due to the difficulty in incorporating task-related prior information. The solution addresses how to introduce large language models into decision-making problems to meet the needs of various practical applications, reduce costs and time, and achieve efficient exploration.

[0028] Specifically, Figure 1 This is a flowchart illustrating a multi-agent reinforcement learning exploration method based on a large language model, provided as an embodiment of this application.

[0029] like Figure 1 As shown, this multi-agent reinforcement learning exploration method based on a large language model includes the following steps: In step S101, a key state discrimination function is generated using a large language model based on at least one preset prompt template.

[0030] As one possible approach, in this embodiment of the application, before exploring based on a large language model, at least one prompt template can be pre-defined for the large language model. For example... Figure 2The diagram shown is a schematic representation of a large language model structure according to an embodiment of this application, including a task prompt section, an answer section, and a feedback section. The task prompt section includes a prompt template and task information; the answer section includes "thinking + definition" and a key state discrimination function; and the feedback section includes self-checking and feedback.

[0031] For each exploration task, this embodiment of the application can populate the prompt template by collecting task descriptions and state forms as task information. To further improve the large language model, this embodiment of the application also designs an answer section in the large language model. That is, when the large language model is queried with a complete task prompt, the large language model will also return a structured answer, which mainly includes, but is not limited to, the thought process, key state definitions, and key state discrimination functions.

[0032] Furthermore, the thought process primarily involves the large language model articulating its understanding of the task and describing, in language, what it perceives as the key states within the task. Here, a key state can be understood as an intermediate state that has explicit semantics and expressive power relevant to the task. For example, if the task requires traveling from point A to point C, and point B is a necessary intermediate state, then B is a key state. The key state discrimination function can then identify and label these key states.

[0033] In addition, to improve the quality of the large language model's answers, this application also employs a self-checking feedback mechanism, including but not limited to allowing the large language model to reflect on whether its current answer is reasonable and ensuring that the discrimination function is executable through code executableness testing.

[0034] The embodiments of this application can utilize a large language model to generate key state discrimination functions, and at the same time design a series of related templates to improve the large language model, effectively improving the reliability of the large language model in this application, thereby ensuring its practical application capability.

[0035] Step S102: Based on the key state discrimination function, find the task-related key states with explicit semantics and expression in the sampled trajectory.

[0036] Based on the descriptions of other embodiments, it is understood that the big oracle model in the embodiments of this application includes a response section, which mainly includes, but is not limited to, the thought process, the definition of key states, and the key state identification function.

[0037] In this context, a key state can be understood as a task-related intermediate state with explicit semantics and expression. The key state identification function serves as the most important content of the answer template. Each discriminant function corresponds to a key state, and the key state identification function can be used to identify and label each key state in the trajectory.

[0038] For example, embodiments of this application can use a critical state discrimination function to find critical states. That is, during the search process, it is determined whether the current state is a critical state. If the current state is determined to be a critical state, a certain label is made. If the current state is identified as a non-critical state, another label can be made.

[0039] The following section further explains the process of using a key state recognition function to find task-related key states with specific semantics and expressions in the sampled trajectory.

[0040] Optionally, in one embodiment of this application, the process of finding task-related key states with explicit semantics and expression in the sampled trajectory based on the key state discrimination function includes: receiving the state at any time step as an input variable and outputting a Boolean value indicating whether the current input state belongs to its corresponding key state, so as to identify and label each state in the trajectory and obtain identification and labeling results; and determining the key state based on the identification and labeling results.

[0041] In actual implementation, when using the key state identification function to find key states in the sampling trajectory, the embodiments of this application may, but are not limited to, use Python functions to design key state discrimination functions to identify and label key states.

[0042] Specifically, Python's critical state discrimination function can accept the state at a certain time step as an input variable. After identifying the current input state, it can output a boolean value (1 or 0) to indicate whether the current input state belongs to its corresponding critical state. Finally, through these discrimination functions, each state in the trajectory can be identified and labeled to determine whether the input state is a critical state.

[0043] For example, if B is a key state identified by the large language model, then its corresponding recognition function determines the state at each time step. If it belongs to B, return 1; otherwise, return 0.

[0044] The embodiments of this application can use a key state discrimination function to identify and label key states, thereby finding key states that have display semantics and expression related to the task, which facilitates subsequent exploration using key states.

[0045] Step S103: The key states are used as prior information to obtain the multi-agent reinforcement learning exploration results.

[0046] In some embodiments, after the key state discrimination function generated by the large language model is used to automatically find key states in the sampled trajectory and identify and label them, the embodiments of this application can also use the labeled key states as prior information to guide efficient exploration and finally obtain the multi-agent reinforcement learning exploration results.

[0047] The following section further explains how this application utilizes key states as prior information to obtain multi-agent reinforcement learning exploration results.

[0048] Optionally, in one embodiment of this application, the key states are used as prior information to obtain the multi-agent reinforcement learning exploration results, including: introducing intrinsic rewards at each time step of the sub-trajectories of the trajectory to guide the training of the multi-agent based on the intrinsic rewards, and combining extrinsic rewards to obtain subspace-based hindsight intrinsic rewards; using a tree structure to record the transition relationships between key states based on subspace-based hindsight intrinsic rewards for memory-based exploration.

[0049] Optionally, in one embodiment of this application, the expression for the intrinsic reward can be: , in, Indicates that time step t is based on critical states. intrinsic reward. I Indicates an intrinsic reward marker. represent The state at any given moment, It is a label for a certain sub-track. This is the key state corresponding to this sub-trajectory. It represents a distance metric, such as Manhattan distance; A subspace mapping function maps the complete state space to a subspace consisting of a subset of its elements. Among them This represents an element in the state space. This represents a subspace spanned by some elements in the state space. For example, if the state space has 5 dimensions, and only the 2nd and 3rd dimensions are related to the reward, then only these two dimensions are selected to calculate the distance. This represents the subspace mapping function.

[0050] Optionally, in one embodiment of this application, the hindsight intrinsic reward can be represented as: , in, This represents the intrinsic reward after time step t. Indicates that time step t is based on critical states. intrinsic reward. It refers to extrinsic reward. , These represent scaling factors for extrinsic and intrinsic rewards, respectively.

[0051] In actual implementation, when using the marked key states as prior information to guide efficient exploration to obtain multi-agent exploration results, the embodiments of this application can also use the subspace-based hindsight intrinsic reward method for efficient exploration.

[0052] For example, such as Figure 3 The diagram illustrates an efficient exploration method for key state guidance in one embodiment of this application. After key states are marked on the trajectory, each trajectory is naturally divided into a series of sub-trajectories, such as... It can be divided into In this embodiment, the last state of each sub-trajectory can be considered as a key state, such as... Critical state , Critical state The subsequent time step is considered the objective of the sub-trajectory, and an intrinsic reward is introduced for each time step in the sub-trajectory based on this objective. The intrinsic reward equation can be expressed as:

[0053] in, Indicates that time step t is based on critical states. intrinsic reward. I Indicates an intrinsic reward marker. represent The state at any given moment, It is a label for a certain sub-track. This is the key state corresponding to this sub-trajectory. It represents a distance metric, such as Manhattan distance; A subspace mapping function maps the complete state space to a subspace consisting of a subset of its elements. Among them This represents an element in the state space. This represents a subspace spanned by some elements in the state space. For example, if the state space has 5 dimensions, and only the 2nd and 3rd dimensions are related to the reward, then only these two dimensions are selected to calculate the distance. This represents the subspace mapping function.

[0054] Furthermore, in the embodiments of this application, the intrinsic reward is positive if the target is closer after a move, otherwise a negative reward is obtained.

[0055] The introduction of hindsight intrinsic rewards can make rewards denser, better guiding the agent's training. This, combined with extrinsic rewards from the environment itself, further enhances the effect. The final hindsight intrinsic reward function can be expressed as:

[0056] in, This represents the intrinsic reward after time step t. Indicates that time step t is based on critical states. intrinsic reward. It refers to extrinsic reward. , These represent scaling factors for extrinsic and intrinsic rewards, respectively.

[0057] Optionally, in one embodiment of this application, based on the hindsight intrinsic reward of the subspace, a tree structure is used to record the transformation relationship between key states for memory-based exploration, including: based on the key state chain of the trajectory, searching for the corresponding branch in the key state memory tree to sample the next key state that satisfies the preset most likely condition; taking the next key state as the target of the last sub-trajectory in the trajectory to apply the hindsight intrinsic reward to the entire trajectory.

[0058] To further optimize the process, this application also proposes a Key States Memory Tree (KSMT), which uses a tree structure to record the transition relationships between key states to better organize memory-based exploration.

[0059] Based on the critical state memory tree, this application proposes a novel hybrid random exploration method: introducing two exploration strategies with high and low randomness respectively. When the reached critical state is a leaf node in the critical state memory tree, the agent can use the high randomness exploration strategy to attempt to find new critical states, thus deepening the critical state memory tree. When a non-leaf node is reached, the agent can choose the high randomness exploration strategy with a certain probability p to expand the tree's width, or choose the low randomness exploration strategy with a probability 1-p to successfully proceed to the next critical state.

[0060] In addition, such as Figure 3The diagram shown illustrates an efficient exploration method for key state guidance in one embodiment of this application. This application also proposes a novel planning method based on a key state memory tree: using the key state chain identified by the trajectory to search for corresponding branches in the key state memory tree. Here, the key state chain can be understood as the branch identified in the trajectory... Critical state , Critical state Then extract As a chain of critical states, the most likely next critical state can then be sampled as the target of the last sub-trajectory in the trajectory. For example, from Sampled from the child nodes of this branch As the last sub-trajectory in the trajectory The goal is to apply subspace-based hindsight intrinsic rewards to the entire trajectory, thereby enhancing the exploration guidance effect.

[0061] It should be noted that the construction of the key state memory tree in this embodiment is similar to that of a normal tree structure, that is, starting from a root node, the key state chain extracted from the new trajectory is used for updating.

[0062] The embodiments of this application can effectively promote efficient multi-agent exploration by enhancing the density of rewards and more organized exploration based on the hindsight intrinsic reward and key state memory tree of the subspace.

[0063] The multi-agent reinforcement learning exploration method based on a large language model proposed in this application can design prompt templates and use the large language model to generate key state discriminant functions to identify key states in the trajectory. Combined with a subspace-based hindsight intrinsic reward method, a key state memory tree, and corresponding exploration and planning methods, the multi-agent exploration is completed, yielding accurate exploration results. This achieves the generation of key state discriminant functions in a single round of dialogue using a large language model for subsequent key state identification. It allows the introduction of linguistic knowledge from the large language model into decision-making tasks, significantly reducing the cost associated with frequent calls to the large language model. The subspace-based hindsight intrinsic reward and key state memory tree effectively promote efficient multi-agent exploration by enhancing reward density and enabling more organized exploration. This solves the problems of related exploration methods, which either require human expertise to design sub-goals or face difficulties in identifying useful sub-goals, making it difficult to integrate task-related information into sub-goals and resulting in low exploration efficiency due to the difficulty in incorporating task-related prior information. This addresses the challenges of introducing large language models into decision-making problems to meet the application needs of various practical scenarios, reducing costs and time while achieving efficient exploration.

[0064] Next, referring to the accompanying drawings, a multi-agent reinforcement learning exploration device based on a large language model, according to an embodiment of this application, is described.

[0065] Figure 4 This is a schematic diagram of the structure of a multi-agent reinforcement learning exploration device based on a large language model according to an embodiment of this application.

[0066] like Figure 4 As shown, the multi-agent reinforcement learning exploration device 10 based on a large language model includes: a generation module 100, a search module 200, and an exploration module 300.

[0067] The generation module 100 is used to generate a key state discrimination function based on at least one preset prompt template using a large language model.

[0068] The search module 200 is used to search for task-related key states with explicit semantics and expressions in the sampled trajectory based on the key state discrimination function.

[0069] The exploration module 300 is used to obtain the exploration results of multi-agent reinforcement learning by taking key states as prior information.

[0070] Optionally, in one embodiment of this application, the searching module 200 includes a receiving unit and a determining unit.

[0071] The receiving unit is used to receive the state at any time step as an input variable and output a Boolean value indicating whether the current input state belongs to its corresponding key state, so as to identify and label each state in the trajectory and obtain the identification and labeling results.

[0072] The determination unit is used to determine the key states based on the identification and labeling results.

[0073] Optionally, in one embodiment of this application, the exploration module 300 includes a training unit and a recording unit.

[0074] The training unit is used to introduce intrinsic rewards at each time step of the sub-trajectories of the trajectory, so as to guide the training of multi-agents according to the intrinsic rewards, and to obtain the subspace-based hindsight intrinsic rewards by combining the extrinsic rewards.

[0075] The recording unit is used for subspace-based hindsight intrinsic rewards, and utilizes a tree structure to record the transition relationships between key states for memory-based exploration.

[0076] Optionally, in one embodiment of this application, the expression for the intrinsic reward can be: , in, Indicates that time step t is based on critical states. intrinsic reward. I Indicates an intrinsic reward marker. represent The state at any given moment, It is a label for a certain sub-track. This is the key state corresponding to this sub-trajectory. It represents a distance metric, such as Manhattan distance; A subspace mapping function maps the complete state space to a subspace consisting of a subset of its elements. Among them This represents an element in the state space. This represents a subspace spanned by some elements in the state space. For example, if the state space has 5 dimensions, and only the 2nd and 3rd dimensions are related to the reward, then only these two dimensions are selected to calculate the distance. This represents the subspace mapping function.

[0077] Optionally, in one embodiment of this application, the hindsight intrinsic reward can be represented as: , in, This represents the intrinsic reward after time step t. Indicates that time step t is based on critical states. intrinsic reward. It refers to extrinsic reward. , These represent scaling factors for extrinsic and intrinsic rewards, respectively.

[0078] Optionally, in one embodiment of this application, the recording unit includes: a search subunit and a determination subunit.

[0079] The search sub-unit is used to search for the corresponding branch in the key state memory tree based on the key state chain of the trajectory, so as to sample the next key state that satisfies the preset most likely conditions.

[0080] Determine the sub-units to set the next critical state as the goal of the last sub-trajectory in the trajectory, so that the hindsight intrinsic reward can be applied to the entire trajectory.

[0081] It should be noted that the foregoing explanation of the embodiment of the multi-agent reinforcement learning exploration method based on a large language model also applies to the multi-agent reinforcement learning exploration device based on a large language model in this embodiment, and will not be repeated here.

[0082] The multi-agent reinforcement learning exploration device based on a large language model proposed in this application can design prompt templates and use the large language model to generate key state discriminant functions to identify key states in the trajectory. Combined with a subspace-based hindsight intrinsic reward method, a key state memory tree, and corresponding exploration and planning methods, it completes the multi-agent exploration and obtains accurate exploration results. This achieves the generation of key state discriminant functions in a single round of dialogue using a large language model for subsequent key state identification. It allows the introduction of linguistic knowledge from the large language model into decision-making tasks, significantly reducing the cost associated with frequent calls to the large language model. The subspace-based hindsight intrinsic reward and key state memory tree effectively promote efficient multi-agent exploration by enhancing reward density and enabling more organized exploration. This solves the problems in related technologies where exploration methods either require human expertise to design sub-goals or face difficulties in identifying useful sub-goals, making it difficult to integrate task-related information into sub-goals. Furthermore, the difficulty in incorporating task-related prior information leads to low exploration efficiency. This addresses the challenges of introducing large language models into decision-making problems to meet the needs of various practical applications, reducing costs and time while achieving efficient exploration.

[0083] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.

[0084] When the processor 502 executes the program, it implements the multi-agent reinforcement learning exploration method based on a large language model provided in the above embodiments.

[0085] Furthermore, electronic devices also include: Communication interface 503 is used for communication between memory 501 and processor 502.

[0086] The memory 501 is used to store computer programs that can run on the processor 502.

[0087] Memory 501 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0088] If the memory 501, processor 502, and communication interface 503 are implemented independently, then the communication interface 503, memory 501, and processor 502 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0089] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.

[0090] Processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0091] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described multi-agent reinforcement learning exploration method based on a large language model.

[0092] This application also provides a computer program product that can run computer instructions. When these computer instructions are executed by a processor, they implement the multi-agent reinforcement learning exploration method based on a large language model provided in this application.

[0093] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0094] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0095] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0096] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0097] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0098] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0099] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0100] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A multi-agent reinforcement learning exploration method based on a large language model, characterized in that, Includes the following steps: Based on at least one preset prompt template, a key state discrimination function is generated using a large language model; Based on the key state discrimination function, the task-related key states with explicit semantics and expression are found in the sampled trajectory; The key states are used as prior information to obtain the multi-agent reinforcement learning exploration results; The step of obtaining the multi-agent reinforcement learning exploration results by using the key states as prior information includes: introducing intrinsic rewards at each time step of the sub-trajectories of the trajectory to guide the training of the multi-agent based on the intrinsic rewards, and combining extrinsic rewards to obtain subspace-based hindsight intrinsic rewards; and using a tree structure to record the transition relationships between the key states based on the subspace-based hindsight intrinsic rewards for memory-based exploration. The expression for the intrinsic reward is: , in, Indicates that time step t is based on critical states. intrinsic reward. I Indicates an intrinsic reward marker. represent The state at any given moment, It is a label for a certain sub-track. This is the key state corresponding to this sub-trajectory. Represents a distance metric; A subspace mapping function maps the complete state space to a subspace consisting of a subset of its elements. Among them This represents an element in the state space. This represents a subspace spanned by a subset of elements in the state space. Represents the subspace mapping function; The hindsight intrinsic reward is: , in, This represents the intrinsic reward after time step t. Indicates that time step t is based on critical states. intrinsic reward. It refers to extrinsic reward. , These represent scaling factors for extrinsic and intrinsic rewards, respectively.

2. The method according to claim 1, characterized in that, The process of finding task-relevant key states with explicit semantics and expression in the sampled trajectory based on the key state discrimination function includes: The system receives the state at any time step as an input variable and outputs a Boolean value indicating whether the current input state belongs to its corresponding key state, so as to identify and label each state in the trajectory and obtain the identification and labeling results. The critical state is determined based on the identification and labeling results.

3. The method according to claim 1, characterized in that, The subspace-based hindsight intrinsic reward, utilizing a tree structure to record the transition relationships between key states for memory-based exploration, includes: Based on the key state chain of the trajectory, the corresponding branch is searched in the key state memory tree to sample the next key state that satisfies the preset most likely conditions. The next key state is set as the target of the last sub-trajectory in the trajectory, so that the hindsight intrinsic reward is applied to the entire trajectory.

4. A multi-agent reinforcement learning exploration device based on a large language model, characterized in that, The method employs a multi-agent reinforcement learning exploration method based on a large language model as described in any one of claims 1-3, comprising: The generation module is used to generate key state discrimination functions based on at least one preset prompt template using a large language model; The search module is used to search for task-related key states with explicit semantics and expressions in the sampled trajectory based on the key state discrimination function; The exploration module is used to obtain the exploration results of multi-agent reinforcement learning by taking the key states as prior information.

5. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the multi-agent reinforcement learning exploration method based on a large language model as described in any one of claims 1-3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the multi-agent reinforcement learning exploration method based on a large language model as described in any one of claims 1-3.

7. A computer program product, characterized in that, When the computer program is executed, it is used to implement the multi-agent reinforcement learning exploration method based on a large language model as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Multi-physical field constrained intelligent quick charging method for lithium ion battery

    CN112018465A

  • Intelligent decision-making method and device based on multi-modal data fusion and reinforcement learning

    CN114860893A