Multi-task-oriented agent training method and decision-making method and device
By using hybrid encoder, shared policy network, action correction policy network and action correction module in the training of agents, combined with sparse and dense rewards, the problems of strategy generalization and training inefficiency in multi-task scenarios are solved, and the agent's effective exploration and cross-task generalization in multi-task scenarios are realized.
Patent Information
- Application Number
- CN202510477428.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing reinforcement learning methods have shortcomings in strategy generalization and training efficiency in multi-task scenarios, making it difficult to form a general agent that supports multi-task scenarios, especially in sparse reward environments, and multi-task sharing strategies often lead to strategy degradation due to task goal conflicts.
A multi-tasking training method is proposed, using a hybrid encoder, a shared policy network, an action correction policy network and an action correction module. By obtaining the initial task state and target task state of each training sample in the training sample set, an estimated task feature, preliminary actions, correction actions and next actions are generated, and network parameters are updated based on sparse and dense rewards.
This method can effectively enhance the agent's exploration ability in multi-task scenarios, allowing the agent to adopt a combination of short-term perspectives and long-term perspectives to generate the next action, realize the migration and generalization of cross-tasks, and avoid strategic degradation.
Smart Images

Figure CN119988988A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of artificial intelligence technology, and more specifically, to a multi-task oriented intelligent agent training method and decision-making method and device. Background Art
[0002] In recent years, reinforcement learning technology has made significant progress in the field of intelligent decision-making of agents, especially in single-task scenarios. However, most reinforcement learning research focuses on specific problem scenarios, giving priority to mastering a single task by learning a single strategy, which often comes at the expense of generalization ability. When faced with multi-task scenarios, existing methods still have obvious defects in policy generalization and training efficiency, making it difficult to form a general agent that supports multi-task scenarios. Traditional methods usually use independent policy networks to handle different tasks, resulting in a linear increase in parameter size with the number of tasks, and it is difficult to achieve cross-task knowledge transfer. Although some studies have attempted to improve generalization ability by sharing network structures, such methods often ignore the differences in reward mechanisms between different tasks, especially in dealing with the collaborative optimization problem of dense rewards and sparse rewards. There is a lack of effective mechanisms.
[0003] At present, the main technical bottlenecks are as follows: First, in a sparse reward environment, the agent's training efficiency is low due to the lack of effective exploration signals, and conventional curriculum learning or reward reshaping methods are prone to suboptimal strategy convergence problems. Secondly, multi-task sharing strategies often lead to strategy degradation due to conflicts in task goals. Especially in long-term decision-making scenarios, it is difficult to effectively coordinate the short-sighted immediate reward maximization strategy with the far-sighted global goal achievement requirements. The task-specific knowledge obtained from one task may hinder the overall learning process of other tasks, resulting in excessive attention to the immediate rewards of a single task during multi-task learning, which leads to myopia and hinders the ability to learn general strategies to effectively complete multiple tasks. In addition, existing dynamic weight adjustment methods mostly use fixed ratios or empirical parameter settings, which cannot adapt to changes in task difficulty and training stages, causing the gradients between policy modules to move in opposite directions, and gradient conflicts cannot be effectively alleviated. Summary of the invention
[0004] The embodiments of the present disclosure provide a multi-task oriented intelligent agent training method and decision-making method and device, which can effectively solve at least one of the above problems.
[0005] In a general aspect, a multi-task oriented intelligent agent training method is provided, wherein the intelligent agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the training method includes: obtaining a training sample set of a round, wherein each training sample in the training sample set includes an initial task state and a target task state of a task; for each training sample in the training sample set, performing the following processing: inputting the initial task state in the training sample into the hybrid encoder to obtain an estimated task feature; inputting the estimated task feature into the shared strategy network to generate an estimated preliminary action for the initial task state; inputting the estimated task feature and the estimated preliminary action into the action correction strategy network to generate an estimated task state for the initial task state. The invention relates to a method for performing a training sample task in a training process, comprising: performing a training sample task in a training process, and performing a training process of the training sample. The method comprises: performing a training sample task in a training process, and performing a training process of the training sample. The method comprises: estimating a correction action; inputting the estimated preliminary action and the estimated correction action into the action correction module to obtain an estimated next action for the initial task state; executing the estimated next action to obtain the estimated next task state of the task in the training sample; determining a sparse reward based on the estimated next task state and the target task state; determining a dense reward based on the initial task state, the estimated next task state and the target task state; taking the estimated next task state as the initial task state and returning to the step of obtaining the estimated task feature until the task in the training sample is completed; in response to all training samples having completed the above processing, updating the parameters of the shared strategy network, the action correction strategy network and the hybrid encoder based on all sparse rewards and all dense rewards.
[0006] Optionally, based on the estimated next task state and the target task state, determining a sparse reward includes: mapping the estimated next task state and the target task state to a first latent space; determining a first distance between the estimated next task state and the target task state in the first latent space; and determining the sparse reward based on a relationship between the first distance and a corresponding distance threshold.
[0007] Optionally, based on the initial task state, the estimated next task state and the target task state, a dense reward is determined, including: mapping the estimated next task state and the target task state to a second latent space; mapping the initial task state and the predetermined task state to a third latent space, wherein the predetermined task state is the target task state or the estimated next task state; in the second latent space, determining a second distance between first predetermined information in the estimated next task state and the target task state, wherein the first predetermined information is determined based on the content of the task in the training sample; in the third latent space, determining a third distance between the initial task state and the predetermined task state, wherein the second predetermined information is determined based on the content of the task in the training sample; and determining the dense reward based on the relationship between the second distance, the third distance and the corresponding distance threshold.
[0008] Optionally, the action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule, wherein the estimated preliminary action and the estimated correction action are input into the action correction module to obtain an estimated next action for the initial task state, including: inputting the estimated preliminary action and the estimated correction action into the weighted summation submodule to obtain a weighted summation result; inputting the weighted summation result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain a maximum value within the range of the weighted summation result and the first action boundary; inputting the maximization result into the minimization submodule to obtain an estimated next action, wherein the minimization submodule is used to obtain a minimum value within the range of the maximization result and the second action boundary; wherein the estimated preliminary action, the estimated correction action and the estimated next action are all within the range of the first action boundary and the second action boundary. In another general aspect, a multi-task decision-making method is provided, which is applied to an intelligent agent, wherein the intelligent agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the decision-making method includes: inputting the initial task state of the target task into the hybrid encoder to obtain task characteristics; inputting the task characteristics into the shared strategy network to generate a preliminary action for the initial task state; inputting the task characteristics and the preliminary action into the action correction strategy network to generate a correction action for the initial task state; inputting the preliminary action and the correction action into the action correction module to obtain the next action for the initial task state; executing the next action to obtain the next task state of the target task; using the next task state as the initial task state and returning to the step of obtaining the task characteristics until the target task is completed; wherein the intelligent agent is trained by any of the training methods above.
[0009] Optionally, the action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule, wherein the preliminary action and the correction action are input into the action correction module to obtain the next action for the initial task state, including: inputting the preliminary action and the correction action into the weighted summation submodule to obtain a weighted summation result; inputting the weighted summation result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain the maximum value within the range of the weighted summation result and the first action boundary; inputting the maximization result into the minimization submodule to obtain the next action, wherein the minimization submodule is used to obtain the minimum value within the range of the maximization result and the second action boundary; wherein the preliminary action, the correction action and the next action are all within the range of the first action boundary and the second action boundary.
[0010] In another general aspect, a multi-task oriented intelligent agent training device is provided, the intelligent agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, the training device includes: a first acquisition unit, configured to acquire a round of training sample sets, wherein each training sample in the training sample set includes an initial task state and a target task state of a task; a first execution unit, configured to perform the following processing on each training sample in the training sample set: input the initial task state in the training sample into the hybrid encoder to obtain estimated task features; input the estimated task features into the shared strategy network to generate an estimated preliminary action for the initial task state; input the estimated task features and the estimated preliminary action into the action correction strategy network to generate a target task state for the initial task The invention relates to a method for performing an estimated correction action on a task state; inputting the estimated preliminary action and the estimated correction action into an action correction module to obtain an estimated next action for the initial task state; executing the estimated next action to obtain an estimated next task state for the task in the training sample; determining a sparse reward based on the estimated next task state and the target task state; determining a dense reward based on the initial task state, the estimated next task state and the target task state; taking the estimated next task state as the initial task state and returning to the step of obtaining the estimated task feature until the task in the training sample is completed; an updating unit, configured to update the parameters of a shared strategy network, an action correction strategy network and a hybrid encoder based on all sparse rewards and all dense rewards in response to all training samples having completed the above processing.
[0011] Optionally, the first execution unit is further configured to map the estimated next task state and the target task state to a first latent space; determine a first distance between the estimated next task state and the target task state in the first latent space; and determine a sparse reward based on a relationship between the first distance and a corresponding distance threshold.
[0012] Optionally, the first execution unit is further configured to map the estimated next task state and the target task state to a second latent space; map the initial task state and the predetermined task state to a third latent space, wherein the predetermined task state is the target task state or the estimated next task state; determine, in the second latent space, a second distance between the estimated next task state and the first predetermined information in the target task state, wherein the first predetermined information is determined based on the content of the task in the training sample; determine, in the third latent space, a third distance between the initial task state and the second predetermined information in the predetermined task state, wherein the second predetermined information is determined based on the content of the task in the training sample; and determine a dense reward based on the relationship between the second distance, the third distance and the corresponding distance threshold.
[0013] Optionally, the action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule, wherein the first execution unit is further configured to input the estimated preliminary action and the estimated correction action into the weighted summation submodule to obtain a weighted summation result; input the weighted summation result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain the maximum value within the range of the weighted summation result and the first action boundary; input the maximization result into the minimization submodule to obtain an estimated next action, wherein the minimization submodule is used to obtain the minimum value within the range of the maximization result and the second action boundary; wherein the estimated preliminary action, the estimated correction action and the estimated next action are all within the range of the first action boundary and the second action boundary.
[0014] In another general aspect, a multi-task decision-making device is provided, which is applied to an intelligent agent, wherein the intelligent agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the decision-making device includes: a second acquisition unit, configured to input the initial task state of the target task into the hybrid encoder to obtain task characteristics; a first generation unit, configured to input the task characteristics into the shared strategy network to generate a preliminary action for the initial task state; a second generation unit, configured to input the task characteristics into the action correction strategy network to generate a correction action for the initial task state; a fusion unit, configured to input the preliminary action and the correction action into the action correction module to obtain a next action for the initial task state; a second execution unit, configured to execute the next action, obtain the next task state of the target task, use the next task state as the initial task state and return to the step of obtaining the task characteristics until the target task is completed; wherein the intelligent agent is trained by any of the training methods above.
[0015] Optionally, the action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule, wherein the fusion unit is further configured to input the preliminary action and the correction action into the weighted summation submodule to obtain a weighted summation result; input the weighted summation result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain the maximum value within the range of the weighted summation result and the first action boundary; input the maximization result into the minimization submodule to obtain the next action, wherein the minimization submodule is used to obtain the minimum value within the range of the maximization result and the second action boundary; wherein the preliminary action, the correction action and the next action are all within the range of the first action boundary and the second action boundary.
[0016] In another general aspect, a computer-readable storage medium storing instructions is provided, wherein, when the instructions are executed by at least one computing device, the at least one computing device is prompted to execute any of the multi-task oriented intelligent agent training methods and multi-task oriented decision-making methods described above.
[0017] In another general aspect, a system is provided comprising at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, cause the at least one computing device to execute any of the multi-task oriented intelligent agent training methods and multi-task oriented decision-making methods described above.
[0018] In another general aspect, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement any of the above multi-task oriented intelligent agent training methods and multi-task oriented decision-making methods.
[0019] According to the embodiment of the present disclosure, the intelligent agent training method, decision-making method and device for multi-task are used. The dense reward takes into account the initial task state and the estimated next task state, so that the dense reward is incorporated into the short-term perspective of the action of each task. At the same time, the sparse reward takes into account the target task state and the estimated next task state, pays attention to the overall completion of the task, enhances the exploration ability of the intelligent agent in multi-task scenarios, and enables the intelligent agent to take a long-term perspective to correct the action. Therefore, the intelligent agent trained in the present disclosure can take a combination of short-term and long-term perspectives to generate the next action and achieve cross-task migration and generalization.
[0020] Additional aspects and / or advantages of the present general inventive concept will be set forth in part in the following description and in part will be apparent from the description or may be learned through practice of the present general inventive concept. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other objects and features of the embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings showing the embodiments, in which: Figure 1 is a flowchart illustrating a multi-task-oriented intelligent agent training method according to an embodiment of the present disclosure; Figure 2 is a flow chart showing a multi-task oriented decision making method according to an embodiment of the present disclosure; Figure 3 is a system flow chart showing a multi-task oriented decision-making method according to an embodiment of the present disclosure; Figure 4 is a system architecture diagram showing a multi-task oriented decision method according to an embodiment of the present disclosure; Figure 5 is a schematic diagram showing an application scenario of a multi-task oriented decision-making method according to an embodiment of the present disclosure; Figure 6 is a schematic diagram showing a battle result of applying a multi-task oriented decision method according to an embodiment of the present disclosure; Figure 7is a block diagram showing a multi-task-oriented intelligent agent training device according to an embodiment of the present disclosure; Figure 8 is a block diagram showing a multi-task oriented decision making apparatus according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] The following specific embodiments are provided to help the reader obtain a comprehensive understanding of the methods, devices and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be clear. For example, the order of operations described herein is only an example and is not limited to those orders set forth herein, but can be changed as will be clear after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, for greater clarity and simplicity, the description of features known in the art may be omitted.
[0023] The features described herein can be implemented in different forms and should not be construed as being limited to the examples described herein. Rather, the examples described herein have been provided to illustrate only some of the many possible ways to implement the methods, devices, and / or systems described herein, which will be clear after understanding the disclosure of the present application.
[0024] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more.
[0025] Although terms such as "first", "second", and "third" may be used herein to describe various members, components, regions, layers, or portions, these members, components, regions, layers, or portions should not be limited by these terms. Instead, these terms are only used to distinguish one member, component, region, layer, or portion from another member, component, region, layer, or portion. Therefore, without departing from the teachings of the examples described herein, the first member, first component, first region, first layer, or first portion referred to in the examples may also be referred to as the second member, second component, second region, second layer, or second portion.
[0026] In the specification, when an element (such as a layer, a region, or a substrate) is described as being “on”, “connected to”, or “coupled to” another element, the element may be directly “on”, “connected to”, or “coupled to” another element, or one or more other elements may be present therebetween. Conversely, when an element is described as being “directly on”, “directly connected to”, or “directly coupled to” another element, there may be no other elements present therebetween.
[0027] The terms used herein are only used to describe various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. The terms "comprise", "include" and "have" indicate the presence of the described features, quantities, operations, components, elements and / or combinations thereof, but do not exclude the presence or addition of one or more other features, quantities, operations, components, elements and / or combinations thereof.
[0028] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which the present disclosure belongs after understanding the present disclosure. Unless explicitly defined as such herein, terms (such as those defined in general dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal manner.
[0029] Furthermore, in the description of examples, when it is considered that a detailed description of a well-known related structure or function would cause vague interpretation of the present disclosure, such a detailed description will be omitted.
[0030] The intelligent agent disclosed in the present invention can be applied to unmanned vehicles, drones, humanoid robots, games (such as as a human-machine player), etc., and the present invention does not limit this. Assuming that the intelligent agent is applied to a humanoid robot, it can perform but is not limited to the following tasks: controlling the robot arm to reach the target location to pick up the target object or deliver the target object, controlling the robot arm to open the drawer, and the present invention does not limit this; assuming that the intelligent agent is applied to a drone, it can perform but is not limited to the following tasks: delivering materials to the target location and arranging a fixed formation, and the present invention does not limit this; assuming that the intelligent agent is applied to an unmanned vehicle, it can perform but is not limited to the following tasks: carrying materials to the target location, and the present invention does not limit this. Assuming that the intelligent agent is used as a human-machine player in the game, it can perform but is not limited to the following tasks: confronting the enemy, conquering several enemy cities, etc., and the present invention does not limit this.
[0031] The multi-task intelligent agent training method and decision-making method and device disclosed in the present invention are described in detail below with reference to the accompanying drawings.
[0032] This paper proposes a multi-task intelligent agent training method. Figure 1 is a flow chart showing a multi-task-oriented agent training method according to an embodiment of the present disclosure. Figure 1 The agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the multi-task agent training method includes the following steps: In step S101, a training sample set of a round is obtained, wherein each training sample in the training sample set includes an initial task state and a target task state of a task.
[0033] As an example, a training sample set of one round can be randomly selected from the sample pool, and the tasks contained in the two training samples in the training sample set can be different or the same, which is not limited in this disclosure. However, the tasks contained in all the training samples in the training sample set are generally multiple tasks, which can improve the richness of the training samples.
[0034] As an example, after obtaining the training samples, the agent can be initialized before starting training, such as initializing the number of tasks of the agent, that is, the number of tasks that the agent can perform, including but not limited to the number of tasks that the agent can perform simultaneously and the number of types of tasks that the agent can accept; for example, initializing the virtual expected budget, that is, the expected value to be achieved in each step of each task; for example, initializing the parameters of the hybrid encoder, shared strategy network, motion correction strategy network and motion correction module; it can also be any other information that needs to be initialized, which is not limited by the present disclosure.
[0035] In step S102, for each training sample in the training sample set, the following processing is performed: the initial task state in the training sample is input into the hybrid encoder to obtain the estimated task features; the estimated task features are input into the shared strategy network to generate an estimated preliminary action for the initial task state; the estimated task features and the estimated preliminary action are input into the action correction strategy network to generate an estimated corrected action for the initial task state; the estimated preliminary action and the estimated corrected action are input into the action correction module to obtain the estimated next action for the initial task state; the estimated next action is executed to obtain the estimated next task state of the task in the training sample; based on the estimated next task state and the target task state, a sparse reward is determined; based on the initial task state, the estimated next task state and the target task state, a dense reward is determined; the estimated next task state is used as the initial task state and returns to the step of obtaining the estimated task features until the task in the training sample is completed.
[0036] As an example, after initializing the agent, a training sample can be randomly selected from the training sample set B, the initial task state in the training sample can be input into the hybrid encoder to obtain the estimated task features, and then the estimated task features can be input into the shared policy network to generate the estimated preliminary action for the initial task state. , the estimated task features and the estimated preliminary actions are input into the action correction strategy network to generate the estimated correction action for the initial task state , and then input the estimated preliminary action and the estimated correction action into the action correction module to obtain the estimated next action for the initial task state After obtaining the estimated next action, execute the estimated next action to obtain the estimated next task state of the task in the training sample , and then, based on the estimated next task status and the target task state, determine the sparse reward ; Based on the initial task state , Estimate the next task status and target task status, determine the intensive reward ; Finally, it will include the initial task state (each subsequent step will be replaced by the corresponding next task state), the estimated next action , Estimate the next task status , sparse rewards and intensive rewards The new samples are added to the training sample set, that is , taking the estimated next task state as the initial task state and returning to the step of obtaining the estimated task features until the tasks in the training sample are completed.
[0037] It should be noted that , denote the shared policy network and the action correction policy network respectively, , denote the learnable network parameters in the shared strategy network and the action correction strategy network, respectively. Represents the function corresponding to the action correction module, Represents the tasks in the training sample i The sparse reward of Represents the tasks in the training sample i Intensive rewards.
[0038] As an example, the hybrid encoder may include 4 encoders, 1 multilayer perceptron and 1 attention network, which is not limited in the present disclosure. The shared strategy network and the action correction strategy network may have the same structure, such as being composed of 3 layers of multilayer perceptrons with 400 neurons, the activation function between layers is ReLu, and the parameters are initialized to a standard Gaussian distribution, which is not limited in the present disclosure.
[0039] According to an embodiment of the present disclosure, determining a sparse reward based on an estimated next task state and a target task state may include: mapping the estimated next task state and the target task state to a first latent space; determining a first distance between the estimated next task state and the target task state in the first latent space; and determining a sparse reward based on a relationship between the first distance and a corresponding distance threshold. Through this embodiment, the distance between the next task state and the target task state is taken into account, so that the sparse reward can better focus on the overall completion of the task, thereby better correcting the action from a long-term perspective. As an example, the above sparse reward , that is, the task in the training sample i The sparse reward brought by the next task state can be specifically expressed as:
[0040] in, represents the task state of the agent at the current moment (that is, the initial task state mentioned above, which will be replaced by the next task state in each subsequent step), Indicates the target task status, It represents a function that maps the target task state and the current task state of the agent to the latent space, and is used to calculate the distance between the target task state and the current task state of the agent in the latent space. At this time, it can be defined as , Indicates the reward value for reaching the target task state, Indicates the distance threshold, which can be set as needed.
[0041] As an example, the above Generally, a constant of 1 can be used to formalize sparse rewards. Only when the task is completed will a reward of 1 be obtained, that is, the reward value is increased by 1.
[0042] According to an embodiment of the present disclosure, based on the initial task state, the estimated next task state and the target task state, determining the intensive reward may include: mapping the estimated next task state and the target task state to a second latent space; mapping the initial task state and the scheduled task state to a third latent space, wherein the scheduled task state is the target task state or the estimated next task state; in the second latent space, determining the second distance between the first predetermined information in the estimated next task state and the target task state, wherein the first predetermined information is determined based on the content of the task in the training sample; in the third latent space, determining the third distance between the initial task state and the second predetermined information in the scheduled task state, wherein the second predetermined information is determined based on the content of the task in the training sample; determining the intensive reward based on the relationship between the second distance, the third distance and the corresponding distance threshold. Through this embodiment, the distance between the next task state and the initial task state is taken into account, so that the intensive reward can better focus on the immediate completion of the task, so that more accurate actions can be obtained from a short-term perspective. As an example, for different tasks, the relationship between the second distance, the third distance and the corresponding distance threshold is different, as illustrated below: Assumption Task To control the manipulator to reach the target location, the dense reward can be expressed as:
[0043]
[0044]
[0045] in, Indicates the location of the target location. Indicates the task status The position of the manipulator in represents the initial position of the robot, Represents the calculation of the bi-norm.
[0046] Assumption Task To control the robot to open the drawer, the dense reward can be expressed as:
[0047] in, Indicates the initial state of the drawer. Indicates the task status The state of the drawer in Indicates the target state of the drawer, Indicates the task status The position of the manipulator in Indicates the initial position of the robot.
[0048] According to an embodiment of the present disclosure, the action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule, wherein the estimated preliminary action and the estimated correction action are input into the action correction module to obtain the estimated next action for the initial task state, which may include: inputting the estimated preliminary action and the estimated correction action into the weighted summation submodule to obtain a weighted summation result; inputting the weighted summation result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain the maximum value within the weighted summation result and the first action boundary range; inputting the maximization result into the minimization submodule to obtain the estimated next action, wherein the minimization submodule is used to obtain the minimum value within the maximization result and the second action boundary range; wherein the estimated preliminary action, the estimated correction action and the estimated next action are all within the range of the first action boundary and the second action boundary. Through this embodiment, the estimated preliminary action and the estimated correction action are fused within the range of the first action boundary and the second action boundary, so that a relatively accurate next action can be obtained.
[0049] As an example, a fixed value A may be selected as the action boundary, the first action boundary may be set to -A, and the second action boundary may be set to A, which is not limited in the present disclosure.
[0050] As an example, assuming that A is selected as the action boundary, the next action can be determined as follows :
[0051] in, , Indicates finding the minimum and maximum values. represents the action boundary, .
[0052] In step S103, in response to all training samples having completed the above processing, the parameters of the shared policy network, the action correction policy network and the hybrid encoder are updated based on all sparse rewards and all dense rewards. As an example, after obtaining all the estimated next actions, all the estimated next task states, all the sparse rewards, and all the dense rewards of all the training samples, the loss functions of the shared strategy network and the action correction strategy network can be determined as follows:
[0053]
[0054] in, represents the loss function of the shared policy network; represents the loss function of the action correction policy network; represents the mathematical expectation of all possible states and actions, represents the mathematical expectation of the task state s under the probability distribution D, Indicates estimated initial action In a shared policy network The mathematical expectation under the strategy probability, Indicates the task The distribution probability of all tasks The mathematical expectation under Indicates the estimated status of the next task On Task The state transition probability The mathematical expectation under Indicates estimated corrective action Correction strategies in action The mathematical expectation under the strategy probability; D represents the playback buffer; Indicates tasks; Represents the distribution probability of all tasks; represents the discount factor; Indicates the task The state transition probability of represents the state estimation value of the shared strategy network, and in this embodiment, the Euclidean distance may be used; Represents the Lagrange multiplier, which is used to adjust the size of the loss function; represents the estimated value of the state-action of the action correction module; is the distance function; Represents the learning rate.
[0055] Then, the gradient of the Lagrange multiplier is estimated using the sparse reward and the virtual expected budget, as follows: ; in, Indicates The value of the sparse reward after the step action, represents the number of action steps required to reach the target task state from the initial task state, Represents the virtual expected budget for each step. It should be noted that for a task, the virtual expected budget can usually be set to a fixed value, which is mainly used to balance the training intensity of the shared strategy network and the action correction strategy network. Different virtual expected budgets can be set for different tasks.
[0056] Then, the shared policy network and the action correction policy network can be updated as follows: , , in, , Respectively , The gradient of , in this embodiment, is calculated using the gradient descent method, Represents the learning rate.
[0057] It should be noted that after the training using the current training sample set is completed, another training sample can be selected to continue training until the termination state is reached or the maximum number of iterations is completed. In addition, the hybrid encoder can be updated along with the update of the shared strategy network and the action correction strategy network, and the specific update method is not limited in this disclosure.
[0058] This disclosure also proposes a multi-task decision-making method. Figure 2 is a flowchart showing a multi-task decision-making method according to an embodiment of the present disclosure. Figure 2 The method is applied to an agent trained by any of the above training methods, the agent comprising a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the multi-task oriented decision method comprises the following steps: In step S201, the initial task state of the target task is input into the hybrid encoder to obtain the task features; In step S202, the task features are input into the shared policy network to generate preliminary actions for the initial task state; In step S203, the task features and the preliminary action are input into the action correction strategy network to generate a correction action for the initial task state; In step S204, the preliminary action and the correction action are input into the action correction module to obtain the next action for the initial task state; Optionally, the action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule, wherein the preliminary action and the correction action are input into the action correction module to obtain the next action for the initial task state, including: inputting the preliminary action and the correction action into the weighted summation submodule to obtain a weighted summation result; inputting the weighted summation result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain the maximum value within the range of the weighted summation result and the first action boundary; inputting the maximization result into the minimization submodule to obtain the next action, wherein the minimization submodule is used to obtain the minimum value within the range of the maximization result and the second action boundary; wherein the preliminary action, the correction action and the next action are all within the range of the first action boundary and the second action boundary. Through this embodiment, the preliminary action and the correction action are fused within the range of the first action boundary and the second action boundary, so that a relatively accurate next action can be obtained.
[0059] In step S205, the next step is performed to obtain the next task status of the target task; In step S206, the process returns to the task feature acquisition step and replaces the initial task state with the next task state until the target task is completed.
[0060] In order to facilitate the understanding of the above embodiments, Figure 3 and Figure 4 Provide a description of the system.
[0061] Figure 3 A system flow chart showing the multi-task oriented decision-making approach, Figure 4 The system architecture diagram of the multi-task decision-making method is shown, such as Figure 3 and Figure 4 As shown, the system flow includes the following steps: S301, initializing agent parameters, such as initializing the number of tasks n, virtual expected budget, etc.; S302, inputting the current task state of the target task into the hybrid encoder of the agent to extract all task features in the current task state; S303, generating preliminary actions of the agent based on the shared strategy network, that is, inputting the extracted task features into the shared strategy network to obtain preliminary actions; S304, generating a correction action of the agent based on the action correction strategy network, that is, inputting the extracted task features and the preliminary action into the action correction strategy network to obtain the correction action; S305, based on the action correction function (i.e. the above-mentioned action correction module), the initial action and the correction action are integrated to output the next action of the intelligent agent, that is, the initial action and the correction action are input into the action correction function to obtain the next action. After obtaining the next task, the next action can be executed to obtain the next task state of the target task, use the next task state as the new current task state, and return to step S302 until the target task is completed.
[0062] It should be noted that the executor of the next action can be a third party. For example, when the intelligent agent is a human-computer player in the game and can perform battle tasks, the intelligent agent can be loaded on a computer device. At this time, the executor is the computer device, which will inform the intelligent agent of the game progress. The intelligent agent will give the most accurate next action based on the game progress and return it to the game application on the computer device to control the corresponding character of the game application to perform the next action. Figure 5 As shown, Figure 5 (A) shows the decision of 3 friendly fighters against 5 enemy fighters. Figure 5(B) shows the decision of 2 friendly fighters against 64 enemy fighters. Figure 5 (C) in the figure shows a battle between 6 friendly soldiers and 24 enemy soldiers. After receiving the task status, the agent will determine the next action of the friendly soldiers according to the task status, thereby controlling the friendly soldiers to take the next action until the battle task is completed. The battle results are shown in Figure 2. Figure 6 As shown, the intelligent agent trained by the training method of the present invention has a relatively high average winning rate; for another example, when the intelligent agent is applied to a humanoid robot, the intelligent agent can give the next action according to the state of the humanoid robot, and the humanoid robot receives and executes the next action.
[0063] In summary, the present disclosure proposes an intelligent decision-making method for an intelligent agent that supports multi-task migration, aiming to give full play to the advantages of short-term perspective and long-term perspective, that is, to combine the shared strategy network and the action correction network, and to incorporate the dense reward into each task to generate the action of the short-term perspective, so as to ignore the overall completion of the task. At the same time, the sparse reward adopts a longer decision range, focusing on the overall completion of the task, so as to enhance the exploration ability of the intelligent agent in the multi-task scenario, so that the intelligent agent can take a long-term perspective to generate the correction action, thereby realizing the generalization of the intelligent agent across tasks, and then the intelligent agent can quickly decide the best action to support the multi-task scenario. Specifically, the intelligent agent parameters are initialized first, and the task state is input into the mixed encoder in the intelligent agent to extract all task features, and then based on the shared strategy network, the initial action of the intelligent agent is generated, and based on the action correction strategy network, the correction action of the intelligent agent is generated, and then based on the action correction function, the initial action and the correction action are fused to output the next action. It can be seen that the present disclosure combines the information of the short-term perspective and the long-term perspective, which can support the decision-making of the best action in the multi-task scenario, so that the intelligent agent has higher cross-task autonomy and strategy migration.
[0064] Figure 7 is a block diagram showing a multi-task agent training apparatus according to an embodiment of the present disclosure, Figure 7 As shown, the agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the device includes a first acquisition unit 70, a first execution unit 72 and an update unit 74.
[0065] The first acquisition unit 70 is configured to acquire a round of training sample sets, wherein each training sample in the training sample set contains an initial task state and a target task state of a task; the first execution unit 72 is configured to perform the following processing on each training sample in the training sample set: input the initial task state in the training sample into the hybrid encoder to obtain the estimated task features; input the estimated task features into the shared strategy network to generate an estimated preliminary action for the initial task state; input the estimated task features and the estimated preliminary action into the action correction strategy network to generate an estimated corrected action for the initial task state; input the estimated preliminary action and the estimated corrected action into the action correction model block, obtains the estimated next action for the initial task state; executes the estimated next action to obtain the estimated next task state of the task in the training sample; determines the sparse reward based on the estimated next task state and the target task state; determines the dense reward based on the initial task state, the estimated next task state and the target task state; uses the estimated next task state as the initial task state and returns to the step of obtaining the estimated task feature until the task in the training sample is completed; the updating unit 74 is configured to update the parameters of the shared strategy network, the action correction strategy network and the hybrid encoder based on all the sparse rewards and all the dense rewards in response to all the training samples having completed the above processing.
[0066] According to an embodiment of the present disclosure, the first execution unit 72 is also configured to map the estimated next task state and the target task state to a first latent space; determine a first distance between the estimated next task state and the target task state in the first latent space; and determine a sparse reward based on a relationship between the first distance and a corresponding distance threshold.
[0067] According to an embodiment of the present disclosure, the first execution unit 72 is also configured to map the estimated next task state and the target task state to a second latent space; map the initial task state and the predetermined task state to a third latent space, wherein the predetermined task state is the target task state or the estimated next task state; in the second latent space, determine a second distance between the first predetermined information in the estimated next task state and the target task state, wherein the first predetermined information is determined based on the content of the task in the training sample; in the third latent space, determine a third distance between the initial task state and the second predetermined information in the predetermined task state, wherein the second predetermined information is determined based on the content of the task in the training sample; and determine a dense reward based on the relationship between the second distance, the third distance and the corresponding distance threshold.
[0068] According to an embodiment of the present disclosure, the action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule, wherein the first execution unit 72 is further configured to input the estimated preliminary action and the estimated correction action into the weighted summation submodule to obtain a weighted summation result; input the weighted summation result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain the maximum value within the range of the weighted summation result and the first action boundary; input the maximization result into the minimization submodule to obtain an estimated next action, wherein the minimization submodule is used to obtain the minimum value within the range of the maximization result and the second action boundary; wherein the estimated preliminary action, the estimated correction action and the estimated next action are all within the range of the first action boundary and the second action boundary.
[0069] Figure 8 is a block diagram showing a multi-task decision-making device according to an embodiment of the present disclosure, such as Figure 8 As shown, the agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the device includes a second acquisition unit 80, a first generation unit 82, a second generation unit 84, a fusion unit 86 and a second execution unit 88.
[0070] The second acquisition unit 80 is configured to input the initial task state of the target task into the hybrid encoder to obtain task features; the first generation unit 82 is configured to input the task features into the shared strategy network to generate a preliminary action for the initial task state; the second generation unit 84 is configured to input the task features into the action correction strategy network to generate a correction action for the initial task state; the fusion unit 86 is configured to input the preliminary action and the correction action into the action correction module to obtain the next action for the initial task state; the second execution unit 88 is configured to execute the next action, obtain the next task state of the target task, use the next task state as the initial task state and return to the step of obtaining task features until the target task is completed; wherein the intelligent agent is trained by any of the training methods above.
[0071] According to an embodiment of the present disclosure, the action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule, wherein the fusion unit 86 is further configured to input the preliminary action and the correction action into the weighted summation submodule to obtain a weighted summation result; input the weighted summation result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain the maximum value within the range of the weighted summation result and the first action boundary; input the maximization result into the minimization submodule to obtain the next action, wherein the minimization submodule is used to obtain the minimum value within the range of the maximization result and the second action boundary; wherein the preliminary action, the correction action and the next action are all within the range of the first action boundary and the second action boundary.
[0072] According to an embodiment of the present disclosure, a computer-readable storage medium storing instructions is provided, wherein, when the instructions are executed by at least one computing device, the at least one computing device is prompted to execute a multi-task-oriented intelligent agent training method and a multi-task-oriented decision-making method as described in any of the above embodiments.
[0073] According to an embodiment of the present disclosure, a system is provided comprising at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, prompt the at least one computing device to execute a multi-task-oriented intelligent agent training method and a multi-task-oriented decision-making method as described in any of the above-mentioned embodiments.
[0074] According to an embodiment of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement any of the above multi-task oriented intelligent agent training methods and multi-task oriented decision-making methods.
[0075] Although some embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that modifications may be made to the embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the claims and their equivalents.
Claims
1. A multi-task agent training method, characterized in that: The agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the training method includes: Acquire a round of training sample sets, wherein each training sample in the training sample set includes an initial task state and a target task state of a task; For each training sample in the training sample set, the following processing is performed: Inputting the initial task state in the training sample into the hybrid encoder to obtain estimated task features; Inputting the estimated task features into the shared policy network to generate an estimated preliminary action for the initial task state; Inputting the estimated task features and the estimated preliminary action into the action correction strategy network to generate an estimated correction action for the initial task state; Inputting the estimated preliminary action and the estimated correction action into the action correction module to obtain the estimated next action for the initial task state; Executing the estimated next action to obtain the estimated next task state of the task in the training sample; Determining a sparse reward based on the estimated next task state and the target task state; Determining a dense reward based on the initial task state, the estimated next task state, and the target task state; Taking the estimated next task state as the initial task state and returning to the step of obtaining the estimated task features until the tasks in the training sample are completed; In response to all training samples having completed the above processing, the parameters of the shared policy network, the action correction policy network and the hybrid encoder are updated based on all sparse rewards and all dense rewards.
2. The agent training method according to claim 1, characterized in that: The determining of a sparse reward based on the estimated next task state and the target task state includes: Mapping the estimated next task state and the target task state to a first latent space; Determining, in the first latent space, a first distance between the estimated next task state and the target task state; The sparse reward is determined based on a relationship between the first distance and a corresponding distance threshold.
3. The agent training method according to claim 1, characterized in that: The determining of the intensive reward based on the initial task state, the estimated next task state and the target task state includes: Mapping the estimated next task state and the target task state to a second latent space; Mapping the initial task state and the predetermined task state to a third latent space, wherein the predetermined task state is the target task state or the estimated next task state; In the second latent space, determining a second distance between the estimated next task state and first predetermined information in the target task state, wherein the first predetermined information is determined based on content of the task in the training sample; In the third latent space, determining a third distance between the initial task state and second predetermined information in the predetermined task state, wherein the second predetermined information is determined based on content of the task in the training sample; The intensive reward is determined based on a relationship between the second distance, the third distance and corresponding distance thresholds.
4. The agent training method according to claim 1, characterized in that: The action correction module includes a weighted summation submodule, a maximization submodule and a minimization submodule. The step of inputting the estimated preliminary action and the estimated correction action into the action correction module to obtain the estimated next action for the initial task state includes: Inputting the estimated preliminary action and the estimated corrective action into the weighted summation submodule to obtain a weighted summation result; Inputting the weighted sum result into the maximization submodule to obtain a maximization result, wherein the maximization submodule is used to obtain a maximum value within the range of the weighted sum result and the first action boundary; Inputting the maximization result into the minimization submodule to obtain the estimated next action, wherein the minimization submodule is used to find the minimum value within the range of the maximization result and the second action boundary; The estimated preliminary action, the estimated correction action and the estimated next action are all within the range of the first action boundary and the second action boundary.
5. A multi-task decision-making method, characterized in that: Applied to an intelligent agent, the intelligent agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the decision method includes: Inputting the initial task state of the target task into the hybrid encoder to obtain task features; Inputting the task features into the shared policy network to generate preliminary actions for the initial task state; Inputting the task characteristics and the preliminary action into the action correction strategy network to generate a correction action for the initial task state; Inputting the preliminary action and the correction action into the action correction module to obtain the next action for the initial task state; Execute the next step action to obtain the next task status of the target task; Taking the next task state as the initial task state and returning to the step of obtaining task features until the target task is completed; Wherein, the intelligent agent is trained by the training method as described in any one of claims 1 to 4 above.
6. A multi-task agent training device, characterized in that: The agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the training device includes: A first acquisition unit is configured to acquire a round of training sample sets, wherein each training sample in the training sample set includes an initial task state and a target task state of a task; The first execution unit is configured to perform the following processing on each training sample in the training sample set: Inputting the initial task state in the training sample into the hybrid encoder to obtain estimated task features; Inputting the estimated task features into the shared policy network to generate an estimated preliminary action for the initial task state; Inputting the estimated task features and the estimated preliminary action into the action correction strategy network to generate an estimated correction action for the initial task state; Inputting the estimated preliminary action and the estimated correction action into the action correction module to obtain the estimated next action for the initial task state; Executing the estimated next action to obtain the estimated next task state of the task in the training sample; Determining a sparse reward based on the estimated next task state and the target task state; Determining a dense reward based on the initial task state, the estimated next task state, and the target task state; Taking the estimated next task state as the initial task state and returning to the step of obtaining the estimated task features until the tasks in the training sample are completed; An updating unit is configured to update the parameters of the shared policy network, the action correction policy network and the hybrid encoder based on all sparse rewards and all dense rewards in response to all training samples having completed the above processing.
7. A decision-making device for multiple tasks, characterized in that: Applied to an intelligent agent, the intelligent agent includes a hybrid encoder, a shared strategy network, an action correction strategy network and an action correction module, and the decision device includes: A second acquisition unit is configured to input the initial task state of the target task into the hybrid encoder to acquire the task feature; A first generating unit is configured to input the task feature into the shared strategy network to generate a preliminary action for the initial task state; A second generating unit is configured to input the task feature into the action correction strategy network to generate a correction action for the initial task state; a fusion unit configured to input the preliminary action and the correction action into the action correction module to obtain the next action for the initial task state; A second execution unit is configured to execute the next step action, obtain the next task state of the target task, use the next task state as the initial task state and return to the step of obtaining task features until the target task is completed; Wherein, the intelligent agent is trained by the training method as described in any one of claims 1 to 4 above.
8. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the multi-task oriented intelligent agent training method as described in any one of claims 1 to 4 and the multi-task oriented decision-making method as described in claim 5.
9. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that: When the instructions are executed by the at least one computing device, they prompt the at least one computing device to execute the multi-task-oriented intelligent agent training method as described in any one of claims 1 to 4 and the multi-task-oriented decision-making method as described in claim 5.
10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by the processor, the multi-task-oriented intelligent agent training method as described in any one of claims 1 to 4 and the multi-task-oriented decision-making method as described in claim 5 are implemented.
Citation Information
Patent Citations
Safety reinforcement learning and safety control method and device, intelligent agent and storage medium
CN116415651A
Intelligent agent learning method and device, equipment and storage medium
CN117035049A
Unmanned aerial vehicle group confrontation control method and system
CN117311392A
Data-driven robot control
WO2021048434A1