Action mask-based agent training method and system

By using a dynamically generated action mask mechanism based on modulated Sigmoid functions in reinforcement learning, the limitations of static and dynamic masks in reinforcement learning are solved, and the training efficiency and autonomous decision-making ability of the agent are improved.

CN120124707APending Publication Date: 2025-06-10NANKAI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510609081.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When using action masks in reinforcement learning in prior art, static masks may too limit the action space, resulting in rigid learning; dynamic masks require manual design and adjustment, which increases complexity and computing overhead, and may affect the exploration of global optimal solutions of the agent in the later stage of training.

Method used

The dynamically generated action mask mechanism based on modulated Sigmoid functions is adopted to dynamically adjust the mask value through the test winning rate of the agent, thereby limiting the action space in the early stage of training, and gradually lifting the restrictions in the later stage to promote the agent's independent learning and exploration.

Benefits of technology

It improves the training efficiency and learning effect of the agent in complex tasks, enhances its independent decision-making ability and adaptability, and avoids the limitations of static and dynamic masks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124707A_ABST
    Figure CN120124707A_ABST
Patent Text Reader

Abstract

The invention provides an agent training method and system based on an action mask, and relates to the field of artificial intelligence, and the method comprises the steps: recording decision data of an expert model in a confrontation process as teaching data; based on a strategy and value function collaborative simulation mechanism, performing mapping learning from a state space to an action space on a strategy network of the intelligent agent according to the teaching data so as to endow the intelligent agent with initial intelligence; performing a confrontation test on the intelligent agent, recording a test winning rate, and entering a subsequent reinforcement learning stage under the condition that the winning rate reaches the standard; inputting the test winning rate into a modulation Sigmoid function to generate an action mask; and updating the intelligent agent by using a reinforcement learning method under the action of the action mask by using a near-end strategy optimization algorithm. According to the method, by designing a degradation mechanism of an action mask, an intelligent agent efficiently avoids sampling of illegal actions at the initial stage of training and explores a strategy space more boldly at the later stage, so that the training efficiency and the final decision performance are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an agent training method and system based on action masking. Background Art

[0002] In reinforcement learning, an agent needs to select an action to execute at each state. The state-action space refers to the set of actions that an agent can choose in a specific state. The action masking technique removes invalid actions from the state-action space, enabling the agent to only select legal and valid actions in the current state. In some problems, the action space may be extremely large, containing a large number of action choices. By using the action masking technique, the complexity of the action space can be reduced to a reasonable range, allowing the agent to learn and make decisions more effectively. Therefore, using action masking helps improve training efficiency because the agent no longer needs to spend time and resources trying invalid actions. By restricting the action space, the agent focuses only on actions that are likely to generate positive rewards and can thus explore effective strategies more quickly. In the process of implementing the present invention, the applicant found that previous research mainly focused on how to use action masking to optimize the reinforcement learning process. Many methods restrict the selection of invalid actions through static or dynamic masking strategies. For example, early research usually used static masking at the initial stage of training to mask illegal actions and help the agent focus on effective learning. Subsequently, some research introduced a dynamic masking mechanism that dynamically adjusts the masking range according to the learning progress of the agent, in order to gradually reduce the dependence on masking in the later stage of training and promote broader exploration by the agent. Although these methods have achieved positive results, there are still some limitations: First, static masking may overly restrict the action space at the initial stage of the agent, resulting in an overly rigid learning process. Second, although dynamic masking can be adjusted flexibly, it still requires manual design and adjustment, increasing the complexity and computational overhead of the training process. In addition, strategies that overly rely on action masking may cause the agent to be unable to effectively explore the global optimal solution in the later stage of training, thus affecting the final decision-making ability. Therefore, how to effectively use action masking for reinforcement learning to improve the autonomous decision-making ability of the agent has become a technical problem to be solved. Summary of the Invention

[0003] The present invention aims to at least solve one of the technical problems existing in the prior art or related technologies, and discloses an agent training method and system based on action masking, enabling the agent to learn effective strategies more stably and quickly, and improving the autonomous decision-making ability of the agent.

[0004] The first aspect of the present invention discloses an agent training method based on action masks, including: expert teaching: constructing an expert model and an opponent model, and recording the decision-making data of the expert model as teaching data during the confrontation between the expert model and the opponent model; imitation learning: constructing an agent to be trained, and based on the collaborative imitation mechanism of the policy and value function, mapping learning from the state space to the action space is performed on the policy network of the agent according to the teaching data to endow the agent with initial intelligence; intelligent testing: performing confrontation testing on the agent with initial intelligence and recording the test winning rate. If the winning rate meets the standard, enter the subsequent reinforcement learning stage. If the winning rate does not meet the standard, continue with imitation learning; or performing confrontation testing on the agent trained through reinforcement learning and recording the test winning rate; dynamically generating action masks: inputting the test winning rate into a modulated Sigmoid function to generate action masks. The higher the winning rate, the larger the generated mask value, and the smaller the restrictive effect of the mask on the action space; updating the agent: using the proximal policy optimization algorithm, updating the agent using the reinforcement learning method under the action of the action mask, and performing intelligent testing on the updated agent and generating new action masks to implement cyclic training.

[0005] In this technical solution, action masks are dynamically generated based on the modulated Sigmoid function. Under the action of the action masks, real-time interaction data under the action of the action masks is collected through parallel sampling, and the agent is updated on the real-time interaction trajectory through the proximal policy optimization method; subsequently, the winning rate of the new agent is tested and new action masks are generated until the loop stops when the training requirements are met. The action masks are used to filter illegal actions. The mask values generated in the early stage of training are small, and sampling of illegal actions is avoided as much as possible. The mask values generated in the later stage of training are relatively large. The agent is equivalent to sampling in the original action space, boldly exploring the feasibility of various strategies. As the winning rate increases, the mask gradually degrades, promoting the improvement of its autonomous learning and exploration capabilities until the agent training requirements are met. In complex scenarios such as high-fidelity air combat games, the mask automatic degradation mechanism is used for agent training, improving the performance of the agent in complex tasks and its autonomous game confrontation ability.

[0006] According to the agent training method based on action masks disclosed by the present invention, preferably, it further includes: in the imitation learning stage, the Monte Carlo method is used to iteratively optimize the value function.

[0007] According to the agent training method based on action masks disclosed by the present invention, preferably, the specific process of iterative optimization includes:

[0008]

[0009]

[0010] Among them, is the loss of Monte Carlo optimization, are the parameters of the value network, denotes taking the expectation of the subsequent expression under the constraint of the policy denotes the value predicted by the value network at state value, is the discounted cumulative return, represents the discount factor, is the learning rate, is the value after the update of the value network parameters.

[0011] According to the agent training method based on action masking disclosed in the present invention, preferably, it further includes: regarding the test win rate as a time series, and using the exponentially weighted moving average method to filter the test win rate to make the time series smoother and reduce the influence of gradient jumps.

[0012] According to the agent training method based on action masking disclosed in the present invention, preferably, the calculation process of the exponentially weighted moving average method is specifically divided into two steps. First, smooth the win rate at the current moment, and then, in order to offset the underestimation bias caused by smoothing in the early stage, perform normalization correction on the smoothed value. The update formula includes:

[0013]

[0014]

[0015] where is the win rate of the agent during the test at moment, represents the unnormalized smoothed value at moment, is the smoothing factor, is the win rate after normalization correction.

[0016] According to the agent training method based on action masking disclosed in the present invention, preferably, the steps of dynamically generating the action mask specifically include:

[0017]

[0018] wherein, represents the win rate after normalization correction, b and d are hyperparameters, and m represents the generated action mask.

[0019] According to the agent training method based on action masking disclosed in the present invention, preferably, the steps of updating the agent specifically include:

[0020]

[0021]

[0022]

[0023]

[0024] Among them, represents the distance between the new and old policies, and clip is a truncation function. is the truncation coefficient. is the traditional loss for proximal policy optimization. is the action-value function. is the value function.

[0025] The second aspect of the present invention discloses an agent training system based on action masking, including: a memory for storing program instructions; a processor for calling the program instructions stored in the memory to implement the agent training method based on action masking as described in any of the above technical solutions.

[0026] The beneficial effects of the present invention at least include: by introducing a mask auto-degradation mechanism and cooperating with a win rate feedback mask generation mechanism, the training efficiency and learning effect of the agent in the face of complex tasks are effectively improved. By using action masks to shield invalid actions, the agent can quickly focus on effective strategies. As the training progresses (the win rate gradually increases), the action masks automatically degrade, gradually lifting the restrictions on the agent's action selection, and promoting the improvement of the agent's autonomous learning and exploration capabilities. The adaptability and decision-making ability of the agent in complex environments are improved, providing an efficient and comprehensive solution for agent learning in various fields that need to deal with complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 shows a schematic flowchart of an agent training method based on action masking according to an embodiment of the present invention.

[0028] Figure 2 shows a schematic block diagram of an agent training system based on action masking according to an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to more clearly understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the present invention is not limited to the limitations of the specific embodiments disclosed below.

[0030] According to an embodiment of the present invention, an agent training method based on action masks is disclosed, including:

[0031] Expert demonstration: Construct an expert model and an opponent model. During the confrontation between the expert model and the opponent model, record the decision-making data of the expert model as demonstration data;

[0032] Imitation learning: Construct an agent to be trained. Based on the collaborative imitation mechanism of policy and value function, perform mapping learning from the state space to the action space on the policy network of the agent according to the demonstration data, and use the Monte Carlo method to iteratively optimize the value function to endow the agent with initial intelligence;

[0033] Intelligent testing: Conduct confrontation testing on the agent with initial intelligence and record the test winning rate. If the winning rate meets the standard, enter the subsequent reinforcement learning stage; if the winning rate does not meet the standard, continue with imitation learning; or conduct confrontation testing on the agent trained through reinforcement learning and record the test winning rate;

[0034] Winning rate data filtering: Regard the test winning rate as a time series, and use the exponentially weighted moving average method to filter the test winning rate to make the time series smoother;

[0035] Dynamically generate action masks: Input the test winning rate into the modulated Sigmoid function to generate action masks. The higher the winning rate, the larger the generated mask value, and the smaller the restrictive effect of the mask on the action space;

[0036] Update the agent: Use the proximal policy optimization algorithm to update the agent using the reinforcement learning method under the action of the action mask, and conduct intelligent testing on the updated agent and generate new action masks to implement cyclic training.

[0037] As Figure 1 shown, in the high-fidelity dogfight game scenario, the specific implementation process of the agent training method based on action masks in the above embodiment is as follows:

[0038] S1: Collect demonstration data using expert experience;

[0039] S2: Introduce the collaborative imitation mechanism of policy and value function to endow the agent with initial intelligence;

[0040] S3: Set up a test module to determine whether to continue with demonstration;

[0041] S4: Filter the test winning rate;

[0042] S5: Dynamically generate action masks using the modulated Sigmoid function;

[0043] S6: Update the agent using proximal policy optimization under the action of the action mask.

[0044] Step S1 specifically includes: constructing two adversarial strategies using rules, one as an expert demonstration strategy and the other as an opponent strategy. During the confrontation, record the decision-making data of each step of the expert strategy: .

[0045] Among them, is the state of the expert strategy at time , where represents the expert strategy. represents the action, represents the reward, represents the parameter of the strategy under

[0046] which can receive the current state and return the probability of a certain action.

[0047]

[0048]

[0049] Among them, is the loss of cloning learning, are the parameters of the policy network, means taking the expectation of the following formula, represents the distribution of expert demonstration data, the policy network selects the action at the state with probability , is the learning rate used to control the step size during each update, is the updated amount of the policy network parameters.

[0050]

[0051]

[0052] Among them, is the loss of Monte Carlo optimization, are the parameters of the value network, means taking the expectation of the subsequent formula under the constraint of the policy , is the discounted cumulative return, represents the value predicted by the value network at the state , is the learning rate used to control the step size at each update, is the gradient of the loss value that guides the direction of parameter update, is the value after the update of the value network parameters.

[0053] Step S3 specifically includes: using the current policy to play against the opponent for 40 games as a test task. The winning rate of the test task is used as the intelligent evaluation index of the agent. In this embodiment, the winning rate index is set to 0.3. When the current winning rate of the agent is greater than 0.3, the policy is updated by combining reinforcement learning with action masking, otherwise the policy and value function co-imitation stage continues.

[0054] Step S4 specifically includes: using the exponential weighted moving average method to smooth the stage winning rate of the agent to reduce the interference of single fluctuations on the overall trend. This process is divided into two steps: first, weight and smooth the winning rate at the current moment; subsequently, in order to correct the underestimation bias caused by the lack of historical data in the initial stage, the smoothed value is normalized and corrected. The update formula is as follows:

[0055]

[0056]

[0057] where, is the winning rate of the agent during the test process at time is the smoothing factor, which is taken as 0.9 in this example, represents the non-normalized exponentially weighted average at time -1, then represents the non-normalized exponentially weighted average at time is the normalization coefficient corresponding to time t, is the finally smoothed winning rate after bias correction.

[0058] Step S5 specifically includes: introducing a mask generation mechanism based on winning rate feedback, processing the smoothed test winning rate through a modulated Sigmoid function to generate the corresponding action mask m, and the generation formula is as follows:

[0059]

[0060] where, the hyperparameter controls the translation position of the function and affects the starting point of the curve. The hyperparameter controls the slope or steepness of the function, the larger it is, the closer it is to the step function. They jointly determine the shape of the function. In this embodiment, through the trial-and-error method, and Are set to 1613.369 and 9.817 respectively, and are used to map the input parameters to an output within an interval . represents the winning rate after normalization correction, and m is the finally generated action mask.

[0061] Step S6 specifically includes: using the action mask to update the current policy of the agent to , and using to compete with the opponent. During the competition, record for each step of decision-making data, that is . Use the proximal policy optimization algorithm to update the policy , and the update formula is as follows:

[0062]

[0063]

[0064]

[0065]

[0066] where represents the probability distribution of all actions of the original policy in state , is the mask for the action space in the current state , represents element-wise multiplication, is the probability distribution of all actions of the masked policy in state . measures the distance between the old and new policies and , is a truncation function used to control the input x between min and max, is the truncation coefficient. is the traditional loss for proximal policy optimization, is the action value calculated for state and action under policy , is the value estimated for state under policy , represents taking the expectation of the subsequent expression under the constraint of policy . is the weight of the new policy network, are the weights of the old policy network, is the update step size, represents the gradient of the loss. After iteration, the new policy is passed into the test module to continue generating new dynamic masks, and this process is looped until the training requirements are met. In this embodiment, the training stop requirement is 400,000 training iterations.

[0067] As Figure 2 shown, according to another embodiment of the present invention, an agent training system 200 based on action masks is also disclosed, including: a memory 201 for storing program instructions; a processor 202 for calling the program instructions stored in the memory to implement the agent training method based on action masks as described in the above embodiment.

[0068] In summary, the present invention provides an agent training method and system with mask automatic degradation for complex tasks. By making full use of the advantages of action masks and introducing a win rate feedback mask generation mechanism, the method effectively adapts to the ability growth characteristics of the agent during the training process. By gradually degrading the mask dependence, the agent can efficiently avoid sampling illegal actions in the initial stage and explore the strategy space more boldly in the later stage, thus significantly improving the training efficiency and the final decision-making performance.

[0069] All or part of the steps in the various methods of the above embodiments can be completed by a program controlling the relevant hardware. The program can be stored in a readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, a magnetic disk memory, a tape memory, or any other readable medium capable of carrying or storing data.

[0070] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An agent training method based on action mask, characterized in that: include: Expert teaching: constructing an expert model and an opponent model, and in the process of confrontation between the expert model and the opponent model, recording the decision data of the expert model as teaching data; Imitation learning: construct an intelligent agent to be trained, and based on the collaborative imitation mechanism of strategy and value function, learn the mapping from state space to action space of the intelligent agent's strategy network according to the teaching data to give the intelligent agent initial intelligence; Intelligent testing: Conduct adversarial testing on the intelligent agents with initial intelligence and record the test winning rate. If the winning rate meets the standard, it will enter the subsequent reinforcement learning stage. If the winning rate does not meet the standard, it will continue to carry out imitation learning. Or conduct adversarial testing on agents trained with reinforcement learning and record the test win rate; Dynamically generate action masks: Input the test win rate into the modulated Sigmoid function to generate action masks. The higher the win rate, the larger the generated mask value, and the less restrictive effect of the mask on the action space; Updating the intelligent agent: using the proximal policy optimization algorithm, using the reinforcement learning method to update the intelligent agent under the action mask, and performing intelligent testing on the updated intelligent agent and generating a new action mask to implement cyclic training.

2. The method for training an intelligent agent based on action mask according to claim 1, characterized in that: Also includes: In the imitation learning stage, the Monte Carlo method is used to iteratively optimize the value function.

3. The method for training an intelligent agent based on action mask according to claim 2, characterized in that: The iterative optimization process specifically includes: in, is the loss of Monte Carlo optimization, are the parameters of the value network, Representation strategy The expected value under the constraints, is the discounted cumulative return, Represents the state predicted by the value network The value of is the learning rate used to control the step size at each update, is the loss value gradient that guides the direction of parameter update, is the updated value of the value network parameter.

4. The method for training an intelligent agent based on action mask according to claim 1, characterized in that: Also includes: The test winning rate is regarded as a time series, and the test winning rate is filtered using an exponentially weighted moving average method to make the time series smoother.

5. The method for training an intelligent agent based on action mask according to claim 4, characterized in that: The calculation process of the exponentially weighted moving average method includes: in, Is the agent in the testing process The winning rate of the moment, is the smoothing factor, express The non-normalized exponentially weighted average at time -1, It represents The non-normalized exponentially weighted average of the time, is the normalization coefficient corresponding to time t, is the final smoothed winning rate after deviation correction.

6. The method for training an intelligent agent based on action mask according to any one of claims 1 to 5, characterized in that: The step of dynamically generating an action mask specifically includes: Among them, m represents the action mask, b and d are constant parameters, Represents the normalized corrected win rate.

7. The method for training an intelligent agent based on action mask according to any one of claims 1 to 5, characterized in that: The steps to update the agent include: in, Indicates that the original policy is in state The probability distribution of all actions is: is in the current state The mask of the action space is represents element-wise multiplication, Is the policy after masking in state The probability distribution of all actions is: Weighing old and new strategies and The distance between represents the truncation function, is the cutoff coefficient, The traditional loss optimized for proximal strategies, It is in strategy Next pair status and actions Calculate the action value, It’s a strategy Next for status Estimated value, Indicated in strategy Under the constraint of , we can find the expectation of the subsequent formula. is the weight of the new policy network, is the weight of the old policy network, is the update step size, Represents the gradient of the loss.

8. An agent training system based on action mask, characterized in that: include: A memory for storing program instructions; A processor, configured to call the program instructions stored in the memory to implement the action mask-based agent training method as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Graphical user interface agent training and prediction method and device, program and medium

    CN121598986A

  • Intelligent digital culture creative content generation method based on reinforcement learning

    CN121958579A