A game agent design method and system based on deep reinforcement learning

By combining the agent design method of deep neural network and Monte Carlo tree search algorithm, the problem of inefficient search in complex games by traditional Monte Carlo tree search algorithm is solved, and more efficient search and decision-making capabilities are achieved.

CN118444887BActive Publication Date: 2025-05-16JILIN ANIMATION INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410798318.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-05-16
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

The traditional Monte Carlo tree search algorithm is inefficient in searching in agent games and is difficult to achieve practical application in complex games.

Method used

The game agent design method based on deep reinforcement learning is adopted, combined with deep neural networks and Monte Carlo tree search algorithm, and predict the value of state and action through deep learning models, and the Monte Carlo tree search method is fused for action search and selection.

Benefits of technology

It improves the search efficiency and decision-making accuracy of the agent in complex gaming environments, and improves the game winning rate and problem-solving speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118444887B_ABST
    Figure CN118444887B_ABST
Patent Text Reader

Abstract

A method and system for designing a game agent based on deep reinforcement learning, which relates to the field of deep reinforcement learning of agents. To solve the problem that the Monte Carlo tree search algorithm used in existing agent games has low search efficiency and is difficult to apply in practice, the present invention provides the following solution: using a deep learning algorithm to model the game state space, and using a neural network to estimate the state value function; using an improved Monte Carlo tree search algorithm to search the game, and using the state value function to evaluate; using a reinforcement learning algorithm to learn the game to learn the optimal strategy; finally, evaluating and optimizing the training results to improve the performance and stability of the game system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent agent deep reinforcement learning. Background Art

[0002] In the field of artificial intelligence, reinforcement learning is an important learning method that enables intelligent agents to learn the optimal behavior strategy through continuous trial and error and reward and punishment mechanisms. Unlike supervised learning and unsupervised learning, the goal of reinforcement learning is not to learn the mapping relationship between input and output, but to learn how to make the best decision in an environment to maximize long-term cumulative rewards.

[0003] In reinforcement learning, the agent learns by interacting with the environment. Its behavior will have a certain impact on the environment. The environment will give a reward or punishment signal based on the agent's behavior, and the agent will adjust its behavior strategy based on the reward or punishment signal. Reinforcement learning mainly includes three elements: value function, strategy, and environment. Among them, the value function refers to the expected value of the long-term cumulative reward that the agent can obtain by taking a certain action in a certain state, the strategy refers to the probability distribution of which action the agent takes in a certain state, and the environment includes the state of the agent and the results of the agent's actions.

[0004] In the field of game theory, the Monte Carlo Tree Search algorithm is a common search algorithm that predicts the optimal action plan by simulating the opponent's behavior and randomly selecting the next action. The basic idea of ​​the Monte Carlo Tree Search algorithm is to start from the root node and continuously expand the tree nodes by simulating the opponent's behavior and randomly selecting the next action until a certain search depth or time limit is reached. In this way, the Monte Carlo Tree Search algorithm can efficiently search for the optimal action plan, especially in complex games, which can effectively improve the winning rate of the intelligent agent.

[0005] However, the traditional Monte Carlo tree search algorithm has problems such as too large search space and low search efficiency. In complex games such as Go, the search space can even reach 10 170 ,The traditional Monte Carlo tree search algorithm needs to continuously simulate opponents for each node to obtain a reliable value estimate.,This makes the search efficiency of the traditional Monte Carlo tree search algorithm very low and ,difficult to implement in practical applications. Summary of the invention

[0006] The present invention aims at the problem that the search efficiency of the Monte Carlo tree search algorithm used in the agent game in the prior art is very low and it is difficult to be used in practical applications. To solve the above technical problems, the present invention is implemented by the following technical solutions:

[0007] The present invention proposes a game agent design method based on deep reinforcement learning, the method comprising:

[0008] Follow these steps:

[0009] S1. Define and abstract the game environment and tasks; game agent, used to maximize its long-term reward by learning strategies in the game environment;

[0010] S2. Initialize the game environment and tasks and construct a deep neural network as a strategy value network, wherein the strategy value network includes state values ​​and different strategies under a certain state;

[0011] S3. By interacting with the game environment described in S1, the game agent collects a sequence of states, actions, and rewards as game data, uses the collected game data to train the strategy value network described in S2, uses strategy iteration to implement deep reinforcement learning training, and evaluates the state value function;

[0012] S4, integrating the Monte Carlo tree search method to search and select actions based on the current game state;

[0013] S5. Based on the results of the Monte Carlo tree search and the output of the policy value network, select the optimal action to execute;

[0014] S6, the game agent executes the selected action and determines the feedback from the game environment; updates the deep neural network to optimize the prediction accuracy of the policy and value function;

[0015] S7. Repeat S3 to S6 until the predetermined training termination condition is met, and the design of the game agent is completed.

[0016] Furthermore, a preferred implementation is provided, in which the deep reinforcement learning training implemented using policy iteration in S3 includes two sub-steps: policy evaluation and policy improvement.

[0017] Further, a preferred implementation is provided, wherein the two sub-steps of strategy evaluation and strategy improvement include the following steps:

[0018] S21. In the strategy evaluation phase, determine the state value function or the state and action value function under the current strategy;

[0019] S22. In the strategy improvement phase, a new strategy is generated based on the state value function.

[0020] Further, a preferred embodiment is provided, in which the searching and selecting of actions based on the current game state in S4 comprises the following steps:

[0021] S41. Starting from the root node, recursively select child nodes until reaching a leaf node; each selection is based on a specific strategy.

[0022] S42, after reaching the leaf node, if this node is not the end point of the game, generate one or at least one successor child node;

[0023] S43, starting from the newly expanded node, simulating the game to the end in a random manner to determine the outcome of the game;

[0024] S44. Update the statistical information of all nodes on the path from the current leaf node to the root node according to the simulation results, repeat the above steps until the preset search time or search number limit is reached, and select the move with the highest winning rate or the best specific value.

[0025] Further, a preferred implementation is provided, wherein S5 includes:

[0026] S51, select the optimal action, adjust the selection strategy in combination with the strategy output provided by the deep neural network, and the deep neural network outputs the probability of each action; adjust the probability of each action by the number of visits of the Monte Carlo tree search method, that is, weight the action probability output by the network by the logarithm of the number of visits;

[0027] S52. The game agent selects an action based on the adjusted action probability and selects the action with the highest probability as the final decision.

[0028] Furthermore, a preferred implementation is provided, in which the prediction accuracy of the optimization strategy and the value function in S6 is achieved by using an experience replay mechanism.

[0029] Further, a preferred implementation is provided, wherein S6 includes:

[0030] S61, the game agent performs an action in the game environment, moves from the current state to the new state, and receives an immediate reward associated with the action;

[0031] S62. Update the policy value network based on the observed rewards and new states;

[0032] S63. If the policy gradient method is used for policy update, the policy value network is optimized to directly output the probability distribution of the action.

[0033] Solution 2: A game agent design system based on deep reinforcement learning, the system comprising:

[0034] An input module, used to define and abstract the game environment and tasks, and a game agent, used to maximize its long-term reward by learning strategies in the game environment;

[0035] Construct a strategy value network module to initialize the game environment and tasks and construct a deep neural network as a strategy value network, which includes state values ​​and different strategies under a certain state. ;

[0036] A state value function evaluation module is used to collect the sequence of states, actions and rewards as game data by the game agent through interaction with the game environment described in the definition module, use the collected game data to train the strategy value network described in the strategy value network construction module, use strategy iteration to implement deep reinforcement learning training, and evaluate the state value function;

[0037] The integrated Monte Carlo Tree Search module is used to integrate the Monte Carlo Tree Search method to search and select actions based on the current game state;

[0038] The execution module is used to select the optimal action to execute based on the results of the Monte Carlo tree search and the output of the policy value network;

[0039] A prediction module that optimizes the policy and value functions for the agent to perform the selected actions and determine the feedback from the game environment; updates the deep neural network to optimize the prediction accuracy of the policy and value functions;

[0040] The output module is used to repeatedly evaluate the state-value function model to optimize the strategy and value function prediction module until the predetermined training termination conditions are met to complete the design of the game agent.

[0041] Solution three: A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes any one of the methods described in Solution one.

[0042] Solution 4: A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of Solution 1.

[0043] The present invention is beneficial in that:

[0044] By integrating deep learning and reinforcement learning strategies, the present invention enables the agent to learn complex patterns and dynamics in the game environment more efficiently, thereby demonstrating higher performance in various game tasks. This performance improvement is not only reflected in the winning rate of the game, but also in the speed and efficiency of the agent in solving problems.

[0045] The present invention combines the use of deep neural networks with Monte Carlo tree search, so that the intelligent agent can more comprehensively consider various possible actions and their results when making decisions, making the decision process more accurate and reasonable. This enhanced decision-making ability enables the intelligent agent to show higher adaptability and flexibility when facing complex decision-making environments.

[0046] The present invention is also suitable for applying deep reinforcement learning of intelligent agents in technical fields such as game theory. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flowchart of a method for designing a game agent based on deep reinforcement learning as described in Implementation Method 1.

[0048] Figure 2 This is a flowchart of the game agent training in the game agent design method based on deep reinforcement learning described in implementation mode eleven. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the implementation methods of the present application clearer, the technical solutions in the implementation methods of the present application will be clearly and completely described below in conjunction with the drawings in the implementation methods of the present application. Obviously, the described implementation methods are only part of the implementation methods of the present application, not all of the implementation methods.

[0050] Implementation method 1: This implementation method provides a game agent design method based on deep reinforcement learning.

[0051] The method comprises the following steps:

[0052] S1. Define and abstract the game environment and tasks; game agent, used to maximize its long-term reward by learning strategies in the game environment;

[0053] S2. Initialize the game environment and tasks and construct a deep neural network as a strategy value network, wherein the strategy value network includes state values ​​and different strategies under a certain state;

[0054] S3. By interacting with the game environment described in S1, the game agent collects a sequence of states, actions, and rewards as game data, uses the collected game data to train the strategy value network described in S2, uses strategy iteration to implement deep reinforcement learning training, and evaluates the state value function;

[0055] S4, integrating the Monte Carlo tree search method to search and select actions based on the current game state;

[0056] S5. Based on the results of the Monte Carlo tree search and the output of the policy value network, select the optimal action to execute;

[0057] S6, the game agent executes the selected action and determines the feedback from the game environment; updates the deep neural network to optimize the prediction accuracy of the policy and value function;

[0058] S7. Repeat S3 to S6 until the predetermined training termination condition is met, and the design of the game agent is completed.

[0059] The deep learning described in this embodiment is a learning method that can efficiently process large-scale data, and can effectively solve the problem of low search efficiency in the Monte Carlo tree search algorithm. Deep learning mainly uses neural networks to model and predict input data, and has achieved important results in the fields of image recognition, speech recognition, natural language processing, etc. In reinforcement learning, deep learning can be used to model value functions and policy functions to improve the efficiency and accuracy of reinforcement learning. By using a deep learning model to predict the value of states and actions, the number of searches of the Monte Carlo tree search algorithm can be reduced, thereby improving search efficiency and accuracy.

[0060] Therefore, the Monte Carlo tree search reinforcement learning algorithm based on deep learning described in this embodiment mainly includes two parts: one is to use the deep learning model to predict the value of each action in the current state; the other is to simulate the opponent's behavior through the Monte Carlo tree search algorithm to obtain the next possible action, and use the deep learning model to predict the value of each possible action and select the optimal action. By continuously executing the selected action and rewarding or punishing according to the execution results, the parameters of the deep learning model and the Monte Carlo tree search algorithm can be updated to adjust the strategy. In this way, the Monte Carlo tree search reinforcement learning algorithm combined with deep learning can be applied in fields such as games to improve the winning rate and game experience of the intelligent agent.

[0061] Implementation method 2. This implementation method is a further limitation of the game agent design method based on deep reinforcement learning described in implementation method 1. In S3, strategy iteration is used to implement deep reinforcement learning training, including two sub-steps: strategy evaluation and strategy improvement.

[0062] Implementation method 3: This implementation method is a further limitation of the game agent design method based on deep reinforcement learning described in implementation method 2. The two sub-steps of strategy evaluation and strategy improvement include the following steps:

[0063] S21. In the strategy evaluation phase, determine the state value function or the state and action value function under the current strategy;

[0064] S22. In the strategy improvement phase, a new strategy is generated based on the state value function.

[0065] Implementation 4: This implementation is a further limitation of the game agent design method based on deep reinforcement learning described in Implementation 1. The search and selection of actions based on the current game state in S4 includes the following steps:

[0066] S41. Starting from the root node, recursively select child nodes until reaching a leaf node; each selection is based on a specific strategy.

[0067] S42, after reaching the leaf node, if this node is not the end point of the game, generate one or at least one successor child node;

[0068] S43, starting from the newly expanded node, simulating the game to the end in a random manner to determine the outcome of the game;

[0069] S44. Update the statistical information of all nodes on the path from the current leaf node to the root node according to the simulation results, repeat the above steps until the preset search time or search number limit is reached, and select the move with the highest winning rate or the best specific value.

[0070] Implementation 5: This implementation is a further limitation of the game agent design method based on deep reinforcement learning described in Implementation 1, and S5 includes:

[0071] S51, select the optimal action, adjust the selection strategy in combination with the strategy output provided by the deep neural network, and the deep neural network outputs the probability of each action; adjust the probability of each action by the number of visits of the Monte Carlo tree search method, that is, weight the action probability output by the network by the logarithm of the number of visits;

[0072] S52. The game agent selects an action based on the adjusted action probability and selects the action with the highest probability as the final decision.

[0073] Implementation method six: This implementation method is a further limitation of the game agent design method based on deep reinforcement learning described in implementation method one. The prediction accuracy of the optimization strategy and value function in S6 is achieved by the agent using the experience replay mechanism.

[0074] Implementation 7: This implementation is a further limitation of the game agent design method based on deep reinforcement learning described in Implementation 1, and S6 includes:

[0075] S61, the game agent performs an action in the game environment, moves from the current state to the new state, and receives an immediate reward associated with the action;

[0076] S62. Update the policy value network based on the observed rewards and new states;

[0077] S63. If the policy gradient method is used for policy update, the policy value network is optimized to directly output the probability distribution of the action.

[0078] Implementation 8: This implementation proposes a game agent design system based on deep reinforcement learning.

[0079] System, the system comprising:

[0080] An input module, used to define and abstract the game environment and tasks, and a game agent, used to maximize its long-term reward by learning strategies in the game environment;

[0081] Construct a strategy value network module to initialize the game environment and tasks and construct a deep neural network as a strategy value network, which includes state values ​​and different strategies under a certain state. ;

[0082] A state value function evaluation module is used to collect the sequence of states, actions and rewards as game data by the game agent through interaction with the game environment described in the definition module, use the collected game data to train the strategy value network described in the strategy value network construction module, use strategy iteration to implement deep reinforcement learning training, and evaluate the state value function;

[0083] The integrated Monte Carlo Tree Search module is used to integrate the Monte Carlo Tree Search method to search and select actions based on the current game state;

[0084] The execution module is used to select the optimal action to execute based on the results of the Monte Carlo tree search and the output of the policy value network;

[0085] A prediction module that optimizes the policy and value functions for the game agent to perform the selected actions and determine the feedback from the game environment; updates the deep neural network to optimize the prediction accuracy of the policy and value functions;

[0086] The output module is used to repeatedly evaluate the state-value function model to optimize the strategy and value function prediction module until the predetermined training termination conditions are met to complete the design of the game agent.

[0087] Embodiment 9: This embodiment proposes a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the method described in any one of embodiments 1 to 7.

[0088] Embodiment 10: This embodiment proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method described in any one of Embodiments 1 to 7 are implemented.

[0089] Implementation eleven: The examples proposed in this implementation are used to explain the above implementations one to ten, and specifically include the following contents:

[0090] See also Figure 1 In this embodiment, a method for designing a game agent based on deep reinforcement learning is provided in this embodiment, comprising the following steps:

[0091] S1. Define and abstract the game environment and tasks, clarify the game rules, goals T, state space, possible action sets D, and the representation method of the game state S. The game agent needs to maximize its long-term rewards by learning strategies in this environment;

[0092] S2. Initialize and construct a deep neural network as a policy value network, which mainly consists of state value V and different strategies P in a given state;

[0093] S3. Through interaction with the environment, the game agent collects game data by collecting sequences of states, actions, and rewards, and uses policy iteration to implement deep reinforcement learning training. The collected experience is used to train the policy value network, evaluate the state value function, and optimize the match between its predicted policies and values ​​and the actual results.

[0094] S4, integrated Monte Carlo Tree Search (MCTS), MCTS explores the action space through a large number of random simulations, and searches and selects actions based on the current game state;

[0095] S5. Based on the results of MCTS and the output of the strategy value network, select the optimal action to execute;

[0096] S6, the game agent performs the selected action and observes the feedback from the environment, including new states and rewards; updates the deep neural network to optimize the prediction accuracy of the policy and value function;

[0097] S7. Repeat S3 to S6 until a predetermined training termination condition is met.

[0098] In one embodiment, the above S1 includes: defining and abstracting the game environment and tasks, clarifying the game rules, the goal T, the state space, the possible action set D, and the representation method of the game state S. The game agent needs to maximize its long-term reward by learning strategies in this environment;

[0099] This step is the basis for designing a game agent based on deep reinforcement learning. This step requires the following key elements to be clarified: Game rules, which describe the basic rules of the game in detail, including the conditions for starting, progressing, and ending the game, as well as the operations that players can perform in the game; Goal T, which clearly defines the ultimate goal of the game agent in the game. This goal is usually to win the game or achieve a specific achievement, which will directly affect the learning direction and strategy selection of the game agent; State space, which defines all possible game states, each of which represents the status of the game at a certain moment; Action set D, which lists the actions that the game agent can take in any given state; Representation method of game state S, which determines how to represent the game state in the game agent. This usually involves converting the game state into a data format that the agent can handle, such as vectors, matrices, or more complex data structures; Maximizing its long-term reward. The goal of the game agent is not only to win in a single game, but also requires the game agent to be able to evaluate the rewards that different actions may bring in the future and make decisions based on these evaluations;

[0100] In one embodiment, the above S2 includes: initializing and constructing a deep neural network as a policy value network, which is mainly composed of a state value V and different strategies P under a given state.

[0101] Deep neural networks are trained to output the probability distribution of taking different actions in a given state and the state value. This training process is mainly achieved through policy iteration, including two sub-steps: policy evaluation and policy improvement.

[0102] S21. In the strategy evaluation phase, the goal is to estimate the current strategy The state value function under Or the state-action value function For the state value function , which can be calculated by the Bellman equation:

[0103]

[0104] in It is in strategy Down state s the value of is in state s Next select action a probability; From the state s Take Action a Transfer to state s' And get rewards r probability; is the discount factor, ranging between 0 and 1.

[0105] For the state-action value function , its Bellman equation is:

[0106]

[0107] S22. The goal of the strategy improvement phase is to generate a better strategy based on the value function. A common strategy improvement method is greedy strategy improvement, which selects Q Generate a new policy by taking actions that maximize the value:

[0108]

[0109] In practical applications, deep neural networks are used to approximate or , and strategies The parameters of the network are updated via a gradient descent algorithm to minimize the gap between the predicted value and the actual value obtained, or to maximize the expected reward.

[0110] In one embodiment, the above S3 includes: through interaction with the environment, the game agent collects game data by collecting sequences of states, actions and rewards, and uses policy iteration to implement deep reinforcement learning training. The collected experience is used to train the policy value network, evaluate the state value function, and optimize the matching degree between its predicted strategy and value and the actual result.

[0111] The game agent collects data by performing actions and observing the results during the interaction with the game environment. Each interaction includes the current state S, the action A taken, the immediate reward R obtained, and the new state S' transferred to. This sequence of data reflects the dynamic process of the game and the decision-making results of the agent. The strategy is continuously optimized through an iterative process. , in order to achieve the goal of maximizing long-term rewards. Policy iteration usually includes two main steps: policy evaluation and policy improvement. In the policy evaluation step, the algorithm evaluates the current policy The state-value function V(s) under ; in the policy improvement step, the algorithm updates the policy to improve the expected reward in each state.

[0112] In one embodiment, the above S4 includes: integrating Monte Carlo tree search, MCTS explores the action space through a large number of random simulations, and searches and selects actions based on the current game state;

[0113] S41. Starting from the root node, recursively select child nodes until reaching a leaf node. Each selection is based on a specific strategy, such as UCB1 (Upper Confidence Bound 1), which aims to balance exploration and exploitation:

[0114]

[0115] in, is the value of the node, C is the exploration parameter, is the number of visits to the parent node, is the number of times a particular action is chosen;

[0116] S42, after reaching the leaf node, if this node is not the end point of the game, one or more possible successor child nodes are generated;

[0117] S43, starting from the newly expanded node, simulating the game to the end using a random method or a certain heuristic strategy to achieve a winning or losing result of the game;

[0118] S44, according to the simulation results, update the statistical information of all nodes on the path from the current leaf node to the root node, such as the number of wins and the number of visits. If a win occurs, the number of wins is increased on the corresponding node, and the number of visits to the node increases by 1 for each visit.

[0119]

[0120] This process is repeated until the preset search time or search number limit is reached, and then the move with the highest win rate or the best UCB1 value is selected.

[0121] In one embodiment, the above S5 includes: selecting the optimal action to execute based on the result of MCTS and the output of the policy value network;

[0122] S51. Select the best action and adjust the selection strategy based on the strategy output provided by the deep neural network to better balance exploration and exploitation. The deep neural network outputs the probability of each action, expressed as ,in s is the current state. A common approach is to adjust these probabilities based on the number of visits by MCTS, that is, to weight the action probabilities output by the network by the logarithm of the number of visits to increase the selection probability of actions that appear more promising after MCTS exploration:

[0123]

[0124] in is the adjusted action probability, It's action a The number of visits, is a temperature parameter that adjusts the degree of exploration. The value will make the selection more concentrated on actions with higher visit counts, while higher Lower values ​​encourage more exploration.

[0125] S52, the game agent takes action according to the adjusted probability To select an action, usually the action with the highest probability is selected as the final decision. This method combines the exploration advantage of MCTS with the predictive ability of deep neural networks, allowing the intelligent agent to make more optimized action choices in complex decision-making environments.

[0126] In one embodiment, the above S6 includes: the agent performs the selected action and observes the feedback of the environment, including the new state and reward; updates the deep neural network to optimize the prediction accuracy of the strategy and value function;

[0127] S61. The game agent performs an action in the environment , from the current state Transition to a new state , and receive the immediate reward associated with this action r ;

[0128] S62. Update the policy value network based on the observed rewards and new states to better predict the value of each action and the strategy for selecting actions. Value Function Aims to estimate from the state To begin with, the expected return that can be obtained by adopting the current strategy, the update of the strategy value network depends on the following formula:

[0129]

[0130] in, is the value estimate of the current state, is the reward received, is the value estimate of the new state, is the discount factor indicating the present value of future rewards, is the learning rate, which determines how quickly new information overwrites old information.

[0131] S63. If the policy gradient method is used for policy update, the goal is to optimize the probability distribution of the policy network's direct output actions to maximize the total reward. Update the parameters of the policy network , using the following gradient ascent formula:

[0132]

[0133] in It is in strategy Next, in the state Select Action The probability of is the gradient of the policy, indicating how to adjust the parameters To add selected actions probability.

[0134] S64, game agents use experience replay mechanism to improve learning efficiency and stability. Stored in a dataset, the learning process can randomly sample from this dataset to update the network, reducing the problems caused by correlation between samples.

[0135] In one embodiment, the above S7 includes: repeating S3 to S6 until a predetermined training termination condition is met.

[0136] When the training termination condition does not reach the number of learning rounds set before training or the preset performance target, or further learning can still significantly improve the performance of the agent, the training continues and steps S3 to S6 are repeated; when the termination condition is reached, the operation is completed and the optimal strategy is obtained.

[0137] Figure 1 Any process or method description in the flowchart described in or otherwise described herein can be understood as representing a module, fragment or portion of a code including one or more executable instructions for implementing the steps of a custom logic function or process, and the scope of the preferred embodiment of the present invention includes other implementations, in which the functions may be performed in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by a person skilled in the art of the present invention. The logic and / or steps represented in the flowchart or otherwise described herein illustrate the possible implementation architecture, functions and operations of the apparatus and methods according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It is also noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions. For example, it can be considered as an ordered list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or used in conjunction with these instruction execution systems, devices or apparatuses.

[0138] As far as the tenth embodiment of this specification is concerned, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use in an instruction execution system, device or equipment or in combination with these instruction execution systems, devices or equipment. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) with one or N wirings, a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting or processing in other suitable ways as necessary, and then stored in a computer memory. It should be understood that the various parts of the present invention can be implemented in hardware, software, firmware or a combination thereof. In the above embodiment, N steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0139] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined or coupled in a variety of ways, even if such combinations or couplings are not explicitly described in the present disclosure. In particular, the various embodiments and of the present disclosure may be combined in a variety of ways without departing from the spirit and teachings of the present disclosure. All of these combinations fall within the scope of the present disclosure.

[0140] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concept is known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A method for designing a game agent based on deep reinforcement learning, characterized in that: The method comprises the following steps: S1. Define and abstract the game environment and tasks; game agent, used to maximize its long-term reward by learning strategies in the game environment; S2. Initialize the game environment and tasks and construct a deep neural network as a strategy value network, wherein the strategy value network includes state values ​​and different strategies under a certain state; S3. By interacting with the game environment described in S1, the game agent collects a sequence of states, actions, and rewards as game data, uses the collected game data to train the strategy value network described in S2, uses strategy iteration to implement deep reinforcement learning training, and evaluates the state value function; S4, integrating the Monte Carlo tree search method to search and select actions based on the current game state; S5. Based on the results of the Monte Carlo tree search and the output of the policy value network, select the optimal action to execute; S6, the game agent executes the selected action and determines the feedback from the game environment; updates the deep neural network to optimize the prediction accuracy of the policy and value function; S7, repeat S3 to S6 until the predetermined training termination condition is met, and the game agent design is completed; S5 includes: S51, select the optimal action, adjust the selection strategy in combination with the strategy output provided by the deep neural network, and the deep neural network outputs the probability of each action; adjust the probability of each action by the number of visits of the Monte Carlo tree search method, that is, weight the action probability output by the network by the logarithm of the number of visits; S52. The game agent selects an action based on the adjusted action probability and selects the action with the highest probability as the final decision.

2. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: The use of policy iteration in S3 to implement deep reinforcement learning training includes two sub-steps: policy evaluation and policy improvement.

3. The method for designing a game agent based on deep reinforcement learning according to claim 2, characterized in that: The two sub-steps of strategy evaluation and strategy improvement include the following steps: S21. In the strategy evaluation phase, determine the state value function or the state and action value function under the current strategy; S22. In the strategy improvement phase, a new strategy is generated based on the state value function.

4. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: The search and selection of actions based on the current game state in S4 includes the following steps: S41. Starting from the root node, recursively select child nodes until reaching a leaf node; each selection is based on a specific strategy. S42, after reaching the leaf node, if this node is not the end point of the game, generate one or at least one successor child node; S43, starting from the newly expanded node, simulating the game to the end in a random manner to determine the outcome of the game; S44. Update the statistical information of all nodes on the path from the current leaf node to the root node according to the simulation results, repeat the above steps until the preset search time or search number limit is reached, and select the move with the highest winning rate or the best specific value.

5. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: The prediction accuracy of the optimization strategy and value function in S6 is achieved by the energy body using the experience replay mechanism.

6. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: S6 includes: S61, the game agent performs an action in the game environment, moves from the current state to the new state, and receives an immediate reward associated with the action; S62. Update the policy value network based on the observed rewards and new states; S63. If the policy gradient method is used for policy update, the policy value network is optimized to directly output the probability distribution of the action.

7. A game agent design system based on deep reinforcement learning, characterized in that: The system comprises: An input module, used to define and abstract the game environment and tasks, and a game agent, used to maximize its long-term reward by learning strategies in the game environment; Constructing a strategy value network module, which is used to initialize the game environment and tasks and construct a deep neural network as a strategy value network, wherein the strategy value network includes state values ​​and different strategies under a certain state; A state value function evaluation module is used to collect the sequence of states, actions and rewards as game data by the game agent through interaction with the game environment described in the definition module, use the collected game data to train the strategy value network described in the strategy value network construction module, use strategy iteration to implement deep reinforcement learning training, and evaluate the state value function; The integrated Monte Carlo Tree Search module is used to integrate the Monte Carlo Tree Search method to search and select actions based on the current game state; The execution module is used to select the optimal action to execute based on the results of the Monte Carlo tree search and the output of the policy value network; A prediction module that optimizes the policy and value functions for the game agent to perform the selected actions and determine the feedback from the game environment; updates the deep neural network to optimize the prediction accuracy of the policy and value functions; The output module is used to repeatedly evaluate the state value function model to optimize the prediction module of the strategy and value function until the predetermined training termination condition is met to complete the design of the game agent; The execution module includes: Select the best action and adjust the selection strategy based on the strategy output provided by the deep neural network. The deep neural network outputs the probability of each action. The probability of each action is adjusted by the number of visits of the Monte Carlo tree search method, that is, the action probability output by the network is weighted by the logarithm of the number of visits. The game agent selects actions based on the adjusted action probabilities and chooses the action with the highest probability as the final decision.

8. A computer device comprising a memory and a processor, characterized in that A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Chess game method based on opponent modeling and Monte Carlo reinforcement learning

    CN116128060A