Tree search device, tree search method, and tree search program
By implementing a backpropagation calculation unit and Bayesian estimation for action value calculation in Monte Carlo tree search, the inefficiencies in initial stage searches are addressed, enhancing the overall search efficiency.
Patent Information
- Application Number
- PCT/JP2023/042415
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-05
AI Technical Summary
Monte Carlo tree search experiences low search efficiency in the initial stages due to undervalued action values estimated by backpropagating reward values from long random rollouts, which often yield smaller rewards than optimal strategies.
Incorporating a backpropagation calculation unit to accumulate reward values, a storage unit to store these values, a posterior distribution calculation unit for Bayesian estimation of action values, and an action value estimation unit to calculate estimated action values using posterior distributions.
This approach improves the efficiency of Monte Carlo tree search by accurately estimating action values even in the initial stages, leading to more effective search strategies.
Smart Images

Figure JP2023042415_05062025_PF_FP_ABST
Abstract
Description
Tree search device, tree search method, and tree search program
[0001] The present invention relates to a tree search device, a tree search method, and a tree search program.
[0002] In recent years, search algorithms have been actively researched for applications such as action selection for game AI in competitive games and searching for optimal actions for robots. Among these search algorithms, Monte Carlo tree search is a method with a wide range of applications to practical problems.
[0003] Levente Kocsis and Csaba Szepesvari, "Bandit based Monte-Carlo Planning", European conference on machine learning. Berlin, Heidelberg: Springer Berlin Heidelberg, 2006.Mern, John, et al., "Bayesian Optimized Monte Carlo Planning", Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 35. No. 13. 2021.
[0004] Monte Carlo tree search is an effective method in environments where high-speed search is possible, but in environments where search can only be performed slowly, the problem is how to improve search efficiency.
[0005] Here, in Monte Carlo tree search, if the evaluation value (action value) when an action is taken from a certain state cannot be directly obtained, the action value is estimated, for example, by performing a random rollout and backpropagating the reward value obtained in the final state.
[0006] In this case, in the early stages when the Monte Carlo tree search is not yet fully performed, the length of the random rollout becomes long, and the reward value obtained by the random rollout is often smaller than the reward value obtained by adopting the optimal strategy. As a result, in the early stages of the tree search, the action value estimated by backpropagating the reward value is underestimated compared to the true action value, resulting in low search efficiency. Therefore, an object of the present invention is to solve the above-mentioned problem and improve the efficiency of the Monte Carlo tree search.
[0007] In order to solve the above-mentioned problems, the present invention is characterized by comprising a backpropagation calculation unit that, when an action value, which is an evaluation value when a certain action is taken from a certain state, rolls out the Monte Carlo tree search and calculates the backpropagation amount of the reward value obtained in the final state; a memory unit that stores the set of backpropagation amounts calculated so far in the Monte Carlo tree search; a posterior distribution calculation unit that calculates the distribution of the posterior probability of the action value by Bayesian estimation using the prior distribution of the action value and the likelihood of the set of backpropagation amounts for the action value; and an action value estimation unit that calculates an estimate of the action value using the distribution of the posterior probability of the action value.
[0008] According to the present invention, the efficiency of the Monte Carlo tree search can be improved.
[0009] FIG. 1 is a diagram for explaining states, actions, and policies. FIG. 2 is a diagram for explaining a tree in Monte Carlo tree search. FIG. 3 is a diagram for explaining the flow of Monte Carlo tree search. FIG. 4 is a diagram showing an example of the configuration of a Monte Carlo tree search device. FIG. 5 is a diagram showing an example of the configuration of an action value calculation unit. FIG. 6 is a flowchart showing an example of the processing procedure executed by the Monte Carlo tree search device. FIG. 7 is a diagram showing a computer that executes a Monte Carlo tree search program.
[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, a description will be given of an embodiment of the present invention with reference to the drawings, but the present invention is not limited to the embodiment.
[0011] First, an overview of the tree search performed by the Monte Carlo tree search device (tree search device) of this embodiment will be described. In the following description, state s indicates an environmental state encountered in the environment to be searched. Action a indicates an action to be executed in a certain state. Policy p(a|s) is a probability value indicating the likelihood that each action a should be executed in state s. Figure 1 shows an example of state s, action a, and policy p in tic-tac-toe.
[0012] A tree to be searched in a Monte Carlo tree search device is made up of nodes and edges, for example, as shown in Fig. 2. A node indicates a state, and an edge indicates an action.
[0013] Next, the procedure for the Monte Carlo tree search will be explained with reference to Fig. 3. The Monte Carlo tree search is performed by repeating the following steps 1 to 4 a predetermined number of times.
[0014] 1. Selection: The computer advances through states from the root state of the tree to the leaf nodes of the tree, selecting actions according to the acquisition function. The acquisition function can be, for example, UCB (Upper Confidence Bounds). The computer calculates an evaluation value for each action by substituting the action value Q(s, a) of each action and the number of times the action is selected into the acquisition function, and selects an action based on the calculated evaluation value.
[0015] 2. Expansion: If a certain criterion is met, such as the number of times an action is selected being greater than or equal to a predetermined value, the computer adds a new leaf node to the tip of the leaf node and transitions to that state.
[0016] 3. Simulation: The computer randomly takes actions from the leaf nodes of the tree, progresses to an evaluable state (especially the final state), and obtains a reward value. This process of the computer randomly taking actions and progressing through the states to the final state is called random rollout / random playout.
[0017] 4. Backpropagation: The computer propagates the reward value obtained in step 3 backwards and updates the action value Q(s, a) of the action assigned to each edge that has been followed so far. The action value Q(s, a) of an action is, for example, the expected total reward (discounted reward) that can be obtained by taking that action.
[0018] [Overview] Next, an overview of the Monte Carlo tree search device of this embodiment will be described. The Monte Carlo tree search device uses Bayesian estimation to estimate the action value Q(s, a) assigned to the edge (action) of the tree from the past backpropagation amount r' in the tree search. The Monte Carlo tree search device then proceeds with the tree search based on the estimated action value Q(s, a).
[0019] For example, the Monte Carlo tree search device calculates the posterior probability of the action value Q(s, a) using the following formula (1): i=1 N is the set of backpropagation values calculated so far in the Monte Carlo tree search. For example, if N=21, it is the set of backpropagation values r' for the past 21 iterations. p({r'} i=1 N |Q(s,a)) is the function of {r'} given Q(s,a). i=1 N is the probability (likelihood) that Q(s,a) occurs. p(Q(s,a)) is the probability (prior distribution) that Q(s,a) appears.
[0020]
[0021] The Monte Carlo tree search device then calculates an estimate of the action value Q(s, a) from the calculated posterior probability of the action value Q(s, a). For example, the Monte Carlo tree search device estimates the action value Q(s, a) by taking the average of the posterior probabilities of the action value Q(s, a) (point estimation). Note that the Monte Carlo tree search device may also estimate the action value Q(s, a) by taking the posterior probability of the action value Q(s, a) itself (the distribution of the posterior probability) (distribution estimation).
[0022] According to the Monte Carlo tree search device, the action value Q(s, a) is estimated using a prior distribution (prior knowledge) of the value of the action value Q(s, a), so that the action value Q(s, a) is less likely to be underestimated even in the early stages of the tree search, thereby improving the efficiency of the Monte Carlo tree search.
[0023] Furthermore, according to the above Monte Carlo tree search device, when the prior distribution of the action value Q(s, a) and the action value Q(s, a) are given, {r'} i=1 N The designer can set the probability (likelihood) of occurrence, whether to estimate the action value Q(s, a) by point estimation or distribution estimation, etc. Therefore, the designer can flexibly design according to the situation in the Monte Carlo tree search.
[0024] [Configuration Example] Next, a configuration example of the Monte Carlo tree search device 10 will be described with reference to Fig. 4. Note that each unit shown in Fig. 4 is realized by a CPU (Central Processing Unit) in the Monte Carlo tree search device 10 executing a program in a storage device.
[0025] When the Monte Carlo tree search device 10 receives an input of a state s0, it constructs a tree that outputs a policy p for the state s0.
[0026] The Monte Carlo tree search device 10 performs Monte Carlo tree search using, for example, a tree configuration calculation unit 11, an acquisition function calculation unit 12, a backpropagation calculation unit 13, an action value calculation unit 14, and an update unit 15 to construct a tree.
[0027] The tree configuration calculation unit 11 calculates the tree configuration. The acquisition function calculation unit 12 calculates the evaluation value of each action (edge) of the tree using the acquisition function. When the Monte Carlo tree search fails to obtain an evaluation value (action value) for taking a certain action from a certain state, the backpropagation calculation unit 13 rolls out the Monte Carlo tree search and calculates the backpropagation amount of the reward value obtained in the final state. The backpropagation calculation unit 13 stores the calculated backpropagation amount in a memory unit (backpropagation amount memory unit 141 shown in FIG. 5 ) within the Monte Carlo tree search device 10.
[0028] The action value calculation unit 14 estimates the action value Q(s, a) based on the set of backpropagation amounts r' for action a from state s calculated by the backpropagation calculation unit 13. Details of the action value calculation unit 14 will be described later using FIG. 5. The update unit 15 updates the action value Q(s, a) using the action value Q(s, a) estimated by the action value calculation unit 14. Thereafter, the Monte Carlo tree search device 10 proceeds with the tree search using the updated action value Q(s, a).
[0029] Next, the action value calculation unit 14 will be described in detail with reference to Fig. 5. The action value calculation unit 14 includes, for example, a backpropagation amount storage unit 141, a likelihood calculation unit 142, a prior distribution calculation unit 143, a posterior distribution calculation unit 144, and an estimation unit 145, as shown in Fig. 5.
[0030] The backpropagation amount storage unit 141 is a memory that stores the backpropagation amount r' of the reward value obtained by the random rollout and backpropagation of the tree search so far. The backpropagation amount storage unit 141 stores a set of backpropagation amounts r' calculated so far for the action a from the state s described in the target tree of the Monte Carlo tree search.
[0031] The likelihood calculation unit 142 calculates the set of backpropagation quantities {r'} i=1 N Likelihood p({r'} i=1 N |Q(s, a)) The prior distribution calculation unit 143 calculates the prior distribution p(Q(s, a)) of the action value Q(s, a).
[0032] The posterior distribution calculation unit 144 calculates p({r'} i=1 N |Q(s, a)) and p(Q(s, a)) calculated by the prior distribution calculation unit 143. For example, the posterior distribution calculation unit 144 calculates the posterior probability distribution of the action value Q(s, a) using the above-mentioned formula (1).
[0033] The estimation unit 145 calculates an estimate of the action value Q(s, a) from the distribution of the posterior probability of the action value Q(s, a) calculated by the posterior distribution calculation unit 144. For example, the estimation unit 145 estimates the action value Q(s, a) by taking the average value of the posterior probability of the action value Q(s, a) (point estimation). Alternatively, the estimation unit 145 may use the posterior probability of the action value Q(s, a) itself (the distribution of the posterior probability) as the estimate of the action value Q(s, a) (distribution estimation).
[0034] [Example of Processing Procedure] Next, an example of the processing procedure executed by the Monte Carlo tree search device 10 will be described with reference to FIG. 6. In the process of tree search executed by the Monte Carlo tree search device 10, every time the backpropagation calculation unit 13 calculates a backpropagation amount r' (S1), the backpropagation amount r' is stored in the backpropagation amount storage unit 141. By the above processing, a set of backpropagation amounts r' ({r'} i=1 N ) is stored in the backpropagation amount storage unit 141 (S2).
[0035] Next, when the action value Q(s, a) is given, the likelihood calculation unit 142 calculates the set of backpropagation amounts r' ({r'} i=1 N ) occurs (likelihood: p({r'} i=1 N |Q(s, a)) (S3). The prior distribution calculation unit 143 also calculates the prior distribution p(Q(s, a)) of the action value Q(s, a) (S4).
[0036] Next, the posterior distribution calculation unit 144 calculates the likelihood (p({r'} i=1 N The distribution of the posterior probability of the action value Q(s, a) is calculated by Bayesian estimation using the prior distribution (p(Q(s, a))) of the action value Q(s, a) calculated in S4 and the prior distribution (p(Q(s, a))) of the action value Q(s, a) calculated in S5).
[0037] Next, the estimation unit 145 calculates an estimate of the action value Q(s, a) from the distribution of the posterior probabilities calculated in S5 (S6). After that, the update unit 15 updates the value of the action value Q(s, a) of the tree using the estimate of the action value Q(s, a) calculated in S6 (S7). Thereafter, the Monte Carlo tree search device 10 proceeds with the tree search using the updated action value Q(s, a).
[0038] By having the Monte Carlo tree search device 10 execute the above process, it is possible to improve the efficiency of the Monte Carlo tree search.
[0039] [System Configuration, etc.] The components of each unit shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU and a program executed by the CPU, or can be realized as hardware using wired logic.
[0040] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method.In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified.
[0041] [Program] The Monte Carlo tree search device 10 can be implemented by installing a program (Monte Carlo tree search program) as package software or online software on a desired computer. For example, by executing the program on an information processing device, the information processing device can function as the Monte Carlo tree search device 10. The information processing device referred to here includes mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone Systems), as well as terminals such as PDAs (Personal Digital Assistants).
[0042] 7 is a diagram showing an example of a computer that executes a Monte Carlo tree search program. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0043] The memory 1010 includes a read-only memory (ROM) 1011 and a random access memory (RAM) 1012. The ROM 1011 stores a boot program such as a basic input / output system (BIOS). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0044] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the programs that define the processes executed by the Monte Carlo tree search device 10 are implemented as program modules 1093 in which computer-executable code is written. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, the program modules 1093 for executing processes similar to those of the functional configuration of the Monte Carlo tree search device 10 are stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0045] Data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. The CPU 1020 then reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary and executes them.
[0046] The program module 1093 and program data 1094 may not necessarily be stored in the hard disk drive 1090, but may also be stored in a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.
[0047] REFERENCE SIGNS LIST 10 Monte Carlo tree search device 11 Tree configuration calculation unit 12 Acquisition function calculation unit 13 Backpropagation calculation unit 14 Action value calculation unit 15 Update unit 141 Backpropagation amount storage unit 142 Likelihood calculation unit 143 Prior distribution calculation unit 144 Posterior distribution calculation unit 145 Estimation unit
Claims
1. In Monte Carlo tree search, when the action value, which is the evaluation value when taking a certain action from a certain state, cannot be obtained, a backpropagation calculation unit that performs rollout of the Monte Carlo tree search and calculates the backpropagation amount of the reward value obtained in the final state; a storage unit that stores the set of the backpropagation amounts calculated so far in the Monte Carlo tree search; a posterior distribution calculation unit that calculates the distribution of the posterior probability of the action value by Bayesian estimation using the prior distribution of the action value and the likelihood of the set of the backpropagation amounts for the action value; and an action value estimation unit that calculates an estimated value of the action value using the distribution of the posterior probability of the action value. A tree search device characterized by comprising these components.
2. The tree search device according to claim 1, characterized in that the Monte Carlo tree search is advanced based on the estimated action value.
3. A tree search method executed by a tree search device, the method including: in Monte Carlo tree search, when the action value, which is the evaluation value when taking a certain action from a certain state, cannot be obtained, performing rollout of the Monte Carlo tree search and calculating the backpropagation amount of the reward value obtained in the final state; storing the set of the backpropagation amounts calculated so far in the Monte Carlo tree search in a storage unit; calculating the distribution of the posterior probability of the action value by Bayesian estimation using the prior distribution of the action value and the likelihood of the set of the backpropagation amounts for the action value; and calculating an estimated value of the action value using the distribution of the posterior probability of the action value. A tree search method characterized by including these steps.
4. A tree search program for causing a computer to execute: in Monte Carlo tree search, when the action value, which is the evaluation value when taking a certain action from a certain state, cannot be obtained, performing rollout of the Monte Carlo tree search and calculating the backpropagation amount of the reward value obtained in the final state; storing the set of the backpropagation amounts calculated so far in the Monte Carlo tree search in a storage unit; calculating the distribution of the posterior probability of the action value by Bayesian estimation using the prior distribution of the action value and the likelihood of the set of the backpropagation amounts for the action value; and calculating an estimated value of the action value using the distribution of the posterior probability of the action value.