R-Tree construction method for global structure perception based on reinforcement learning
Through the global structure perception method based on reinforcement learning, the training agent automatically selects child nodes during the R-Tree construction process, solving the problem of poor query performance in dynamic data update scenarios, and achieving efficient R-Tree construction and query performance improvement.
Patent Information
- Application Number
- CN202510189196.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
When facing dynamic data update scenarios, the existing R-Tree construction method is difficult to build a global optimal tree structure, resulting in poor query performance and relying on fixed rules and parameters determined manually, and cannot adapt to different data distributions.
A global structure perception R-Tree construction method based on reinforcement learning is designed. Through Markov decision-making process, Actor-Critic framework and proximal strategy optimization algorithm, combined with the self-gaming mechanism, reinforcement learning agents are trained to automatically select children's nodes during the data insertion process to build R-Tree.
It significantly improves the overall query performance of R-Tree, can maintain efficient query performance in dynamic data update scenarios, reduces manual intervention, and reduces the cost of system maintenance and optimization.
Smart Images

Figure CN120123338A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of spatial data processing, and particularly relates to a method for constructing an R-Tree with global structure awareness based on reinforcement learning. Background Art
[0002] With the mature development of spatial information acquisition technology, the scale of spatial data has expanded rapidly. In order to improve the retrieval efficiency of spatial data, many spatial indexing methods have been proposed. Among various types of spatial indexing algorithms, the spatial indexing algorithms represented by R-Tree have always been the main direction and important research content of spatial indexing development due to their good flexibility and compatibility. They are widely used in the underlying database spatial indexing structure of major GIS systems to help manage and quickly retrieve massive spatial data. The query performance provided by R-Tree depends to a large extent on the finally constructed tree structure. In order to improve the query performance of R-Tree, different types of variants have been developed for different application scenarios, mostly algorithms implemented by fixed heuristic rules. With the prosperous development of machine learning technology and its successful applications in various fields, it has gradually begun to be combined to improve the performance of the index.
[0003] The construction methods of R-Tree are divided into constructing by inserting data one by one and constructing by batch loading data. The construction of R-Tree using reinforcement learning also starts from these two tree construction methods. Regarding the construction of R-Tree by batch loading data, whether it is a traditional variant or a construction based on reinforcement learning, it is not suitable for dynamically growing data sets. It is mostly a one-time construction and cannot be applied to application scenarios with real-time data updates. Under the method of constructing R-Tree by inserting data one by one, traditional variants mostly adopt fixed heuristic rules and various parameters. These fixed rules and parameters rely on manual determination and do not have the best performance in all cases. That is, more suitable variants may need to be selected for data with different distributions. In addition, its algorithm operation has the property of local optimality, that is, the algorithm adopted each time when selecting a subtree for insertion depends on local metrics related to the layer where the node is located. The accumulation of local optimal operations will not make the overall tree structure construct towards the global optimal direction, but will cause the problem of global sub-optimality. Similarly, in the existing incremental R-Tree construction based on reinforcement learning, the basic framework has focused on using the action selection of the RL agent to replace the fixed rules of traditional variants to make decisions, but there is still a problem of local optimal selection, that is, traditional metrics are still used in the state design, and the vision is limited to the layer where the node is located. Due to the local optimality of the selection algorithm or the state on which the RL agent's decision depends in the existing incremental R-Tree construction method, the global tree structure cannot be well considered. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides an R-Tree construction method based on global structure perception of reinforcement learning, including:
[0005] Design a Markov decision process for R-Tree construction, where the Markov decision process for R-Tree construction includes state design, action design, and reward design;
[0006] Apply a reinforcement learning agent algorithm framework, and select to construct a reinforcement learning agent based on the Actor-Critic framework and the proximal policy optimization algorithm;
[0007] Adopt a self-play mechanism to train the reinforcement learning agent for incremental R-Tree construction;
[0008] During the process of inserting data one by one through the trained reinforcement learning agent, replace the traditional variant's locally optimal child selection algorithm to automatically construct the R-Tree.
[0009] Preferably, the state design includes:
[0010] Extract information associated with each child subtree of the current node, and construct a state vector for use as a reference basis when the agent makes decisions according to the information associated with each child subtree of the current node;
[0011] Among them, the information associated with each child subtree of the current node includes subtree data proximity, layer area ratio, and layer perimeter ratio.
[0012] Preferably, the action design includes: selecting the position of the child node and performing a data insertion operation; assuming the node capacity is M, the output action space is {insert(1), inert(2),..., insert(M)}.
[0013] Preferably, the reward design includes:
[0014] Taking the degree of improvement in query performance after inserting the test data as the reward and distributing it to the actions made by all agents during the data insertion process; among them, the degree of improvement in query performance is represented by the difference from the query performance tested in the past.
[0015] Preferably, the process of taking the degree of improvement in query performance after inserting the test data as the reward and distributing it to the actions made by all agents during the data insertion process includes
[0016] Use a self-play mechanism to synchronously construct a competitor identical to the GSAR-Tree. After data insertion, perform exactly the same random range queries on the two tree structures, calculate the reward R. If the reward R is positive, it indicates that the constructed GSAR-Tree has better query efficiency; otherwise, it is inferior to the competitor.
[0017] Secondly, the formula expression of the reward R is:
[0018]
[0019] In the formula, R is the number of node accesses after performing a random query.
[0020] Preferably, applying the reinforcement learning agent algorithm framework, the process of selecting a reinforcement learning agent based on the Actor-Critic framework and the proximal policy optimization algorithm includes:
[0021] Use the Actor-Critic algorithm to train the policy network Actor and the value network Critic simultaneously to learn the R-Tree construction process; among them, the training policy network Actor is used to fit the policy function π θ (a|s) to directly give the action a according to the current state s; the value network Critic is used to fit the state value function V θ (s) to evaluate the quality of the current policy.
[0022] Optimize the objective function directly through the gradient policy algorithm of reinforcement learning to update the parameters of the policy network, and then optimize the loss of the objective function to update the Actor network. Calculate the mean square error loss between the predicted value of the state and the TD error to update the Critic network.
[0023] Preferably, using the Actor-Critic algorithm to train the policy network Actor and the value network Critic simultaneously to learn the R-Tree construction process includes:
[0024] Calculate the TD error of each time step according to the collected interaction data through the Actor-Critic algorithm; among them, the TD error is the actual return value r at the current time step t +γV(s t+1 ), and the difference between the predicted value V(s t ), where r t is the immediate reward at the current time step, γ is the discount factor, representing the weight of future rewards;
[0025] δ t =r t +γV(s t+1 )-V(s t )
[0026] Calculate the Generalized Advantage Estimation (GAE) based on the TD error, and the formula expression is:
[0027] A t = δ t +(γλ)δ t+1 +(γλ) 2 δ t+2 +...
[0028] Preferably, the process of directly updating the parameters of the policy network by optimizing the objective function through the gradient policy algorithm of reinforcement learning includes:
[0029] Control the amplitude of policy update by clipping the objective function, and the final objective function expression is:
[0030] J(θ) = ∑min(r t (θ)A t , clip(r t (θ), 1 - ∈, 1 + ∈)A t )
[0031]
[0032] where r t (θ) is the ratio of the current policy to the old policy.
[0033] Preferably, the process of training the reinforcement learning agent to perform incremental R-Tree construction using the self-play mechanism includes:
[0034] Initialize a GSAR-Tree and an identical copy as a competitor. Insert all the data in the dataset one by one completely to construct an R-Tree once as a round. In each round, the GSAR-Tree and the competitor are executed synchronously. After inserting the data, perform multiple identical random range queries on the GSAR-Tree and the competitor respectively, calculate the reward R. If the reward is positive, synchronize the GSAR-Tree to the competitor.
[0035] After multiple rounds of testing, construct the R-Tree using the trained reinforcement learning agent. Replace the traditional algorithm for selecting subtrees with the trained reinforcement learning agent. Starting from the root node from top to bottom, extract the current node state and input it to the agent, and execute the output action.
[0036] Preferably, the process of synchronously executing the GSAR-Tree and the competitor in each round when inserting all the data in the dataset one by one completely to construct an R-Tree once as a round includes:
[0037] Taking all the data in the dataset and inserting them completely one by one to construct an R-Tree once is regarded as one round. In each round, first reset the GSAR-Tree and the competitor. Each time a data is taken out and starts from the root node. If a leaf node is encountered, the data is directly inserted. Otherwise, the state vector corresponding to the node is extracted as the input of the agent, the given action is obtained and executed, a child is selected for insertion, and then the iterative operation continues layer by layer until the data is inserted into the leaf node;
[0038] The GSAR-Tree and the competitor are executed synchronously each time. When inserting data to construct a partial tree structure and storing state-action pairs through the agent, at this time, the same random range queries are executed multiple times on the GSAR-Tree and the competitor respectively, and then the reward R is calculated. The same reward R is assigned to the state-action pairs. If the calculated reward is positive, it indicates that the existing structure of the GSAR-Tree has higher query efficiency, and the neural network component in the agent has learned a better strategy. Then the GSAR-Tree is synchronized to the competitor, so that the GSAR-Tree can continue to compete with the competitor in the tree construction performance during the training process.
[0039] Compared with the prior art, the present invention has the following advantages and technical effects:
[0040] By designing unique state vectors, including elements such as SDP, LAR, and LPR, the present invention incorporates the subtree structure information of all children under the current node into the decision-making vision of the intelligent agent, enabling the reinforcement learning intelligent agent to optimize the construction process of the R-Tree from a global perspective, avoiding the disadvantages of local optimal operations, and thus significantly improving the overall query performance of the R-Tree.
[0041] The self-play mechanism and reinforcement learning algorithm framework (such as Actor-Critic and PPO) adopted by the present invention enable the intelligent agent to continuously learn and optimize strategies during the dynamic data insertion process. By continuously comparing with the historical optimal version (Competitor), the intelligent agent can dynamically adjust decisions to ensure that the R-Tree always maintains high query performance in real-time data update scenarios.
[0042] The PPO algorithm and self-play mechanism adopted by the present invention not only improve the stability of the training process, but also significantly enhance the training efficiency. The PPO algorithm directly updates the parameters of the policy network by optimizing the objective function, avoiding instability and overfitting problems during the training process. The self-play mechanism provides continuous optimization motivation for the agent by comparing with the historical optimal version, enabling the trained agent to better adapt to different data distributions and query patterns, and thus showing better performance in practical applications.
[0043] The present invention has high flexibility and scalability in state design and algorithm selection. Elements such as SDP, LAR, and LPR in the state vector can be adjusted or extended according to specific application scenarios. For example, other metrics related to the subtree structure (such as the overlapping area between layers, the layer-to-layer ratio of some layers, etc.) can be introduced. In addition, variants of PPO (such as PPO-Clip, PPO-Penalty) or other reinforcement learning algorithms (such as A2C, A3C, SAC, etc.) can also be adopted in algorithm selection. Through these variants and extension schemes, the performance and adaptability of R-Tree construction can be further improved.
[0044] Through experimental verification, the R-Tree (GSAR-Tree) constructed by the present invention is significantly superior to traditional R-Tree variants and existing reinforcement learning-based construction methods in terms of query efficiency. Especially in large-scale dynamic data scenarios, GSAR-Tree can better organize and optimize the intermediate-level structure, reduce multi-path problems and dead zone areas during the query process, thus greatly improving the efficiency of spatial data retrieval. This technical effect is of great significance for practical applications in fields such as geographic information systems (GIS) and spatial databases.
[0045] The present invention reduces the need for manual intervention through the automated decision-making of the reinforcement learning agent, making the construction process of the R-Tree more intelligent and automated, and reducing the costs of system maintenance and optimization.
[0046] In summary, through the state design with global structure awareness, the efficient reinforcement learning algorithm framework, and the self-play training mechanism, the present invention significantly improves the construction performance and query efficiency of the R-Tree. At the same time, it has good flexibility and scalability, can meet the real-time update requirements in dynamic data scenarios, and provides an innovative technical solution for the fields of geographic information systems and spatial data management. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0048] Figure 1Schematic flowchart of the method according to an embodiment of the present invention;
[0049] Figure 2 Schematic diagram of the state view according to an embodiment of the present invention. Detailed implementation manners
[0050] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0051] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0052] In the incremental construction of the R-Tree, most existing traditional R-Tree variants adopt fixed heuristic rules and various parameters, sometimes relying on manual selection and determination, unable to cope with data with different distributions, and the tree construction operation has the property of local optimality. For the existing incremental construction of the R-Tree based on reinforcement learning, its state design still refers to traditional local metrics, unable to estimate the global structure of the tree, and the tree construction process design cannot fully explore better R-Tree structures.
[0053] As Figure 1-2 shown, this embodiment provides a method for constructing a globally structure-aware R-Tree based on reinforcement learning, including:
[0054] Design a Markov decision process for R-Tree construction, where the Markov decision process for R-Tree construction includes state design, action design, and reward design;
[0055] Apply a reinforcement learning agent algorithm framework, and select to construct a reinforcement learning agent based on the Actor-Critic framework and the proximal policy optimization algorithm;
[0056] Adopt a self-play mechanism to train the reinforcement learning agent for incremental construction of the R-Tree;
[0057] During the process of inserting data one by one through the trained reinforcement learning agent, replace the locally optimal child selection algorithm of the traditional variant to automatically construct the R-Tree.
[0058] Furthermore, this embodiment proposes a model GSAR-Tree (Global Structure-Aware R-Tree) for incrementally constructing an R-Tree based on reinforcement learning technology (RL), which mainly includes four parts: the design of the MDP (Markov Decision Process) for R-Tree construction, the application combination of the RL agent algorithm framework, the training process design, and the application test.
[0059] Among them, the MDP design mainly includes state design, action design, and reward design. That is, when the RL agent constructs an R-Tree, the tree environment state serves as its input, the child selection action serves as its output, and the reward design feedback from the tree environment to the RL agent. Among them, the state design abandons the traditional locally optimal metrics and uses the information associated with each child subtree of the current node to represent the state of the tree structure. The state composition includes elements such as SDP (Subtree-Data Proximity), LAR (Layer Area Ratio), and LPR (Layer Perimeter ratio), which are used as reference bases for the agent's decision-making, and can avoid the disadvantages of locally optimal operations. The entire MDP design completely replaces the traditional heuristic child selection algorithm.
[0060] More specifically, the state design in the MDP design of the R-Tree includes:
[0061] Extract the information associated with each child subtree of the current node, and construct a state vector for use as a reference basis for the agent's decision-making according to the information associated with each child subtree of the current node; wherein, the information associated with each child subtree of the current node includes subtree data proximity, layer area ratio, and layer perimeter ratio.
[0062] For constructing an R-Tree using an RL agent, it is natural to think of the environment faced by the agent, which is the tree structure itself. However, in most cases, when inserting one data or even multiple data into the leaves, the overall change of the R-Tree is not obvious, so it cannot be directly used to represent the environmental state.
[0063] The entire MDP design of this embodiment is as follows: The RL agent takes a state vector associated with the current node as input. The selection logic is the potential pattern learned by the internal neural network component during exploration, and the output is also the selected child position. However, the traditional algorithm only focuses on the node layer where it is located because the algorithm itself requires a fixed linear selection logic from input to output. The RL agent, combined with the function fitting ability of the neural network, can handle functional relationships beyond linearity. In fact, at each layer, the entire subtree structure with each node as the root node can be considered as the scope of information input for the agent. Therefore, there is an opportunity to take into account the existing tree structure and data distribution information during the process of inserting data one by one to build the R-tree, and avoid making locally optimal decisions in the selection of children of the current node. As Figure 2 shown, in the middle, the metrics used in traditional data insertion are limited to the information related to the children of the current node, and on the right, the state vector view of this embodiment is associated with the subtree information of all children of the current node.
[0064] The construction of the state vector mainly consists of three elements: When making a child selection at a node, it is hoped that the nodes and data distributions under the MBRs of each child can be directly seen, and these subtree-related information is associated with the data to be inserted to better support the decision-making of the agent. Therefore, the present invention proposes a metric, SDP (Subtree-Data Proximity), to express the distance between the data to be inserted and the subtree data of the node. That is, for each child of the node, calculate the sum of the distances between the MBRs of all nodes in its subtree and the center of the rectangular data to be inserted as part of the state vector. For the subtree T i with the i-th child of the current node as the root, the SDP calculation formula is as follows:
[0065]
[0066] where N i and O i respectively represent the node MBR and the data in the leaf of the subtree T i , center(rec) is the center coordinate of the rectangle, and d is the Euclidean distance between two points.
[0067] The incremental construction of the R-Tree actually organizes and optimizes the intermediate hierarchical structure of the tree. For each intermediate layer, it is always desired to have a smaller dead zone and overlapping area, that is, to increase the space utilization rate as much as possible and reduce the multi-path query problem. This embodiment proposes a measurement index related only to the subtree structure as one of the components of the state vector, LAR (Layer Area Ratio) and LPR (Layer Perimeter Ratio). Assuming that the subtree rooted at a node N has n layers, and node N is at the 0th layer, define f t (x), t ∈ {area, perimeter} as the sum of calculating the feature t of all nodes in the xth layer of the subtree. Then the calculation formulas for LAR and LPR are as follows:
[0068]
[0069] The occupancy rate OR (Occupancy Rate) of the children of a node can also be used as the basis for the agent's decision-making state. However, it is not necessary to consider the entire subtree. Just consider two layers downward. That is, for each child, calculate the average of its own occupancy rate and the occupancy rates of all its children, that is, the node occupancy rate OR considering two layers downward 2 .
[0070] When the agent faces the decision of choosing a child to insert at node N, for the ith child of node N, the state vector corresponding to this child is composed of (OR 2 , SDP, LAR, LPR) i . Finally, for node N, the state vectors of all children are arranged to form state N as the state vector of node N. For nodes that are not fully filled, the state values of the existing nodes are circularly filled into the vacancies.
[0071] The action design includes: for each subtree selection, the output of the traditional algorithm is the position of a child. Then, for the output of the agent model, we can also define it as the position of the selected child, and then perform the data insertion operation. Assuming that the node capacity is M, the output action space is {insert(1), inert(2),..., insert(M)}.
[0072] The reward design includes: after the agent makes a decision at a node in each layer and inserts it into the child nodes of the next layer, after such an action is taken, what kind of reward should the tree environment feedback to the agent? A simple idea is to obtain the improvement degree of its query efficiency as the action reward, that is, it is hoped that during the exploration and learning process of the agent, it can build the R-Tree in the direction of continuously improving the query efficiency.
[0073] In this embodiment, the improvement degree of the query performance is tested once after a small batch of data is inserted, and it is used as a reward to be distributed to all the actions made by the agents during the insertion process of this small batch of data. The improvement of the query performance is represented by the difference in the query performance tested with the past self.
[0074] First, use the self-play mechanism to synchronously construct a competitor that is the same as the GSAR-Tree. After a small batch of data is inserted, perform exactly the same random range query on the two tree structures. The reward calculation formula is
[0075]
[0076] In the formula, R is the number of node accesses after performing the random query. If R is regular, it indicates that the constructed GSAR-Tree has better query efficiency, otherwise it is inferior to the competitor. The competitor is equivalent to the GSAR-Tree with the optimal model currently. The GSAR-tree will also regularly determine whether to update the competitor or fall back to the competitor according to R.
[0077] Furthermore, applying the reinforcement learning agent algorithm framework, the process of selecting to construct a reinforcement learning agent based on the Actor-Critic framework and the proximal policy optimization algorithm includes:
[0078] Use the Actor-Critic algorithm to train the policy network Actor and the value network Critic simultaneously to learn the R-Tree construction process; among them, the trained policy network Actor is used to fit the policy function π θ (a|s) to directly give the action a according to the current state s; the value network Critic is used to fit the state value function V θ (s) to evaluate the quality of the current policy;
[0079] Update the parameters of the policy network directly by optimizing the objective function through the gradient policy algorithm of reinforcement learning, and then optimize the loss of the objective function to update the Actor network. Calculate the mean square error loss between the predicted value of the state and the TD error to update the Critic network.
[0080] More specifically, the application combination of the RL agent algorithm framework, that is, the algorithm selection adopted by the RL agent, the selected algorithm is the Actor-Critic framework and the PPO (Proximal Policy Optimaization) algorithm. The PPO algorithm is more suitable for the above state design because the PPO does not depend on the experience replay mechanism. Each time it uses the latest interaction data to learn and does not save it, that is, it only focuses on optimizing the policy for the current state.
[0081] During the training process, the self-play mechanism is adopted to train the agent to construct the R-Tree incrementally. Two identical models are compared, one as the GSAR-Tree and the other as the Competitor. The Competitor always remains the current optimal GSAR-Tree model. On this basis, the GSAR-Tree itself is continuously compared with the Competitor during the construction process, so as to construct in the direction of a tree structure with higher query efficiency.
[0082] The GSAR-Tree model consists of two main parts: the RL agent and the basic R-Tree model. The goal of the training process is to train an agent that can make better child selection decisions based on the states obtained from the R-Tree nodes. In actual test applications, this agent model is used to automatically construct the R-Tree by replacing the locally optimal child selection algorithm of the traditional variant during the process of inserting data one by one.
[0083] The application combination of the RL agent algorithm framework;
[0084] In order to enable the agent to better explore and learn how to make better child selection decision actions according to the input state, in terms of its RL algorithm selection, the Actor-Critic and PPO (Proximal Policy Optimization) algorithms are used.
[0085] The Actor-Critic algorithm trains the policy network Actor and the value network Critic simultaneously to learn the R-Tree construction process. Actor is used to fit a policy function π θ (a|s) to directly give the action a according to the current state s, while Critic is used to fit a state value function V θ (s) to evaluate the quality of the current policy.
[0086] In the Actor-Critic algorithm, first calculate the TD error (Temporal Difference Error) at each time step according to a batch of collected interaction data. The TD error is the difference between the actual return value r t +γV(s t+1 ) and the estimated value V(s t ), where r t is the immediate reward at the current time step, and γ is the discount factor, representing the weight of future rewards;
[0087] δ t =r t +γV(s t+1 )-V(s t )
[0088] Furthermore, the Generalized Advantage Estimation (GAE) is calculated based on the TD error. The advantage function measures how good an action is relative to other actions. If the advantage at the current time step is calculated as positive, the action in the state at that time step is better than the average level.
[0089] A t = δ t + (γλ)δ t+1 + (γλ) 2 δ t+2 +...
[0090] The Proximal Policy Optimization (PPO) algorithm is a gradient-based policy algorithm for reinforcement learning, aiming to directly update the parameters of the policy network by optimizing the objective function. In this paper, a clipped objective function is used to control the magnitude of its policy update, thus avoiding instability during the training process. The final objective function is as follows, where r t (θ) is the ratio of the current policy to the old policy.
[0091] J(θ) = ∑min(r t (θ)A t , clip(r t (θ), 1 - ∈, 1 + ∈)A t )
[0092]
[0093] Finally, PPO directly optimizes the loss of this objective function to update the Actor network, and calculates the mean squared error loss between the predicted value of the state and the TD error to update the Critic network.
[0094] Furthermore, the process of training the reinforcement learning agent to perform incremental R-Tree construction using the self-play mechanism includes:
[0095] Initially, a GSAR-Tree is initialized, which includes an Actor-Critic agent using the PPO algorithm and the R-Tree model itself. Then, an identical copy is initialized as a competitor, which also includes an R-Tree model and an agent, and is used to synchronously construct with the GSAR-Tree. Since the number of splits during the construction of an R-Tree is much less than the number of insert operations, the split operation consistent with the traditional variant is adopted, and the neural network components of the agent are randomly initialized with parameters.
[0096] Inserting all the data in a dataset one by one completely to construct an R-Tree once is regarded as an episode. In each episode, first reset the GSAR-Tree and the competitor. Each time a data is taken out and starts from the root node. If a leaf node is encountered, directly insert the data. Otherwise, extract the corresponding state vector of this node as the input of the agent, obtain the action given by it and execute it, that is, select a child node for insertion, and then continue to iterate and operate layer by layer downward until the data is inserted into the leaf node. For the above operations, the GSAR-Tree and the competitor are executed synchronously each time. After inserting a small batch of data, a part of the tree structure has been constructed and a small number of state-action pairs have also been stored in the agent. At this time, perform the same random range queries on the GSAR-Tree and the competitor multiple times respectively, and then calculate the reward R, and assign the same reward R to this small batch of state-action pairs. Such an approach is common in reinforcement learning. Each time a child node is selected once, and even when a data is completely inserted into the leaf node, multiple child node selections are performed. The effect on the tree structure cannot be immediately shown. Therefore, delayed rewards are adopted to let the impact of the insertion of a small batch of data on the tree structure be manifested, and then the rewards are uniformly assigned.
[0097] If the calculated reward is positive, it indicates that the existing structure of the GSAR-Tree has higher query efficiency, indicating that the neural network component in its agent has learned a better strategy. Then synchronize the GSAR-Tree to the competitor, and let the GSAR-Tree continue to compete with the competitor in the tree construction performance during the training process. That is, the competitor always maintains the current optimal version of the GSAR-Tree, and the GSAR-Tree always compares the query efficiency with its own historical optimal version.
[0098] After multiple rounds of testing, the agent can be placed in a larger test dataset to construct the R-Tree. There is no need to calculate the rewards and construct the competitor anymore. Directly use the agent to replace the traditional algorithm of selecting subtrees. Starting from the root node from top to bottom, extract the current node state and input it to the agent, and execute the output action.
[0099] This embodiment regarding the incremental construction of the R-Tree using reinforcement learning includes three aspects: state design in MDP design, Actor-Critic and PPO algorithms in algorithm selection, and Self-Play application mechanism in training.
[0100] Among them, the state design, especially the three basic components of SDP, LAR, and LPR as state vectors, takes into account the subtree structure information of all children under the current node, further explores the global structure of the R-Tree, and provides it to the RL agent to avoid the local-optimality nature of the selection decision. From the very beginning, the RL agent can pay attention to the existing tree structure during construction, enabling it to continuously explore the potential dependencies between the tree structure under the current node and child selection during exploration and learning, thereby constructing a more efficient R-Tree for queries.
[0101] As a supplementary embodiment, for a similar state design involving node subtrees, such as when SDP needs to traverse the subtree to obtain associated information, it may not be the central distance between the new data and the data of each node in the subtree. It may also be to obtain the perimeter or area information of all nodes in the subtree, or traverse part of the subtree, including the LR series. It may be LOR (Layer Overlap Ratio), calculating the inter-layer overlap area, or the inter-layer ratio of some layers, or layer calculation values, or performing summation or averaging on the layer calculation values / layer inter-layer ratios, etc. They are all similar ideas, that is, traversing the subtree to perform similar calculations and traversing each layer of the subtree to calculate indicators related to similar layer structure information.
[0102] Regarding algorithm selection, variant algorithms of Actor-Critic and PPO may be chosen. For example, A2C, A3C, SAC, etc. are all variant algorithm frameworks of Actor-Critic that can be applied to discrete action spaces, such as PPO-Clip (used in this invention), PPO-Penalty, etc. These variant algorithms essentially make the training process more stable or more effective. However, combined with a similar state design and Self-Play, it is still the R-Tree construction idea based on reinforcement learning adopted in this invention.
[0103] The above is only a preferred specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A global structure-aware R-Tree construction method based on reinforcement learning, characterized in that: include: Designing a Markov decision process constructed by an R-Tree, wherein the Markov decision process constructed by the R-Tree includes a state design, an action design, and a reward design; Apply the reinforcement learning agent algorithm framework and choose to build a reinforcement learning agent based on the Actor-Critic framework and the proximal strategy optimization algorithm; The reinforcement learning agent is trained by a self-playing mechanism to perform R-Tree incremental construction; The trained reinforcement learning agent automatically constructs the R-Tree by inserting data one by one in the construction process, replacing the traditional variant local optimal child selection algorithm.
2. The method according to claim 1, characterized in that The state design includes: Extracting information associated with each child subtree of the current node, and constructing a state vector used as a reference for the agent's decision-making based on the information associated with each child subtree of the current node; The information associated with each child subtree of the current node includes subtree data proximity, layer area ratio and layer perimeter ratio.
3. The method according to claim 1, characterized in that The action design includes: selecting the position of the child node and performing the data insertion operation; assuming that the node capacity is M, the output action space is {insert(1), inert(2), ..., insert(M)}.
4. The method according to claim 1, characterized in that The reward design includes: The degree of improvement in query performance after the test data is inserted is distributed as a reward to the actions taken by all agents during the data insertion process; wherein the degree of improvement in query performance is represented by the difference in query performance from past tests.
5. The method according to claim 4, characterized in that The process of assigning the improvement in query performance after inserting test data as a reward to all actions taken by all agents during the data insertion process includes: Use the self-playing mechanism to synchronously build a competitor identical to GSAR-Tree. After data is inserted, perform the exact same random range query on the two tree structures and calculate the reward R. If the reward R is positive, it indicates that the query efficiency of the constructed GSAR-Tree is better; otherwise, it is not as good as the competitor. Secondly, the formula expression of the reward R is: Where R is the number of node visits after executing random query.
6. The method according to claim 1, characterized in that Applying the reinforcement learning agent algorithm framework, the process of choosing to build a reinforcement learning agent based on the Actor-Critic framework and the proximal policy optimization algorithm includes: The Actor-Critic algorithm is used to simultaneously train the policy network Actor and the value network Critic to learn the R-Tree construction process; wherein the trained policy network Actor is used to fit the policy function π θ (a|s) to directly give action a according to the current state s; the value network Critic is used to fit the state value function V θ (s) to evaluate the quality of the current strategy; The objective function is optimized by the gradient strategy algorithm of reinforcement learning to directly update the parameters of the policy network, and then the objective function loss is optimized to update the Actor network, and the mean square error loss between the predicted value of the state and the TD error is calculated to update the Critic network.
7. The method according to claim 6, characterized in that Using the Actor-Critic algorithm to simultaneously train the policy network Actor and the value network Critic to learn the R-Tree construction process includes: The TD error of each time step is calculated based on the collected interaction data through the Actor-Critic algorithm; wherein the TD error is the actual reward value r of the current time step t +γV(s t+1 ), and the estimated value V(s t ), where r t is the immediate reward of the current time step, γ is the discount factor, and represents the weight of future rewards; δ t =r t +γV(s t+1 )-V(s t ) The generalized advantage estimate GAE is calculated based on the TD error, and the formula is: A t =d t +(gl)d t+1 +(cl) 2 d t+2 +...。 8. The method according to claim 6, characterized in that The process of directly updating the parameters of the policy network by optimizing the objective function through the gradient policy algorithm of reinforcement learning includes: By clipping the objective function to control the magnitude of the policy update, the final objective function expression is: J(θ)=Σmin(r t (i)A t ,clip(r t (θ),1-∈,1+∈)A t ) Among them, r t (θ) is the ratio of the current policy to the old policy.
9. The method according to claim 1, characterized in that: The process of using the self-playing mechanism to train the reinforcement learning agent to perform R-Tree incremental construction includes: Initialize a GSAR-Tree and an identical copy as a competitor, insert all the data in the dataset one by one to build an R-Tree as a round. In each round, the GSAR-Tree and the competitor are executed synchronously. After the data is inserted, perform multiple identical random range queries on the GSAR-Tree and the competitor respectively, calculate the reward R, and if the reward is positive, synchronize the GSAR-Tree to the competitor. After multiple rounds of testing, the trained reinforcement learning agent is used to construct an R-Tree. The trained reinforcement learning agent is used to replace the traditional subtree selection algorithm, starting from the root node from top to bottom, extracting the current node state and inputting it to the agent, and executing the output action.
10. The method according to claim 9, characterized in that Insert all data in the dataset one by one to build an R-Tree as a round. In each round, the process of synchronous execution of GSAR-Tree and competitor includes: Insert all the data in the data set one by one to build an R-Tree as a round. In each round, first reset the GSAR-Tree and competitor, take out one data at a time and start from the root node. If a leaf node is encountered, insert the data directly. Otherwise, extract the state vector corresponding to the node as the input of the agent, obtain the given action and execute it, select a child to insert, and then continue to iterate layer by layer until the data is inserted into the leaf node; GSAR-Tree and competitor are executed synchronously each time. When data is inserted to build a partial tree structure and the state-action pairs are stored through the agent, the same random range query is performed multiple times on GSAR-Tree and competitor respectively, and then the reward R is calculated. The same reward R is assigned to the state-action pair. If the calculated reward is positive, it means that the existing structure query of GSAR-Tree is more efficient and the neural network component in the agent has learned a better strategy. Then GSAR-Tree is synchronized to the competitor, so that GSAR-Tree continues to compete with the competitor in tree building performance during training.
Citation Information
Cited By
Online three-dimensional boxing method and device based on filling layout tree
CN121684216A
An online three-dimensional packing method and device based on packing layout tree
CN121684216B