Index optimization method for big data analysis
Patent Information
- Application Number
- CN202610450371.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-08
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-04-08
AI Technical Summary
[0007]鉴于上述的分析,本发明实施例旨在提供一种面向大数据分析的索引调优方法,用以解决现有无法实现对索引选择与空间分配的兼顾处理影响调优效果的问题
[0018]与现有技术相比,本发明至少可实现如下有益效果之一:
Smart Images

Figure CN121979891B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of database indexing technology, and in particular to an index tuning method for big data analysis. Background Technology
[0002] With the development of data-intensive applications such as big data analytics and artificial intelligence, the ever-growing scale of massive amounts of data has made accessing and processing external data a new performance bottleneck. Furthermore, these cloud-based applications typically deploy storage and computing resources separately to achieve independent scaling of various resources, further exacerbating the cost of cloud data access. Specifically, these applications usually utilize virtual machines provided by cloud vendors to build computing clusters as the computing side, deploying big data analytics and machine learning applications on them. Simultaneously, they use object storage services as the storage side to store data required for analysis, model training, and other tasks. The massive data scale and the separate resource deployment model introduce additional network read costs. While improving resource scalability, this also results in high data read overhead, severely impacting system performance.
[0003] To reduce data access overhead, such data-intensive applications typically use columnar data storage formats, thereby reducing data access costs by avoiding access to irrelevant attribute columns during task execution. Furthermore, data is horizontally partitioned into data blocks to further alleviate data read overhead; these data blocks typically serve as the smallest unit of I / O and analytical processing. To improve data read efficiency, these applications often combine data skipping indexes, which avoid reading irrelevant data blocks to enhance data access efficiency.
[0004] Skipping indexes requires a trade-off between storage overhead and filtering performance. This involves improving data filtering performance to reduce data access costs while avoiding high maintenance costs. Therefore, this technique is crucial for optimizing system performance in emerging data-intensive scenarios such as big data analytics. Compared to fine-grained index structures like B+ trees, this type of technique often has adjustable storage costs, allowing for adjustments based on specific scenarios and user needs.
[0005] To avoid wasting storage resources, after determining the index structure, the system needs to select appropriate attribute columns and create indexes for them (i.e., index tuning).
[0006] Existing index tuning techniques focus on optimizing index tuning effects or computational efficiency, but they are limited to indexing techniques such as B+ trees with fixed storage costs. They cannot balance index selection and space allocation, and it is difficult to skip indexes to select appropriate storage space for data with variable storage costs, thus affecting the tuning effect. Summary of the Invention
[0007] Based on the above analysis, the embodiments of the present invention aim to provide an index tuning method for big data analysis, in order to solve the problem that existing methods cannot achieve a balance between index selection and space allocation, thus affecting the tuning effect.
[0008] On the one hand, embodiments of the present invention provide an index optimization method for big data analysis, including the following steps: Extract global state data; the global state data includes a workload set, a candidate index set, and data block distribution characteristics. Construct a hierarchical reinforcement learning model, which includes a high-level policy network for index selection and a low-level policy network for allocating space for the selected index. The hierarchical reinforcement learning model is trained based on the global state data to obtain a trained hierarchical reinforcement learning model. The state data to be tuned is input into the trained hierarchical reinforcement learning model to obtain the selected index and allocate storage space for the selected index.
[0009] Based on the above method, a further improvement is made to train the low-level policy network in the following way: S311. Randomly select a workload from the workload set as the current workload; randomly select an index from the candidate index set as the current index; S312. Construct the current first state data based on the current workload and current index; S313. Input the current first state data into the low-level policy network to obtain the first action, calculate the reward value and obtain the first state data for the next time step; S314. If the calculated reward value is the same as the reward value of the previous k time steps, determine whether the training of the low-level policy network has reached the preset number of times. If yes, end the training; otherwise, complete the space allocation of the current index and return to step S311. If the calculated reward value is different from the reward value of the previous k time steps, use the first state data of the next time step as the current first state data and return to step S313.
[0010] Based on the further improvement of the above method, the reward obtained by the low-level policy network at time step t is calculated using the following formula. : ; ; in, This indicates the relative reduction in load. This represents the overhead of executing the current workload on the index set created at time step t-1 of the lower-level policy network. This represents the cost of executing the current workload on the index set already created at time step t in the lower-level policy network. This represents the set of indexes created at time step t in the lower-level policy network. This represents the set of indexes created at time step t-1 of the lower-level policy network. Indicates the initial cost. This indicates the change in storage overhead. This represents the amount of storage space obtained from other indexes when the index is created at time step t in the low-level policy network. This represents the penalty coefficient.
[0011] Based on further improvements to the above method, the current first state data includes the load characteristics related to the current index, the distribution characteristics of the data block where the current index is located, the attribute column where the current index is located, the current storage resource information, and the currently generated index information.
[0012] Based on further improvements to the above method, the action space of the low-level policy network includes: increasing storage space for the current index, decreasing storage space for the current index, and reallocating storage space between indexes.
[0013] Based on the above method, a further improvement is made to train the high-level policy network using the following approach: S321. Randomly select a workload from the workload set as the current workload; S322. Construct current second state data based on the current workload; input the current second state data into the higher-level policy network to obtain the second action; S323. Based on the trained low-level policy network, obtain the storage space allocated for the current index corresponding to the second action, and calculate the reward value; S324. If no effective action can be executed, determine whether the training of the high-level policy network has reached the preset number of training iterations. If it has, end the training; otherwise, return to step S321.
[0014] Based on the further improvement of the above method, the reward value of the high-level policy network is calculated using the following formula: ; ; in, This indicates the relative reduction in load. This represents the overhead of executing the current workload on the index set created at time step t-1 of the high-level policy network. This represents the overhead of executing the current workload on the set of indexes already created at time step t in the high-level policy network. Indicates the initial cost. This represents the amount of storage space obtained from other indexes when the index is created at time step t in the high-level policy network; Indicates the penalty coefficient. This indicates the change in storage overhead.
[0015] Based on a further improvement of the above method, the high-level policy network adopts a PPO model with an attention layer; The current workload's load characteristics are used as the query for the attention layer; the selection characteristics of the current workload on each data block are concatenated and used as the key and value for the attention layer to perform attention calculation.
[0016] Based on the above method, the selection characteristics of the current workload on each data block are obtained in the following way: Based on the number of partitions in each attribute column and the value range of each partition, the bit column of the current workload in each attribute column is determined, and a histogram of the current workload is obtained. Perform an AND operation between the histogram of the current workload and the histogram of each data block to obtain the selection characteristics of the current workload on each data block.
[0017] Based on further improvements to the above method, the distribution characteristics of the data blocks include the maximum value, minimum value, mean value, and histogram of each data block.
[0018] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: By decoupling the variable-overhead index tuning problem into two sub-problems—index selection and space allocation—effective tuning of data skipping indexes can be achieved. This ultimately alleviates the data read overhead of data-intensive applications and improves the execution efficiency of cloud-based analytics tasks.
[0019] By extracting the distribution characteristics of each data block, the high-level policy network can effectively perceive the differences in data distribution between data blocks, thus laying the foundation for selecting appropriate data to skip the index and allocating appropriate storage resources to it.
[0020] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0021] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Figure 1 This is a flowchart of an index optimization method for big data analysis according to an embodiment of the present invention. Detailed Implementation
[0022] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0023] A specific embodiment of the present invention discloses an index optimization method for big data analysis, such as... Figure 1 As shown, it includes the following steps: S1. Extract global state data; the global state data includes a workload set, a candidate index set, and data block distribution characteristics; S2. Construct a hierarchical reinforcement learning model, which includes a high-level policy network for index selection and a low-level policy network for allocating space for the selected index. S3. Train the hierarchical reinforcement learning model based on the global state data to obtain a trained hierarchical reinforcement learning model; S4. Obtain the state data to be optimized, input it into the trained hierarchical reinforcement learning model, obtain the selected index, and allocate storage space for the selected index.
[0024] Reinforcement learning can gradually approach the optimal solution by continuously interacting with a complex and dynamic external environment to obtain rewards (or penalties) for actions. This allows reinforcement learning to capture the dynamic characteristics of the environment, making it suitable for complex big data analysis, large-scale modeling, and other application scenarios. In this invention, spatial allocation decisions mainly rely on the currently created indexes, while index selection decisions require global index and storage information. Therefore, this invention proposes a spatially variable index tuning framework based on hierarchical reinforcement learning, using the divide-and-conquer approach.
[0025] Compared with existing technologies, the index tuning method for big data analysis provided in this embodiment decouples the variable storage overhead index tuning problem into two sub-problems: index selection and space allocation by constructing a hierarchical reinforcement learning model. The hierarchical reinforcement learning model is trained by constructing training samples using global state data. The state data to be tuned is input into the trained hierarchical reinforcement learning model to generate an index and allocate appropriate storage space for the generated index. This approach takes into account both index selection and space allocation issues, alleviates the data reading overhead of data-intensive applications, improves the execution efficiency of cloud analysis tasks, and achieves effective tuning of data that skips indexes.
[0026] Before building and training a hierarchical reinforcement learning model, it is necessary to extract global state data, including building a workload set and a candidate index set, and extracting the distribution characteristics of data blocks.
[0027] In implementation, historical user queries can be collected, and a subset of queries can be randomly selected to form a workload, thus constructing a workload set. A workload is a set of query statements.
[0028] During implementation, the workload is converted into a vector representation using an existing query encoding model, serving as the workload characteristics. For example, the BOO (Bag of Operators) model can be used to obtain the workload characteristics.
[0029] During implementation, all indexable attribute columns of all data blocks in the database are collected as the basic elements of candidate indexes, and then arranged and combined to generate a candidate index set.
[0030] It should be noted that this invention considers indexes composed of different attribute columns and indexes composed of the same attribute columns on different data blocks as different indexes. An index can be composed of a single column on a data block, called a one-dimensional index, or it can be composed of multiple columns on a data block, called a multi-dimensional index. For example, if data block 1 includes three attribute columns a, b, and c, then there are three one-dimensional indexes (a, b), (a, c), (b, a), (b, c), (c, a), and (c, b), six two-dimensional indexes (a, b, c), (a, c, b), (b, a, c), (b, c, a), (c, a, b), and six three-dimensional indexes (c, b, a). These are all different operations. Similarly, the index on attribute column 'a' in data block 1 and the index on attribute column 'a' in data block 2 are also different operations. The high-level policy network selects the index to be generated from the candidate index set.
[0031] During implementation, for each data block in the database, the distribution characteristics of that data block are extracted.
[0032] Existing index tuning algorithms typically represent the load uniformly. However, in scenarios where data is stored in blocks, differences in data distribution between blocks often lead to varying filtering performance of the data skipping index across different blocks. Therefore, representing each data block uniformly makes it difficult for the algorithm to effectively distinguish the differences in data distribution between blocks, thus hindering the accurate evaluation of the data skipping index's filtering performance across different blocks and ultimately impacting the tuning results.
[0033] The differences in data distribution between data blocks can affect the allocation of storage resources for the index. Therefore, it is necessary to represent the features of each data block and extract its distribution features so that the high-level policy network can effectively perceive the differences in data distribution between data blocks. This will lay the foundation for selecting appropriate data to skip the index and allocating appropriate storage resources to it.
[0034] Specifically, the data block distribution characteristics include the maximum, minimum, mean, and histogram of each data block. The maximum value of a data block is the vector formed by the maximum values of each numerical attribute column; the minimum value of a data block is the vector formed by the minimum values of each numerical attribute column; and the mean of a data block is the vector formed by the mean values of each numerical attribute column.
[0035] Specifically, the histogram of the data block is obtained using the following method: For each numeric attribute column, determine the number of partitions for that attribute column and the range of values for each partition; For each data block, the bit columns of the data block on each attribute column are determined based on the number of partitions in each attribute column and the value range of each partition. The bit columns of the data block on all attribute columns constitute the histogram of the data block.
[0036] It should be noted that the number of buckets (partitions) is the same for each attribute column.
[0037] For example, a numeric attribute column A defines 7 partitions, with the value range of each partition as shown in Table 1. Data block information is shown in Table 2. For data block 1, column A contains values within the ranges of partitions 2, 3, and 5, with the corresponding bit value being 0; it does not contain values within the ranges of partitions 1, 4, 6, and 7, with the corresponding bit value also being 0. Therefore, the bit representation of data block 1 in attribute column A is "0110100". Similarly, the bit representation of data block 2 in attribute column A is "0000011", and the bit representation of data block 3 in attribute column A is "0000100".
[0038] The bit columns of all numeric attribute columns in data block 1 are merged together to form a histogram of data block 1.
[0039] Table 1 Partition Information Table
[0040] Table 2 Data Block Information Table
[0041] It should be noted that the distribution characteristics of the data blocks only apply to numerical attribute columns.
[0042] The constructed hierarchical reinforcement learning model includes a high-level policy network and a low-level policy network.
[0043] The high-level policy network is used for index selection, that is, choosing indexes from the candidate index set and determining which attribute columns of which data blocks to generate the indexes on. The low-level policy network is used for space allocation of the selected indexes. In implementation, the low-level and high-level policy networks are trained in stages.
[0044] In implementation, the low-level policy network can adopt existing network structures, such as multilayer perceptrons, to build agents.
[0045] The Markov Decision Process (MDP) for low-level policy network space allocation can be represented by tuples. ,in and These represent the state space and action space for the spatial allocation MDP process, respectively. For the reward function, Represents the transition matrix. Then represents the discount factor, where, and These are the parameters learned by the low-level policy network.
[0046] Specifically, the low-level policy network is trained using the following method: S311. Randomly select a workload from the workload set as the current workload; randomly select an index from the candidate index set as the current index; S312. Construct the current first state data based on the current workload and current index; S313. Input the current first state data into the low-level policy network to obtain the first action, calculate the reward value and obtain the first state data for the next time step; S314. If the calculated reward value is the same as the reward value of the previous k time steps, determine whether the training of the low-level policy network has reached the preset number of times. If yes, end the training; otherwise, complete the space allocation of the current index and return to step S311. If the calculated reward value is different from the reward value of the previous k time steps, use the first state data of the next time step as the current first state data and return to step S313.
[0047] During implementation, a workload is randomly selected from the workload set as the current workload. An index is randomly selected from the candidate index set as the current index, and the low-level policy network allocates storage space for the current index through interactions over multiple time steps. If the training does not reach the preset number of iterations, workloads and indices are selected again for training of the low-level policy network.
[0048] Specifically, the current first-state data, constructed based on the current workload and current indexes, includes the load characteristics related to the current index, the distribution characteristics of the data block containing the current index, the attribute column containing the current index, the current storage resource information, and the information of currently generated indexes. The current storage resource information includes the current remaining storage resource size, and the information of currently generated indexes includes the currently created indexes and their corresponding storage space sizes.
[0049] It should be noted that the load characteristics related to the current index are obtained by filtering out query statements related to the current index from the current workload (i.e., query statements whose predicates contain the attribute columns corresponding to the current index). The filtered query statements are converted into vector representations through the query statement encoding model, which are the load characteristics related to the current index.
[0050] In a low-level policy network, the action space at each time step This includes adding storage space to the current index, reducing storage space for the current index, and reallocating storage space between indexes.
[0051] Add storage space to the current index, that is, take s storage units from the remaining storage resources as s storage units for the current index.
[0052] Reduce the storage space of the current index by subtracting s storage units from the current index's storage space.
[0053] Inter-index space reallocation involves taking *s* storage units from the storage space of an existing index and using them as the *s* storage units for the current index. The storage space of the current index increases by *s* units, while the storage space of the index that was taken decreases by *s* units. S represents the preset maximum number of storage units to be redistributed.
[0054] In implementation, the low-level policy network uses the Proximal Policy Optimization (PPO) algorithm to obtain the optimal policy, i.e., the action. .
[0055] The low-level policy network is based on the current first state data. Execute the first action Receive rewards and the first state data of the next time step .
[0056] Specifically, the reward obtained by the low-level policy network at time step t is calculated using the following formula. : ; ; in, This indicates the relative reduction in load. This represents the overhead of executing the current workload on the index set created at time step t-1 of the lower-level policy network. This represents the overhead of performing the current workload on the index set created at time step t in the lower-level policy network, including time overhead and I / O data transfer overhead. This represents the set of indexes created at time step t in the lower-level policy network. This represents the set of indexes created at time step t-1 of the lower-level policy network. This represents the initial overhead, which is the cost of executing the current workload without any indexes. The cost of executing the current workload includes time overhead and I / O data transfer overhead.
[0057] This indicates the change in storage overhead. , This represents the memory overhead of the index created at time step t in the low-level policy network. This represents the memory overhead of the index created at time step t-1 of the lower-level policy network.
[0058] This represents the amount of storage space obtained from other indexes when the index is created at time step t in the low-level policy network. This represents the penalty coefficient.
[0059] The reason for choosing to design the reward function using relative load overhead is that absolute load overhead, by not considering the impact of load type and required storage overhead, can lead to significant differences in the index combinations created by similar actions under different loads, thus affecting the training results. It is worth noting that, compared to traditional index tuning problems, the storage resources allocated by the strategy of this invention may come from other indexes, resulting in a constant total storage overhead. Specifically, the actions of this invention include obtaining storage space from other indexes. In this case, although the overall system performance may change, the total storage overhead remains unchanged (i.e., ...). To make the model prioritize using unused storage resources, the denominator of the reward function is replaced with... ( ),in The penalty coefficient is... This is the storage space size obtained from other indexes.
[0060] The low-level policy network uses existing technology to optimize policies based on reward values.
[0061] For the current index, if the reward value remains unchanged for k consecutive time steps, then the space allocated to that current index is obtained, which completes one round of the low-level policy network. One round is one training session. In practice, k can be set according to the efficiency and accuracy requirements of the training.
[0062] After the low-level policy network has been trained for a preset number of times, training is stopped, and the trained low-level policy network is obtained.
[0063] The parameters of the trained low-level policy network are fixed, and the high-level policy network is trained based on the trained low-level policy network.
[0064] The primary function of the high-level policy network is to select the most suitable index for creation. The MDP for the index selection process can be represented by tuples. ,in and This indicates the state space and action space of the index selection MDP. This represents the reward function. The reward function of the higher-level policy network is the same as that of the lower-level policy network. Represents the transition matrix. Then represents the discount factor, where and Parameters for learning high-level policy networks.
[0065] Specifically, the high-level policy network is trained using the following methods: S321. Randomly select a workload from the workload set as the current workload; S322. Construct current second state data based on the current workload; input the current second state data into the higher-level policy network to obtain the second action; S323. Based on the trained low-level policy network, obtain the storage space allocated for the current index corresponding to the second action, and calculate the reward value; S324. If no effective action can be executed, determine whether the training of the high-level policy network has reached the preset number of training iterations. If it has, end the training; otherwise, return to step S321.
[0066] During implementation, a workload is randomly selected from the workload set as the current workload.
[0067] The second state data includes the load characteristics of the current workload, the distribution characteristics of all data blocks, the current storage resource information, and the generated index information.
[0068] The action space of the index selection strategy in the high-level policy network is discrete. Each action in the action space corresponds to a candidate index in the candidate index set; that is, the action space is the set of all candidate indexes. It should be noted that this invention considers indexes composed of different attribute columns, as well as indexes composed of the same attribute columns on different data blocks, as different actions. This allows the invention to select and create appropriate indexes based on the data distribution characteristics of the data blocks.
[0069] At each time step of training the high-level policy network, based on the current second state... Generate a policy - second action In essence, after selecting an index as the current index to be generated, the load characteristics related to the current index, the distribution characteristics of the data block where the current index is located, the attribute columns of the current index, the current storage resource information, and the information of the currently generated indexes are input into the low-level policy network. The low-level policy network then allocates storage space for the current index, and the high-level policy network calculates the reward. Next, it is determined whether there is a valid action to be executed. If all indexes in the candidate index set have been created or the remaining storage resources are less than a preset value, then there is no valid action to be executed. In this case, one round of the high-level policy network ends, and one iteration of training of the high-level policy network is completed. Then, it is determined whether the training of the high-level policy network has reached the preset number of training times. If it has, the training ends; otherwise, it returns to step S321.
[0070] It should be noted that after one round of the high-level policy network, the workload set, candidate index set, current storage resource information, and currently generated index information are all reset to their initial values before network training.
[0071] During implementation, the reward value of the high-level policy network is calculated using the following formula: ; ; in, This indicates the relative reduction in load. This represents the overhead of executing the current workload on the index set created at time step t-1 of the high-level policy network. This represents the overhead of executing the current workload on the set of indexes already created at time step t in the high-level policy network. Indicates the initial cost. This represents the amount of storage space obtained from other indexes when the index is created at time step t in the high-level policy network. Indicates the penalty coefficient. This indicates the change in storage overhead. , This represents the memory overhead of the index created at time step t in the high-level policy network. This represents the memory overhead of the index created at time step t-1 of the high-level policy network.
[0072] High-level policy networks use existing technologies to optimize policies based on reward values.
[0073] Since the index selection agent needs to acquire complex global information to make discrete decisions, the high-level policy network uses a PPO model to learn appropriate index selection decisions under complex data distribution and load characteristics. Meanwhile, to avoid the massive input data resulting from global block-level load and data distribution characteristics affecting the model's training efficiency and processing performance, the high-level policy network of this invention adopts a PPO model with an attention layer, that is, an attention layer is inserted before the fully connected layer in the traditional PPO model.
[0074] Specifically, this attention network layer takes a query, a set of keys, and a set of values as input. By evaluating the relevance between the query and the keys (e.g., using dot product (DP)), this network layer performs a weighted summation operation on the values and outputs the final result, where the weights are the relevance between the query and each key learned by the model.
[0075] During implementation, the load characteristics of the current workload are used as the query of the attention layer; the selection characteristics of the current workload on each data block are concatenated and used as the key and value of the attention layer for attention calculation.
[0076] Specifically, the selection characteristics of the current workload on each data block are obtained in the following way: Based on the number of partitions in each attribute column and the value range of each partition, the bit column of the current workload in each attribute column is determined, and a histogram of the current workload is obtained. Perform an AND operation between the histogram of the current workload and the histogram of each data block to obtain the selection characteristics of the current workload on each data block.
[0077] The query statement for the current workload is, for example, "select The query "from table where A between 3000 and 3500" refers to the method for determining the bit column of each attribute column in the aforementioned data block. The predicate of the current workload hits partition 5 of attribute column A; therefore, the bit column in attribute column A is represented as "0000100". The bits of the current workload in each attribute column are concatenated to obtain the histogram of the current workload.
[0078] Using the above information, the attention layer can calculate the relevance of each data block to the current workload. Its final output, the embedding information of each data block, is obtained by weighted summation of the workload embedding features of each data block based on this relevance. Intuitively, the attention-based neural network assigns greater weight to data blocks that are more suitable for indexing under the current workload, thereby attracting more attention from the agent and guiding the agent to select a more appropriate index.
[0079] The time steps of high-level policy networks differ from those of low-level policy networks; a single time step in a high-level policy network contains multiple time steps from the low-level policy network.
[0080] It should be noted that, to constrain model training and processing costs, this invention uses an invalid action masking technique when training both low-level and high-level policy networks. This technique temporarily disables some actions in the current state to accelerate model training efficiency. Specifically, during low-level policy network training, actions are masked primarily from the perspective of indexing and storage overhead. For example, actions exceeding the upper bound of indexing and storage overhead, or exceeding storage overhead constraints, are masked.
[0081] For high-level policy networks, invalid action masks are implemented by excluding candidate indexes that are irrelevant to the workload or that have or lack preceding indexes. Based on heuristics commonly used in existing work, multi-column indexes are only considered valid actions if a single-column index has already been built on its preceding column.
[0082] After obtaining the trained hierarchical reinforcement learning model, the state data to be tuned is acquired and input into the trained hierarchical reinforcement learning model to generate an index and allocate storage space for the generated index. Specifically, the current second-state data (load characteristics of the current workload, distribution characteristics of all data blocks, current storage resource information, and information on generated indexes) is input into the high-level policy network to generate the current index. The load characteristics related to the current index, the distribution characteristics of the data block containing the current index, the attribute column containing the current index, the current storage resource information, and the information on currently generated indexes are input into the low-level policy network to obtain the storage space for the current index. By continuously repeating these two processes, all data that needs to be created is selected, the index is skipped, and the corresponding storage overhead is allocated to it, thus obtaining the end-to-end load overhead for the corresponding state.
[0083] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0084] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. An index optimization method for big data analysis, characterized in that, Includes the following steps: Extract global state data; the global state data includes a workload set, a candidate index set, and data block distribution characteristics. Construct a hierarchical reinforcement learning model, which includes a high-level policy network for index selection and a low-level policy network for allocating space for the selected index. Based on the global state data, the high-level policy network and low-level policy network of the hierarchical reinforcement learning model are trained in stages. First, the low-level policy network is trained, and then the high-level policy network is trained based on the trained low-level policy network, thus obtaining the trained hierarchical reinforcement learning model. The state data to be tuned is input into the trained hierarchical reinforcement learning model to obtain the selected index and allocate storage space for the selected index. The action space of the low-level policy network includes: adding storage space to the current index, reducing storage space to the current index, and inter-index reallocation; adding storage space to the current index means taking s storage units from the remaining storage resources as s storage units for the current index; reducing storage space to the current index means subtracting s storage units from the storage space of the current index; inter-index reallocation means taking s storage units from the storage space of the created index as s storage units for the current index, increasing the storage space of the current index by s storage units, and reducing the storage space of the index that was taken from the index by s storage units. The reward obtained by the low-level policy network at time step t is calculated using the following formula. : ; ; in, This indicates the relative reduction in load. This represents the overhead of executing the current workload on the index set created at time step t-1 of the lower-level policy network. This represents the cost of executing the current workload on the index set created at time step t in the lower-level policy network. This represents the set of indexes created at time step t in the lower-level policy network. This represents the set of indexes created at time step t-1 of the lower-level policy network. Indicates the initial cost. This indicates the change in storage overhead. This represents the amount of storage space obtained from other indexes when the index is created at time step t in the low-level policy network. This represents the penalty coefficient.
2. The index optimization method for big data analysis according to claim 1, characterized in that, The low-level policy network is trained using the following method: S311. Randomly select a workload from the workload set as the current workload; randomly select an index from the candidate index set as the current index; S312. Construct the current first state data based on the current workload and current index; S313. Input the current first state data into the low-level policy network to obtain the first action, calculate the reward value and obtain the first state data for the next time step; S314. If the calculated reward value is the same as the reward value of the previous k time steps, determine whether the training of the low-level policy network has reached the preset number of times. If yes, end the training; otherwise, complete the space allocation of the current index and return to step S311. If the calculated reward value is different from the reward value of the previous k time steps, use the first state data of the next time step as the current first state data and return to step S313.
3. The index optimization method for big data analysis according to claim 2, characterized in that, The current first state data includes the load characteristics related to the current index, the distribution characteristics of the data block where the current index is located, the attribute column where the current index is located, the current storage resource information, and the currently generated index information.
4. The index optimization method for big data analysis according to claim 1, characterized in that, The high-level policy network is trained using the following method: S321. Randomly select a workload from the workload set as the current workload; S322. Construct current second state data based on the current workload; input the current second state data into the higher-level policy network to obtain the second action; S323. Based on the trained low-level policy network, obtain the storage space allocated for the current index corresponding to the second action, and calculate the reward value; S324. If no effective action can be executed, determine whether the training of the high-level policy network has reached the preset number of training iterations. If it has, end the training; otherwise, return to step S321.
5. The index tuning method for big data analysis according to claim 4, characterized in that, The reward value of the high-level policy network is calculated using the following formula: ; ; in, This indicates the relative reduction in load. This represents the overhead of executing the current workload on the index set created at time step t-1 of the high-level policy network. This represents the overhead of executing the current workload on the set of indexes already created at time step t in the high-level policy network. Indicates the initial cost. This represents the amount of storage space obtained from other indexes when the index is created at time step t in the high-level policy network; Indicates the penalty coefficient. This indicates the change in storage overhead.
6. The index tuning method for big data analysis according to claim 4, characterized in that, The high-level policy network adopts the PPO model with an attention layer; The current workload's load characteristics are used as the query for the attention layer; the selection characteristics of the current workload on each data block are concatenated and used as the key and value for the attention layer to perform attention calculation.
7. The index optimization method for big data analysis according to claim 6, characterized in that, The selection characteristics of the current workload on each data block are obtained in the following way: Based on the number of partitions in each attribute column and the value range of each partition, the bit column of the current workload in each attribute column is determined, and a histogram of the current workload is obtained. Perform an AND operation between the histogram of the current workload and the histogram of each data block to obtain the selection characteristics of the current workload on each data block.
8. The index tuning method for big data analysis according to claim 1, characterized in that, The distribution characteristics of the data blocks include the maximum value, minimum value, mean value, and histogram of each data block.
Citation Information
Patent Citations
Adaptive joint segmentation federated learning method based on HRL in combination with PPO algorithm
CN121638383A