Dynamic fitting, evaluation and learning index construction method and equipment
Through dynamic fitting and evaluation methods, the data key mapping address in the learning index is adjusted, which solves the problems of high index growth and reduced efficiency, realizes efficient data query and insertion operations, and reduces the storage consumption of the index.
Patent Information
- Application Number
- CN202510068870.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
As the data volume increases, the indexes increase highly, resulting in reduced query and insertion efficiency, and traditional index reconstruction methods will lead to a decrease in the overall index throughput.
A dynamic fitting, evaluation, and learning index construction method is proposed. By calculating the average difference value of the data key and the difference value of the data pair, the data is divided into two groups, and the corresponding data points are stored using the minimum heap and the maximum heap, the gap terms are inserted to adjust the mapping address of the data key, and finally the fitted line parameters are calculated using the least squares method.
It effectively improves the fitting performance of the index model, significantly reduces the overall height of the index, improves the overall throughput of the index, and achieves space savings and conflict reduction to a large extent.
Smart Images

Figure CN120067103A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and particularly relates to a method and device for dynamically fitting, evaluating, and constructing a learning index. Background Art
[0002] As one of the important technologies supporting efficient data reading in databases, indexes have become increasingly important in the era of massive data. Currently, indexes based on tree structures (such as B+ trees) are widely used in the database field. By organizing data into multi-way trees according to certain rules, query efficiency is improved. However, with the explosive growth of data volume, traditional tree-structured indexes have exposed defects such as a relatively large average search length and excessive memory occupation, and it is necessary to further improve their query performance.
[0003] To solve the problems of traditional indexes, Kraska et al. proposed the concept of "learning index" and adopted the method of Recursive Model Index (RMI). The recursive model index establishes a hierarchical structure of models, and different models can be selected at each level to better fit the data distribution. However, RMI can only build indexes relying on existing data and does not support index updates and insertions.
[0004] To support efficient insertion operations, Ding et al. proposed the Adaptive Learned index ALEX. ALEX distributes data into multiple different partitions and performs independent model training on each partition. These models all use linear regression models. The leaf nodes of ALEX use Gapped Array (GA) to store keys, and gaps are filled between adjacent keys. When a new key is inserted, ALEX uses the model to predict the insertion position. If the position is exactly a gap, it is inserted directly. Otherwise, it uses exponential search to find the nearest gap to the current position and inserts. However, as the data volume increases, the gaps will be gradually filled, and at this time, a large number of key movements are required. In extreme cases, this movement will consume a large amount of time.
[0005] LIPP (Updatable Learned Index with Precise Positions) solves the problem of inaccurate model prediction in previous learned indexes. LIPP adopts an unbalanced tree structure, and each node uses a separate linear regression model. LIPP divides the elements in the node into three types: DATA (indicating the stored data item), NULL (indicating the void item), and NODE (indicating the pointer to the child node in the next layer). In the insertion operation, prediction is made through the model of the current node, and corresponding processing is carried out according to the prediction result. Specifically: 1) If the corresponding position is NULL, the data is directly inserted; 2) If the corresponding position is DATA, a new node is constructed using the data item at the current position and the newly inserted data item, and the element type at the current position is set to a pointer, pointing to the newly created node; 3) If the corresponding position is NODE, access is made through the pointer to enter the next layer of nodes and continue the query.
[0006] However, since a node in the next layer is created every time a prediction conflict occurs, as the amount of data and the number of conflicts increase, the index height of LIPP will gradually increase, and in extreme cases, a structure similar to a linked list will be formed, as Figure 1 shown. This will lead to a gradual decrease in the efficiency of querying and insertion. Although the article proposes using index reconstruction to reduce the impact, the reconstruction process will still result in a decrease in the overall throughput of the index. In addition, LIPP adopts the FMCD (Fastest Minimum Confilict Degree) linear fitting algorithm it proposed for index reconstruction, but the conflict degree of this algorithm is closely related to the magnification factor of the capacity of the new node and the original node. To reduce the conflict degree, LIPP adopts a magnification factor of 6. However, the magnification factor is determined empirically and has no adaptability to the data distribution, which may lead to a large number of void items inside the index, thus wasting a large amount of space. Summary of the Invention
[0007] In view of the deficiencies of the prior art, the present invention proposes a method and device for constructing a dynamically fitted, evaluated, and learned index to ensure an efficient data query process.
[0008] A dynamic fitting method designed by the present invention, in the construction of a learned index, first calculates the average difference value of the data key set , and then calculates the difference values of data pairs in sequence. According to the size relationship between them and the average difference value, the data is divided into two groups: The data points with difference values greater than enter group , and the elements in the group are stored using a minimum heap; the data points with difference values less than enter group , the elements within the group are stored using a max heap; among them, a data pair refers to two data keys; Traverse the data key set and select the smallest element in the group, insert a gap item between the data pairs, then update the average difference value and continue to select pairs until all the gap items are inserted, and remove the data points with the smallest difference value in the group B data pairs, where is a system parameter set to , where represents the number of elements in the data key set; Using the data key set formed by the processed group A and group B, calculate the fitting line parameters using the least squares method to obtain the fitting model.
[0009] Preferably, in the process of calculating the fitting line parameters by the least squares method, if data points are represented as: , the linear function is used to fit data points, and are the parameters to be solved; define the objective function , where From to , respectively, take the partial derivatives of and and set the partial derivatives to to obtain the linear equation system:
[0010] Solve the linear equation system to obtain the analytical formula of the least squares method:
[0011] Calculate the slope and intercept of the best fitting line through the analytical formula of the least squares method.
[0012] Based on the same inventive concept, the present invention also designs an evaluation method for evaluating the dynamic fitting method, and the process is as follows: The fitting indicators include: Conflict_ratio represents the situation of conflicts when data is inserted according to the mapped address of the fitting model, and Empty_ratio represents the number of addresses without stored data; Introduce the data fitting evaluation index T T = Conflict_ratio + Empty_ratio
[0013] where Conflict_num element_num is the number of elements in the data key set; empty_num represents the number of empty items. This metric emphasizes that when evaluating the model, not only the conflict rate of the data keys needs to be considered, but also the waste of storage space. The smaller T is, the lower the conflict rate and space waste rate of the model, and the better the model performance.
[0014] Based on the same inventive concept, the present invention also designs a learning index construction method based on dynamic fitting, In the tree-structured data storage, nodes are divided into buffer nodes and formal nodes. Buffer nodes are used to temporarily store a limited number of data elements, and formal nodes are used to store the fitted data elements. The fitting uses the dynamic fitting method of the present invention to fit the data. Formal nodes store data items, pointer items, and empty items in the form of an array; When inserting data, buffer nodes are inserted directly. In formal nodes, empty items are inserted directly, pointer items enter the lower-level nodes for insertion, and for data items, a new buffer node is created for data insertion. The successfully inserted data key is written into the Bloom filter, and buffer nodes with a data volume exceeding the threshold are adjusted to formal nodes; When querying, after the Bloom filter returns a true value, the node type is judged, and then the data access method is determined according to the node type.
[0015] Further, the buffer nodes internally adopt a tightly arranged manner, store data items in the form of an array or a linked list, and perform queries through binary search. Further, the query process is specifically as follows: If the node is a buffer node, binary search is used inside the node to find the data key and obtain the corresponding data value; If the current node is a formal node, first use the fitting model in the formal node to calculate the position of the data key in the node, and then access the element at that position: if the position is a data item, directly return the corresponding data value; if the position is a pointer item, enter the lower-level element pointed to by the pointer to continue the query; if the position is an empty item, the query ends.
[0016] Further, the retraining of the fitting model is determined by the following factors: The current node type is a buffer node, and the number of elements in the node exceeds a predetermined threshold; The current node type is a formal node, and the number of conflicts that have occurred inside the node satisfies the following formula condition,
[0017] where represents the number of conflicts that have occurred in the current node and its subtree, Represents the total number of elements of a node and its subtree, Represents the initial number of elements of a node when it is established, which is a preset parameter.
[0018] Furthermore, during the retraining process, first traverse and collect all data key - value pairs on the current node and its child nodes, calculate a new fitting model using the dynamic fitting method, insert these data key - value pairs into a new node whose length is greater than the length of the original node and then point the pointer of the original node to the new node.
[0019] Based on the same inventive concept, the present invention also designs an electronic device, including: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the learning index construction method based on dynamic fitting.
[0020] Based on the same inventive concept, the present invention also designs a computer - readable medium, on which a computer program is stored, and when the program is executed by a processor, the learning index construction method based on dynamic fitting is implemented.
[0021] The advantages of the present invention are as follows: 1. The present invention proposes a data fitting algorithm based on the differential mean of data points, which can more accurately perceive the distribution of the data key set, and use the empty items in the index to adjust the mapping address of the data key, effectively improving the fitting performance of the index model. At the same time, through the top - down index construction method for local index reconstruction, more data can enter lower - level nodes, thus significantly reducing the overall height of the index and improving the overall throughput of the index.
[0022] 2. The present invention proposes a data fitting algorithm evaluation strategy based on the conflict rate and the vacancy rate. This strategy emphasizes considering both the conflict situation of data keys and the waste of storage space. The data fitting algorithm designed based on this evaluation strategy can achieve significant space savings and conflict reduction to a large extent, thereby improving the overall performance of the learning index and effectively reducing its storage consumption.
[0023] 3. The present invention designs a request filtering module and a data storage module. The request filtering module can efficiently filter invalid query requests, thereby improving the overall read throughput; the data storage module as a whole adopts a tree - type structure. By dividing into formal nodes and buffer nodes, and cooperating with the corresponding data fitting algorithm and local index reconstruction strategy, it effectively reduces the increase in index height caused by data conflicts and further optimizes the index performance. Brief Description of the Drawings
[0024] Figure 1 It is an example of the deterioration of the updatable exact fitting index LIPP.
[0025] Figure 2 It is the index structure of the learning index construction method based on dynamic fitting of the present invention.
[0026] Figure 3 It is a schematic diagram of the algorithm principle of the fitting process of the present invention. Specific embodiments
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0028] Embodiment 1 This embodiment discloses a dynamic fitting method. In the construction of a learning index, this method optimizes the fitting effect by dynamically adjusting the positions of data keys. Specifically, consider a set of ordered data keys , and the goal is to place these data keys on an array Arr of length L in a certain order . The set of positions of the data keys is , satisfying . Convert the one-dimensional data keys into two-dimensional data points , and perform fitting. It is necessary to find the optimal set of positions P such that the most data points fall on the same straight line. That is, if the cosine similarity of the connection lines between adjacent points is higher, then the possibility of these points falling on the same straight line is higher.
[0029] When the length L of the array exceeds the number n of elements in the data key set, after address mapping processing, some addresses that do not store data will be generated, and this address is called a "void item". In the fitting algorithms used in existing learning indexes, some do not utilize the void items and directly perform data fitting, and some only evenly insert the void items between the data keys and then perform data fitting. However, the above fitting methods ignore the relationship between the void items and the distribution of data points, and the effect is not ideal. For example Figure 3As shown, by reasonably inserting gap items, the present invention can adjust the slope of the segmented line segments, making the overall distribution of data points approximate to a straight line. In addition, the difference values between data keys can reveal the degree of data fluctuation. If the difference values are closer, then the fluctuation of the data keys is smaller, the cumulative distribution function of the data is closer to a straight line, and the possibility of fitting the data set into a straight line is greater. Therefore, the key to the fitting method proposed by the present invention lies in determining how to insert gaps by observing the data difference values, so as to improve the fitting effect.
[0030] The dynamic fitting method proposed in this embodiment is as shown in the appendix Figure 1 First, calculate the average difference value of the data keys , and then calculate the difference values of data pairs in turn. A data pair refers to every two data keys. According to the size relationship between the difference value of the data pair and the average difference value , the data is divided into two groups: the data points with the difference value of the data pair greater than enter group A, and the elements in the group are stored using a min-heap; the data points with the difference value of the data pair less than enter group B, and the elements in the group are stored using a max-heap. Then, perform the gap item insertion process. The number of gap items is determined by the upper limit of the node capacity set by the system. The number of gaps + the number of data keys = the upper limit of the node capacity. In fact, no matter where the gap item is inserted, the impact on the average difference value is the same. As shown in Equation (1), the overall average difference value of the data key set after inserting the gap item is equal to the difference between the maximum data key and the minimum data key divided by the address space length after inserting the gap. In the formula, new_avg_gap represents the average difference value after insertion, represents the data key set, max_key represents the maximum value of the data key set, min_key represents the minimum value of the data key set, and array_length represents the number of data keys.
[0031] (1) However, when the gap item is inserted at different positions, the improvement of the local data distribution is different. As shown in Equation (2): the local slope between data point pairs is equal to the difference between data keys divided by the number of gap items between the two points. If the local slope is closer to the average slope of the data key set, then the two data points are closer to the overall straight line. In the formula, partial_slope represents the local slope, Arr represents the data key set, and gaps_nums represents the number of gaps.
[0032] (2) In a real scenario, there may be some data pairs with difference values much larger than the average difference value, which are called "outliers". If a gap item is inserted in such a pair, the local slope may still be far from the overall average slope, thus wasting a large number of gap items. Therefore, during traversal, the smallest element in group A is selected each time. The slope of the line connecting this data point and its previous data point is slightly larger than the overall average slope. A gap item is inserted between this pair, that is, the mapped address of the latter data point is moved one position backward, making the slope of the line between the two data points gentler. This insertion strategy makes it easier for the local slope to approach the average slope. Then, the average difference value is updated using Equation (1) and pairs are continued to be selected. This process continues until all gap items are inserted. Then, the data point with the smallest difference value in group B is removed to avoid the influence of "outliers" with too small difference values and further improve the overall data distribution. is a system parameter and is set to , where represents the number of elements in the data key set. Finally, least squares fitting is performed based on the improved data distribution.
[0033] Least squares method is a classic data fitting algorithm, and its goal is to minimize the sum of the squares of the errors between the predicted values and the actual values. Specifically, assume there are data points: , then a linear function needs to be found to fit data points. Among them, and are the parameters to be solved.
[0034] To achieve this goal, first define the objective function , where ranges from to . Then, take the partial derivatives of and respectively, and set the partial derivatives to , obtaining a system of linear equations as shown in Equation (3): (3) Solving this system of linear equations can obtain the analytical formula of the least squares method, as shown in Equation (4): (4) Through the above formula, the slope and intercept of the best fitting line are calculated. and represent the average values of x and y respectively. Through the above process, a model with a better fitting effect can be obtained, thus improving the insertion and query efficiency of the index.
[0035] Example 2 Based on the same inventive concept, this embodiment also discloses an evaluation method for evaluating the dynamic fitting method in Embodiment 1. For the sorted data key set K, a function needs to be found according to its distribution characteristics , which maps any key to an array of length L, such that . This process is called the data fitting process. The function must be a monotonically non-decreasing function, and the higher the degree of , the higher the computational complexity usually is. Therefore, the present invention selects the most commonly used linear function of degree one as the fitting function. However, in actual scenarios, it is difficult to find a perfect for a given data key set. Currently, for the concept of conflict degree, only the number of key conflicts at a single position is generally described. To solve this problem, the present invention proposes a data fitting evaluation index T for evaluating the quality of the linear fitting model in the learning index scenario. As shown in Equation (6), the fitting index consists of two parts, represents the situation of conflicts occurring after the data is inserted according to the mapped address of the fitting model, and Empty_ratio represents the number of addresses without stored data.
[0036] T = Conflict_ratio + Empty_ratio (6) In the formula, Conflict_num element_num is the number of elements in the data key set; empty_num represents the number of empty items; this index emphasizes that when evaluating the model, not only the conflict rate of data keys needs to be considered, but also the waste of storage space needs to be concerned. The smaller T is, the lower the conflict rate and space waste rate of the model are, and the better the model effect is.
[0037] Example 3 Based on the same inventive concept, this embodiment also discloses a learning index construction method based on dynamic fitting. As Figure 2As shown in the figure, the learning index consists of two parts: a request filtering module and a data storage module. The request filtering module is an optional part implemented by a Bloom filter, which can quickly determine whether an element exists, and the time and space complexity of writing and querying are within a constant range, thus improving the index throughput rate in extreme cases. The data storage module adopts a tree structure and contains two types of nodes: buffer nodes and formal nodes. Buffer nodes are used to temporarily store a small amount of data. Internally, they are tightly arranged, storing data items in the form of an array or a linked list, and querying through binary search. This design helps prevent the rapid growth of the tree height caused by data conflicts, thereby improving the read and write throughput. Formal nodes are used to store data elements fitted by the dynamic fitting method described in Embodiment 1. The specific fitting process has been elaborated in detail in Embodiment 1 and will not be repeated here. Formal nodes store data items, pointer items, and empty items in the form of an array. Among them, data items are used to store data keys and corresponding data values; pointer items only point to child nodes and do not store data; empty items represent the empty spaces reserved when the node is created to reduce the probability of conflicts during subsequent insertions.
[0038] The internal elements of buffer nodes are arranged in order according to the data keys. When the number of keys exceeds the threshold, the node will be adjusted to a formal node. Formal nodes need to additionally store a bit array as a flag to distinguish element types and store model parameters to predict the data key position. The index of the present invention is constructed from top to bottom in a tree structure to ensure accurate prediction of the data key position by the model within the node and avoid additional time consumption caused by prediction errors. In addition, the tree structure will be dynamically adjusted according to internal data changes to achieve higher throughput efficiency.
[0039] The learning index construction method designed by the present invention can support both data insertion and query operations. When data is inserted into the index, the index method uses the data fitting algorithm described in Embodiment 1 for model training and performs local retraining and local index reconstruction at appropriate times to ensure an efficient data query process. The remaining part of this section will describe these key processes in detail in sequence.
[0040] Data insertion and retraining: When new data enters, it is accessed level by level starting from the root node. If the node type is a buffer node, the new data will be written into the node in order; if the node type is a formal node, the placement position of the new data is calculated according to the node model. If the position is empty, it is directly inserted; if the position is a pointer item, it enters the lower-level node for insertion; if the position is a data item and a data conflict occurs, a new buffer node needs to be created, and the old and new data are placed in order, and at the same time, the position is updated to a pointer pointing to the buffer node. After successful insertion, the new data key is written into the Bloom filter for query request filtering.
[0041] The present invention implements a node buffer to handle node creation and destruction operations frequently caused by data conflicts. During the initialization process, a batch of nodes are pre-created in memory. When a new node is needed, it is preferentially obtained from the buffer; when a node needs to be destroyed, only the internal data of the node is cleared, the memory is not released, and it is returned to the buffer.
[0042] As data is inserted, the model parameters within the node may gradually become invalid, and data conflicts increase. At this time, model retraining and node reconstruction are required. The index structure of the present invention performs model retraining in the following two cases: (1) The current node type is a buffer node, and the number of elements within the node exceeds a predetermined threshold ; (2) The current node type is a formal node, but the number of conflicts that have occurred within the node satisfies the condition of Equation (7).
[0043] (7) In the formula, represents the number of conflicts that have occurred in the current node and its subtree, represents the total number of elements in the node and its subtree, represents the initial number of elements when the node is established, is a preset parameter.
[0044] During the retraining process, first traverse and collect all data key values on the current node and its child nodes, calculate a new model using a fitting algorithm, and insert these data key value pairs into a new node with a length greater than the length of the original node . Then, point the pointer of the original node to the new node. The new model can better fit the data key set, most elements will be placed in the new node, and only a small part of the data enters the subtree, thereby improving the read and insert efficiency.
[0045] Retraining and tree reconstruction are the most time-consuming parts of the index insertion process. Through the time complexity analysis of each step of the algorithm, it can be seen that the time complexity of the learning index construction method based on dynamic fitting is , where represents the number of elements to be fitted, Indicates the number of inserted gap items. To reduce the impact of retraining and tree reconstruction on the overall performance, the learning index construction method based on dynamic fitting adopts two key designs: First, buffer nodes are introduced. When the amount of data is small, a tightly packed array is used instead of the subtree structure, which avoids retraining caused by the rapid growth of subtree height and frequent conflicts, thus significantly reducing the frequency of tree reconstruction and model training. Second, a bottom-up retraining mechanism is designed, that is, starting from the leaf nodes to traverse the formal nodes, and tree reconstruction and model retraining are only performed when necessary, minimizing the scope of model retraining and significantly reducing the performance overhead. Experimental results show that these two designs effectively avoid the negative impact of model retraining, making the overall throughput of the learning index construction method based on dynamic fitting significantly higher than that of existing index structures.
[0046] Query, as one of the core functions of the index, aims to quickly find the location of the input data key and return the corresponding data value according to it.
[0047] First, input the data key into the request filtering component BL. If BL returns False, it indicates that the data key does not exist in the current index, and the query process terminates immediately, thus avoiding the reduction of query throughput caused by invalid requests. If BL returns True, then enter the storage module for data query.
[0048] The query process starts from the index root node and performs the following operations: (1) If the node is a BN, use binary search inside the node to find the data key and obtain the corresponding data value.
[0049] (2) If the current node is an FN, first use the model in the FN to calculate the position of the data key in the node, and then access the element at that position. If the position is a data item, directly return the corresponding data value; if the position is a pointer item, enter the lower-level element pointed to by the pointer to continue the query; if the position is a gap item, it indicates that the key does not exist in the index structure and the query ends. Since the index adopts a tree structure and data insertion is performed in a top-down manner during the construction process, the accuracy of the model output position can be guaranteed.
[0050] Embodiment 4 Based on the same inventive concept, the present invention also provides an electronic device, including one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in Embodiment 1.
[0051] Since the device described in the fourth embodiment of the present invention is the electronic device used to implement the learning index construction method in the third embodiment of the present invention, based on the method described in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the electronic device, so it will not be elaborated here. Any electronic device used in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.
[0052] Embodiment Five Based on the same inventive concept, the present invention also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first embodiment is implemented.
[0053] Since the device described in the fifth embodiment of the present invention is the computer-readable medium used to implement the learning index construction in the third embodiment of the present invention, based on the method described in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the electronic device, so it will not be elaborated here. Any electronic device used in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.
[0054] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A dynamic fitting method for learning index construction, characterized in that: Calculate the average difference value of the data key set, and then calculate the difference value of the data pair in turn. According to the relationship between the difference value of the data pair and the average difference value, the data is divided into two groups: Data points whose difference between pairs of data is greater than the average difference are grouped together. , the elements in the group are stored in a minimum heap; the data points whose difference value of the data pair is less than the average difference value enter the group , the elements in the group are stored using a maximum heap, where a data pair refers to two data keys; Traverse the data key collection and select the group The smallest element in the data pair is inserted into a gap item between the data pairs, and then the average difference value is updated and the pairs are selected again until all the gap items are inserted. The pair with the smallest difference value in group B is removed. data points, of which is the system parameter, set to ,in Indicates the number of elements in the data key set; The data key set formed by the processed group A and group B is used to calculate the fitting straight line parameters using the least squares method to obtain the fitting model.
2. The dynamic fitting method according to claim 1, characterized in that: The data point is the point of the data key in the two-dimensional coordinate system. The least squares method is used to calculate the parameters of the fitting line. If The data points are represented as: , linear function For fitting data points, and is the parameter to be determined; define the objective function ,in from arrive , respectively and Find the partial derivative and set it to , we get the linear equation system: Solve the linear equations and get the analytical expression of the least squares method: Calculate the slope of the best fit line using the least squares method and the intercept , and represent the average of x and y respectively.
3. An evaluation method for evaluating the dynamic fitting method according to claim 1 or 2, characterized in that: The fitting indicators include: Conflict_ratio indicates the situation where conflicts occur after data is inserted according to the mapping address of the fitting model, and Empty_ratio indicates the number of addresses where no data is stored; Introduce data fitting evaluation index T, T = Conflict_ratio + Empty_ratio In the formula, conflict_num indicates the number of conflicts, element_num indicates the number of elements in the data key set, and empty_num indicates the number of empty items.
4. A learning index construction method based on dynamic fitting, characterized in that: In the tree structure data storage, nodes are divided into buffer nodes and formal nodes, the buffer nodes are used to temporarily store a limited number of data elements, and the formal nodes are used to store fitted data elements, wherein the fitting utilizes the dynamic fitting method described in claim 1 or 2, and the formal nodes use arrays to store data items, pointer items, and gap items; When inserting data, the buffer node is directly inserted. In the formal node, the gap item is directly inserted. The pointer item is inserted into the lower node. For the data item, a new buffer node is created for data insertion. The key of the successfully inserted data is written into the Bloom filter. The buffer node with data volume exceeding the threshold is adjusted to a formal node. When querying, after the Bloom filter returns a true value, the node type is determined, and then the data access method is determined based on the node type.
5. The method for constructing a learning index based on dynamic fitting according to claim 4, characterized in that: The buffer node is tightly arranged inside, and data items are stored in array or linked list form, and are queried through binary search.
6. The method for constructing a learning index based on dynamic fitting according to claim 4, characterized in that: The query process is as follows: If the node is a buffer node, use binary search inside the node to find the data key and obtain the corresponding data value; If the current node is a formal node, first use the fitting model in the formal node to calculate the position of the data key in the node, and then access the element at that position: if the position is a data item, directly return the corresponding data value; if the position is a pointer item, enter the lower-level element pointed to by the pointer to continue the query; if the position is a gap item, the query ends.
7. The method for constructing a learning index based on dynamic fitting according to claim 4, characterized in that: If the current node type is a buffer node and the number of elements in the node exceeds the preset threshold, the fitting model is retrained; The current node type is a formal node. If the number of conflicts that have occurred within the node satisfies the following conditions, the fitting model is retrained: In the formula, Indicates the number of conflicts that have occurred in the current node and its subtree. Represents the total number of elements in a node and its subtree, Indicates the initial number of elements when the node is created. are preset parameters.
8. The method for constructing a learning index based on dynamic fitting according to claim 7, characterized in that: During the retraining process, first traverse and collect all the data key values on the current node and its child nodes, use the dynamic fitting method to calculate the new fitting model, and insert the data key value pair into a node with a length greater than the original node length. The pointer of the original node points to the new node.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the learning index construction method as described in any one of claims 4-7.
10. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the learning index construction method as described in any one of claims 4 to 7 is implemented.