Data processing method and device, equipment, readable storage medium and program product
By mapping data sets into integers and using tree structure parameters for data processing, the problem of inefficiency of B+ trees and cardinality trees in big data scenarios is solved, and efficient data operation and memory management are achieved.
Patent Information
- Application Number
- CN202510373562.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-25
AI Technical Summary
In scenarios where data volume is large and insertion or deletion are frequent, the query efficiency is not high and the calculation overhead is too large, so it cannot effectively balance space and time use.
By obtaining the data quantity information of the data set and the target information of the target tree, determining the tree structure parameters of the target tree, mapping the target data in the data set into integers, data processing is performed based on the mapping and tree structure parameters, and data operations are realized using the index target tree.
It realizes simple and efficient calculations during data processing, reduces system overhead, and the time complexity of data operations is linear, and there is no need to adjust the inode and target tree structure when inserting and deleting data.
Smart Images

Figure CN120371830A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data storage, and particularly to a data processing method, apparatus, device, readable storage medium, and program product. Background Art
[0002] In a software system, the fast storage and query of data are important indicators of the software system performance. When the amount of data is huge, how to balance the use of space and time is an important issue in software technology development. Currently, in distributed storage, data is usually queried and stored through the tree structure of a B+ tree or a radix tree.
[0003] However, the B+ tree has a fast data query speed, but additional overhead is required for inserting and deleting data, which is not suitable for scenarios with a large amount of data and frequent data insertion or deletion. For a radix tree, the corresponding data structure needs to be selected according to the number of non-empty child nodes during each insertion and deletion operation, which is computationally complex and increases additional computational overhead, resulting in low query efficiency. Summary of the Invention
[0004] Based on this, it is necessary to provide a data processing method, apparatus, device, readable storage medium, and program product that can reduce system overhead for the above technical problems.
[0005] In a first aspect, this application provides a data processing method, including:
[0006] Obtain the data volume information of a data set, and obtain the target information of a target tree for storing the data set;
[0007] Determine the tree structure parameters of the target tree according to the data volume information and the target information;
[0008] Map the target data in the data set to an integer, and perform data processing on the target data based on the mapped integer and the tree structure parameters.
[0009] In the above embodiment, first, obtain the data volume information of the data set and the target information of the target tree for storing the data set. Then, determine the tree structure parameters of the target tree according to the data volume information and the target information. Finally, map the target data in the data set to an integer, and perform data processing on the target data based on the mapped integer and the tree structure parameters. In this way, the structure of the target tree for storing the data set is determined through the tree structure parameters. The target data in the data set is mapped to an integer, and the mapped integer is used as an index keyword. The data processing operation on the target data is realized by indexing the target tree. The index position of the target data can be determined only through numerical calculation, which is simple and efficient in calculation and reduces system overhead.
[0010] In one embodiment, the data volume information includes the maximum number of target data in the data set and the memory occupancy of the target data, and the target information includes the maximum memory space of the index nodes of the target tree and the maximum memory space of the data nodes of the target tree; according to the data volume information and the target information, determine the tree structure parameters of the target tree, including:
[0011] According to the maximum memory space of the index nodes, determine the maximum number of indexes stored in the index nodes;
[0012] According to the maximum memory space of the data nodes and the memory occupancy of the target data, determine the maximum number of data stored in the data nodes;
[0013] According to the maximum number of target data in the data set, the maximum number of indexes stored in the index nodes, and the maximum number of data stored in the data nodes, determine the depth of the target tree;
[0014] According to the maximum number of indexes stored in the index nodes, the maximum number of data stored in the data nodes, and the depth of the target tree, determine the step size of each layer of the target tree;
[0015] Take the maximum number of indexes stored in the index nodes, the maximum number of data stored in the data nodes, the depth of the target tree, and the step size of each layer of the target tree as the tree structure parameters of the target tree.
[0016] In the above embodiment, the tree structure parameters of the target tree can be obtained through calculation in the initialization stage, and the target tree can be constructed according to the tree structure parameters.
[0017] In one embodiment, perform data processing on the target data based on the integer obtained by mapping and the tree structure parameters, including:
[0018] Based on the integer obtained by mapping, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree, determine the data node corresponding to the target data;
[0019] According to the data node corresponding to the target data, perform data processing on the target data.
[0020] In the above embodiment, since the target data is mapped to an integer, the amount of data stored in the index nodes is small, and the step size of each layer is large, so the calculated target tree has a smaller tree depth and faster data processing speed.
[0021] In one embodiment, based on the integer obtained by mapping, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree, determine the data node corresponding to the target data, including:
[0022] Index the target tree layer by layer. During the indexing process of each layer, determine whether the currently indexed layer is less than the difference between the depth of the target tree and the preset threshold;
[0023] If so, determine the first index value of the target data in the currently indexed layer according to the integer obtained by mapping and the divisor of the step size corresponding to the currently indexed layer, determine the second index value of the target data in the currently indexed layer according to the integer obtained by mapping and the remainder of the step size corresponding to the currently indexed layer, determine the index node of the target data in the currently indexed layer according to the first index value and the memory address information corresponding to the root node of the target tree, and update the integer obtained by mapping according to the second index value;
[0024] If not, determine the data index value of the target data according to the integer obtained by mapping, and determine the data node corresponding to the target data according to the index node of the currently indexed layer and the data index value.
[0025] In the above embodiments, during the process of indexing data, the data node of the target data can be indexed through simple numerical calculations. The calculation process is simple and efficient, reducing the computational overhead of the system.
[0026] In one of the embodiments, data processing is performed on the target data according to the data node corresponding to the target data, including:
[0027] If the data processing is to add data, save the target data to the memory space of the data node corresponding to the target data;
[0028] If the data processing is to query data, return the data in the memory space of the data node corresponding to the target data;
[0029] If the data processing is to modify data, modify the data in the memory space of the data node corresponding to the target data according to the target data;
[0030] If the data processing is to delete data, delete the data in the memory space of the data node corresponding to the target data.
[0031] In this embodiment, operations such as adding data, querying data, modifying data, and deleting data all have a linear time complexity, and there is no need to adjust the index node data structure and the target tree structure when adding data and deleting data, resulting in higher efficiency.
[0032] In one of the embodiments, after deleting the data in the memory space of the data node corresponding to the target data, the method further includes:
[0033] Judge whether the upper-layer index node of the data node corresponding to the target data further includes other data nodes;
[0034] If not, delete the upper-layer index node of the data node corresponding to the target data.
[0035] In the above embodiments, after deleting the data of the data node, the data node and the index node without data are recycled by means of recursive traversal to realize the recycling of memory and improve the memory utilization rate. At the same time, after all the data nodes of the index node are deleted, the memory of the data node and the index node is recycled, reducing the number of memory application and release times.
[0036] In one of the embodiments, mapping the target data in the data set to an integer includes:
[0037] Mapping the target data to an integer through a bitmap;
[0038] After deleting the data in the memory space of the data node corresponding to the target data, the method further includes:
[0039] Reusing the position of the deleted target data in the bitmap data when adding data next time.
[0040] In a second aspect, the present application further provides a data processing device, including:
[0041] An acquisition module, configured to acquire the data volume information of the data set and acquire the target information of the target tree for storing the data set;
[0042] A determination module, configured to determine the tree structure parameters of the target tree according to the data volume information and the target information;
[0043] A processing module, configured to map the target data in the data set to an integer, and perform data processing on the target data based on the mapped integer and the tree structure parameters.
[0044] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the data processing method according to any one of the first aspects above.
[0045] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data processing method according to any one of the first aspects above.
[0046] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the data processing method according to any one of the first aspects above.
[0047] The above data processing method, apparatus, device, readable storage medium, and program product first obtain the data volume information of the data set and the target information of the target tree for storing the data set. Then, based on the data volume information and the target information, the tree structure parameters of the target tree are determined. Finally, the target data in the data set is mapped to an integer, and the target data is processed based on the mapped integer and the tree structure parameters. In this way, the structure of the target tree for storing the data set is determined by the tree structure parameters, the target data in the data set is mapped to an integer, and the mapped integer is used as the index keyword. The data processing operation of the target data is implemented by indexing the target tree. Only numerical calculations are required to determine the index position of the target data, the calculation is simple and efficient, and the system overhead is reduced. Description of the Drawings
[0048] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0049] Figure 1 It is a flowchart of the data processing method in an embodiment;
[0050] Figure 2 It is a flowchart of the steps for determining the tree structure parameters of the target tree in an embodiment;
[0051] Figure 3 It is a schematic diagram of the structure of the target tree in an embodiment;
[0052] Figure 4 It is a flowchart of the data processing method in another embodiment;
[0053] Figure 5 It is a flowchart of the data processing method in another embodiment;
[0054] Figure 6 It is a flowchart of the data processing method in another embodiment;
[0055] Figure 7 It is an index schematic diagram of the target tree in the data processing method in an embodiment;
[0056] Figure 8 It is a flowchart of the process of adding data in the data processing method in an embodiment;
[0057] Figure 9 It is a structural block diagram of the data processing apparatus in an embodiment;
[0058] Figure 10 It is the internal structure diagram of a computer device in an embodiment. Specific implementation manners
[0059] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] In an exemplary embodiment, as Figure 1 shown, a data processing method is provided. Taking the application of this method to a server as an example for illustration, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server, including the following steps:
[0061] Step 101, obtain the data volume information of the data set, and obtain the target information of the target tree for storing the data set.
[0062] Among them, the data set is a set composed of metadata that the system needs to process. The data volume information of the data set may include the number of metadata in the data set, the size of the metadata, etc. The target tree is a data structure for storing and indexing the data set. Optionally, the target tree can be a radix tree. The target tree may include index nodes and data nodes, and the target information may be the memory space information of each node in the target tree, such as the memory space information of each index node and the memory space information of each data node.
[0063] Step 102, determine the tree structure parameters of the target tree according to the data volume information and the target information.
[0064] Among them, the tree structure parameters of the target tree may include the maximum number of indexes stored in each index node, the maximum number of data stored in each data node, the depth of the target tree, and the step size of each layer of the target tree. The depth of the target tree is also the number of layers included in the target tree. According to the data volume information of the data set and the target information of the target tree, each tree structure parameter of the target tree can be obtained through calculation, and the target tree can be initialized according to the tree structure parameters.
[0065] Step 103, map the target data in the data set to an integer, and perform data processing on the target data based on the integer obtained by the mapping and the tree structure parameters.
[0066] The target data, namely each piece of metadata in the data set, is mapped to an integer. That is, a piece of metadata in the data set is mapped to obtain an integer, and different metadata can be mapped to different integers. The obtained integers are unique. Optionally, methods such as bitmaps can be used to map the target data to an integer, and the obtained integer is used as an index keyword to perform indexing in the target tree constructed according to the tree structure parameters, so as to implement data processing on the target data. Optionally, the target data can include multiple fields, such as immutable fields and data fields. An integer is mapped according to the immutable fields of the target data, and the change of the specific data part of the target data does not affect the mapped integer.
[0067] Optionally, the data processing operations can include adding data, that is, storing the target data into the target tree, and can also include modifying data, that is, indexing to the corresponding data node according to the integer to find the information of the target data, and modifying the content corresponding to the data node. It can also include querying data, that is, returning the content of the data node, and deleting data, that is, deleting the corresponding data node.
[0068] In the above embodiments, first, obtain the data volume information of the data set and obtain the target information of the target tree for storing the data set. Then, according to the data volume information and the target information, determine the tree structure parameters of the target tree. Finally, map the target data in the data set to an integer, and perform data processing on the target data based on the obtained integer and the tree structure parameters. In this way, the structure of the target tree for storing the data set is determined through the tree structure parameters, the target data in the data set is mapped to an integer, and based on the obtained integer as an index keyword, data processing operations on the target data are implemented through indexing the target tree. Only through numerical calculation can the index position of the target data be determined, the calculation is simple and efficient, and the system overhead is reduced.
[0069] In the embodiments of the present application, the data volume information includes the maximum number of target data in the data set and the memory occupancy of the target data, and the target information includes the maximum memory space of the index nodes of the target tree and the maximum memory space of the data nodes of the target tree. The above step 102 is as Figure 2 shown, and includes:
[0070] Step 201, determine the maximum number of indexes that the index node can store according to the maximum memory space of the index node.
[0071] Among them, the maximum memory space of the index node, that is, the maximum memory allowed to be used by a single index node, can be expressed in bytes. According to the maximum memory space of a single index node, the maximum number of indexes that a single index node can store is calculated according to the following formula.
[0072]
[0073] Among them, maxIndexNum is the maximum number of indexes stored in the index node, and maxIndexMem is the maximum memory space of the index node.
[0074] Step 202: Determine the maximum number of data stored in the data node according to the maximum memory space of the data node and the memory occupancy of the target data.
[0075] Among them, the maximum memory space of the data node, that is, the maximum memory allowed to be used by a single data node, can be expressed in bytes, and the memory occupancy of the target data is the memory space occupied by the target data, which can be expressed in bytes. The maximum number of data stored in a single data node is calculated according to the following formula.
[0076]
[0077] Among them, maxDataNum is the maximum number of data stored in the data node, maxDataMem is the maximum memory space of the data node, and singleMem is the memory occupancy of the target data.
[0078] Step 203: Determine the target tree depth according to the maximum number of target data in the data set, the maximum number of indexes stored in the index node, and the maximum number of data stored in the data node.
[0079] Among them, the target tree depth, that is, the number of layers of the target tree, is calculated according to the following formula. The target tree depth is the minimum value of i that satisfies the if conditional statement in the following formula.
[0080]
[0081]
[0082] Among them, is the target tree depth, is the maximum number of target data.
[0083] Step 204: Determine the step size of each layer of the target tree according to the maximum number of indexes stored in the index node, the maximum number of data stored in the data node, and the target tree depth.
[0084] Among them, the target tree structure can be as Figure 3 shown. The 0th layer of the target tree is the root node, and the step size of the root node is not calculated. steps[0] is used to represent the step size of the first layer, steps[1] is used to represent the step size of the second layer, and steps[depth - 2] is used to represent the step size of the last layer of data nodes. The step size of each layer can be calculated according to the following formula.
[0085]
[0086]
[0087] Step 205: Use the maximum number of indexes stored in the inode, the maximum number of data stored in the data node, the target tree depth, and the step size of each layer of the target tree as the tree structure parameters of the target tree.
[0088] Use the maximum number of indexes stored in the inode, the maximum number of data stored in the data node, the target tree depth, and the step size of each layer of the target tree calculated above as the tree structure parameters of the target tree. The target tree can be initialized according to the tree structure parameters. The initialization is only completed with the root node. Memory is allocated for each inode and data node during the process of adding data to the target tree. The root node and the inodes of each layer can index maxIndexNum addresses, and the data nodes of the last layer can index maxDataNum data.
[0089] In the above embodiments, the tree structure parameters of the target tree can be obtained through calculation during the initialization phase, and the target tree can be constructed according to the tree structure parameters.
[0090] In one embodiment, the steps of data processing on the target data based on the integer obtained by mapping and the tree structure parameters are as Figure 4 shown, and may include:
[0091] Step 401: Based on the integer obtained by mapping, the memory address information corresponding to the root node of the target tree, the target tree depth, and the step size of each layer of the target tree, determine the data node corresponding to the target data.
[0092] Among them, the memory address information corresponding to the root node of the target tree is the memory address of the root node. During the process of data processing on the target data, first use the integer obtained by mapping the target data as the index keyword, and index the data node corresponding to the target data according to the memory address information corresponding to the root node of the target tree, the target tree depth, and the step size of each layer of the target tree.
[0093] Optionally, the indexing process of the data node can be as Figure 5 shown, and includes:
[0094] Step 501: Index the target tree layer by layer. During the indexing of each layer, determine whether the currently indexed layer is less than the difference between the target tree depth and the preset threshold.
[0095] Among them, the preset threshold is 2. Start indexing the target tree layer by layer from the root node of layer 0. When indexing each layer, determine whether the currently indexed layer is less than the target tree depth minus 2. If it is less, execute step 502. If it is not less, it is the upper layer index node layer of the data node layer, and execute step 503.
[0096] Step 502: If so, determine the first index value of the target data in the layer corresponding to the current index based on the integer obtained by mapping and the divisor of the step size corresponding to the layer of the current index, determine the second index value of the target data in the layer corresponding to the current index based on the integer obtained by mapping and the remainder of the step size corresponding to the layer of the current index, determine the index node of the target data in the layer corresponding to the current index based on the first index value and the memory address information corresponding to the root node of the target tree, and update the integer obtained by mapping according to the second index value.
[0097] When indexing the 0th layer, at this time, the integer obtained by mapping is the integer obtained by mapping the target data. Determine the first index value of the target data in the layer corresponding to the current index based on the integer obtained by mapping and the divisor of steps[0], that is, the step size corresponding to the first layer. The index node of the first layer can be determined according to the first index value. If the index node of the first layer is empty, apply for memory for the index node according to the memory address information corresponding to the root node of the target tree and update the index node. The index node stores pointer information for indicating the next index address.
[0098] Then, determine the second index value of the target data in the layer corresponding to the current index based on the integer obtained by mapping and the remainder of steps[0], and update the integer obtained by mapping according to the second index value. When indexing the next layer, the integer obtained by mapping is the integer updated in the previous layer. After the update is completed, increment the layer number by one, perform indexing of the next layer, and continue to repeat the judgment process of step 501.
[0099] Step 503: If not, determine the data index value of the target data according to the integer obtained by mapping, and determine the data node corresponding to the target data according to the index node of the layer corresponding to the current index and the data index value.
[0100] Optionally, if the layer corresponding to the current index is not less than the target tree depth minus 2, determine the data index value of the target data according to the integer obtained by mapping updated in the above indexing process, determine the data node corresponding to the target data according to the data index value, and apply for memory for the data node if the data node is empty.
[0101] In the above embodiment, during the process of indexing data, the data node of the target data can be indexed through simple numerical calculations. The calculation process is simple and efficient, reducing the calculation overhead of the system.
[0102] Step 402: Process the target data according to the data node corresponding to the target data.
[0103] The actual location where the target data is stored can be determined according to the corresponding data node, so that data processing operations can be performed on the target data.
[0104] In the above embodiments, since the target data is mapped to an integer, the amount of data stored in the index node is small, and the step size of each layer is large. Therefore, the calculated target tree has a smaller tree depth, and the data processing speed is faster.
[0105] In an embodiment of the present application, the data processing operation may include adding data, querying data, modifying data, and deleting data. According to the data node corresponding to the target data, data processing on the target data includes the following situations.
[0106] In the first situation, if the data processing is to add data, the target data is saved to the memory space of the data node corresponding to the target data.
[0107] In the second situation, if the data processing is to query data, the data in the memory space of the data node corresponding to the target data is returned.
[0108] In the third situation, if the data processing is to modify data, the data in the memory space of the data node corresponding to the target data is modified according to the target data.
[0109] In the fourth situation, if the data processing is to delete data, the data in the memory space of the data node corresponding to the target data is deleted.
[0110] Among them, as described in the above embodiments, the target data can be mapped to an integer through a bitmap. After deleting the data in the memory space of the data node corresponding to the target data, the method further includes: reusing the position of the deleted target data in the bitmap data when adding data next time.
[0111] After deleting the data, the bitmap data during the mapping of the target data is modified at the same time, so that this position can be reused when adding data next time, thereby realizing the reusability of the memory of the data node and saving memory space.
[0112] In this embodiment, operations such as adding data, querying data, modifying data, and deleting data all have a linear time complexity, and there is no need to adjust the index node data structure and the target tree structure when adding data and deleting data, so the efficiency is higher.
[0113] In one embodiment, after deleting the data in the memory space of the data node corresponding to the target data, the method further includes: determining whether the upper-layer index node corresponding to the data node of the target data further includes other data nodes; if not, deleting the upper-layer index node corresponding to the data node of the target data.
[0114] After deleting the target data, recursively traverse the data node corresponding to the target data, that is, determine whether the index node on the upper layer of the data node of the target data still includes other data nodes. If all the data nodes of the index node on the upper layer are deleted, that is, no other data nodes are included, then delete the upper layer index node, and at the same time recursively traverse the index node on the upper layer of this index node for judgment. If all the nodes under all the index nodes are deleted, then all the index nodes can be deleted to release all data memory.
[0115] In the above embodiment, after deleting the data of the data node, the data node and index node without data are recycled by means of recursive traversal to realize the recycling of memory and improve the memory utilization rate. At the same time, when all the data nodes of the index node are deleted, the memory of the data node and index node is recycled, reducing the number of memory application and release times.
[0116] In the embodiment of the present application, as Figure 6 shown, a data processing method is provided, including:
[0117] Step 601, obtain the data volume information of the data set, and obtain the target information of the target tree for storing the data set.
[0118] Step 602, determine the maximum number of indexes that the index node can store according to the maximum memory space of the index node.
[0119] Step 603, determine the maximum number of data that the data node can store according to the maximum memory space of the data node and the memory occupancy of the target data.
[0120] Step 604, determine the depth of the target tree according to the maximum number of target data in the data set, the maximum number of indexes that the index node can store, and the maximum number of data that the data node can store.
[0121] Step 605, determine the step size of each layer of the target tree according to the maximum number of indexes that the index node can store, the maximum number of data that the data node can store, and the depth of the target tree.
[0122] Step 606, take the maximum number of indexes that the index node can store, the maximum number of data that the data node can store, the depth of the target tree, and the step size of each layer of the target tree as the tree structure parameters of the target tree.
[0123] Step 607, determine the data node corresponding to the target data based on the integer obtained by mapping, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree.
[0124] Step 608, perform data processing on the target data according to the data node corresponding to the target data.
[0125] In an embodiment of the present application, a data processing method is provided. As Figure 7 shown, it is a schematic flowchart of adding data in this data processing method. Figure 8 It is a schematic diagram of the addresses that each index node can index and the data that the data node can index. The root node and the index nodes of each layer can index maxIndexNum addresses, and the data nodes of the last layer can index maxDataNum data. The indexing process of num obtained according to the mapping can calculate the index value of each layer according to the following method.
[0126] (1) Root node index = num / steps[0]
[0127] (2) First layer index = (num % steps[0]) / steps[1]
[0128] (3) The calculation method of the middle layer is the same
[0129] Index of the i-th layer = (num % steps[0] % steps[1]… % steps[depth - 3]) / steps[depth - 2]
[0130] (4) Data layer index = num % steps[0] % steps[1]… % steps[depth - 2]
[0131] In this embodiment, the calculated target tree has a relatively small tree depth. Operations such as adding data, querying data, modifying data, and deleting data all have a linear time complexity, and there is no need to adjust the index node data structure and the target tree structure when adding and deleting data, so the efficiency is higher.
[0132] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0133] Based on the same inventive concept, an embodiment of this application further provides a data processing apparatus for implementing the data processing method involved above. The solution provided by this apparatus for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the data processing apparatus provided below can refer to the limitations on the data processing method in the foregoing, and will not be elaborated herein.
[0134] In an exemplary embodiment, as Figure 9 shown, a data processing apparatus 900 is provided, including: an acquisition module 901, a determination module 902, and a processing module 903, where:
[0135] The acquisition module 901 is configured to acquire the data volume information of the data set and acquire the target information of the target tree for storing the data set;
[0136] The determination module 902 is configured to determine the tree structure parameters of the target tree according to the data volume information and the target information;
[0137] The processing module 903 is configured to map the target data in the data set to an integer, and perform data processing on the target data based on the mapped integer and the tree structure parameters.
[0138] In one of the embodiments, the data volume information includes the maximum number of target data in the data set and the memory occupancy of the target data, and the target information includes the maximum memory space of the index nodes of the target tree and the maximum memory space of the data nodes of the target tree; specifically, the determination module 902 is configured to determine the maximum number of indexes stored in the index nodes according to the maximum memory space of the index nodes; determine the maximum number of data stored in the data nodes according to the maximum memory space of the data nodes and the memory occupancy of the target data; determine the depth of the target tree according to the maximum number of target data in the data set, the maximum number of indexes stored in the index nodes, and the maximum number of data stored in the data nodes; determine the step size of each layer of the target tree according to the maximum number of indexes stored in the index nodes, the maximum number of data stored in the data nodes, and the depth of the target tree; and use the maximum number of indexes stored in the index nodes, the maximum number of data stored in the data nodes, the depth of the target tree, and the step size of each layer of the target tree as the tree structure parameters of the target tree.
[0139] In one of the embodiments, specifically, the processing module 903 is configured to determine the data node corresponding to the target data based on the mapped integer, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree; and perform data processing on the target data according to the data node corresponding to the target data.
[0140] In one embodiment, the processing module 903 is specifically configured to perform layer-by-layer indexing on the target tree. During the indexing of each layer, it is determined whether the currently indexed layer is less than the difference between the depth of the target tree and the preset threshold. If so, according to the integer obtained by mapping and the divisor of the step size corresponding to the currently indexed layer, the first index value of the target data in the currently indexed layer is determined. According to the integer obtained by mapping and the remainder of the step size corresponding to the currently indexed layer, the second index value of the target data in the currently indexed layer is determined. According to the first index value and the memory address information corresponding to the root node of the target tree, the index node of the target data in the currently indexed layer is determined. According to the second index value, the integer obtained by mapping is updated. If not, according to the integer obtained by mapping, the data index value of the target data is determined, and according to the index node of the currently indexed layer and the data index value, the data node corresponding to the target data is determined.
[0141] In one embodiment, the processing module 903 is specifically configured to, if the data processing is adding data, save the target data to the memory space of the data node corresponding to the target data; if the data processing is querying data, return the data in the memory space of the data node corresponding to the target data; if the data processing is modifying data, modify the data in the memory space of the data node corresponding to the target data according to the target data; if the data processing is deleting data, delete the data in the memory space of the data node corresponding to the target data.
[0142] In one embodiment, the apparatus further includes a deletion module, configured to determine whether the upper-layer index node of the data node corresponding to the target data further includes other data nodes; if not, delete the upper-layer index node of the data node corresponding to the target data.
[0143] In one embodiment, the processing module is specifically configured to map the target data to an integer through a bitmap; when adding data next time, reuse the positions of the deleted target data in the bitmap data.
[0144] Each module in the above data processing apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0145] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 10As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data in the data set. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a data processing method.
[0146] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0147] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: obtaining the data volume information of the data set, and obtaining the target information of the target tree for storing the data set; determining the tree structure parameters of the target tree according to the data volume information and the target information; mapping the target data in the data set to an integer, and performing data processing on the target data based on the mapped integer and the tree structure parameters.
[0148] In one embodiment, the data volume information includes the maximum number of target data in the data set and the memory occupancy of the target data, and the target information includes the maximum memory space of the index nodes of the target tree and the maximum memory space of the data nodes of the target tree; when the processor executes the computer program, the following steps are further implemented: determining the maximum number of indexes stored in the index nodes according to the maximum memory space of the index nodes; determining the maximum number of data stored in the data nodes according to the maximum memory space of the data nodes and the memory occupancy of the target data; determining the depth of the target tree according to the maximum number of target data in the data set, the maximum number of indexes stored in the index nodes, and the maximum number of data stored in the data nodes; determining the step size of each layer of the target tree according to the maximum number of indexes stored in the index nodes, the maximum number of data stored in the data nodes, and the depth of the target tree; and using the maximum number of indexes stored in the index nodes, the maximum number of data stored in the data nodes, the depth of the target tree, and the step size of each layer of the target tree as the tree structure parameters of the target tree.
[0149] In one embodiment, when the processor executes the computer program, the following steps are further implemented: determining the data node corresponding to the target data based on the integer obtained by mapping, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree; and performing data processing on the target data according to the data node corresponding to the target data.
[0150] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing layer-by-layer indexing on the target tree, and during the indexing of each layer, determining whether the current indexed layer is less than the difference between the depth of the target tree and the preset threshold; if so, determining the first index value of the target data in the current indexed layer according to the integer obtained by mapping and the divisor of the step size corresponding to the current indexed layer, determining the second index value of the target data in the current indexed layer according to the integer obtained by mapping and the remainder of the step size corresponding to the current indexed layer, determining the index node of the current indexed layer of the target data according to the first index value and the memory address information corresponding to the root node of the target tree, and updating the integer obtained by mapping according to the second index value; if not, determining the data index value of the target data according to the integer obtained by mapping, and determining the data node corresponding to the target data according to the index node of the current indexed layer and the data index value.
[0151] In one embodiment, when the processor executes the computer program, the following steps are further implemented: if the data processing is to add data, saving the target data to the memory space of the data node corresponding to the target data; if the data processing is to query data, returning the data in the memory space of the data node corresponding to the target data; if the data processing is to modify data, modifying the data in the memory space of the data node corresponding to the target data according to the target data; and if the data processing is to delete data, deleting the data in the memory space of the data node corresponding to the target data.
[0152] In one embodiment, when the processor executes the computer program, the following steps are further implemented: determining whether there are other data nodes in the upper-level index node corresponding to the target data; if not, deleting the upper-level index node corresponding to the target data.
[0153] In one embodiment, when the processor executes the computer program, the following steps are further implemented: mapping the target data to an integer through a bitmap; and reusing the position of the deleted target data in the bitmap data when adding data next time.
[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining the data volume information of the data set, and obtaining the target information of the target tree for storing the data set; determining the tree structure parameters of the target tree according to the data volume information and the target information; mapping the target data in the data set to an integer, and performing data processing on the target data based on the mapped integer and the tree structure parameters.
[0155] In one embodiment, the data volume information includes the maximum number of target data in the data set and the memory occupancy of the target data, and the target information includes the maximum memory space of the index nodes of the target tree and the maximum memory space of the data nodes of the target tree; when the computer program is executed by the processor, the following steps are further implemented: determining the maximum number of indexes stored in the index node according to the maximum memory space of the index node; determining the maximum number of data stored in the data node according to the maximum memory space of the data node and the memory occupancy of the target data; determining the depth of the target tree according to the maximum number of target data in the data set, the maximum number of indexes stored in the index node, and the maximum number of data stored in the data node; determining the step size of each layer of the target tree according to the maximum number of indexes stored in the index node, the maximum number of data stored in the data node, and the depth of the target tree; and using the maximum number of indexes stored in the index node, the maximum number of data stored in the data node, the depth of the target tree, and the step size of each layer of the target tree as the tree structure parameters of the target tree.
[0156] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented: determining the data node corresponding to the target data based on the mapped integer, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree; and performing data processing on the target data according to the data node corresponding to the target data.
[0157] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: index the target tree layer by layer. During the indexing of each layer, determine whether the currently indexed layer is less than the difference between the depth of the target tree and a preset threshold; if so, determine the first index value of the target data in the currently indexed layer according to the integer obtained by mapping and the divisor of the step size corresponding to the currently indexed layer, determine the second index value of the target data in the currently indexed layer according to the integer obtained by mapping and the remainder of the step size corresponding to the currently indexed layer, determine the index node of the target data in the currently indexed layer according to the first index value and the memory address information corresponding to the root node of the target tree, and update the integer obtained by mapping according to the second index value; if not, determine the data index value of the target data according to the integer obtained by mapping, and determine the data node corresponding to the target data according to the index node of the currently indexed layer and the data index value.
[0158] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: if the data processing is to add data, save the target data to the memory space of the data node corresponding to the target data; if the data processing is to query data, return the data in the memory space of the data node corresponding to the target data; if the data processing is to modify data, modify the data in the memory space of the data node corresponding to the target data according to the target data; if the data processing is to delete data, delete the data in the memory space of the data node corresponding to the target data.
[0159] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: determine whether the upper-layer index node of the data node corresponding to the target data further includes other data nodes; if not, delete the upper-layer index node of the data node corresponding to the target data.
[0160] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: map the target data to an integer through a bitmap; reuse the position of the deleted target data in the bitmap data when adding data next time.
[0161] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor, implements the following steps: obtain the data volume information of a data set, and obtain the target information of a target tree for storing the data set; determine the tree structure parameters of the target tree according to the data volume information and the target information; map the target data in the data set to an integer, and perform data processing on the target data based on the integer obtained by mapping and the tree structure parameters.
[0162] In one embodiment, the data volume information includes the maximum number of target data in the data set and the memory occupancy of the target data, and the target information includes the maximum memory space of the index nodes of the target tree and the maximum memory space of the data nodes of the target tree; when the computer program is executed by a processor, the following steps are further implemented: determining the maximum number of indexes stored by the index nodes according to the maximum memory space of the index nodes; determining the maximum number of data stored by the data nodes according to the maximum memory space of the data nodes and the memory occupancy of the target data; determining the depth of the target tree according to the maximum number of target data in the data set, the maximum number of indexes stored by the index nodes, and the maximum number of data stored by the data nodes; determining the step size of each layer of the target tree according to the maximum number of indexes stored by the index nodes, the maximum number of data stored by the data nodes, and the depth of the target tree; and using the maximum number of indexes stored by the index nodes, the maximum number of data stored by the data nodes, the depth of the target tree, and the step size of each layer of the target tree as the tree structure parameters of the target tree.
[0163] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: determining the data node corresponding to the target data based on the integer obtained by mapping, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree; and performing data processing on the target data according to the data node corresponding to the target data.
[0164] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing layer-by-layer indexing on the target tree, and during the indexing of each layer, determining whether the current indexed layer is less than the difference between the depth of the target tree and a preset threshold; if so, determining the first index value of the target data in the current indexed layer according to the integer obtained by mapping and the divisor of the step size corresponding to the current indexed layer, determining the second index value of the target data in the current indexed layer according to the integer obtained by mapping and the remainder of the step size corresponding to the current indexed layer, determining the index node of the current indexed layer of the target data according to the first index value and the memory address information corresponding to the root node of the target tree, and updating the integer obtained by mapping according to the second index value; if not, determining the data index value of the target data according to the integer obtained by mapping, and determining the data node corresponding to the target data according to the index node of the current indexed layer and the data index value.
[0165] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: if the data processing is to add data, saving the target data to the memory space of the data node corresponding to the target data; if the data processing is to query data, returning the data in the memory space of the data node corresponding to the target data; if the data processing is to modify data, modifying the data in the memory space of the data node corresponding to the target data according to the target data; and if the data processing is to delete data, deleting the data in the memory space of the data node corresponding to the target data.
[0166] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: determining whether there are other data nodes in the upper-layer index node corresponding to the target data; if not, deleting the upper-layer index node corresponding to the target data.
[0167] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: mapping the target data to an integer through a bitmap; and reusing the position of the deleted target data in the bitmap data when adding data next time.
[0168] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0169] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0170] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0171] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining the data volume information of the data set, and obtaining the target information of the target tree for storing the data set; Determining the tree structure parameters of the target tree according to the data volume information and the target information; Mapping the target data in the data set to an integer, and performing data processing on the target data based on the mapped integer and the tree structure parameters.
2. The method according to claim 1, wherein The data volume information includes the maximum number of the target data in the data set and the memory occupancy of the target data, and the target information includes the maximum memory space of the index nodes of the target tree and the maximum memory space of the data nodes of the target tree; The determining the tree structure parameters of the target tree according to the data volume information and the target information includes: Determining the maximum number of indexes stored in the index node according to the maximum memory space of the index node; Determining the maximum number of data stored in the data node according to the maximum memory space of the data node and the memory occupancy of the target data; Determining the depth of the target tree according to the maximum number of the target data in the data set, the maximum number of indexes stored in the index node, and the maximum number of data stored in the data node; Determining the step size of each layer of the target tree according to the maximum number of indexes stored in the index node, the maximum number of data stored in the data node, and the depth of the target tree; Taking the maximum number of indexes stored in the index node, the maximum number of data stored in the data node, the depth of the target tree, and the step size of each layer of the target tree as the tree structure parameters of the target tree.
3. The method according to claim 2, wherein The performing data processing on the target data based on the mapped integer and the tree structure parameters includes: Determining the data node corresponding to the target data based on the mapped integer, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree; Performing data processing on the target data according to the data node corresponding to the target data.
4. The method according to claim 3, characterized in that The determining the data node corresponding to the target data based on the mapped integer, the memory address information corresponding to the root node of the target tree, the depth of the target tree, and the step size of each layer of the target tree includes: Performing layer-by-layer indexing on the target tree, and during the indexing of each layer, determining whether the current index layer is less than the difference between the depth of the target tree and the preset threshold; If so, determining the first index value of the target data in the current index layer according to the mapped integer and the divisor of the step size corresponding to the current index layer, determining the second index value of the target data in the current index layer according to the mapped integer and the remainder of the step size corresponding to the current index layer, determining the index node of the current index layer of the target data according to the first index value and the memory address information corresponding to the root node of the target tree, and updating the mapped integer according to the second index value; If not, determining the data index value of the target data according to the mapped integer, and determining the data node corresponding to the target data according to the index node and the data index value of the current index layer.
5. The method according to claim 3, wherein Performing data processing on the target data according to the data node corresponding to the target data includes: If the data processing is adding data, saving the target data to the memory space of the data node corresponding to the target data; If the data processing is querying data, returning the data in the memory space of the data node corresponding to the target data; If the data processing is modifying data, modifying the data in the memory space of the data node corresponding to the target data according to the target data; If the data processing is deleting data, deleting the data in the memory space of the data node corresponding to the target data.
6. The method according to claim 5, wherein After deleting the data in the memory space of the data node corresponding to the target data, the method further includes: Determining whether the upper-level index node of the data node corresponding to the target data further includes other data nodes; If not, deleting the upper-level index node of the data node corresponding to the target data.
7. The method according to claim 5, wherein Mapping the target data in the data set to an integer includes: Mapping the target data to an integer through a bitmap; After deleting the data in the memory space of the data node corresponding to the target data, the method further includes: Reusing the position of the deleted target data in the bitmap data when adding data next time.
8. A data processing device, characterized in that, The apparatus includes: An acquisition module, configured to acquire the data volume information of the data set and acquire the target information of the target tree for storing the data set; A determination module, configured to determine the tree structure parameters of the target tree according to the data volume information and the target information; A processing module, configured to map the target data in the data set to an integer, and perform data processing on the target data based on the mapped integer and the tree structure parameters.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.