Data storage method and device and storage medium
By specifying multiple tags as sorting keys and determining the storage order of data, the problem of poor local order in the database is solved, and query efficiency is improved.
Patent Information
- Application Number
- PCT/CN2024/105142
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2024-07-12
- Publication Date
- 2025-05-30
AI Technical Summary
The local order of the data stored in the database is poor, resulting in inefficient query.
By specifying a plurality of tags as sorting keys, and determining the storage order of data based on the order of the sorting keys, identification information of the set of tag values to which the tag values of the first tag belong in each piece of data, and label values of the second tag, thereby improving the local order of data.
It improves the orderly storage of data in the database, reduces the data range required for query, and improves query efficiency.
Smart Images

Figure CN2024105142_30052025_PF_FP_ABST
Abstract
Description
Data storage method, device and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202311558370.2, filed on November 21, 2023, entitled “A Quantitative Cluster Indexing Method,” the entire contents of which are hereby incorporated by reference into this application. This application also claims priority to Chinese Patent Application No. 202410372304.4, filed on March 27, 2024, entitled “Data Storage Method, Apparatus, and Storage Medium,” the entire contents of which are hereby incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communications, and in particular to a method, device, and storage medium for storing data. Background Art
[0003] A database is a warehouse for storing data. The storage space of a database is often very large and can store millions, tens of millions, or even hundreds of millions of data items.
[0004] Each piece of data stored in the database includes a set of tag values for multiple tags. For a particular tag among these multiple tags, the tag value may include a large number of different values. When the same tag value is stored in a relatively scattered manner, the local order of the data stored in the database will be poor.
[0005] Poor local ordering of data stored in a database can affect data processing operations. For example, if a database query contains data containing a tag value, the poor local ordering of the data will result in a large data range being scanned, making the query inefficient. Therefore, improving the orderliness of data stored in a database is an urgent issue.
[0006] Summary of the Invention
[0007] This application provides a method, device, and storage medium for storing data to improve the orderliness of data stored in a database. The technical solution is as follows:
[0008] In a first aspect, the present application provides a method for storing data, in which a plurality of data pieces are received, wherein each data piece includes tag values of a plurality of tags, the plurality of tags including a first tag and a second tag designated as a sort key, the first tag being in an order before the second tag in the sort key. The tag value set to which the tag value of the first tag in each data piece belongs is determined. The storage order of the plurality of data pieces is determined based on the order of the sort keys, identification information of the tag value set to which the tag value of the first tag in each data piece belongs, and the tag value of the second tag in each data piece. The plurality of data pieces are stored based on the storage order of the plurality of data pieces.
[0009] For each piece of data, the tag value set to which the tag value of the first tag in the data belongs includes not only the tag value but also other tag values of the first tag, so that the cardinality of the tag value set to which the tag value of the first tag in each piece of data belongs is much smaller than the cardinality of the tag value of the first tag in each piece of data. Since the order of the first tag in the sorting key is before the second tag, the storage order of the multiple pieces of data is determined based on the order of the sorting key, the identification information of the tag value set to which the tag value of the first tag in each piece of data belongs, and the tag value of the second tag in each piece of data. After storing the multiple pieces of data in the storage order of the multiple pieces of data, the local orderliness of the tag values of the second tags in the multiple pieces of data can be improved, that is, the orderliness of the data stored in the database can be improved.
[0010] In one possible implementation, data structure information is received, the data structure information being used to specify a first tag and a second tag as sort keys. The data structure information includes quantization identification information, the quantization identification information being used to indicate the first tag. Based on the quantization identification information, it is possible to determine which tag value in each piece of data needs to be quantized, thereby determining the tag value set to which the tag value belongs.
[0011] In another possible implementation, based on the granularity of the tag value set and the tag value of the first tag in each piece of data, the tag value set to which the tag value of the first tag in each piece of data belongs is determined, wherein the quantitative identification information is also used to indicate the granularity of the tag value set, or the granularity of the tag value set is determined based on the characteristics of the tag value of the first tag in the stored data, thereby improving the flexibility of obtaining the granularity of the tag value set.
[0012] In another possible implementation, tags requiring quantization are determined. Tags requiring quantization are tags whose tag value cardinality satisfies a high-cardinality condition among the multiple tags. When a first tag is a tag requiring quantization, the tag value set to which the tag value of the first tag in each piece of data belongs is determined. Because tags whose tag value cardinality satisfies the high-cardinality condition are considered tags requiring quantization, tags whose tag value cardinality does not meet the high-cardinality condition are not considered tags requiring quantization. This reduces the number of tags requiring quantization and conserves computing resources.
[0013] In another possible implementation, the tag value set to which the tag value of the first tag in each piece of data belongs is a tag value range that includes the tag value of the first tag in each piece of data, and the granularity of the tag value set is the length of the tag value range.
[0014] In another possible implementation, index information is constructed for the multiple pieces of data. The index information includes identification information for the tag value set to which the tag value of the first tag in at least one selected piece of data belongs, wherein two adjacent selected pieces of data are separated by one or more pieces of data in the storage order. In this way, when querying data, the index information can be used to determine the data range to be scanned, and queries within this data range can be performed, which can improve query efficiency compared to searching the entire database.
[0015] In another possible implementation, a query request is received, the query request including a query condition related to a first tag. Data to be scanned from the plurality of data items is determined based on the query condition and identification information of the tag value set to which the tag value of the first tag in the index information belongs. The data to be scanned is queried to obtain query results that meet the query condition. Determining the data to be scanned from the plurality of data items based on the query condition and identification information of the tag value set to which the tag value of the first tag in the index information belongs can reduce the scope of data to be scanned, thereby improving query efficiency.
[0016] In a second aspect, the present application provides a device for storing data, configured to execute the method in the first aspect or any possible implementation of the first aspect. Specifically, the device includes a unit configured to execute the method in the first aspect or any possible implementation of the first aspect.
[0017] In a third aspect, the present application provides a computing device cluster, the computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory;
[0018] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method in the first aspect or any possible implementation manner of the first aspect.
[0019] In a fourth aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes the computing device cluster to execute the method in the first aspect or any possible implementation of the first aspect.
[0020] In a fifth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method in the first aspect or any possible implementation of the first aspect.
[0021] In a sixth aspect, the present application provides a chip comprising a memory and a processor, wherein the memory is used to store computer instructions, and the processor is used to call and run the computer instructions from the memory to execute the method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] FIG1 is a schematic diagram of a network architecture provided by an embodiment of the present application;
[0023] FIG2 is a flow chart of a method for storing data provided in an embodiment of the present application;
[0024] FIG3 is a schematic diagram of a first editing interface provided in an embodiment of the present application;
[0025] FIG4 is a schematic diagram of another second editing interface provided in an embodiment of the present application;
[0026] FIG5 is a flow chart of a method for querying data provided by an embodiment of the present application;
[0027] FIG6 is a schematic diagram of the structure of a data storage device provided in an embodiment of the present application;
[0028] FIG7 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0029] FIG8 is a schematic diagram of a cluster structure for storing data provided in an embodiment of the present application;
[0030] FIG9 is a schematic diagram of another cluster structure for storing data provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] 1 , an embodiment of the present application provides a network architecture 100 , which includes a terminal device 101 and a database system 102 , and the terminal device 101 communicates with the database system 102 .
[0032] A technician may configure multiple tags and data structure information of the database on the terminal device 101 , where the data structure information is used to specify at least two tags among the multiple tags as sorting keys.
[0033] That is, the sort key includes multiple tags that need to be sorted, and the multiple tags are in order in the sort key. For example, the multiple tags in the sort key include a first tag and a second tag, and the first tag is in an order before the second tag in the sort key.
[0034] In some embodiments, the data structure information further includes quantization identification information, which is used to indicate the first tag. Optionally, the quantization identification information indicates that the tag value of the first tag needs to be quantized to obtain identification information of the tag value set to which the tag value of the first tag belongs.
[0035] The terminal device 101 is used to obtain the configured multiple tags and the data structure information, and send a database creation request to the database system 102, where the database creation request includes the multiple tags and the data structure information.
[0036] The database system 102 is configured to receive the database creation request and create a database based on the multiple tags included in the database creation request.
[0037] The database is used to store multiple pieces of data, each piece of data including tag values of the configured multiple tags.
[0038] In some embodiments, the database may be a data table that uses a column storage method to store the multiple pieces of data. For each piece of data in the multiple pieces of data, the piece of data may be a row of data in the data table.
[0039] The data table includes a plurality of columns, and the plurality of columns included in the data table correspond to the plurality of configured tags.
[0040] Optionally, the multiple configured tags correspond to the multiple columns one-to-one, and for each tag in the multiple configured tags, the column corresponding to the tag is used to store one or more tag values belonging to the tag.
[0041] In some embodiments, the database may be a file system, and the database may use files to store the plurality of data.
[0042] In some embodiments, the terminal device 101 may also send a storage request to the database system 102 , where the storage request includes multiple pieces of data to be stored. The database system 102 receives the storage request and saves the multiple pieces of data into the database.
[0043] In some embodiments, the data structure information defines a sort key, which includes multiple tags to be sorted. This allows for multi-level sorting according to the order of the multiple tags defined by the sort key. After receiving the storage request, the database system 102 performs multi-level sorting on the multiple data items based on the tag values of each of the multiple tags to be sorted included in each of the multiple data items, thereby obtaining a storage order for the multiple data items. Based on the storage order of the multiple data items, the multiple data items are stored in the database.
[0044] Specifying a tag as a sort key for data structure information means sorting the tag values of that tag. Specifying which tags are sort keys means, after receiving multiple pieces of data to be stored, determining the storage order of those pieces of data based on the tag values of those tags serving as sort keys. In implementation, since a piece of data is a whole, sorting the tag values of the tag serving as the sort key actually sorts the multiple pieces of data as a whole. Furthermore, when the sort key has multiple tags, the order of the tags in the sort key determines which tag value is used to sort the multiple pieces of data first.
[0045] When the sorting key has multiple labels that need to be sorted, multi-level sorting is performed in the order of the multiple labels. For example, the order of the first label in the sorting key is before the second label. After receiving multiple pieces of data to be stored, the database system 102 sorts based on the label value of the first label included in each piece of data, obtains the storage order of the multiple pieces of data, and classifies the multiple pieces of data including the label value of the same first label into one category after sorting. When sorting based on the label value of the second label included in each piece of data, the multiple pieces of data classified into one category (including multiple pieces of data with the same label value of the first label) are sorted, and the relative order between different categories does not change. Therefore, when performing multi-level sorting, first perform a first-level sorting according to the label value of the first sorting key, and then perform a second-level sorting according to the label value of the second sorting key. The second level is to sort the data in the class that has been sorted in the first level again.
[0046] For example, the number of tags to be sorted included in the sort key may be n, where n is an integer greater than or equal to 1.
[0047] For the first of the n tags, the database system 102 performs a primary sort on each piece of data based on the tag value of the first tag to be sorted included in each piece of data, thereby grouping data with the same tag value of the first tag into a single category. For the second of the n tags, the database system 102 performs a secondary sort on each piece of data in each category obtained through the primary sorting, based on the tag value of the second tag to be sorted included in each piece of data, thereby grouping data with the same tag value of the second tag within the same category of data obtained through the primary sorting into a single category. ... For the nth of the n tags, the database system 102 performs an nth sort on each piece of data in each category obtained through the n-1th level sorting, based on the tag value of the nth tag to be sorted included in each piece of data, thereby grouping data with the same tag value of the nth tag within the same category of data obtained through the n-1th level sorting into a single category, thereby determining the storage order of the multiple pieces of data.
[0048] In some embodiments, the database creation request also includes identification information of the database (such as the name of the database, etc.).
[0049] For example, the database system 102 receives a database creation request as shown below.
[0050] Create Table information(Operator, Direction, Address, Time, Bps)
[0051] SORTKEY Operator, Time, Address.
[0052] This database creation request includes five tags in the "information" database, which correspond one-to-one to the five columns of the "information" database. The first column of the "information" database stores the tag value for the "Operator" tag, the second column stores the tag value for the "Direction" tag, the third column stores the tag value for the "Address" tag, the fourth column stores the tag value for the "Time" tag, and the fifth column stores the tag value for the "Bps" tag.
[0053] The database creation request also includes data structure information, which defines three tags that need to be sorted, and the order of sorting the three tags is Operator, Time, and Address. That is to say, for multiple data to be saved in the database, when sorting the multiple data, first perform a first-level sort on each data based on the tag value of the tag "Operator" included in each data. Then, perform a second-level sort based on the tag value of the tag "Time" included in each data after sorting. Finally, perform a third-level sort based on the tag value of the tag "Address" included in each data after sorting to obtain the storage order of each data. Each data is saved in the database according to the storage order of each data.
[0054] For example, the database system 102 receives a storage request as shown below, which includes twenty pieces of data.
[0055] Insert into information(Operator, Direction, Address, Time, Bps)Value(
[0056] Operator 1, out, 192.163.13.1, 1:03:53, 60;
[0057] Operator 1, out, 192.163.13.5, 1:03:52, 40;
[0058] Operator 1, out, 192.163.13.2, 1:03:55, 59;
[0059] Operator 1, out, 192.163.13.1, 1:03:56, 69;
[0060] Operator 2, out, 192.163.13.4, 7:02:40, 70;
[0061] Operator 2, out, 192.163.16.4, 7:22:40, 30;
[0062] Operator 1,in,192.163.13.2,9:03:49,102;
[0063] Operator 1, in, 192.163.12.2, 9:03:49, 122;
[0064] Carrier 1, out, 192.163.13.2, 1:03:50, 160;
[0065] Operator 1, out, 192.163.23.2, 1:03:50, 60;
[0066] Operator 2, out, 192.163.13.4, 3:02:52, 59;
[0067] Operator 2, out, 192.163.11.4, 3:02:52, 39;
[0068] Carrier 2, out, 192.163.13.5, 7:02:29, 80;
[0069] Operator 2, out, 192.163.33.5, 7:02:39, 85;
[0070] Operator 1, out, 192.163.13.1, 1:03:43, 20;
[0071] Operator 1, out, 192.163.19.1, 1:03:43, 30;
[0072] Operator 1,in,192.163.13.1,1:03:42,60;
[0073] Operator 1,in,192.163.113.1,1:03:42,20;
[0074] Operator 2, out, 192.163.13.4, 2:02:45, 90;
[0075] Operator 2, out, 192.163.33.4, 2:02:45, 80;
[0076] ).
[0077] Database system 102 first performs a primary sort on the 20 data items based on the "Operator" tag value included in each of the 20 data items, resulting in the following results. During the primary sort, data items with the same "Operator" tag value are clustered together to form a cluster. As shown below, during the primary sort, the 12 data items including operator 1 are clustered together to form cluster 1, and the 8 data items including operator 2 are clustered together to form cluster 2.
[0078] Operator 1, out, 192.163.13.1, 1:03:53, 60;
[0079] Operator 1, out, 192.163.13.5, 1:03:52, 40;
[0080] Operator 1, out, 192.163.13.2, 1:03:55, 59;
[0081] Operator 1, out, 192.163.13.1, 1:03:56, 69;
[0082] Operator 1,in,192.163.13.2,9:03:49,102;
[0083] Operator 1, in, 192.163.12.2, 9:03:49, 122;
[0084] Carrier 1, out, 192.163.13.2, 1:03:50, 160;
[0085] Operator 1, out, 192.163.23.2, 1:03:50, 60;
[0086] Operator 1, out, 192.163.13.1, 1:03:43, 20;
[0087] Operator 1, out, 192.163.19.1, 1:03:43, 30;
[0088] Operator 1,in,192.163.13.1,1:03:42,60;
[0089] Operator 1,in,192.163.113.1,1:03:42,20;
[0090] Operator 2, out, 192.163.13.4, 7:02:40, 70;
[0091] Operator 2, out, 192.163.16.4, 7:22:40, 30;
[0092] Operator 2, out, 192.163.13.4, 3:02:52, 59;
[0093] Operator 2, out, 192.163.11.4, 3:02:52, 39;
[0094] Carrier 2, out, 192.163.13.5, 7:02:29, 80;
[0095] Operator 2, out, 192.163.33.5, 7:02:39, 85;
[0096] Operator 2, out, 192.163.13.4, 2:02:45, 90;
[0097] Operator 2, out, 192.163.33.4, 2:02:45, 80.
[0098] Database system 102 performs a secondary sort on the 20 pieces of data based on the "Time" tag value included in each of the 20 pieces of data. During the secondary sort, the relative order between class 1 and class 2 remains unchanged. Instead, the secondary sort is performed on the 12 pieces of data in class 1 based on the "Time" tag value included in each of the 12 pieces of data. That is to say, the two data items including "1:03:42" will be clustered into class 11, the two data items including "1:03:43" will be clustered into class 12, the two data items including "1:03:50" will be clustered into class 13, the one data item including "1:03:52" will be clustered into class 14, the one data item including "1:03:53" will be clustered into class 15, the two data items including "1:03:55" will be clustered into class 16, the two data items including "1:03:56" will be clustered into class 17, and the two data items including "9:03:49" will be clustered into class 18.
[0099] Based on the label value of "Time" included in each of the eight data items in class 2, the eight data items are sorted in a secondary order. That is, the two data items including "2:02:45" are clustered into class 21, the two data items including "3:02:52" are clustered into class 22, the one data item including "7:02:29" is clustered into class 23, the one data item including "7:02:39" is clustered into class 24, the one data item including "7:02:40" is clustered into class 25, and the one data item including "7:22:40" is clustered into class 26, resulting in the following results.
[0100] Operator 1,in,192.163.13.1,1:03:42,60;
[0101] Operator 1,in,192.163.113.1,1:03:42,20;
[0102] Operator 1, out, 192.163.13.1, 1:03:43, 20;
[0103] Operator 1, out, 192.163.19.1, 1:03:43, 30;
[0104] Carrier 1, out, 192.163.13.2, 1:03:50, 160;
[0105] Operator 1, out, 192.163.23.2, 1:03:50, 60;
[0106] Operator 1, out, 192.163.13.5, 1:03:52, 40;
[0107] Operator 1, out, 192.163.13.1, 1:03:53, 60;
[0108] Operator 1, out, 192.163.13.2, 1:03:55, 59;
[0109] Operator 1, out, 192.163.13.1, 1:03:56, 69;
[0110] Operator 1,in,192.163.13.2,9:03:49,102;
[0111] Operator 1, in, 192.163.12.2, 9:03:49, 122;
[0112] Operator 2, out, 192.163.13.4, 2:02:45, 90;
[0113] Operator 2, out, 192.163.33.4, 2:02:45, 80;
[0114] Operator 2, out, 192.163.13.4, 3:02:52, 59;
[0115] Operator 2, out, 192.163.11.4, 3:02:52, 39;
[0116] Carrier 2, out, 192.163.13.5, 7:02:29, 80;
[0117] Operator 2, out, 192.163.33.5, 7:02:39, 85;
[0118] Operator 2, out, 192.163.13.4, 7:02:40, 70;
[0119] Operator 2, out, 192.163.16.4, 7:22:40, 30.
[0120] The database system 102 performs a three-level sort on the twenty pieces of data based on the label value of "Address" included in each of the twenty pieces of data. When performing the three-level sort, the relative order between class 11, class 12, class 13, class 14, class 15, class 16, class 17, class 18, class 21, class 22, class 23, class 24, class 25 and class 26 remains unchanged. And when performing three-level sorting, based on the tag value of "Address" included in the two data in class 11, the two data are sorted in three levels; based on the tag value of "Address" included in the two data in class 12, the two data are sorted in three levels; based on the tag value of "Address" included in the two data in class 13, the two data are sorted in three levels; based on the tag value of "Address" included in the two data in class 16, the two data are sorted in three levels; based on the tag value of "Address" included in the two data in class 17, the two data are sorted in three levels; based on the tag value of "Address" included in the two data in class 18, the two data are sorted in three levels; based on the tag value of "Address" included in the two data in class 21, the two data are sorted in three levels; based on the tag value of "Address" included in the two data in class 22, the two data are sorted in three levels. The storage order of the twenty data obtained after the three-level sorting is as follows.
[0121] Operator 1,in,192.163.13.1,1:03:42,60;
[0122] Operator 1,in,192.163.113.1,1:03:42,20;
[0123] Operator 1, out, 192.163.13.1, 1:03:43, 20;
[0124] Operator 1, out, 192.163.19.1, 1:03:43, 30;
[0125] Carrier 1, out, 192.163.13.2, 1:03:50, 160;
[0126] Operator 1, out, 192.163.23.2, 1:03:50, 60;
[0127] Operator 1, out, 192.163.13.5, 1:03:52, 40;
[0128] Operator 1, out, 192.163.13.1, 1:03:53, 60;
[0129] Operator 1, out, 192.163.13.2, 1:03:55, 59;
[0130] Operator 1, out, 192.163.13.1, 1:03:56, 69;
[0131] Operator 1, in, 192.163.12.2, 9:03:49, 122;
[0132] Operator 1,in,192.163.13.2,9:03:49,102;
[0133] Operator 2, out, 192.163.13.4, 2:02:45, 90;
[0134] Operator 2, out, 192.163.33.4, 2:02:45, 80;
[0135] Operator 2, out, 192.163.11.4, 3:02:52, 39;
[0136] Operator 2, out, 192.163.13.4, 3:02:52, 59;
[0137] Carrier 2, out, 192.163.13.5, 7:02:29, 80;
[0138] Operator 2, out, 192.163.33.5, 7:02:39, 85;
[0139] Operator 2, out, 192.163.13.4, 7:02:40, 70;
[0140] Operator 2, out, 192.163.16.4, 7:22:40, 30.
[0141] The database system 102 saves the twenty pieces of data into the database "information" as shown in Table 1 below based on the storage order of the twenty pieces of data.
[0142] Table 1
[0143] For certain tags in a database, the cardinality of their tag values satisfies the high-cardinality condition. This condition refers to a tag with a large number of deduplicated tag values. Typically, the number of deduplicated tag values for this tag may exceed a threshold. For example, the number of deduplicated tag values for this tag may be in the thousands, tens of thousands, millions, tens of millions, or even hundreds of millions. A high-cardinality column in a database refers to a column of tag values in the database that satisfies the high-cardinality condition.
[0144] When there are multiple sort keys, multi-level sorting is required. When a column is preceded by a high-base column, first sorting the data by the tag value of the high-base column will result in poor ordering of the subsequent columns. For example, in the case of the first and second tags described above, the first tag precedes the second tag in the sort key, and the tag value of the column corresponding to the first tag is a high-base column. Because the data is sorted by the tag value of the high-base column first and then sorted by the tag value of the column corresponding to the second tag, the order of the tag value of the column corresponding to the second tag is poor.
[0145] For example, for the database shown in Table 1 above, the tag "Address" and the tag "Time" in the database, the deduplicated tag value of the tag "Address" is a large number of addresses, which may have hundreds of millions or billions of different addresses. The deduplicated tag value of the tag "Time" is a large number of timestamps, which may have hundreds of millions or billions of different timestamps. Therefore, each piece of data is sorted based on the timestamp and address in each piece of data. Refer to Table 1. For the first piece of data with sequence number 1, the third piece of data with sequence number 3, the fifth piece of data with sequence number 5, the seventh piece of data with sequence number 7, the ninth piece of data with sequence number 9, the tenth piece of data with sequence number 10, and the twelfth piece of data with sequence number 12, these data are all addresses within the 192.163.13 network segment, but these data are scattered and stored in the data table shown in Table 1. The 13th data item with sequence number 13, the 16th data item with sequence number 16, the 17th data item with sequence number 17, and the 19th data item with sequence number 19 all have addresses within the 192.163.13 network segment, but these addresses are scattered across the database shown in Table 1. As a result, the addresses corresponding to the "Address" tag included in each sorted data item have poor local ordering. Therefore, if you need to query data for a specific consecutive address within the 192.163.13 network segment, the scanned data range will be very large, increasing the query difficulty and reducing query efficiency.
[0146] In order to improve the local orderliness of data stored in a database, data can be stored in the database using any of the following embodiments. The database can then be processed to improve the efficiency of database processing. For example, querying data from the database can improve query efficiency.
[0147] 2 , an embodiment of the present application provides a method 200 for storing data. The method 200 is applied to the network architecture 100 shown in FIG1 and includes the following process.
[0148] Step 201: The terminal device sends a database creation request to the database system. The database creation request includes multiple tags and data structure information. The data structure information is used to specify at least two tags among the multiple tags as sort keys.
[0149] The at least two tags include a first tag and a second tag, the first tag being before the second tag in the sort key.
[0150] In some embodiments, the data structure information further includes quantization identification information, which is used to indicate the first tag. Optionally, the quantization identification information indicates that the tag value of the first tag needs to be quantized to obtain identification information of the tag value set to which the tag value of the first tag belongs.
[0151] In some embodiments, for a first tag among the at least two tags that requires quantization, a tag may be added to the first tag, and the quantization identification information includes the tag added to the first tag. Optionally, the tag may be the granularity of the tag value set to which the tag value of the first tag belongs, or the tag may be an automatic tag that automatically determines the granularity of the tag value set to which the tag value of the first tag belongs.
[0152] Optionally, the database creation request may further include identification information of the database. Optionally, the identification information of the database may include the name of the database, etc.
[0153] In some embodiments, for the tag value set to which the tag value of the first tag belongs, the tag value set includes multiple tag values belonging to the first tag. For example, assuming that the first tag is a timestamp, the granularity of the tag value set is 1 minute. For the timestamp 7:05:40, the tag value set to which the timestamp 7:05:40 belongs is the set greater than or equal to 7:05:01 and less than or equal to 7:05:59. After quantizing the timestamp 7:05:40, the tag value set obtained is the tag value set whose identification information can be "7:05". For the timestamp 7:06:25, the tag value set to which the timestamp 7:06:25 belongs is the set greater than or equal to 7:06:01 and less than or equal to 7:06:59. After quantizing the timestamp 7:06:25, the tag value set obtained is the tag value set whose identification information can be "7:06".
[0154] For another example, assuming the first label is a timestamp and the granularity of the label value set is 10 minutes, for timestamp 7:15:40, the label value set to which timestamp 7:15:40 belongs is the set greater than or equal to 7:10:01 and less than or equal to 7:19:59. Quantizing timestamp 7:15:40 yields a label value set whose identification information can be "7:10." For timestamp 7:25:25, the label value set to which timestamp 7:25:25 belongs is the set greater than or equal to 7:20:01 and less than or equal to 7:29:59. Quantizing timestamp 7:25:25 yields a label value set whose identification information can be "7:20."
[0155] In some embodiments, the first tag indicated by the quantitative identification information may include one tag or multiple tags.
[0156] In step 201, the terminal device may display a first editing interface, and the user may enter a database creation statement in the first editing interface. The terminal device obtains the database creation statement from the first editing interface and uses the database creation statement as the database creation request.
[0157] For example, the terminal device can display the first editing interface as shown in Figure 3, and the user can enter the database creation statement 1 shown below in the first editing interface. The database creation statement 1 includes the name of the database "information", five tags and data structure information. The five tags are "Operator", "Direction", "Address", "Time" and "Bps". The data structure information is used to indicate that the three tags that need to be sorted are sorting keys, and the order in which the three tags need to be sorted is "Operator", "Time", and "Address". In one embodiment, the data structure information also includes quantification identification information. The quantification identification information can be the granularity "1m" of the tag value set marked on the first tag "Time". The quantification identification information indicates that the first tag that needs to be quantified is "Time". Therefore, the granularity "1m" not only indicates that the tag that needs to be quantified is "Time", but also indicates the granularity.
[0158] Database creation statement 1:
[0159] Create Table information(Operator, Direction, Address, Time, Bps)
[0160] SORTKEY Operator, direction, Time (1m), Address.
[0161] The terminal device can obtain the database creation statement 1 from the first editing interface shown in FIG. 3 , and use the database creation statement 1 as the database creation request.
[0162] For another example, referring to the first editing interface shown in FIG4 , a user can enter the following database creation statement 2 in the first editing interface. Compared to the database creation statement 1 in the first editing interface shown in FIG3 , the first tag "Time" in the database creation statement 2 in the first editing interface shown in FIG4 is marked with an automatic tag "auto." The automatic tag "auto" is used to instruct the database system to determine the granularity of the tag value set for the first tag "Time" based on the characteristics of the tag value of the first tag "Time."
[0163] Database creation statement 2:
[0164] Create Table information(Operator, Direction, Address, Time, Bps)
[0165] SORTKEY Operator, direction, Time (auto), Address.
[0166] For the first tag marked with automatic marking, the database system may determine the granularity based on the characteristics of the first tag data stored statistically, for example, based on the distribution characteristics of the first tag data stored statistically.
[0167] Step 202: The database system receives the database creation request and creates a database based on the multiple tags included in the database creation request.
[0168] In step 202, a database creation request includes the multiple tags, and a database is created based on the multiple tags. Optionally, if the database is a data table, the multiple tags correspond one-to-one to multiple columns included in the data table, and the columns corresponding to each tag are used to store the tag value of each tag.
[0169] In some embodiments, each label is column identification information of a column corresponding to each label.
[0170] In some embodiments, the database creation request further includes identification information of the database. After creating the database, the database system sets the identification information of the database to the identification information of the database included in the database creation request.
[0171] For example, the database system receives the aforementioned database creation statement 1 or database creation statement 2 and creates the database "information" based on the five tags included in database creation statement 1 or the five tags included in database creation statement 2. The database "information" includes five columns: the first column corresponding to the tag "Operator", the second column corresponding to the tag "Direction", the third column corresponding to the tag "Address", the fourth column corresponding to the tag "Time", and the fifth column corresponding to the tag "Bps". "Operator", "Direction", "Address", "Time", and "Bps" are the column names of the five columns.
[0172] After creating the database, users can save data to the database. The detailed implementation process is as follows.
[0173] Step 203: The terminal device sends a storage request to the database system, where the storage request includes M pieces of data. For each piece of data, the data includes the tag values of the multiple tags, where the multiple tags include a first tag and a second tag designated as sort keys, where the first tag is before the second tag in the sort key, and M is an integer greater than 1.
[0174] In some embodiments, the storage request may be a data insert statement.
[0175] The user can input a data insertion statement into the second editing interface displayed on the terminal device, where the data insertion statement includes the identification information of the database and M pieces of data. The terminal device obtains the data insertion statement from the second editing interface and sends the data insertion statement to the database system.
[0176] For example, the user may enter a data insertion statement in the second editing interface displayed on the terminal device. The data insertion statement is as shown in the example below.
[0177] Example of data insertion statement:
[0178] Insert into information(Operator, Direction, Address, Time, Bps)Value(
[0179] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60;
[0180] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40;
[0181] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59;
[0182] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69;
[0183] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70;
[0184] Data 6: Carrier 2, out, 192.163.16.4, 7:22:40, 30;
[0185] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102;
[0186] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122;
[0187] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160;
[0188] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60;
[0189] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59;
[0190] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39;
[0191] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80;
[0192] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85;
[0193] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20;
[0194] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30;
[0195] Data 17: Carrier 1, in, 192.163.13.1, 1:03:42, 60;
[0196] Data 18: Carrier 1, in, 192.163.113.1, 1:03:42, 20;
[0197] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90;
[0198] Data 20: Carrier 2, out, 192.163.33.4, 2:02:45, 80;
[0199] ).
[0200] The data insertion statement includes database identification information "information" and twenty pieces of data (for ease of explanation, in this example of the data insertion statement, these twenty pieces of data are referred to as data 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20), i.e., M = 20. The terminal device obtains the data insertion statement from the second editing interface, uses the data insertion statement as a storage request, and sends the storage request to the database system.
[0201] Step 204: The database system receives the storage request and obtains the tag value of the first tag included in each piece of data in the storage request.
[0202] In step 204 , for each piece of data included in the storage request, the tag value of the first tag in the piece of data is determined, thereby obtaining the tag value of the first tag included in the piece of data.
[0203] The first label is the label indicated by the quantified identification information. The database system determines the first label based on the quantified identification information, determines the label value of the first label in the data piece, and obtains the label value of the first label included in the data piece. Alternatively, the database system determines the first label to be quantified, where the first label is a label whose cardinality among the multiple labels satisfies a high cardinality condition, determines the label value of the first label in the data piece, and obtains the label value of the first label included in the data piece.
[0204] In some embodiments, for the first tag that needs to be quantified, the database system can also count the frequency of using the first tag as a filtering condition in query requests received in the past. If the frequency is high, that is, the frequency exceeds the frequency threshold, the first tag can be used as a tag that needs to be quantified.
[0205] For example, the database system receives the data insertion statement listed in the above example, which includes data 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20, and the first label indicated by the quantitative identification information is "Time".
[0206] For data 1, obtain the tag value "1:03:53" for the first tag "Time" included in data 1. For data 2, obtain the tag value "1:03:52" for the first tag "Time" included in data 2. For data 3, obtain the tag value "1:03:55" for the first tag "Time" included in data 3. ..., for data 20, obtain the tag value "2:02:45" for the first tag "Time" included in data 20.
[0207] Step 205: The database system determines the tag value set to which the tag value of the first tag included in each data belongs based on the tag value of the first tag included in each data, and obtains N tag value sets, where N is an integer greater than or equal to 1 and less than M.
[0208] In step 205 , based on the tag value of the first tag included in each piece of data and the granularity of the tag value set of the first tag, it is determined whether the tag of the first tag included in each piece of data belongs to the tag value set.
[0209] In some embodiments, the granularity of the tag value set of the first tag may be the granularity indicated by the quantization identification information.
[0210] In some embodiments, the granularity of the tag value set of the first tag may be a granularity determined based on a feature of the tag value of the first tag.
[0211] In some embodiments, the database system may have multiple quantization methods, and may select a corresponding quantization method based on the first label and granularity. The quantization method may be used to quantify the label of the first label included in each piece of data, and the label value of the first label included in each piece of data may belong to the label value set.
[0212] For example, assuming the tag value of the first tag to be quantized is the timestamp 7:05:40, if the first tag is the timestamp "Time" and the granularity is 1 minute, the database system can select a corresponding quantization method based on the first tag "Time" and the granularity "1 minute", and use this quantization method to quantize the timestamp 7:05:40 to obtain the tag value set "7:05". assuming the tag value of the first tag to be quantized is the timestamp 7:15:40, if the first tag is the timestamp "Time" and the granularity is 10 minutes, the database system can select a corresponding quantization method based on the first tag "Time" and the granularity "10 minutes", and use this quantization method to quantize the timestamp 7:15:40 to obtain the tag value set "7:10".
[0213] For another example, suppose that the label value of the first label that needs to be quantized is the address 192.168.13.4. If the first label is the address "Address" and the subnet mask has a granularity of 255.255.255.0, the database system can select a corresponding quantization method based on the first label "Address" and the subnet mask, and use the quantization method to mask the address 192.168.13.4 to obtain a label value set of "192.168.13".
[0214] The process of quantizing the label value of the first label is to divide the amplitude of the entire first label into a set of finite small amplitudes (quantization steps), classify the samples falling within a certain step into one category, and assign the same quantization value.
[0215] Therefore, in step 205, the database system divides all tag values of the first tag into intervals, classifies tag values within the same interval into a tag value set, and assigns identification information to the tag value set. For each tag value of the first tag included in each piece of data, the database system can determine the tag value set to which the tag value belongs from the divided tag value sets, thereby obtaining the identification information of the tag value set.
[0216] For example, assuming the first label is a timestamp and the granularity of the label value set is 1 minute, the amplitude of the entire timestamp is divided into multiple steps to obtain multiple label value sets, and the identification information of the multiple label value sets is assigned. The identification information of the multiple label value sets obtained by the assignment is 7:01, 7:02, 7:03, 7:04, 7:05, etc. For the timestamp 7:05:40, the timestamp 7:05:40 is quantized to obtain the label value set with the identification information "7:05", which includes the timestamp 7:05:40. Therefore, the label value set obtained by the quantization of the timestamp 7:05:40 is the label value set with the identification information "7:05".
[0217] In some embodiments, there may be multiple first labels that need to be quantified. For each piece of data, the label value set to which the label value of each first label included in the data belongs is determined based on the label value of each first label included in the data and the granularity corresponding to each first label.
[0218] For example, assuming that the first tag is "Time", in the above data 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20, the tag value corresponding to the first tag "Time" is a timestamp. The characteristic of the timestamp is that the seconds in the timestamp are constantly changing. Therefore, the granularity of the tag value set of the first tag "Time" is determined to be 1m.
[0219] Assuming that the granularity of the label value set of the first label "Time" is 1m, for the above-mentioned data 1, based on the label value "1:03:53" of the first label "Time" included in data 1 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:53" of the first label "Time" included in data 1 belongs.
[0220] For the above-mentioned data 2, based on the label value "1:03:52" of the first label "Time" included in data 2 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:52" of the first label "Time" included in data 2 belongs.
[0221] For the above-mentioned data 3, based on the label value "1:03:55" of the first label "Time" included in data 3 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:55" of the first label "Time" included in data 3 belongs.
[0222] For the above-mentioned data 4, based on the label value "1:03:56" of the first label "Time" included in data 4 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:56" of the first label "Time" included in data 4 belongs.
[0223] For the above-mentioned data 5, based on the label value "7:02:40" of the first label "Time" included in data 5 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "7:02" to which the label value "7:02:40" of the first label "Time" included in data 5 belongs.
[0224] For the above-mentioned data 6, based on the label value "7:22:40" of the first label "Time" included in data 6 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "7:22" to which the label value "7:22:40" of the first label "Time" included in data 6 belongs.
[0225] For the above-mentioned data 7, based on the label value "9:03:49" of the first label "Time" included in data 7 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "9:03" to which the label value "9:03:40" of the first label "Time" included in data 7 belongs.
[0226] For the above-mentioned data 8, based on the label value "9:03:49" of the first label "Time" included in data 8 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "9:03" to which the label value "9:03:49" of the first label "Time" included in data 8 belongs.
[0227] For the above-mentioned data 9, based on the label value "1:03:50" of the first label "Time" included in data 9 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:50" of the first label "Time" included in data 9 belongs.
[0228] For the above-mentioned data 10, based on the label value "1:03:50" of the first label "Time" included in the data 10 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:50" of the first label "Time" included in the data 10 belongs.
[0229] For the above-mentioned data 11, based on the label value "3:02:52" of the first label "Time" included in the data 11 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "3:02" to which the label value "3:02:52" of the first label "Time" included in the data 11 belongs.
[0230] For the above-mentioned data 12, based on the label value "3:02:52" of the first label "Time" included in the data 12 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "3:02" to which the label value "3:02:52" of the first label "Time" included in the data 12 belongs.
[0231] For the above-mentioned data 13, based on the label value "7:02:29" of the first label "Time" included in the data 13 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "7:02" to which the label value "7:02:29" of the first label "Time" included in the data 13 belongs.
[0232] For the above-mentioned data 14, based on the label value "7:02:39" of the first label "Time" included in the data 14 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "7:02" to which the label value "7:02:39" of the first label "Time" included in the data 14 belongs.
[0233] For the above-mentioned data 15, based on the label value "1:03:43" of the first label "Time" included in data 15 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:43" of the first label "Time" included in data 15 belongs.
[0234] For the above-mentioned data 16, based on the label value "1:03:43" of the first label "Time" included in data 16 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:43" of the first label "Time" included in data 16 belongs.
[0235] For the above-mentioned data 17, based on the label value "1:03:42" of the first label "Time" included in data 17 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:42" of the first label "Time" included in data 17 belongs.
[0236] For the above-mentioned data 18, based on the label value "1:03:42" of the first label "Time" included in the data 18 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:42" of the first label "Time" included in the data 18 belongs.
[0237] For the above-mentioned data 19, based on the label value "2:02:45" of the first label "Time" included in data 19 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "2:02" to which the label value "2:02:45" of the first label "Time" included in data 19 belongs.
[0238] For the above-mentioned data 20, based on the label value "2:02:45" of the first label "Time" included in the data 20 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "2:02" to which the label value "2:02:45" of the first label "Time" included in the data 20 belongs.
[0239] As shown below, the tag value set to which the first tag included in each of the twenty pieces of data belongs is as follows:
[0240] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60; 1:03;
[0241] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40; 1:03;
[0242] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59; 1:03;
[0243] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69; 1:03;
[0244] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70; 7:02;
[0245] Data 6: Carrier 2, out, 192.163.16.4, 7:22:40, 30; 7:22;
[0246] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102; 9:03;
[0247] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122; 9:03;
[0248] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160; 1:03;
[0249] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60; 1:03;
[0250] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59; 3:02;
[0251] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39; 3:02;
[0252] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80; 7:02;
[0253] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85; 7:02;
[0254] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20; 1:03;
[0255] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30; 1:03;
[0256] Data 17: Carrier 1, in, 192.163.13.1, 1:03:42, 60; 1:03;
[0257] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20; 1:03;
[0258] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90; 2:02;
[0259] Data 20: Operator 2, out, 192.163.33.4, 2:02:45, 80; 2:02.
[0260] Step 206: The database system determines the storage order of the M pieces of data based on the order of the sort keys, identification information of the tag value set to which the tag value of the first tag included in each piece of data belongs, and the tag value of the second tag included in each piece of data.
[0261] In step 206, the data structure information defines at least two tags that need to be sorted. If the order of the at least two tags included in the sort key is earlier than the tag of the first tag, each piece of data is sorted based on the tag value included in each piece of data. Then, based on the identification information of the tag value set to which the tag value of the first tag included in each piece of data belongs, each piece of data is sorted based on the sorted data. Then, based on the tag value of the second tag included in each piece of data, each piece of data is sorted based on the sorted data, to obtain the storage order of the M pieces of data.
[0262] For example, the data structure information defines at least two tags that require sorting: "Operator," "Time," and "Address." For the 20 data items listed above, the "Operator" tag precedes the "Time" tag in the sort key. The database system performs a primary sort on these 20 data items based on the "Operator" tag value included in each data item. During the primary sort, the 12 data items including operator 1 are clustered into cluster 1, and the 8 data items including operator 2 are clustered into cluster 2. The resulting sorting results are as follows.
[0263] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60; 1:03;
[0264] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40; 1:03;
[0265] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59; 1:03;
[0266] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69; 1:03;
[0267] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102; 9:03;
[0268] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122; 9:03;
[0269] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160; 1:03;
[0270] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60; 1:03;
[0271] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20; 1:03;
[0272] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30; 1:03;
[0273] Data 17: Carrier 1, in, 192.163.13.1, 1:03:42, 60; 1:03;
[0274] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20; 1:03;
[0275] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70; 7:02;
[0276] Data 6: Carrier 2, out, 192.163.16.4, 7:22:40, 30; 7:22;
[0277] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59; 3:02;
[0278] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39; 3:02;
[0279] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80; 7:02;
[0280] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85; 7:02;
[0281] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90; 2:02;
[0282] Data 20: Operator 2, out, 192.163.33.4, 2:02:45, 80; 2:02.
[0283] In the sorting key, the order of the label "Time" is before the label "Address". The database system performs secondary sorting on the twenty pieces of data based on the identification information of the label value set to which the label value of the first label "Time" included in each piece of data belongs, based on the above sorting results. When performing secondary sorting, the relative order between class 1 and class 2 remains unchanged, but the twelve pieces of data in class 1 are sorted at the secondary level based on the identification information of the label value set to which the label value of "Time" included in each piece of data belongs. That is, data 1, 2, 3, 4, 9, 10, 15, 16, 17 and 18 including the identification information "1:03" of the label value set to which the label value of "Time" belongs are clustered to form class 11, and data 7 and 8 including the identification information "9:03" of the label value set to which the label value of "Time" belongs are clustered to form class 12.
[0284] Based on the identification information of the tag value set to which the tag value of "Time" included in each of the eight data in class 2 belongs, the eight data are sorted at the secondary level. That is, data 19 and 20 including the identification information "2:02" of the tag value set to which the tag value of "Time" belongs are clustered to form class 21, data 11 and 12 including the identification information "3:02" of the tag value set to which the tag value of "Time" belongs are clustered to form class 22, data 5, 13 and 14 including the identification information "7:02" of the tag value set to which the tag value of "Time" belongs are clustered to form class 23, and data 6 including the identification information "7:22" of the tag value set to which the tag value of "Time" belongs are clustered to form class 24. The sorting results are as follows.
[0285] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60; 1:03;
[0286] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40; 1:03;
[0287] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59; 1:03;
[0288] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69; 1:03;
[0289] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160; 1:03;
[0290] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60; 1:03;
[0291] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20; 1:03;
[0292] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30; 1:03;
[0293] Data 17: Carrier 1, in, 192.163.13.1, 1:03:42, 60; 1:03;
[0294] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20; 1:03;
[0295] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102; 9:03;
[0296] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122; 9:03;
[0297] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90; 2:02;
[0298] Data 20: Carrier 2, out, 192.163.33.4, 2:02:45, 80; 2:02;
[0299] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59; 3:02;
[0300] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39; 3:02;
[0301] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70; 7:02;
[0302] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80; 7:02;
[0303] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85; 7:02;
[0304] Data 6: Operator 2, out, 192.163.16.4, 7:22:40, 30; 7:22.
[0305] Then, based on the tag value of the first tag "Address" included in each piece of data, the database system performs a three-level sort on the twenty pieces of data on the basis of the above sorting results. When performing the three-level sorting, the relative order between class 11, class 12, class 21, class 22, class 23, and class 24 remains unchanged. And when performing the three-level sorting, based on the tag value of "Address" included in the ten pieces of data in class 11, the ten pieces of data are sorted in three levels; based on the tag value of "Address" included in the two pieces of data in class 12, the two pieces of data are sorted in three levels; based on the tag value of "Address" included in the two pieces of data in class 21, the two pieces of data are sorted in three levels; based on the tag value of "Address" included in the two pieces of data in class 22, the two pieces of data are sorted in three levels; based on the tag value of "Address" included in the three pieces of data in class 23, the three pieces of data are sorted in three levels. After performing the three-level sorting, the storage order of the twenty pieces of data is as follows.
[0306] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60; 1:03;
[0307] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69; 1:03;
[0308] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20; 1:03;
[0309] Data 17: Carrier 1, in, 192.163.13.1, 1:03:42, 60; 1:03;
[0310] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59; 1:03;
[0311] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160; 1:03;
[0312] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40; 1:03;
[0313] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30; 1:03;
[0314] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60; 1:03;
[0315] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20; 1:03;
[0316] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122; 9:03;
[0317] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102; 9:03;
[0318] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90; 2:02;
[0319] Data 20: Carrier 2, out, 192.163.33.4, 2:02:45, 80; 2:02;
[0320] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39; 3:02;
[0321] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59; 3:02;
[0322] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70; 7:02;
[0323] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80; 7:02;
[0324] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85; 7:02;
[0325] Data 6: Operator 2, out, 192.163.16.4, 7:22:40, 30; 7:22.
[0326] Step 207: The database system stores the M pieces of data in the database based on the storage order of the M pieces of data.
[0327] For example, based on the storage order of the twenty pieces of data, the twenty pieces of data are saved in a database as shown in Table 2 below.
[0328] Table 2
[0329] Compared with database 1, database 2 tries to group together data of addresses belonging to the same network segment. The addresses in the "address" column in database 2 are more ordered than those in the "address" column in database 1.
[0330] Step 208: The database system constructs index information of the M pieces of data, where the index information includes identification information of the label value set to which the label value of the first label in at least one selected piece of data belongs, wherein two adjacent selected pieces of data are separated by one or more pieces of data in the storage order.
[0331] Optionally, the selected at least one piece of data may include the first piece of data and the last piece of data among the M pieces of data.
[0332] Assume that the number of the one or more interval data is x, where x is an integer greater than or equal to 1. In step 208, the database system selects at least one data from the M data stored in the database, with an interval of x data between two adjacent selected data.
[0333] For example, the database system selects data 1, data 17, data 2, data 18, data 19, data 11, and data 6 from the twenty pieces of data stored in the database as shown in Table 2. Any two adjacent selected pieces of data are separated by two pieces of data, that is, x=2.
[0334] In some embodiments, the constructed index information may be identification information of a tag value set and a correspondence between tag values and storage locations. In step 208, the database system stores the identification information of the tag value set to which the tag value of the first tag included in each selected data item belongs, the tag values of other tags included in each selected data item, and the storage location of each selected data item in the database, in the identification information of the tag value set and the correspondence between tag values and storage locations. The other tags include tags in the sort key other than the first tag.
[0335] For example, the storage position of the selected data 1 in the database shown in Table 2 is sequence number 1, the storage position of the selected data 17 in the database shown in Table 2 is sequence number 4, the storage position of the selected data 2 in the database shown in Table 2 is sequence number 7, the storage position of the selected data 18 in the database shown in Table 2 is sequence number 10, the storage position of the selected data 19 in the database shown in Table 2 is sequence number 13, the storage position of the selected data 11 in the database shown in Table 2 is sequence number 16, and the storage position of the selected data 11 in the database shown in Table 2 is sequence number 19.
[0336] The database system saves operator 1 included in data 1, identification information "1:03" of the label value set, 192.163.13.1 included in data 1, and storage location "serial number 1" in the corresponding relationship between identification information of the label value set, label value, and storage location shown in Table 3 below.
[0337] The database system saves operator 1, label value set identification information "1:03" included in data 17, 192.163.13.1 and storage location "serial number 4" included in data 17 in the corresponding relationship between label value set identification information, label value and storage location shown in Table 3 below.
[0338] The database system saves operator 1 included in data 2, identification information "1:03" of the label value set, 192.163.13.5 included in data 2, and storage location "serial number 7" in the corresponding relationship between identification information of the label value set, label value, and storage location shown in Table 3 below.
[0339] The database system saves operator 1 included in data 18, identification information "1:03" of the label value set, 192.163.113.1 included in data 18, and storage location "serial number 10" in the corresponding relationship between identification information of the label value set, label value, and storage location shown in Table 3 below.
[0340] The database system saves operator 2 included in data 19, identification information "2:02" of the label value set, 192.163.13.4 included in data 19, and storage location "serial number 13" in the corresponding relationship between identification information of the label value set, label value, and storage location shown in Table 3 below.
[0341] The database system saves the operator 2 included in data 11, the identification information "3:02" of the label value set, 192.163.13.4 included in data 11, and the storage location "serial number 16" in the corresponding relationship between the identification information of the label value set, the label value, and the storage location as shown in Table 3 below.
[0342] The database system saves operator 2 included in data 6, identification information "7:22" of the label value set, 192.163.16.4 included in data 6, and storage location "serial number 19" in the corresponding relationship between identification information of the label value set, label value and storage location shown in Table 3 below.
[0343] Table 3
[0344] Among them, the first label is a label whose cardinality meets the high cardinality condition. The number of label values of the first label after deduplication is very large. The order of the first label in the sorting key is before the second label. Therefore, a first-level sort is performed based on the label value of the first label included in each piece of data to be saved, and then a second-level sort is performed based on the label value of the second label included in each piece of data to be saved. This will reduce the local orderliness of a column of label values corresponding to the second label in the database.
[0345] However, in an embodiment of the present application, the tag value set to which the tag value of the first tag included in each piece of data to be saved belongs is determined, and identification information of the tag value set to which the tag value of the first tag included in each piece of data belongs is obtained, and the number of tag value sets is much smaller than the number of tag values after deduplication of the first tag. In this way, a primary sort is performed based on the identification information of the tag value set to which the tag value of the first tag included in each piece of data to be saved belongs, and a secondary sort is performed based on the tag value of the second tag included in each piece of data to be saved, which will improve the local orderliness of the column of tag values corresponding to the second tag in the database.
[0346] In an embodiment of the present application, a database system receives a database creation request, which includes multiple tags and data structure information, and the data structure information is used to define the first tag and the second tag among the multiple tags as sorting keys. When the database system receives a storage request including M data, it determines the tag value set to which the tag value of the first tag included in each data belongs. The first tag may be a tag with an extremely large cardinality of tag values, and the cardinality of the tag value set of the first tag is much smaller than the cardinality of the tag value of the first tag. Based on the order of the sorting key, the identification information of the tag value set to which the tag value of the first tag included in each data belongs and the tag value of the second tag included in each data determine the storage order of each data. Storing each data in the order in which it is stored can increase the local orderliness of a column of tag values corresponding to the second tag stored in the data table.
[0347] 5 , an embodiment of the present application provides a method 500 for querying data. The method 500 is applied to the network architecture 100 shown in FIG1 and includes the following process.
[0348] Step 501: The terminal device sends a query request to the database system, where the query request includes a query condition related to a first tag.
[0349] The query condition includes a tag value of the first tag, or the query condition includes a tag value set of the first tag, and the tag value set of the first tag includes multiple tag values of the first tag. That is, the query condition includes at least one tag value of the first tag.
[0350] In some embodiments, the query request may also include tag values of other tags.
[0351] In step 501, the terminal device may display a third editing interface, and the user may enter a data query statement in the third editing interface. The terminal device obtains the data query statement from the third editing interface and uses the data query statement as a query request.
[0352] For example, suppose a user enters the following data query statement into the third editing interface displayed on the terminal device, where the query conditions are Operator = Operator 1, Time = 1:03:56, Address = 192.163.13.1. The terminal device receives the data query statement, uses it as a query request, and sends the query request to the database system.
[0353] Data query statement:
[0354] Select from information
[0355] Where Operator = Operator 1, Time = 1:03:56, Address = 192.163.13.1.
[0356] Step 502: The database system receives the query request and determines the data to be scanned based on the query condition and the identification information of the tag value set to which the tag value of the first tag in the index information belongs.
[0357] In step 502 , the data to be scanned may be determined through the following operations 5021 - 5022 .
[0358] 5021: The database system receives the query request and obtains at least one storage location from the index information based on the identification information of the target tag value set in which at least one tag value of the first tag is located. The at least one storage location includes the storage location of the first data and the storage location of the last data in the data to be scanned.
[0359] In step 5021, the database system determines a target tag value set based on at least one tag value of the first tag, where the target tag value set includes the at least one tag value; or determines multiple target tag value sets based on at least one tag value of the first tag, where each target tag value set includes some tag values of the at least one tag value. The maximum identification information and the minimum identification information are selected from the identification information of the determined target tag value sets.
[0360] Select first data from the index information, where the first data includes identification information of a tag value set to which the tag value of the first tag belongs that is less than or equal to the minimum identification information, but no data preceding the first data in the index information includes identification information of a tag value set to which the tag value of the first tag belongs that is greater than or equal to the minimum identification information. Then, obtain a first storage location corresponding to the first data from the index information, where the first storage location is the storage location of the first piece of data in the data to be scanned.
[0361] Second data is selected from the index information. The second data includes identification information of a tag value set to which the tag value of the first tag belongs that is greater than or equal to the maximum identification information, but no data following the first data in the index information includes identification information of a tag value set to which the tag value of the first tag belongs that is less than or equal to the maximum identification information. Then, a second storage location corresponding to the second data is obtained from the index information. The second storage location is the storage location of the last piece of data in the data to be scanned.
[0362] In some embodiments, if the storage request also includes tag values of other tags, the other tags are tags other than the first tag in the sort key, the maximum identification information and the minimum identification information are selected from the identification information of the determined target tag value set, and the maximum tag value and the minimum tag value are selected from the tag values of the other tags.
[0363] First data is selected from the index information, where the first data includes identification information of a tag value set to which the tag value of the first tag belongs that is less than or equal to the minimum identification information, and tag values of other tags are less than or equal to the minimum tag value, but no data preceding the first data in the index information includes identification information of a tag value set to which the tag value of the first tag belongs that is greater than or equal to the minimum identification information, and tag values of other tags are greater than or equal to the minimum tag value. Then, a first storage location corresponding to the first data is obtained from the index information, where the first storage location is the storage location of the first piece of data in the data to be scanned.
[0364] Second data is selected from the index information, where the second data includes identification information of a tag value set to which the tag value of the first tag belongs that is greater than or equal to the maximum identification information, and tag values of other tags are greater than or equal to the maximum tag value, but no data following the first data in the index information includes identification information of a tag value set to which the tag value of the first tag belongs that is less than or equal to the maximum identification information, and tag values of other tags are less than or equal to the maximum tag value. Then, a second storage location corresponding to the second data is obtained from the index information, where the second storage location is the storage location of the last piece of data in the data to be scanned.
[0365] For example, the query conditions in the storage request are Operator = Operator 1, Time = 1:03:56, Address = 192.163.13.1. Based on the tag value "1:03:56" of the first tag "Time", the identification information of the target tag value set is determined to be "1:03". Based on "Operator = Operator 1", the identification information "1:03" of the target tag value set and "Address = 192.163.13.1", the first data and the second data are selected from the index information shown in Table 3. The first data includes operator 1, the identification information "1:03" of the tag value set, the tag value "192.163.13.1" of Address, and the storage location of the first data is serial number 1. The first data includes operator 1, the identification information "1:03" of the tag value set, the tag value "192.163.13.5" of Address, and the storage location of the second data is serial number 7.
[0366] 5022: The database system determines the data to be scanned in the M pieces of data based on the at least one storage location.
[0367] The database system obtains data located between the first storage location and the second storage location from the database as data to be scanned.
[0368] For example, the first storage location is serial number 1, and the second storage location is serial number 7. The database system can obtain 7 pieces of data between serial number 1 and serial number 7 from the database shown in Table 2 as the data to be scanned. The 7 pieces of data are data 1, data 4, data 15, data 17, data 3, data 9, and data 2.
[0369] Step 503: The database system queries the data to be scanned to obtain query results that meet the query conditions.
[0370] In step 503, the query condition includes a tag value of the first tag, or the query condition includes a tag value set of the first tag, the tag value set of the first tag includes multiple tag values of the first tag, and the database system selects target data from the data to be scanned based on the one or multiple tag values, and the tag value of the first tag included in the target data is the one tag value or one of the multiple tag values.
[0371] In some embodiments, if the storage request also includes tag values of other tags, target data is selected from the data to be scanned based on the one tag value or the multiple tag values, and the tag values of the other tags, the target data includes the tag value of the other tags, and the tag value of the included first tag is the one tag value or one of the multiple tag values.
[0372] For example, the database system selects target data from data 1, data 4, data 15, data 17, data 3, data 9, and data 2 based on Operator = Operator 1, Time = 1:03:56, Address = 192.163.13.1, and the selected target data is data 4. Since the database sorts and classifies data based on the identification information of the tag value set to which the tag value of Time in each piece of data belongs, multiple pieces of data with the same identification information of the tag value set to which the tag value of Time belongs are grouped into one category. For each piece of data grouped into one category, the data is sorted based on the tag value of Address included in each piece of data, thereby improving the local orderliness of each tag value corresponding to Address in the database, thereby reducing the range of data to be scanned to only 7 pieces of data, thereby improving query efficiency.
[0373] Step 504: The database system sends a query response to the terminal device, where the query response includes the query result.
[0374] The query result is the target data obtained by the above selection. After receiving the query response, the terminal device displays the target data included in the query response.
[0375] In an embodiment of the present application, a database system receives a query request including a query condition indicating at least one tag value of a first tag. Based on identification information of a target tag value set in which the at least one tag value resides, at least one storage location is retrieved from index information. Based on the at least one storage location, data to be scanned is determined from the database, thereby narrowing the scope of the data to be scanned. Then, based on the at least one tag value, target data including some or all of the tag values in the at least one tag value is selected from the data to be scanned, thereby improving the efficiency of querying the target data.
[0376] 6 , an embodiment of the present application provides a device 600 for storing data. The device 600 may be deployed in the database system 102 in the network architecture 100 shown in FIG1 , or in the database system 102 in the method 200 shown in FIG2 , or in the database system 102 in the method 500 shown in FIG5 . The device 600 includes:
[0377] A receiving unit 601 is configured to receive a plurality of data pieces, wherein each data piece includes tag values of a plurality of tags, the plurality of tags including a first tag and a second tag designated as a sort key, the first tag being before the second tag in the sort key;
[0378] The processing unit 602 is configured to determine the tag value set to which the tag value of the first tag in each piece of data belongs;
[0379] The processing unit 602 is further configured to determine a storage order of the plurality of data pieces based on the order of the sort keys, identification information of the tag value set to which the tag value of the first tag in each piece of data belongs, and the tag value of the second tag in each piece of data;
[0380] The processing unit 602 is further configured to store the multiple pieces of data based on the storage order of the multiple pieces of data.
[0381] Optionally, the detailed implementation process of the receiving unit 601 receiving multiple data can refer to the relevant content of step 204 of the method 200 shown in Figure 2, and will not be described in detail here.
[0382] Optionally, the detailed implementation process of the processing unit 602 determining the tag value set to which the tag value of the first tag in each piece of data belongs can be referred to the relevant content of step 205 of the method 200 shown in FIG. 2 , which will not be described in detail here.
[0383] Optionally, the detailed implementation process of the processing unit 602 determining the storage order of the multiple data can refer to the relevant content of step 206 of the method 200 shown in FIG2 , which will not be described in detail here.
[0384] Optionally, the processing unit 602 stores the multiple pieces of data based on the storage order of the multiple pieces of data. For a detailed implementation process, reference may be made to the relevant content of step 207 of the method 200 shown in FIG. 2 , which will not be described in detail here.
[0385] Optionally, the receiving unit 601 is further configured to:
[0386] Data structure information is received, where the data structure information is used to specify a first tag and a second tag as sorting keys, and the data structure information includes quantization identification information, where the quantization identification information is used to indicate the first tag.
[0387] Optionally, the detailed implementation process of the receiving unit 601 receiving the data structure information can refer to the relevant content of step 204 of the method 200 shown in FIG2 , which will not be described in detail here.
[0388] Optionally, the processing unit 602 is configured to:
[0389] Based on the granularity of the tag value set and the tag value of the first tag in each piece of data, determine the tag value set to which the tag value of the first tag in each piece of data belongs, wherein the quantitative identification information is also used to indicate the granularity of the tag value set, or the granularity of the tag value set is determined based on the characteristics of the tag value of the first tag in the stored data.
[0390] Optionally, the processing unit 602 determines the tag value set to which the tag value of the first tag in each piece of data belongs based on the granularity of the tag value set and the tag value of the first tag in each piece of data. For the detailed implementation process, please refer to the relevant content of step 205 of method 200 shown in Figure 2, which will not be described in detail here.
[0391] Optionally, the processing unit 602 is further configured to:
[0392] Determine the tags that need to be quantified. The tags that need to be quantified are the tags whose cardinality meets the high cardinality condition among the multiple tags.
[0393] When the first label is a label that needs to be quantified, the label value set to which the label value of the first label in each piece of data belongs is determined.
[0394] Optionally, the detailed implementation process of the processing unit 602 determining the labels that need to be quantified can refer to the relevant content of step 204 of the method 200 shown in FIG. 2 , which will not be described in detail here.
[0395] Optionally, when the first label is a label that needs to be quantified, the processing unit 602 determines the label value set to which the label value of the first label in each piece of data belongs. The detailed implementation process can be referred to the relevant content of step 205 of method 200 shown in Figure 2, which is not described in detail here.
[0396] Optionally, the tag value set to which the tag value of the first tag in each piece of data belongs is a tag value range that includes the tag value of the first tag in each piece of data, and the granularity of the tag value set is the length of the tag value range.
[0397] Optionally, the processing unit 602 is further configured to:
[0398] Index information of the multiple pieces of data is constructed, where the index information includes identification information of a label value set to which a label value of a first label in at least one selected piece of data belongs, wherein two adjacent selected pieces of data are separated by one or more pieces of data in a storage order.
[0399] Optionally, the detailed implementation process of the processing unit 602 constructing the index information of the multiple data can refer to the relevant content of step 208 of the method 200 shown in FIG2 , which will not be described in detail here.
[0400] Optionally, the receiving unit 601 is further configured to receive a query request, where the query request includes a query condition related to the first tag;
[0401] The processing unit 602 is further configured to determine data to be scanned from the plurality of pieces of data according to the query condition and identification information of the tag value set to which the tag value of the first tag in the index information belongs;
[0402] The processing unit 602 is further configured to query the data to be scanned to obtain query results that meet the query conditions.
[0403] Optionally, the detailed implementation process of the receiving unit 601 receiving the query request can refer to the relevant content of step 502 of the method 500 shown in FIG5 , and will not be described in detail here.
[0404] Optionally, the detailed implementation process of the processing unit 602 determining the data to be scanned among the multiple pieces of data can refer to the relevant content of step 502 of the method 500 shown in FIG5 , which will not be described in detail here.
[0405] Optionally, the detailed implementation process of the processing unit 602 querying the data to be scanned to obtain the query results that meet the query conditions can be found in the relevant content of step 503 of the method 500 shown in FIG5 , which will not be described in detail here.
[0406] In an embodiment of the present application, for each piece of data, the tag value set to which the tag value of the first tag in the data belongs includes not only the tag value but also other tag values of the first tag, so that the cardinality of the tag value set to which the tag value of the first tag in each piece of data belongs is much smaller than the cardinality of the tag value of the first tag in each piece of data. Since the order of the first tag in the sorting key is before the second tag, the processing unit determines the storage order of the multiple pieces of data based on the order of the sorting key, the identification information of the tag value set to which the tag value of the first tag in each piece of data belongs, and the tag value of the second tag in each piece of data. After the processing unit stores the multiple pieces of data in the storage order of the multiple pieces of data, the local orderliness of the tag values of the second tags in the multiple pieces of data can be improved, that is, the orderliness of the data stored in the database can be improved.
[0407] 7 , an embodiment of the present application provides a computing device 700. For example, the computing device 700 may be a device in the database system 102 included in the network architecture 10 shown in FIG1 , or the computing device 700 may be a device in the database system in the method 200 shown in FIG2 or the method 500 shown in FIG5 .
[0408] As shown in Figure 7, computing device 700 includes a bus 702, a processor 704, a memory 706, and a communication interface 708. Processor 704, memory 706, and communication interface 708 communicate with each other via bus 702. Computing device 700 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 700.
[0409] Bus 702 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG7 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 702 may include a path for transmitting information between various components of computing device 700 (e.g., processor 704, memory 706, and communication interface 708).
[0410] The processor 704 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0411] The memory 706 may include volatile memory, such as random access memory (RAM), or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0412] Referring to FIG7 , the memory 706 stores executable program code, and the processor 704 executes the executable program code to respectively implement the functions of the receiving unit 601 and the processing unit 602 in the apparatus 600 shown in FIG6 , thereby implementing the method provided by any of the above embodiments. In other words, the memory 706 stores instructions for executing the method provided by any of the above embodiments. Alternatively,
[0413] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or a communication network.
[0414] Embodiments of the present application also provide a data storage cluster. The data storage cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0415] As shown in Figure 8, the data storage cluster includes at least one computing device 700. The memory 706 of one or more computing devices 700 in the data storage cluster may store the same instructions for executing the method provided in any of the above embodiments.
[0416] In some possible implementations, the memory 706 of one or more computing devices 700 in the data storage cluster may also store partial instructions for executing the above-mentioned data storage method. In other words, the combination of one or more computing devices 700 can jointly execute instructions for executing the method provided in any of the above-mentioned embodiments.
[0417] In some possible implementations, one or more computing devices in the data storage cluster may be connected via a network. The network may be a wide area network (WAN), a local area network (LAN), or the like. FIG9 illustrates one possible implementation. As shown in FIG9 , two computing devices 700A and 700B are connected via a network. Specifically, each computing device is connected to the network via a communication interface within the computing device.
[0418] In this type of possible implementation, the memory 706 in the computing device 700A stores instructions for executing the functions of the receiving unit 601 in the embodiment shown in Figure 6. Simultaneously, the memory 706 in the computing device 700B stores instructions for executing the functions of the processing unit 602 in the embodiment shown in Figure 6.
[0419] It should be understood that the functionality of the computing device 700A shown in FIG9 may also be implemented by multiple computing devices 700. Similarly, the functionality of the computing device 700B may also be implemented by multiple computing devices 700.
[0420] Embodiments of the present application also provide another data storage cluster. The connection relationship between the computing devices in this data storage cluster can be similar to the connection method of the data storage cluster described in Figure 9. The difference is that the memory 706 of one or more computing devices 700 in this data storage cluster can store the same instructions for executing the methods provided in any of the above embodiments.
[0421] In some possible implementations, the memory 706 of one or more computing devices 700 in the data storage cluster may also store partial instructions for executing the method provided in any of the above embodiments. In other words, a combination of one or more computing devices 700 can jointly execute instructions for executing the method provided in any of the above embodiments.
[0422] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method provided in any of the above embodiments.
[0423] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the method provided in any of the above embodiments.
[0424] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0425] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for storing data, characterized in that: The method comprises: Receiving a plurality of pieces of data, wherein each piece of data includes a label value of a plurality of labels, the plurality of labels including a first label and a second label designated as a sort key, the first label being before the second label in the sort key; Determine the tag value set to which the tag value of the first tag in each piece of data belongs; Determine the storage order of the plurality of data based on the order of the sorting keys, the identification information of the tag value set to which the tag value of the first tag in each piece of data belongs, and the tag value of the second tag in each piece of data; The plurality of pieces of data are stored based on a storage order of the plurality of pieces of data.
2. The method according to claim 1, characterized in that Before receiving the plurality of pieces of data, the method further includes: Data structure information is received, where the data structure information is used to specify the first label and the second label as sorting keys, and the data structure information includes quantization identification information, where the quantization identification information is used to indicate the first label.
3. The method according to claim 2, characterized in that The determining the tag value set to which the tag value of the first tag in each piece of data belongs includes: Based on the granularity of the label value set and the label value of the first label in each piece of data, determine the label value set to which the label value of the first label in each piece of data belongs, wherein the quantitative identification information is also used to indicate the granularity of the label value set, or the granularity of the label value set is determined based on the characteristics of the label value of the first label in the stored data.
4. The method according to claim 1, characterized in that: The method further comprises: Determining a label that needs to be quantified, where the label that needs to be quantified is a label whose cardinality of label value satisfies a high cardinality condition among the multiple labels; When the first label is a label that needs to be quantified, determine the label value set to which the label value of the first label in each piece of data belongs.
5. The method according to claim 3, characterized in that: The tag value set to which the tag value of the first tag in each piece of data belongs is a tag value range that includes the tag value of the first tag in each piece of data, and the granularity of the tag value set is the length of the tag value range.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Constructing index information of the plurality of data, the index information comprising identification information of a tag value set to which a tag value of a first tag in at least one selected data belongs, wherein two adjacent selected data are separated by one or more data in storage order.
7. The method according to claim 6, characterized in that The method further comprises: receiving a query request, wherein the query request includes a query condition related to the first tag; Determine the data to be scanned among the multiple pieces of data according to the query condition and identification information of the tag value set to which the tag value of the first tag in the index information belongs; The data to be scanned is queried to obtain query results that meet the query conditions.
8. A device for storing data, characterized in that: The device comprises: A receiving unit, configured to receive a plurality of pieces of data, wherein each piece of data includes label values of a plurality of labels, the plurality of labels including a first label and a second label designated as a sorting key, the first label being prior to the second label in the sorting key; a processing unit, configured to determine a tag value set to which a tag value of the first tag in each piece of data belongs; The processing unit is further configured to determine a storage order of the plurality of pieces of data based on an order of the sorting keys, identification information of a label value set to which a label value of a first label in each piece of data belongs, and a label value of a second label in each piece of data; The processing unit is further configured to store the multiple pieces of data based on a storage order of the multiple pieces of data.
9. The device according to claim 8, characterized in that The receiving unit is further used for: Receive data structure information, the data structure information is used to specify the first label and the second label as sorting keys, the data structure information is used to specify the first label and the second label as sorting keys, The structure information includes quantified identification information, and the quantified identification information is used to indicate the first label.
10. The device according to claim 9, characterized in that The processing unit is used for: Based on the granularity of the label value set and the label value of the first label in each piece of data, determine the label value set to which the label value of the first label in each piece of data belongs, wherein the quantitative identification information is also used to indicate the granularity of the label value set, or the granularity of the label value set is determined based on the characteristics of the label value of the first label in the stored data.
11. The device according to claim 8, characterized in that The processing unit is further used for: Determining a label that needs to be quantified, where the label that needs to be quantified is a label whose cardinality of label value satisfies a high cardinality condition among the multiple labels; When the first label is a label that needs to be quantified, determine the label value set to which the label value of the first label in each piece of data belongs.
12. The device according to claim 11, characterized in that The tag value set to which the tag value of the first tag in each piece of data belongs is a tag value range that includes the tag value of the first tag in each piece of data, and the granularity of the tag value set is the length of the tag value range.
13. The device according to any one of claims 8 to 12, characterized in that The processing unit is further used for: Constructing index information of the plurality of data, the index information comprising identification information of a tag value set to which a tag value of a first tag in at least one selected data belongs, wherein two adjacent selected data are separated by one or more data in storage order.
14. The device according to claim 13, characterized in that The receiving unit is further configured to receive a query request, wherein the query request includes a query condition related to the first tag; The processing unit is further configured to determine the data to be scanned from among the plurality of pieces of data according to the query condition and identification information of the tag value set to which the tag value of the first tag in the index information belongs; The processing unit is further used to query the data to be scanned to obtain query results that meet the query conditions.
15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.
17. A computer program product comprising instructions, characterized in that When the instruction is executed by the computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data storage method and device and storage medium
CN120030012A
Index construction method, data query method and related equipment
CN116126864A
Automatic results caching for dynamically generated queries
EP4141691A1
Methods, systems, and computer program products for storing data in collections of tagged data pieces
US20020112116A1
Method and System for Providing Pre-Approved A / A Data Buckets
US20190057108A1