Data storage method and device and storage medium

By specifying multiple tags as sorting keys and determining the storage order of data based on these information, the problem of poor local order in the data in the database is solved, and query efficiency is improved.

CN120030012APending Publication Date: 2025-05-23HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410372304.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-03-27
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The local order of the data stored in the database is poor, resulting in inefficient query.

Method used

The storage order of data is determined by specifying a plurality of tags as sorting keys, and based on the order of the sorting keys, the identification information of the set of tag values ​​to which the tag values ​​of the first tag in each piece of data belong, and the tag values ​​of the second tag in each piece of data.

Benefits of technology

Improve the orderly storage of data in the database, thereby improving query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030012A_ABST
    Figure CN120030012A_ABST
Patent Text Reader

Abstract

The invention discloses a data storage method and device and a storage medium, and belongs to the field of communication. The method includes: receiving a plurality of pieces of data, each piece of data including tag values of a plurality of tags, the plurality of tags including a first tag and a second tag designated as a sorting key, the order of the first tag in the sorting key being before the order of the second tag; determining a label value set to which the label value of the first label in each piece of data belongs; determining a storage sequence of the multiple pieces of data based on a sequence of the sorting keys, identification information of a label value set to which a label value of a first label in each piece of data belongs, and a label value of a second label in each piece of data; and storing the multiple pieces of data based on the storage sequence of the multiple pieces of data. According to the method, the storage orderliness of the data in the database can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese patent application No. 202311558370.2, filed on November 21, 2023, and entitled “A Quantitative Cluster Indexing Method”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of communications, and in particular to a method, device and storage medium for storing data. Background Art

[0003] A database is a warehouse for storing data. The storage space of a database is often very large and can store millions, tens of millions, or even hundreds of millions of data.

[0004] For each piece of data stored in the database, the data includes a set of tag values ​​of multiple tags. For a certain tag among the multiple tags, the tag value of the tag may include a large number of different tag values. When the same tag value is stored in a relatively scattered manner, the local orderliness of the data stored in the database will be poor.

[0005] The local orderliness of the data stored in the database is poor, which may affect the operation of processing the data. For example, if you need to query the data including the tag value from the database, due to the poor local orderliness of the data, the required data range is large and the efficiency of querying the data is very low. Therefore, how to improve the orderliness of data stored in the database is an urgent problem to be solved. Summary of the invention

[0006] The present application provides a method, device and storage medium for storing data to improve the orderliness of data stored in a database. The technical solution is as follows:

[0007] In a first aspect, the present application provides a method for storing data, in which a plurality of pieces of data are received, wherein each piece of data includes a label value of a plurality of labels, the plurality of labels including a first label and a second label designated as a sorting key, the first label being in an order before the second label in the sorting key. The label value set to which the label value of the first label in each piece of data belongs is determined. Based on the order of the sorting keys, identification information of the label value set to which the label value of the first label in each piece of data belongs, and the label value of the second label in each piece of data, the storage order of the plurality of pieces of data is determined. Based on the storage order of the plurality of pieces of data, the plurality of pieces of data are stored.

[0008] For each piece of data, the label value set to which the label value of the first label in the data belongs includes not only the label value but also other label values ​​of the first label, so that the cardinality of the label value set to which the label value of the first label in each piece of data belongs is much smaller than the cardinality of the label value of the first label in each piece of data. Since the order of the first label in the sorting key is before the second label, the storage order of the multiple pieces of data is determined based on the order of the sorting key, the identification information of the label value set to which the label value of the first label in each piece of data belongs, and the label value of the second label in each piece of data. After storing the multiple pieces of data in the storage order of the multiple pieces of data, the local orderliness of the label values ​​of the second labels in the multiple pieces of data can be improved, that is, the orderliness of the data stored in the database can be improved.

[0009] In a possible implementation, data structure information is received, the data structure information is used to specify the first label and the second label as sorting keys, and the data structure information includes quantization identification information, the quantization identification information is used to indicate the first label. In this way, based on the quantization identification information, it can be known which label value in each piece of data needs to be quantized to determine the label value set to which the label value belongs.

[0010] In another possible implementation, based on the granularity of the label value set and the label value of the first label in each piece of data, the label value set to which the label value of the first label in each piece of data belongs is determined, wherein the quantitative identification information is also used to indicate the granularity of the label value set, or the granularity of the label value set is determined based on the characteristics of the label value of the first label in the stored data, thereby improving the flexibility of obtaining the granularity of the label value set.

[0011] In another possible implementation, a label that needs to be quantified is determined, and the label that needs to be quantified is a label whose cardinality of label value satisfies a high cardinality condition among the multiple labels. When the first label is a label that needs to be quantified, the label value set to which the label value of the first label in each piece of data belongs is determined. Since the label whose cardinality of label value satisfies the high cardinality condition is taken as the label that needs to be quantified, the label whose cardinality of label value does not satisfy the high cardinality condition will not be taken as the label that needs to be quantified, thereby reducing the number of labels that need to be quantified and saving computing resources.

[0012] In another possible implementation, the tag value set to which the tag value of the first tag in each piece of data belongs is a tag value range that includes the tag value of the first tag in each piece of data, and the granularity of the tag value set is the length of the tag value range.

[0013] In another possible implementation, index information of the multiple pieces of data is constructed, and the index information includes identification information of the tag value set to which the tag value of the first tag in the at least one selected piece of data belongs, wherein two adjacent selected pieces of data are separated by one or more pieces of data in the storage order. Thus, when querying data, the data range to be scanned can be determined through the index information, and querying within the data range can improve query efficiency compared to querying from the entire database.

[0014] In another possible implementation, a query request is received, the query request including a query condition related to the first tag. According to the query condition and the identification information of the tag value set to which the tag value of the first tag in the index information belongs, the data to be scanned in the multiple pieces of data are determined. The data to be scanned are queried to obtain a query result that satisfies the query condition. Among them, according to the query condition and the identification information of the tag value set to which the tag value of the first tag in the index information belongs, the data to be scanned in the multiple pieces of data are determined, which can reduce the range of data to be scanned, thereby improving the query efficiency.

[0015] In a second aspect, the present application provides a device for storing data, which is used to execute the method in the first aspect or any possible implementation of the first aspect. Specifically, the device includes a unit for executing the method in the first aspect or any possible implementation of the first aspect.

[0016] In a third aspect, the present application provides a computing device cluster, the computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory;

[0017] The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method in the first aspect or any possible implementation manner of the first aspect.

[0018] In a fourth aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes the computing device cluster to execute the method in the first aspect or any possible implementation manner of the first aspect.

[0019] In a fifth aspect, the present application provides a computer-readable storage medium, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method in the first aspect or any possible implementation of the first aspect.

[0020] In a sixth aspect, the present application provides a chip comprising a memory and a processor, the memory being used to store computer instructions, and the processor being used to call and run the computer instructions from the memory to execute the method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a schematic diagram of a network architecture provided by an embodiment of the present application;

[0022] Figure 2 is a flow chart of a method for storing data provided by an embodiment of the present application;

[0023] Figure 3 is a schematic diagram of a first editing interface provided in an embodiment of the present application;

[0024] Figure 4 is a schematic diagram of another second editing interface provided in an embodiment of the present application;

[0025] Figure 5 It is a flow chart of a method for querying data provided by an embodiment of the present application;

[0026] Figure 6 is a schematic diagram of the structure of a data storage device provided in an embodiment of the present application;

[0027] Figure 7 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0028] Figure 8 This is a schematic diagram of a cluster structure for storing data provided in an embodiment of the present application;

[0029] Fig. 9 This is a schematic diagram of another cluster structure for storing data provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] See also Figure 1 An embodiment of the present application provides a network architecture 100, which includes a terminal device 101 and a database system 102, and the terminal device 101 communicates with the database system 102.

[0031] A technician may configure multiple tags and data structure information of the database on the terminal device 101 , where the data structure information is used to specify at least two tags among the multiple tags as sorting keys.

[0032] That is, the sort key includes multiple tags that need to be sorted, and the multiple tags are in order in the sort key. For example, the multiple tags in the sort key include a first tag and a second tag, and the first tag is before the second tag in the sort key.

[0033] In some embodiments, the data structure information further includes quantization identification information, and the quantization identification information is used to indicate the first tag. Optionally, the quantization identification information indicates that the tag value of the first tag needs to be quantized to obtain identification information of the tag value set to which the tag value of the first tag belongs.

[0034] The terminal device 101 is used to obtain the configured multiple tags and the data structure information, and send a database creation request to the database system 102, where the database creation request includes the multiple tags and the data structure information.

[0035] The database system 102 is used to receive the database creation request, and create a database based on the multiple tags included in the database creation request.

[0036] The database is used to store multiple pieces of data, each piece of data includes the configured tag values ​​of the multiple tags.

[0037] In some embodiments, the database may be a data table, and the data table uses a column storage method to store the multiple pieces of data. For each piece of data in the multiple pieces of data, the piece of data may be a row of data in the data table.

[0038] The data table is a data table including multiple columns, and the multiple columns included in the data table correspond to the multiple configured tags.

[0039] Optionally, the multiple configured tags correspond to the multiple columns one-to-one, and for each tag in the multiple configured tags, the column corresponding to the tag is used to store one or more tag values ​​belonging to the tag.

[0040] In some embodiments, the database may be a file system, and the database may use files to store the plurality of data.

[0041] In some embodiments, the terminal device 101 may also send a storage request to the database system 102 , where the storage request includes multiple pieces of data to be stored. The database system 102 receives the storage request and saves the multiple pieces of data into the database.

[0042] In some embodiments, the data structure information defines a sort key, and the sort key includes multiple tags that need to be sorted, so that multi-level sorting can be performed according to the order of the multiple tags defined by the sort key. After receiving the storage request, the database system 102 performs multi-level sorting on the multiple data based on the tag value of each tag in the multiple tags that need to be sorted included in each of the multiple data to obtain the storage order of the multiple data. Based on the storage order of the multiple data, the multiple data are stored in the database.

[0043] For data structure information, specifying a tag as a sort key means sorting the tag values ​​of the tag. Specifying which tags are sort keys means that after receiving multiple pieces of data to be stored, the storage order of the multiple pieces of data is determined by the tag values ​​of these tags that serve as sort keys. In implementation, since a piece of data is a whole, sorting the tag values ​​of the tag that serves as the sort key is actually sorting the multiple pieces of data as a whole. Then, when the sort key has multiple tags, the order of the multiple tags in the sort key determines which tag value to sort the multiple pieces of data first.

[0044] When the sort key has multiple labels that need to be sorted, multi-level sorting is performed in the order of the multiple labels. For example, the order of the first label in the sort key is before the second label. After receiving multiple pieces of data to be stored, the database system 102 sorts based on the label value of the first label included in each piece of data to obtain the storage order of the multiple pieces of data, and classifies the multiple pieces of data including the same label value of the first label into one category after sorting. When sorting based on the label value of the second label included in each piece of data, the multiple pieces of data classified into one category (including multiple pieces of data with the same label value of the first label) are sorted, and the relative order between different categories does not change. Therefore, when performing multi-level sorting, first-level sorting is performed according to the label value of the first sort key, and then second-level sorting is performed according to the label value of the second sort key. The second level is to sort the data in the class that has been sorted in the first level again.

[0045] For example, the number of tags to be sorted included in the sort key may be n, where n is an integer greater than or equal to 1.

[0046] For the first tag among the n tags, the database system 102 performs a primary sort on each piece of data based on the tag value of the first tag to be sorted included in each piece of data among the multiple pieces of data, so as to classify the data including the same tag value of the first tag into one category. For the second tag among the n tags, the database system 102 performs a secondary sort on each piece of data in each class obtained through the primary sorting based on the tag value of the second tag to be sorted included in each piece of data, so as to classify the data including the same tag value of the second tag in the same class of data obtained through the primary sorting into one category. ... For the nth tag among the n tags, the database system 102 performs an n-level sort on each piece of data in each class obtained through the n-1-level sorting based on the tag value of the nth tag to be sorted included in each piece of data, so as to classify the data including the same tag value of the nth tag in the same class of data obtained through the n-1-level sorting into one category, and obtain the storage order of the multiple pieces of data.

[0047] In some embodiments, the database creation request also includes identification information of the database (such as the name of the database, etc.).

[0048] For example, the database system 102 receives a database creation request as shown below.

[0049] Create Table information(Operator, Direction, Address, Time, Bps)

[0050] SORTKEY Operator, Time, Address.

[0051] The database creation request includes five tags in the database "information", and the five tags correspond one to one with the five columns of the database "information". The first column of the database "information" is used to store the tag value belonging to the tag "Operator", the second column is used to store the tag value belonging to the tag "Direction", the third column is used to store the tag value belonging to the tag "Address", the fourth column is used to store the tag value belonging to the tag "Time", and the fifth column is used to store the tag value belonging to the tag "Bps".

[0052] The database creation request also includes data structure information, which defines three tags that need to be sorted, and the order of sorting the three tags is Operator, Time, and Address. That is to say, for multiple pieces of data to be saved in the database, when sorting the multiple pieces of data, first sort each piece of data based on the tag value of the tag "Operator" included in each piece of data. Then, based on the tag value of the tag "Time" included in each piece of data after sorting, perform a second-level sort. Finally, based on the tag value of the tag "Address" included in each piece of data after sorting, perform a third-level sort to obtain the storage order of each piece of data. Each piece of data is saved in the database according to the storage order of each piece of data.

[0053] For example, the database system 102 receives a storage request as shown below, which includes twenty pieces of data.

[0054] Insert into information(Operator, Direction, Address, Time, Bps)Value(

[0055] Operator 1,out,192.163.13.1,1:03:53,60;

[0056] Operator 1, out, 192.163.13.5, 1:03:52, 40;

[0057] Operator 1,out,192.163.13.2,1:03:55,59;

[0058] Operator 1,out,192.163.13.1,1:03:56,69;

[0059] Operator 2, out, 192.163.13.4, 7:02:40, 70;

[0060] Operator 2, out, 192.163.16.4, 7:22:40, 30;

[0061] Operator 1,in,192.163.13.2,9:03:49,102;

[0062] Operator 1,in,192.163.12.2,9:03:49,122;

[0063] Operator 1, out, 192.163.13.2, 1:03:50, 160;

[0064] Carrier 1,out,192.163.23.2,1:03:50,60;

[0065] Operator 2, out, 192.163.13.4, 3:02:52, 59;

[0066] Operator 2, out, 192.163.11.4, 3:02:52, 39;

[0067] Carrier 2, out, 192.163.13.5, 7:02:29, 80;

[0068] Carrier 2, out, 192.163.33.5, 7:02:39, 85;

[0069] Operator 1, out, 192.163.13.1, 1:03:43, 20;

[0070] Operator 1,out,192.163.19.1,1:03:43,30;

[0071] Operator 1,in,192.163.13.1,1:03:42,60;

[0072] Operator 1,in,192.163.113.1,1:03:42,20;

[0073] Operator 2, out, 192.163.13.4, 2:02:45, 90;

[0074] Operator 2, out, 192.163.33.4, 2:02:45, 80;

[0075] ).

[0076] The database system 102 first performs a primary sort on the 20 pieces of data based on the label value of "Operator" included in each of the 20 pieces of data, and obtains the following results. In the primary sort, data including the same label value of "Operator" are clustered together to form a class. As shown below, in the primary sort, the twelve pieces of data including operator 1 are clustered together to form class 1, and the eight pieces of data including operator 2 are clustered together to form class 2.

[0077] Operator 1,out,192.163.13.1,1:03:53,60;

[0078] Operator 1, out, 192.163.13.5, 1:03:52, 40;

[0079] Operator 1,out,192.163.13.2,1:03:55,59;

[0080] Operator 1,out,192.163.13.1,1:03:56,69;

[0081] Operator 1,in,192.163.13.2,9:03:49,102;

[0082] Operator 1,in,192.163.12.2,9:03:49,122;

[0083] Operator 1, out, 192.163.13.2, 1:03:50, 160;

[0084] Carrier 1,out,192.163.23.2,1:03:50,60;

[0085] Operator 1, out, 192.163.13.1, 1:03:43, 20;

[0086] Operator 1,out,192.163.19.1,1:03:43,30;

[0087] Operator 1,in,192.163.13.1,1:03:42,60;

[0088] Operator 1,in,192.163.113.1,1:03:42,20;

[0089] Operator 2, out, 192.163.13.4, 7:02:40, 70;

[0090] Operator 2, out, 192.163.16.4, 7:22:40, 30;

[0091] Operator 2, out, 192.163.13.4, 3:02:52, 59;

[0092] Operator 2, out, 192.163.11.4, 3:02:52, 39;

[0093] Carrier 2, out, 192.163.13.5, 7:02:29, 80;

[0094] Carrier 2, out, 192.163.33.5, 7:02:39, 85;

[0095] Operator 2, out, 192.163.13.4, 2:02:45, 90;

[0096] Operator 2, out, 192.163.33.4, 2:02:45, 80.

[0097] The database system 102 performs secondary sorting on the twenty pieces of data based on the label value of "Time" included in each of the twenty pieces of data. When performing the secondary sorting, the relative order between class 1 and class 2 remains unchanged, but the secondary sorting is performed on the twelve pieces of data in class 1 based on the label value of "Time" included in each of the twelve pieces of data. That is to say, the two data including "1:03:42" will be clustered into class 11, the two data including "1:03:43" will be clustered into class 12, the two data including "1:03:50" will be clustered into class 13, the one data including "1:03:52" will be clustered into class 14, the one data including "1:03:53" will be clustered into class 15, the two data including "1:03:55" will be clustered into class 16, the two data including "1:03:56" will be clustered into class 17, and the two data including "9:03:49" will be clustered into class 18.

[0098] Based on the label value of "Time" included in each of the eight pieces of data in class 2, the eight pieces of data are sorted at the second level. That is, the two pieces of data including "2:02:45" are clustered into class 21, the two pieces of data including "3:02:52" are clustered into class 22, the one piece of data including "7:02:29" is clustered into class 23, the one piece of data including "7:02:39" is clustered into class 24, the one piece of data including "7:02:40" is clustered into class 25, and the one piece of data including "7:22:40" is clustered into class 26, and the following results are obtained.

[0099] Operator 1,in,192.163.13.1,1:03:42,60;

[0100] Operator 1,in,192.163.113.1,1:03:42,20;

[0101] Operator 1, out, 192.163.13.1, 1:03:43, 20;

[0102] Operator 1,out,192.163.19.1,1:03:43,30;

[0103] Operator 1, out, 192.163.13.2, 1:03:50, 160;

[0104] Carrier 1,out,192.163.23.2,1:03:50,60;

[0105] Operator 1, out, 192.163.13.5, 1:03:52, 40;

[0106] Operator 1, out, 192.163.13.1, 1:03:53, 60;

[0107] Operator 1,out,192.163.13.2,1:03:55,59;

[0108] Operator 1,out,192.163.13.1,1:03:56,69;

[0109] Operator 1,in,192.163.13.2,9:03:49,102;

[0110] Operator 1,in,192.163.12.2,9:03:49,122;

[0111] Operator 2, out, 192.163.13.4, 2:02:45, 90;

[0112] Operator 2, out, 192.163.33.4, 2:02:45, 80;

[0113] Operator 2, out, 192.163.13.4, 3:02:52, 59;

[0114] Operator 2, out, 192.163.11.4, 3:02:52, 39;

[0115] Carrier 2, out, 192.163.13.5, 7:02:29, 80;

[0116] Carrier 2, out, 192.163.33.5, 7:02:39, 85;

[0117] Operator 2, out, 192.163.13.4, 7:02:40, 70;

[0118] Operator 2, out, 192.163.16.4, 7:22:40, 30.

[0119] The database system 102 sorts the twenty pieces of data into three levels based on the label value of "Address" included in each of the twenty pieces of data. When performing the three-level sorting, the relative order between class 11, class 12, class 13, class 14, class 15, class 16, class 17, class 18, class 21, class 22, class 23, class 24, class 25 and class 26 remains unchanged. And when performing three-level sorting, based on the label value of "Address" included in the two data in class 11, the two data are sorted in three levels; based on the label value of "Address" included in the two data in class 12, the two data are sorted in three levels; based on the label value of "Address" included in the two data in class 13, the two data are sorted in three levels; based on the label value of "Address" included in the two data in class 16, the two data are sorted in three levels; based on the label value of "Address" included in the two data in class 17, the two data are sorted in three levels; based on the label value of "Address" included in the two data in class 18, the two data are sorted in three levels; based on the label value of "Address" included in the two data in class 21, the two data are sorted in three levels; based on the label value of "Address" included in the two data in class 22, the two data are sorted in three levels. The storage order of the twenty data obtained after the three-level sorting is as follows.

[0120] Operator 1,in,192.163.13.1,1:03:42,60;

[0121] Operator 1,in,192.163.113.1,1:03:42,20;

[0122] Operator 1, out, 192.163.13.1, 1:03:43, 20;

[0123] Operator 1,out,192.163.19.1,1:03:43,30;

[0124] Operator 1, out, 192.163.13.2, 1:03:50, 160;

[0125] Carrier 1,out,192.163.23.2,1:03:50,60;

[0126] Operator 1, out, 192.163.13.5, 1:03:52, 40;

[0127] Operator 1, out, 192.163.13.1, 1:03:53, 60;

[0128] Operator 1,out,192.163.13.2,1:03:55,59;

[0129] Operator 1,out,192.163.13.1,1:03:56,69;

[0130] Operator 1,in,192.163.12.2,9:03:49,122;

[0131] Operator 1,in,192.163.13.2,9:03:49,102;

[0132] Operator 2, out, 192.163.13.4, 2:02:45, 90;

[0133] Operator 2, out, 192.163.33.4, 2:02:45, 80;

[0134] Operator 2, out, 192.163.11.4, 3:02:52, 39;

[0135] Operator 2, out, 192.163.13.4, 3:02:52, 59;

[0136] Carrier 2, out, 192.163.13.5, 7:02:29, 80;

[0137] Carrier 2, out, 192.163.33.5, 7:02:39, 85;

[0138] Operator 2, out, 192.163.13.4, 7:02:40, 70;

[0139] Operator 2, out, 192.163.16.4, 7:22:40, 30.

[0140] The database system 102 saves the twenty pieces of data into the database "information" as shown in Table 1 below based on the storage order of the twenty pieces of data.

[0141] Table 1

[0142] Serial number Operator Direction Address Time Bps 1 Operator 1 in 192.163.13.1 1:03:42 60 2 Operator 1 in 192.163.113.1 1:03:42 20 3 Operator 1 out 192.163.13.1 1:03:43 20 4 Operator 1 out 192.163.19.1 1:03:43 30 5 Operator 1 out 192.163.13.2 1:03:50 160 6 Operator 1 out 192.163.23.2 1:03:50 60 7 Operator 1 out 192.163.13.5 1:03:52 40 8 Operator 1 out 192.163.13.1 1:03:53 60 9 Operator 1 out 192.163.13.2 1:03:55 59 10 Operator 1 out 192.163.13.1 1:03:56 69 11 Operator 1 in 192.163.12.2 9:03:49 122 12 Operator 1 in 192.163.13.2 9:03:49 102 13 Operator 2 out 192.163.13.4 2:02:45 90 14 Operator 2 out 192.163.33.4 2:02:45 80 15 Operator 2 out 192.163.11.4 3:02:52 39 16 Operator 2 out 192.163.13.4 3:02:52 59 17 Operator 2 out 192.163.13.5 7:02:29 80 18 Operator 2 out 192.163.33.5 7:02:39 85 19 Operator 2 out 192.163.13.4 7:02:40 70 20 Operator 2 out 192.163.16.4 7:22:40 30

[0143] For some tags in the database, the tag is a tag whose cardinality of the tag value satisfies the high cardinality condition. The so-called tag whose cardinality of the tag value satisfies the high cardinality condition means that the tag has a large number of deduplicated tag values, and usually the number of deduplicated tag values ​​of the tag may exceed the quantity threshold. For example, the number of deduplicated tag values ​​of the tag may be thousands, tens of thousands, millions, tens of millions, or hundreds of millions. A high cardinality column in the database refers to a column of tag values ​​in the database whose cardinality satisfies the high cardinality condition.

[0144] In the case of multiple sorting keys, multi-level sorting is required. At this time, when there is a high base column in front of a column, after sorting the data by the label value of the high base column, the order of the latter column will be very poor when sorting the latter column. As mentioned above, the first label and the second label, the order of the first label in the sorting key is before the second label, and the label value of the column corresponding to the first label is a high base column. Since the data is sorted by the label value of the high base column first, and then sorted by the label value of the column corresponding to the second label, the order of the label value of the column corresponding to the second label is very poor.

[0145] For example, for the database shown in Table 1 above, the tag "Address" and the tag "Time" in the database, the tag value of the tag "Address" after deduplication is a large number of addresses, and there may be hundreds of millions or billions of different addresses. The tag value of the tag "Time" after deduplication is a large number of timestamps, and there may be hundreds of millions or billions of different timestamps. Therefore, each piece of data is sorted based on the timestamp and address in each piece of data. Refer to Table 1. For the first piece of data with sequence number 1, the third piece of data with sequence number 3, the fifth piece of data with sequence number 5, the seventh piece of data with sequence number 7, the ninth piece of data with sequence number 9, the tenth piece of data with sequence number 10, and the twelfth piece of data with sequence number 12, these data are all addresses within the 192.163.13 network segment, but these data are scattered and stored in the data table shown in Table 1. The 13th data with sequence number 13, the 16th data with sequence number 16, the 17th data with sequence number 17 and the 19th data with sequence number 19 are also addresses in the 192.163.13 network segment, but these data are scattered and stored in the database shown in Table 1. Therefore, the local order of the address corresponding to the label "Address" included in each data after sorting is poor, so if you need to query the data of a certain continuous address in the 192.163.13 network segment, the scanned data range will be very large, which will increase the difficulty of query and reduce the query efficiency.

[0146] In order to improve the local orderliness of data stored in the database, data is stored in the database by any of the following embodiments. Then, the database is processed to improve the efficiency of processing the database. For example, querying data from the database can improve query efficiency.

[0147] See also Figure 2 The present application embodiment provides a method 200 for storing data, wherein the method 200 is applied to Figure 1 The network architecture 100 shown includes the following processes.

[0148] Step 201: The terminal device sends a database creation request to the database system. The database creation request includes multiple tags and data structure information. The data structure information is used to specify at least two tags among the multiple tags as sorting keys.

[0149] The at least two tags include a first tag and a second tag, and the first tag is before the second tag in the sort key.

[0150] In some embodiments, the data structure information further includes quantization identification information, which is used to indicate the first tag. Optionally, the quantization identification information indicates that the tag value of the first tag needs to be quantized to obtain identification information of the tag value set to which the tag value of the first tag belongs.

[0151] In some embodiments, for a first tag among the at least two tags that needs to be quantized, a mark may be added to the first tag, and the quantization identification information includes the mark added to the first tag. Optionally, the mark may be the granularity of the tag value set to which the tag value of the first tag belongs, or the mark may be an automatic mark that automatically determines the granularity of the tag value set to which the tag value of the first tag belongs.

[0152] Optionally, the database creation request may also include identification information of the database. Optionally, the identification information of the database may include the name of the database, etc.

[0153] In some embodiments, for the tag value set to which the tag value of the first tag belongs, the tag value set includes multiple tag values ​​belonging to the first tag. For example, assuming that the first tag is a timestamp, the granularity of the tag value set is 1 minute. For the timestamp 7:05:40, the tag value set to which the timestamp 7:05:40 belongs is a set greater than or equal to 7:05:01 and less than or equal to 7:05:59. The timestamp 7:05:40 is quantized, and the tag value set obtained is the tag value set whose identification information can be "7:05". For the timestamp 7:06:25, the tag value set to which the timestamp 7:06:25 belongs is a set greater than or equal to 7:06:01 and less than or equal to 7:06:59. The timestamp 7:06:25 is quantized, and the tag value set obtained is the tag value set whose identification information can be "7:06".

[0154] For another example, still assuming that the first label is a timestamp, the granularity of the label value set is 10 minutes. For the timestamp 7:15:40, the label value set to which the timestamp 7:15:40 belongs is a set greater than or equal to 7:10:01 and less than or equal to 7:19:59. After the timestamp 7:15:40 is quantized, the label value set obtained is the label value set whose identification information can be "7:10". For the timestamp 7:25:25, the label value set to which the timestamp 7:25:25 belongs is a set greater than or equal to 7:20:01 and less than or equal to 7:29:59. After the timestamp 7:25:25 is quantized, the label value set obtained is the label value set whose identification information can be "7:20".

[0155] In some embodiments, the first tag indicated by the quantitative identification information may include one tag or multiple tags.

[0156] In step 201, the terminal device may display a first editing interface, and the user may enter a database creation statement in the first editing interface. The terminal device obtains the database creation statement from the first editing interface and uses the database creation statement as the database creation request.

[0157] For example, the terminal device may display Figure 3 In the first editing interface shown, the user can enter the database creation statement 1 shown below in the first editing interface. The database creation statement 1 includes the name of the database "information", five tags and data structure information. The five tags are "Operator", "Direction", "Address", "Time" and "Bps". The data structure information is used to indicate that the three tags that need to be sorted are sorting keys, and the order in which the three tags need to be sorted is "Operator", "Time", and "Address". In one embodiment, the data structure information also includes quantification identification information. The quantification identification information can be the granularity "1m" of the tag value set marked on the first tag "Time". The quantification identification information indicates that the first tag that needs to be quantified is "Time". Therefore, the granularity "1m" indicates that the tag that needs to be quantified is "Time" and also indicates the granularity.

[0158] Database creation statement 1:

[0159] Create Table information(Operator, Direction, Address, Time, Bps)

[0160] SORTKEY Operator, direction, Time (1m), Address.

[0161] The terminal device can be Figure 3 In the first editing interface shown, a database creation statement 1 is obtained, and the database creation statement 1 is used as the database creation request.

[0162] For another example, see Figure 4 In the first editing interface shown in FIG. 1 , the user can enter the database creation statement 2 shown below in the first editing interface. Figure 3 The database creation statement 1 in the first editing interface shown, Figure 4 The first tag "Time" in the database creation statement 2 in the first editing interface shown is marked with an automatic tag "auto", which is used to indicate that the database system determines the granularity of the tag value set of the first tag "Time" based on the characteristics of the tag value of the first tag "Time".

[0163] Database creation statement 2:

[0164] Create Table information(Operator, Direction, Address, Time, Bps)

[0165] SORTKEY Operator, direction, Time (auto), Address.

[0166] For the first tag marked with automatic marking, the database system may determine the granularity according to the characteristics of the first tag data stored by statistics. For example, the granularity may be determined according to the distribution characteristics of the first tag data stored by statistics.

[0167] Step 202: The database system receives the database creation request and creates a database based on a plurality of tags included in the database creation request.

[0168] In step 202, the database creation request includes the multiple tags, and a database is created based on the multiple tags. Optionally, when the database is a data table, the multiple tags correspond to multiple columns included in the data table, and the columns corresponding to each tag are used to store the tag value of each tag.

[0169] In some embodiments, each tag is column identification information of a column corresponding to each tag.

[0170] In some embodiments, the database creation request also includes identification information of the database. After creating the database, the database system sets the identification information of the database to the identification information of the database included in the database creation request.

[0171] For example, the database system receives the above database creation statement 1 or database creation statement 2, and creates a database "information" based on the five tags included in the database creation statement 1 or the five tags included in the database creation statement 2. The database "information" includes five columns, namely, the first column corresponding to the tag "Operator", the second column corresponding to the tag "Direction", the third column corresponding to the tag "Address", the fourth column corresponding to the tag "Time" and the fifth column corresponding to the tag "Bps". "Operator", "Direction", "Address", "Time" and "Bps" are the column names of the five columns respectively.

[0172] After creating the database, users can save data to the database. The detailed implementation process is as follows.

[0173] Step 203: The terminal device sends a storage request to the database system, where the storage request includes M pieces of data. For each piece of data, the data includes tag values ​​of the multiple tags, where the multiple tags include a first tag and a second tag designated as sort keys, where the first tag is before the second tag in the sort key, and M is an integer greater than 1.

[0174] In some embodiments, the storage request may be a data insert statement.

[0175] The user can input a data insertion statement into the second editing interface displayed by the terminal device, and the data insertion statement includes the identification information of the database and M pieces of data. The terminal device obtains the data insertion statement from the second editing interface and sends the data insertion statement to the database system.

[0176] For example, the user may input a data insertion statement in the second editing interface displayed on the terminal device, and the data insertion statement is a schematic example as shown below.

[0177] An example of a data insert statement:

[0178] Insert into information(Operator, Direction, Address, Time, Bps)Value(

[0179] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60;

[0180] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40;

[0181] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59;

[0182] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69;

[0183] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70;

[0184] Data 6: Carrier 2, out, 192.163.16.4, 7:22:40, 30;

[0185] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102;

[0186] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122;

[0187] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160;

[0188] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60;

[0189] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59;

[0190] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39;

[0191] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80;

[0192] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85;

[0193] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20;

[0194] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30;

[0195] Data 17: Operator 1, in, 192.163.13.1, 1:03:42, 60;

[0196] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20;

[0197] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90;

[0198] Data 20: Carrier 2, out, 192.163.33.4, 2:02:45, 80;

[0199] ).

[0200] The data insertion statement includes the identification information "information" of the database and twenty pieces of data (for the sake of convenience, in the example of the data insertion statement, the twenty pieces of data are referred to as data 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20), that is, M = 20. The terminal device obtains the data insertion statement from the second editing interface, uses the data insertion statement as a storage request, and sends the storage request to the database system.

[0201] Step 204: The database system receives the storage request and obtains the tag value of the first tag included in each piece of data in the storage request.

[0202] In step 204, for each piece of data included in the storage request, a tag value belonging to the first tag in the piece of data is determined, thereby obtaining the tag value of the first tag included in the piece of data.

[0203] The first label is the label indicated by the quantified identification information, and the database system determines the first label based on the quantified identification information, determines the label value of the first label in the data, and obtains the label value of the first label included in the data. Alternatively, the first label to be quantified is determined, the first label is a label whose cardinality among the multiple labels satisfies a high cardinality condition, determines the label value of the first label in the data, and obtains the label value of the first label included in the data.

[0204] In some embodiments, for the first tag that needs to be quantified, the database system can also count the frequency of using the first tag as a filtering condition in query requests received in the past. If the frequency is high, that is, the frequency exceeds the frequency threshold, the first tag can be used as a tag that needs to be quantified.

[0205] For example, the database system receives the data insertion statement listed in the above example, which includes data 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, and the first label indicated by the quantitative identification information is "Time".

[0206] For data 1, obtain the tag value "1:03:53" of the first tag "Time" included in data 1. For data 2, obtain the tag value "1:03:52" of the first tag "Time" included in data 2. For data 3, obtain the tag value "1:03:55" of the first tag "Time" included in data 3. ..., for data 20, obtain the tag value "2:02:45" of the first tag "Time" included in data 20.

[0207] Step 205: The database system determines the tag value set to which the tag value of the first tag included in each data belongs based on the tag value of the first tag included in each data, and obtains N tag value sets, where N is an integer greater than or equal to 1 and less than M.

[0208] In step 205, based on the tag value of the first tag included in each piece of data and the granularity of the tag value set of the first tag, it is determined that the tag of the first tag included in each piece of data belongs to the tag value set.

[0209] In some embodiments, the granularity of the tag value set of the first tag may be the granularity indicated by the quantization identification information.

[0210] In some embodiments, the granularity of the tag value set of the first tag may be a granularity determined based on a feature of the tag value of the first tag.

[0211] In some embodiments, the database system may have multiple quantization methods, and may select a corresponding quantization method based on the first label and granularity, and use the quantization method to quantify the label of the first label included in each piece of data to obtain a label value set belonging to the first label included in each piece of data.

[0212] For example, assuming that the label value of the first label that needs to be quantified is timestamp 7:05:40, if the first label is timestamp "Time" and the granularity is 1 minute, the database system can select the corresponding quantization method based on the first label "Time" and the granularity "1 minute", and the label value set obtained by quantizing the timestamp 7:05:40 using the quantization method is "7:05". Assuming that the label value of the first label that needs to be quantified is timestamp 7:15:40, if the first label is timestamp "Time" and the granularity is 10 minutes, the database system can select the corresponding quantization method based on the first label "Time" and the granularity "10 minutes", and the label value set obtained by quantizing the timestamp 7:15:40 using the quantization method is "7:10".

[0213] For another example, suppose that the label value of the first label that needs to be quantized is the address 192.168.13.4. If the first label is the address "Address" and the subnet mask with a granularity of 255.255.255.0, the database system can select a corresponding quantization method based on the first label "Address" and the subnet mask, and use the quantization method to mask the address 192.168.13.4 to obtain a label value set of "192.168.13".

[0214] The process of quantizing the label value of the first label is to divide the amplitude of the entire first label into a set of finite small amplitudes (quantization steps), classify the sample values ​​falling within a certain step into one category, and assign the same quantization value.

[0215] Therefore, in step 205, the database system divides all the tag values ​​of the first tag into levels, classifies the tag values ​​in the same level into a tag value set, and assigns the tag value set to obtain the identification information of the set. For the tag value of the first tag included in each piece of data, the database system can determine the tag value set to which the tag value belongs from the divided tag value sets, thereby obtaining the identification information of the tag value set.

[0216] For example, assuming that the first label is a timestamp, the granularity of the label value set is 1 minute, the amplitude of the entire timestamp is divided into multiple steps, and multiple label value sets are assigned to obtain identification information of the multiple label value sets. The identification information of the multiple label value sets obtained by assignment is 7:01, 7:02, 7:03, 7:04, 7:05, etc. For the timestamp 7:05:40, the timestamp 7:05:40 is quantized, and the label value set with the identification information "7:05" includes the timestamp 7:05:40, so the label value set obtained by quantizing the timestamp 7:05:40 is the label value set with the identification information "7:05".

[0217] In some embodiments, there may be multiple first labels that need to be quantified. For each piece of data, the label value set to which the label value of each first label included in the data belongs is determined based on the label value of each first label included in the data and the granularity corresponding to each first label.

[0218] For example, assuming that the first tag is "Time", in the above data 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, the tag value corresponding to the first tag "Time" is a timestamp. The characteristic of the timestamp is that the seconds in the timestamp are constantly changing. Therefore, the granularity of the tag value set of the first tag "Time" is determined to be 1m.

[0219] Assuming that the granularity of the label value set of the first label "Time" is 1m, for the above-mentioned data 1, based on the label value "1:03:53" of the first label "Time" included in data 1 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:53" of the first label "Time" included in data 1 belongs.

[0220] For the above-mentioned data 2, based on the label value "1:03:52" of the first label "Time" included in data 2 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:52" of the first label "Time" included in data 2 belongs.

[0221] For the above-mentioned data 3, based on the label value "1:03:55" of the first label "Time" included in data 3 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:55" of the first label "Time" included in data 3 belongs.

[0222] For the above-mentioned data 4, based on the label value "1:03:56" of the first label "Time" included in data 4 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:56" of the first label "Time" included in data 4 belongs.

[0223] For the above data 5, based on the label value "7:02:40" of the first label "Time" included in data 5 and the granularity "1m" of the label value set of the first label "Time", determine that the label value "1:03:56" of the first label "Time" included in data 5 belongs to the label value set "7:02".

[0224] For the above-mentioned data 6, based on the tag value "7:22:40" of the first tag "Time" included in data 6 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "7:22" to which the tag value "7:22:40" of the first tag "Time" included in data 6 belongs.

[0225] For the above data 7, based on the label value "9:03:49" of the first label "Time" included in data 7 and the granularity "1m" of the label value set of the first label "Time", determine that the label value "7:22:40" of the first label "Time" included in data 7 belongs to the label value set "9:03".

[0226] For the above data 8, based on the tag value "9:03:49" of the first tag "Time" included in data 8 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "9:03" to which the tag value "9:03:49" of the first tag "Time" included in data 8 belongs.

[0227] For the above-mentioned data 9, based on the tag value "1:03:50" of the first tag "Time" included in data 9 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "1:03" to which the tag value "1:03:50" of the first tag "Time" included in data 9 belongs.

[0228] For the above-mentioned data 10, based on the label value "1:03:50" of the first label "Time" included in the data 10 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:50" of the first label "Time" included in the data 10 belongs.

[0229] For the above-mentioned data 11, based on the tag value "3:02:52" of the first tag "Time" included in data 11 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "3:02" to which the tag value "3:02:52" of the first tag "Time" included in data 11 belongs.

[0230] For the above-mentioned data 12, based on the tag value "3:02:52" of the first tag "Time" included in data 12 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "3:02" to which the tag value "3:02:52" of the first tag "Time" included in data 12 belongs.

[0231] For the above-mentioned data 13, based on the tag value "7:02:29" of the first tag "Time" included in data 13 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "7:02" to which the tag value "7:02:29" of the first tag "Time" included in data 13 belongs.

[0232] For the above-mentioned data 14, based on the tag value "7:02:39" of the first tag "Time" included in data 14 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "7:02" to which the tag value "7:02:39" of the first tag "Time" included in data 14 belongs.

[0233] For the above-mentioned data 15, based on the label value "1:03:43" of the first label "Time" included in data 15 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "1:03" to which the label value "1:03:43" of the first label "Time" included in data 15 belongs.

[0234] For the above-mentioned data 16, based on the tag value "1:03:43" of the first tag "Time" included in data 16 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "1:03" to which the tag value "1:03:43" of the first tag "Time" included in data 16 belongs.

[0235] For the above-mentioned data 17, based on the tag value "1:03:42" of the first tag "Time" included in data 17 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "1:03" to which the tag value "1:03:42" of the first tag "Time" included in data 17 belongs.

[0236] For the above-mentioned data 18, based on the tag value "1:03:42" of the first tag "Time" included in data 18 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "1:03" to which the tag value "1:03:42" of the first tag "Time" included in data 18 belongs.

[0237] For the above-mentioned data 19, based on the label value "2:02:45" of the first label "Time" included in data 19 and the granularity "1m" of the label value set of the first label "Time", determine the label value set "2:02" to which the label value "2:02:45" of the first label "Time" included in data 19 belongs.

[0238] For the above-mentioned data 20, based on the tag value "2:02:45" of the first tag "Time" included in the data 20 and the granularity "1m" of the tag value set of the first tag "Time", determine the tag value set "2:02" to which the tag value "2:02:45" of the first tag "Time" included in the data 20 belongs.

[0239] As shown below, the tag value set to which the tag of the first tag included in each of the twenty pieces of data belongs is as follows:

[0240] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60; 1:03;

[0241] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40; 1:03;

[0242] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59; 1:03;

[0243] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69; 1:03;

[0244] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70; 7:02;

[0245] Data 6: Carrier 2, out, 192.163.16.4, 7:22:40, 30; 7:22;

[0246] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102; 9:03;

[0247] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122; 9:03;

[0248] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160; 1:03;

[0249] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60; 1:03;

[0250] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59; 3:02;

[0251] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39; 3:02;

[0252] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80; 7:02;

[0253] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85; 7:02;

[0254] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20; 1:03;

[0255] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30; 1:03;

[0256] Data 17: Operator 1, in, 192.163.13.1, 1:03:42, 60; 1:03;

[0257] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20; 1:03;

[0258] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90; 2:02;

[0259] Data 20: Operator 2, out, 192.163.33.4, 2:02:45, 80; 2:02.

[0260] Step 206: The database system determines the storage order of the M pieces of data based on the order of the sorting keys, identification information of the label value set to which the label value of the first label included in each piece of data belongs, and the label value of the second label included in each piece of data.

[0261] In step 206, at least two tags that need to be sorted are defined for the data structure information. If the order of the at least two tags included in the sort key is earlier than the tag of the first tag, then each piece of data is sorted based on the tag value of the tag included in each piece of data. Then, based on the identification information of the tag value set to which the tag value of the first tag value included in each piece of data belongs, each piece of data is sorted on the basis of the sorted data. Then, based on the tag value of the second tag included in each piece of data, each piece of data is sorted on the basis of the sorted data to obtain the storage order of the M pieces of data.

[0262] For example, the data structure information defines that at least two tags that need to be sorted include tags "Operater", "Time", and "Address". For the above 20 pieces of data, the order of the tag "Operator" in the sort key is before the tag "Time". The database system performs a primary sort on the 20 pieces of data based on the tag value of the tag "Operater" included in each piece of data. When performing the primary sort, the 12 pieces of data including operator 1 are clustered together to form class 1, and the eight pieces of data including operator 2 are clustered together to form class 2. The sorting results are as follows.

[0263] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60; 1:03;

[0264] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40; 1:03;

[0265] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59; 1:03;

[0266] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69; 1:03;

[0267] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102; 9:03;

[0268] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122; 9:03;

[0269] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160; 1:03;

[0270] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60; 1:03;

[0271] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20; 1:03;

[0272] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30; 1:03;

[0273] Data 17: Operator 1, in, 192.163.13.1, 1:03:42, 60; 1:03;

[0274] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20; 1:03;

[0275] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70; 7:02;

[0276] Data 6: Carrier 2, out, 192.163.16.4, 7:22:40, 30; 7:22;

[0277] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59; 3:02;

[0278] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39; 3:02;

[0279] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80; 7:02;

[0280] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85; 7:02;

[0281] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90; 2:02;

[0282] Data 20: Operator 2, out, 192.163.33.4, 2:02:45, 80; 2:02.

[0283] In the sorting key, the order of the label "Time" is before the label "Address". The database system performs secondary sorting on the twenty pieces of data based on the identification information of the label value set to which the label value of the first label "Time" included in each piece of data belongs, based on the above sorting results. When performing secondary sorting, the relative order between class 1 and class 2 remains unchanged, but based on the identification information of the label value set to which the label value of "Time" included in each of the twelve pieces of data in class 1 belongs, the twelve pieces of data are sorted at the secondary level. That is to say, data 1, 2, 3, 4, 9, 10, 15, 16, 17 and 18 including the identification information "1:03" of the label value set to which the label value of "Time" belongs are clustered to form class 11, and data 7 and 8 including the identification information "9:03" of the label value set to which the label value of "Time" belongs are clustered to form class 12.

[0284] Based on the identification information of the label value set to which the label value of "Time" included in each of the eight data in class 2 belongs, the eight data are sorted at the second level. That is, data 19 and 20 including the identification information "2:02" of the label value set to which the label value of "Time" belongs are clustered to form class 21, data 11 and 12 including the identification information "3:02" of the label value set to which the label value of "Time" belongs are clustered to form class 22, data 5, 13 and 14 including the identification information "7:02" of the label value set to which the label value of "Time" belongs are clustered to form class 23, and data 6 including the identification information "7:22" of the label value set to which the label value of "Time" belongs is clustered to form class 24. The sorting results are as follows.

[0285] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60; 1:03;

[0286] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40; 1:03;

[0287] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59; 1:03;

[0288] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69; 1:03;

[0289] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160; 1:03;

[0290] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60; 1:03;

[0291] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20; 1:03;

[0292] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30; 1:03;

[0293] Data 17: Operator 1, in, 192.163.13.1, 1:03:42, 60; 1:03;

[0294] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20; 1:03;

[0295] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102; 9:03;

[0296] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122; 9:03;

[0297] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90; 2:02;

[0298] Data 20: Carrier 2, out, 192.163.33.4, 2:02:45, 80; 2:02;

[0299] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59; 3:02;

[0300] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39; 3:02;

[0301] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70; 7:02;

[0302] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80; 7:02;

[0303] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85; 7:02;

[0304] Data 6: Operator 2, out, 192.163.16.4, 7:22:40, 30; 7:22.

[0305] Then, the database system performs a three-level sorting on the twenty pieces of data based on the label value of the first label "Address" included in each piece of data and the above sorting results. When performing the three-level sorting, the relative order between class 11, class 12, class 21, class 22, class 23 and class 24 remains unchanged. And when performing the three-level sorting, the ten pieces of data in class 11 are sorted in three levels based on the label value of "Address" included in the ten pieces of data; the two pieces of data in class 12 are sorted in three levels based on the label value of "Address" included in the two pieces of data; the two pieces of data in class 21 are sorted in three levels based on the label value of "Address" included in the two pieces of data; the two pieces of data in class 22 are sorted in three levels based on the label value of "Address" included in the two pieces of data; and the three pieces of data in class 23 are sorted in three levels based on the label value of "Address" included in the three pieces of data. After the three-level sorting, the storage order of the twenty pieces of data is as follows.

[0306] Data 1: Carrier 1, out, 192.163.13.1, 1:03:53, 60; 1:03;

[0307] Data 4: Carrier 1, out, 192.163.13.1, 1:03:56, 69; 1:03;

[0308] Data 15: Carrier 1, out, 192.163.13.1, 1:03:43, 20; 1:03;

[0309] Data 17: Operator 1, in, 192.163.13.1, 1:03:42, 60; 1:03;

[0310] Data 3: Carrier 1, out, 192.163.13.2, 1:03:55, 59; 1:03;

[0311] Data 9: Carrier 1, out, 192.163.13.2, 1:03:50, 160; 1:03;

[0312] Data 2: Carrier 1, out, 192.163.13.5, 1:03:52, 40; 1:03;

[0313] Data 16: Carrier 1, out, 192.163.19.1, 1:03:43, 30; 1:03;

[0314] Data 10: Carrier 1, out, 192.163.23.2, 1:03:50, 60; 1:03;

[0315] Data 18: Operator 1, in, 192.163.113.1, 1:03:42, 20; 1:03;

[0316] Data 8: Carrier 1, in, 192.163.12.2, 9:03:49, 122; 9:03;

[0317] Data 7: Carrier 1, in, 192.163.13.2, 9:03:49, 102; 9:03;

[0318] Data 19: Carrier 2, out, 192.163.13.4, 2:02:45, 90; 2:02;

[0319] Data 20: Carrier 2, out, 192.163.33.4, 2:02:45, 80; 2:02;

[0320] Data 12: Carrier 2, out, 192.163.11.4, 3:02:52, 39; 3:02;

[0321] Data 11: Carrier 2, out, 192.163.13.4, 3:02:52, 59; 3:02;

[0322] Data 5: Carrier 2, out, 192.163.13.4, 7:02:40, 70; 7:02;

[0323] Data 13: Carrier 2, out, 192.163.13.5, 7:02:29, 80; 7:02;

[0324] Data 14: Carrier 2, out, 192.163.33.5, 7:02:39, 85; 7:02;

[0325] Data 6: Operator 2, out, 192.163.16.4, 7:22:40, 30; 7:22.

[0326] Step 207: The database system stores the M pieces of data in the database based on the storage order of the M pieces of data.

[0327] For example, based on the storage order of the twenty pieces of data, the twenty pieces of data are saved in a database as shown in Table 2 below.

[0328] Table 2

[0329] Serial number data Operator Direction Address Time Bps 1 Data 1 Operator 1 out 192.163.13.1 1:03:53 60 2 Data 4 Operator 1 out 192.163.13.1 1:03:56 69 3 Data 15 Operator 1 out 192.163.13.1 1:03:43 20 4 Data 17 Operator 1 in 192.163.13.1 1:03:42 60 5 Data 3 Operator 1 out 192.163.13.2 1:03:55 59 6 Data 9 Operator 1 out 192.163.13.2 1:03:50 160 7 Data 2 Operator 1 out 192.163.13.5 1:03:52 40 8 Data 16 Operator 1 out 192.163.19.1 1:03:43 30 9 Data 10 Operator 1 out 192.163.23.2 1:03:50 60 10 Data 18 Operator 1 in 192.163.113.1 1:03:42 20 11 Data 8 Operator 1 in 192.163.12.2 9:03:49 122 12 Data 7 Operator 1 in 192.163.13.2 9:03:49 102 13 Data 19 Operator 2 out 192.163.13.4 2:02:45 90 14 Data 20 Operator 2 out 192.163.33.4 2:02:45 80 15 Data 12 Operator 2 out 192.163.11.4 3:02:52 39 16 Data 11 Operator 2 out 192.163.13.4 3:02:52 59 17 Data 5 Operator 2 out 192.163.13.4 7:02:40 70 18 Data 13 Operator 2 out 192.163.13.5 7:02:29 80 19 Data 14 Operator 2 out 192.163.33.5 7:02:39 85 20 Data 6 Operator 2 out 192.163.16.4 7:22:40 30

[0330] Compared with database 1, database 2 tries to group together data of addresses belonging to the same network segment, and the addresses in the "address" column in database 2 are more ordered than those in the "address" column in database 1.

[0331] Step 208: The database system constructs index information of the M data, the index information including identification information of the tag value set to which the tag value of the first tag in at least one selected data belongs, wherein two adjacent selected data are separated by one or more data in storage order.

[0332] Optionally, the selected at least one piece of data may include the first piece of data and the last piece of data among the M pieces of data.

[0333] Assume that the number of the one or more interval data is x, and x is an integer greater than or equal to 1. In step 208, the database system selects at least one data from the M data stored in the database, and there is an interval of x data between two adjacent selected data.

[0334] For example, the database system selects data 1, data 17, data 2, data 18, data 19, data 11 and data 6 from the twenty pieces of data stored in the database as shown in Table 2, and any two adjacent selected pieces of data are separated by two pieces of data, that is, x=2.

[0335] In some embodiments, the constructed index information may be identification information of a tag value set, and a correspondence between a tag value and a storage location. In step 208, the database system stores the identification information of the tag value set to which the tag value of the first tag included in each selected piece of data belongs, the tag values ​​of other tags included in each selected piece of data, and the storage location of each selected piece of data in the database in correspondence with the identification information of the tag value set, and the correspondence between the tag value and the storage location, wherein the other tags include tags other than the first tag in the sort key.

[0336] For example, the storage position of the selected data 1 in the database shown in Table 2 is sequence number 1, the storage position of the selected data 17 in the database shown in Table 2 is sequence number 4, the storage position of the selected data 2 in the database shown in Table 2 is sequence number 7, the storage position of the selected data 18 in the database shown in Table 2 is sequence number 10, the storage position of the selected data 19 in the database shown in Table 2 is sequence number 13, the storage position of the selected data 11 in the database shown in Table 2 is sequence number 16, and the storage position of the selected data 11 in the database shown in Table 2 is sequence number 19.

[0337] The database system saves operator 1 included in data 1, identification information "1:03" of the label value set, 192.163.13.1 included in data 1, and storage location "serial number 1" in the corresponding relationship between identification information of the label value set, label value and storage location as shown in Table 3 below.

[0338] The database system saves operator 1 included in data 17, identification information "1:03" of the label value set, 192.163.13.1 included in data 17, and storage location "serial number 4" in the corresponding relationship between identification information of the label value set, label value and storage location as shown in Table 3 below.

[0339] The database system saves operator 1 included in data 2, identification information "1:03" of the label value set, 192.163.13.5 included in data 2, and storage location "serial number 7" in the corresponding relationship between identification information of the label value set, label value and storage location as shown in Table 3 below.

[0340] The database system saves the operator 1 included in data 18, the identification information "1:03" of the label value set, 192.163.113.1 included in data 18, and the storage location "serial number 10" in the corresponding relationship between the identification information of the label value set, the label value and the storage location as shown in Table 3 below.

[0341] The database system saves the operator 2 included in data 19, the identification information "2:02" of the label value set, 192.163.13.4 included in data 19, and the storage location "serial number 13" in the corresponding relationship between the identification information of the label value set, the label value and the storage location as shown in Table 3 below.

[0342] The database system saves the operator 2 included in data 11, the identification information "3:02" of the label value set, 192.163.13.4 included in data 11, and the storage location "serial number 16" in the corresponding relationship between the identification information of the label value set, the label value and the storage location as shown in Table 3 below.

[0343] The database system saves operator 2 included in data 6, identification information "7:22" of the label value set, 192.163.16.4 included in data 6, and storage location "serial number 19" in the corresponding relationship between identification information of the label value set, label value and storage location as shown in Table 3 below.

[0344] Table 3

[0345] Operator Identification information of the Time tag value set Address Storage Location Operator 1 1:03 192.163.13.1 No. 1 Operator 1 1:03 192.163.13.1 No. 4 Operator 1 1:03 192.163.13.5 No.7 Operator 1 1:03 192.163.113.1 No. 10 Operator 2 2:02 192.163.13.4 No. 13 Operator 2 3:02 192.163.13.4 No. 16 Operator 2 7:02 192.163.33.5 No. 19

[0346] Among them, the first label is a label whose cardinality satisfies the high cardinality condition. The number of label values ​​of the first label after deduplication is very large. The order of the first label in the sorting key is before the second label. Therefore, a primary sort is performed based on the label value of the first label included in each data to be saved, and a secondary sort is performed based on the label value of the second label included in each data to be saved. This will reduce the local orderliness of a column of label values ​​corresponding to the second label in the database.

[0347] However, in the embodiment of the present application, the tag value set to which the tag value of the first tag included in each piece of data to be saved belongs is determined, and identification information of the tag value set to which the tag value of the first tag included in each piece of data belongs is obtained, and the number of tag value sets is much smaller than the number of tag values ​​of the first tag after deduplication. In this way, a primary sorting is performed based on the identification information of the tag value set to which the tag value of the first tag included in each piece of data to be saved belongs, and a secondary sorting is performed based on the tag value of the second tag included in each piece of data to be saved, which will improve the local orderliness of a column of tag values ​​corresponding to the second tag in the database.

[0348] In an embodiment of the present application, a database system receives a database creation request, which includes multiple tags and data structure information, and the data structure information is used to define the first tag and the second tag among the multiple tags as sorting keys. When the database system receives a storage request including M data, it determines the tag value set to which the tag value of the first tag included in each data belongs. The first tag may be a tag with an extremely large cardinality of tag values, and the cardinality of the tag value set of the first tag is much smaller than the cardinality of the tag value of the first tag. Based on the order of the sorting key, the identification information of the tag value set to which the tag value of the first tag included in each data belongs and the tag value of the second tag included in each data determine the storage order of each data. Storing each data in the order in which it is stored can increase the local orderliness of a column of tag values ​​corresponding to the second tag stored in the data table.

[0349] See also Figure 5 The present application embodiment provides a method 500 for querying data, wherein the method 500 is applied to Figure 1 The network architecture 100 shown includes the following processes.

[0350] Step 501: The terminal device sends a query request to a database system, where the query request includes a query condition related to a first tag.

[0351] The query condition includes a tag value of the first tag, or the query condition includes a tag value set of the first tag, and the tag value set of the first tag includes multiple tag values ​​of the first tag. That is, the query condition includes at least one tag value of the first tag.

[0352] In some embodiments, the query request may also include tag values ​​of other tags.

[0353] In step 501, the terminal device may display a third editing interface, and the user may enter a data query statement in the third editing interface. The terminal device obtains the data query statement from the third editing interface and uses the data query statement as a query request.

[0354] For example, suppose the user enters the following data query statement into the third editing interface displayed by the terminal device, where the query conditions are Operator = Operator 1, Time = 1:03:56, Address = 192.163.13.1. The terminal device obtains the data query statement, uses the data query statement as a query request, and sends the query request to the database system.

[0355] Data query statement:

[0356] Select from information

[0357] Where Operator = Operator 1, Time = 1:03:56, Address = 192.163.13.1.

[0358] Step 502: The database system receives the query request and determines the data to be scanned according to the query condition and the identification information of the tag value set to which the tag value of the first tag in the index information belongs.

[0359] In step 502, the data to be scanned may be determined through the following operations 5021-5022.

[0360] 5021: The database system receives the query request, and based on the identification information of the target tag value set in which at least one tag value of the first tag is located, obtains at least one storage location from the index information, wherein the at least one storage location includes the storage location of the first data and the storage location of the last data in the data to be scanned.

[0361] In step 5021, the database system determines a target tag value set based on at least one tag value of the first tag, the target tag value set including the at least one tag value; or determines multiple target tag value sets based on at least one tag value of the first tag, each target tag value set including some tag values ​​in the at least one tag value. The maximum identification information and the minimum identification information are selected from the identification information of the determined target tag value sets.

[0362] Select the first data from the index information, where the identification information of the tag value set to which the tag value of the first tag belongs is less than or equal to the minimum identification information, but there is no data before the first data in the index information where the identification information of the tag value set to which the tag value of the first tag belongs is greater than or equal to the minimum identification information. Then, obtain the first storage position corresponding to the first data from the index information, where the first storage position is the storage position of the first data in the data to be scanned.

[0363] Select the second data from the index information, the second data including identification information of the tag value set to which the tag value of the first tag belongs is greater than or equal to the maximum identification information, but there is no data in the index information after the first data, the identification information of the tag value set to which the tag value of the first tag belongs is less than or equal to the maximum identification information. Then, obtain the second storage position corresponding to the second data from the index information, the second storage position being the storage position of the last piece of data in the data to be scanned.

[0364] In some embodiments, if the storage request also includes tag values ​​of other tags, the other tags are tags other than the first tag in the sort key, the maximum identification information and the minimum identification information are selected from the identification information of the determined target tag value set, and the maximum tag value and the minimum tag value are selected from the tag values ​​of other tags.

[0365] Select the first data from the index information, the first data including the identification information of the tag value set to which the tag value of the first tag belongs is less than or equal to the minimum identification information, and the tag values ​​of other tags are less than or equal to the minimum tag value, but there is no data before the first data in the index information in which the identification information of the tag value set to which the tag value of the first tag belongs is greater than or equal to the minimum identification information, and the tag values ​​of other tags are greater than or equal to the minimum tag value. Then, obtain the first storage position corresponding to the first data from the index information, the first storage position being the storage position of the first data in the data to be scanned.

[0366] Select the second data from the index information, the second data including the identification information of the tag value set to which the tag value of the first tag belongs is greater than or equal to the maximum identification information, and the tag values ​​of other tags are greater than or equal to the maximum tag value, but there is no data in the index information after the first data in which the identification information of the tag value set to which the tag value of the first tag belongs is less than or equal to the maximum identification information, and the tag values ​​of other tags are less than or equal to the maximum tag value. Then, obtain the second storage position corresponding to the second data from the index information, the second storage position being the storage position of the last piece of data in the data to be scanned.

[0367] For example, the query conditions in the storage request are Operator = Operator 1, Time = 1:03:56, Address = 192.163.13.1. Based on the tag value "1:03:56" of the first tag "Time", the identification information of the target tag value set is determined to be "1:03". Based on "Operator = Operator 1", the identification information "1:03" of the target tag value set, and "Address = 192.163.13.1", the first data and the second data are selected from the index information shown in Table 3. The first data includes operator 1, the identification information "1:03" of the tag value set, the tag value "192.163.13.1" of Address, and the storage location of the first data is sequence number 1. The first data includes operator 1, the identification information "1:03" of the tag value set, the tag value "192.163.13.5" of Address, and the storage location of the second data is sequence number 7.

[0368] 5022: The database system determines the data to be scanned in the M pieces of data based on the at least one storage location.

[0369] The database system obtains data located between the first storage location and the second storage location from the database as data to be scanned.

[0370] For example, the first storage location is serial number 1, and the second storage location is serial number 7. The database system can obtain 7 data between serial number 1 and serial number 7 from the database shown in Table 2 as the data to be scanned. The 7 data are data 1, data 4, data 15, data 17, data 3, data 9 and data 2.

[0371] Step 503: The database system queries the data to be scanned to obtain query results that meet the query conditions.

[0372] In step 503, the query condition includes a tag value of the first tag, or the query condition includes a tag value set of the first tag, the tag value set of the first tag includes multiple tag values ​​of the first tag, and the database system selects target data from the data to be scanned based on the one or the multiple tag values, and the tag value of the first tag included in the target data is the one tag value or one of the multiple tag values.

[0373] In some embodiments, if the storage request also includes tag values ​​of other tags, target data is selected from the data to be scanned based on the one tag value or the multiple tag values, and the tag values ​​of the other tags, the target data includes the tag values ​​of the other tags, and the tag value of the included first tag is one of the one tag value or the multiple tag values.

[0374] For example, the database system selects target data from data 1, data 4, data 15, data 17, data 3, data 9 and data 2 based on Operator = operator 1, Time = 1:03:56, Address = 192.163.13.1, and the selected target data is data 4. Among them, because the database is sorted and classified based on the identification information of the tag value set to which the tag value of Time in each data belongs, multiple data with the same identification information of the tag value set to which the tag value of Time belongs are classified into one category. For each data classified into one category, the tag value of Address included in each data is sorted based on the tag value, so as to improve the local orderliness of each tag value corresponding to Address in the database, thereby reducing the determined data range to be scanned to only 7 data, thereby improving the query efficiency.

[0375] Step 504: The database system sends a query response to the terminal device, where the query response includes the query result.

[0376] The query result is the target data obtained by the above selection. After receiving the query response, the terminal device displays the target data included in the query response.

[0377] In an embodiment of the present application, a database system receives a query request, the query request includes a query condition, and the query condition is used to indicate at least one tag value of a first tag. Based on the identification information of the target tag value set where the at least one tag value is located, at least one storage location is obtained from the index information. Based on the at least one storage location, the data to be scanned is determined from the database, thereby narrowing the range of the data to be scanned that needs to be queried, and then based on the at least one tag value, the target data including some or all of the tag values ​​in the at least one tag value is selected from the data to be scanned, thereby improving the efficiency of querying the target data.

[0378] See also Figure 6 The present application embodiment provides a device 600 for storing data, which can be deployed in Figure 1 The database system 102 in the network architecture 100 shown in FIG. 1 may alternatively be deployed in Figure 2 The database system 102 in the method 200 shown in FIG. 200 may alternatively be deployed in Figure 5 The database system 102 in the method 500 is shown. The apparatus 600 includes:

[0379] A receiving unit 601 is configured to receive a plurality of pieces of data, wherein each piece of data includes label values ​​of a plurality of labels, the plurality of labels including a first label and a second label designated as a sorting key, and the first label is prior to the second label in the sorting key;

[0380] The processing unit 602 is used to determine the tag value set to which the tag value of the first tag in each piece of data belongs;

[0381] The processing unit 602 is further configured to determine a storage order of the plurality of data pieces based on the order of the sorting keys, identification information of a tag value set to which the tag value of the first tag in each piece of data belongs, and a tag value of the second tag in each piece of data;

[0382] The processing unit 602 is further configured to store the multiple pieces of data based on the storage order of the multiple pieces of data.

[0383] Optionally, the detailed implementation process of receiving unit 601 receiving multiple data can be found in Figure 2 The relevant contents of step 204 of the method 200 are not described in detail here.

[0384] Optionally, the detailed implementation process of the processing unit 602 determining the tag value set to which the tag value of the first tag in each piece of data belongs can be found in Figure 2 The relevant contents of step 205 of the method 200 are not described in detail here.

[0385] Optionally, the detailed implementation process of the processing unit 602 determining the storage order of the plurality of data can be found in Figure 2 The relevant contents of step 206 of the method 200 are not described in detail here.

[0386] Optionally, the processing unit 602 stores the plurality of data in a storage order. The detailed implementation process of storing the plurality of data can be found in Figure 2 The relevant contents of step 207 of the method 200 are not described in detail here.

[0387] Optionally, the receiving unit 601 is further configured to:

[0388] Data structure information is received, where the data structure information is used to specify a first tag and a second tag as sorting keys, and the data structure information includes quantified identification information, where the quantified identification information is used to indicate the first tag.

[0389] Optionally, the detailed implementation process of receiving unit 601 receiving data structure information can be found in Figure 2 The relevant contents of step 204 of the method 200 are not described in detail here.

[0390] Optionally, the processing unit 602 is configured to:

[0391] Based on the granularity of the label value set and the label value of the first label in each piece of data, determine the label value set to which the label value of the first label in each piece of data belongs, wherein the quantitative identification information is also used to indicate the granularity of the label value set, or the granularity of the label value set is determined based on the characteristics of the label value of the first label in the stored data.

[0392] Optionally, the processing unit 602 determines the tag value set to which the tag value of the first tag in each piece of data belongs based on the granularity of the tag value set and the tag value of the first tag in each piece of data. For a detailed implementation process, see Figure 2 The relevant contents of step 205 of the method 200 are not described in detail here.

[0393] Optionally, the processing unit 602 is further configured to:

[0394] Determine the labels that need to be quantified. The labels that need to be quantified are labels whose cardinality of label values ​​among the multiple labels meets the high cardinality condition.

[0395] When the first label is a label that needs to be quantified, determine the label value set to which the label value of the first label in each piece of data belongs.

[0396] Optionally, the detailed implementation process of the processing unit 602 determining the label to be quantified can be found in Figure 2 The relevant contents of step 204 of the method 200 are not described in detail here.

[0397] Optionally, when the first label is a label that needs to be quantified, the processing unit 602 determines the label value set to which the label value of the first label in each piece of data belongs. For a detailed implementation process, see Figure 2 The relevant contents of step 205 of the method 200 are not described in detail here.

[0398] Optionally, the tag value set to which the tag value of the first tag in each piece of data belongs is a tag value range that includes the tag value of the first tag in each piece of data, and the granularity of the tag value set is the length of the tag value range.

[0399] Optionally, the processing unit 602 is further configured to:

[0400] Index information of the multiple pieces of data is constructed, the index information including identification information of a label value set to which a label value of a first label in at least one selected piece of data belongs, wherein two adjacent selected pieces of data are separated by one or more pieces of data in storage order.

[0401] Optionally, the detailed implementation process of the processing unit 602 constructing the index information of the multiple data can be found in Figure 2 The relevant contents of step 208 of the method 200 are not described in detail here.

[0402] Optionally, the receiving unit 601 is further configured to receive a query request, where the query request includes a query condition related to the first tag;

[0403] The processing unit 602 is further configured to determine the data to be scanned from the plurality of pieces of data according to the query condition and the identification information of the tag value set to which the tag value of the first tag in the index information belongs;

[0404] The processing unit 602 is further configured to query the data to be scanned to obtain query results that meet the query conditions.

[0405] Optionally, the detailed implementation process of receiving the query request by the receiving unit 601 can be found in Figure 5 The relevant contents of step 502 of the method 500 are not described in detail here.

[0406] Optionally, the detailed implementation process of the processing unit 602 determining the data to be scanned among the multiple data can be found in Figure 5 The relevant contents of step 502 of the method 500 are not described in detail here.

[0407] Optionally, the processing unit 602 queries the data to be scanned to obtain the query results that meet the query conditions. For details on the implementation process, see Figure 5 The relevant contents of step 503 of the method 500 are not described in detail here.

[0408] In an embodiment of the present application, for each piece of data, the tag value set to which the tag value of the first tag in the data belongs includes not only the tag value but also other tag values ​​of the first tag, so that the cardinality of the tag value set to which the tag value of the first tag in each piece of data belongs is much smaller than the cardinality of the tag value of the first tag in each piece of data. Since the order of the first tag in the sorting key is before the second tag, the processing unit determines the storage order of the multiple pieces of data based on the order of the sorting key, the identification information of the tag value set to which the tag value of the first tag in each piece of data belongs, and the tag value of the second tag in each piece of data. And after the processing unit stores the multiple pieces of data in the storage order of the multiple pieces of data, the local orderliness of the tag values ​​of the second tags in the multiple pieces of data can be improved, that is, the orderliness of the data stored in the database can be improved.

[0409] See also Figure 7 , the present application embodiment provides a computing device 700. For example, the computing device 700 may be Figure 1 The network architecture 10 shown includes a device in the database system 102, or the computing device 700 can be Figure 2 The method 200 or Figure 5Devices in the database system in the method 500 shown, etc.

[0410] like Figure 7 As shown, computing device 700 includes: bus 702, processor 704, memory 706 and communication interface 708. Processor 704, memory 706 and communication interface 708 communicate through bus 702. Computing device 700 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 700.

[0411] The bus 702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus 702 may include a path for transmitting information between various components of the computing device 700 (eg, the processor 704, the memory 706, and the communication interface 708).

[0412] The processor 704 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0413] The memory 706 may include a volatile memory, such as a random access memory (RAM). The memory 706 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0414] See also Figure 7 , the memory 706 stores executable program codes, and the processor 704 executes the executable program codes to respectively implement Figure 6The functions of the receiving unit 601 and the processing unit 602 in the device 600 shown in the figure are implemented to realize the method provided by any of the above embodiments. That is, the memory 706 stores instructions for executing the method provided by any of the above embodiments. Or,

[0415] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or communication networks.

[0416] The embodiment of the present application also provides a cluster for storing data. The cluster for storing data includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0417] like Figure 8 As shown, the data storage cluster includes at least one computing device 700. The memory 706 in one or more computing devices 700 in the data storage cluster may store the same instructions for executing the method provided in any of the above embodiments.

[0418] In some possible implementations, the memory 706 of one or more computing devices 700 in the data storage cluster may also store partial instructions for executing the above-mentioned method for storing data. In other words, a combination of one or more computing devices 700 may jointly execute instructions for executing the method provided in any of the above-mentioned embodiments.

[0419] In some possible implementations, one or more computing devices in the data storage cluster may be connected via a network, which may be a wide area network or a local area network. Figure 8 A possible implementation is shown. Figure 8 As shown, two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network via a communication interface in each computing device.

[0420] In this type of possible implementation, the memory 706 in the computing device 700A stores the execution Figure 6 Instructions for the functions of the receiving unit 601 in the embodiment shown. Meanwhile, the memory 706 in the computing device 700B stores instructions for executing the following Figure 6 Instructions for the functionality of processing unit 602 in the illustrated embodiment.

[0421] It should be understood that Fig. 9The functions of the computing device 700A shown in FIG. 7 may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B may also be completed by multiple computing devices 700.

[0422] The present application embodiment also provides another cluster for storing data. The connection relationship between the computing devices in the cluster for storing data can be similarly referred to as Fig. 9 The connection mode of the cluster storing data is different in that the memory 706 in one or more computing devices 700 in the cluster storing data may store the same instructions for executing the method provided in any of the above embodiments.

[0423] In some possible implementations, the memory 706 of one or more computing devices 700 in the data storage cluster may also store partial instructions for executing the method provided in any of the above embodiments. In other words, the combination of one or more computing devices 700 may jointly execute instructions for executing the method provided in any of the above embodiments.

[0424] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method provided in any of the above embodiments.

[0425] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the method provided in any of the above embodiments.

[0426] A person skilled in the art will appreciate that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.

[0427] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for storing data, characterized in that: The method comprises: Receiving a plurality of pieces of data, wherein each piece of data includes a label value of a plurality of labels, the plurality of labels including a first label and a second label designated as a sort key, the first label being before the second label in the sort key; Determine the tag value set to which the tag value of the first tag in each piece of data belongs; Determine the storage order of the plurality of data based on the order of the sorting keys, the identification information of the tag value set to which the tag value of the first tag in each piece of data belongs, and the tag value of the second tag in each piece of data; The plurality of pieces of data are stored based on a storage order of the plurality of pieces of data.

2. The method according to claim 1, characterized in that Before receiving the plurality of pieces of data, the method further includes: Data structure information is received, where the data structure information is used to specify the first label and the second label as sorting keys, and the data structure information includes quantization identification information, where the quantization identification information is used to indicate the first label.

3. The method according to claim 2, characterized in that The determining the tag value set to which the tag value of the first tag in each piece of data belongs includes: Based on the granularity of the label value set and the label value of the first label in each piece of data, determine the label value set to which the label value of the first label in each piece of data belongs, wherein the quantitative identification information is also used to indicate the granularity of the label value set, or the granularity of the label value set is determined based on the characteristics of the label value of the first label in the stored data.

4. The method according to claim 1, characterized in that: The method further comprises: Determining a label that needs to be quantified, where the label that needs to be quantified is a label whose cardinality of label value satisfies a high cardinality condition among the multiple labels; When the first label is a label that needs to be quantified, determine the label value set to which the label value of the first label in each piece of data belongs.

5. The method according to claim 3, characterized in that: The tag value set to which the tag value of the first tag in each piece of data belongs is a tag value range that includes the tag value of the first tag in each piece of data, and the granularity of the tag value set is the length of the tag value range.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Constructing index information of the plurality of data, the index information comprising identification information of a tag value set to which a tag value of a first tag in at least one selected data belongs, wherein two adjacent selected data are separated by one or more data in storage order.

7. The method according to claim 6, characterized in that The method further comprises: receiving a query request, wherein the query request includes a query condition related to the first tag; Determine the data to be scanned among the multiple pieces of data according to the query condition and identification information of the tag value set to which the tag value of the first tag in the index information belongs; The data to be scanned is queried to obtain query results that meet the query conditions.

8. A device for storing data, characterized in that: The device comprises: A receiving unit, configured to receive a plurality of pieces of data, wherein each piece of data includes label values ​​of a plurality of labels, the plurality of labels including a first label and a second label designated as a sorting key, the first label being prior to the second label in the sorting key; a processing unit, configured to determine a tag value set to which a tag value of the first tag in each piece of data belongs; The processing unit is further configured to determine a storage order of the plurality of pieces of data based on an order of the sorting keys, identification information of a label value set to which a label value of a first label in each piece of data belongs, and a label value of a second label in each piece of data; The processing unit is further configured to store the multiple pieces of data based on a storage order of the multiple pieces of data.

9. The device according to claim 8, characterized in that The receiving unit is further used for: Data structure information is received, where the data structure information is used to specify the first label and the second label as sorting keys, and the data structure information includes quantization identification information, where the quantization identification information is used to indicate the first label.

10. The device according to claim 9, characterized in that The processing unit is used for: Based on the granularity of the label value set and the label value of the first label in each piece of data, determine the label value set to which the label value of the first label in each piece of data belongs, wherein the quantitative identification information is also used to indicate the granularity of the label value set, or the granularity of the label value set is determined based on the characteristics of the label value of the first label in the stored data.

11. The device according to claim 8, characterized in that The processing unit is further used for: Determining a label that needs to be quantified, where the label that needs to be quantified is a label whose cardinality of label value satisfies a high cardinality condition among the multiple labels; When the first label is a label that needs to be quantified, determine the label value set to which the label value of the first label in each piece of data belongs.

12. The device according to claim 11, characterized in that The tag value set to which the tag value of the first tag in each piece of data belongs is a tag value range that includes the tag value of the first tag in each piece of data, and the granularity of the tag value set is the length of the tag value range.

13. The device according to any one of claims 8 to 12, characterized in that The processing unit is further used for: Constructing index information of the plurality of data, the index information comprising identification information of a tag value set to which a tag value of a first tag in at least one selected data belongs, wherein two adjacent selected data are separated by one or more data in storage order.

14. The device according to claim 13, characterized in that The receiving unit is further configured to receive a query request, wherein the query request includes a query condition related to the first tag; The processing unit is further configured to determine the data to be scanned from among the plurality of pieces of data according to the query condition and identification information of the tag value set to which the tag value of the first tag in the index information belongs; The processing unit is further used to query the data to be scanned to obtain query results that meet the query conditions.

15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.

17. A computer program product comprising instructions, characterized in that When the instruction is executed by the computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Data storage method and device and storage medium

    EP4797113A1

  • Data storage method and device and storage medium

    WO2025107662A1