Index structure, data query method, file creation method and data processing method

By using file list information and two-dimensional bitmap to record the index column value status in the index structure, the problem of valuelist index taking up a large memory space and low query efficiency is solved, and more efficient data query and simplified index maintenance are achieved.

CN120353758APending Publication Date: 2025-07-22ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411003150.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When using Valuelist index to query data, metadata files in the prior art occupy a large amount of memory space, affecting memory performance and reducing query efficiency, and frequent index maintenance changes have a great impact on machine performance.

Method used

It provides an index structure, including file list information and column value distribution information, uses a two-dimensional bitmap to record the status of index column values in the data file, reduces cache space usage, and updates and maintains by modifying the file information and bit record status in the index structure.

Benefits of technology

It reduces the use of memory cache space, ensures the overall performance of the host, improves data query efficiency, and simplifies the update and maintenance process of index structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353758A_ABST
    Figure CN120353758A_ABST
Patent Text Reader

Abstract

The invention provides an index structure, a data query method, a file creation method and a data processing method, the index structure comprises file list information and column value distribution information, the file list information comprises file information of a plurality of data files, and the file information comprises file identifiers; the column value distribution information comprises an index column value domain and a two-dimensional bit map, and bits in the two-dimensional bit map, file identifiers of the data files and index column values in the index column value domain have corresponding relations; the bit is used for recording the state of the index column value corresponding to the bit in the data file indicated by the file identifier corresponding to the bit. Therefore, when data query is carried out, whether a certain index column value exists in the data file or not can be determined by loading the index structure into the memory and querying the two-dimensional bit map in the index structure, the occupation of the cache space of the memory can be reduced, the overall performance of a host can be ensured, and the query efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of indexing technology, and in particular, to an index structure, a data query method, a file creation method, and a data processing method. Background Art

[0002] Data skipping is a technology for performing Structured Query Language (SQL) analysis on structured data. Data skipping stores column summary statistics in each object (or file, etc.), such as Valuelist indexes. If the column index value range does not contain the data specified in the query, the query engine will skip the file during scanning, thereby reducing the amount of data scanned in the query, accelerating the query speed, and reducing resource costs.

[0003] In the related art, when using a Valuelist index for data query, the metadata file storing the Valuelist index is usually fully loaded into memory, and then data query is performed according to the data skipping technology. However, when the metadata file is large in size, fully loading the metadata file into memory will occupy a large amount of memory space, which not only affects the memory performance but also reduces the query efficiency. Summary of the Invention

[0004] This application provides an index structure, a data query method, a file creation method, and a data processing method, which are used to solve the problem that in the related art, when using a Valuelist index for data query, it will occupy a large amount of memory space, which not only affects the memory performance but also reduces the query efficiency.

[0005] To solve the above technical problems, this application is implemented as follows:

[0006] In a first aspect, an index structure is provided, including file list information and column value distribution information, where:

[0007] The file list information includes file information of multiple data files, and the file information includes a file identifier;

[0008] The column value distribution information includes an index column value range and a two-dimensional bitmap, and there is a corresponding relationship between the bit positions in the two-dimensional bitmap, the file identifiers of the data files, and the index column values in the index column value range. For any bit position, the bit position is used to record the status of the index column value corresponding to the bit position in the data file indicated by the file identifier corresponding to the bit position.

[0009] In a second aspect, a data query method based on the index structure is provided, where the index structure includes the index structure described in the first aspect, and the method includes:

[0010] Receive a query request;

[0011] Obtain an index file corresponding to multiple data files, where the index structure is included in the index file;

[0012] Determine valid files from the multiple data files according to the query request and the index structure;

[0013] Query for target data corresponding to the query request in the valid files.

[0014] In a third aspect, a method for creating a data file based on an index structure is provided. The index structure includes the index structure as described in the first aspect. The method includes:

[0015] Receive a file creation request for a first data file;

[0016] Determine the file information of the first data file, and determine whether there is an index column value in the index column value range in the index structure in the first data file. The file information of the first data file includes the file identifier of the first data file;

[0017] Store the file information of the first data file in the file list information of the index structure, and record the status of the index column value in the index column value range in the first data file in the two-dimensional bitmap of the index structure;

[0018] Create the first data file according to the file creation request.

[0019] In a fourth aspect, a data processing method based on an index structure is provided. The index structure includes the index structure according to any one of claims 1 to 9. The method includes:

[0020] Receive a data processing request for a second data file;

[0021] Perform data processing on the second data file according to the data processing request to obtain a third data file;

[0022] Determine the file information of the third data file, and determine whether there is an index column value in the index column value range in the index structure in the third data file. The file information of the third data file includes the file identifier of the third data file;

[0023] Store the file information of the third data file in the file list information of the index structure, and record the status of the index column value in the index column value range in the third data file in the two-dimensional bitmap of the index structure.

[0024] In a fifth aspect, there is provided an electronic device, comprising:

[0025] a processor;

[0026] a memory for storing executable instructions of the processor;

[0027] wherein the processor is configured to execute the instructions to implement the method as described in the second aspect or the third aspect or the fourth aspect.

[0028] In a sixth aspect, there is provided a computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method as described in the second aspect or the third aspect or the fourth aspect.

[0029] In a seventh aspect, there is provided a computer program product, the computer program product comprising a non-transitory computer-readable storage medium storing a computer program, the computer program being operable to cause a computer to execute some or all of the steps of the method as described in the second aspect, or execute some or all of the steps of the method as described in the third aspect, execute some or all of the steps of the method as described in the fourth aspect.

[0030] In the embodiments of the present application, since the content stored in the index structure is the file information of multiple data files, the index column value range, and the two-dimensional bitmap, the cache space occupied by the index structure is relatively small. Since the bit positions in the two-dimensional bitmap of the index structure can be used to record the status of the index column value in the data file, when performing data query, by loading the index structure into the memory and querying the two-dimensional bitmap in the index structure, it is possible to determine whether a certain index column value exists in the data file. Compared with the related art of loading all the metadata files storing the Valuelist index into the memory for data query, it can not only reduce the occupation of the memory cache space, ensure the overall performance of the host, but also improve the query efficiency.

[0031] Based on the index structure provided in the embodiments of the present application, when creating a data file or performing an update process on the data in the data file, by modifying the file information stored in the index structure and the status recorded by the bit positions in the two-dimensional bitmap, the update and maintenance of the index structure can be realized, without frequently loading and updating the metadata file, with less impact on the performance of the machine and being easy to realize the update and maintenance of the index structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the implementation examples or the prior art descriptions. Obviously, the drawings in the following descriptions are only some implementation examples recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0033] Figure 1 is a schematic diagram of data query according to the Valuelist index in the related art;

[0034] Figure 2 is a schematic diagram of a schematic system architecture provided by an embodiment of the present application;

[0035] Figure 3 is a schematic diagram of the index structure of an embodiment of the present application;

[0036] Figure 4 is a schematic diagram of the index structure of another embodiment of the present application;

[0037] Figure 5 is a schematic flowchart of a data query method based on the index structure in an embodiment of the present application;

[0038] Figure 6 is a schematic diagram of the expression tree obtained by parsing the query conditions in an embodiment of the present application;

[0039] Figure 7 is a schematic diagram of operating on multiple rows of bit positions to obtain the target bit string in an embodiment of the present application;

[0040] Figure 8 is a schematic diagram of determining valid files according to the target bit string in an embodiment of the present application;

[0041] Figure 9 is a schematic flowchart of a data file creation method based on the index structure in an embodiment of the present application;

[0042] Figure 10 is a schematic flowchart of a data file creation method based on the index structure in another embodiment of the present application;

[0043] Figure 11 is a schematic flowchart of a data processing method based on the index structure in an embodiment of the present application;

[0044] Figure 12 is a schematic flowchart of a data modification method based on the index structure in an embodiment of the present application;

[0045] Figure 13It is a schematic flowchart of a method for deleting data files based on an index structure according to an embodiment of the present application;

[0046] Figure 14 It is a schematic flowchart of a method for deleting data files based on an index structure according to another embodiment of the present application;

[0047] Figure 15 It is a schematic structural diagram of an electronic device according to an embodiment of the present application;

[0048] Figure 16 It is a schematic structural diagram of a data query device based on an index structure according to an embodiment of the present application;

[0049] Figure 17 It is a schematic structural diagram of an electronic device according to an embodiment of the present application;

[0050] Figure 18 It is a schematic structural diagram of a data file creation device based on an index structure according to an embodiment of the present application;

[0051] Figure 19 It is a schematic structural diagram of an electronic device according to an embodiment of the present application;

[0052] Figure 20 It is a schematic structural diagram of a data processing device based on an index structure according to an embodiment of the present application. Detailed implementation manners

[0053] Data statistical information plays a key role in improving query efficiency. The specific statistical data used determines the effectiveness and cost of the query. An ideal statistic is small, easy to maintain, and supports the evaluation of rich query predicates. The Valuelist index is based on the statistical information commonly used for low-cardinality columns. In the related art, when performing data query, data skipping can be performed according to the Valuelist index to improve query efficiency. As Figure 1 shown, when the client queries the data related to the two cities of "Beijing" and "Nanjing" from Table A, the following steps may be included:

[0054] S1: Load all metadata files, parse out the Valuelist index information of the field "city" saved in the metadata files, and then load it into the memory to generate a Map based on the index field;

[0055] S2: Traverse the Valuelist index information of the Map one by one, and determine whether the conditional values in the predicate (i.e., "Beijing" and "Nanjing") exist in each record;

[0056] S3: Screen out the records that meet the predicate conditions, and intercept the corresponding FileName information (for example Figure 1Files File0000, File0001, and File0005 that record "Beijing" or "Nanjing");

[0057] S4: Cache the FileName and merge it with other predicate filtering conditions (perform OR / AND operations).

[0058] In the above steps, before querying data according to the Valuelist index, all metadata files need to be loaded into memory. In some scenarios (such as in the scenario of a large amount of data in a lakehouse), when the volume of metadata files is very large, sufficient memory is required to cache the metadata files, which will reduce the overall performance of the host, and the time to list the metadata files will also be relatively long. In addition, when querying data according to the Valuelist index, since the metadata files are loaded in full volume, therefore, it takes a certain amount of calculation to determine whether the Value in the predicate condition exists or to optimize the memory structure (for accelerating the judgment, perform data reorganization) and then make a judgment, which will reduce the data query efficiency. In addition, index change maintenance also requires rewriting the metadata files, and frequent changes have a greater impact on the machine performance.

[0059] It can be seen that in the related art, how to reduce the memory occupation during the use of the Valuelist index, avoid the long time of the ListFile file in scenarios such as lakehouse integration, optimize the index maintenance ability, reduce the impact on the existing system, and improve the query performance are urgent problems to be solved.

[0060] The embodiment of the present application provides an index structure, a data query method, a file creation method, and a data processing method. The index structure is used to store the Valuelist index, specifically including file list information and column value distribution information. The file list information includes the file information of multiple data files, and the file information includes file identifiers; the column value distribution information includes an index column value range and a two-dimensional bitmap. The bits in the two-dimensional bitmap have a corresponding relationship with the file identifiers of the data files and the index column values in the index column value range. For any bit, this bit is used to record the status of the index column value corresponding to this bit in the data file indicated by the file identifier corresponding to this bit.

[0061] Since the content stored in the index structure is the file information of multiple data files, the value range of the index column, and the two-dimensional bitmap, the cache space occupied by the index structure is relatively small. Since the bits in the two-dimensional bitmap of the index structure can be used to record the status of the index column value in the data file, when performing data query, by loading the index structure into the memory and querying the two-dimensional bitmap in the index structure, it is possible to determine whether a certain index column value exists in the data file. Compared with the related art in which the metadata file storing the Valuelist index is fully loaded into the memory for data query, it can not only reduce the occupation of the memory cache space, ensure the overall performance of the host, but also improve the query efficiency.

[0062] Based on the index structure provided by the embodiments of the present application, when creating a data file or performing an update process on the data in the data file, only by modifying the file information stored in the index structure and the status recorded by the bits in the two-dimensional bitmap, the update and maintenance of the index structure can be realized, without frequently loading and updating the metadata file, which has less impact on the performance of the machine and is easy to realize the update and maintenance of the index structure.

[0063] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the present application with reference to the accompanying drawings in one or more embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0064] The terms "first", "second", etc. in the present application and the claims are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the present application and the claims means at least one of the connected objects, and the character " / " generally means that the related objects before and after are in an "or" relationship.

[0065] Figure 2 It is a schematic system architecture diagram provided by the embodiments of the present application.

[0066] Figure 2The application environment of the system architecture shown can be logically divided into a computing layer 21 and a storage layer 22. The computing layer 21 manages and accesses data based on the storage layer 22, and its TableFormat is an optional deployment. The storage layer 12 provides data storage capabilities, and the underlying layer uses formats such as Parquet for storage. The computing layer 21 and the storage layer 22 can communicate through a network connection. Among them, TableFormat is a table format that enables the data lake to have the ability of ACID (atomicity, consistency, isolation, and durability, that is, atomicity (or indivisibility), consistency, isolation (or independence), and durability) transactions and enhances the data management ability. Open-source systems such as Iceberg, Hudi, and DeltaLake are all open-source products of TableFormat.

[0067] In actual deployment, according to the application scenario, the deployment of the computing engine can include but is not limited to the following scenarios (1) to (5):

[0068] (1) A separate big data operating environment (not based on the TableFormat system);

[0069] (2) A separate big data operating environment (based on the TableFormat system);

[0070] (3) A separate data warehouse operating environment (based on the TableFormat system);

[0071] (4) A separate data warehouse operating environment (not based on the TableFormat system);

[0072] (5) A lakehouse integrated operating environment (co-deploying a big data and data warehouse engine based on TableFormat).

[0073] The operating environment described in any one of the above (1) to (5) can include but is not limited to a physical machine environment and a cloud environment, and no specific limitation is made here.

[0074] The technical solution provided by the embodiments of the present application can be applied to data storage devices such as databases and data warehouses, as well as scenarios such as big data processing.

[0075] The following will describe in detail the technical solution provided by the embodiments of the present application in conjunction with the accompanying drawings.

[0076] Figure 3 It is a schematic diagram of the index structure of an embodiment of the present application.

[0077] Figure 3The index structure shown includes file list information and column value distribution information. The file list information includes the file information of multiple data files, and the multiple data files may be data files under multiple table partitions. The file information of the data files may include the file identifier of the data files. The column value distribution information includes an index column value range and a two-dimensional bitmap. The index column value range includes multiple index column values. The two-dimensional bitmap includes multiple bits, and these bits have a corresponding relationship with the file identifier of the data file and the index column values in the index column value range. For any bit in the two-dimensional bitmap, this bit can be used to record the status of the index column value corresponding to this bit in the data file indicated by the file identifier corresponding to this bit. The status may be that the data file indicated by the file identifier corresponding to this bit exists (or includes) the index column value corresponding to this bit, or the data file indicated by the file identifier corresponding to this bit does not exist (or does not include) the index column value corresponding to this bit.

[0078] Based on the index structure provided by the embodiments of the present application, since the content stored in the index structure is the file information of multiple data files, the index column value range, and the two-dimensional bitmap, the cache space occupied by the index structure is relatively small. Since the bits in the two-dimensional bitmap of the index structure can be used to record the status of the index column value in the data file, when performing data query, by loading the index structure into the memory and querying the two-dimensional bitmap in the index structure, it is possible to determine whether a certain index column value exists in the data file. Compared with the related art in which the metadata file storing the Valuelist index is fully loaded into the memory for data query, it can not only reduce the occupation of the memory cache space and ensure the overall performance of the host, but also improve the query efficiency.

[0079] In some embodiments, the file identifier of the data file may include a file sequence number, and the file sequence number may be determined according to the generation order of the data files. For example, the file sequence number may start from 0, and for each newly added data file, the file sequence number is incremented by 1. Among them, for any data file, there is a one-to-one correspondence between the file sequence number of the data file and the bit order of the bit corresponding to the data file (this correspondence is equivalent to the correspondence between the file identifier of the data file and the bit). Optionally, in some embodiments, the file sequence number of the data file may specifically be equal to the bit order of the bit corresponding to the file identifier of the data file or equal to the sum of this bit order and a specified offset value. For example, if the bit order of the bit corresponding to the file identifier of the data file is 19, then the file sequence number of this data file may be 19 or 10019 (the sum of 19 and the specified offset value 10000, and 10000 may be the maximum value of the bit order).

[0080] By using the file serial number as the file identifier of the data file, and the file serial number can be equal to the bit order of the bit corresponding to the data file or the sum of the bit order and the specified offset value. Therefore, it is not only convenient to distinguish different data files, but also convenient to establish a one-to-one correspondence between the data file and the bit according to the file serial number. When performing data query, the corresponding file serial number can be determined according to the bit order of the bit, and then the corresponding data file can be determined. Then, according to the record of the bit, it can be determined whether a certain index column value exists in the data file. Thus, it can be quickly determined whether the data file is the target file to be queried, improving the query efficiency.

[0081] In some embodiments, the file list information of the index structure may include a multi-dimensional file list, and the multi-dimensional file list can be used to store the file information of multiple data files. That is to say, the file information of multiple data files can be stored in the form of a multi-dimensional file list. Among them, the dimension of the multi-dimensional file list can be set flexibly, such as column names, etc., which are not specifically limited here.

[0082] Taking the file information including the file serial number as an example, when storing the file serial numbers of data files in the form of a multi-dimensional file list, the storage method can be referred to Table 1. When the file list shown in Table 1 stores the file serial numbers of data files 0 to 6 respectively, with the column name "country" of the data file as the dimension, the file serial numbers are stored in multiple dimensions according to different countries. Among them, the data files with file serial numbers 0, 2, 4, and 6 include the column name "China", and the data files with file serial numbers 1, 3, and 5 include the column name "America". Therefore, when storing the file serial numbers, the column name "China" can be stored corresponding to the file serial numbers 0, 2, 4, and 6 (stored in the same row in the file list), and the column name "America" can be stored corresponding to the file serial numbers 1, 3, and 5.

[0083] Table 1

[0084]

[0085]

[0086] Since the file information of multiple data files can be stored in the form of a multi-dimensional file list, flexible storage of file information can be achieved. In addition, when performing data query, the corresponding data file can also be queried according to the dimension of the file list, thereby improving the query efficiency.

[0087] In some embodiments, the file information of the data file may further include at least one of the file path, file name, and file status. Among them, the file status may include normal or invalid. For a certain data file, if the file status of the data file is normal, then when performing data query, the data file can be used as the data file to be queried. If the file status of the data file is invalid, then when performing data query, the data file will not be used as the data file to be queried (it should be noted that the reason why the data file has normal and invalid statuses is that in actual applications, the data file is dynamically updated and not kept unchanged. After updating a certain data file, the old data file becomes invalid, but the old data file will not be immediately deleted. At this time, the file status is needed to distinguish the old data file from the updated data file). Optionally, in some embodiments, the file status may be represented by 0 or 1. For example, 0 represents that the status of the data file is normal, 1 represents that the status of the data file is invalid, or 1 represents that the status of the data file is normal, and 0 represents that the status of the data file is invalid.

[0088] Optionally, when the file information of the data file includes the file path, file name, and file status, the multi-dimensional file list shown in Table 1 may be as shown in Table 2 (the file status 0 in Table 2 represents valid).

[0089] Table 2

[0090]

[0091] By recording the file path and / or file name of the data file in the file list information of the index structure, it is convenient to quickly locate the target data file according to the file path and / or file name of the valid file after performing data query and querying the valid file, thereby improving the data query efficiency. By recording the file status of the data file in the file list information of the index structure, it is convenient to distinguish normal files and invalid files, and when performing data query, it is possible to avoid the problem of query failure caused by using invalid files as the files to be queried.

[0092] In some embodiments, in the case where multiple data files include data files under multiple table partitions, the file list information may further include a first data dictionary in addition to the file information of the data file. The first data dictionary can be used to store the value range of the table partition fields. For example, the first data dictionary can store table partitions A and B, and different table partitions can be determined according to the first data dictionary. Optionally, the first data dictionary may be a forward-sorted data dictionary or a reverse-sorted data dictionary, which is not specifically limited here.

[0093] It should be noted that in other possible implementation manners, the value range of the table partition field may also be stored as other data structures other than the data dictionary, and no specific limitation is made here.

[0094] In some implementation manners, the column value distribution information may include a second data dictionary, and the second data dictionary may be used to store the value range of the index column. That is to say, the value range of the index column in the column value distribution information may be stored in the form of a data dictionary. In this way, when performing data query, it is possible to quickly determine whether the value range of the index column includes the index column value to be queried according to the second data dictionary in the column value distribution information, thereby improving the query efficiency. Optionally, the second data dictionary may be a forward-sorted data dictionary or a reverse-sorted data dictionary, and no specific limitation is made here.

[0095] There is a corresponding relationship between the index column values in the value range of the index column, the bit positions in the two-dimensional bit map, and the file identifiers of the data files. In some implementation manners, the corresponding relationship may be that in the two-dimensional bit map, each row of bit positions corresponds to an index column value, and each column of bit positions corresponds to a file identifier of a data file (the number of rows and columns of the two-dimensional bit map may be determined according to the actual scenario, and no specific limitation is made here). That is to say, each bit position in the two-dimensional bit map corresponds to an index column value and a file identifier. The index column values corresponding to the bit positions in the same row are the same, and the file identifiers corresponding to the bit positions in the same column are the same. For any bit position in the two-dimensional bit map, the state recorded by this bit position may include that the data file indicated by the file identifier corresponding to this bit position exists (or includes) the index column value corresponding to this bit position, or the data file indicated by the file identifier corresponding to this bit position does not exist (or does not include) the index column value corresponding to this bit position.

[0096] For any bit position, the bit position value of 0 may indicate that the data file indicated by the file identifier corresponding to the bit position does not exist (or does not include) the index column value corresponding to the bit position, and the bit position value of 1 indicates that the data file indicated by the file identifier corresponding to the bit position exists (or includes) the index column value corresponding to the bit position. Alternatively, it may also be that the bit position value of 1 indicates that the data file indicated by the file identifier corresponding to the bit position does not exist (or does not include) the index column value corresponding to the bit position, and the bit position value of 0 indicates that the data file indicated by the file identifier corresponding to the bit position exists (or includes) the index column value corresponding to the bit position. In the embodiments of the present application, an example is given in which the bit position value of 0 indicates that the data file indicated by the file identifier corresponding to the bit position does not exist (or does not include) the index column value corresponding to the bit position, and the bit position value of 1 indicates that the data file indicated by the file identifier corresponding to the bit position exists (or includes) the index column value corresponding to the bit position.

[0097] In some embodiments, the two-dimensional bit map in the column value distribution information can be stored in the form of a hexadecimal byte string. In this way, when the index structure is loaded into memory for data query, the memory cache space occupied by the two-dimensional bit map can be reduced, and thus the memory cache space occupied by the index structure can be reduced. Optionally, in other possible embodiments, the two-dimensional bit map can also be stored in other forms. For example, the two-dimensional bit map can be stored in the form of a binary string. Here, no further examples will be given one by one.

[0098] The index column values in the embodiments of the present application can support different data types. In some embodiments, the data types of the index column values can include but are not limited to integer type, floating-point type, character type, and date type. It should be noted that since the Valuelist index is based on the common statistical information of low-cardinality columns, in some more specific embodiments, the data types of the index column values can be applicable to the data types of low-cardinality fields, such as the integer type, floating-point type, character type, and date type, etc., which conform to the data types of low-cardinality fields.

[0099] To facilitate understanding of the index structure provided by the embodiments of the present application, the following will use Figure 4 a more specific embodiment of the index structure shown as an example for description.

[0100] Figure 4 The index structure shown includes file list information and column value distribution information. The file list information includes a partition dictionary and a multi-dimensional file list. The partition dictionary includes table partition A and table partition B. Table partition A includes 6 files, and table partition B includes 10 files. The multi-dimensional file list includes the file identifier of the data file (i.e., Figure 4 the fileid shown), the file name, and the file status ( Figure 4 the specific file name and file status are not shown). Among them, the file identifier is equal to the bit order of the bit corresponding to the file identifier.

[0101] The column value distribution information includes a column value dictionary and a two-dimensional bit map. The column value dictionary includes index column value A, index column value B, and index column value C ( Figure 4 abbreviated as column value A, column value B, and column value C). The two-dimensional bit map includes 3×16 bits. Among them, the first row of bits corresponds to index column value A, the second row of bits corresponds to index column value B, and the third row of bits corresponds to index column value C. The bit order of the first column of bits is 0, corresponding to fileid0, the bit order of the second column of bits is 1, corresponding to fileid1,..., and the bit order of the last column of bits is 15, corresponding to fileid15. In actual storage, the two-dimensional bit map can be stored in the form of a hexadecimal byte string. For example, a row of bits corresponding to index column value A can be stored as 0x96BA.

[0102] For any bit, each bit has a corresponding relationship with an index column value and a file identifier. Each bit is used to record the status of the index column value corresponding to the bit in the data file indicated by the file identifier corresponding to the bit. Taking the first bit in the first row and the first bit in the second row as an example, if the value of the first bit in the first row is 1 and the corresponding index column value is A and the fileid is 0, it can indicate that the index column value A exists (or is included) in the data file with fileid 0. If the value of the first bit in the second row is 0 and the corresponding index column value is B and the fileid is 0, it can indicate that the index column value B does not exist (or is not included) in the data file with fileid 0.

[0103] Figure 4 For the shown index structure, for 16 data files with fileids from 0 to 15, it is possible to record whether the index column value A, index column value B, and index column value C exist (or are included) in these 16 data files through a two-dimensional bit map. In this way, when querying data related to a certain index column value (such as index column value A), it is possible to determine which data files include the index column value to be queried according to the two-dimensional bit map in the index structure, thereby determining the target data file. Then, according to the file list information in the index structure, locate the valid file and perform data query in the valid file. During the entire query process, it is not necessary to load all the metadata files storing the Valuelist index into the memory, which can reduce the occupation of the memory cache space, ensure the overall performance of the host, and improve the query efficiency.

[0104] For the index structure provided by the embodiments of the present application, since the content stored in the index structure is the file information of multiple data files, the index column value range, and the two-dimensional bit map, the cache space occupied by the index structure is relatively small. Since the bits in the two-dimensional bit map of the index structure can be used to record the status of the index column value in the data file, during data query, by loading the index structure into the memory and querying the two-dimensional bit map in the index structure, it is possible to determine whether a certain index column value exists in the data file. Compared with the related technology of loading all the metadata files storing the Valuelist index into the memory for data query, it can not only reduce the occupation of the memory cache space, ensure the overall performance of the host, but also improve the query efficiency.

[0105] Based on the index structure provided by the embodiments of the present application, the embodiments of the present application also provide a data query method, a data file creation method, and a data processing method based on the index structure. The following will use Figures 5 to 14 the shown embodiments as examples for illustration.

[0106] Figure 5FIG. 0 is a flowchart of a data query method based on an index structure according to an embodiment of the present application. The index structure includes the index structure provided by the embodiment of the present application. For details, reference can be made to Figure 3 and Figure 4 the embodiments shown, which will not be repeated here. Figure 5 The data query method shown in Figure 2 can be executed by the big data execution engine in the computing layer 21 shown in Figure 5 The data query method shown in

[0107] Step S502: Receive a query request.

[0108] When an external client or other system has a data query requirement, it can send a query request to the big data execution engine. At this time, the big data query engine can receive the query request sent by the external client or system. Among them, the query request may include an SQL statement, and the SQL query statement includes query conditions.

[0109] Step S504: Obtain index files corresponding to multiple data files. The index files include an index structure.

[0110] The multiple data files are the objects to be queried. That is, when performing data query, the target data required is queried from multiple data files. In this embodiment, for multiple data files, the index structure provided by the embodiment of the present application can be determined (or constructed) in advance according to the data in the data files, and the index structure is stored in the index file. When performing data query, the index files corresponding to multiple data files can be obtained. Among them, multiple data files may correspond to one index file, and the index file includes the index structure provided by the embodiment of the present application.

[0111] In some embodiments, the index structure provided by the embodiment of the present application can be determined in the following manner:

[0112] Determine the file information of multiple data files. The file information includes file identifiers;

[0113] According to the index column names in the index column configuration information, determine the index column values corresponding to the index column names in multiple data files to obtain an index column value range;

[0114] Generate file list information according to the file information of multiple data files, and generate column value distribution information according to the index column value range and whether each data file has the index column values in the index column value range;

[0115] Construct an index structure according to the file list information and the column value distribution information.

[0116] Multiple data files can be data files under multiple table partitions. The file information of each data file includes at least the file identifier of the data file. Optionally, it can also include at least one of the file name, file path, and file status of the data file. For the relevant explanations of the file information of the data file, reference can be made to Figure 3 the corresponding content in the illustrated embodiment, which will not be described in detail here.

[0117] When determining the file information of multiple data files, the basic information of the multiple data files can be obtained first, such as the generation time or the order of generation, the name of the data file, the storage location of the data file, whether the data file is a normal file or a failed file, the partition to which the data file belongs, etc. Then, the file information of the data file can be determined based on this basic information. For example, the file sequence numbers of the multiple data files can be determined according to the generation time or the order of generation of the multiple data files (the file sequence number can start from 0, and for each newly added data file, the file sequence number is incremented by 1), and the file sequence number can be used as the file identifier of the multiple data files. For another example, the status of the multiple data files (including normal and failed) can be determined according to whether the multiple data files are normal files or failed files, and the file path of the data file can be determined according to the storage location of the data file.

[0118] The index column configuration information can be pre-configured. The index column configuration information can include multiple index column names, and the multiple index column names can be obtained by counting the index column names in the multiple data files, or can also be set according to the actual situation, which is not specifically limited here. Optionally, the index column configuration information can be Valuelist_Index_Column={...,...} / / Valuelist index column configuration information, separated by Western commas.

[0119] For each index column name in the index column configuration information, the index column values corresponding to the index column name in the multiple data files can be determined, that is, find and count which index column values the index column name specifically corresponds to in the multiple data files, and the index column value range can be obtained based on these index column values. For example, if the index column name is "city", and the index column values corresponding to "city" in the multiple data files include "Beijing", "Nanjing", "Shanghai",... "Ningbo", then the index column value range {"Beijing", "Nanjing", "Shanghai",... "Ningbo"} can be obtained.

[0120] After obtaining the file information and index column value ranges of multiple data files, file list information and index column value information can be further generated. Specifically, file list information can be generated based on the file information of multiple data files. The file list information includes the file information of multiple data files, and the file information includes at least a file identifier. Optionally, it may further include at least one of a file name, a file path, and a file status. When generating column value distribution information, for each data file, it can be determined whether each index column value in the index column value range exists (or is included) in the data file. According to the determination result, the value of the bit position corresponding to the file identifier of the data file in a two-dimensional bitmap (the number of rows and columns of the two-dimensional bitmap can be pre-configured, and each bit position in the two-dimensional bitmap has a corresponding relationship with the index column value and the file identifier of the data file) is determined (for example, for a certain index column value, if the index column value exists in the data file, the value of the bit position corresponding to the index column value and the file identifier of the data file is 1, otherwise it is 0). Then, column value distribution information is generated based on the index column value range and the two-dimensional bitmap.

[0121] It should be noted that when generating the file list information and column value distribution information above, it can be divided into two methods: online generation and offline generation, which are not specifically limited here.

[0122] After generating the file list information and column value distribution information, the index structure provided by the embodiments of the present application can be generated based on the file list information and column value distribution information. After generating the index structure, the index structure can be stored in an index file. When performing data queries, the index file can be obtained, and then the index structure can be obtained. Among them, when storing the index structure, the storage methods include but are not limited to json, Parquet, etc., which are not specifically limited here.

[0123] Step S506: Determine valid files from multiple data files according to the query request and the index structure.

[0124] After obtaining the index files corresponding to multiple data files, invalid files can be filtered out from multiple data files according to the query request and the index structure in the index file to obtain valid files. The target data to be queried is stored in the valid files, and the number of valid files can be one or more.

[0125] In some embodiments, determining valid files from multiple data files according to the query request and the index structure may include the following steps:

[0126] Determine target index column values that meet the query conditions from the index column value range of the index structure according to the query conditions carried in the query request;

[0127] Determine a target bit string according to a target index column value and a two-dimensional bit map in an index structure. The target bit string includes multiple bits, and the multiple bits are used to indicate whether the target index column value is included in multiple data files;

[0128] Determine valid files from multiple data files according to the target bit string.

[0129] A query request usually carries query conditions. When determining valid files from multiple data files according to the query request and the index structure, the index column value(s) that meet the query conditions can be determined from the index column value range of the index structure according to the query conditions first, and then further queries can be performed according to the index column value. Among them, for the convenience of distinction, the index column value that meets the query conditions can be represented as the target index column value.

[0130] In some embodiments, determining the target index column value that meets the query conditions from the index column value range of the index structure according to the query conditions carried in the query request may include the following steps:

[0131] Parse the query conditions to obtain the predicate, predicate operator, and index column value of the predicate in the query conditions;

[0132] Determine whether the index column value range of the predicate is included in the index column value range of the index structure;

[0133] In the case where the index column value range of the predicate is included in the index column value range of the index structure, perform an operation on the index column value of the predicate according to the predicate operator, and determine the target index column value that meets the query conditions from the index column value range of the index structure according to the operation result.

[0134] The query conditions usually include a predicate, a predicate operator, and an index column value of the predicate. When determining the target index column value, the query conditions can be parsed first, and the predicate, predicate operator, and index column value carried in the query conditions can be obtained according to the parsing result. For example, the query conditions can be parsed into expression trees, and the predicate, predicate operator, and index column value carried in the query conditions can be obtained from the expression trees.

[0135] Taking the query condition "Select * from customer where Nation = 'China' and City in ('Beijing', 'Nanjing','Shanghai')" as an example, when parsing the query conditions, the obtained expression tree can be as Figure 6 shown, based on Figure 6For the expression tree shown, the predicates in the query condition can be obtained as "nation" and "city", the predicate operators are "=" and "in", and the index column values of the predicates are "China", "Beijing", "Nanjing", and "Shanghai".

[0136] After parsing the predicates, predicate operators, and index column values of the predicates in the query condition, it is possible to further determine whether the index column value range in the index structure includes the index column values of the predicates. For example, if the index column values of the predicates are "China", "Beijing", "Nanjing", and "Shanghai", it can be determined whether the index column value range in the index structure includes "China", "Beijing", "Nanjing", and "Shanghai".

[0137] When the index column value range in the index structure includes the index column values of the predicates, the index column values of the predicates can be further operated according to the predicate operators in the query condition, and the target index column values that meet the query condition can be determined from the index column value range in the index structure according to the operation results. For example, if the index column values of the predicates are "China", "Beijing", "Nanjing", and "Shanghai", and the predicate operators are "=" and "in", if the index column value range includes these index column values of the predicates, then after operating on these index column values according to the predicate operators, the target index column values (i.e., the column value result set that meets the query condition) can be obtained as follows:

[0138] Nation:China

[0139] City: Beijing, Nanjing, Shanghai.

[0140] Optionally, in some embodiments, the index column value range in the index structure may not include the index column values of the predicates. In this case, information indicating query failure can be returned, or data query can be performed according to the data query method in the related art, which is not specifically limited here. The embodiments of the present application are described by taking the case where the index column value range in the index structure includes the index column values of the predicates as an example.

[0141] After determining the target index column values, the target bitstring can be determined according to the target index column values and the two-dimensional bit map in the index structure. The target bitstring can include multiple bits, and these multiple bits are used to indicate whether the target index column values are included in multiple data files.

[0142] In some embodiments, when the number of target index column values is one, in this case, determining the target bitstring according to the target index column values and the two-dimensional bit map in the index structure may include the following steps:

[0143] According to the target index column value, determine a row of bit positions corresponding to the target index column value from the two-dimensional bit map;

[0144] Determine a row of bit positions as the target bit string.

[0145] For example, taking Figure 4 the index structure shown as an example, assuming the target index column value is column value B, then the second row of bit positions "0100100101001100" corresponding to column value B can be determined as the target bit string.

[0146] In some embodiments, the number of target index column values is multiple. In this case, according to the target index column values and the two-dimensional bit map in the index structure, determining the target bit string may include the following steps:

[0147] For each target index column value, determine a row of bit positions corresponding to the target index column value from the two-dimensional bit map, obtaining multiple rows of bit positions corresponding to multiple target index column values;

[0148] According to the relationship between multiple target index column values, perform an operation on the multiple rows of bit positions to obtain the target bit string.

[0149] Taking the target index column values "Nation:China", "City: Beijing, Nanjing, Shanghai" as an example, assuming that in the two-dimensional bit map, the row of bit positions corresponding to "Nation:China" is "001010101101010111", the row of bit positions corresponding to "City: Beijing" is "001010100001010001", the row of bit positions corresponding to "City: Nanjing" is "000010100100010100", and the row of bit positions corresponding to "City: Shanghai" is "001000100101000010". According to the query condition Select * from customer where Nation = 'China' and City in ('Beijing', 'Nanjing','Shanghai'), the relationship between "Nation:China" and "City: Beijing, Nanjing, Shanghai" is "AND", and the relationship between "City: Beijing", "City: Nanjing" and "City: Shanghai" is "OR". Then, when performing an operation on the four rows of bit strings corresponding to these four target index column values, it can be as Figure 7 shown, and the finally obtained target bit string is Figure 7 the "001010100101010111" shown.

[0150] After obtaining the target bit string, valid files can be determined from multiple data files according to the target bit string. The number of valid files can be one or more. In some embodiments, determining valid files from multiple data files according to the target bit string may include the following steps:

[0151] Determine the value of each bit in the target bit string. Wherein, for each bit in the target bit string, a bit value of 0 indicates that the target index column value does not exist (or is not included) in the data file indicated by the file identifier corresponding to the bit, and a bit value of 1 indicates that the target index column value exists (or is included) in the data file indicated by the file identifier corresponding to the bit;

[0152] Determine the target data files indicated by the file identifiers corresponding to the bits with a value of 1;

[0153] Determine the target data files as valid files.

[0154] In this embodiment, for each bit in the two-dimensional bit map, a bit value of 0 indicates that the index column value corresponding to the bit does not exist (or is not included) in the data file indicated by the file identifier corresponding to the bit, and a bit value of 1 indicates that the index column value corresponding to the bit exists (or is included) in the data file indicated by the file identifier corresponding to the bit. Since the target bit string is a row of bits in the two-dimensional bit map or a row of bits obtained by operating on multiple rows of bits in the two-dimensional bit map, therefore, the bits in the target two-dimensional bit string have the same attributes as the bits in the two-dimensional bit map, that is, for each bit in the target bit string, a bit value of 0 indicates that the target index column value does not exist (or is not included) in the data file indicated by the file identifier corresponding to the bit, and a bit value of 1 indicates that the target index column value exists (or is included) in the data file indicated by the file identifier corresponding to the bit. In this way, when determining valid files from multiple data files, the value of each bit in the target bit string can be determined first, that is, determine which bits in the target bit string have a value of 1 and which bits have a value of 0. Since the target index column value exists (or is included) in the data file indicated by the file identifier corresponding to the bit with a value of 1, therefore, the data files indicated by the file identifiers corresponding to the bits with a value of 1 can be determined as target data files, and the target data files are determined as valid files.

[0155] In some embodiments, since the status of the data files includes normal and invalid, and the data files in the invalid status cannot be used as the files to be queried for data query, therefore, after determining the target data files according to the method described above, these files can be further screened according to their status to ensure that the final valid files obtained are the data files with the status of normal. Specifically, the following steps can be included:

[0156] Obtain the file status of the target data file from the index structure;

[0157] For each target data file, determine whether the file status of the target data file is normal;

[0158] When the file status of the target data file is normal, determine the target data file as a valid file.

[0159] The file information of multiple data files stored in the file list information of the index structure can include the file status of multiple data files, and the file status includes normal or invalid. After determining the target data files, the file status of the target data files can be obtained from the index structure, and then for each target data file, it is judged whether the status of the target data file is normal or invalid. If the status of the target data file is normal, the target data file can be determined as a valid file; if the status of the target data file is invalid, the target data file can be determined as an invalid file to be filtered. By making the above judgments on multiple target data files, one or more valid files with the status of normal can be finally obtained. Among them, when judging whether the status of the target data file is normal or invalid, it can be judged by judging whether the status of the target data file recorded in the index structure is 0 or 1. 0 indicates the status is normal, and 1 indicates the status is invalid, so that the status of each target data file can be obtained.

[0160] To facilitate understanding of the specific implementation manner of determining valid files from multiple data files according to the target bit string, reference can be made to Figure 8 . Figure 8 Taking the target bit string as Figure 7Taking the "001010100101010111" shown as an example, when determining valid files from multiple data files according to the target bit string, it is known that the values of the 3rd, 5th, 7th, 10th, 12th, 14th, and 16th to 18th bit positions in the bit string are 1, and the bit orders of these bit positions are 2, 4, 6, 9, 11, 13, 15, 16, and 17 respectively. Therefore, the data files with fileids 2, 4, 6, 9, 11, 13, 15, 16, and 17 corresponding to these bit orders can be determined as the target data files. Further, since the file statuses of these target data files are all 0, that is, the file statuses are all in the normal state, these target data files can be determined as valid files. Then, according to the file paths and names stored in the file list information of the index structure, a valid file list can be obtained.

[0161] It should be noted that in the process of determining valid files according to the index structure above, the predicate operators are illustrated with "=" and "in" as examples. In some embodiments, the predicate operators can also be other operators. Specifically, the predicate operators supported by the embodiments of the present application can include at least one of =, in or not in, >, ≥, <, ≤, betweenrange, like. The following shows how to determine the target bit string and then determine the valid file list for these 8 different predicate operators:

[0162] (1) = (Equality);

[0163] Directly obtain the bitstring mapping the index column value and fileid in the two-dimensional bit map of the column value distribution information, and obtain the Fileid list after parsing.

[0164] (2) in or not in;

[0165] Obtain the index column values included in the in statement, loop to obtain the bitstrings mapping the index column value and fileid in the two-dimensional bit map of the column value distribution information to get multiple bitstrings, perform a vectorized or operation on the obtained multiple bitstrings to obtain a new bitstring, and parse this bitstring to obtain the Fileid list.

[0166] (3) > (Greater than);

[0167] 1) Obtain the set of column values stored in the column value dictionary in the column value distribution information, and judge the result set that conforms to the > operation operator in the set of column values;

[0168] (2) Use the result set to loop through the two-dimensional bitmap of column value distribution information to obtain the bitstrings that map the index column values to fileids. Get multiple bitstrings, perform a vectorized OR operation on the obtained multiple bitstrings to obtain a new bitstring, and parse this bitstring to obtain the Fileid list.

[0169] (4) ≥ (Greater than or equal);

[0170] Same operation as the > operator.

[0171] (5) < (Less than);

[0172] Same operation as the > operator.

[0173] (6) ≤ (Less than or equal);

[0174] Same operation as the > operator.

[0175] (7) between range;

[0176] Same operation as the > operator.

[0177] (8) like.

[0178] Same operation as the > operator.

[0179] In the related art, it is necessary to load all the metadata files storing the Valuelist index into the memory and perform operator calculations one by one. However, the method of the embodiment of the present application is to use the value list with Key = all_Valuelists for calculation, filter out the value set that meets the operator conditions, and then use Key to obtain the corresponding Fileid, reducing the number of calculations to the lowest level and greatly improving the usage efficiency and operation range of the Valuelist index.

[0180] Step S508: Query the target data corresponding to the query request in the valid files.

[0181] After determining the valid files based on step S506, the target data corresponding to the query request can be queried in the valid files to achieve the purpose of data query.

[0182] In some embodiments, the file information of multiple data files in the file list information of the index structure may further include at least one of the file paths and file names of the multiple data files. In this way, when querying the target data corresponding to the query request in the valid files, the following steps may be included:

[0183] Obtain at least one of the file path and file name of the valid file from the index structure;

[0184] Obtain the valid file according to at least one of the file path and file name of the valid file, and query the target data corresponding to the query request in the valid file.

[0185] In this way, since after determining the valid file, the valid file can be quickly located through the file path and / or file name stored in the index structure and data query can be performed in the valid file, the query efficiency can be improved.

[0186] When performing a query, the data query method provided by the embodiment of the present application can obtain index files corresponding to multiple data files, filter out invalid files in the data files according to the index structure in the index files to obtain valid files, and perform data query in the valid files. Among them, since the content stored in the index structure is the file information, index column value range and two-dimensional bit map of multiple data files, the cache space occupied by the index structure is small. Since the bit positions in the two-dimensional bit map of the index structure can be used to record the status of the index column values in the data files, when performing data query, by loading the index structure into the memory and querying the two-dimensional bit map in the index structure, it can be determined whether a certain index column value exists in the data file. Compared with the related technology of loading all the metadata files storing the Valuelist index into the memory for data query, it can not only reduce the occupation of the memory cache space, ensure the overall performance of the host, but also improve the query efficiency.

[0187] Figure 9 It is a method for creating a data file based on an index structure according to an embodiment of the present application. The index structure includes the index structure provided by the embodiment of the present application. Specifically, reference can be made to Figure 3 and Figure 4 The embodiments shown, and the description will not be repeated here. Figure 9 The data file creation method shown can be executed by Figure 2 The big data execution engine in the computing layer 21 shown. Figure 9 The method shown includes the following steps.

[0188] Step S902: Receive a file creation request for the first data file.

[0189] The first data file is the file to be created. When creating the first data file, a file creation request for the first data file can be received.

[0190] Step S904: Determine the file information of the first data file, and determine whether there is an index column value in the index column value range in the first data file according to the index column value range in the index structure. The file information of the first data file includes the file identifier of the first data file.

[0191] When creating the first data file, the index structure provided by the embodiments of the present application can be updated and maintained. When updating and maintaining the index structure, the file information of the first data file can be determined, and it can be determined whether there is (or whether it includes) an index column value in the index column value range in the first data file according to the index column value range in the index structure.

[0192] The file information of the first data file can include a file identifier, and the file identifier can specifically be a file serial number. When determining the file serial number of the first data file, it can be determined according to the generation order of the first data file. For example, before creating the first data file, the file serial number of the latest data file is 2000, then the file serial number of the first data file can be determined to be 2001. Optionally, the file information of the first data file can also include at least one of a file name, a file path, and a file status. Among them, the file status includes normal and invalid. Since the first data file is a newly created file, the file status of the first data file can be the normal status.

[0193] When determining whether there is (or whether it includes) an index column value in the index column value range in the first data file according to the index column value range in the index structure, it can first be determined which index column values are specifically included in the index column value range, and then for each index column value, it can be counted whether there is such an index column value in the first data file, so that the status (existence or non-existence) of each index column value in the first data file can be obtained.

[0194] Step S906: Store the file information of the first data file in the file list information of the index structure, and record the status of the index column value in the index column value range in the first data file in the two-dimensional bitmap of the index structure.

[0195] After determining the file information of the first data file, the file information of the first data file can be stored (or recorded) in the file list information of the index structure. After obtaining the status of each index column value in the index column value range in the first data file, the status of the index column value in the index column value range in the first data file can be recorded in the two-dimensional bitmap of the index structure. Specifically, a column of bit positions corresponding to the file identifier of the first data file can be determined in the two-dimensional bitmap. For example, if the file identifier of the first data file is the file serial number and the file serial number is 2001, then it can be determined that a column of bit positions with the bit position order of 2001 in the two-dimensional bitmap is the column of bit positions corresponding to the file identifier of the first data file. Then, the status of each index column value in the index column value range in this column of bit positions can be recorded in the first data file. For example, a column of bit positions corresponding to the file identifier of the first data file corresponds to 3 index column values, namely column value A, column value B, and column value C. It is known that the first data file includes column value A and B and does not include column value C. Then, in this column of bit positions corresponding to the file identifier of the first data file, the bit position corresponding to column value A can be recorded as 1, the bit position corresponding to column value B can be recorded as 1, and the bit position corresponding to column value C can be recorded as 0.

[0196] After storing the file information of the first data file in the file list information of the index structure and recording the status of each index column value in the index structure in the first data file in the two-dimensional bitmap of the index structure, the update and maintenance of the index structure can be realized when creating a new data file.

[0197] Step S908: Create a first data file according to the file creation request.

[0198] After updating and maintaining the index structure through the above steps S904 and S906, a first data file can be created according to the file creation request.

[0199] It should be noted that in the above steps S902 to S908, when creating the first data file, the index structure is first updated and maintained, and then the first data file is created. In other embodiments, the first data file can also be created first and then the index structure is updated and maintained, or the creation operation of the first data file and the update and maintenance operation of the index structure can be executed in parallel. There is no specific limitation here.

[0200] In some embodiments, when creating the first data file, the specific creation process can be as Figure 10 shown. Figure 10After creating a data file in memory, the file list information maintenance module in the index unit (which can be a module unit for updating and maintaining the index structure) can first create file list information and generate a file ID. Then, the column value distribution information maintenance module in the index unit can create column value distribution information based on the file ID and the column value information in the statistical data file. After both of these two pieces of data are created, it notifies the data import program to write the data to the data directory and returns a successful data import response.

[0201] Based on the index structure provided by the embodiments of the present application, when creating a data file, by modifying the file information stored in the index structure and the status of the bit records in the two-dimensional bitmap, the update and maintenance of the index structure can be achieved, without frequently loading and updating the metadata file, which has less impact on the performance of the machine and is easy to implement the update and maintenance of the index structure.

[0202] Figure 11 is a data processing method based on the index structure in an embodiment of the present application. The index structure includes the index structure provided by the embodiments of the present application. Specifically, reference can be made to Figure 3 and Figure 4 the embodiments shown, and the description will not be repeated here. Figure 11 The data processing method shown can be executed by Figure 2 the big data execution engine in the computing layer 21 shown. Figure 11 The method shown includes the following steps.

[0203] Step S112: Receive a data processing request for the second data file.

[0204] When performing data processing on the second data file, a data processing request for the second data file can be received. Among them, the data processing on the second data file can include at least one of the following:

[0205] Adding data to the second data file;

[0206] Deleting data from the second data file;

[0207] Updating data in the second data file.

[0208] Step S114: Perform data processing on the second data file according to the data processing request to obtain a third data file.

[0209] Specifically, the second data file can be loaded into memory, and then the data in the second data file is processed accordingly according to the data processing request to obtain a third data file, which is a data file obtained by processing the data in the second data file.

[0210] Step S116: Determine the file information of the third data file, and determine whether there is an index column value in the index column value range in the third data file according to the index column value range in the index structure. The file information of the third data file includes the file identifier of the third data file.

[0211] After obtaining the third data file, the index structure provided by the embodiments of the present application can be updated and maintained. When updating and maintaining the index structure, the file information of the third data file can be determined, and it can be determined whether there is (or includes) an index column value in the index column value range in the third data file according to the index column value range in the index structure.

[0212] The file information of the third data file may include a file identifier, and the file identifier may specifically be a file serial number. When determining the file serial number of the third data file, it can be determined according to the generation order of the third data file. For example, before generating the third data file, the file serial number of the latest data file is 2001, then the file serial number of the third data file can be determined to be 2002. Optionally, the file information of the third data file may further include at least one of a file name, a file path, and a file status. Among them, the file name and file path of the third data file can be determined according to the basic information of the first data file. The file status includes normal and invalid. Since the third data file is a file obtained by processing the second data file, the file status of the third data file can be a normal status.

[0213] When determining whether there is (or includes) an index column value in the index column value range in the third data file according to the index column value range in the index structure, it can be first determined which index column values are specifically included in the index column value range, and then for each index column value, it can be counted whether there is (the index column value in the third data file, so that the status (existence or non-existence) of each index column value in the third data file can be obtained.

[0214] Step S118: Store the file information of the third data file in the file list information of the index structure, and record the status of the index column value in the index column value range in the third data file in the two-dimensional bitmap of the index structure.

[0215] After determining the file information of the third data file, the file information of the third data file can be stored (or recorded) in the file list information of the index structure. After obtaining the status of each index column value in the index column value range in the third data file, the status of the index column value in the index column value range in the third data file can be recorded in the two-dimensional bitmap of the index structure. Specifically, a column of bit positions corresponding to the file identifier of the third data file in the two-dimensional bitmap can be determined first. For example, if the file identifier of the third data file is the file serial number and the file serial number is 2002, then it can be determined that the column of bit positions with the bit position sequence of 2002 in the two-dimensional bitmap is the column of bit positions corresponding to the file identifier of the third data file. Then, the status of each index column value in the index column value range in the third data file is recorded in this column of bit positions. For example, a column of bit positions corresponding to the file identifier of the third data file corresponds to 3 index column values, which are column value A, column value B, and column value C respectively. It is known that the third data file includes column value A and does not include column value B and C. Then, in this column of bit positions corresponding to the file identifier of the third data file, the bit position corresponding to column value A is recorded as 1, the bit position corresponding to column value B is recorded as 0, and the bit position corresponding to column value C is recorded as 0.

[0216] After storing the file information of the third data file in the file list information of the index structure and recording the status of each index column value in the index structure in the third data file in the two-dimensional bitmap of the index structure, it is possible to realize the update and maintenance of the index structure when processing data of the data file.

[0217] Optionally, in some embodiments, when processing the second data file to obtain the third data file, for the second data file, it may include at least one of the following:

[0218] Delete the file information of the second data file in the file list information of the index structure and delete the records in the bit positions corresponding to the file identifier of the second data file in the two-dimensional bitmap of the index structure;

[0219] Modify the file status of the second data file in the file list information of the index structure from normal to invalid.

[0220] Since the second data file becomes an old file after processing the second data file to obtain the third data file, the records related to the second data file can be deleted in the index structure, that is, delete the file information of the second data file in the file list information of the index structure and the records in the bit positions corresponding to the file identifier of the second data file in the two-dimensional bitmap of the index structure, so as to reduce the space occupied by the index structure.

[0221] Since the second data file becomes an old file after the third data file is obtained by processing the second data file, the file status of the second data file changes from the normal state to the invalid state. At this time, the file status of the second data file recorded in the file list of the index structure can be modified from normal to invalid. In this way, when performing data queries subsequently, it can be determined that the second data file is an invalid file based on the file status, thereby avoiding the problem of incorrect query results caused by using the second data file as the file to be queried.

[0222] In some embodiments, when modifying the data in the data file, the specific operation process can be as Figure 12 shown, and specifically may include the following steps:

[0223] Step 1: First, obtain the data file where the data to be modified is located, load the data file into memory and modify the corresponding content, and generate a new file in memory.

[0224] Step 2: Notify the index unit to first modify the file list information, delete the fileid information of the old file, and allocate a new fileid for the newly created file.

[0225] The index unit is the same as the index unit in the Figure 10 illustrated embodiment, and can be a module unit for updating and maintaining the index structure, and may include a file list information maintenance module and a column value distribution information maintenance module. The file list information maintenance module can be used to modify the file list information, delete the fileid information of the old file, and allocate a new fileid for the newly created file.

[0226] Step 3: Use the new fileid and the column value information in the newly counted data file to continue creating the column value distribution information, and at the same time set the file status of the old file to 1 (1 indicates invalid).

[0227] Step 3 can be executed by the column value distribution information maintenance module in the index unit.

[0228] Step 4: After the above steps 2 and 3 are completed, notify the data import program to write the data to the data directory and return a data import success response.

[0229] Step 5: The old file will be marked as the delete state.

[0230] The file in the delete state is the file to be deleted. When deleting a file, the specific file deletion process is the expired / junk file cleaning process. In some embodiments, it can be as Figure 13 shown. Figure 13In it, when deleting an old file, first delete the file ID of the file to be deleted in the file list information, then set the bit position corresponding to the file ID in the column value distribution information to 0, and then notify the deletion program to perform the data file deletion operation.

[0231] In some embodiments, when deleting an old file and maintaining the index structure, the simplest operation only needs to maintain the file status corresponding to the old file in the file list information, and the column value distribution information does not need to be modified. The specific process can be as Figure 14 shown. When performing data query, after determining the target data file, it is only necessary to further determine whether the target data file is a valid file according to the file status recorded in the file list information in the index structure, which is not only simple in operation but also convenient for data traceback processing.

[0232] Based on the index structure provided by the embodiments of the present application, when updating the data in the data file, by modifying the file information stored in the index structure and the status recorded by the bits in the two-dimensional bitmap, the update and maintenance of the index structure can be realized without frequently loading and updating the metadata file, which has less impact on the performance of the machine and is easy to realize the update and maintenance of the index structure.

[0233] In summary, the embodiments of the present application provide a new storage format for column value statistical information, which occupies less space than the metadata file storing the Valuelist index in the related art, and can effectively reduce the system memory occupancy and IO.

[0234] Based on the new index storage format, a method for obtaining a query result set is proposed, which can support vectorized operations between single indexes and multiple indexes at the same time, greatly accelerating the index calculation efficiency. At the same time, the applicable range of the index operator is extended to support more query expressions, including but not limited to Equality, IN / Not IN, Greater than, Greater than or equal, Less than or equal, Between range, Like operations, making the application range of the column value index technology wider. During the process of using the Valuelist index, according to the predicates and corresponding conditions existing in the expression tree, obtain the set of column values that meet the query conditions in the column value dictionary, Get the bitstring content corresponding to the column value in the two-dimensional bitmap, and directly obtain the file list containing the query data after calculation. This process does not need to cache any valuelist index, and finally use the file ID to obtain the actual data file list in the file list information, thereby reducing the memory system space occupied by loading the Valuelist index data, and can extract all index data at one time for vector operation, improving the calculation efficiency and the usage efficiency of the host.

[0235] Based on the new index storage format, in terms of index maintenance, the file list information and column value distribution information in the new index storage format are loosely coupled. For the generation and deletion of data files, in the best case, only the file status of the files corresponding to the file list information needs to be maintained, and the column value distribution information does not need to be changed. This can greatly improve the efficiency of index changes and make index maintenance more convenient.

[0236] When performing data queries, query optimization includes but is not limited to scenarios such as partition pruning, file filtering, and cost-based optimization (Cost-Based Optimization, CBO) cost optimization. Among them:

[0237] During the query process of big data execution engines (such as spark, hive, flink, presto, etc.), the data query method of the embodiments of the present application can be used to improve query efficiency;

[0238] During the query process of data warehouse engines (such as Greeplum, Doris, etc.), the data query method of the embodiments of the present application can be used to improve query efficiency;

[0239] During the query process of database engines (such as Postgresql, etc.), the data query method of the embodiments of the present application can be used to improve query efficiency;

[0240] Under the lakehouse architecture, during the query process of big data execution engines (such as spark, hive, flink, presto, etc.) based on TableFormat, the data query method of the embodiments of the present application can be used to improve query efficiency;

[0241] Under the lakehouse architecture, during the query process of data warehouse engines (such as Greeplum, Doris, etc.) based on TableFormat, the data query method of the embodiments of the present application can be used to improve query efficiency;

[0242] During the query process of cloud-native data warehouse engines (such as SelectDB Cloud, etc.) under the cloud-native architecture, the data query method of the embodiments of the present application can be used to improve query efficiency.

[0243] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0244] Figure 15 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 15 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0245] The processor, the network interface, and the memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 15 only a bidirectional arrow is used in

[0246] Memory, for storing programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0247] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a data query device based on an index structure at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0248] Receive a query request;

[0249] Obtain index files corresponding to multiple data files, and the index files include the index structure;

[0250] Determine valid files from the multiple data files according to the query request and the index structure;

[0251] Query for target data corresponding to the query request in the valid files.

[0252] The above is as in the present application Figure 15The method executed by the data query device based on the index structure disclosed in the illustrated embodiment can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute various methods, steps, and logic block diagrams disclosed in this application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with this application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0253] The electronic device can also execute Figure 5 the method and implement the functions of the data query device based on the index structure in Figure 5 the illustrated embodiment, which will not be elaborated herein.

[0254] Of course, in addition to the software implementation, the electronic device of this application does not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0255] This application also proposes a computer-readable storage medium. The computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by a portable electronic device including multiple application programs, can enable the portable electronic device to execute Figure 5 the method of the illustrated embodiment and specifically be used to perform the following operations:

[0256] Receive a query request;

[0257] Obtain an index file corresponding to multiple data files, where the index structure is included in the index file;

[0258] Determine valid files from the multiple data files according to the query request and the index structure;

[0259] Query for target data corresponding to the query request in the valid files.

[0260] Figure 16 It is a schematic structural diagram of a data query device 160 based on an index structure according to an embodiment of the present application. Please refer to Figure 16 In a software implementation manner, the data query device 160 based on the index structure may include: a receiving module 161, an obtaining module 162, a determining module 163, and a query module 164, where:

[0261] The receiving module 161 receives a query request;

[0262] The obtaining module 162 obtains an index file corresponding to multiple data files, where the index structure is included in the index file;

[0263] The determining module 163 determines valid files from the multiple data files according to the query request and the index structure;

[0264] The query module 164 queries for target data corresponding to the query request in the valid files.

[0265] In some embodiments, the determining module 163 determines valid files from the multiple data files according to the query request and the index structure, including:

[0266] Determine a target index column value that meets the query condition from the index column value range of the index structure according to the query condition carried in the query request;

[0267] Determine a target bit string according to the target index column value and the two-dimensional bit map in the index structure, where the target bit string includes multiple bits, and the multiple bits are used to indicate whether the target index column value is included in the multiple data files;

[0268] Determine valid files from the multiple data files according to the target bit string.

[0269] In some embodiments, the determining module 163 determines a target index column value that meets the query condition from the index column value range of the index structure according to the query condition carried in the query request, including:

[0270] Parse the query condition to obtain the predicate, predicate operator, and index column value of the predicate in the query condition;

[0271] Determine whether the index column value range of the index structure includes the index column value of the predicate;

[0272] When the index column value range of the index structure includes the index column value range of the predicate, perform an operation on the index column value of the predicate according to the predicate operator, and determine the target index column value that meets the query condition from the index column value range of the index structure according to the operation result.

[0273] In some embodiments, the predicate operator includes at least one of the following:

[0274] =; in or not in; >; ≥; <; ≤; between range; like.

[0275] In some embodiments, the number of the target index column values is one or more; the determining module 163 determines a target bit string according to the target index column value and the two-dimensional bit map in the index structure, including:

[0276] When the number of the target index column values is one, determine a row of bit positions corresponding to the target index column value from the two-dimensional bit map according to the target index column value; determine the row of bit positions as the target bit string;

[0277] When the number of the target index column values is multiple, for each target index column value, determine a row of bit positions corresponding to the target index column value from the two-dimensional bit map to obtain multiple rows of bit positions corresponding to the multiple target index column values; perform an operation on the multiple rows of bit positions according to the relationship between the multiple target index column values to obtain the target bit string.

[0278] In some embodiments, the determining module 163 determines valid files from the multiple data files according to the target bit string, including:

[0279] Determine the value of each bit position in the target bit string, where for each bit position in the target bit string, the bit position value of 0 indicates that the target index column value does not exist in the data file indicated by the file identifier corresponding to the bit position, and the bit position value of 1 indicates that the target index column value exists in the data file indicated by the file identifier corresponding to the bit position;

[0280] Determine the target data file indicated by the file identifier corresponding to the bit position with a value of 1;

[0281] Determine the target data file as a valid file.

[0282] In some embodiments, in the index structure, the file information of the multiple data files further includes the file status of the multiple data files, and the file status includes normal or invalid; after determining the target data file, the determining module 163 further includes:

[0283] Obtain the file status of the target data file from the index structure;

[0284] For each target data file, determine whether the file status of the target data file is normal;

[0285] In the case where the file status of the target data file is normal, determine the target data file as a valid file.

[0286] In some embodiments, the file information of the multiple data files further includes at least one of the file path and file name of the multiple data files; the query module 164 queries the target data corresponding to the query request from the valid files, including:

[0287] Obtain at least one of the file path and file name of the valid file from the index structure;

[0288] Obtain the valid file according to at least one of the file path and file name of the valid file, and query the target data corresponding to the query request in the valid file.

[0289] In some embodiments, the obtaining module 162 determines the index structure in the following manner:

[0290] Determine the file information of the multiple data files, where the file information includes file identifiers;

[0291] According to the index column names in the index column configuration information, determine the index column values corresponding to the index column names in the multiple data files to obtain an index column value range;

[0292] Generate file list information according to the file information of the multiple data files, and generate column value distribution information according to the index column value range and whether each data file has an index column value in the index column value range;

[0293] Construct the index structure according to the file list information and the column value distribution information.

[0294] The data query device 160 based on the index structure provided by this application can also execute Figure 5method, and implement the functions of the data query device 160 based on the index structure in Figure 5 the embodiments shown, which are not elaborated herein.

[0295] Figure 17 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 17 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0296] The processor, network interface, and memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 17 only a bidirectional arrow is used in

[0297] Memory, for storing programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0298] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, and forms a data file creation device based on the index structure at the logical level. The processor executes the program stored in the memory, and specifically is used to perform the following operations:

[0299] Receive a file creation request for the first data file;

[0300] Determine the file information of the first data file, and determine whether there is an index column value in the index column value range in the first data file according to the index column value range in the index structure, and the file information of the first data file includes the file identifier of the first data file;

[0301] Store the file information of the first data file in the file list information of the index structure, and record the status of the index column value in the index column value range in the first data file in the two-dimensional bitmap of the index structure;

[0302] Create the first data file according to the file creation request.

[0303] As described in this application Figure 17 The method executed by the data file creation device based on the index structure disclosed in the above embodiments of the present application can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in this application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with this application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0304] The electronic device can also execute Figure 9 the method, and implement the functions of the data file creation device based on the index structure in Figure 9 the embodiments shown. Details are not described herein again in this application.

[0305] Of course, in addition to the software implementation manner, the electronic device of this application does not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.

[0306] The present application also provides a computer-readable storage medium storing one or more programs including instructions that, when executed by a portable electronic device including a plurality of application programs, enable the portable electronic device to execute Figure 9 the method of the embodiment shown in

[0307] and specifically used to perform the following operations:

[0308] Receiving a file creation request for a first data file;

[0309] Determining the file information of the first data file, and determining whether there is an index column value in the index column value range in the first data file according to the index column value range in the index structure, where the file information of the first data file includes the file identifier of the first data file;

[0310] Storing the file information of the first data file in the file list information of the index structure, and recording the status of the index column value in the index column value range in the first data file in the two-dimensional bitmap of the index structure;

[0311] Figure 18 FIG. 18 is a schematic structural diagram of a data file creation device 180 based on an index structure according to an embodiment of the present application. Please refer to Figure 18 , in a software implementation, the data file creation device 180 based on the index structure may include: a receiving module 181, a determining module 182, an index maintenance module 183, and a file creation module 184, where:

[0312] The receiving module 181 receives a file creation request for a first data file;

[0313] The determining module 182 determines the file information of the first data file, and determines whether there is an index column value in the index column value range in the first data file according to the index column value range in the index structure, where the file information of the first data file includes the file identifier of the first data file;

[0314] The index maintenance module 183 stores the file information of the first data file in the file list information of the index structure, and records the status of the index column value in the index column value range in the first data file in the two-dimensional bitmap of the index structure;

[0315] The file creation module 184 creates the first data file according to the file creation request.

[0316] The data file creation device 180 based on the index structure provided by this application can also execute Figure 9 the method, and implement the functions of the data file creation device 180 based on the index structure in Figure 9 the embodiments shown. This application will not elaborate further here.

[0317] Figure 19 FIG. is a schematic structural diagram of an electronic device according to an embodiment of this application. Please refer to Figure 19 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include internal memory, such as high-speed random access memory (Random-Access Memory, RAM), and may also include non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0318] The processor, network interface, and memory can be interconnected through the internal bus. The internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 19 only a bidirectional arrow is used in, but it does not mean that there is only one bus or one type of bus.

[0319] The memory is used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include internal memory and non-volatile memory, and provide instructions and data to the processor.

[0320] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a data processing device based on the index structure at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0321] Receive a data processing request for the second data file;

[0322] Perform data processing on the second data file according to the data processing request to obtain a third data file;

[0323] Determine the file information of the third data file, and determine whether there is an index column value in the index column value range in the third data file according to the index column value range in the index structure. The file information of the third data file includes the file identifier of the third data file.

[0324] Store the file information of the third data file in the file list information of the index structure, and record the status of the index column value in the index column value range in the third data file in the two-dimensional bitmap of the index structure.

[0325] The above as in this application Figure 19 The method executed by the data processing device based on the index structure disclosed in the embodiments shown in this application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in this application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with this application can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0326] The electronic device can also execute Figure 11 the method, and implement the functions of the data processing device based on the index structure in Figure 11 the embodiments shown. This application will not elaborate here.

[0327] Of course, in addition to the software implementation, the electronic device of the present application does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.

[0328] The present application also provides a computer-readable storage medium storing one or more programs, where the one or more programs include instructions that, when executed by a portable electronic device including a plurality of application programs, can cause the portable electronic device to execute Figure 11 the method of the illustrated embodiment, and specifically used to perform the following operations:

[0329] Receive a data processing request for a second data file;

[0330] Perform data processing on the second data file according to the data processing request to obtain a third data file;

[0331] Determine the file information of the third data file, and determine whether there is an index column value in the index column value range in the third data file according to the index column value range in the index structure. The file information of the third data file includes the file identifier of the third data file;

[0332] Store the file information of the third data file in the file list information of the index structure, and record the status of the index column value in the index column value range in the two-dimensional bitmap of the index structure in the third data file.

[0333] Figure 20 FIG. is a schematic structural diagram of a data processing apparatus 200 based on an index structure according to an embodiment of the present application. Please refer to Figure 20 , in a software implementation manner, the data processing apparatus 200 based on the index structure may include: a receiving module 201, a processing module 202, a determining module 203, and an index maintenance module 204, where:

[0334] The receiving module 201 receives a data processing request for a second data file;

[0335] The processing module 202 performs data processing on the second data file according to the data processing request to obtain a third data file;

[0336] The determining module 203 determines the file information of the third data file, and determines whether there is an index column value in the index column value range in the third data file according to the index column value range in the index structure. The file information of the third data file includes the file identifier of the third data file;

[0337] The index maintenance module 204 stores the file information of the third data file in the file list information of the index structure, and records the status of the index column values in the index column value range in the two-dimensional bitmap of the index structure in the third data file.

[0338] In some embodiments, the data processing includes at least one of the following:

[0339] Adding data to the second data file;

[0340] Deleting data from the second data file;

[0341] Updating data in the second data file.

[0342] In some embodiments, the index maintenance module 204 further includes any one of the following:

[0343] Deleting the file information of the second data file in the file list information of the index structure, and deleting the records in the bit positions corresponding to the file identifier of the second data file in the two-dimensional bitmap of the index structure;

[0344] Changing the file status of the second data file from normal to invalid in the file list information of the index structure.

[0345] The data processing device 200 based on the index structure provided by this application can also execute Figure 11 the method, and implement the functions of the data query device 200 based on the index structure in Figure 11 the embodiments shown. This application will not elaborate here.

[0346] In summary, the above are only the preferred embodiments of this application, and are not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

[0347] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0348] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0349] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0350] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

Claims

1. An index structure, comprising file list information and column value distribution information, wherein: The file list information includes file information of multiple data files, and the file information includes a file identifier; The column value distribution information includes an index column value range and a two-dimensional bitmap, and there is a corresponding relationship between the bits in the two-dimensional bitmap, the file identifier of the data file, and the index column values in the index column value range. For any bit, the bit is used to record the status of the index column value corresponding to the bit in the data file indicated by the file identifier corresponding to the bit.

2. The index structure according to claim 1, wherein the file list information includes a multi-dimensional file list for storing the file information of the multiple data files.

3. The index structure according to claim 1 or 2, wherein the file information further includes at least one of a file path, a file name, and a file status, and the file status includes normal or invalid.

4. The index structure according to claim 1, wherein the file identifier includes a file serial number, and the file serial numbers of the multiple data files are determined according to the generation order of the multiple data files; Among them, For any data file, the file serial number of the data file is equal to the bit order of the bit corresponding to the file identifier of the data file, or equal to the sum of the bit order and a specified offset value.

5. The index structure according to claim 1, wherein the file list information further includes a first data dictionary for storing a table partition field value range, and the multiple data files include data files under multiple table partitions.

6. The index structure according to claim 1, wherein the column value distribution information includes a second data dictionary for storing the index column value range.

7. The index structure according to claim 1, in the two-dimensional bitmap, each row of bits corresponds to an index column value, and each column of bits corresponds to a file identifier of a data file; the status includes that there is an index column value corresponding to the bit in the data file indicated by the file identifier corresponding to the bit, or there is no index column value corresponding to the bit in the data file indicated by the file identifier corresponding to the bit; Among them, For any bit, the bit takes a value of 0 indicating that there is no index column value corresponding to the bit in the data file indicated by the file identifier corresponding to the bit, and the bit takes a value of 1 indicating that there is an index column value corresponding to the bit in the data file indicated by the file identifier corresponding to the bit.

8. The index structure according to claim 1 or 7, wherein the two-dimensional bitmap is stored in the form of a hexadecimal byte string.

9. The index structure according to claim 1, 6 or 7, wherein the data type of the index column values in the index column value range includes integer type, floating point type, character type, and date type.

10. A data query method based on an index structure, the index structure includes the index structure according to any one of claims 1 to 9, and the method includes: Receiving a query request; Obtain an index file corresponding to multiple data files, where the index structure is included in the index file; Determine valid files from the multiple data files according to the query request and the index structure; Query for target data corresponding to the query request in the valid files.

11. The method according to claim 10, where the determining valid files from the multiple data files according to the query request and the index structure includes: Determine a target index column value that meets the query conditions from the index column value range of the index structure according to the query conditions carried in the query request; Determine a target bit string according to the target index column value and the two-dimensional bit map in the index structure, where the target bit string includes multiple bits, and the multiple bits are used to indicate whether the multiple data files include the target index column value; Determine valid files from the multiple data files according to the target bit string.

12. The method according to claim 11, where the determining a target index column value that meets the query conditions from the index column value range of the index structure according to the query conditions carried in the query request includes: Parse the query conditions to obtain the predicate, predicate operator, and index column value of the predicate in the query conditions; Determine whether the index column value range of the index structure includes the index column value of the predicate; When the index column value range of the index structure includes the index column value range of the predicate, perform an operation on the index column value of the predicate according to the predicate operator, and determine a target index column value that meets the query conditions from the index column value range of the index structure according to the operation result.

13. The method according to claim 12, where the predicate operator includes at least one of the following: =; in or not in; >; ≥; <; ≤; between range; like.

14. The method according to claim 11, where the number of target index column values is one or more; the determining a target bit string according to the target index column value and the two-dimensional bit map in the index structure includes: When the number of target index column values is one, determine a row of bits corresponding to the target index column value from the two-dimensional bit map according to the target index column value; and determine the row of bits as the target bit string; When the number of target index column values is multiple, for each target index column value, determine a row of bits corresponding to the target index column value from the two-dimensional bit map to obtain multiple rows of bits corresponding to the multiple target index column values; and perform an operation on the multiple rows of bits according to the relationship between the multiple target index column values to obtain the target bit string.

15. The method according to claim 11, where the determining valid files from the multiple data files according to the target bit string includes: Determine the value of each bit in the target bit string. For each bit in the target bit string, a value of 0 for the bit indicates that the target index column value does not exist in the data file indicated by the file identifier corresponding to the bit, and a value of 1 for the bit indicates that the target index column value exists in the data file indicated by the file identifier corresponding to the bit; Determine the target data files indicated by the file identifiers corresponding to the bits with a value of 1; Determine the target data files as valid files.

16. The method according to claim 15, wherein in the index structure, the file information of the multiple data files further includes the file status of the multiple data files, and the file status includes normal or invalid; after determining the target data files, the method further includes: Obtain the file status of the target data files from the index structure; For each of the target data files, determine whether the file status of the target data file is normal; In the case where the file status of the target data file is normal, determine the target data file as a valid file.

17. The method according to claim 10, wherein in the index structure, the file information of the plurality of data files further includes at least one of the file paths and file names of the plurality of data files; The querying for target data corresponding to the query request from the valid files includes: Obtain at least one of the file path and file name of the valid files from the index structure; Obtain the valid files according to at least one of the file path and file name of the valid files, and query for the target data corresponding to the query request in the valid files.

18. The method according to any one of claims 10 to 17, wherein the index structure is determined by the following method: Determine the file information of the multiple data files, where the file information includes file identifiers; Determine the index column values corresponding to the index column names in the multiple data files according to the index column names in the index column configuration information to obtain an index column value range; Generate file list information according to the file information of the multiple data files, and generate column value distribution information according to the index column value range and whether the index column values in the index column value range exist in each data file; Construct the index structure according to the file list information and the column value distribution information.

19. A method for creating a data file based on an index structure, where the index structure includes the index structure according to any one of claims 1 to 9, and the method includes: Receive a file creation request for a first data file; Determine the file information of the first data file, and determine whether the index column values in the index column value range exist in the first data file according to the index column value range in the index structure. The file information of the first data file includes the file identifier of the first data file; Store the file information of the first data file in the file list information of the index structure, and record the status of the index column values in the index column value range in the first data file in the two-dimensional bit map of the index structure; Create the first data file according to the file creation request.

20. A data processing method based on an index structure, where the index structure includes the index structure described in any one of claims 1 to 9, and the method includes: Receiving a data processing request for a second data file; Performing data processing on the second data file according to the data processing request to obtain a third data file; Determining the file information of the third data file, and determining whether there is an index column value in the index column value range in the third data file according to the index column value range in the index structure, where the file information of the third data file includes the file identifier of the third data file; Storing the file information of the third data file in the file list information of the index structure, and recording the status of the index column value in the index column value range in the third data file in the two-dimensional bitmap of the index structure.

21. The method according to claim 20, where the data processing includes at least one of the following: Adding data to the second data file; Deleting data from the second data file; Updating data in the second data file.

22. The method according to claim 20 or 21, and the method further includes any one of the following: Deleting the file information of the second data file in the file list information of the index structure, and deleting the records in the bit positions corresponding to the file identifier of the second data file in the two-dimensional bitmap of the index structure; Modifying the file status of the second data file from normal to invalid in the file list information of the index structure.

23. An electronic device, including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the method described in any one of claims 10 to 22.

24. A computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in any one of claims 10 to 22.

25. A computer program product, where the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps in the method described in any one of claims 10 to 22.