Data processing method, data query method, and data processing and query system
By dividing the column maximum value and column minimum value of the data table, a column statistical range set with discontinuous intervals is generated, which solves the problem of low data query efficiency in the existing technology, and realizes more efficient data skipping and querying.
Patent Information
- Application Number
- PCT/CN2024/143606
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-17
AI Technical Summary
In the prior art, the data skipping technology based on the column minimum value and the column maximum value has a large range, resulting in more data being scanned during data query, which reduces query efficiency.
By partitioning the data range composed of column maximum value and column minimum value, a column statistical range set of multiple discontinuous intervals is generated to describe the data distribution characteristics more accurately and make finer-grained data skip judgments during data query.
Effectively filter out invalid tables, partitions and files, reduce the scanned data range, and improve data query efficiency.
Smart Images

Figure CN2024143606_17072025_PF_FP_ABST
Abstract
Description
Data processing method, data query method, data processing and query system
[0001] Cross-references
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on January 12, 2024, with application number 202410053441.1 and application name “Data processing method, data query method, data processing and query system”. The entire contents of this application are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of data processing technology, and in particular to a data processing method, a data query method, and a data processing and query system. Background Art
[0004] When storing data in a database, statistics are usually collected for each column of data to obtain summary information such as the maximum and minimum values of the column. This allows data queries to be performed based on data skipping technology when a query request is received. Data skipping is a technology that performs Structured Query Language (SQL) analysis on structured data. Data skipping stores summary information such as the minimum and maximum values of columns in each object (or file, etc.). When performing data queries, the minimum and maximum values of columns can be used as column indexes. If the column index value range does not overlap with the data range specified by the query, the query engine will skip the object when scanning, thereby reducing the amount of data scanned by the query, speeding up query efficiency, and reducing resource costs.
[0005] However, in actual applications, the range between the maximum value and the minimum value of a column is usually large. Therefore, when performing data queries based on data skipping technology, the amount of skipped data will be small, which in turn will result in more data being scanned during data scanning, reducing query efficiency. Summary of the Invention
[0006] The present application provides a data processing method, a data query method, and a data processing and query system.
[0007] This application is implemented as follows:
[0008] In a first aspect, a data processing method is provided, including: obtaining first statistical information, the first statistical information including column statistical information obtained by performing statistics on columns of a data table, the column statistical information including a column maximum value and a column minimum value; generating second statistical information for data query based on the first statistical information, the second statistical information including a column statistical range set, the column statistical range set including a plurality of discontinuous intervals, the plurality of discontinuous intervals being obtained by dividing a data range constituted by a column maximum value and a column minimum value in the column statistical information into intervals.
[0009] In a second aspect, a data query method is provided, including: receiving a query request; obtaining second statistical information of a data table, wherein the second statistical information is obtained based on the data processing method described in the first aspect; determining a valid object from the data table based on the query request and the second statistical information, wherein the valid object includes at least one of a valid table, a valid partition, and a valid file; and querying target data corresponding to the query request from the valid object.
[0010] According to a third aspect, a data processing and query system is provided, comprising a storage layer and a computing layer, wherein: the storage layer obtains first statistical information, the first statistical information comprises column statistical information obtained by performing statistics on columns of a data table, the column statistical information comprises column maximum and column minimum values; second statistical information for data query is generated based on the first statistical information, the second statistical information comprises a column statistical range set, the column statistical range set comprises a plurality of discontinuous intervals, the plurality of discontinuous intervals being obtained by dividing a data range constituted by the column maximum and column minimum values in the column statistical information; the computing layer receives a query request; obtains the second statistical information; determines a valid object from the data table based on the query request and the second statistical information, the valid object comprising at least one of a valid table, a valid partition and a valid file; and queries the target data corresponding to the query request from the valid object.
[0011] In a fourth aspect, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method described in the first aspect or the second aspect.
[0012] In a fifth aspect, a computer-readable storage medium is provided, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in this application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0014] FIG1 is a schematic diagram of data skipping using column minimum and column maximum values in the related art;
[0015] FIG2 is a schematic diagram of a schematic system architecture provided in an embodiment of the present application;
[0016] FIG3 is a flow chart of a data processing method according to an embodiment of the present application;
[0017] FIG4 is a schematic diagram of a data structure corresponding to first statistical information according to an embodiment of the present application;
[0018] FIG5 is a schematic diagram of generating file-level column statistics according to an embodiment of the present application;
[0019] FIG6 is a schematic diagram of generating partition-level column statistics according to an embodiment of the present application;
[0020] FIG7 is a schematic diagram of generating table-level column statistics according to an embodiment of the present application;
[0021] FIG8 is a schematic diagram of a data structure corresponding to second statistical information according to an embodiment of the present application;
[0022] FIG9 is a flow chart of a data query method according to an embodiment of the present application;
[0023] FIG10 is a schematic diagram of determining and storing a column predicate range set according to an embodiment of the present application;
[0024] FIG11 is a flow chart of a data query method according to another embodiment of the present application;
[0025] FIG12 is a schematic diagram of data skipping according to a column statistics range set according to an embodiment of the present application;
[0026] FIG13 is a schematic structural diagram of an electronic device according to an embodiment of the present application;
[0027] FIG14 is a schematic structural diagram of a data processing device according to an embodiment of the present application;
[0028] FIG15 is a schematic structural diagram of an electronic device according to another embodiment of the present application;
[0029] FIG16 is a schematic structural diagram of a data query device according to an embodiment of the present application.
[0030] FIG17 is a schematic diagram of the structure of a data processing and query system according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] In an increasing number of scenarios, users expect to quickly and interactively analyze data. These scenarios range from traditional batch queries to exploratory, needle-in-a-haystack queries (such as point queries), and the size of data objects ranges from several GB to hundreds of PB. Linearly scanning these large datasets with extensive clusters is prohibitively expensive. In practical applications, multidimensional analysis often involves filtering, and the analysis results may only represent a small subset of the original record set. In theory, it is possible to skip all irrelevant data during data reading, reading only the minimal required data. Data skipping is a technique for performing SQL (Structured Query Language) analysis on structured data. Data skipping stores column summary statistics (such as minimum and maximum values) for each object (or file) as a column index. If the column index value range does not overlap with the query-specified data range, the query engine will skip the file during scanning, thereby reducing the amount of data scanned for a given query, speeding up queries, and lowering resource costs.
[0032] In the related art, database technologies represented by Oracle and data lake technologies represented by Iceberg, Hudi, etc., all use column minimums and column maximums to eliminate partitions on tables with predicates. In an exemplary embodiment, please refer to Figure 1, the minimum and maximum column values from the ZoneMap zone mapping of the partition are aggregated or summarized and associated with the partition, that is, the two zone maps shown in Figure 1 are associated with partitions F1 and F2 respectively, and the column maximum and column minimum values in the zone map associated with F1 are obtained by aggregating the column values in F1, and the column maximum and column minimum values in the zone map associated with F2 are obtained by aggregating the column values in F2. When performing a data query, if the query request carries a filter condition for a certain column, and the minimum and maximum values of the column have been aggregated for the partition, then the query engine can determine whether the entire partition can be omitted from the access path for processing the query based on the column value (or range of column values) in the filter condition and the aggregated minimum and maximum values. As shown in Figure 1, the client's query request contains the filter condition "salary > 500 and salary < 1000." In the region mapping associated with partition F1, the maximum value for the salary column is 20,000 and the minimum value is 1. In the region mapping associated with partition F2, the maximum value for the salary column is 25,000 and the minimum value is 140. These overlap with the filter conditions. Therefore, partitions F1 and F2 cannot be skipped, requiring a scan of both. Similarly, to achieve faster scan planning, Iceberg uses column-level minimum and maximum values to eliminate partitions and files that don't match the query predicate.
[0033] However, in real applications, the range between the minimum column value (MIN) and the maximum column value (MAX) can be very large. For example, in some scenarios, if data is not clustered or sorted on a column, the maximum and minimum column values may be close to or equal to the entire range of values in the column in the table. This coarse-grained data statistics makes it impossible to effectively filter partitions or files when skipping data based on the minimum and maximum column values, resulting in low query efficiency. For example, in the first region mapping shown in Figure 1, the range between the minimum column value 1 and the maximum column value 20,000 for salary is 20,000. In the second region mapping, the range between the minimum column value 140 and the maximum column value 25,000 for salary is 24,860. These data ranges are large, and when querying data, they overlap with the query conditions. Therefore, both partitions F1 and F2 must be scanned, meaning that partitions F1 and F2, or the files within them, cannot be skipped, resulting in low query efficiency.
[0034] The embodiments of the present application provide a data processing method, a data query method, and a data processing and query system, and specifically propose a data structure of a column statistical range set based on multiple discontinuous interval sets. The data structure is obtained by dividing the data range constituted by the column maximum value and the column minimum value into intervals. This can narrow the interval between the column maximum value and the column minimum value, and achieve a finer-grained division of the column statistical data, so that the data distribution characteristics of the data table can be described more accurately and in a smaller range. After obtaining the column statistical range set, since data queries can be performed based on the column statistical range set, a finer-grained judgment can be made on whether to skip data, thereby effectively filtering out invalid tables, partitions, and files, reducing the data range that needs to be scanned, and improving data query efficiency.
[0035] In order to help those skilled in the art better understand the technical solutions of this application, the following will clearly and completely describe the technical solutions of this application in conjunction with the drawings of one or more embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0036] The terms "first," "second," and the like in this application and the claims are used to distinguish similar objects and are not used to describe a particular order or precedence. It should be understood that such terms are interchangeable where appropriate so that this application can be implemented in sequences other than those illustrated or described herein. In addition, the term "and / or" in this application and the claims refers to at least one of the connected objects, and the character " / " generally indicates that the connected objects are in an "or" relationship.
[0037] FIG2 is a schematic diagram of a schematic system architecture provided in an embodiment of the present application.
[0038] The application environment of the system architecture shown in Figure 2 can be logically divided into a computing layer 21 and a storage layer 22. The computing layer 21 performs data management and access based on the storage layer 22, and the computing layer 21 and the storage layer 22 can be connected through a network. TableFormat can be optionally deployed in the computing layer 21. TableFormat is a table format that enables the data lake to have ACID (atomicity, consistency, isolation and durability) transactions and enhances data management capabilities. Open source systems such as Iceberg, Hudi, and DeltaLake are all open source products of TableFormat. In actual deployment, the deployment of the computing layer 21 can include but is not limited to the following scenarios (1) to (4):
[0039] (1) A separate big data runtime environment (not based on the TableFormat system), that is, only the big data execution engine is deployed in the computing layer 21, without deploying a data warehouse or a TableFormat system;
[0040] (2) A separate data warehouse operating environment (based on the TableFormat system), that is, only the data warehouse is deployed in the computing layer 21, without deploying the big data execution engine, but the TableFormat system is deployed;
[0041] (3) A separate data warehouse operating environment (not based on the TableFormat system), that is, only the data warehouse is deployed in the computing layer 21, without deploying the big data execution engine or the TableFormat system;
[0042] (4) Lake-warehouse integrated operating environment (big data execution engine and data warehouse based on TableFormat are deployed together), that is, the big data execution engine and data warehouse are deployed simultaneously in the computing layer 11, and the TableFormat system is deployed.
[0043] The operating environment described in any one of (1) to (4) above may be a physical machine environment or a cloud environment, etc., and is not specifically limited here.
[0044] The storage layer 22 provides data storage capabilities. The underlying storage format can be Parquet. The storage layer 22 can be cloud storage, local storage, or distributed storage, etc., which is not specifically limited here.
[0045] In the related art, when storing data, the storage layer 22 can perform statistics on each column of stored data to obtain the maximum and minimum values of the column. When performing data queries, the computing layer 21 can skip data based on the maximum and minimum values of the column to reduce the amount of data scanned by the query. However, the data range between the maximum and minimum values of the column counted in the related art is usually relatively large. When performing data queries, when skipping data based on the maximum and minimum values of the column, the data range between the maximum and minimum values of the column is likely to overlap with the data range specified in the query condition. This will result in a small amount of data being skipped during data queries, which in turn will require scanning more data during data scanning, reducing query efficiency. Based on the technical solution provided in the embodiment of the present application, the storage layer 22 can further divide the data range formed by the maximum and minimum values of the column into intervals based on the statistics of the maximum and minimum values of the column to obtain multiple discontinuous intervals, and then store the column statistical range set formed by the multiple discontinuous intervals. When the computing layer 21 performs data queries, it can perform data queries based on the column statistical range set stored by the storage layer 22. In this way, since the data range composed of the maximum value and the minimum value of the column can be divided into intervals to obtain multiple discontinuous intervals, the interval between the maximum value and the minimum value of the column can be narrowed, and a finer-grained division of the column statistical data can be achieved, so that the data distribution characteristics of the data table can be described more accurately and in a smaller range. When performing data queries, since data queries can be performed based on the column statistical range set composed of multiple discontinuous intervals, a more fine-grained judgment can be made on whether to skip data, thereby effectively filtering out invalid tables, partitions and files, reducing the data range that needs to be scanned, and improving data query efficiency.
[0046] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0047] Figure 3 is a flow chart of a data processing method according to an embodiment of the present application. The data processing method shown in Figure 3 can be executed by the storage layer 22 shown in Figure 2. The data processing method shown in Figure 3 is as follows.
[0048] S302: Obtain first statistical information, where the first statistical information includes column statistical information obtained by performing statistics on columns of a data table, and the column statistical information includes a maximum column value and a minimum column value.
[0049] When data is stored in a database in the form of a data table, first statistical information related to the data table can be obtained. The first statistical information can be column statistical information obtained by counting data in a column of the data table, and the column statistical information includes at least a maximum column value and a minimum column value.
[0050] Optionally, in some embodiments, when storing data in the form of a data table, the storage granularity of the data table may include at least one of a table, a partition, and a file. When obtaining the first statistical information, the obtained first statistical information may include at least one of table-level column statistics, partition-level column statistics, and file-level column statistics. Tables have the largest storage granularity, followed by partitions, and files have the smallest storage granularity. Table-level column statistics may be column statistics obtained by performing statistics on columns of a data table according to the table, partition-level column statistics may be column statistics obtained by performing statistics on columns of a data table according to the partition, and file-level column statistics may be column statistics obtained by performing statistics on columns of a data table according to the file.
[0051] It should be noted that, in other possible implementations, the storage granularity of the data table may also be other granularities, such as buckets, and accordingly, the first statistical information obtained may also include bucket-level column statistical information, etc. This embodiment of the application is only described by taking the storage granularity of table, partition, and / or file as an example, and the first statistical information including table-level column statistical information, partition-level column statistical information, and / or file-level column statistical information.
[0052] The first statistical information can be generated synchronously when the data is written, or it can be generated through an asynchronous mechanism, which is not specifically limited here. The specific implementation method of generating the first statistical information can refer to the specific implementation in the relevant technology, which will not be described in detail here. After generating the first statistical information, the first statistical information can be stored in a metadata table or metadata file. In this way, when obtaining the first statistical information, the first statistical information can be obtained from the metadata table or metadata file through the application programming interface (Application Programming Interface, API).
[0053] To help understand how to obtain the first statistical information, let's use Iceberg as an example. Iceberg is a universal table format (data organization format) that provides read / write and metadata management capabilities for big data engines. Its first statistical information is typically written to the metadata system as a system table or metadata, accessed by the query engine for query optimization. The Iceberg metadata format is shown in Figure 4.
[0054] As shown in Figure 4, Iceberg collects statistics on the column data in a data table at the table, partition, and file levels, generating table-level, partition-level, and file-level column statistics. It then manages these column statistics as metadata. These column statistics can be stored in a logical structure table containing the column name, minimum column value, and maximum column value.
[0055] As shown in Figure 4, table-level column statistics can be maintained in Iceberg's table metadata files, which track snapshots of the table schema, partition configuration, and content. A snapshot represents the state of a table at a specific point in time and is used to access the complete set of data files in the table. Partition-level column statistics can be stored in Iceberg's manifest lists, which store manifest metadata, including partition statistics and data file counts. These statistics are used to avoid reading and manipulating unnecessary manifests. File-level column statistics can be tracked by one or more manifest files, which contain partition data and column-level statistics for each data file in the table. Data files can be stored in the Parquet columnar file format. The data portion of a Parquet file is segmented into multiple row groups (RowGroups). Each row group stores multiple columns, and each column can be divided into multiple Pages. The final data portion is stored at the page level. The footer of the file stores the metadata of the Parquet file. Column statistics within the row group include the maximum and minimum values of the column.
[0056] Based on the metadata shown in Figure 4, when obtaining the first statistical information, if the first statistical information includes table-level column statistics, the first statistical information can be obtained from the table metadata file; if the first statistical information includes partition-level column statistics, the first statistical information can be obtained from the manifest list; if the first statistical information includes file-level column statistics, the first statistical information can be obtained from the manifest file. The first statistical information can be obtained from the table metadata file, manifest list, or manifest file through an API.
[0057] S304: Generate second statistical information for data query based on the first statistical information, the second statistical information includes a column statistical range set, the column statistical range set includes multiple discontinuous intervals, and the multiple discontinuous intervals are obtained by dividing the data range composed of the column maximum value and the column minimum value in the column statistical information into intervals.
[0058] After obtaining the first statistical information, second statistical information can be generated based on the first statistical information. The second statistical information can be used for data query.
[0059] When generating the second statistical information based on the first statistical information, the data range formed by the column maximum value and the column minimum value in the first statistical information can be divided into intervals to obtain multiple discontinuous intervals. The set formed by the multiple discontinuous intervals can be represented as a column statistical range set. After obtaining the column statistical range set, the column statistical range set can be determined as the second statistical information. Thus, the second statistical information can be generated based on the first statistical information.
[0060] Intervals can be used to define the range boundaries of continuous spans, and this continuous span is a comparable type that is used to retrieve a set of numbers / strings within a specific range. After partitioning the data range formed by the column maximum and minimum values, the resulting multiple discontinuous intervals are non-empty and non-intersecting intervals, which can include at least one of open, closed, half-open, unbounded, and unbounded intervals. That is, each interval can be open, closed, half-open, unbounded, or unbounded. The data type of each interval's interval value includes, but is not limited to, at least one of integer, floating-point, and string types. The set (union) of multiple discontinuous intervals can be represented as a column statistics range set, meaning that a column statistics range set can include multiple discontinuous intervals. Optionally, range sets can be described using Guava's RangeSet, and intervals can be described using Range. Guava is an open-source Java library that provides the Range type for defining the range boundaries of continuous spans. Furthermore, Guava also provides RangeSet for maintaining multiple range boundary sets.
[0061] It should be noted that the second statistical information can be generated in real time or asynchronously. The generation process of the second statistical information can overlap with the generation process of the first statistical information. If the first statistical information is generated in real time, the second statistical information can be generated in real time after the first statistical information is generated in real time. If the first statistical information is generated asynchronously, the second statistical information can be generated after the first statistical information is generated.
[0062] Optionally, in some embodiments, the first statistical information may include at least one of table-level column statistics, partition-level column statistics, and file-level column statistics. On this basis, when generating the second statistical information based on the first statistical information, the column statistics range set in the generated second statistical information may include at least one of the table-level column statistics range set, partition-level column statistics range set, and file-level column statistics range set. The file-level column statistics range set may be a column value range set of the underlying data file, which may be obtained by performing interval division on the data range composed of the maximum column value and the minimum column value included in the file-level column statistics. The partition-level column statistics range set may be a column value range set of the partition, which may be obtained by performing interval division on the data range composed of the maximum column value and the minimum column value included in the partition-level column statistics. The table-level column statistics range set may be a column value range set of the table, which may be obtained by performing interval division on the data range composed of the maximum column value and the minimum column value included in the table-level column statistics.
[0063] As an embodiment, generating the second statistical information according to the first statistical information may include the following steps:
[0064] In a case where the first statistical information includes file-level column statistical information, generating at least one of a file-level column statistical range set and a partition-level column statistical range set according to the file-level column statistical information, and the second statistical information includes at least one of the file-level column statistical range set and the partition-level column statistical range set;
[0065] In a case where the first statistical information includes partition-level column statistical information, a table-level column statistical range set is generated according to the partition-level column statistical information, and the second statistical information includes the table-level column statistical range set.
[0066] In an exemplary embodiment, when generating second statistical information based on first statistical information, if the first statistical information includes file-level column statistical information, the generated second statistical information may be a file-level column statistical range set, a partition-level column statistical range set, or a file-level column statistical range set and a partition-level column statistical range set. If the first statistical information includes partition-level column statistical information, the generated second statistical information may be a table-level column statistical range set. If the first statistical information includes file-level column statistical information and partition-level column statistical information, the generated second statistical information may be at least one of a table-level column statistical range set, a partition-level column statistical range set, and a file-level column statistical range set.
[0067] In some implementations, the file-level column statistics information may include the maximum column value and the minimum column value of each column in each row group of each file. Based on this, when generating a file-level column statistics range set based on the file-level column statistics information, the following steps may be included:
[0068] For each file, do the following:
[0069] Get the maximum and minimum values of each column in each row group in the file;
[0070] For each column in each row group, determine a closed interval based on the column maximum and column minimum values;
[0071] For each column, multiple closed intervals corresponding to the column in multiple row groups of the file are merged, and the multiple discontinuous intervals obtained by the merging are determined as the file-level column statistical range set corresponding to the column.
[0072] In an exemplary embodiment, the file-level column statistics may be column statistics for each underlying file. Each underlying file may include multiple row groups, each row group may include multiple columns (the columns in each row group have the same column name). Accordingly, the file-level column statistics may include column statistics for each row group in the multiple underlying files, and the column statistics for each row group may include the maximum and minimum values of each column in the row group. When generating a file-level column statistics range set based on file-level column statistics, for each file (i.e., the underlying file), the column maximum and column minimum values of each column in each row group in the file can be determined based on the file-level column statistics information. Then, for each column in each row group, a closed interval can be determined based on the column maximum and column minimum values of the column. The closed interval can be an interval formed by the column maximum and column minimum values, i.e., [column minimum, column maximum]. Finally, for each column, multiple closed intervals corresponding to the column in multiple row groups of the file can be merged to obtain multiple discontinuous intervals. The set formed by the multiple discontinuous intervals is the file-level column statistics range set corresponding to the column. Based on the same method, the file-level column statistics range set corresponding to each column in the file can be obtained.
[0073] For a file, after obtaining the file-level column statistics range set for each column in the file based on the above method, the file-level column statistics range set for each column of other files in the database can be obtained based on the same method, and finally the file-level column statistics range sets of multiple files can be obtained.
[0074] For easier understanding, please refer to Figure 5.
[0075] Figure 5 uses the Parquet format as an example. The file footer stores metadata and some statistical information for the Parquet file, and serves as the entry point for the entire file. In Figure 5, generating a file-level column statistics range set based on file-level column statistics can include the following steps:
[0076] Step 1: Read the underlying data file DataFile through the API and read the file's Footer information.
[0077] The file's footer information includes file-level column statistics information, which includes the maximum and minimum values of each column (only columns A and B are shown in FIG5 ) in each row group (RowGroup) of each underlying file.
[0078] Step 2: For each file, obtain the maximum and minimum values of all columns in each row group in the file one by one. For each column, perform the following processing:
[0079] Step 2.1: Combine the maximum and minimum values of the column in each row group into a closed interval, obtaining multiple closed intervals corresponding to multiple row groups, with one closed interval corresponding to each row group. For example, if the minimum value of column A in row group 1 is 1 and the maximum value is 5, the closed interval [1, 5] is obtained. If the minimum value of column A in row group 2 is 7 and the maximum value is 10, the closed interval [7, 10] is obtained. If the minimum value of column A in row group 3 is 8 and the maximum value is 15, the closed interval [8, 15] is obtained.
[0080] Step 2.2: Merge the multiple closed intervals corresponding to the column in multiple row groups to obtain the file-level column statistics range set corresponding to the column. During merging, all connected intervals are merged, and empty intervals are ignored. Ultimately, a set of discontinuous intervals is obtained, which is the file-level column statistics range set corresponding to the column. For example, after merging the three closed intervals [1, 5], [7, 10], and [8, 15] corresponding to column A, multiple discontinuous intervals [1, 5] and [7, 15] are obtained, and the corresponding column statistics range set is [1, 5] ∪ [7, 15].
[0081] Step 2.3: Repeat steps 2.1 and 2.2 for each column until all columns are processed.
[0082] After generating file-level column statistics for a file based on step 2 above, you can generate file-level column statistics range sets for other files in the same way, and ultimately obtain file-level column statistics range sets for multiple files.
[0083] FIG5 shows the file-level column statistics range set of n files in a partition, and the logical structure of each file-level column statistics range set is the same. Taking file 1 as an example, the column statistics range set of file 1 may include the column maximum value, column minimum value, and column statistics range set of each column in file 1 (FIG5 only shows two columns, column A and column B). Among them, the column minimum value of column A can be the minimum value of column A in all row groups of file 1, the column maximum value of column A can be the maximum value of column A in all row groups of file 1, and the column statistics range value of column A is a plurality of discontinuous intervals obtained by merging the closed intervals corresponding to the column maximum value and column minimum value of column A in each row group of file 1.
[0084] In some implementations, the file-level column statistics information may include the maximum and minimum values of each column of each file in each partition. Based on this, when generating a partition-level column statistics range set based on the file-level column statistics information, the following steps may be included:
[0085] For each partition, do the following:
[0086] Get the maximum and minimum values of each column in each file in the partition;
[0087] For each column in each file, determine a closed interval based on the maximum and minimum values of the column;
[0088] For each column, multiple closed intervals corresponding to the column in multiple files of the partition are merged, and the multiple discontinuous intervals obtained by the merger are determined as the partition-level column statistics range set corresponding to the column.
[0089] In an exemplary embodiment, the file-level column statistics may be column statistics for each underlying file in a partition. Each partition may include multiple underlying files, and each underlying file may include multiple columns (the columns of each file in each partition have the same column name). Accordingly, the file-level column statistics may include column statistics for each file in the multiple partitions, and the column statistics for each file may include the maximum and minimum column values for each column in the file. When generating a partition-level column statistics range set based on the file-level column statistics, the maximum and minimum column values for each column included in each file in the partition may be determined based on the file-level column statistics. Then, for each column in each file, a closed interval may be determined based on the maximum and minimum column values of the column. The closed interval may be an interval consisting of the maximum and minimum column values of the column, i.e., [minimum column value, maximum column value]. Finally, for each column, the multiple closed intervals corresponding to the column in multiple files in the partition may be merged to obtain multiple discontinuous intervals. The set consisting of the multiple discontinuous intervals is the partition-level column statistics range set corresponding to the column. Based on the same method, a partition-level column statistics range set corresponding to each column in the partition may be obtained.
[0090] For a partition, after obtaining the partition-level column statistics range set for each column in the partition based on the above method, the partition-level column statistics range set for each column in other partitions in the database can be obtained based on the same method. Ultimately, the partition-level column statistics range sets for multiple partitions can be obtained.
[0091] For easier understanding, please refer to Figure 6.
[0092] In FIG6 , when generating a partition-level column statistics range set based on file-level column statistics, the following steps may be included:
[0093] Step 1: Get file-level column statistics for all files under multiple partitions from the manifest file.
[0094] Each partition includes multiple files. The file-level column statistics of each file include the maximum and minimum values of each column in the file.
[0095] Step 2: For each partition, obtain the maximum and minimum values of all columns of each file in the partition one by one. For each column, perform the following processing:
[0096] Step 2.1: Combine the maximum and minimum values of the column in each file into a closed interval, obtaining multiple closed intervals corresponding to multiple files, with one closed interval corresponding to each file. For example, if the minimum value of column A in file 1 is 1 and the maximum value is 15, then the closed interval is [1, 15]. If the minimum value of column A in file 2 is 5 and the maximum value is 30, then the closed interval is [5, 30]. If the minimum value of column A in file 3 is 36 and the maximum value is 110, then the closed interval is [36, 110].
[0097] Step 2.2: Merge the multiple closed intervals corresponding to the column in multiple files to obtain the partition-level column statistics range set corresponding to the column. During the merging process, all connected intervals are merged, and empty intervals are ignored. Ultimately, a set of discontinuous intervals is obtained, which is the partition-level column statistics range set corresponding to the column. For example, the three closed intervals corresponding to column A above, [1, 15], [5, 30], and [36, 110], are merged into multiple discontinuous intervals [1, 30] and [36, 110], and the corresponding column statistics range set is [1, 30] ∪ [36, 110].
[0098] Step 2.3: Repeat steps 2.1 and 2.2 for each column until all columns are processed.
[0099] After generating partition-level column statistics for a partition based on step 2 above, you can generate partition-level column statistics range sets for other partitions in the same way, and ultimately obtain partition-level column statistics range sets for multiple partitions.
[0100] Figure 6 shows the partition-level column statistics range sets for two partitions, and the logical structure of each partition-level column statistics range set is the same. Taking partition 1 as an example, the column statistics range set for partition 1 can include the column maximum value, column minimum value, and column statistics range set for each column in partition 1 (Figure 6 only shows two columns, column A and column B). Among them, the column minimum value of column A can be the minimum value of column A in all files in partition 1, the column maximum value of column A can be the maximum value of column A in all files in partition 1, and the column statistics range value of column A is a plurality of discontinuous intervals obtained by merging the closed intervals corresponding to the column maximum value and column minimum value of column A in each file in partition 1.
[0101] In some implementations, partition-level column statistics may include the maximum and minimum values of each column in each partition. Based on this, when generating a table-level column statistics range set based on the partition-level column statistics, the following steps may be included:
[0102] For each table, do the following:
[0103] Get the maximum and minimum values of each column in each partition of the table;
[0104] For each column in each partition, a closed interval is determined based on the maximum and minimum values of the column;
[0105] For each column, multiple closed intervals corresponding to the column in multiple partitions of the table are merged, and the multiple discontinuous intervals obtained by the merging are determined as the table-level column statistical range set corresponding to the column.
[0106] In an exemplary embodiment, table-level column statistics may be column statistics for each partition in a table. Each table may include multiple partitions, each partition may include multiple columns (columns in each partition may have the same column name). Accordingly, the table-level column statistics may include column statistics for each partition in the multiple tables, and the column statistics for each partition may include the maximum and minimum column values for each column in the partition. When generating a table-level column statistics range set based on the partition-level column statistics, the maximum and minimum column values for each column in each partition of the table may be determined based on the partition-level column statistics. Then, for each column in each partition, a closed interval may be determined based on the maximum and minimum column values of the column. The closed interval may be an interval consisting of the maximum and minimum column values of the column, i.e., [minimum column value, maximum column value]. Finally, for each column, the multiple closed intervals corresponding to the column in multiple partitions of the table may be merged to obtain multiple discontinuous intervals. The set consisting of the multiple discontinuous intervals is the table-level column statistics range set corresponding to the column. Based on the same method, a table-level column statistics range set corresponding to each column in the table may be obtained.
[0107] For a table, after obtaining the table-level column statistics range set for each column in the table based on the above method, the table-level column statistics range set for each column of other tables in the database can be obtained based on the same method, and finally the table-level column statistics range sets for multiple tables can be obtained.
[0108] For easier understanding, please refer to Figure 7.
[0109] In FIG7 , when generating a table-level column statistics range set based on partition-level column statistics, the following steps may be included:
[0110] Step 1: Get partition-level column statistics for all partitions in multiple tables from the list.
[0111] Each table includes multiple partitions. The partition-level column statistics of each partition include the maximum and minimum values of each column in the partition.
[0112] Step 2: For each table, obtain the maximum and minimum values of all columns in each partition of the table one by one. For each column, perform the following processing:
[0113] Step 2.1: Combine the maximum and minimum values of the column in each partition into a closed interval, obtaining multiple closed intervals corresponding to the multiple partitions, with one closed interval corresponding to each partition. For example, if the minimum value of column A in partition 1 is 1 and the maximum value is 30, the closed interval [1, 30] is obtained. If the minimum value of column A in partition 2 is 40 and the maximum value is 100, the closed interval [40, 100] is obtained. If the minimum value of column A in partition 3 is 36 and the maximum value is 200, the closed interval [36, 200] is obtained.
[0114] Step 2.2: Merge the multiple closed intervals corresponding to the column in multiple partitions to obtain the table-level column statistics range set corresponding to the column. During the merging process, all connected intervals are merged, and empty intervals are ignored. Ultimately, a set of discontinuous intervals is obtained, which is the table-level column statistics range set corresponding to the column. For example, the three closed intervals corresponding to column A above, [1, 30], [40, 100], and [36, 200], are merged into multiple discontinuous intervals [1, 30] and [36, 200], and the corresponding column statistics range set is [1, 30] ∪ [36, 200].
[0115] Step 2.3: Repeat steps 2.1 and 2.2 for each column until all columns are processed.
[0116] After generating table-level column statistics for a table based on step 2 above, you can generate table-level column statistics range sets for other tables in the same way, and ultimately obtain table-level column statistics range sets for multiple tables.
[0117] FIG7 shows the table-level column statistics range sets of two tables, and the logical structure of each table-level column statistics range set is the same. Taking Table 1 as an example, the column statistics range set of Table 1 can include the column maximum value, column minimum value, and column statistics range set of each column in Table 1 (FIG7 only shows two columns, Column A and Column B). Among them, the column minimum value of Column A can be the minimum value of Column A in all partitions of Table 1, the column maximum value of Column A can be the maximum value of Column A in all partitions of Table 1, and the column statistics range value of Column A is a plurality of discontinuous intervals obtained by merging the closed intervals corresponding to the column maximum value and column minimum value of Column A in each partition of Table 1.
[0118] Optionally, in some embodiments, when generating a file-level column statistics range set based on file-level column statistics information, a partition-level column statistics range set can also be generated based on the file-level column statistics range set. That is, when generating a column statistics range set based on file-level column statistics information, the file-level column statistics range set can be progressively generated based on the file-level column statistics information first, and then the partition-level column statistics range set can be generated based on the file-level column statistics range set (in this case, there is no need to perform the step of generating the partition-level column statistics range set based on the file-level column statistics information). This bottom-up, layer-by-layer convergence method for generating column statistics range sets has strong universality, high efficiency in constructing column statistics range sets, low modification cost for the execution engine, and no impact on upper-layer applications. In addition, since the file-level column statistics range set can divide column statistics data at a finer granularity than file-level column statistics information, the partition-level column statistics range set generated based on the file-level column statistics range set can describe the data distribution characteristics of the data table more accurately and in a smaller range than the partition-level column statistics range generated based on file-level column statistics information. When performing data queries, a finer-grained judgment can be made on whether to skip data, thereby effectively filtering out invalid partitions, thereby reducing the data range that needs to be scanned and improving data query efficiency.
[0119] When generating a partition-level column statistics range set based on a file-level column statistics range set, the following steps may be included:
[0120] For each partition, do the following:
[0121] Get the file-level column statistics range set corresponding to each column included in each file in the partition;
[0122] For each column, multiple file-level column statistics range sets corresponding to the column in multiple files of the partition are merged, and multiple discontinuous intervals obtained by the merging are determined as the partition-level column statistics range set corresponding to the column.
[0123] In an exemplary embodiment, each partition may include multiple files, and each file may include multiple columns. Accordingly, the file-level column statistics range set may include the column statistics range sets of the multiple files, and the column statistics range set of each file may include the file-level column statistics range set corresponding to each column in the file. When generating a partition-level column statistics range set based on the file-level column statistics range set, for each partition, the column statistics range set corresponding to each column of each file in the partition may be determined based on the file-level column statistics range set of the partition. Then, for each column, the column statistics range sets corresponding to the column in each file may be merged. The range set consisting of the multiple discontinuous intervals obtained by merging is the partition-level column statistics range set corresponding to the column. Based on the same method, the partition-level column statistics range set corresponding to each column may be obtained.
[0124] For example, partition 1 includes 3 files, each of which includes column A. The file-level column statistics range set corresponding to column A in file 1 is [1, 30] ∪ [36, 100], the file-level column statistics range set corresponding to column A in file 2 is [5, 20] ∪ [50, 150], and the file-level column statistics range set corresponding to column A in file 3 is [-10, 15] ∪ [100, 300]. Then, when determining the partition-level column statistics range of column A, When the partition-level column statistics range set is set, the column statistics range sets corresponding to column A in file 1, file 2 and file 3 can be merged, that is, [1, 30] ∪ [36, 100], [5, 20] ∪ [50, 150] and [-10, 40] ∪ [100, 300] can be merged to obtain multiple discontinuous intervals [-10, 30] and [36, 300]. The partition-level column statistics range set of column A is [-10, 30] ∪ [36, 300].
[0125] Optionally, in some embodiments, when a partition-level column statistics range set is generated based on file-level column statistics information or a partition-level column statistics range set is generated based on a file-level column statistics range set, a table-level column statistics range set can also be generated based on the partition-level column statistics range set. That is, after generating the partition-level column statistics range set, a table-level column statistics range set can be further generated based on the partition-level column statistics range set (in this case, there is no need to perform the step of generating a table-level column statistics range set based on the partition-level column statistics information). This bottom-up, layer-by-layer, convergent method for generating column statistics range sets has strong universality, high efficiency in constructing column statistics range sets, low execution engine modification cost, and no impact on upper-layer applications. In addition, since the partition-level column statistics range set can divide column statistics data at a finer granularity than the partition-level column statistics information, the table-level column statistics range set generated based on the partition-level column statistics range set can describe the data distribution characteristics of the data table more accurately and in a smaller scope than the table-level column statistics range set generated based on the partition-level column statistics information. When performing data queries, a finer granularity judgment can be made on whether to skip data, thereby effectively filtering out invalid tables and improving data query efficiency.
[0126] When generating a table-level column statistics range set based on a partition-level column statistics range set, the following steps can be included:
[0127] For each table, do the following:
[0128] Get the partition-level column statistics range set corresponding to each column included in each partition in the table;
[0129] For each column, multiple partition-level column statistics range sets corresponding to the column in multiple partitions of the table are merged, and multiple discontinuous intervals obtained by the merging are determined as the table-level column statistics range set corresponding to the column.
[0130] In an exemplary embodiment, each table may include multiple partitions, and each partition may include multiple columns. Accordingly, the partition-level column statistics range set may include the column statistics range sets of each of the multiple partitions, and the column statistics range set of each partition may include the partition-level column statistics range set corresponding to each column in the partition. When generating a table-level column statistics range set based on the partition-level column statistics range set, for each table, the column statistics range set corresponding to each column in each partition of the table may be determined based on the partition-level column statistics range set of the table. Then, for each column, the column statistics range sets corresponding to the column in each partition may be merged. The range set consisting of the multiple discontinuous intervals obtained by merging is the table-level column statistics range set corresponding to the column. Based on the same method, the table-level column statistics range set corresponding to each column may be obtained.
[0131] For example, Table 1 includes two partitions, each of which includes column A. The partition-level column statistics range set corresponding to column A in partition 1 is [-10, 30]∪[36, 300], and the partition-level column statistics range set corresponding to column A in partition 2 is [-100, 0]∪[50, 500]∪[600, 1000]. Then, when determining the table-level column statistics range set for column A, the corresponding ranges of column A in partitions 1 and 2 can be used. The column statistics range set of column A is merged, that is, [-10, 30]∪[36, 300] and [-100, 0]∪[50, 500]∪[600, 1000] are merged to obtain multiple discontinuous intervals [-100, 30], [36, 500] and [600, 1000]. The table-level column statistics range set of column A is [-100, 30]∪[36, 500]∪[600, 1000].
[0132] The above details how to generate the second statistical information based on the first statistical information. The generation methods include generating a column statistical range set based on the column statistical information and generating a column statistical range set based on the column statistical range set. It should be noted that in order to improve the construction efficiency of the column statistical range set and reduce the cost of modifying the execution engine, the column statistical range set can be generated by a bottom-up layer-by-layer aggregation method, and optionally generating the column statistical range set based on the column statistical information. For example, when the first statistical information includes file-level column statistical information and partition-level column statistical information, a file-level column statistical range set is generated based on the file-level column statistical information, and then a partition-level column statistical range set is generated based on the file-level column statistical range set. Finally, a table-level column statistical range set is generated based on the partition-level column statistical range set. Optionally, a file-level column statistical range set and a partition-level column statistical range set are generated based on the file-level column statistical information, and a table-level column statistical range set is generated based on the partition-level column statistical information. It should also be noted that when generating the column statistical range set, the type of column statistical range set to be generated can be determined based on the first statistical information obtained, or it can be determined based on actual needs. No specific limitations are made here. For example, if the first statistical information includes partition-level column statistics, then when generating a column statistics range set, a table-level column statistics range set may be generated. If the first statistical information includes file-level column statistics, then when generating a column statistics range set, at least one of the file-level column statistics range set, partition-level column statistics range set, and table-level column statistics range set may be generated. For another example, in an actual query scenario, if it is necessary to filter partitions, then when generating a column statistics range set, only a partition-level column statistics range set may be generated. If it is necessary to filter partitions and files, then when generating a column statistics range set, both a file-level column statistics range set and a partition-level column statistics range set may be generated.
[0133] Optionally, in some embodiments, after generating the second statistical information based on the first statistical information, the second statistical information may be stored. The second statistical information may be used for data querying. In an exemplary embodiment, when performing a data query, the second statistical information may be used to determine whether a data table, partition, and / or file meets the query filtering conditions, thereby determining whether to skip the data. The specific implementation of data querying based on the second statistical information can be seen in the embodiment shown in FIG9 and will not be described in detail here.
[0134] When storing the second statistical information, in some implementations, the second statistical information may be stored as extended information of the first statistical information. For example, if the first statistical information includes file-level column statistics, partition-level column statistics, and table-level column statistics, and the second statistical information includes a file-level column statistics range set, a partition-level column statistics range set, and a table-level column statistics range set, the file-level column statistics range set may be stored as extended information of the file-level column statistics, the partition-level column statistics range set may be stored as extended information of the partition-level column statistics, and the table-level column statistics range set may be stored as extended information of the table-level column statistics.
[0135] By storing the second statistical information as extended information of the first statistical information, the impact on the existing system can be reduced, and storage and maintenance costs can be saved.
[0136] Taking Lceberg as an example, when the second statistical information is stored as extended information of the first statistical information, the storage method can be as shown in Figure 8. In Figure 8, the table-level column statistical information in the first statistical information is stored in the table metadata file, the partition-level column statistical information is stored in the list list, and the file-level column statistical information is stored in the list file. After the second statistical information is generated based on the first statistical information, when storing the second statistical information, the table-level column statistical range set in the second statistical information can be stored in the table metadata file as extended information of the table-level column statistical information (in Figure 8, the table-level column statistical information and the table-level column statistical range set are uniformly represented as table-level column statistical data), the partition-level column statistical range set is stored in the list list as extended information of the partition-level column statistical information (in Figure 8, the partition-level column statistical information and the partition-level column statistical range set are uniformly represented as partition-level column statistical data), and the file-level column statistical range set is stored in the list file as extended information of the file-level column statistical information (in Figure 8, the file-level column statistical information and the file-level column statistical range set are uniformly represented as file-level column statistical data).
[0137] When storing second statistical information, different types of second statistical information can have the same storage structure. For example, when storing a table-level column statistical range set, a column of data can be added to the table-level column statistical information. This added column of data becomes the table-level column statistical range set. Similarly, when storing partition-level column statistical range sets and file-level column statistical range sets, storage can also be based on the same method. As shown in Figure 8, when storing a column statistical range set, a new column can be added to the right of the column statistical information, and the column statistical range set can be stored in the new column.
[0138] In other embodiments, when storing the second statistical information, the second statistical information may also be stored independently of the first statistical information, that is, the second statistical information and the first statistical information may be stored separately. The storage location of the second statistical information may be the same as or different from the storage location of the first statistical information, and may be determined based on actual needs, and is not specifically limited here. For example, if the first statistical information is stored in a metadata table or a metadata file, then when storing the second statistical information, the second statistical information may also be stored in the metadata table or the metadata file, but it may be stored independently of the first statistical information. Alternatively, the second statistical information may be stored in a location other than the metadata table or the metadata file.
[0139] Optionally, in some embodiments, when the first statistical information changes, the second statistical information needs to be updated to ensure its accuracy. During the update, to improve update efficiency, the first statistical information that has changed can be determined, and then the corresponding second statistical information can be generated based on the changed first statistical information. The specific implementation of generating the second statistical information based on the changed first statistical information can be found in S304 above and will not be repeated here.
[0140] Based on the data processing method provided in the embodiment of the present application, since the data range formed by the column maximum value and the column minimum value in the first statistical information can be divided into intervals to obtain multiple discontinuous intervals, the interval between the column maximum value and the column minimum value can be narrowed, and a finer-grained division of the column statistical data can be achieved, so that the data distribution characteristics of the data table can be described more accurately and in a smaller range; after generating multiple discontinuous intervals, since the column statistical range set formed by the multiple discontinuous intervals can be stored and used for subsequent data queries, a finer-grained judgment can be made on whether to skip data, so that invalid tables, partitions and files can be effectively filtered out, thereby reducing the data range that needs to be scanned and improving data query efficiency.
[0141] Based on the data processing method provided in the embodiment of the present application, the embodiment of the present application also provides a data query method, as shown in Figure 9. Figure 9 is a schematic flow chart of a data query method according to one embodiment of the present application. The data query method shown in Figure 9 can be executed by the big data execution engine in the computing layer 21 shown in Figure 2. The data query method shown in Figure 9 is described as follows.
[0142] S902: Receive a query request.
[0143] When an external client or other system has a data query requirement, it can send a query request to the big data execution engine, which can receive the query request. The query request can include an SQL statement, and the SQL query statement can include query filter conditions.
[0144] S904: Obtain second statistical information of the data table.
[0145] The second statistical information of the data table can be generated in advance based on the first statistical information of the data table. The specific implementation method of generating the second statistical information based on the first statistical information can be referred to the embodiment shown in Figure 3, which is not specifically limited here. The second statistical information can be obtained when a query request is received. The second statistical information includes at least one item from the file-level column statistical range set, the partition-level column statistical range set, and the table-level column statistical range set. For details, please refer to the embodiment shown in Figure 3, which will not be described in detail here.
[0146] S906: Determine a valid object from the data table according to the query request and the second statistical information. The valid object includes at least one of a valid table, a valid partition, and a valid file.
[0147] After receiving a query request and obtaining the second statistical information of the data table, invalid objects can be filtered out from the data table based on the query request and the second statistical information to obtain valid objects. An object can be at least one of a table, a partition, and a file. A valid object can be an object that stores target data that matches the query criteria, and an invalid object can be an object that does not store target data that matches the query criteria. In an embodiment of the present application, a valid object can include at least one of a valid table, a valid partition, and a valid file, and an invalid object can include at least one of an invalid table, an invalid partition, and an invalid file.
[0148] When determining a valid object from a data table according to a query request and second statistical information, the following steps may be included:
[0149] Parse and analyze the query request to determine the logical plan corresponding to the query request, which includes the query conditions;
[0150] Convert the query condition into a query interval;
[0151] The second statistical information is queried according to the query interval, and valid objects are determined from the data table according to the query result.
[0152] The query request includes an SQL query statement. When parsing and analyzing the query request, a SQL parser can be used to perform lexical analysis on the SQL statement in the query request to obtain a lexical analysis general symbol stream, and then perform syntactic analysis on the symbol stream to construct a syntax tree. Afterwards, the SQL parser can be used to perform in-depth analysis on the unparsed relationships in the syntax tree and generate a resolved logical plan (Resolved Logical Plan). The specific implementation methods of using the SQL parser for lexical analysis and using the SQL parser for syntactic analysis can be found in the specific implementation of the related art and will not be described in detail here.
[0153] The logical plan includes query conditions. After obtaining the logical plan, the query conditions can be converted into query ranges. For example, if the SQL statement in the query request includes salary>500and salary<1000, when converted into a logical plan, it can be converted into Greater Than(salary,500)AND Lower Than(salary,1000), and further into the query range (500,1000).
[0154] After the query interval is obtained, the second statistical information may be queried according to the query interval, and valid objects may be determined from the data table according to the query result.
[0155] In some embodiments, the second statistical information may include at least one of a table-level column statistical range set, a partition-level column statistical range set, and a file-level column statistical range set. Thus, when querying the second statistical information based on the query interval and determining valid objects based on the query results, at least one of the following (1) to (3) may be included:
[0156] (1) When the second statistical information includes a table-level column statistical range set, the table-level column statistical range set is queried according to the query interval, and a first target range set matching the query interval is determined according to the query result; and a table corresponding to the first target range set is determined as a valid table.
[0157] When querying the table-level column statistics range set based on the query interval, the query interval can be compared with multiple discontinuous intervals in the table-level column statistics range set of each table to determine which discontinuous interval in the table-level column statistics range set matches the query interval, and the table-level column statistics range set corresponding to the matching discontinuous interval is determined as the first target range set, and the table corresponding to the first target range set is determined as the valid table.
[0158] For example, the query interval is (500, 1000), and the corresponding column field is salary. The table-level column statistics range set corresponding to the salary column in Table 1 is [1, 200] ∪ [1500, 2000], and the table-level column statistics range set corresponding to the salary column in Table 2 is [100, 600] ∪ [800, 3000]. After querying the table-level column statistics range sets corresponding to the salary column in Tables 1 and 2 based on the query interval, the query result is: the discontinuous intervals in the table-level column statistics range set corresponding to the salary column in Table 1 do not overlap with the query interval, and the two do not match. The discontinuous intervals in the table-level column statistics range set corresponding to the salary column in Table 2 overlap with the query interval, that is, (500, 600] ∪ [800, 1000], and the two match. Based on the query result, Table 1 can be determined as an invalid table, the table-level column statistics range set of Table 2 can be determined as the first target range set, and Table 2 can be determined as a valid table.
[0159] (2) When the second statistical information includes a partition-level column statistical range set, the partition-level column statistical range set or the partition-level column statistical range sets of multiple partitions in the valid table are queried according to the query interval to determine a second target range set that matches the query interval; and the partition corresponding to the second target range set is determined as a valid partition.
[0160] When querying the partition-level column statistical range set based on the query interval, two implementation methods can be included. One is to query the partition-level column statistical range set of all partitions based on the query interval. The other is to query the partition-level column statistical range set of each partition in the valid table based on the query interval based on a known valid table. The valid table can be the valid table determined according to the query interval and the table-level column statistical range set as mentioned above, or it can be a valid table determined by other methods, such as a valid table obtained by skipping data based on table-level column statistical information. No specific limitation is made here.
[0161] When querying the partition-level column statistics range set of all partitions or querying the partition-level column statistics range set of each partition in a valid table based on the query interval, the query interval can be compared with multiple discontinuous intervals in the partition-level column statistics range set of each partition to determine which discontinuous interval in the partition-level column statistics range set matches the query interval, and the partition-level column statistics range set corresponding to the matched discontinuous interval is determined as the second target range set, and the partition corresponding to the second target range set is determined as the valid partition.
[0162] Take the query interval for each partition in the valid table as an example. Assume that the query interval is (500,1000), the corresponding column field is salary, the valid table is Table 2, and the partition-level column statistics range set corresponding to the salary column of partition 1 in Table 2 is [100,200]∪[1500,3000], and the partition-level column statistics range set corresponding to the salary column of partition 2 is [400,600]∪[800,2000]. Then, according to the query interval, the partition-level column statistics range set corresponding to the salary column of partition 1 and partition 2 is After performing the query separately, the query results are as follows: the discontinuous interval in the partition-level column statistical range set corresponding to the column salary of partition 1 does not overlap with the query interval, and the two do not match. The discontinuous interval in the partition-level column statistical range set corresponding to the column salary of partition 2 overlaps with the query interval, that is, (500,600]∪[800,1000), and the two match. Based on the query results, partition 1 can be determined as an invalid partition, the partition-level column statistical range set of partition 2 can be determined as the second target range set, and partition 2 can be determined as a valid partition.
[0163] (3) When the second statistical information includes a file-level column statistical range set, the file-level column statistical range set or the file-level column statistical range set of multiple files in the valid partition is queried according to the query interval to determine a third target range set that matches the query interval; and the files corresponding to the third target range set are determined as valid files.
[0164] When querying the file-level column statistical range set based on the query interval, two implementation methods can be included. One is to query the file-level column statistical range set of all files based on the query interval, and the other is to query the file-level column statistical range set of each file in the valid partition based on the query interval on the basis of a known valid partition. The valid partition can be the valid partition determined according to the query interval and the partition-level column statistical range set as mentioned above, or it can be a valid partition determined by other methods, such as a valid partition obtained by skipping data according to partition-level column statistical information. No specific limitation is made here.
[0165] When querying the file-level column statistical range set of all files or the file-level column statistical range set of each file in a valid partition based on a query interval, the query interval can be compared with multiple discontinuous intervals in the file-level column statistical range set of each file to determine which discontinuous interval in the file-level column statistical range set matches the query interval, and the file-level column statistical range set corresponding to the matched discontinuous interval can be determined as a third target range set, and the file corresponding to the third target range set can be determined as a valid file.
[0166] Take the example of querying the file-level column statistics range set of each file in the valid partition according to the query interval. Assume that the query interval is (500,1000), the corresponding column field is salary, the valid partition is partition 2, and the file-level column statistics range set corresponding to the salary column of file 1 in partition 2 is [400,500]∪[1000,2000], and the file-level column statistics range set corresponding to the salary column of file 2 is [500,600]∪[800,1500]. Then, when querying the file-level column statistics range set corresponding to the salary column of file 1 and file 2 according to the query interval, After querying the surrounding sets separately, the query results are as follows: the discontinuous interval in the file-level column statistical range set corresponding to the column salary of file 1 does not overlap with the query interval, and the two do not match. The discontinuous interval in the file-level column statistical range set corresponding to the column salary of file 2 overlaps with the query interval, that is, (500,600]∪[800,1000), and the two match. Therefore, based on the query results, file 1 can be determined as an invalid file, the file-level column statistical range set of file 2 can be determined as the third target range set, and file 2 can be determined as a valid file.
[0167] The above details how to determine valid objects based on the second statistical information. It should be noted that, in actual applications, if the second statistical information includes any one of the table-level column statistical range set, the partition-level column statistical range set, and the file-level column statistical range set, then when determining valid objects based on the second statistical information, the valid table, valid partition, or valid file can be determined based on the method described in any one of (1) to (3) above. For example, if the second statistical information only includes the partition-level column statistical range set, then when determining valid objects based on the second statistical information, the valid partition can be determined based on the method described in item (2) above. If the second statistical information includes at least two of the table-level column statistical range set, the partition-level column statistical range set, and the file-level column statistical range set, then when determining valid objects based on the second statistical information, at least two of the valid tables, valid partitions, and valid files can be determined step by step from top to bottom to improve query efficiency and accuracy. For example, if the second statistical information includes a table-level column statistical range set, a partition-level column statistical range set, and a file-level column statistical range set, then when determining valid objects based on the second statistical information, you can first determine the valid table based on the table-level column statistical range set, then determine the valid partition from the valid table based on the partition-level column statistical range set, and finally determine the valid file from the valid partition based on the file-level column statistical range set.
[0168] Optionally, in some embodiments, when determining valid objects based on the second statistical information, the first statistical information can also be used to jointly determine valid objects. For example, when determining valid objects, the valid table can be first determined based on the table-level column statistical range set in the second statistical information, and then the valid partition can be determined from the valid table based on the partition-level column statistical range set in the second statistical information. Finally, the valid file can be determined from the valid partition based on the file-level column statistical information in the first statistical information. In actual applications, whether to determine the valid table, valid partition, or valid file in combination with the first statistical information can be determined based on actual conditions, and no specific limitation is made here. For example, if the second statistical information only includes the partition-level column statistical range set, then after determining the valid partition based on the partition-level column statistical range set, the valid table can be determined from the valid partition in combination with the file-level column statistical information in the first statistical information.
[0169] Optionally, in some embodiments, a column predicate range set can be determined in advance based on historical query information. After valid objects are determined based on the second statistical information, the valid objects can be further screened in combination with the column predicate range set to filter out more invalid objects, thereby improving query efficiency. The column predicate range set has the same structure as the column statistics range set, and both are sets consisting of multiple discontinuous intervals. The multiple discontinuous intervals can include at least one of an open interval, a closed interval, a half-open interval, an upper bound interval, and a lower bound interval. The data type of the interval value of each interval can include at least one of an integer, a floating point type, and a string type. The column predicate range set can include at least one of a table-level column predicate range set, a partition-level column predicate range set, and a file-level column predicate range set. The table-level column predicate range set can be used to screen valid tables, the partition-level column predicate range set can be used to screen valid partitions, and the file-level column predicate range set can be used to screen valid files.
[0170] When valid objects are filtered according to the column predicate range set, optionally, in some embodiments, at least one of the following (1) to (3) may be included:
[0171] (1) When a valid table is determined, determine whether a first query function corresponding to the valid table is in an available state; when the first query function is in an available state, query a table-level column predicate range set corresponding to the valid table according to the query interval, and determine a fourth target range set matching the query interval from the table-level column predicate range set; and determine the table corresponding to the fourth target range set as the valid table.
[0172] The first query function can be understood as a function of querying table-level column predicate range sets based on a query interval. The first query function can be in an enabled state and an unenabled state. In the enabled state, it can be indicated that querying table-level column predicate range sets based on a query interval is allowed to filter valid tables. In the unenabled state, it can be indicated that querying table-level column predicate range sets based on a query interval is not allowed.
[0173] After determining the valid table based on the query interval and the second statistical information, it can be further determined whether the first query function corresponding to the valid table is in an available state. If the first query function is available, further screening of the valid table using the table-level column predicate range set is permitted. In this case, the table-level column predicate range set corresponding to the valid table can be queried based on the query interval to screen the valid table. When querying the table-level column predicate range set corresponding to the valid table based on the query interval, the query interval can be compared with multiple discontinuous intervals in the table-level column predicate range set to determine which discontinuous interval in the table-level column predicate range set matches the query interval. The table-level column predicate range set corresponding to the matching discontinuous interval is then determined as a fourth target range set, and the table corresponding to the fourth target range set is determined as the valid table. This allows further screening of the valid table to filter out more invalid data. If the first query function is unavailable, further screening of the valid table using the table-level column predicate range set is not permitted. In this case, the step of querying the table-level column predicate range set corresponding to the valid table based on the query interval is not required.
[0174] The available state and unavailable state of the first query function can be specified by the user, or can also be determined based on the actual query situation, which is not specifically limited here. In some embodiments, when determining whether the first query function is in an available state, it can be determined whether the first query function meets a preset condition. If the preset condition is met, it can be determined that the first query function is in an available state. If the preset condition is not met, it can be determined that the first query function is in an unavailable state. The preset condition can include at least one of the following:
[0175] Receiving activation instruction information for the first query function;
[0176] The query efficiency is lower than the specified efficiency;
[0177] The data in the table has not changed within the specified time period, or the proportion of changed data is less than the preset proportion.
[0178] The activation instruction information for the first query function can be triggered by the user. In an exemplary embodiment, when the business needs to query the table-level column predicate range set based on the query interval, the user can send an activation instruction information for the first query function to the execution engine to instruct the execution engine to activate the first query function. After receiving the activation instruction information, the execution engine can activate the first query function, and the first query function is now in an available state. If the execution engine does not receive the activation instruction information, there is no need to activate (or deactivate) the first query function.
[0179] The query efficiency may be the query efficiency when the data query is performed this time. The specified efficiency may be set according to actual conditions and is not specifically limited here. When the query efficiency is lower than the specified efficiency, it may indicate that the number of valid tables determined according to the query interval and the second statistical information is large, resulting in less invalid data skipped in this query, and the query efficiency is low. It is necessary to enable the first query function and further screen the valid tables. When the query efficiency is higher than the specified efficiency, it may indicate that the number of valid tables determined according to the query interval and the second statistical information is small, and more invalid data is skipped during this query, and the query efficiency is high. It is not necessary to enable the first query function and further screen the valid tables. When the query is actually performed, the execution engine may determine whether the query efficiency is lower than the specified efficiency. If so, the first query function may be enabled, and the first query function is now in an available state. If not, there is no need to enable (or disable) the first query function, and the first query function is now in an unavailable state.
[0180] The table-level column predicate range set includes multiple discontinuous intervals, which are obtained by statistically analyzing the data in the table. If the data in the table changes, the table-level column predicate range set may no longer correspond to the data in the table. When filtering the valid table based on the table-level column predicate range set, valid data may be filtered out, resulting in incorrect query results. Therefore, when performing an actual query, it is possible to determine whether the data in the table has changed within a specified time period. If no changes have occurred or the proportion of changed data is less than a preset proportion (i.e., the changed data is within an acceptable range and has a minor impact on the query results), the first query function can be enabled and the first query function can be enabled. If the data has changed or the proportion of changed data is greater than or equal to the preset proportion (i.e., the changed data is not within an acceptable range and has a significant impact on the query results), the first query function can be disabled (or closed) and the first query function can be disabled. The specified time period and preset proportion can be set according to actual circumstances and are not specifically limited here.
[0181] When determining whether the first query function is in an available state, if at least one of the above preset conditions is met, it can be determined that the first query function is in an available state; if none of the above three preset conditions are met, it can be determined that the first query function is in an unavailable state.
[0182] (2) When a valid partition is determined, determine whether a second query function corresponding to the valid partition is in an available state; when the second query function is in an available state, query the partition-level column predicate range set corresponding to the valid partition according to the query interval, and determine a fifth target range set matching the query interval from the partition-level column predicate range set; and determine the partition corresponding to the fifth target range set as a valid partition.
[0183] The second query function can be understood as a function of querying the partition-level column predicate range set according to the query interval. The second query function includes an available state and an unavailable state. In the available state, it can be explained that the partition-level column predicate range set is allowed to be queried according to the query interval to filter the valid partitions. In the unavailable state, it can be explained that the partition-level column predicate range set is not allowed to be queried according to the query interval.
[0184] After determining valid partitions based on the query interval and the second statistical information, it is possible to further determine whether the second query function corresponding to the valid partition is available. If the second query function is available, this indicates that further filtering of valid partitions using the partition-level column predicate range set is permitted. In this case, the partition-level column predicate range set corresponding to the valid partition can be queried based on the query interval to filter the valid partition. When querying the partition-level column predicate range set corresponding to the valid partition based on the query interval, the query interval can be compared with multiple discontinuous intervals in the partition-level column predicate range set to determine which discontinuous interval in the partition-level column predicate range set matches the query interval. The partition-level column predicate range set corresponding to the matching discontinuous interval is then determined as a fifth target range set, and the partition corresponding to the fifth target range set is determined as the valid partition. This allows further filtering of the valid partition and eliminates more invalid data. If the second query function is unavailable, this indicates that further filtering of valid partitions using the partition-level column predicate range set is not permitted. In this case, the step of querying the partition-level column predicate range set corresponding to the valid partition based on the query interval is unnecessary.
[0185] The availability and unavailability of the second query function can be user-specified or determined based on the actual query situation, and are not specifically limited here. In some embodiments, when determining whether the second query function is available, the specific implementation method can be the same as the specific implementation method for determining whether the first query function is available, and will not be described in detail here.
[0186] (3) When a valid file is determined, determine whether the third query function corresponding to the valid file is in an available state; when the third query function is in an available state, query the file-level column predicate range set corresponding to the valid file according to the query interval, and determine a sixth target range set that matches the query interval from the file-level column predicate range set; determine the file corresponding to the sixth target range set as a valid file.
[0187] The third query function can be understood as a function of querying the file-level column predicate range set according to the query interval. The first query function includes an available state and an unavailable state. In the available state, it can be explained that the file-level column predicate range set is allowed to be queried according to the query interval to screen valid files. In the unavailable state, it can be explained that the file-level column predicate range set is not allowed to be queried according to the query interval.
[0188] When valid files are determined based on the query interval and the second statistical information, it is possible to further determine whether the third query function corresponding to the valid files is in an available state. If the third query function is in an available state, it can be indicated that the use of the file-level column predicate range set to further screen valid files is permitted. In this case, the file-level column predicate range set corresponding to the valid files can be queried based on the query interval to screen the valid files. When querying the file-level column predicate range set corresponding to the valid files based on the query interval, the query interval can be compared with multiple discontinuous intervals in the file-level column predicate range set to determine which discontinuous interval in the file-level column predicate range set matches the query interval, and the file-level column predicate range set corresponding to the matching discontinuous interval is determined as the sixth target range set. The files corresponding to the sixth target range set are then determined as valid files, thereby enabling further screening of valid files and filtering out more invalid data. If the third query function is in an unavailable state, it can be indicated that the use of the file-level column predicate range set to further screen valid files is not permitted. In this case, the step of querying the file-level column predicate range set corresponding to the valid files based on the query interval is unnecessary.
[0189] The availability and unavailability of the third query function can be user-specified or determined based on the actual query situation, and are not specifically limited here. In some embodiments, when determining whether the third query function is available, the specific implementation method can be the same as the specific implementation method for determining whether the first query function is available, and will not be described in detail here.
[0190] The aforementioned table-level column predicate range sets, partition-level column predicate range sets, and file-level column predicate range sets can be determined based on historical query information, and the determination methods for different column predicate range sets can be the same. The following uses the partition-level column predicate range set as an example to illustrate how to determine the partition-level column predicate range set based on historical query information.
[0191] Alternatively, in some implementations, the partition-level column predicate range set may be determined by:
[0192] Obtaining historical query information obtained when performing a historical query based on the second statistical information;
[0193] determining, based on the historical query information, whether a data selection rate of the partition is greater than a first threshold and whether the second query function is in an available state when performing a historical query, wherein the data selection rate of the partition is equal to a ratio between the number of partitions selected during the historical query and the total number of partitions;
[0194] When the data selectivity of the partition is greater than the first threshold and the second query function is in an available state, a partition-level column predicate range set is determined according to historical query conditions during historical queries.
[0195] The historical query information may be the query information obtained when multiple historical queries are performed based on the second statistical information in the last time or in the past period of time (which can be flexibly set according to actual conditions). The historical query information may include the historical query interval converted according to the historical query request when performing the historical query, the valid partitions determined according to the historical query interval and the second statistical information, the filtered invalid partitions, whether the column predicate range set is used for the query (that is, whether the query function corresponding to the column predicate range set is in an available state when performing the historical query), the query results, the query efficiency and other information. After performing a historical query and obtaining the target data, the information related to the historical query can be stored, so that the pre-stored historical query information can be obtained when performing a subsequent data query.
[0196] After obtaining the historical query information, the data selectivity of the partition when performing the historical query can be determined based on the historical query information. In an embodiment of the present application, given a data set D and a query Q, a row r in D is related to Q, Dr represents the set of related rows in D, and D is stored as a data block (file, row group, page and / or record, etc.), then the data selectivity can be expressed as the ratio of the relevant rows scanned, that is, Dr / D. In this way, when determining the data selectivity of the partition, the ratio between the number of partitions selected during the historical query and the total number of partitions can be determined as the data selectivity of the partition. The number of selected partitions can be the number of partitions scanned when performing the historical query, which can be specifically equal to the number of valid partitions determined based on the historical query interval and the second statistical information. The total number of partitions can be the total number of partitions included in the table to be queried. Generally speaking, the smaller the data selectivity of the partition, the more data skipped during the query, the less data scanned, and the higher the query efficiency. The larger the data selectivity of the partition, the less data skipped during the query, the more data scanned, and the lower the query efficiency.
[0197] After obtaining the data selectivity of the partition, it can be determined whether the data selectivity of the partition is greater than the first threshold and whether the second query function corresponding to the partition is in an available state when performing historical queries. If the data selectivity of the partition is greater than the first threshold and the second query function is in an available state, it can be explained that the data skipped by this historical query is small, the data scanned is large, and the query efficiency is low. In this case, in order to facilitate subsequent queries to skip more invalid data, reduce the scanned data range, and thus improve query efficiency, the partition-level column predicate range set can be determined based on the historical query conditions during historical queries. If the data selectivity of the partition is less than or equal to the first threshold, or the second query function is in an unavailable state, it can be explained that the data skipped by the query or the query efficiency is within an acceptable range. In this case, there is no need to determine the partition-level column predicate range set based on the historical query conditions during historical queries. Among them, the first threshold can be set according to actual conditions and is not specifically limited here.
[0198] When determining a partition-level column predicate range set based on historical query conditions during historical queries, optionally, in some implementations, the following steps may be included:
[0199] Convert the historical query conditions into historical query intervals, and determine the complementary interval corresponding to the historical query intervals;
[0200] For each partition that was selected during the historical query and is invalid, perform the following operations:
[0201] For columns in the partition that are related to historical query results, a closed interval is determined based on the maximum and minimum values of the column.
[0202] The closed interval and the complement interval are intersected, and the obtained multiple discontinuous intervals are determined as the partition-level column predicate range set corresponding to the columns related to the historical query results.
[0203] For ease of understanding, the following description uses the partition shown in Figure 10 as an example. The statistical range set shown in Figure 10 is the partition-level column statistical range set for partition 1, and the predicate range set shown in Figure 10 is the partition-level column predicate range set for partition 1. Both belong to the partition-level statistical information of partition 1.
[0204] In Figure 10, during a historical query, the SQL query statement is SELECT * FROM employees WHERE salary>500and salary<1000 , indicating a search for records where the salary value is greater than 500 but less than 1000. The corresponding query range is (500, 1000). During the data query, partition 1 is an invalid partition, meaning that all records in partition 1 have salary values outside the range (500, 1000). Therefore, the column predicate range set for the salary column in partition 1 can be as follows:
[0205] Step 1: Determine the complement interval (-∞, 500]∪[1000, +∞) of the historical query interval (500, 1000).
[0206] Step 2: Obtain a closed interval [1,3000] based on the maximum and minimum values of the salary column in partition 1.
[0207] Step 3: Intersect the complement interval (-∞, 500] ∪ [1000, +∞) and the closed interval [1, 3000] to obtain multiple discontinuous intervals [1, 500] and [1000, 3000]. The set [1, 500] ∪ [1000, 3000] formed by these two intervals is the partition-level column predicate range set corresponding to the salary column of partition 1.
[0208] According to the partition-level column predicate range set corresponding to the salary column, the skippable data interval corresponding to the salary column is expanded from the original (-∞, 1) ∪ (3000, +∞) to (-∞, 1) ∪ (500, 1000) ∪ (3000, +∞). The skippable data interval is expanded to 150% of the original value, which can more accurately filter out invalid intervals.
[0209] Optionally, in some implementations, if, when determining the partition-level column predicate range set for a column, the column already has a partition-level column predicate range set, the existing partition-level column predicate range set can be merged with the multiple discontinuous intervals determined based on the closed interval and the complement interval, and the merged result is determined as the partition-level column predicate range set corresponding to the column. For example, if the original partition-level column predicate range set for partition 1 in FIG10 is [-100, 100], [-100, 100] can be merged with [1, 500] and [1000, 3000], and the resulting merged range set [-100, 500] ∪ [1000, 3000] can be used as the partition-level column predicate range set corresponding to the column salary.
[0210] After obtaining a partition-level column predicate range set corresponding to a column in a partition based on the method described above, the partition-level column predicate range set corresponding to the column can be stored. When storing, the partition-level column predicate range set corresponding to the column can optionally be stored as extended information of the partition-level column statistics range set corresponding to the column. As shown in Figure 10, after obtaining the partition-level column predicate range set corresponding to the column salary, the partition-level column predicate range set can be stored as extended information of the partition-level column statistics range set for the column salary in partition 1. That is, a new column is added to the right of the partition-level column statistics range set for the column salary to store the partition-level column predicate range set for the column salary.
[0211] The above describes in detail how to determine the partition-level column predicate range set based on historical query information. Based on the same method, table-level column predicate range sets and file-level column predicate range sets can be generated (replacing the partition in the partition-level column predicate range set generation process with a table generates a table-level column predicate range set, and replacing the partition in the partition-level column predicate range set generation process with a file generates a file-level column predicate range set). The detailed description of how to generate the table-level column predicate range set and the file-level column predicate range set is omitted here. Optionally, after generating the table-level column predicate range set and the file-level column predicate range set, they can also be stored based on the storage method for the partition-level column predicate range set. This is also omitted here.
[0212] After generating table-level, partition-level, and file-level column predicate range sets, when performing data queries, after valid objects are determined based on the column statistics range set, the column predicate range set can be used to further filter valid objects, provided the query function corresponding to the column predicate range set is available. This complementary relationship between the column statistics range set and the column predicate range set allows for a more accurate portrayal of partition data distribution, filtering out invalid data and improving query efficiency.
[0213] S908: Query target data corresponding to the query request from valid objects.
[0214] After valid objects are determined based on the method described in S906 , the valid objects may be scanned, and target data corresponding to the query request may be searched from the valid objects.
[0215] When querying target data corresponding to a query request from a valid object, the following steps may be included:
[0216] Convert the logical plan into an executable physical plan;
[0217] Query the target data corresponding to the query request from the valid objects according to the physical plan.
[0218] The specific implementation of converting a logical plan into an executable physical plan can be found in related art and will not be detailed here. After obtaining the physical plan, the execution engine can execute it. Upon completion, the target data can be retrieved from valid objects. After the target data is retrieved, it can be returned to the sender of the query request.
[0219] Optionally, in some embodiments, after obtaining the target data through query, relevant query information can be determined based on the situation of the current data query, and at least one of the table-level column range set, partition-level column predicate range set, and file-level column predicate range can be updated based on the query information. The query information for the current data query may include the data selectivity for the table, the data selectivity for the partition, the data selectivity for the file, whether the table-level column range set, partition-level column predicate range set, and file-level column predicate range were used during the query, query efficiency, etc. The specific implementation method for updating the column predicate range set based on the query information can refer to the specific implementation method for determining the column predicate range set based on historical query information recorded in S906 above, which will not be described in detail here.
[0220] To facilitate understanding of how to perform data queries based on column statistics range sets and column predicate range sets, a more specific implementation is provided below (see Figure 11).
[0221] In the embodiment shown in FIG11 , when performing a data query, it is necessary to filter partitions and files to obtain valid partitions and valid files. Specifically, the following steps may be included:
[0222] Step 1: Use the SQL parser to parse the SQL query statement in the query request.
[0223] Step 2: Use SQL analyzer to analyze the parsing results.
[0224] Step 3: Generate and optimize the logical plan based on the analysis results to obtain the query range.
[0225] Step 4: Filter the partitions to obtain valid partitions.
[0226] When filtering partitions, you first obtain the query conditions, convert them into a query range, then query the partition-level column statistics range set based on the query range. Based on the query results, you filter out invalid partitions to obtain valid partitions. You can then determine whether the query function corresponding to the partition is available. If so, you can query the partition-level column predicate range set based on the query range to further filter valid partitions. If not, you do not need to query the partition-level column predicate range set based on the query range.
[0227] Step 5: Filter the files in the valid partition to obtain valid files.
[0228] The valid partition here can be the valid partition obtained by filtering based on the column statistics range set in the previous step, or the valid partition obtained by filtering based on the column statistics range set and column predicate range set in the previous step. When filtering files in the valid partition, you can first obtain the query conditions, convert the query conditions into a query interval, then query the file-level column statistics range set based on the query interval, and filter the invalid files in the valid partition based on the query results to obtain valid files. Afterwards, you can determine whether the query function corresponding to the file is in an available state. If so, you can query the file-level column predicate range set based on the query interval to further filter valid files. If not, you do not need to query the file-level column predicate range set based on the query interval.
[0229] Step 6: Generate an executable physical plan.
[0230] Step 7: Execute the physical plan and query the target data in the valid file.
[0231] Optionally, after obtaining the target data, the data selectivity for partitions and files may be determined based on the relevant information of this query, and the partition-level column predicate range set and / or the file-level column predicate range set may be updated.
[0232] Step 8: Return the execution result.
[0233] The specific implementation of the above steps 1 to 8 can refer to the specific implementation of the corresponding steps in S902 to S908 above, and will not be repeated here.
[0234] Data query based on the column statistics range set or column statistics range set and column predicate range set provided in the embodiment of the present application can filter out more invalid data, thereby improving query efficiency. For ease of understanding, taking data query based on column statistics range set as an example, please refer to Figure 12.
[0235] In Figure 12, there are 3 files under a partition, and the MIN~MAX of the 3 files are [1, 35], [100, 135], and [6000, 65535] respectively. In related technologies, such as area mapping, the statistical data [1, 65535] is used to describe the data distribution of the partition, while the embodiment of the present application uses a column statistical range set (abbreviated as range set in Figure 12) composed of three discontinuous intervals [1, 35], [100, 135], and [6000, 65535] to describe the data distribution of the partition. The interval range between the discontinuous intervals is the data skip area. There is no column value distribution in the data skip area, and partition pruning can be performed.
[0236] The column statistics range set generated by the embodiment of the present application is [1, 35] ∪ [100, 135] ∪ [6000, 65535]. When performing a data query, if the range of the query filter condition value is in the data skip zone (-∞, 1) ∪ (35, 100) ∪ (135, 6000) ∪ (65535, +∞), the partition or file can be directly skipped. The data skip zone of the related art is (-∞, 1) ∪ (65535, +∞). Obviously, the data skip zone of the present application has a range twice as large as that of the data skip zone of the related art, which can effectively reduce data file scanning during the query process and further improve query efficiency. For example, if the range of the query filter condition value is (500, 1000), then the data skip zone based on the present application can skip the partition, while the data skip zone based on the related art cannot skip the partition. It can be seen that the column statistics range set generated based on this application can skip more invalid data, thereby reducing the scanned data range and improving query efficiency.
[0237] Based on the technical solution provided in the embodiments of the present application, since the data range formed by the column maximum value and the column minimum value in the first statistical information can be divided into intervals to obtain multiple discontinuous intervals, the interval between the column maximum value and the column minimum value can be narrowed, and a finer-grained division of the column statistical data can be achieved, so that the data distribution characteristics of the data table can be described more accurately and in a smaller range; after generating multiple discontinuous intervals, when performing data query, since data query can be performed through the column statistical range set formed by the multiple discontinuous intervals, a finer-grained judgment can be made on whether to skip data, so that invalid tables, partitions and files can be effectively filtered out, thereby reducing the data range that needs to be scanned and improving data query efficiency.
[0238] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0239] FIG13 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Referring to FIG13 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage device. Of course, the electronic device may also include hardware required for other services.
[0240] The processor, network interface, and memory can be interconnected via an internal bus, such as an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. These buses can be classified as address buses, data buses, and control buses. For ease of illustration, FIG13 shows only one bidirectional arrow, but this does not imply that there is only one bus or only one type of bus.
[0241] The memory is used to store the program. In an exemplary embodiment, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0242] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a data processing device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0243] Obtaining first statistical information, where the first statistical information includes column statistical information obtained by performing statistics on columns of the data table, where the column statistical information includes a maximum value and a minimum value of the column;
[0244] Second statistical information for data query is generated based on the first statistical information. The second statistical information includes a column statistical range set. The column statistical range set includes multiple discontinuous intervals. The multiple discontinuous intervals are obtained by dividing the data range composed of the column maximum value and the column minimum value in the column statistical information into intervals.
[0245] The method performed by the data processing device disclosed in the embodiment shown in Figure 13 of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps and logic block diagrams disclosed in this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0246] The electronic device can also execute the method of FIG3 and realize the functions of the data processing device in the embodiment shown in FIG3 , which will not be described in detail in this application.
[0247] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0248] The present application also provides a computer-readable storage medium storing one or more programs. The one or more programs include instructions. When executed by a portable electronic device including multiple application programs, the instructions enable the portable electronic device to perform the method of the embodiment shown in FIG. 3 , and are specifically configured to perform the following operations:
[0249] Obtaining first statistical information, where the first statistical information includes column statistical information obtained by performing statistics on columns of the data table, where the column statistical information includes a maximum value and a minimum value of the column;
[0250] Second statistical information for data query is generated based on the first statistical information. The second statistical information includes a column statistical range set. The column statistical range set includes multiple discontinuous intervals. The multiple discontinuous intervals are obtained by dividing the data range composed of the column maximum value and the column minimum value in the column statistical information into intervals.
[0251] FIG14 is a schematic diagram of the structure of a data processing device 140 according to an embodiment of the present application. Referring to FIG14 , in a software implementation, the data processing device 140 may include: an acquisition module 141 and a generation module 142, wherein:
[0252] An acquisition module 141 acquires first statistical information, where the first statistical information includes column statistical information obtained by performing statistics on columns of a data table, where the column statistical information includes a maximum column value and a minimum column value.
[0253] Generation module 142 generates second statistical information for data query based on the first statistical information, the second statistical information includes a column statistical range set, the column statistical range set includes multiple discontinuous intervals, and the multiple discontinuous intervals are obtained by dividing the data range composed of the column maximum value and the column minimum value in the column statistical information into intervals.
[0254] The data processing device 140 provided in the present application can also execute the method of FIG3 and realize the functions of the data processing device 140 in the embodiment shown in FIG3 , which will not be described in detail in the present application.
[0255] FIG15 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Referring to FIG15 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage device. Of course, the electronic device may also include hardware required for other services.
[0256] The processor, network interface, and memory can be interconnected via an internal bus, such as an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. These buses can be classified as address buses, data buses, and control buses. For ease of illustration, FIG15 uses only one bidirectional arrow, but this does not imply that there is only one bus or only one type of bus.
[0257] The memory is used to store the program. In an exemplary embodiment, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0258] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a data query device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0259] receiving query requests;
[0260] Get the second statistical information of the data table;
[0261] determining a valid object from the data table according to the query request and the second statistical information, the valid object comprising at least one of a valid table, a valid partition, and a valid file;
[0262] Target data corresponding to the query request is searched from the valid objects.
[0263] The method performed by the data query device disclosed in the embodiment shown in Figure 15 of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in this application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in this application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0264] The electronic device can also execute the method of FIG9 and realize the functions of the data query device in the embodiment shown in FIG9 , which will not be described in detail in this application.
[0265] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0266] The present application also provides a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions. When executed by a portable electronic device including multiple application programs, the instructions enable the portable electronic device to perform the method of the embodiment shown in FIG. 9 , and are specifically configured to perform the following operations:
[0267] receiving query requests;
[0268] Get the second statistical information of the data table;
[0269] determining a valid object from the data table according to the query request and the second statistical information, the valid object comprising at least one of a valid table, a valid partition, and a valid file;
[0270] Target data corresponding to the query request is searched from the valid objects.
[0271] FIG16 is a schematic diagram of the structure of a data query device 160 according to an embodiment of the present application. Referring to FIG16 , in a software implementation, the data query device 160 may include: a receiving module 161, an acquisition module 162, a determination module 163, and a query module 164, wherein:
[0272] Receiving module 161, receiving a query request;
[0273] An acquisition module 162 acquires second statistical information of the data table;
[0274] A determination module 163 is configured to determine a valid object from the data table according to the query request and the second statistical information, where the valid object includes at least one of a valid table, a valid partition, and a valid file.
[0275] The query module 164 queries the target data corresponding to the query request from the valid objects.
[0276] The data query device 160 provided in the present application can also execute the method of FIG. 9 and realize the functions of the data query device 160 in the embodiment shown in FIG. 9 , which will not be described in detail in the present application.
[0277] FIG17 is a schematic diagram of the structure of a data processing and query system 170 according to an embodiment of the present application. Referring to FIG17 , the data processing and query system 170 may include a storage layer 171 and a computing layer 172, wherein:
[0278] The storage layer 171 obtains first statistical information, the first statistical information including column statistical information obtained by performing statistics on columns of a data table, the column statistical information including a maximum column value and a minimum column value; generates second statistical information for data query based on the first statistical information, the second statistical information including a column statistical range set, the column statistical range set including a plurality of discontinuous intervals, the plurality of discontinuous intervals being obtained by dividing a data range formed by the maximum column value and the minimum column value in the column statistical information into intervals;
[0279] The computing layer 172 receives a query request; obtains the second statistical information; determines a valid object from the data table based on the query request and the second statistical information, the valid object including at least one of a valid table, a valid partition, and a valid file; and queries the target data corresponding to the query request from the valid object.
[0280] In the embodiment of the present application, the specific implementation of each step performed by the storage layer 171 can refer to the specific implementation of the corresponding steps in the embodiment shown in FIG3 , and can achieve the same technical effects, so it will not be described in detail here. The specific implementation of each step performed by the computing layer 172 can refer to the specific implementation of the corresponding steps in the embodiment shown in FIG9 , and can achieve the same technical effects, so it will not be described in detail here.
[0281] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0282] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0283] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0284] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0285] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment.
Claims
1. A data processing method, comprising: Obtaining first statistical information, where the first statistical information includes column statistical information obtained by statistically analyzing columns of a data table, and the column statistical information includes a column maximum value and a column minimum value; Generating second statistical information for data query according to the first statistical information, where the second statistical information includes a set of column statistical ranges, and the set of column statistical ranges includes a plurality of discontinuous intervals, and the plurality of discontinuous intervals are obtained by partitioning the data range formed by the column maximum value and the column minimum value in the column statistical information.
2. The data processing method according to claim 1, where the storage granularity of the data table includes at least one of a table, a partition, and a file, and the first statistical information includes at least one of table-level column statistical information, partition-level column statistical information, and file-level column statistical information; Among them, The table-level column statistical information is the column statistical information obtained by statistically analyzing columns of the data table by table; The partition-level column statistical information is the column statistical information obtained by statistically analyzing columns of the data table by partition; The file-level column statistical information is the column statistical information obtained by statistically analyzing columns of the data table by file.
3. The data processing method according to claim 1 or 2, where the obtaining of the first statistical information includes: Obtaining the first statistical information from a metadata table or a metadata file through an application programming interface (API).
4. The data processing method according to claim 2, where the generating of the second statistical information for data query according to the first statistical information includes: When the first statistical information includes the file-level column statistical information, generating at least one of a file-level column statistical range set and a partition-level column statistical range set according to the file-level column statistical information, and the second statistical information includes at least one of the file-level column statistical range set and the partition-level column statistical range set; When the first statistical information includes the partition-level column statistical information, generating a table-level column statistical range set according to the partition-level column statistical information, and the second statistical information includes the table-level column statistical range set.
5. The data processing method according to claim 4, wherein the file-level column statistical information includes the column maximum value and the column minimum value of each column in each row group of each file; The generating of the file-level column statistical range set according to the file-level column statistical information includes: For each file, perform the following operations: Obtaining the column maximum value and the column minimum value of each column included in each row group in the file; For each column in each row group, determining a closed interval according to the column maximum value and the column minimum value; For each column, merging the multiple closed intervals corresponding to the column in multiple row groups of the file, and determining the multiple discontinuous intervals obtained by the merging as the file-level column statistical range set corresponding to the column.
6. The data processing method according to claim 4, wherein the file-level column statistical information includes the column maximum value and the column minimum value of each column of each file in each partition; The generating of the partition-level column statistical range set according to the file-level column statistical information includes: For each partition, perform the following operations: Obtaining the column maximum value and the column minimum value of each column included in each file in the partition; For each column in each file, determining a closed interval according to the column maximum value and the column minimum value; For each column, merge the multiple closed intervals corresponding to the column in multiple files of the partition, and determine the multiple discontinuous intervals obtained by the merge as the partition-level column statistical range set corresponding to the column.
7. The data processing method according to claim 4, wherein the partition-level column statistical information includes the column maximum value and the column minimum value of each column in each partition; The generating the table-level column statistical range set according to the partition-level column statistical information includes: For each table, perform the following operations: Obtain the column maximum value and column minimum value of each column included in each partition of the table; For each column in each partition, determine a closed interval according to the column maximum value and the column minimum value; For each column, merge the multiple closed intervals corresponding to the column in multiple partitions of the table, and determine the multiple discontinuous intervals obtained by the merge as the table-level column statistical range set corresponding to the column.
8. The data processing method according to any one of claims 4 to 7, the method further includes: In the case of generating the file-level column statistical range set according to the file-level column statistical information, generate the partition-level column statistical range set according to the file-level column statistical range set; In the case of generating the partition-level column statistical range set according to the file-level column statistical information or the file-level column statistical range set, generate the table-level column statistical range set according to the partition-level column statistical range set.
9. The data processing method according to claim 8, wherein the file-level column statistical range set includes column statistical range sets of multiple files, and the column statistical range set of each file includes the file-level column statistical range set corresponding to each column in the file; The generating the partition-level column statistical range set according to the file-level column statistical range set includes: For each partition, perform the following operations: Obtain the file-level column statistical range set corresponding to each column included in each file in the partition; For each column, merge the multiple file-level column statistical range sets corresponding to the column in multiple files of the partition, and determine the multiple discontinuous intervals obtained by the merge as the partition-level column statistical range set corresponding to the column.
10. The data processing method according to claim 8, wherein the partition-level column statistical range set includes column statistical range sets of multiple partitions, and the column statistical range set of each partition includes a partition-level column statistical range set corresponding to each column in the partition; The generating the table-level column statistical range set according to the partition-level column statistical range set includes: For each table, perform the following operations: Obtain the partition-level column statistical range set corresponding to each column included in each partition of the table; For each column, merge the multiple partition-level column statistical range sets corresponding to the column in multiple partitions of the table, and determine the multiple discontinuous intervals obtained by the merge as the table-level column statistical range set corresponding to the column.
11. The data processing method according to any one of claims 1, 5 to 7, 9 to 10, the multiple discontinuous intervals include at least one of an open interval, a closed interval, a semi-open interval, an interval without an upper bound, and an interval without a lower bound, and the data type of the interval value of each interval includes at least one of an integer type, a floating-point type, and a string type.
12. The data processing method according to claim 1, after generating the second statistical information for data query according to the first statistical information, the method further includes: Store the second statistical information as extended information of the first statistical information; Or, Store the second statistical information independently of the first statistical information.
13. A data query method, including: Receive a query request; Obtain the second statistical information of the data table, where the second statistical information is obtained based on the data processing method according to any one of claims 1 to 12; Determine valid objects from the data table according to the query request and the second statistical information, where the valid objects include at least one of valid tables, valid partitions, and valid files; Query target data corresponding to the query request from the valid objects.
14. The data query method according to claim 13, wherein the determining valid objects from the data table according to the query request and the second statistical information includes: Perform statement parsing and analysis on the query request to determine a logical plan corresponding to the query request, where the logical plan includes query conditions; Convert the query conditions into a query range; Query the second statistical information according to the query range, and determine valid objects from the data table according to the query result.
15. The data query method according to claim 14, wherein the second statistical information includes at least one of a table-level column statistical range set, a partition-level column statistical range set, and a file-level column statistical range set; the querying the second statistical information according to the query range and determining valid objects from the data table according to the query result includes at least one of the following: Query the table-level column statistical range set according to the query range to determine a first target range set matching the query range; determine the table corresponding to the first target range set as a valid table; Query the partition-level column statistical range set or the partition-level column statistical range sets of multiple partitions in the valid table according to the query range to determine a second target range set matching the query range; Determine the partition corresponding to the second target range set as a valid partition; Query the file-level column statistical range set or the file-level column statistical range sets of multiple files in the valid partition according to the query range to determine a third target range set matching the query range; determine the file corresponding to the third target range set as a valid file.
16. The data query method according to claim 15 further includes at least one of the following: When the valid table is determined, it is judged whether the first query function corresponding to the valid table is in an available state; When the first query function is in an available state, query the table-level column predicate range set corresponding to the valid table according to the query range, and determine a fourth target range set matching the query range from the table-level column predicate range set; Determine the table corresponding to the fourth target range set as a valid table; When determining a valid partition, determine whether the second query function corresponding to the valid partition is in an available state; When the second query function is in an available state, query the partition-level column predicate range set corresponding to the valid partition according to the query range, and determine a fifth target range set matching the query range from the partition-level column predicate range set; Determine the partition corresponding to the fifth target range set as a valid partition; When determining a valid file, determine whether the third query function corresponding to the valid file is in an available state; When the third query function is in an available state, query the file-level column predicate range set corresponding to the valid files according to the query range, and determine a sixth target range set that matches the query range from the file-level column predicate range set; determine the files corresponding to the sixth target range set as valid files.
17. The data query method according to claim 16, wherein the first query function is in an available state when a preset condition is satisfied; the preset condition includes at least one of the following: Receiving an opening indication message for the first query function; The query efficiency is lower than a specified efficiency; The data in the table has not changed within a specified time period or the proportion of the changed data is less than a preset proportion.
18. The data query method according to claim 16, wherein the partition-level column predicate range set is determined by the following method: Obtain historical query information obtained when performing a historical query according to the second statistical information; According to the historical query information, determine whether the data selection rate of the partition is greater than a first threshold and whether the second query function is in an available state when performing a historical query, where the data selection rate of the partition is equal to the ratio between the number of selected partitions and the total number of partitions during the historical query; When the data selection rate of the partition is greater than the first threshold and the second query function is in an available state, determine the partition-level column predicate range set according to the historical query conditions during the historical query.
19. The data query method according to claim 18, wherein determining the partition-level column predicate range set according to the historical query conditions during the historical query includes: Convert the historical query conditions into a historical query range, and determine a complement range corresponding to the historical query range; For each partition that was selected and invalid during the historical query, perform the following operations: For the columns related to the historical query result in the partition, determine a closed range according to the column maximum value and the column minimum value; Take the intersection of the closed range and the complement range, and determine the obtained multiple discontinuous ranges as the partition-level column predicate range set corresponding to the column.
20. The data query method according to claim 19, after obtaining the partition-level column predicate range set corresponding to the column, the data query method further includes: Store the partition-level column predicate range set corresponding to the column as extended information of the partition-level column statistical range set corresponding to the column.
21. The data query method according to any one of claims 14 to 16, wherein querying the target data corresponding to the query request from the valid objects includes: Convert the logical plan into an executable physical plan; Query the target data corresponding to the query request from the valid objects according to the physical plan.
22. A data processing and query system, including a storage layer and a computing layer, wherein: The storage layer obtains first statistical information, where the first statistical information includes column statistical information obtained by statistically analyzing the columns of a data table, and the column statistical information includes a column maximum value and a column minimum value; Generate second statistical information for data query according to the first statistical information, where the second statistical information includes a column statistical range set, and the column statistical range set includes a plurality of discontinuous intervals obtained by partitioning the data range formed by the column maximum value and the column minimum value in the column statistical information; The computing layer receives a query request; Obtain the second statistical information; determine valid objects from the data table according to the query request and the second statistical information, where the valid objects include at least one of a valid table, a valid partition, and a valid file; Query target data corresponding to the query request from the valid objects.
23. An electronic device, comprising: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the method according to any one of claims 1 to 21.
24. A computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 21.
Citation Information
Patent Citations
Indexing method of distributed column storage system
CN106250523A
Data statistics method, device and equipment and storage medium
CN109325031A
HDFS-oriented split layer indexing method and device
CN110019084A
Extended synopsis pruning in database management systems
US20230385282A1
Cited By
Data query method and device, electronic equipment and storage medium
CN120929480A