Extensible Data Skip
By introducing scalable APIs and user-defined functions in the data skipping system, the application restriction of the prior art for semi-structured or unstructured data is solved, and flexible and efficient skipping of various types of data is achieved, improving the scalability and performance of the system.
Patent Information
- Application Number
- CN202080023334.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-31
- Filing Date
- 2020-03-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-03-20
AI Technical Summary
Existing data skipping techniques are limited by structured data, cannot be effectively applied to semi-structured or unstructured data, and lack user-defined and scalable capabilities.
Provides an extensible data skipping system and method that allows developers to add new types of profile metadata by defining appropriate application programming interfaces (APIs) and combines arbitrary predicates for data skipping. The system is able to process structured, semi-structured or unstructured data and supports user-defined functions and complex SQL expressions.
It realizes flexible and efficient data skipping for structured, semi-structured or unstructured data, enhances the scalability of the system and user customization capabilities, and improves the performance and cost-effectiveness of data analysis.
Smart Images

Figure CN113632072B_ABST
Abstract
Description
Background Art
[0001] The present invention relates to providing data skipping techniques that can be user-defined and extensible.
[0002] Data skipping is a technique for big data analysis of structured data. For example, tabular data (including multiple rows) may involve storing summary metadata for a subset of rows. This summary metadata can be used to determine that a subset of rows is irrelevant to a query (such as in SQL) and can thus be skipped during query processing. This results in significant performance improvements and cost reduction.
[0003] Conventionally, data skipping is applied to structured data and is typically stored per column and works when, for example, the column is a number such as an integer or a floating point and the query predicate is one of {<, ≤, >, ≥, =}. In this case, the summary metadata used is the minimum and maximum values of the column. Another example is when the column is a number or a string and the query predicate is one of {=, IN}. In this case, the summary metadata used is a list of values or a Bloom filter.
[0004] Conventional data skipping can be useful, but is limited to a relatively limited set of uses. There is a need for techniques that can provide data skipping in additional environments, which can be user-defined and extensible, and applicable to structured, semi-structured, or unstructured data. Summary of the Invention
[0005] Embodiments of the present system and method can provide techniques for data skipping that can be user-defined and extensible. For example, extensible data skipping can define appropriate application programming interfaces (APIs) that enable developers to add new types of summary metadata that can be combined with arbitrary predicates for data skipping. Such predicates can be built into SQL. For example, {LIKE, <=}, etc., or they can be user-defined functions (UDFs). Extensible data skipping can provide queries involving UDFs or built-in predicates other than those listed above {<, <=, >, >=, =, IN}.
[0006] In addition, the extensible data skipping API can allow specifying how combinations of predicates (e.g., using AND / OR / NOT) should be mapped to combinations of summary metadata types. This can provide data skipping using arbitrarily complex predicates.
[0007] For example, in an embodiment, the method may include receiving a query at a computer system, the computer system including a processor, a memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor, modifying the received query at the computer system to evaluate at least one criterion of the query by using at least one data skip index, wherein the at least one data skip index includes an index on at least one attribute of data, the at least one attribute of the data excluding at least a portion of those data items that do not satisfy the at least one criterion, and wherein at least one of the at least one data skip index and / or a mapping from the at least one criterion to the at least one data skip index is generated based on information received from an application programming interface, and evaluating the query at the computer system.
[0008] In an embodiment, at the computer system, at least one data skip index may be generated by receiving information defining a data skip index type from an application programming interface, receiving information at the computer system that interprets the at least one criterion, generating metadata at the computer system related to the defined data skip index type and the defined at least one criterion, and generating the at least one data skip index at the computer system based on the generated metadata. The received query may be represented as an expression tree, and the expression tree is modified by labeling at least one node of the expression tree with an optimization rule to use a skip requirement representing the at least one criterion and a clause referencing at least one data skip index. The at least one criterion may be a Structured Query Language (SQL) predicate. At least one data skip index may be generated by receiving from an application programming interface, receiving at the computer system information defining a plurality of data skip index types, receiving at the computer system information defining a plurality of criteria, combining at the computer system metadata related to each defined data skip index type and each defined criterion to form metadata related to the plurality of defined data skip index types and the plurality of defined criteria, and generating the at least one data skip index at the computer system based on the generated metadata. The received query may be represented as an expression tree, and the expression tree is modified by labeling each of a plurality of nodes of the expression tree with a plurality of optimization rules to use a skip requirement representing a criterion and a clause referencing a data skip index. Each criterion may be a Structured Query Language (SQL) predicate.
[0009] In an embodiment, the system may include a processor, a memory accessible by the processor, and instructions stored in the memory and executable by the processor to perform: receiving a query, modifying the received query to evaluate at least one criterion of the query using at least one data skip index, wherein the at least one data skip index includes an index on at least one attribute of data, the at least one attribute of the data excluding at least a portion of those data items that do not satisfy the at least one criterion, and wherein at least one of the at least one data skip index or a mapping from the at least one criterion to the at least one data skip index is generated based on information received from an application programming interface, and evaluating the query.
[0010] In one embodiment, a computer program product may include a non-transitory computer-readable storage device having program instructions embodied therein, the program instructions executable by a computer to cause the computer to perform a method that includes: receiving a query, modifying the received query to evaluate at least one criterion of the query using at least one data skip index, wherein the at least one data skip index includes an index on at least one attribute of data, the at least one attribute of the data excluding at least a portion of those data items that do not satisfy the at least one criterion, and wherein at least one of the at least one data skip index or a mapping from the at least one criterion to the at least one data skip index is generated based on information received from an application programming interface, and evaluating the query. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The details of the present invention (both as to its structure and operation) may best be understood by reference to the accompanying drawings, in which like reference numerals and designations refer to like elements.
[0012] Figure 1 An exemplary system is shown in which embodiments of the present system and method may be implemented.
[0013] Figure 2 is an exemplary diagram of a scalable data skip interface according to an embodiment of the present system and method.
[0014] Figure 3 is an exemplary diagram of a predicate specification interface according to an embodiment of the present system and method.
[0015] Figure 4 is an exemplary flowchart of an operational process according to an embodiment of the present system and method.
[0016] Figure 5 An example of the definition of a query expression according to an embodiment of the present system and method is shown.
[0017] Figure 6Shows an example of the modification of a query expression according to an embodiment of the present system and method.
[0018] Figure 7 Shows an example of the modification of a query expression according to an embodiment of the present system and method.
[0019] Figure 8 Shows an example of a clause that can be generated according to an embodiment of the present system and method.
[0020] Figure 9 Shows an example of the translation of a clause for forming a metadata storage representation according to an embodiment of the present system and method.
[0021] Figure 10 Shows an instance of a MinMax filter according to an embodiment of the present system and method.
[0022] Figure 11 Shows an instance of a ValueList filter according to an embodiment of the present system and method.
[0023] Figure 12 Shows an example of a GeoBox filter according to an embodiment of the present system and method.
[0024] Figure 13 Is an exemplary block diagram of a computer system in which the processes involved in the embodiments described herein can be implemented.
[0025] Figure 14 Is an exemplary flowchart of a process for index creation based on an existing data stream according to an embodiment of the present system and method.
[0026] Figure 15 Is an exemplary flowchart of a process for index creation based on an ingested data stream according to an embodiment of the present system and method.
[0027] Figure 16 Is an exemplary flowchart of a process for a query processing flow according to an embodiment of the present system and method. Detailed Description
[0028] Embodiments of the present system and method can provide techniques for data skipping, which can be user-defined and extensible.
[0029] Embodiments of the present system and method can provide techniques for data skipping, which can be user-defined and extensible. For example, extensible data skipping can define appropriate application programming interfaces (APIs) that enable developers to add new types of summary metadata that can be combined with arbitrary predicates for data skipping. Such predicates can be built into SQL. For example, LIKE, <=, etc., or they can be user-defined functions (UDFs). Extensible data skipping can provide queries involving UDFs or built-in predicates other than those listed above (<, <=, >, >=, =, IN).
[0030] This technique can provide extensible data skipping for any SQL predicate. Examples can include supporting AND / OR / NOT predicates, supporting user-defined functions (UDFs), supporting additional predicates (such as LIKE), and supporting complex SQL expressions. This technique can define a general mechanism for specifying data skipping optimization rules, which can provide the ability to easily add optimization rules and can be extended by users to, for example, handle new UDFs. This technique can provide an abstraction level between the metadata store and the skip framework, which provides the ability to be agnostic to the metadata store and can allow for easy integration of new metadata stores.
[0031] Figure 1 Exemplary system 100 in which embodiments of the present system and method can be implemented is shown. In this example, system 100 can include a data processing system 102, a data store 104, an extensible data skip framework 106, an application 108, other systems 110, a user 112, a data ingestion system 114, and a metadata store 116. The data processing system 102 can include hardware and software for processing structured big data. The data processing system can be or can include one or more database management systems (DBMSs) and additional hardware and software for processing structured big data. The DBMS is only one example of a system that can be included. Other types of systems such as Hadoop, Spark, Presto, etc. are additional examples of system types that can be used by and include the data processing system 102. Embodiments of the present system and method can be advantageously applied to directly execute database or other queries on big data sets that can reside in an object store or a file system (not in a database system). The data skip index can reduce the amount of data that needs to be scanned in such cases and can turn infeasible queries into feasible queries. The data processing system 102 can include query / index processing 118, which can use the data skip index to process databases or other processing queries when processing queries. The data store 104 can include hardware and software for storing data created, retrieved, updated, managed, or otherwise processed by the data processing system 102.
[0032] The extensible data skip framework 106 can provide developers with the following capabilities: adding support for new types of summary metadata to support new data skip index types, and specifying how predicates used in queries can be mapped to summary metadata to evaluate whether data can be skipped. The extensible data skip framework 106 can include registration 120, expression tree processing 122, abstract clause generation 124, transformation 126, and index generation and maintenance 128. Registration 120 can provide the ability to extend the system by adding support for new UDFs and metadata index types in the extensible data skip framework 106 using the extensible data skip API. Expression tree processing 122 can accept query expressions in, for example, expression tree form but also in SQL form, and perform metadata processing of the query expressions such that abstract clause generation 124 can generate an abstract representation of clauses that reference relevant data skip indexes. Transformation 126 can transform the abstract clauses to form a metadata representation of clauses that reference relevant data skip indexes. The metadata representation can be stored in the metadata memory 116. Index generation and maintenance 128 can generate and maintain data skip indexes.
[0033] Predicate specifications can ultimately output filters that can use data skipping to provide filtering of data. For example, predicate specifications can allow specifying how combinations of predicates (such as using AND / OR / NOT) should be mapped to combinations of summary metadata types. This can provide data skipping using arbitrarily complex predicates. For example, the extensible data skip framework 106 can provide the following capabilities: adding support for new types of summary metadata to support new data skip index types, and specifying how predicates used in queries are mapped to summary metadata to evaluate whether data can be skipped.
[0034] The extensible data skip framework 106 can be implemented to be accessible to the application 108, other systems 110, and the user 112. The extensible data skip framework 106 can be used to generate data skip indexes using index generation and maintenance 128. Data skip indexes can be in the form of metadata when generated, and the metadata can be stored in the data memory 104 together with the data 122 itself or separately in the data memory 104. Compared to a conventional database table index that includes index information for each data item in a data table, a data skip index includes summary index information for groups of data items and can be used to determine that these data items do not meet certain criteria.
[0035] The data ingestion system 114 can receive data to be processed from the application 108, other systems 110, and the user 112, process the received data, and store it in the data memory 104.
[0036] The metadata store 116 can store metadata that defines data types and data skip indexes and can support pluggable configurations for different data types and indexes. For example, an Elasticsearch (ES) search engine can be used to implement the metadata store 116, but support for additional metadata stores can be added in an extensible manner.
[0037] To generate an effective data skip index, the summary metadata should be significantly smaller than the data, and the mapping should have no false negatives (whenever data is skipped, there should definitely be no relevant rows). However, false positives are acceptable. For example, for skipped geospatial data, the developer can add bounding box summary metadata that summarizes a collection of geospatial points in a subset of the rows. Note that the bounding box metadata is much smaller than the data itself. In this instance, a mapping of a UDF (e.g., ST_distance (geospatial distance)) to the bounding box summary metadata can be defined.
[0038] For example, consider a query to find data points within 10 kilometers of LaGuardia Airport. (This query can use the ST_distance UDF). This specifies a circle with a radius of 10 km around the point. The circle can be square, and bounding box summary metadata that does not overlap with the square can be generated. The corresponding subset of rows can be safely skipped (note that there are no false negatives here).
[0039] By using these APIs, support can be added for application-specific UDFs and / or data types without any changes to the analysis engine or data skip library. For example, new data skip indexes can be added for 1) geospatial data, 2) astronomical data, 3) genomic data, 4) image data, etc. Conventionally, data skipping is not supported for queries involving UDFs or for application-specific data types. In an embodiment, the present system and method can be applied to structured, semi-structured, or unstructured data. Data skipping can be extended to any type of data, whether organized in a database or other data collection. In one embodiment, the data can be partitioned into objects with associated metadata.
[0040] For example, embodiments of the present system and method can be applied to image data, such as a photo collection, where the user-defined metadata is the thumbnail of each photo. In such an embodiment, the query can provide the ability to skip photos where processing the thumbnail can ensure that the original photo does not have the property being sought. For example, the query can search for photos of cats, so when processing the thumbnail can determine that the original image is not a cat, etc., the embodiment can provide the ability to skip other animals / objects.
[0041] Conventional systems that allow the addition of optimization rules in an extensible manner do not provide support for data skipping, so adding data skipping support for a collection of specific data types / UDFs is a difficult task. Supporting data skipping in an extensible manner includes important features of data skipping indexes, such as no false negatives. Such properties are different from those of conventional database indexes (which are typically as large as the data itself and thus not usable for data skipping).
[0042] The present technology can provide the ability to independently implement data skipping APIs and combine their inputs in some cases where the input of each alone may not be sufficient to allow data skipping. For example, if a data set has both a geospatial column and a time series column, different developers can independently implement data skipping APIs for each of these data types. Using the present technology, a query involving the predicate "P1 OR P2" (where P1 involves the geospatial column and P2 involves the time series column) can skip data by combining the two API implementations.
[0043] Figure 2 An example of an interface that can be included in the extensible data skipping framework 106 is shown. In this example, the interface can include one or more sets of extensible data type interfaces 202, one or more extensible data skipping specification interfaces 204, and can utilize the metadata memory 116. The extensible data type interfaces 202 can provide the ability to define one or more data types for which data skipping can be implemented. For example, the extensible data type interfaces 202 can provide the ability to specify a new MetaDataType, which is a logical representation of summary metadata. Examples of such data types can include geospatial data (such as GPS coordinates, street addresses, etc.), asterisk data (such as astronomical coordinates, etc.), and genomic data (such as DNA / protein sequences, etc.).
[0044] The extensible data skip specification interface 204 can provide the ability to add a new data skip index type (MetadataType) for the data types included in the extensible data type 202. For example, geospatial data can use a 2D bounding box data skip index type, asterisk data can use a 3D structure data skip index type, and genomic data can use a bloom filter / trie tree search tree filter for sequence data. The extensible data skip specification interface 204 can include a metadata generation interface that provides the ability to convert a representation of a subset of rows, such as a sparse data frame, into a MetaDataType defined by the extensible data type interface 202. Additionally, the extensible data skip interface 204 can include a conversion interface where, depending on the specific metadata store used, the logical representation of the summary metadata can be converted into a physical representation and stored in the metadata store 206. Once the above data is defined, a data skip index for a given data set can be created.
[0045] An example of clause generation 124 is shown in Figure 3 . Clause generation 124 can be used to specify how the predicates used in a query are mapped to summary metadata to evaluate whether data can be skipped. Clause generation 124 can include API interfaces such as clause interface 302, SQL-to-clause interface 304, SQL mapping interface 306, and clause mapping interface 308. A term "clause" can be defined that applies to an object (e.g., a file) and returns a boolean value. A clause c "represents" (boolean value) a query expression e if whenever an object contains a row that satisfies e, then the object satisfies c. This means that if an object (a subset of its rows) does not satisfy clause c, then the object can be skipped. More formally, a clause is a boolean function c: U → {0, 1}, where U represents the universe of possible files. Assume e represents an expression (more specifically, a boolean expression). Clause c represents e (denoted by c e) if the following holds. If a file contains a row that satisfies e, then the file satisfies c. This ensures that if a file does not satisfy c, then it does not contain relevant data for the query. Thus, if clauses are constructed that represent query expressions, then all files that do not satisfy the clause can be safely skipped.
[0046] For example, a clause can be defined as MaxClause clause(age, >, 5), which represents age > 5. Whenever an object has a row with age > 5, its maximum age > 5. Then, if the maximum age of an object does not exceed 5, then the object can be skipped. This property can be used to ensure no false negatives.
[0047] The clause interface 302 may provide the ability to define appropriate clauses for the metadata. The SQL-to-clause interface 304 may provide the ability to map SQL query expressions to clauses representing them. These query expressions may include user-defined functions (UDFs). The SQL mapping interface 306 may provide the ability to specify how to recursively map complex SQL query expressions based on simpler expressions. For example, the SQL mapping interface 306 may be used to specify how to handle AND, OR, and NOT. Default embodiments such as AND and OR may be provided, but the default embodiments may be overridden. Generally, NOT may be implemented where there are also no false positives, but this is not always the case.
[0048] This recursive construction depends on the available metadata indexes of the data set. It is possible that more than one metadata index is available and can be applied. This increases the chance of data skipping because data can be skipped due to any single metadata index that provides skip tuning.
[0049] The clause mapping interface 308 may provide the ability to map from clauses to code / query or search specifications in the corresponding metadata store (such as ES). In the case of ES, this will result in searching for objects that can be skipped.
[0050] The predicate specification interface 116 may be used to define all parts where data skipping can be applied to complex query expressions, involving those UDFs specified in the SQL-to-clause interface 304 in SQL, and to handle those complex expressions specified in the SQL mapping interface 306. As described above, this data skipping can be applied relative to the available data skipping indexes of the data set.
[0051] This technique may provide multiple optimization rules (filters). In this context, an "optimization rule" (filter) is an algorithm that operates on an expression and tags the expression tree nodes with clauses such that each node it tags is tagged with the clause representing that node. This list may be changed at runtime based on external rule registration (via a dedicated registration API) and based on existing metadata (rules that require metadata not yet collected will not be applied). During query time, all filters may be applied to the query expression. The results from all filters may be combined to create an abstract clause. Ensuring that this clause represents the query expression means that as long as the optimization rules are implemented correctly, objects containing relevant data will never be skipped. Thus, to add an optimization rule, all that needs to be done is to implement the rule logic itself - all other work is done by the skip framework.
[0052] An exemplary process of operation 400 of this system and method is shown in Figure 4 which is shown in. Process 400 begins with the definition of a query expression 402, as Figure 5shown in the form of an expression tree at 502 and in SQL form at 504. In this example, it is assumed that the existing data skip index includes a MinMax index on salary and a ValueList index on name. At 404, executable metadata processing can be performed to generate a representation of clause 406 that represents the query expressions 502, 504. That is, the query expression 402 can be modified to form a clause by modifying the query expression 402 to reference the relevant data skip index. For example, as Figure 6 shown, the query expression can be modified by a filter 602 marked to reference the MinMax index on salary. Similarly, as Figure 7 shown, the query expression can be modified by a filter 702 marked to reference the ValueList index on name. Figure 8 An example of the clause 406 that can be generated is shown in
[0053] Returning to Figure 4 , at 408, the generated clause can be transformed to form a metadata store representation 410 of clause 406. Figure 9 An example of translating clause 406 to form a metadata store representation 410 using, for example, an ES metadata store is shown in Figure 1 . When executing a query corresponding to the query expression 402, the metadata store representation 410 can be used to execute the query using the data skip index that the query expression has been modified to reference. The processing shown at 402 to 410 can be performed, for example, by Figure 1 the expression tree processing 122, abstract clause generation 124, transformation 126, and index generation and maintenance 128 shown in
[0054] To generate a data skip index, a query expression (e.g., 402) can be evaluated relative to a representation of a subset of rows 412 (e.g., a sparse data frame). The metadata generation interface 414 can provide the ability to transform a representation of the subset of rows 412 (e.g., a sparse data frame) into a metadata type 416 defined by the extensible data type interface 202 shown in Figure 2 . The transformation interface 416 can have the ability to transform a logical representation of the summary metadata into a physical representation 420 depending on the particular metadata store used, and can be stored in the Figure 2 shown metadata store 116. Once the above data is defined, a data skip index for a given data set can be created. Once created, the data skip index can be referenced by a filter that can be used to modify the query expression 402 to execute the query using the data skip index that the query expression has been modified to reference.
[0055] This technique can use any type of desired filter. Examples of such filters are shown in Figures 10 - 12 Figure 10An example of a MinMax filter is shown, including Max filter embodiment 1002 and Min filter embodiment 1004, assuming a MinMax index on a selected attribute. Figure 11 An example of a ValueList filter is shown, including In filter implementation 1102 and Equals filter implementation 1404, assuming a ValueList index on a selected attribute. Figure 12 An example of a GeoBox filter 1202 is shown, assuming a GeoBox index in latitude and longitude, where the GeoBox filter 1202 looks for a unified GeoBox bounding box in a direct join path.
[0056] In addition to the predefined filters, the present technology can be used with custom filters, which can be defined, for example, by extending an abstract filter class with code that implements the desired filter functionality. This can provide the ability to capture any structure in the tree and add abstract clauses, which can be useful, for example, for UDFs and SQL predicates LIKE.
[0057] It should be noted that embodiments of the present system and method can be applied to any type of structured, semi-structured, or unstructured data. Additionally, for embodiments applied to tabular data, such embodiments can utilize indexes related to multiple columns. For example, GeoBox bounding box data can utilize indexes on multiple columns, such as both latitude and longitude. Similarly, for any type of data that can utilize a "feature" index, the feature index can be related to multiple columns.
[0058] Figure 14 An exemplary process 1400 for index creation based on an existing data stream is shown. In conjunction with Figure 1 Best seen. Process 1400 can utilize and interact with application 108, other systems 110, user 112, and data processing system 102. Process 1400 begins at 1402, where index processing can be performed. Index processing 1402 can provide the ability to transform a representation of a subset of rows into a defined MetaDataType. At 1404, at least a portion of the corresponding data set can be read from data memory 104. At 1406, index generation and maintenance 128 can generate indexes that implement data skipping for the defined data set. At 1408, transformation 126 can transform the logical representation of the summary metadata into a physical representation depending on the particular metadata store used, and at 1410, the logical representation of the summary metadata can be stored in metadata memory 116.
[0059] Figure 15 An exemplary process 1500 for query processing flow is shown. In conjunction with Figure 1Best seen. Process 1500 can utilize and interact with application 108, other systems 110, and user 112. Process 1500 begins at 1502, where data ingestion can be performed. Data ingestion 1502 can utilize data ingestion system 114 to receive and process new data. At 1504, index generation and maintenance 128 can generate an index that implements defined data skipping for the newly ingested data. At 1506, transformation 126 can transform the logical representation of the summary metadata into a physical representation depending on the particular metadata store being used, and at 1508, the logical representation of the summary metadata can be stored in metadata store 116. At 1510, the ingested data can be stored in data store 104, as indexed by the newly generated index.
[0060] Figure 16 An exemplary process 1600 for index creation based on an existing data stream is shown in conjunction with Figure 1 Best seen. Process 1600 can utilize and interact with application 108, other systems 110, user 112, and data processing system 102. Process 1600 begins at 1602, where query processing can be performed. Query processing 1602 can perform database or other processing of queries, as well as the creation, maintenance, and processing of indexes that can be used to process queries. At 1604, data store 104 can be accessed to obtain a list. At 1606, expression tree processing 122 can accept a query expression, such as in the form of an expression tree but also in SQL form, and perform metadata processing of the query expression. At 1608, abstract clause generation 124 can generate an abstract representation of a clause that references a relevant data skipping index. At 1610, transformation 126 can transform the logical representation of the summary metadata into a physical representation depending on the particular metadata store being used, and at 1612, the logical representation of the summary metadata can be stored in metadata store 116. At 1614, at least a portion of the corresponding data set can be read from data store 104.
[0061] Registration 120 can provide the ability to extend the system by adding support for new UDFs and metadata index types in the extensible data skipping framework 106 using the extensible data skipping API. Such an extension can utilize and interact with application 108, other systems 110, user 112, and data processing system 102.
[0062] Figure 13An exemplary block diagram of a computer system / calculation device 1300 is shown, in which the processes involved in the embodiments described herein can be implemented. The computer system / calculation device 1300 can be implemented using one or more programmed general computer systems (such as embedded processors, system-on-chips, personal computers, workstations, server systems, and minicomputers or mainframes), mobile devices (such as smart phones or tablets), or in a distributed networked computing environment. The computer system / calculation device 1300 can include one or more processors (CPUs) 1302A - 1302N, input / output circuitry 1304, a network adapter 1306, and a memory 1308. The CPUs 1302A - 1302N execute program instructions to perform the functions of this communication system and method. Generally, the CPUs 1302A - 1302N are one or more microprocessors, such as a processor or a processor. Figure 13 An embodiment is shown in which the computer system / calculation device 1302 is implemented as a single multi-processor computer system / calculation device, where multiple processors 1302A - 1302N share system resources such as the memory 1308, input / output circuitry 1304, and network adapter 1306. However, this communication system and method also includes embodiments in which the computer system / calculation device 1302 is implemented as multiple networked computer systems, and the multiple networked computer systems can be single-processor computer systems / calculation devices, multi-processor computer systems / calculation devices, or a mixture thereof.
[0063] The input / output circuitry 1304 provides the ability to input data to or output data from the computer system / calculation device 1302. For example, the input / output circuitry can include input devices (such as keyboards, mice, touch pads, trackballs, scanners, analog-to-digital converters, etc.), output devices (such as video adapters, monitors, printers, biometric information acquisition devices, etc.), and input / output devices (such as modems, etc.). The network adapter 1306 interfaces the device 1300 with the network 1310. The network 1310 can be any public or private LAN or WAN, including but not limited to the Internet.
[0064] The memory 1308 stores program instructions executed by the CPU 1302 and data used and processed by the CPU 1302 to perform the functions of the computer system / computing device 1302. The memory 1308 can include, for example, electronic memory devices (such as random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.), and electromechanical memories such as disk drives, tape drives, and optical disk drives, which can use an Integrated Drive Electronics (IDE) interface or a variant or enhancement thereof, such as Enhanced IDE (EIDE) or Ultra Direct Memory Access (UDMA) or Small Computer System Interface (SCSI)-based interfaces, or variants or enhancements thereof, such as Fast SCSI, Wide SCSI, Fast and Wide SCSI, Serial Advanced Technology Attachment (SATA) or variants or enhancements thereof, or Fibre Channel Arbitrated Loop (FC-AL) interfaces.
[0065] The contents of the memory 1308 can vary depending on the functions that the computer system / computing device 1302 is programmed to perform. In Figure 13 the example shown, exemplary memory contents representing routines and data for embodiments of the above processes are shown. For example, in Figure 13 it is shown the memory contents of both a computer system and a database management system including an extensible data skip API. However, those skilled in the art will recognize that, based on well-known engineering considerations, these routines and the memory contents associated with these routines may not be included on one system or device, but may be distributed among multiple systems or devices. This communication system and method can include any and all such arrangements.
[0066] In Figure 13In the example shown, the memory 1308 may include software code and data for the scalable data skip framework 1312 and for the database management system 1316. The scalable data skip framework 1312 may include a scalable data skip interface routine 1318 and a predicate specification interface routine 1320. The scalable data skip interface routine 1318 may include software routines for implementing support for new types of summary metadata and for new data skip index types, and may output data skip indexes and associated metadata. The scalable data skip framework 106 may include registration 1320, expression tree processing 1320, abstract clause generation 1322, transformation 1324, and index generation and maintenance 1326. Registration 1320 may provide the ability to extend the system by adding support for new UDFs and metadata index types in the scalable data skip framework 106 using the scalable data skip API. Expression tree processing 1320 may accept query expressions, for example in the form of an expression tree but also in SQL form, and perform metadata processing of the query expression such that abstract clause generation 1322 may generate an abstract representation of a clause that references a relevant data skip index. Transformation 1324 may transform the abstract clause to form a metadata representation of a clause that references a relevant data skip index. The metadata representation may be stored in a metadata memory. Index generation and maintenance 1326 may generate and maintain data skip indexes.
[0067] The data processing system 1316 may include data query / index processing 1328, which may use data skip indexes to process databases or other processing that processes queries when processing queries. The operating system 1330 may provide overall system functionality.
[0068] As Figure 13 shown, the present communication system and method may include implementations on one or more systems that provide multi-processor, multi-task, multi-process, and / or multi-threaded computing, as well as implementations on systems that provide only single-processor, single-threaded computing. Multi-processor computing involves using more than one processor to perform a computation. Multi-task computing involves using more than one operating system task to perform a computation. A task is an operating system concept that refers to a combination of a program being executed and the bookkeeping information used by the operating system. Whenever a program is executed, the operating system creates a new task for it. A task is similar to an envelope for a program because it identifies the program with a task number and attaches other bookkeeping information to the program. Many operating systems (including Linux, and )An operating system that can run many tasks simultaneously is called a multitasking operating system. Multitasking is the ability of an operating system to execute more than one executable file at the same time. Each executable file runs in its own address space, which means that executable files cannot share any memory in their memory. This has the advantage that no program can corrupt the execution of any other program running on the system. However, programs have no way to exchange any information except through the operating system (or by reading files stored on the file system). Multiprocessing computing is similar to multitasking computing because the terms task and process are often used interchangeably, although some operating systems make a distinction between the two.
[0069] The present invention can be a system, method, and / or computer program product at any possible level of integrated technical detail. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention. A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device.
[0070] A computer-readable storage medium can be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as punched cards or raised structures in grooves having instructions recorded thereon), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0071] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0072] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit so as to perform aspects of the present invention.
[0073] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0074] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, which, when executed by the processor of the computer or other programmable data processing apparatus, creates a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having the instructions stored therein includes an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0075] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, other programmable apparatus or other devices that cause a series of operational steps to be performed on a computer, to produce a computer-implemented process such that the instructions executed on the computer, other programmable apparatus or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0076] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart, and combinations of blocks in the block diagrams and / or flowchart, can be implemented by special-purpose hardware-based systems that perform the specified functions or actions or implement a combination of special-purpose hardware and computer instructions.
[0077] Although specific embodiments of the present invention have been described, those skilled in the art will understand that there are other embodiments equivalent to the described embodiments. Therefore, it should be understood that the present invention is not limited by the specific embodiments described, but only by the scope of the appended claims.
Claims
1. A method for data skipping, comprising: receiving a query at a computer system, the computer system including a processor, a memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor; modifying, at the computer system, the received query to evaluate at least one criterion of the query by using at least one data skipping index, wherein the at least one data skipping index includes an index on at least one attribute of data, the at least one attribute of the data excluding at least a portion of those data items that do not satisfy the at least one criterion, wherein the at least one data skipping index relates to a newly added data skipping index type, and wherein at least one of the at least one data skipping index or a mapping from the at least one criterion to the at least one data skipping index is generated based on information received from an application programming interface; and evaluating, at the computer system, the query, wherein the at least one data skipping index is generated by the following steps: receiving, at the computer system, information defining a data skipping index type from the application programming interface; receiving, at the computer system, information interpreting the at least one criterion from the application programming interface; generating, at the computer system, metadata related to the defined data skipping index type and the defined at least one criterion; and generating, at the computer system, the at least one data skipping index based on the generated metadata.
2. The method according to claim 1, wherein, the received query is represented as an expression tree, and the expression tree is modified by using optimization rules to label at least one node of the expression tree with a clause representing a skip requirement of the at least one criterion and referencing at least one data skipping index.
3. The method according to claim 2, wherein the at least one criterion is a Structured Query Language (SQL) predicate.
4. The method according to claim 1, wherein the at least one data skipping index is generated by the following steps: receiving, at the computer system, information defining a plurality of data skipping index types from the application programming interface; receiving, at the computer system, information defining a plurality of criteria from the application programming interface; generating, at the computer system, metadata related to each defined data skipping index type and each defined criterion; combining, at the computer system, the metadata related to each defined data skipping index type and each defined criterion to form metadata related to the plurality of defined data skipping index types and the plurality of defined criteria; and generating, at the computer system, the at least one data skipping index based on the generated metadata.
5. The method according to claim 4, wherein, the received query is represented as an expression tree, and the expression tree is modified by using a plurality of optimization rules to label each of a plurality of nodes of the expression tree with a clause representing a skip requirement of a criterion and referencing a data skipping index.
6. The method according to claim 5, wherein each criterion is a Structured Query Language (SQL) predicate.
7. A system for data skipping, comprising a processor, a memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor to: Receive a query; Modify the received query to evaluate at least one criterion of the query by using at least one data skipping index, wherein The at least one data skipping index includes an index on at least one attribute of data, the at least one attribute of the data excluding at least a portion of those data items that do not satisfy the at least one criterion, wherein the at least one data skipping index relates to a newly added data skipping index type, and wherein at least one of the at least one data skipping index or a mapping from the at least one criterion to the at least one data skipping index is generated based on information received from an application programming interface; and Evaluate the query, wherein the at least one data skipping index is generated by the steps of: Receiving information from the application programming interface that interprets the data skipping index type; Receiving information from the application programming interface that defines the at least one criterion; Generating metadata related to the defined data skipping index type and the defined at least one criterion; And Generating the at least one data skipping index based on the generated metadata.
8. The system according to claim 7, wherein The received query is represented as an expression tree, and the expression tree is modified by using optimization rules to label at least one node of the expression tree with a clause that represents the skipping requirement of the at least one criterion and references at least one data skipping index.
9. The system according to claim 8, wherein the at least one criterion is a Structured Query Language (SQL) predicate.
10. The system according to claim 7, wherein the at least one data skipping index is generated by the steps of: Receiving information from the application programming interface that defines multiple data skipping index types; Receiving information from the application programming interface that defines multiple criteria; Generating metadata related to each defined data skipping index type and each defined criterion; Combining the metadata related to each defined data skipping index type and each defined criterion to form metadata related to the multiple defined data skipping index types and the multiple defined criteria; And Generating the at least one data skipping index based on the generated metadata.
11. The system according to claim 10, wherein The received query is represented as an expression tree, and the expression tree is modified by using multiple optimization rules to label each of multiple nodes of the expression tree with a clause that represents the skipping requirement of a criterion and references a data skipping index.
12. The system according to claim 11, wherein each criterion is a Structured Query Language (SQL) predicate.
13. A computer program product comprising a non-transitory computer-readable storage device having program instructions embodied therein, the program instructions being executable by a computer to cause the computer to perform a method, the method comprising: receiving a query; modifying the received query to evaluate at least one criterion of the query by using at least one data skip index, wherein the at least one data skip index includes an index on at least one attribute of data, the at least one attribute of the data excluding at least a portion of those data items that do not satisfy the at least one criterion, wherein the at least one data skip index relates to a newly added data skip index type, and wherein at least one of the at least one data skip index or a mapping from the at least one criterion to the at least one data skip index is generated based on information received from an application programming interface; and evaluating the query, wherein the at least one data skip index is generated by the following steps: receiving information defining a data skip index type from the application programming interface; receiving information interpreting the at least one criterion from the application programming interface; generating metadata related to the defined data skip index type and the defined at least one criterion; and generating the at least one data skip index based on the generated metadata.
14. The computer program product according to claim 13, wherein, the received query is represented as an expression tree, and the expression tree is modified by using optimization rules to label at least one node of the expression tree with a clause representing a skip requirement of the at least one criterion and referencing at least one data skip index.
15. The computer program product according to claim 13, wherein the at least one data skip index is generated by the following steps: receiving information defining a plurality of data skip index types from the application programming interface; receiving information defining a plurality of criteria from the application programming interface; generating metadata related to each defined data skip index type and each defined criterion; combining the metadata related to each defined data skip index type and each defined criterion to form metadata related to the plurality of defined data skip index types and the plurality of defined criteria; and and generating the at least one data skip index based on the generated metadata.
16. The computer program product according to claim 15, wherein, the received query is represented as an expression tree, and the expression tree is modified by using a plurality of optimization rules to label each of a plurality of nodes of the expression tree with a clause representing a skip requirement of a criterion and referencing a data skip index.
17. The computer program product according to claim 16, wherein each criterion is a Structured Query Language (SQL) predicate.
Citation Information
Patent Citations
System, method, and apparatus for parallelizing query optimization
US20110047144A1
Extensible RDF databases
US20120150922A1
Computer Relational Database Method and System Having Role Based Access Control
US20130138666A1