Secure query processing
By executing plans in stages in distributed databases, separating user-defined functions and accessing secure tables, solving the problem of difficulty in achieving security in multi-user environments, and achieving efficient and secure query processing.
Patent Information
- Application Number
- CN202380068852.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-26
- Filing Date
- 2023-09-25
- Publication Date
- 2025-05-06
AI Technical Summary
In distributed databases, it is difficult to achieve security without sacrificing performance and flexibility, especially in multi-user environments where user-defined functions can pose security risks.
By generating and executing a phased execution plan, the execution of user-defined functions is separated from access to security tables, ensuring that incompatible risk classifications are performed in different execution stages, and risk is controlled through separation and tagging of executors.
Effectively prevent user-defined functions from being used as Trojan horses, obtaining unauthorized access data, improving the security of distributed databases while maintaining performance and flexibility.
Smart Images

Figure CN119948473A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Patent Application No. 17 / 953,038, entitled “SECURE QUERY PROCESSING,” filed on September 26, 2022, the entire contents of which are incorporated herein by reference and for all purposes. For the United States, which recognizes different types of priority, this application claims priority to and is a continuation of U.S. Patent Application No. 17 / 953,038. The identification of this international application as a continuation application only with respect to the United States is not intended to change the priority claim with respect to any other member state or jurisdiction. Background Art
[0003] Distributed databases are increasingly used in a variety of applications, including those where performance, flexibility, and security are important factors. Distributed databases are also increasingly used in multi-user environments. In these and other environments, security is often difficult to achieve without sacrificing performance and flexibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Various techniques will be described with reference to the accompanying drawings, in which:
[0005] Figure 1 illustrating an example of a distributed database system according to at least one embodiment;
[0006] Figure 2 An example of a distributed database system assigning portions of a query plan to stages of an execution plan according to at least one embodiment is illustrated;
[0007] Figure 3 An example of a distributed database system assigning stages of an execution plan to computing nodes for execution according to at least one embodiment is illustrated;
[0008] Figure 4 illustrating an example of maintaining a pool of compute node executors according to at least one embodiment;
[0009] Figure 5 illustrates an example process for processing a database query by a distributed database system according to at least one embodiment;
[0010] Figure 6 illustrates an example process for generating a secure execution plan for a distributed database system according to at least one embodiment;
[0011] Figure 7 illustrating an example process for executing a query using a distributed database that separates operations associated with incompatible risk classifications according to at least one embodiment; and
[0012] Figure 8 Illustrate systems in which various embodiments may be implemented. DETAILED DESCRIPTION
[0013] In one example, a distributed database system processes queries in a secure manner by dividing query operations into stages based on risk classifications associated with the stages. For example, a query may include both a user-defined function and access to a security table, the security of which may be compromised by the execution of the user-defined function. In an example, the distributed database processes queries by generating an execution plan in which the stages of executing the user-defined function are separated from the stages of accessing the security table. The stages of the execution plan are then executed by separate executors. An executor is selected to execute a given stage so that any executor currently being used to execute a user function or previously used to execute a user function is not used to access the security table. Similarly, an executor currently used to access the security table is not used to execute a user-defined function. This approach prevents a user function from being used as, for example, a Trojan horse, which may then obtain access to data from the security table that the user function is not granted access to.
[0014] In the previous and following descriptions, various techniques are described. For the purpose of explanation, specific configurations and details are set forth in order to provide a thorough understanding of possible ways to implement the techniques. However, it will also be apparent that the techniques described below may be practiced in different configurations without the specific details. In addition, well-known features may be omitted or simplified to avoid obscuring the described techniques.
[0015] Figure 1 An example of a distributed database system is illustrated in accordance with at least one embodiment. In the example embodiment, a distributed database 100 includes a query engine 102 that processes a query 106 using executors 120a-c.
[0016] A query, such as depicted query 106, may include or correspond to instructions for inserting, updating, deleting, or reading data stored in a distributed database. In at least one embodiment, query 106 is embodied as textual data, which may include, but is not limited to, structured query language ("SQL") statements or other programming languages. In some embodiments, query 106 may be expressed as a natural language. Query 106 may also be embodied in code, such as as a series of application programming interface ("API") calls.
[0017] A distributed database, such as depicted distributed database 100, may include any of a variety of computing systems that use multiple computing nodes to store partitioned data to store and retrieve data. The computing nodes may include any of a variety of computing devices that include at least one processor, a memory device for storing instructions to be processed by the at least one processor, and a storage device on which a portion of the partitioned data is stored. Queries of the distributed database are processed by performing insert, update, delete, and / or read operations on two or more of the multiple computing nodes that make up the distributed database.
[0018] An executor, such as any of the depicted executors 120a-c, may include one of the computing nodes that make up the distributed database 100. The executors assist in processing queries by each executing one or more phases of an execution plan. The phases may include operations such as accessing a table, merging data from different tables, sorting data, executing a user-defined function, and the like.
[0019] Distributed database 100 may include a query engine 102. A query engine, such as the depicted query engine 102, may include software and / or hardware to perform the functions described and attributed thereto herein. Coordinator 104 interacts with query engine 102 via an API to obtain a query 106 and generates a query plan 108 from the query.
[0020] A query plan, such as the depicted query plan 108, may include a set of instructions indicating which operations are to be performed in order to process the query. The query plan may have a tree structure consisting of various nodes, where each node represents one or more of these operations. These operations may include, but are not limited to, operations such as reading data from a table, writing data to a table, executing a user-defined function, merging data, sorting data, and the like.
[0021] The coordinator 104 cleans the query plan 108 to form a cleaned query plan 110. A cleaned query plan, such as the depicted cleaned query plan 110, may include attributes indicating classifications of risks associated with certain nodes and operations represented by those nodes. For example, nodes associated with user-defined functions may be marked with the attribute UserCode=true, while nodes associated with access to secure data may be marked with the attribute Secured=true. The cleaned query plan may also be modified relative to the query plan on which the cleaned query plan is based to ensure that operations with incompatible risk classifications are separable so that the operations may be performed in different stages of the execution plan.
[0022] The coordinator 104 transforms the cleaned query plan 110 into an execution plan 112. An execution plan, such as the depicted execution plan 112, is a set of executable stages that can be executed by an executor, such as the depicted executors 120a-c. Each stage of an execution plan can correspond to a node of the query plan on which the execution plan is based. The stages can be executed by the executors in an order specified by the execution plan, although some stages may be executed in a varying or indeterminate order, or in parallel.
[0023] Each stage of the execution plan 112 may access one or more data sources, such as the depicted data source 130. A data source, such as the depicted data source 130, may include any computing device or service that maintains a table or partitions of a table. In some cases, the executor itself may store such data, such as when the executor maintains horizontal or vertical partitions of a table. In other cases, the executor accesses data stored on another computing node or service, such as Figure 1 Depicted.
[0024] The distributed database 100 may include the ability to execute user-defined functions. User-defined functions (also referred to as user functions, user code, etc.) may include processor executable code, intermediate code, interpretable code, etc. Examples of user-defined functions include routines that accept one or more parameters and return one or more values as output. User-defined functions provide important and useful flexibility, but may be associated with a certain degree of risk, particularly in a multi-user environment. For example, a user-defined function may contain "Trojan horse" code that is intended to obtain access rights that the user is not granted.
[0025] Distributed database 100 may also include support for secure access to certain types of tables or other data sources, such as tables or other data sources that include data owned by multiple users, where each user may have a different set of permissions. For example, a given user may be permitted to access data owned by that user, but not permitted to access data owned by any other user.
[0026] In order to operate at high speed, embodiments of the distributed database 100 may execute user-defined functions within the executor and implement access control within the query engine 102. However, there are security issues with this approach. The code submitted by the user may potentially have the ability to inspect and / or change the content or behavior of the application program running therein or the system running thereon when executed, the application program or the system including one or more of a process, a virtual machine, or a computing device. This may be the case whether or not it is intended to be used in this way. In addition, the user-defined function may leave running code that may still endanger the system even after the user-defined function is executed.
[0027] In some embodiments, the distributed database system may execute a query according to the following steps. First, a user may submit a query 106 by providing SQL text or calling one or more application programming interfaces ("APIs"). The query 106 is then transformed into a query plan. The query plan may take the form of a tree data structure, where the leaf nodes of the tree typically represent operations that read data from a data source. The ancestors of these nodes may describe various operations that may be performed on the data, such as sorting, filtering, merging, etc. The resulting query plan may then be transformed into an executable plan consisting of a set of stages. The stages may have dependencies, such as one stage requiring input provided by the execution of another stage.
[0028] The coordinator 104 then orchestrates the execution of the phases by sending the phases to the executors 120a-c. The executors then execute the corresponding phases.
[0029] However, implementations that take this approach may encounter certain issues. For example, a single stage may have code that both reads data from a secure source and applies a security filter to the data and executes a user-defined function. This may present a security risk because the user-defined function may execute code that may interfere with the operation of the security filter, or otherwise obtain data that the user-defined function is not granted access to.
[0030] To address these issues, an embodiment may generate an execution plan that prevents user-defined functions from executing on an executor that would access secure data. Secure data may include data subject to a security policy, such as a security policy that limits user access to certain rows in a table. This may include multi-user tables where a given user is only permitted access to data owned by that user. Security policies may include any of a variety of restrictions on data access, including restrictions on reading, writing, updating, or deleting data. Security policies may be embodied in a variety of techniques or algorithms for implementing these restrictions.
[0031] In at least one embodiment, the coordinator 104 includes a cleanup component that takes the query plan Q1 as input and transforms it into a cleaned query plan Q2. The plan Q2 is then used to generate an execution plan.
[0032] To generate Q2, the cleanup component of the coordinator 104 searches Q1 for each node N that has a user-defined function associated therewith. i , and N i The node is marked as UserCode=true. The other nodes are set to UserCode=false. The nodes of the query plan (such as node N i ) may include attributes (such as a UserCode attribute) for indicating the risk classification associated with the node.
[0033] The cleanup component also searches for access to the secure data source T in Q1. s Each node N s This step may include finding s associated data catalog to determine if the data is secure, and / or determine if part of Plan Q1 is necessary to s The obtained data has security filters applied to it. If a node accesses a secure data source, it is marked as Secured = true. Other nodes can be marked as Secured = false.
[0034] In at least one embodiment, the cleaning component is not simply to s is marked as Secure = true, but node N s Replace with s Nodes that access data and apply relevant security filters s ). NodeFilter(N s ) is marked as Secured = true. Node Filter (Ns) can be a node that includes a s The leaf nodes of the data and the subtree of ancestor nodes to which the security filter is applied.
[0035] In at least one embodiment, the planner component of the coordinator 104 then uses the resulting cleaned query plan Q2 to generate an execution plan 112 in which user code execution and secure data access are separated into different execution phases. This can be done by the planner during the generation of the execution plan 112, or as a post-processing step, where the initial version of the execution plan is cleaned up by moving user code execution and secure data access into different phases.
[0036] In at least one embodiment, the technique for generating the execution plan includes identifying each stage in the execution plan that contains nodes with Secured=true, and marking these stages as Secured=true. Similarly, if there is a stage that contains a node with UserCode=true, then mark the stage with UserCode=true. Then, any stage that has both Secured=true and UserCode=true is split into two or more stages, so that no single stage with both UserCode=true and Secured=true is generated. In some cases, these stages can be merged with other stages, provided that there is no stage in the resulting execution plan that sets both properties to true.
[0037] In at least one embodiment, the phases of the execution plan are assigned to executors 120a-c in the following manner. Each executor is initially marked with a tag (such as Sandboxed=false) indicating that it has never executed a phase containing user code. Once the executor is used to execute a phase containing user code, it is also marked with a tag (such as Sandboxed=true).
[0038] When the coordinator 104 schedules a phase for execution, it can assign the UserCode=true phase to any executor if that executor is not running a Secure=true phase. For example, if an executor is currently running another phase Secured=true, the coordinator 104 will not select that executor to run the UserCode=true phase. However, once the Secured=true phase is completed, the executor may be selected.
[0039] When an executor is used to execute a UserCode=true phase, the executor is marked as Sandboxed=true, and the coordinator 104 will no longer assign phases with Secured=true to the executor. The executor does not assign a Secured=true phase to an executor that has Sandboxed=true when assigning it to the executor. In at least one embodiment, the executor will remain Sandboxed=true indefinitely (e.g., until a hardware reset) such that a Secured=true phase is never executed on an executor that is used to execute user code.
[0040] Figure 2 An example of assigning parts of a query plan to stages of an execution plan according to a distributed database system of at least one embodiment is illustrated. In the depicted example 200, a query plan 202 is a tree structure including nodes 206-210. The nodes 206-210 of the query plan 202 represent operations to be performed in order to process a query. The branches between the nodes may represent dependencies between the nodes or a suggested execution order. As depicted in the example 200, the nodes of the query plan 202 may be assigned risk classifications. Some nodes (such as the depicted node 208) may include instructions for executing user-defined functions or other user codes and are assigned to user code risk classifications. Other nodes (such as the depicted node 210) may include instructions for accessing secure data and are assigned to secure data risk classifications. Still other nodes (such as the depicted node 208) may not be associated with operations that are considered risky and may be assigned neutral risk classifications. Once assigned, the risk classification may be indicated by the attributes stored using the corresponding nodes of the query plan 202.
[0041] Distributed databases (such as Figure 1 The distributed database depicted in FIG. 200 may transform the query plan 202 into an execution plan 204. The distributed database may generate the execution plan 204 using the risk classifications assigned to the nodes in the query plan 202. Figure 1 Consistent with the algorithm described above, the distributed database generates an execution plan 204 such that any stage including user code does not include secure data access, and any stage including secure data access does not include user code. Figure 2 As depicted, stage 222 of query execution plan 204 generated by the distributed database includes executing operations corresponding to node 208, which includes user code. However, stage 222 does not include executing any operations involving secure data access, such as those represented by node 210 of query plan 202. Nodes that are not associated with any risk classification, such as node 206, may be assigned to any stage, such as stage 220. The distributed database will consider dependencies between nodes when generating execution plan 204. For example, stage 224 may be executed first because it includes nodes of the query plan that obtain data that are supplied to its ancestor nodes for further processing.
[0042] Figure 3 An example of a distributed database system according to at least one embodiment assigning stages of an execution plan to computing nodes for execution is illustrated. In example 300, execution plan 304 (which may be associated with Figure 2 ) includes three stages 306-310. To execute the instruction set, each stage is assigned to an executor, such as one of the executors 320-322 depicted. The assignment may be performed by a distributed database (including via a connection to Figure 1 320-322). Of the depicted executors 320-322, the secure stage 310 may be assigned by the distributed database only to the unsandboxed executor 322, meaning that the executor has never been used to execute a stage that includes user code. The user code stage 308 may be assigned to any of the executors 320-322, but the distributed database may implement a preference for assigning it to an already sandboxed executor (such as the depicted executor 320). The final stage 306, which contains neither user code nor secure data access, may be assigned by the distributed database to any of the executors 320-322.
[0043] Figure 4An example of maintaining a pool of compute node executors according to at least one embodiment is illustrated. As described with respect to the previous figures, the distributed database can prevent executors used to execute the user code phase from being used to execute the secure data access phase. In at least one embodiment, the process uses a pooling mechanism (such as, Figure 4 In example 400, executors 420-424 are assigned to pools 402-404. These pools may include a pool of sandboxed executors 402 and a pool of non-sandboxed executors 404. Here, sandboxed refers to executors that have been used to execute stages with user code, or executors that are intended to be used specifically to execute stages with user code. Non-sandboxed refers to nodes that have not yet been used to execute user code.
[0044] like Figure 4 As depicted, executor 422 in non-sandboxed pool 404 can be transferred to sandboxed pool 402. This can be done, for example, when an insufficient number of executors are available to execute user code. However, once added to sandboxed pool 402, executors are not transferred back to non-sandboxed pool 404. For example, once executor 420 has been placed in sandboxed pool 402, it will not be transferred back to non-sandboxed pool 404. If additional executors are needed in pool 404, they can be added by pulling from other sources (such as from a newly configured additional executor pool).
[0045] Figure 5 An example process for processing a database query by a distributed database system according to at least one embodiment is illustrated. Figure 5 The process 500 of the present invention is depicted as a certain order of steps, but the depicted order should not be interpreted as limiting the scope of the present disclosure to only those embodiments that conform to the depicted order. For example, unless otherwise specified or clear from the context (e.g., when the output of one step is used as the input of another step), at least some of the depicted steps may be reordered or performed in parallel.
[0046] At 502, the distributed database obtains a query plan. The query plan may be obtained by parsing a received query, analyzing the parsed query, and generating a set of instructions for processing the query based on the analysis. In some embodiments, the query plan may be generated by a query engine (such as, Figure 1 ) generated by the query engine depicted in .
[0047] At 504, the distributed database identifies nodes of the query plan that contain user functions and secure data access. For example, this can be done as described with respect to the previous figures. In at least one embodiment, the query plan is analyzed and the nodes of the query plan can be marked with attributes to indicate the corresponding risk classification of the nodes. In some embodiments, the identification or marking of the nodes includes storing information indicating the risk classification associated with the nodes. The risk classifications may include user code and secure data access, or other classifications, such as operations that may be associated with denial of service or other types of attacks.
[0048] At 506, the distributed database generates an execution plan in which user code operations are separated from secure data access operations. For example, this can be done as described with respect to the previous figures. In at least one embodiment, the execution plan is generated so that incompatible risk classifications are assigned to separate stages, and the stages are marked according to these risk classifications. The stages can be marked, for example, using attributes such as Secured=true or UserCode=true to indicate the risk classification associated with the execution of a given stage.
[0049] At 508, the distributed database identifies compatible executors for each stage of the execution plan. For example, this can be done as described with respect to the previous figures. Each stage of the execution plan can include one or more operations from the query plan and (optionally) be marked with risk classifications associated with those operations. The distributed database generates the execution plan so that for a given stage, the included operations do not contain incompatible risk classifications, such as user code and secure data access. To identify compatible executors, the distributed database can, for example, match stages that require sandboxing (such as, UserCode=true stages) with executors that have been sandboxed previously in history. Conversely, the distributed database can match stages whose risk classification is incompatible with sandboxing (such as, Secured=true stages) with executors that have not been sandboxed. Note that this does not necessarily prevent a user from using hardware that was previously used to run code with an incompatible risk classification, but at least some embodiments of the system will ensure that such hardware has been hard reset, soft reset, or otherwise secured prior to such reuse.
[0050] At 510, the distributed database executes the query using a compatible executor. In at least one embodiment, this includes the distributed database causing each stage of the execution plan to be executed on the compatible executor to which the stage is assigned. For example, this can be done as described with respect to the previous figures.
[0051] Figure 6 An example process for generating a secure execution plan for a distributed database system according to at least one embodiment is illustrated. Figure 6 The process 600 is depicted as a sequence of steps, but the depicted order should not be interpreted as limiting the scope of the present disclosure to only those embodiments that conform to the depicted order. For example, unless otherwise indicated or clear from the context (e.g., when the output of one step is used as the input of another step), at least some of the depicted steps may be reordered or performed in parallel.
[0052] At 602, the distributed database identifies and marks the user code nodes in the query plan. In at least one embodiment, this may include searching one or more data structures corresponding to the query plan. The data structure may include a tree data structure, which includes nodes linked by edges. The nodes may represent operations to be performed to process the query, and the edges may represent dependencies between the corresponding operations. The search of the data structure may include traversal of the nodes via the edges, and review of the attributes associated with the nodes. Relative to box 602, the distributed database may locate nodes with attributes indicating user code operations, and add additional attributes to indicate that the nodes should be considered associated with the user code risk classification.
[0053] At 604, the distributed database identifies the node that performs the secure data access. Similar to block 602, this may include a search for a query plan. Operations accessing data may be checked to determine if they are accessing secure data. In some embodiments, this may be accomplished by identifying the data being accessed, reviewing the data catalog or scheme, and using information from the data catalog or scheme to determine if the data is secure.
[0054] At 606, the distributed database creates a child node to represent the operation of accessing secure data. This may be done, for example, to separate secure data access from non-secure data access. As depicted at 608, the distributed database marks the child node including secure data access with a corresponding attribute (eg, Secured=true).
[0055] At 610, the distributed database generates an execution plan that separates phases containing user code from phases containing secure data access. For example, this can be done as described with respect to the previous figures.
[0056] Those skilled in the art will appreciate from this disclosure that certain embodiments may be able to achieve certain advantages, including improving the security of a distributed database while still allowing user code and access to secure data, including access to data in multi-user tables.
[0057] Figure 7 An example process 700 illustrates a distributed database executing a query using separation of operations associated with incompatible risk classifications in accordance with at least one embodiment. Figure 6The process 600 is depicted as a sequence of steps, but the depicted order should not be interpreted as limiting the scope of the present disclosure to only those embodiments that conform to the depicted order. For example, unless otherwise indicated or clear from the context (e.g., when the output of one step is used as the input of another step), at least some of the depicted steps may be reordered or performed in parallel.
[0058] At 702, the distributed database identifies a first portion of a query plan that is associated with a first risk classification.
[0059] At 704, the distributed database identifies a second portion of the query plan that is associated with a second risk classification.
[0060] At 706 , the distributed database generates an execution plan in which the first portion and the second portion are executed in separate phases.
[0061] At 708, the distributed database identifies an executor for separately executing the first phase and the second phase.
[0062] At 710 , the distributed database executes the first phase and the second phase on the executor.
[0063] At 712, the distributed database generates results for the query based on the execution of the stages.
[0064] In About Figure 7 In an example embodiment of the described process, a system includes at least one processor and at least one memory for storing computer executable instructions, wherein the computer executable instructions, in response to being executed by the at least one processor, cause the system to perform a Figure 7 The operations depicted. The system identifies instructions in a query plan that indicate execution of a user-defined function, and identifies instructions that indicate access to a database table associated with a security policy. The system then generates an execution plan based at least in part on the identification of these instructions, wherein the instructions associated with the user-defined function are to be executed in a first phase, the first phase being separate from a second phase of the execution plan in which the instructions that indicate access to the database table are to be executed. The database table is associated with a security policy that controls access to at least a portion of the database table. The system then causes the first phase of the execution plan to be executed on a first computing node, and causes the second phase of the execution plan to be executed on a second computing node that is different from the first computing node. The system then provides results for the query based on the execution of these phases.
[0065] In the example, execution of the instructions may further cause the system to reserve the first computing node for use in executing stages that include user-defined functions. This may be done in response to it being used to execute stages that include user-defined functions. The stages may be reserved specifically for stages that include user-defined functions, or alternatively may be reserved for stages that include user-defined functions and stages whose security risks are compatible with execution on an executor that is executing or has executed a user-defined function.
[0066] In the described example, execution of the instructions may further cause the system to select a second computing node to execute a second stage of the execution plan based at least in part on determining that the second computing node has not been used to execute the user-defined function.
[0067] In the example described, a database table may store data representing multiple users. Furthermore, a security policy may restrict access by any one user to a subset of the tables associated with that user. Implementation of this policy may be incompatible with execution of a user-defined function on the same executor, and further may be incompatible with execution on an executor that has previously executed the user-defined function.
[0068] In About Figure 7 In another example of the described process, a computer-implemented method for processing a database query includes identifying a first portion of a query plan associated with a first risk classification, and identifying a second portion of the query plan associated with a second risk classification. These risk classifications are risk classifications that are considered incompatible for simultaneous execution on an executor, and further may be considered incompatible for execution on an executor that has been previously used for at least one of the risk classifications. The example method also includes generating an execution plan in which portions of the query plan are to be executed in separate stages of the execution plan. The method also includes executing the execution plan using at least a first computing node and a second computing node to execute the respective stages, and generating results for the query based on the execution of the stages.
[0069] In another aspect of the example method, the first risk classification is associated with a user-defined function and the second risk classification is associated with a table of a database shared by multiple users.
[0070] In another aspect of the example method, a risk classification is determined based at least in part on a review of a database catalog.
[0071] In another aspect of the example method, the example method further comprises reserving at a compute node for executing stages comprising a user-defined function. The compute node may be reserved specifically for stages comprising executing a user-defined function, or may be reserved for stages comprising executing a user-defined function and other operations whose risk profile is compatible with the execution of the user-defined function.
[0072] In another aspect of the example method, the example method further includes determining that the computing node has not been used to execute the user-defined function, and selecting the computing node to execute a phase including a risk profile that is incompatible with the user-defined function based at least in part on the determination. For example, a phase including accessing a security table may be assigned to an executor computing node that has not been previously used to execute the user-defined function.
[0073] In another aspect of the example method, the example method further includes generating a version of the query plan in which portions of the query plan are labeled according to their respective associations with risk classifications.
[0074] In another aspect of the example method, an execution plan is generated based at least in part on assigning operations to phases according to risk classifications associated with the assigned operations. This can be done, for example, based on attributes marking risk classifications associated with nodes of the query plan.
[0075] Figure 8 Illustrate various aspects of the example system 800 for realizing various aspects according to an embodiment.As will be appreciated, although a web-based system has been used for the purpose of explanation, various embodiments can be realized using different systems as appropriate.In one embodiment, the system includes an electronic client device 802, which includes any suitable device that can be operated to send and / or receive requests, messages or information through a suitable network 804 and transmit information back to the device of the user.The example of such client device includes a personal computer, a cellular phone or other mobile phone, a handheld messaging device, a laptop computer, a tablet computer, a set-top box, a personal data assistant, an embedded computer system, an e-book reader, etc.In one embodiment, the network includes any suitable network, including an intranet, the Internet, a cellular network, a local area network, a satellite network or any other such network and / or a combination thereof, and the components for such a system depend at least in part on the type of the network and / or system selected.Many protocols and components for communicating via such a network are well known, and will not be discussed in detail herein.In one embodiment, the communication carried out by the network is realized by wired and / or wireless connection and a combination thereof. In one embodiment, the network includes the Internet and / or other publicly addressable communications network, as the system includes a web server 806 for receiving requests and serving content in response to the requests, but for other networks, alternative devices serving similar purposes may be used, as will be apparent to one of ordinary skill in the art.
[0076] In one embodiment, the illustrative system includes at least one application server 808 and a distributed database 810, and it should be understood that there may be several application servers, layers or other elements, processes or components that may be linked or otherwise configured that may interact to perform tasks such as obtaining data from an appropriate data store. In at least one embodiment, the distributed database 810 corresponds to the distributed databases described herein with respect to the aforementioned figures. The distributed database 810 may include a plurality of computing nodes 812, 814, and 816. The computing nodes 812, 814, and 816 may, for example, be associated with a plurality of computing nodes 812, 814, and 816 as described with respect to Figure 1 Corresponding to the executor and query engine described in the other aforementioned figures.
[0077] In one embodiment, the server is implemented as a hardware device, a virtual computer system, a programming module executed on a computer system, and / or other devices configured with hardware and / or software to receive and respond to communications (e.g., web service application programming interface (API) requests) over a network. As used herein, unless otherwise specified or clear from the context, the term "data storage area" refers to any device or combination of devices capable of storing, accessing, and retrieving data, which may include any combination and any number of data servers, databases, data storage devices, and data storage media in any standard, distributed, virtual, or clustered system. In one embodiment, the data storage area communicates with a block-level and / or object-level interface. The application server may include any appropriate hardware, software, and firmware that are used to integrate with the data storage area as needed to execute various aspects of one or more applications for client devices, thereby handling some or all of the data access and business logic of the application.
[0078] In one embodiment, the application server cooperates with the data store to provide access control services and generates content including but not limited to text, graphics, audio, video and / or other content, which is provided by the web server to a user associated with a client device in the form of Hypertext Markup Language ("HTML"), Extensible Markup Language ("XML"), JavaScript, Cascading Style Sheets ("CSS"), JavaScript Object Notation (JSON) and / or another appropriate client-side structured language or other structured language. In one embodiment, the content transmitted to the client device is processed by the client device to provide the content in one or more forms, including but not limited to forms that the user can perceive by hearing, vision and / or by other senses. In one embodiment, the handling of all requests and responses and the delivery of content between the client device 802 and the application server 808 are handled by the web server in this example using PHP: Hypertext Preprocessor ("PHP"), Python, Ruby, Perl, Java, HTML, XML, JSON and / or another appropriate server-side structured language. In one embodiment, the operations described herein as being performed by a single device are performed jointly by multiple devices forming a distributed and / or virtual system.
[0079] In one embodiment, the distributed database 810 includes several separate data tables, databases, data documents, dynamic data storage schemes, and / or other data storage mechanisms and media for storing data related to specific aspects of the present disclosure. In one embodiment, the exemplified data storage area includes a mechanism for storing production data and user information, which is used to serve content on the production side. In one embodiment, the data storage area is also shown as including a mechanism for storing log data, which is used for reporting, computing resource management, analysis, or other such purposes. In one embodiment, other aspects such as page image information and access rights information (e.g., access control policies or other permission encodings) are stored in the data storage area, in any one of the mechanisms listed above, or in an additional mechanism in the distributed database 810 as appropriate.
[0080] In one embodiment, the distributed database 810 can be operated by logic associated therewith to receive instructions from the application server 808 and obtain, update or otherwise process data in response to the instructions, and the application server 808 provides static data, dynamic data, or a combination of static data and dynamic data in response to the received instructions. In one embodiment, dynamic data such as data used in web logs (blogs), shopping applications, news services, and other such applications are generated by a server-side structured language as described herein or provided by a content management system ("CMS") operating on the application server or under the control of the application server. In one embodiment, a user submits a search request for a certain type of item through a device operated by the user. In this example, the data store accesses user information to verify the identity of the user, accesses directory details to obtain information about the type of item, and returns the information to the user, such as in the form of a result list on a web page viewed by the user via a browser on the user device 802. Continuing with this example, the information of a specific item of interest is viewed in a dedicated page or window of the browser. However, it should be noted that the embodiments of the present disclosure are not necessarily limited to the context of web pages, but are more generally applicable to processing requests in a general manner, where the request is not necessarily a request for content. Example requests include requests to manage and / or interact with computing resources hosted by system 800 and / or another system (such as to launch, terminate, delete, modify, read, and / or otherwise access such computing resources).
[0081] In one embodiment, each server typically includes an operating system that provides executable program instructions for general administration and operation of the server, and includes a computer-readable storage medium (e.g., a hard disk, random access memory, read-only memory, etc.) that stores instructions that, when executed by a processor of the server, cause or otherwise allow the server to perform its intended functions (e.g., functions are performed as a result of one or more processors of the server executing instructions stored on the computer-readable storage medium).
[0082] In one embodiment, system 800 is a distributed and / or virtualized computing system utilizing several computer systems and components that are interconnected using one or more computer networks via communication links (e.g., Transmission Control Protocol (TCP) connections and / or Transport Layer Security (TLS) or other cryptographically protected communication sessions) or direct connections. However, those of ordinary skill in the art will appreciate that such a system may be implemented on a variety of networks having more than 100 nodes. Figure 8 The system can be operated in a system with fewer or more components than those illustrated in the example. Figure 8 The depiction of system 800 should be regarded as illustrative in nature and not limiting the scope of the present disclosure.
[0083] Various embodiments may further be implemented in a variety of operating environments, which in some cases may include one or more user computers, computing devices, or processing devices that can be used to operate any of a plurality of applications. In one embodiment, the user or client device includes any of the following: a plurality of computers (such as desktop computers, laptop computers, or tablet computers running standard operating systems), and cellular (mobile), wireless, and handheld devices running mobile software and capable of supporting a variety of networking and messaging protocols, and such systems also include multiple workstations running a variety of commercially available operating systems and any of other known applications for purposes such as development and database management. In one embodiment, these devices also include other electronic devices, such as virtual terminals, thin clients, gaming systems, and other devices capable of communicating via a network, as well as virtual devices, such as virtual machines, hypervisors, software containers utilizing operating system-level virtualization, and other virtual or non-virtual devices capable of supporting virtualization that can communicate via a network.
[0084] In one embodiment, the system utilizes at least one network that may be familiar to those skilled in the art to support communication using any of a variety of commercially available protocols, such as the Transmission Control Protocol / Internet Protocol ("TCP / IP"), the User Datagram Protocol ("UDP"), protocols operating in the various layers of the Open Systems Interconnection ("OSI") model, the File Transfer Protocol ("FTP"), the Universal Plug and Play ("UPnP"), the Network File System ("NFS"), the Common Internet File System ("CIFS"), and other protocols. In one embodiment, the network is a local area network, a wide area network, a virtual private network, the Internet, an intranet, an extranet, a public switched telephone network, an infrared network, a wireless network, a satellite network, and any combination thereof. In one embodiment, a connection-oriented protocol is used to communicate between network endpoints, so that a connection-oriented protocol (sometimes referred to as a connection-based protocol) can transmit data in an ordered stream. In one embodiment, a connection-oriented protocol can be reliable or unreliable. For example, the TCP protocol is a reliable connection-oriented protocol. Asynchronous Transfer Mode (ATM) and frame relay are unreliable connection-oriented protocols. Connection-oriented protocols are in contrast to packet-oriented protocols, such as UDP, which transmit packets without guaranteeing ordering.
[0085] In one embodiment, the system utilizes a web server that runs one or more of a variety of server or middle-tier applications, including a Hypertext Transfer Protocol ("HTTP") server, an FTP server, a Common Gateway Interface ("CGI") server, a data server, a Java server, an Apache server, and a business application server. In one embodiment, one or more servers are also capable of executing programs or scripts in response to requests from user devices, such as by executing one or more web applications implemented in any programming language (such as C, C# or C++) or any scripting language (such as Ruby, PHP, Perl, Python or TCL and their combinations). In one embodiment, the one or more servers also include a database server, including but not limited to a database server that can be accessed from and Commercially available servers, as well as open source servers such as MySQL, Postgres, SQLite, MongoDB, and any other server capable of storing, retrieving, and accessing structured or unstructured data. In one embodiment, the database server includes a table-based server, a document-based server, an unstructured server, a relational server, a non-relational server, or a combination of these and / or other database servers.
[0086] In one embodiment, the system includes various data storage areas as discussed above, as well as other memories and storage media, which may reside in a variety of locations, such as on storage media local to (and / or residing in) one or more computers, or on storage media remote from any or all computers on a network. In one embodiment, the information resides in a storage area network (SAN) familiar to those skilled in the art, and similarly, any necessary files for performing functions belonging to a computer, server, or other network device are stored locally or remotely as appropriate. In one embodiment, where the system includes computerized devices, each such device may include hardware elements electrically coupled via a bus, including, for example, at least one central processing unit (CPU or "processor"), at least one input device (e.g., a mouse, keyboard, controller, touch screen, or keypad), at least one output device (e.g., a display device, printer, or speaker), at least one storage device (such as a disk drive, optical storage device, and solid-state storage device, such as a random access memory (RAM) or read-only memory (ROM)), as well as removable media devices, memory cards, flash memory cards, etc., and various combinations thereof.
[0087] In one embodiment, such equipment also includes a computer-readable storage medium reader, a communication device (e.g., a modem, a network card (wireless or wired), an infrared communication device, etc.), and a working memory as described above, wherein the computer-readable storage medium reader is connected to or configured to receive a computer-readable storage medium, and the computer-readable storage medium represents a remote, local, fixed and / or removable storage device and a storage medium for temporarily and / or more permanently containing, storing, transmitting and retrieving computer-readable information. In one embodiment, the system and various devices generally also include a number of software applications, modules, services or other elements located within at least one working memory device, including an operating system and application programs, such as a client application or a web browser. In one embodiment, custom hardware is used, and / or specific elements are implemented in hardware, software (including portable software, such as applets), or both. In one embodiment, connections to other computing devices (such as network input / output devices) are employed.
[0088] In one embodiment, storage media and computer-readable media for containing code or portions of code include any suitable media known or used in the art, including storage media and communication media, such as but not limited to volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing and / or transmitting information (such as computer-readable instructions, data structures, program modules or other data), including RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or any other medium that can be used to store the desired information and can be accessed by system devices. Based on the present disclosure and the teachings provided herein, a person of ordinary skill in the art will understand other ways and / or methods of implementing various embodiments.
[0089] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader spirit and scope of the invention as set forth in the claims.
[0090] Other variations are also within the spirit of the present disclosure. Therefore, although the disclosed technology allows for various modifications and alternative configurations, certain embodiments thereof are illustrated in the drawings and have been described in detail above. However, it should be understood that it is not intended to limit the present invention to one or more specific forms disclosed, but on the contrary, the present invention is intended to cover all modifications, alternative configurations and equivalents that fall within the spirit and scope of the present invention as defined in the appended claims.
[0091] Additionally, embodiments of the present disclosure may be described in terms of the following:
[0092] 1. A system comprising:
[0093] at least one processor;
[0094] at least one memory storing computer executable instructions responsive to
[0095] executed by the at least one processor so that the system:
[0096] identifying a first portion of the query plan, the first portion indicating execution of a user-defined function,
[0097] identifying a second portion of the query plan, the second portion indicating accessing a database table associated with a security policy;
[0098] generating an execution plan based at least in part on the identification of the first portion and the second portion, wherein the first portion is to be executed in a first phase separate from a second phase of the execution plan in which the second portion is to be executed;
[0099] Causing the first stage of the execution plan to be executed on a first computing node;
[0100] causing the second phase of the execution plan to be executed on a second computing node; and
[0101] Results of the query are provided based at least in part on the execution of the first phase and the second phase.
[0102] 2. The system of clause 1, wherein the at least one memory further comprises computer executable instructions that, in response to being executed by the at least one processor, cause the system to:
[0103] In response to using the first computing node to execute the first stage including a user-defined function, the first computing node is reserved for executing other stages in other execution plans including the user-defined function or operations compatible with the user-defined function.
[0104] 3. The system of clause 1 or 2, wherein the at least one memory further comprises computer executable instructions that, in response to being executed by the at least one processor, cause the system to:
[0105] The second computing node is selected to execute the second stage of the execution plan based at least in part on a determination that the second computing node has not been used to execute a user-defined function.
[0106] 4. A system as described in any of clauses 1 to 3, wherein the database table stores data representing multiple users and the security policy restricts access by any one user to a subset of the table associated with that one user.
[0107] 5. A computer-implemented method of processing a query of a database, comprising:
[0108] identifying a first portion of the query plan associated with a first risk classification;
[0109] identifying a second portion of the query plan associated with a second risk classification;
[0110] generating an execution plan, wherein the first portion is to be executed in a first phase separate from a second phase of the execution plan in which the second portion is to be executed;
[0111] executing the execution plan using at least a first computing node to execute the first phase and using a second computing node to execute the second phase; and
[0112] Results of the query are generated based at least in part on the execution of the first phase and the second phase.
[0113] 6. The computer-implemented method of clause 5, wherein the first risk classification is associated with a user-defined function and the second risk classification is associated with a plurality of users sharing a table of a database.
[0114] 7. A computer-implemented method as described in clause 5 or 6, wherein the stage of the execution plan includes one or more operations corresponding to one or more parts of the query plan.
[0115] 8. The computer-implemented method of any of clauses 5 to 7, wherein the risk classification is identified based at least in part on a review of at least one of a database catalog or a schema.
[0116] 9. The computer-implemented method of any of clauses 5 to 8, wherein the first computing node is reserved for executing stages that include the first risk classification and executing stages that are compatible with the first risk classification.
[0117] 10. The computer-implemented method of any one of clauses 5 to 9, further comprising:
[0118] determining that the second computing node has not been and is not being used to perform a phase associated with the first risk classification; and
[0119] The second computing node is selected to perform the second stage based at least in part on the determination.
[0120] 11. The computer-implemented method of any one of clauses 5 to 10, further comprising:
[0121] A version of the query plan is generated in which portions of the query plan are labeled according to their respective associations with risk classifications.
[0122] 12. The computer-implemented method of any of clauses 5 to 11, wherein the execution plan is generated based at least in part on assigning operations to phases according to risk classifications associated with the assigned operations.
[0123] 13. A non-transitory computer-readable storage medium having executable instructions stored thereon, the executable instructions, when executed by one or more processors of a computer system, causing the computer system to at least:
[0124] identifying a first portion of the query plan associated with a first risk classification;
[0125] identifying a second portion of the query plan associated with a second risk classification;
[0126] generating an execution plan in which the first portion is to be executed in a first phase separate from a second phase of the execution plan in which the second portion is to be executed; and
[0127] The execution plan is caused to execute using at least a first computing node to execute the first phase and using a second computing node to execute the second phase.
[0128] 14. The non-transitory computer-readable storage medium of clause 13, further comprising: instructions that, due to execution by the one or more processors, cause the computer system to:
[0129] Portions of the query plan are labeled according to risk classifications associated with the portions.
[0130] 15. The non-transitory computer-readable storage medium of clause 13 or 14, further comprising: instructions that, due to execution by the one or more processors, cause the computer system to:
[0131] identifying a portion of the query plan associated with the second risk classification; and
[0132] The identified portion is modified so that the second risk classification is associated with at least one of a sub-portion or a sibling portion of the identified portion.
[0133] 16. A non-transitory computer-readable storage medium as described in any of clauses 13 to 15, wherein the second risk classification is associated with access to database tables storing data representing multiple users and a security policy that restricts access by any one user to a subset of the tables associated with that one user.
[0134] 17. The non-transitory computer-readable storage medium of any of clauses 13 to 16, wherein the first computing node is reserved for executing stages comprising a user-defined function and stages that are compatible with risks associated with executing the user-defined function.
[0135] 18. The non-transitory computer-readable storage medium of any of clauses 13 to 17, further comprising: instructions that, due to execution by the one or more processors, cause the computer system to:
[0136] A computing node determined to be used to perform a phase associated with the second risk classification has not been used to perform a phase associated with the first risk classification.
[0137] 19. The non-transitory computer-readable storage medium of any of clauses 13 to 18, further comprising: instructions that, due to execution by the one or more processors, cause the computer system to:
[0138] The first computing node is selected to perform the first phase based at least in part on a determination that the first computing node is not currently performing a phase associated with the second risk classification.
[0139] 20. The non-transitory computer-readable storage medium of any of clauses 13 to 19, further comprising: instructions that, due to execution by the one or more processors, cause the computer system to:
[0140] identifying an additional portion of the query plan that is not associated with either the first risk classification or the second risk classification; and
[0141] The additional portion is assigned to a selection phase of the execution plan, the selection being made independent of risk classifications associated with other portions assigned to the selection phase.
[0142] Unless otherwise indicated herein or clearly contradicted by context, the use of the terms "a" and "an" and "the" and similar referents in the context of describing the disclosed embodiments (especially in the context of the appended claims) should be interpreted to cover both the singular and the plural. Similarly, the use of the term "or" should be interpreted to mean "and / or" unless explicitly contradicted or contradicted by context. Unless otherwise indicated, the terms "comprising", "having", "including" and "containing" should be interpreted as open terms (i.e., meaning "including but not limited to") when unmodified and referring to a physical connection. The term "connected" when unmodified and referring to a physical connection should be interpreted as partially or completely contained within, attached to, or linked together, even if there are intervening objects. Unless otherwise indicated herein, the recitation of ranges of values herein is intended to be used only as a shorthand method of referring individually to each individual value falling within the range, and each individual value is incorporated into the specification as if it were individually listed herein. Unless otherwise noted or contradicted by context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set that includes one or more members. In addition, unless otherwise noted or contradicted by context, the term "subset" of a corresponding set does not necessarily mean a true subset of the corresponding set, but a subset and a corresponding set may be equal. Unless explicitly stated otherwise or clear from the context, the use of the phrase "based on" means "based at least in part on" and is not limited to "based only on."
[0143] Unless specifically stated otherwise or otherwise clearly contradicted by context, connective language, such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C" (i.e., the same phrase with or without the Oxford comma) is understood within the context of ordinary usage to mean that the item, term, etc. may be any non-empty subset of the set of A or B or C, A and B and C, or any set that includes at least one A, at least one B, or at least one C that is not contradicted by context or otherwise excluded. For example, in the illustrative example of a set with three members, the connective phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A,B}, {A,C}, {B,C}, {A,B,C}, and, unless expressly contradicted or inconsistent with context, any set having {A}, {B}, and / or {C} as a subset (e.g., a set with multiple "A's"). Thus, such connective language is generally not intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C, respectively. Similarly, phrases such as "at least one of A, B, or C" and "at least one of A, B, or C" refer to any of the following sets: {A}, {B}, {C}, {A,B}, {A,C}, {B,C}, {A,B,C}, unless explicitly stated or clearly indicated by the context to mean something different. In addition, unless otherwise indicated or contradicted by the context, the term "plurality" indicates a state of being in plurality (e.g., "plurality of items" indicates a plurality of items). The number of items in a plurality is at least two, but may be more when explicitly indicated or indicated by the context.
[0144] Unless otherwise noted herein or otherwise clearly contradicted by context, the operations of the processes described herein may be performed in any suitable order. In one embodiment, processes such as the processes described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors. In one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program including a plurality of instructions executable by one or more processors. In one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that does not include a transient signal (e.g., a propagating transient electrical or electromagnetic transmission) but includes a non-transitory data storage circuit (e.g., a buffer, a cache, and a queue) within a transceiver of a transient signal. In one embodiment, the code (e.g., executable code or source code) is stored on one or more non-transitory computer-readable storage medium sets, on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., due to being executed), cause the computer system to perform the operations described herein. In one embodiment, the set of non-transitory computer-readable storage media includes multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the multiple non-transitory computer-readable storage media do not contain all the code, and the multiple non-transitory computer-readable storage media collectively store all the code. In one embodiment, the executable instructions are executed so that different instructions are executed by different processors, for example, in one embodiment, the non-transitory computer-readable storage medium stores instructions, and the main CPU executes some instructions, while the graphics processor unit executes other instructions. In another embodiment, different components of the computer system have separate processors, and different processors execute different subsets of instructions.
[0145] Thus, in one embodiment, the computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such computer system is configured with applicable hardware and / or software that enables the operations to be performed. Additionally, in one embodiment of the present disclosure, the computer system is a single device, and in another embodiment is a distributed computer system comprising multiple devices that operate differently such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.
[0146] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate embodiments of the invention and does not impose limitations on the scope of the invention unless otherwise required. No language in this specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[0147] Embodiments of the present disclosure are described herein, including the best mode known to the inventor for carrying out the present invention. After reading the foregoing description, variations of these embodiments may become apparent to those of ordinary skill in the art. The inventor hopes that the technician will adopt such variations as appropriate, and the inventor intends to practice the embodiments of the present disclosure in a manner different from that specifically described herein. Therefore, the scope of the present disclosure includes all modifications and equivalents of the subject matter recited in the appended claims permitted by applicable law. In addition, unless otherwise noted herein or otherwise clearly contradictory to the context, any combination of all possible variations of the above-mentioned elements is covered by the scope of the present disclosure.
[0148] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
Claims
1. A system comprising: at least one processor; at least one memory storing computer executable instructions that, in response to being executed by the at least one processor, cause the system to: identifying a first portion of the query plan, the first portion indicating execution of a user-defined function, identifying a second portion of the query plan, the second portion indicating accessing a database table associated with a security policy; generating an execution plan based at least in part on the identification of the first portion and the second portion, wherein the first portion is to be executed in a first phase separate from a second phase of the execution plan in which the second portion is to be executed; Causing the first stage of the execution plan to be executed on a first computing node; causing the second phase of the execution plan to be executed on a second computing node; as well as Results of the query are provided based at least in part on the execution of the first phase and the second phase.
2. The system of claim 1, wherein the at least one memory further comprises computer executable instructions that, in response to being executed by the at least one processor, cause the system to: In response to using the first computing node to execute the first stage including a user-defined function, the first computing node is reserved for executing other stages in other execution plans including the user-defined function or operations compatible with the user-defined function.
3. The system of claim 1, wherein the at least one memory further comprises computer executable instructions that, in response to being executed by the at least one processor, cause the system to: The second computing node is selected to execute the second stage of the execution plan based at least in part on a determination that the second computing node has not been used to execute a user-defined function.
4. The system of claim 1, wherein the database table stores data representing a plurality of users and the security policy restricts access by any one user to a subset of the table associated with the one user.
5. A computer-implemented method of processing a query of a database, comprising: identifying a first portion of the query plan associated with a first risk classification; identifying a second portion of the query plan associated with a second risk classification; generating an execution plan, wherein the first portion is to be executed in a first phase separate from a second phase of the execution plan in which the second portion is to be executed; executing the execution plan using at least a first computing node to execute the first stage and using a second computing node to execute the second stage; as well as Results of the query are generated based at least in part on the execution of the first phase and the second phase.
6. The computer-implemented method of claim 5, wherein the first risk classification is associated with a user-defined function and the second risk classification is associated with a plurality of users sharing a table of a database.
7. The computer-implemented method of claim 5, wherein the stages of the execution plan include one or more operations corresponding to one or more portions of the query plan.
8. The computer-implemented method of claim 5, wherein the risk classification is identified based at least in part on a review of at least one of a database catalog or a schema.
9. The computer-implemented method of claim 5, further comprising: A version of the query plan is generated in which portions of the query plan are labeled according to their respective associations with risk classifications.
10. The computer-implemented method of claim 5, wherein the execution plan is generated based at least in part on assigning operations to phases according to risk classifications associated with the assigned operations.
11. A non-transitory computer-readable storage medium having executable instructions stored thereon, the executable instructions, when executed by one or more processors of a computer system, causing the computer system to at least: identifying a first portion of the query plan associated with a first risk classification; identifying a second portion of the query plan associated with a second risk classification; generating an execution plan, wherein the first portion is to be executed in a first phase separate from a second phase of the execution plan in which the second portion is to be executed; as well as The execution plan is caused to execute using at least a first computing node to execute the first phase and using a second computing node to execute the second phase.
12. The non-transitory computer-readable storage medium of claim 11, further comprising: instructions, which, when executed by the one or more processors, cause the computer system to: identifying a portion of the query plan associated with the second risk classification; as well as The identified portion is modified so that the second risk classification is associated with at least one of a sub-portion or a sibling portion of the identified portion.
13. The non-transitory computer-readable storage medium of claim 11, wherein the first computing node is reserved for executing stages that include a user-defined function and stages that are compatible with risks associated with executing the user-defined function.
14. The non-transitory computer-readable storage medium of claim 11, further comprising: instructions, which, when executed by the one or more processors, cause the computer system to: The first computing node is selected to perform the first phase based at least in part on a determination that the first computing node is not currently performing a phase associated with the second risk classification.
15. The non-transitory computer-readable storage medium of claim 11, further comprising: instructions, which, when executed by the one or more processors, cause the computer system to: identifying an additional portion of the query plan that is not associated with either the first risk classification or the second risk classification; as well as The additional portion is assigned to a selection phase of the execution plan, the selection being made independent of risk classifications associated with other portions assigned to the selection phase.