Techniques for Deterministic Distributed Caching to Accelerate SQL Queries

JP2024521730A5Active Publication Date: 2025-05-19ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023571975
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-21
Filing Date
2022-05-20
Publication Date
2025-05-19
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

Cloud-based data processing systems face inefficiencies in retrieving data segments for interactive queries due to slow retrieval from global storage, necessitating improved caching strategies within distributed computing clusters.

Method used

Implementing a consistent distributed cache using consistent hashing techniques to assign data segments to specific nodes, ensuring efficient data retrieval by maintaining token boundaries and cache stability across cluster changes, and optimizing query execution through preferred node selection and cache management.

Benefits of technology

Enhances query execution speed and efficiency by reducing the need for data retrieval from remote storage, maintaining even load distribution, and ensuring fast access to cached data segments, thereby improving overall performance in distributed computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Techniques are disclosed for providing an improved distributed cache. The distributed computing system can be implemented in a cluster including a plurality of worker nodes configured to host one or more executors for processing data related to a query. The worker nodes can host a cache accessible to the executors. The data can be processed as a plurality of data segments. The worker nodes can be uniformly assigned a plurality of token boundaries that define ranges of integer token values. A hashing algorithm can be used to compute a token for each data segment associated with a query. Tasks can be preferentially launched on the executors such that tasks processing data segments include tokens within the token boundaries associated with the preferred executors. The executors can review associated caches to identify outlier data segments and direct other nodes in the cluster to notify.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a non-provisional application claiming priority to and benefit of Indian Provisional Patent Application No. 202141022725, filed on May 21, 2021, and entitled "TECHNIQUES FOR DETERMINISTIC DISTRIBUTED CACHE FOR ACCELERATORY SQL QUERIES," the entire contents of which are incorporated herein by reference for all purposes.

[0002] Technical Field The present disclosure relates to cloud computing data analytics. More particularly, the present disclosure is directed to maintaining a distributed cache of data within a computing cluster and executing computing tasks within the cluster that correspond to the location of the cached data. [Background technology]

[0003] background Cloud-based services provide a solution for processing large amounts of data by any number of tenants. Cloud-based data processing services may be implemented by a distributed computing system including a suitable number of computing clusters including nodes that perform operations to handle user requests. High-performance data analysis services can process distributed data across clusters based on interactive, near real-time queries from users, but rely on slow retrieval of data segments from a global storage device. There is a need to improve caching of data segments on cluster nodes to improve the speed and efficiency of data processing for interactive queries. Summary of the Invention

[0004] Quick Overview Embodiments of the present disclosure relate to providing a cache in a cloud computing environment. More particularly, some embodiments provide methods and systems for implementing a distributed cache that is consistent among nodes in a distributed computing environment. Caching can be provided in part using a consistent hashing technique that allows for the configuration of multiple nodes in the distributed computing environment, each with an associated cache. One or more tasks executing on a node can share an associated cache, but have ownership of a unique data segment stored in the cache. Ownership can be based on a consistent hashing technique that calculates a token for each data segment associated with a query processed by the distributed computing system. The correspondence between the task, node, and token for each data segment can enable a particular node to be identified as a preferred node for execution of a task, as determined by a query optimizer. [Means for solving the problem]

[0005] One embodiment is directed to a method performed by a distributed computing system that provides analytical data processing services. The method can include implementing a cluster including a plurality of nodes. The distributed computing system can maintain a cluster state including a plurality of token boundaries uniformly associated with the plurality of nodes. Maintaining the cluster state can include storing a mapping of one or more execution units to the plurality of nodes and assigning the plurality of token boundaries to one or more execution units executing on the plurality of nodes based on the mapping. The method also includes a driver node of the plurality of nodes receiving a query for execution. Based on the query, a set of one or more data segments corresponding to the query can be identified or determined. The set of tokens corresponding to the set of one or more data segments can be computed using a hashing algorithm to determine a hash value that uniquely identifies the one or more data segments. The method also includes initiating a first task on a first execution unit executing on a first worker node of the plurality of nodes to process a first data segment from the set of one or more data segments, the first worker node being selected based at least in part on a first token of the set of tokens corresponding to a first pair of token boundaries associated with the first worker node. Finally, the first worker node can obtain the data segment to be processed by the first task. The data segment can reside in a cache associated with the first worker node or in a data store, database, object storage, or other repository remote to the node.

[0006] According to certain embodiments, a distributed computing system may receive an indication that a cluster has changed. For example, a hardware or software failure may cause a node to become inoperable. The method may include maintaining a cluster state by updating a plurality of token boundaries and uniformly allocating the updated token boundaries to one or more executions executing within the distributed computing system.

[0007] According to some other embodiments, the method can include obtaining a first data segment present in a second cache associated with the second worker node. In these embodiments, the first execution unit can determine that the first data segment does not exist in the first cache associated with the first worker node. The first execution unit can send a request to one or more neighboring nodes of the plurality of nodes. The execution unit executing on the neighboring node can inspect a cache associated with each neighboring node based on the request. The second execution unit executing on a second worker node of the one or more neighboring nodes can determine that the first data segment exists in the second cache, place the data segment in a block manager or other module configured to handle data transfer between nodes in a distributed computing system, and then send an ID of the first data segment to the first execution unit that made the request. The first worker node can then copy the data segment from the second worker node to the first cache associated with the first worker node.

[0008] In some embodiments, the method may include calculating a plurality of token boundaries using a hashing algorithm. The output of the calculation may be an integer value calculated as a hash from identifying information from each data segment associated with the received query, e.g., the file name of the data segment. The hash value of each data segment may be a token corresponding to that data segment. The plurality of token boundaries may include a plurality of integer values ​​that divide a range of possible integer token values. Each pair of token boundaries thus defines a range of consecutive integer values ​​that lie between the values ​​of the pair of token boundaries.

[0009] In some other embodiments, the method may include operations for stabilizing a distributed cache in a distributed computing system. The method may include sending a housekeeping request to a first worker node to prompt a first execution unit to determine one or more outlier data segments present in a first cache associated with the first worker node. The outlier data segments in the cache are determined in part by whether a token associated with the segment is within a token boundary assigned to the first execution unit. The first execution unit may then send an identifier of the outlier data segment to all other nodes. The second worker node may then copy the one or more outlier data segments from the first cache of the first worker node to a second cache of the second worker node.

[0010] In yet another embodiment, the method can include an act of deleting data segments from the distributed cache to save space and other resources. The method can include sending a housekeeping request to a first worker node to prompt a first execution unit to determine a set of valid data segments and one or more invalid data segments present in the first cache. The one or more invalid data segments can be deleted from the first cache. The method can also include determining a current storage availability associated with the first worker node. If the current storage availability is less than a threshold, the first execution unit can determine one or more target data segments present in the first cache based on the set of segment temperatures. The one or more target data segments can then be deleted from the cache.

[0011] Another embodiment is directed to a distributed computing system including one or more processors and one or more memories storing computer-executable instructions that, when executed on the one or more processors, cause the distributed computing system to execute a cluster including a plurality of nodes. The distributed computing system can maintain a cluster state including a plurality of token boundaries uniformly associated with the plurality of nodes. Maintaining the cluster state can include storing a mapping of one or more executions to the plurality of nodes and assigning the plurality of token boundaries to one or more executions executing on the plurality of nodes based on the mapping. The instructions can also cause the distributed computing system to have a driver node of the plurality of nodes receive a query for execution. Based on the query, a set of one or more data segments corresponding to the query can be identified or determined. The set of tokens corresponding to the set of one or more data segments can be calculated using a hashing algorithm to determine a hash value that uniquely identifies the one or more data segments. The instructions may also cause the distributed computing system to launch, on a first execution unit executing on a first worker node of the plurality of nodes, a first task to process a first data segment from the set of one or more data segments, where the first worker node is selected based at least in part on a first token of the set of tokens that corresponds to a first pair of token boundaries associated with the first worker node. Finally, the first worker node may obtain the data segment to be processed by the first task. The data segment may reside in a cache associated with the first worker node or in a data store, database, object storage, or other repository remote to the node.

[0012] According to certain embodiments, a distributed computing system may receive an indication that a cluster has changed (e.g., a hardware or software failure may cause a node to become inoperable), and the instructions may cause the distributed computing system to maintain the cluster state by updating a number of token boundaries and uniformly allocating the updated token boundaries to one or more executions executing within the distributed computing system.

[0013] According to some other embodiments, the instructions may also cause the distributed computing system to obtain the first data segment present in a second cache associated with the second worker node. In these embodiments, the first execution unit may determine that the first data segment does not exist in the first cache associated with the first worker node. The first execution unit may send a request to one or more neighboring nodes of the plurality of nodes. The execution unit executing on the neighboring node may check a cache associated with each neighboring node based on the request. The second execution unit executing on a second worker node of the one or more neighboring nodes may determine that the first data segment exists in the second cache, place the data segment in a block manager or other module configured to handle data transfers between nodes in the distributed computing system, and then send an ID of the first data segment to the first execution unit that made the request. The first worker node may then copy the data segment from the second worker node to the first cache associated with the first worker node.

[0014] Another embodiment is directed to a non-transitory computer-readable medium storing computer-executable instructions that, when executed by one or more processors, cause a computer system to execute a cluster including a plurality of nodes, maintain a cluster state including a plurality of token boundaries uniformly associated with the plurality of nodes, receive a query for execution by a driver node of the plurality of nodes, determine a set of one or more data segments corresponding to the query based on the query, calculate a set of tokens corresponding to the set of one or more data segments using a hashing algorithm to determine a hash value that uniquely identifies the one or more data segments, launch a first task on a first execution unit executing on a first worker node of the plurality of nodes to process a first data segment from the set of one or more data segments, the first worker node being selected based at least in part on a first token of the set of tokens that corresponds to a first pair of token boundaries associated with the first worker node, and further cause the computer system to retrieve the data segment to be processed by the first task. The data segment can reside in a cache associated with the first worker node or in a data store, database, object storage, or other repository remote from the node. Maintaining the cluster state may include storing a mapping of one or more execution units to multiple nodes and assigning multiple token boundaries to one or more execution units executing on the multiple nodes based on the mapping.

[0015] According to certain embodiments, a distributed computing system may receive an indication that a cluster has changed. For example, a hardware or software failure may cause a node to become inoperable. The instructions may also cause the computer system to maintain the cluster state by updating a number of token boundaries and uniformly allocating the updated token boundaries to one or more executing units within the distributed computing system.

[0016] According to some other embodiments, the instructions may also cause the computer system to obtain the first data segment present in a second cache associated with the second worker node. In these embodiments, the first execution unit may determine that the first data segment does not exist in the first cache associated with the first worker node. The first execution unit may send a request to one or more neighboring nodes of the plurality of nodes. The execution unit executing on the neighboring node may check a cache associated with each neighboring node based on the request. The second execution unit executing on a second worker node of the one or more neighboring nodes may determine that the first data segment exists in the second cache, place the data segment in a block manager or other module configured to handle data transfers between nodes in a distributed computing system, and then send an ID of the first data segment to the first execution unit that made the request. The first worker node may then copy the data segment from the second worker node to the first cache associated with the first worker node. [Brief description of the drawings]

[0017] [Figure 1] FIG. 1 illustrates a distributed computing system in a cloud computing environment including a distributed cache of data segments retrieved from an object storage system, according to some embodiments. [Diagram 2]FIG. 1 illustrates a cluster of computing nodes in a distributed computing system that implements deterministic caching techniques, according to some embodiments. [Figure 3A] FIG. 11 is a code fragment illustrating an example mapping of token boundaries to nodes in a cluster according to some embodiments. [Figure 3B] FIG. 13 is another fragment of code illustrating another example of mapping token boundaries to a set of updated nodes in a cluster according to some embodiments. [Figure 4] 1 is a simplified flowchart of an example process for processing queries on a cluster using a deterministic cache, according to some embodiments. [Diagram 5] 1 is another simplified flowchart of an exemplary process for stabilizing a deterministic cache in a cluster when cluster conditions change, according to some embodiments. [Figure 6] 11 is yet another simplified flowchart of an example process for cleaning a deterministic cache according to a timer process. [Figure 7] FIG. 1 is a block diagram illustrating one pattern for implementing a cloud infrastructure as a service system, according to at least one embodiment. [Figure 8] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 9] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 10] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure as a service system in accordance with at least one embodiment. [Figure 11] FIG. 1 is a block diagram illustrating an example computer system in accordance with at least one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] Detailed Description In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of particular embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and descriptions are not intended to be limiting. The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments or designs.

[0019] Distributed computing systems have become increasingly common in data analytics, with the ability to provide fast, reliable, and scalable solutions for processing large amounts of data. Providing distributed computing systems in a cloud computing environment provides these data processing capabilities to multiple different customers (e.g., tenants) of the cloud computing environment. Recent developments in the distributed computing space (e.g., robust query optimization) allow interactive queries over large data sets to be processed in near real-time, with requesting users launching queries and interactively expecting results in a user interface. To ensure independent scalability, data storage and computing resources are separated in cloud computing environments. This separation can result in less than optimal performance when data is retrieved from storage for processing. The technology disclosed herein is directed to methods, systems, and computer-readable storage media for providing a deterministic distributed cache in a distributed computing system to improve performance of analytical data processing.

[0020] A distributed computing system may include a computing cluster of connected nodes (e.g., computers, servers, virtual machines, etc.) that cooperate in a coordinated manner to handle various requests (e.g., storage and retrieval of data in a system that maintains a database, queries, etc.) by any suitable number of tenants. As used herein, a "computing node" (also referred to as a "worker node" and / or simply a "node") may include a server, computing device, virtual machine, or any suitable physical or virtual computing resource configured to perform operations as part of a computing cluster. For example, a computing cluster may include one or more master nodes (also referred to herein as a "driver node") and one or more worker nodes. In some embodiments, the driver node may perform any suitable operation related to task assignment, load balancing, node provisioning, node removal corresponding to one or more worker nodes, or any suitable operation corresponding to management of a computing cluster. The worker nodes may be configured to perform operations corresponding to tasks assigned to the worker nodes by one or more driver nodes. For example, the worker nodes may perform data storage tasks and / or data retrieval tasks associated with a database as part of tasks assigned to the worker nodes by the driver node.

[0021] In a distributed computing system that provides analytical data processing services (e.g., online analytical processing (OLAP) services), data may be stored in object storage data stores, databases, or similar "deep" or "offline" storage. Typically, object storage systems support infrequent access to data and include various non-volatile memories, which may result in delays in operations to read and / or write data to the object storage. Data may be stored as segments of larger data structures that represent the data multidimensionally. A data segment may be a compressed file that includes a dictionary. A data segment may represent the smallest unit of data that is atomically handled and processed by a data processing service. To execute a query, a data segment may be retrieved from the object storage and stored in a local file system or other memory associated with one or more nodes of the distributed computing system. Local copies of the data segments may constitute a distributed cache in the distributed computing system. Because each node in the system may have a cache that stores a different set of data segments, the techniques of this disclosure provide techniques for, among other things, consistent caching of data segments, stabilizing the cache in response to changes to the cluster, and properly handling replication.

[0022] In some embodiments, a driver module (e.g., an Apache Spark® Thrift server) may run on a driver or master node of a cluster in a distributed computing system. The driver module may convert user queries (e.g., SQL queries) into smaller units of execution called tasks. The tasks may be configured to be executed by one or more execution processes (also called “executors”) running on one or more worker nodes in the cluster. Each task may correspond to the processing of one data segment by the executor. The driver module may optimize the execution of queries by, among other things, launching tasks on the preferred worker node of the executor that may “own” the data segment to which the task corresponds (e.g., has an associated cache that stores a cached copy). In other words, a particular data segment may be cached on a particular worker node due to previous query executions in the cluster (e.g., previous queries from a user in an interactive session). Launching a task on a preferred worker node may retrieve the data segment from the associated cache.

[0023] In some embodiments, a driver node (via a driver module) can assign a set of token boundaries to one or more worker nodes in a cluster. The token boundaries can represent a division of a set of integer values ​​that represent the token division of the set. For example, the integer values ​​can range from 0 to 2. 31The token boundary may be an integer between -1 and 1 (i.e., the size of a 32-bit signed integer). The token boundary may be determined at least in part based on the state of the cluster, such as the number of worker nodes and the number of executors configured to run on the worker nodes. For example, for a cluster with three worker nodes hosting two executors each, the token boundary may divide the set of integer values ​​into six contiguous subsets of token partitions. The driver node may assign token boundaries to the worker nodes according to the number of executors configured on each worker node. When the state of the cluster changes (e.g., a node is added or removed), the driver node may determine an updated token boundary and assign the updated token boundary to the worker nodes according to the change in cluster state. In some other embodiments, the driver node may receive an indication that a worker node has failed and assign the token boundary previously assigned to that worker node to one or more worker nodes remaining in the cluster. Unlike conventional consistent hashing techniques that may use a ring-like structure to assign a token partition to the next node in the ring when a node fails, the techniques of the present disclosure may attempt to distribute the token boundary uniformly to the worker nodes and maintain that uniform distribution throughout the lifetime of the cluster.

[0024] To schedule a task within a node of a cluster, in some embodiments, a driver node may compute a token for the data segment to be processed in response to a query. The token may be a hash value computed based on a unique segment key or segment identifier (e.g., file name, etc.). The hash value may be mapped to a particular worker node by searching for a value within an assigned token boundary. The corresponding worker node may be a preferred node for the task, so that a scheduler process responsible for launching the task may preferentially place the task on the preferred node. An executor processing the task may retrieve the data segment from a cache associated with the worker node, or, if the data segment is not present in the cache, retrieve the data segment from another node or object store. Once the data segment is in the cache of the preferred worker node, a subsequent task processing the data segment may be launched on the preferred worker node, increasing the likelihood of the data segment resulting in a cache hit.

[0025] According to some embodiments, caches associated with one or more worker nodes in a cluster may be maintained consistently using various housekeeping techniques. Housekeeping methods may include both active and passive techniques for updating one or more caches. In response to any change in the state of the cluster (e.g., addition or removal of a node, failure of a node, addition or removal of an execution unit, etc.), a driver node in the cluster may send instructions to all execution units in the cluster to scan the caches associated with the worker nodes hosting the execution units. The driver node may assign token boundaries to the worker nodes that are updated when the state of the cluster changes, so that data segments stored in the cache may correspond to different worker nodes. In response to housekeeping requests, the execution units identify these outlier data segments and pass them to a block manager or other process configured to transfer data between worker nodes, which sends identifiers of the outlier data segments to all other execution units in the cluster. The execution units may then retrieve the outlier data segments and store them in their local caches.

[0026] In some other embodiments, the housekeeping method may include operations occurring based on a timer. As changes to data segments in the object store are committed to the object store, over time, cached data segments may become invalid. Additionally, memory resources allocated to each worker node may be limited. Based on the timer, the driver node may send instructions to executors in the cluster to, among other operations, remove invalid data segments from its cache, forward and remove outlier data segments, and remove infrequently accessed data segments from the cache to maintain free space in the cache.

[0027] Maintaining a deterministic distributed cache offers many advantages over conventional techniques. As briefly discussed above, typical consistent hashing techniques may provide poor load balancing in the event of node failure. Tokens that map to a failed node may be remapped to a single node without evaluating the single node's current load. The techniques described herein attempt to maintain a uniform distribution of tokens among all executions in a cluster as cluster state changes. In this way, query execution speed is maintained by not overloading a single node when a node is lost. Assigning token boundaries to worker nodes to create preferred nodes for task execution may limit overlap of cached data segments in a cluster, since node preference may help launch repeated tasks with the same data segment on a preferred node rather than other nodes in the cluster that may need to copy the data segment from the preferred node's cache or object store. Additionally, executions in a cluster may communicate with other executions to see if adjacent nodes have the data segment cached before retrieving the segment from the object store, thereby improving the speed and efficiency of query execution by reducing the number of retrievals from deep store.

[0028] FIG. 1 illustrates a distributed computing system 110 in a cloud computing environment 100, including a distributed cache of data segments retrieved from an object storage system, according to some embodiments. The distributed computing system 110 may be implemented by one or more computing systems executing computer-readable instructions (e.g., code, programs) to implement the distributed computing system. As illustrated in FIG. 1, the distributed computing system 110 includes various systems, including a load balancer 112, a gateway 114 (e.g., a multi-tenant gateway), an application programming interface (API) server 116, and one or more computing cluster(s) 122. Portions of data or information used by or generated by the system illustrated in FIG. 1 may be stored in an object storage system 156. The system illustrated in FIG. 1 may be implemented using software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the computing systems, hardware, or combinations thereof. The software may be stored in a non-transitory storage medium (e.g., memory device).

[0029] The distributed computing system 110 can be implemented in a variety of different configurations. In the embodiment shown in FIG. 1, the distributed computing system 110 can be implemented on one or more servers of a cloud provider network and can provide its data processing and data analysis services to subscribers of the cloud service on a subscription basis. The computing environment 100 including the distributed computing system 110 shown in FIG. 1 is merely an example and is not intended to unduly limit the scope of the claimed embodiments. Those skilled in the art will recognize many possible variations, alternatives, and modifications. For example, in some implementations, the distributed computing system 110 can be implemented using more or fewer systems than those shown in FIG. 1, may combine two or more systems, or may have a different configuration or arrangement of systems.

[0030] In some embodiments, a computing cluster (e.g., cluster 122) may represent a distributed computing engine for processing and analyzing large amounts of data for tenants or customers of distributed computing system 110. Different computing clusters may be associated with one tenant. For example, in the embodiment shown in FIG. 1, computing cluster(s) 122 are associated with tenant 120 of the distributed computing system. One or more different computing clusters may be associated with additional tenants (e.g., tenants 118, 119). A cluster may be configured to utilize any suitable number of computing nodes to perform operations in a coordinated manner. As previously mentioned, a "computing node" (also referred to herein as a "node") may include a server, computing device, virtual machine, or any suitable physical or virtual computing resource configured to perform operations as part of a computing cluster. By way of example, cluster 122 may include multiple nodes including a driver node 124 and one or more worker node(s) 128, both of which are examples of computing nodes. In some embodiments, driver node 124 performs any suitable operation related to task allocation corresponding to worker nodes, such as load balancing, node provisioning, node removal, or any suitable operation corresponding to managing cluster 122. One or more worker node(s) 128 are configured to perform operations corresponding to tasks assigned to the worker node(s) 128 by the driver node 124. As a non-limiting example, the worker node(s) 128 may perform data storage and / or data retrieval tasks associated with a storage system / database at the direction of the driver node 124, which assigns particular storage or retrieval tasks to the worker node(s) 128.

[0031] Resources allocated to tenants of a distributed computing system are selected from a plurality of cloud-based resources arranged hierarchically. For example, as shown in FIG. 1, the resources may be a pool of cloud infrastructure services 140. Cloud infrastructure services 140 may include a key-value database as a service (KaaS) 142, a workflow as a service (WFaaS) 144, a Kubernetes engine (KE) 146, an identity and access management (IAM) service 148, a virtual cloud network (VCN) 150, and a data catalog 152.

[0032] According to some embodiments, a driver node (e.g., driver node 124) in a computing cluster (e.g., cluster 122) may be configured to execute a driver program (also called a driver module, driver process, or simply a driver, i.e., a process that executes an application that is being built on the computing cluster) and may perform operations to create a context for the application. An "application" may refer to a complete executable driver program that runs as an independent process and is coordinated by the context of the application in the driver program running on the driver node 124. The context of the application may be connected to a cluster manager 126 that allocates system resources to all nodes in the cluster. Each worker node (e.g., worker node(s) 128) in the cluster 122 may be managed by one or more executors, which may be processes (execution engines) that are launched on the worker node(s) 128 to execute operations corresponding to tasks that are assigned to the node. Application code is sent from the driver program to the executor, which specifies the context and various tasks to execute. The executor communicates with the driver program for data sharing or interaction. The executor may also perform operations related to computing and storage and caching of data on the node. In a particular implementation, the computing cluster may be implemented using a distributed computing engine (e.g., Apache Spark), and the cluster manager 126 in the computing cluster may be implemented using a container orchestration platform such as Kubernetes, or another cluster management solution such as Mesos or YARN.

[0033] In some embodiments, portions of data or information used or generated by the computing clusters shown in FIG. 1 may be stored in one or more storage systems of the distributed computing system 110. In the embodiment shown in FIG. 1, the storage systems include an object storage system 156. The object storage system 156 may represent a deep or offline storage system for storing data used, analyzed, and processed by different computing clusters associated with different tenants of the distributed computing system 110. As an example, the object storage system 156 may represent a type of storage system that uses object-based storage devices to store data by managing and manipulating the data as separate units called objects. The object storage system 156 may be used for offline storage and / or archiving of data (e.g., used for backup or long-term storage where the data is accessed infrequently).

[0034] In certain embodiments, data processed by a computing cluster may be represented and stored in an object storage system as a "data cube." A "data cube" may refer to a data structure that can be used to represent data along some measure of interest, such as a two-dimensional, three-dimensional, or higher-dimensional representation. A data cube can store large amounts of data while providing users with searchable access to any data point and can execute queries to provide real-time results. In certain examples, at runtime, a computing cluster may cache a portion or the entire data cube index as one or more data segments in a cache memory (e.g., cache 130) of the computing nodes forming the computing cluster, or may store the segments in a connected secondary storage device (e.g., random access memory (RAM), solid-state drive (SSD), or hard disk drive (HDD)) associated with the computing node. As used herein, a "data segment" may refer to an individual dimension of a data cube that can be filtered and analyzed to provide detailed results to customers / tenants of the distributed computing system 110.

[0035] In some embodiments, the distributed computing system 110 may be configured to determine placement of data segments on various computing nodes that make up the cluster. Placing the data segments on a particular computing node in the cluster allows for faster retrieval of the data segments when a particular data segment needs to be read at a node of the computing cluster. The distributed computing system 110 may use various approaches to determine placement of data segments on various computing nodes that make up the cluster. For example, in the approach described herein, the distributed computing system 110 may utilize a particular hashing strategy to identify nodes that can process a particular set of data segments. When a query is submitted by a user of the distributed computing system 110 (e.g., customer(s) 102), a driver node (e.g., driver node 124) in a cluster (e.g., cluster 122) may be configured to identify (e.g., using a consistent hashing strategy) a worker node (e.g., worker node 128) that stores the data segment, which can then be used to execute the query and then send the query to the worker node for execution.

[0036] In a particular embodiment, a user (e.g., a customer 102) may interact with the distributed computing system 110 via a computing device 104 that is communicatively coupled to the distributed computing system 110, possibly via a public network 108 (e.g., the Internet). The computing device 104 may be of various types, including, but not limited to, a mobile phone, a tablet, a desktop computer, and the like. The user may interact with the cloud computing system using a console user interface (UI) (which may be a graphical user interface (GUI)) of an application executed by the computing device, or via API operations provided by the distributed computing system 110. For example, a user may interact with the distributed computing system 110 to create one or more computing clusters, perform interactive queries on existing data stored in a storage system, and obtain results as a result of query processing.

[0037] As an example, a user associated with a tenant 120 of the distributed computing system 110 may interact with the distributed computing system 110 by sending a request to the distributed computing system 110 to create one or more clusters, including cluster 122. The cluster creation request may be received by a load balancer 112 in the distributed computing system 110, which may send the request to a multi-tenant proxy service in the distributed computing system, for example, a gateway 114. The gateway 114 may be responsible for authenticating / authorizing the user's request and routing the request to an API server 116, which may be configured to perform operations to create a computing cluster. In a particular example, the gateway 114 may represent a shared multi-tenant Hypertext Transfer Protocol (HTTP) proxy service that authorizes the user and sends the user's request to the API server 116 to enable the creation of a computing cluster for the tenant. In a particular example, as previously described, the creation of a cluster may include the creation of a pool of nodes, including a set of driver nodes and worker nodes. One or more clusters may be created under a dedicated subnet for the tenant. Thus, the distributed computing system 110 includes functionality to provide isolation between computing clusters belonging to different tenants, for example, tenants 118, 119.

[0038] To accelerate query processing and execution within the cluster, the cluster 122 within the distributed computing system 110 may include the capability to cache data required for query computation within memory (also referred to herein as “cache memory” or “cache”) within different nodes of the cluster. For example, a cache (e.g., cache 130) may represent a small amount of very fast and expensive dynamic random access memory (DRAM) located near the central processing unit of the computing node. In certain embodiments, nodes within the cluster are provided with improved capabilities to perform efficient processing and analysis of data within the cluster by retrieving cached data, including, for example, cached data segments, from a cache associated with the node. In particular, when a worker node within the cluster is assigned a task for execution and the associated data segment is not in the cache, the worker node may determine an adjacent node within the cluster that has the data segment in its cache and retrieve the segment from the adjacent node instead of retrieving the data segment from the object store 156. Additional details of the operations performed by the nodes to retrieve and maintain the data segments in the associated cache are described in more detail in FIG. 2.

[0039] FIG. 2 illustrates a cluster 202 of computing nodes in a distributed computing system 200 implementing deterministic caching techniques, according to some embodiments. The cluster 202 may be similar to the cluster 122 described above with reference to FIG. 1. The cluster 202 may include multiple nodes, including a driver node 204 and worker nodes 206, 208. The nodes in the cluster 202 may host any suitable number of execution units (e.g., execution engines or execution processes) configured to perform operations corresponding to one or more tasks assigned to the node. As illustrated in FIG. 2, the tasks may include tasks 214-226 that are assigned to execution units 210-213. The worker node 206 utilizes two execution units 210, 211, while the worker node 208 utilizes two execution units 212, 213. According to various embodiments, more or fewer nodes with more or fewer execution units are possible. Each execution unit may further comprise an execution unit memory, which may be part of a cache associated with the worker node. For example, the execution unit 210 may have a portion of the cache 228 as its execution unit memory, while the execution unit 211 may have a second portion of the cache 228 as its execution unit memory. Similarly, the execution units 212, 213 running on the worker node 208 may utilize the cache 230. The execution units 210-213 may use cache memories, such as the caches 228, 230, to store data segments used in the tasks 214-226. The caches 228, 230 may be part of a larger memory resource provided to the worker nodes 206, 208 to support additional data operations (e.g., splitting and shuffling data) beyond the dedicated memory resources of each execution unit.

[0040] As previously mentioned, in some embodiments, a driver node (e.g., driver node 204) may be configured to execute a driver 232 (also referred to as a driver program, driver module, or driver process; i.e., a process that executes an application that is built on a computing cluster) and may perform operations to create a context for the application. The application may be an analytical data processing application, such that the driver 232 may include a query optimizer 234 (e.g., Apache Spark Catalyst) for optimizing a query 250. The driver 204 may include a cluster manager 236 that allocates system resources across all nodes in the cluster. The cluster manager 236 may be similar to the cluster manager 126 described with respect to FIG. 1. The cluster manager 236 may be configured to maintain a state of the cluster. For example, the cluster state may include, but is not limited to, a mapping of worker nodes to executors and a mapping of token boundaries to worker nodes. The cluster manager 236 may update the cluster state in response to an indication that the cluster state has changed, such as when a worker node fails. The cluster manager 236 may also update the cluster state when additional executors and / or worker nodes are provisioned into the cluster. The cluster manager 236 can communicate with an external cluster service, such as a Kubernetes service (e.g., the Kubernetes engine 146 of FIG. 1), to maintain the cluster state.

[0041] In some embodiments, the driver 204 may also include a task manager 238, which may be a scheduler for launching one or more tasks, e.g., tasks 214-226, on the executors in the cluster. The scheduling of the tasks may be based on query optimizations provided by the query optimizer 234. The tasks may be preferentially launched on the executors based on consistent hashing techniques described herein, such that the tasks are launched on worker nodes that have caches containing copies of the corresponding data segments. As an example, in response to a query 250, a task 214 may be scheduled on an executor 210 running on a worker node 206 based at least in part on identifying that the data segment corresponding to the task 214 has a hash token value that is within a token boundary assigned to the worker node 206. When the executor 210 fetches the data segment, it is likely that the data segment is present in the local cache 228, resulting in a faster retrieval than if the executor 210 had to fetch the data segment from the object store 256.

[0042] In some embodiments, the driver 204 may include a cache manager 240. The cache manager 240 may be configured to communicate with cache managers 242-248 running on the executors 210-213. The cache managers may be configured to implement a remote procedure call (RPC) endpoint that passes requests between cache managers. For example, during execution of a query 250, the cache manager 246 may invoke a request to the executors 210, 211 running on the neighboring worker node 206 to check the local cache 228 to identify the presence of a data segment required by either the task 220 or the task 222 on the executor 212. The invocation of the request may be based at least in part on the cache manager 246 determining that a hash token value of the data segment is within a token boundary assigned to the worker node 206. In response to the RPC request, the cache manager 242 may locate the data segment in the cache 228, place the data segment in a block manager (e.g., a block storage device or an inter-node storage system in a distributed computing system), and return a block ID to the requesting cache manager 246. The execution unit 212 may then obtain the data segment from the block manager and copy the data segment to the cache 230 .

[0043] In some embodiments, cache manager 240 may communicate with cache managers 242-248 to perform one or more housekeeping operations on caches 228, 230. Housekeeping operations may include active operations to synchronize data stored in the caches. Active housekeeping operations may occur in response to changes in the state of the cluster (e.g., when nodes or executions are added or removed from the cluster). Additionally, housekeeping operations may include passive operations to maintain cache space availability and reduce node resource utilization, such as removing invalid data from local caches and removing infrequently used data from the caches. Specific details of housekeeping operations are described in more detail below with reference to Figures 5 and 6.

[0044] 3A and 3B show code fragments that provide an example of a consistent hashing technique according to embodiments described herein. The particular language, style, and syntax of the code shown is not intended to limit the present disclosure. Those skilled in the art will recognize other languages ​​and syntaxes for implementing the described functionality.

[0045] FIG. 3A is a fragment of code 300 illustrating an example mapping of token boundaries to nodes in a cluster, according to some embodiments. As shown, the mapping represents a cluster including three worker nodes identified as "host_0", "host_1", and "host_2". The worker nodes may be similar to other worker nodes described herein, including worker nodes 206, 208 of FIG. 2. The worker nodes are configured to host two executions, resulting in a total of six executions in the cluster. In other embodiments, the cluster may have more than one, two, or three worker nodes hosting any suitable number of executions. In these embodiments, the mapping of token boundaries will be similar to that shown in FIG. 3A.

[0046] As mentioned above, a token boundary can represent a division of a set of integer values ​​that represent a set of token divisions. For example, integer values ​​between 0 and 2 31 For a cluster corresponding to the mapping of FIG. 3A, the token boundaries may divide the set of integer values ​​into six contiguous subsets of token divisions. Each worker node is mapped to two "TokenBounds" sets corresponding to the range of token divisions between pairs of integer token boundaries. For example, host_1 is mapped to a first pair of token boundaries 302 identified as "TokenBounds(0, 357913941)" and a second pair of token boundaries 304 identified as "TokenBounds(715827882, 1073741823)". Although shown with specific integer values ​​for the described cluster configuration, other values ​​may be derived from other embodiments of the present disclosure.

[0047] In some embodiments, a driver process of a driver node (e.g., driver 232 of driver node 204) can uniformly assign token boundaries to the worker nodes in the cluster. The uniform assignment of token boundaries can balance the number of token splits assigned to each worker node, such that each worker node has a number of token boundaries assigned based on the number of executors it hosts. As shown in FIG. 3A, three worker nodes are assigned two token boundary pairs, one for each of the two executors configured on the worker node.

[0048] FIG. 3B is another fragment of code 310 illustrating another example of a mapping of token boundaries to a set of updated nodes in a cluster, according to some embodiments. The set of updated nodes may be a result of the node updates depicted in FIG. 3A. As shown in the figure, the example mapping corresponds to a cluster with two worker nodes, host_0 and host_2. This mapping may occur due to a failure of the host_1 node. In response to a change in the cluster state, the first pair of token boundaries 302 previously assigned to host_1 may be reassigned by the driver node to the worker node corresponding to host_0. Similarly, the second pair of token boundaries previously assigned to host_1 may be reassigned to the worker node corresponding to host_0. This results in two worker nodes with uniformly distributed token splits.

[0049] In some embodiments, the token boundaries are updated in response to changes in cluster state. For example, if a worker node corresponding to host_1 in FIG. 3A fails, the driver node can update the token boundaries to correspond to a cluster with four executors. The updated token boundaries can then be assigned to the remaining worker nodes in the cluster. The token mapping then has four pairs of token boundaries, with two pairs each assigned to the remaining two worker nodes. Similarly, if worker nodes and / or executors are added to the cluster, the token boundaries can be updated to correspond to the executors and worker nodes being added.

[0050] FIG. 4 is a simplified flowchart of an exemplary process 400 for processing queries on a cluster with a deterministic cache, according to some embodiments. The cluster may be a cluster of distributed computing systems, including any of the distributed computing systems described herein, including cluster 202 of FIG. 2 and cluster 122 and distributed computing system 110 of FIG. 1. Process 400 is illustrated as a logical flowchart, each operation of which represents a sequence of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be omitted or the process may be implemented in any order and / or in any combination in parallel.

[0051] Some, any, or all of process 400 (or any other process described herein, or variations and / or combinations thereof) may be executed under the control of one or more computer systems configured with executable instructions and may be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors. The code may be stored in a computer-readable storage medium, for example, in the form of a computer program that includes a plurality of instructions executable by one or more processors. The computer-readable storage medium may be non-transitory.

[0052] Process 400 begins at block 402, where a distributed computing system (e.g., distributed computing system 110) implements a cluster (e.g., cluster 122 or cluster 202) including multiple nodes (e.g., driver node 204, worker nodes 206, 208). Implementing the cluster may include provisioning any suitable number of computing devices (including, but not limited to, physical computing devices and / or virtual machines running on one or more physical computing devices) as part of the cluster, including deploying appropriate software on the computing devices to support the operation of the computing devices as a cluster (e.g., hosting multiple nodes). At block 404, the distributed computing system maintains a state of the cluster. The state of the cluster may include a mapping of one or more execution units to multiple nodes and assigning multiple token boundaries to one or more execution units running on the multiple nodes based on the mapping. Maintaining the state of the cluster may include updating the mapping in response to changes in the cluster state, such as when nodes or execution units are added or removed within the cluster. Maintaining the cluster state may also include updating multiple token boundaries and uniformly allocating the updated token boundaries to one or more executing entities within the distributed computing system.

[0053] At block 406, the driver node may receive a query for execution. The query may be a query (e.g., query 250) submitted by a user associated with a tenant of the cluster. To execute the query, the driver node may identify a set of one or more data segments to be processed by a worker node of the cluster. The data segments may remain in deep storage, such as, for example, offline storage, object storage, or other similar storage systems. According to certain embodiments, the data segments may also reside in one or more caches associated with the worker nodes in the cluster.

[0054] At block 408, the driver node may compute a set of tokens corresponding to the set of one or more data segments. Each data segment may have a corresponding token that uniquely identifies the data segment. The tokens may be computed using a hashing algorithm, including MurmurHash3 or other suitable hashing algorithm. The input to the hashing algorithm may be identifying information associated with the data segments (e.g., file names).

[0055] At block 412, the driver node, in cooperation with a task manager, scheduler, or other similar process executing on the driver node, may launch a first task on a first executor executing on the first worker node. The first task may correspond to an operation on the first data segment, such that the first executor may retrieve the first data segment, perform the requested operation, and complete the task. The first task may be one task of a plurality of tasks constituting the query execution. The first worker node may be selected based at least in part on a first token of a set of tokens corresponding to a first pair of token boundaries associated with the first worker node. For example, a hash calculation of the first data segment may return a particular integer value of the token. This value may fall within a range of integer values ​​defined by the first pair of token boundaries.

[0056] At block 414, the first worker node may retrieve the first data segment. The first data segment may reside in a cache associated with the first worker node. The first data segment may reside in a cache associated with a second worker node in the cluster, or may reside in a data store, database, object store, deep store, or other repository off the node. Retrieving the first data segment from a cache associated with the first worker node may include the first execution unit identifying the first data segment in the cache and reading the contents into active memory. Retrieving the first data segment from a cache associated with the second worker node may include the first execution unit sending a request to one or more neighboring nodes of the plurality of nodes. The execution unit executing on the neighboring nodes may inspect the cache associated with each neighboring node based on the request. A second execution unit executing on a second worker node of one or more neighboring nodes can determine that the first data segment is present in a second cache, place the data segment in a block manager (e.g., associated with a block storage device or an internode storage system) or other module configured to handle data transfers between nodes in a distributed computing system, and then send an ID of the first data segment to the first execution unit that made the request. The first worker node can then copy the data segment from the second worker node to a first cache associated with the first worker node.

[0057] 5 is another simplified flowchart of an example process 500 for stabilizing a deterministic cache in a cluster when cluster conditions change, according to some embodiments. The cluster may be a cluster of distributed computing systems, including any of the distributed computing systems described herein, including cluster 202 of FIG. 2 and cluster 122 and distributed computing system 110 of FIG. 1.

[0058] Process 500 may begin at block 502 with a driver node (e.g., driver node 232) sending housekeeping requests to one or more worker nodes (e.g., worker nodes 206, 208) in the cluster. Housekeeping requests may be generated and sent in response to changes in the state of the cluster, for example, the addition or removal of worker nodes or executors. As the state of the cluster changes, the new configuration may affect the location of cached data segments in the cluster. Performing housekeeping operations may actively stabilize the cache.

[0059] In response to the housekeeping request, at block 504, one or more execution units running on the worker node may scan a cache associated with the worker node for data segments. For example, a first execution unit running on a first worker node may scan a cache associated with the first worker node to identify data segments. Similarly, another execution unit running on the first worker node may scan a cache associated with the first worker node to identify data segments. Because the cache on the first worker node is shared by the first execution unit and the other execution units, each execution unit may scan all data segments in the cache to determine whether the data segments correspond to a pair of token boundaries assigned to the execution unit.

[0060] At decision 506, the execution unit evaluates whether the token associated with the scanned data segment is within the token boundaries assigned to the execution unit. For example, the first execution unit may be assigned a first pair of token boundaries that define a range of token values ​​that the execution unit "owns." When scanning the cache associated with the first worker node, the first execution unit may determine whether the tokens of the data segments present correspond to values ​​within the token boundaries of the first execution unit. The execution unit may calculate the token value using the same hash algorithm used by the driver node when executing the query. In some embodiments, the token value may be passed by the driver node to the execution unit as part of a segment key table that is hashed for the token. If the data segment token is within the token boundaries of the execution unit, the segment is kept in the cache. If the data segment token is not within the token boundaries of the execution unit, at block 510 the execution unit sends the segment ID to all other execution units in the cluster.

[0061] At block 512, the execution unit receiving the outlier segment ID can copy the outlier segments from the other worker nodes to a local cache. For example, a first execution unit can identify one or more outlier data segments in a cache associated with the first worker node. Based on the current state of the cluster and the current assignment of token boundaries, one or more of these outlier data segments can be "owned" by a second execution unit of a second worker node. The second worker node can copy the data segments from the first worker node to a cache associated with the second worker node. In this way, the data segments can be present in an expected cache depending on the state of the cluster, and the cache can be stabilized after the state of the cluster changes.

[0062] 6 is a simplified flowchart of an example process 600 for cleaning a deterministic cache according to a timer process. The cluster may be a cluster of distributed computing systems, including any of the distributed computing systems described herein, including cluster 202 of FIG. 2 and cluster 122 and distributed computing system 110 of FIG. 1. Process 600 may represent a "passive" housekeeping operation that complements the "active" housekeeping operations described above with respect to FIG. 5.

[0063] At input 601, a driver node (e.g., driver node 232) may receive a timer instruction from a timer process implemented on the driver node. The timer may be configured to trigger operation of process 600 according to a fixed schedule. In some embodiments, the timer interval may vary according to cluster parameters or other conditions related to the cluster (e.g., cluster load, cluster state, etc.). In response to the timer, at block 602, the driver node may send a housekeeping request to one or more worker nodes in the cluster (e.g., worker nodes 206, 208). The housekeeping request may reduce resource usage in the cluster by "cleaning" the distributed cache to remove invalid or infrequently accessed data segments.

[0064] In response to the housekeeping request, at block 604, one or more execution units operating on the worker node may scan a cache associated with the worker node for data segments, similar to block 504 of FIG. 5. At decision 606, the one or more execution units evaluate whether the data segments in the cache are valid. During execution of the query, the data segments may be updated or modified such that the cached data segments no longer represent the current data stored in the object store. Data segments in the distributed cache that no longer match the data in the object store are invalid but may be retained in the cache until deleted. To determine whether the cached data segments are valid, the one or more execution units may build a valid segment list for each data cube in the object store. Based on the valid segment list, the one or more execution units may determine whether the cached data segments are valid. If the data segments are valid, they are retained in the cache at block 608. If the data segments are invalid, they are evicted (e.g., deleted) from the cache at block 610.

[0065] At decision 612, the driver node may determine whether the available cache storage is below a threshold. In some embodiments, the determination that the available storage is below a threshold may be made by one or more executions or cache managers executing on the node. In some embodiments, the threshold may be configured as part of the cluster configuration, and may be, for example, 80% of the total space configured in the cache. If the available cache space is above the threshold, the housekeeping operations may terminate at endpoint 614.

[0066] If the available cache storage space is less than the threshold, process 600 may proceed to decision 616. Similar to the operations described with respect to FIG. 5, at decision 616, one or more executives may evaluate whether a token associated with the scanned data segment is within a token boundary assigned to the executive. If so, at block 618, the segment is retained in the cache. If not, at block 620, the outlier segment may be copied to other worker nodes in the cluster according to the current assignment of token boundaries. At block 622, the outlier data segment may be removed from the local cache. The removal may occur after the data segment is copied to another worker node.

[0067] In decision 624, similar to decision 612, the available cache storage space is again checked to determine if it is below the threshold. If the available space is above the threshold, the housekeeping operation can terminate at endpoint 626. If the cache is still space constrained because the available space is below the set threshold, a third technique can be used to clean the cache. In block 628, the one or more executives can determine the temperature of the data segments stored in the distributed cache. As used herein, "temperature" refers to a measure of the frequency of access to a data segment in the cache, such that a "hot" data segment is accessed frequently over time by various queries executed by the cluster. In contrast, a "cold" data segment is accessed less frequently and resides in the cache. In some embodiments, the temperature of a segment can be classified as "hot," "mild," or "cold." The temperature can be determined by calculating the interval between previous accesses to the data segment. If the interval is decreasing (e.g., recent reads have become more frequent), the temperature level of the segment is upgraded (e.g., from low to mild, or from mild to high). If the interval is increasing (e.g., the frequency of reads has recently decreased), the temperature level of that segment can be reduced (e.g., from hot to mild, or from mild to cold). For example, a data segment may have been first accessed in cache 10 seconds ago, a second time accessed 5 seconds ago, and a third time accessed 2 seconds ago. Thus, the interval between the last accesses is decreasing, and the segment's temperature can be upgraded (e.g., from mild to hot).

[0068] At block 630, based on the temperatures of the segments in the cache, the one or more execution units may remove the coolest segment from the cache. In some embodiments, this may mean eliminating one, some, or all of the data segments that have a "cold" temperature. After removing the coolest data segment, the available cache space is checked again. The process of removing the coolest data segment is repeated until the available cache space exceeds a threshold.

[0069] In some embodiments, additional statistics for a data segment include, but are not limited to, cache hits, cache misses, number of invalidations, size, average access interval, replication factor, and whether the segment was retrieved from object store or inter-node store. Other combinations of data segment statistics can also be used in addition to temperature to determine which data segments to remove from the cache.

[0070] Example Infrastructure as a Service architecture As mentioned above, Infrastructure as a Service (IaaS) is one particular type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, cloud computing providers can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, IaaS providers can also provide various services (e.g., billing, monitoring, logging, load balancing, and clustering, etc.) that accompany these infrastructure components. Thus, these services can be policy-driven, so that IaaS users can potentially implement policies that drive load balancing to maintain application availability and performance.

[0071] In some cases, IaaS customers can access resources and services over a wide area network (WAN) such as the Internet, and can use the cloud provider's services to install the remaining elements of their application stack. For example, a user can log into an IaaS platform and create virtual machines (VM(s)), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on the VMs. The customer can then use the provider's services to perform a variety of functions such as balancing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.

[0072] In most cases, cloud computing models may require the participation of a cloud provider, which may be, but does not have to be, a third-party service that specializes in providing (e.g., providing, renting, selling) IaaS. An entity may also choose to deploy a private cloud and become its own infrastructure service provider.

[0073] In some examples, IaaS deployment is the process of placing a new application, or a new version of an application, onto a prepared application server, etc. This may also include the process of preparing the server (e.g., installing libraries, daemons, etc.). This is often managed by the cloud provider below the hypervisor layer (e.g., servers, storage, network hardware, and virtualization). Thus, the customer may be responsible for handling things like (OS), middleware, and / or application deployment (e.g., self-service virtual machines (e.g., that can be spun up on demand)).

[0074] In some instances, IaaS provisioning may even refer to obtaining computers or virtual hosts for use and installing desired libraries or services on them. In most cases, deployment does not include provisioning, so it may be desirable to perform provisioning first.

[0075] In some cases, there are two distinct challenges with IaaS provisioning. First, there is the initial challenge of provisioning an initial set of infrastructure before anything can run. Second, there is the challenge of evolving the existing infrastructure after everything has been provisioned (e.g. adding new services, modifying services, removing services, etc.). In some cases, these two challenges can be addressed by allowing the configuration of the infrastructure to be specified declaratively. In other words, the infrastructure (e.g. which components are required and how they interact) can be specified by one or more configuration files. Thus, the overall topology of the infrastructure (e.g. which resources depend on which resources and how each works together, etc.) can be described declaratively. In some cases, once the topology is specified, workflows can be generated that create and / or manage the various components described in the configuration files.

[0076] In some examples, the infrastructure may have many elements that are interconnected. For example, there may be one or more virtual private clouds (VPC(s)), also known as a core network (e.g., a potentially on-demand pool of configurable and / or shared computing resources). In some examples, there may also be one or more inbound / outbound traffic group rules and one or more virtual machines (VM(s)) that are provisioned to dictate how the network's inbound and / or outbound traffic is configured. Other infrastructure elements such as load balancers, databases, etc. may also be provisioned. The infrastructure may evolve in stages as more infrastructure elements are desired or added.

[0077] In some cases, continuous deployment techniques may be used to enable deployment of infrastructure code across various virtual computing environments. Additionally, the described techniques enable infrastructure management within these environments. In some examples, a service team may write code that is desirable to deploy to one or more, but often many, different production environments (e.g., across various different geographic locations, possibly even across the globe). However, in some examples, it may be desirable to first set up the infrastructure to which the code will be deployed. In some cases, provisioning may be done manually, and a provisioning tool may be utilized to provision resources and / or a deployment tool may be utilized to deploy the code after the infrastructure has been provisioned.

[0078] FIG. 7 is a block diagram 700 illustrating an example IaaS architecture pattern, according to at least one embodiment. A service operator 702 can be communicatively coupled to a secure host tenant 704, which can include a virtual cloud network (VCN) 706 and a secure host subnet 708. In some examples, the service operator 702 can use one or more client computing devices, which can be portable handheld devices (e.g., iPhone, mobile phone, iPad, computing tablet, personal digital assistant (PDA)) or wearable devices (e.g., Google Glass head mounted display), running software such as Microsoft Windows Mobile, and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, PalmOS, and other enabled communication protocols. Alternatively, the client computing devices can be general purpose personal computers, including, for example, personal computers and / or laptop computers running various versions of Microsoft Windows, Apple Macintosh, and / or Linux operating systems. The client computing device may be a workstation computer running any of a variety of commercially available UNIX or UNIX-like operating systems, including, but not limited to, various GNU / Linux operating systems such as Google Chrome OS.Alternatively, or in addition, the client computing device may be any other electronic device, such as a thin-client computer, an Internet-enabled gaming system (e.g., a Microsoft Xbox game console with or without a Kinect® gesture input device), and / or a personal messaging device capable of communicating over a network with access to the VCN 706 and / or the Internet.

[0079] The VCN 706 may include a local peering gateway (LPG) 710 that may be communicatively coupled to a secure shell (SSH) VCN 712 via an LPG 710 that is included in the SSH VCN 712. The SSH VCN 712 may include an SSH subnet 714, which may be communicatively coupled to a control plane VCN 716 via an LPG 710 that is included in the control plane VCN 716. The SSH VCN 712 may also be communicatively coupled to a data plane VCN 718 via the LPG 710. The control plane VCN 716 and the data plane VCN 718 may be included in a service tenant 719, which may be owned and / or operated by the IaaS provider.

[0080] The control plane VCN 716 may include a control plane demilitarized zone (DMZ) tier 720 that serves as a perimeter network (e.g., a portion of an enterprise network between the enterprise intranet and an external network). DMZ-based servers have limited responsibility and may help prevent breaches. Additionally, the DMZ tier 720 may include one or more load balancer (LB) subnet(s) 722, a control plane app tier 724 that may include app subnet(s) 726, a control plane data tier 728, which may include database (DB) subnet(s) 730 (e.g., front-end DB subnet(s) and / or back-end DB subnet(s)). The LB subnet(s) 722 included in the control plane DMZ tier 720 can be communicatively coupled to app subnet(s) 726 included in the control plane app tier 724 and an Internet gateway 734 that may be included in the control plane VCN 716, and the app subnet(s) 726 can be communicatively coupled to DB subnet(s) 730 included in the control plane data tier 728, as well as to a service gateway 736 and a network address translation (NAT) gateway 738. The control plane VCN 716 can include the service gateway 736 and the NAT gateway 738.

[0081] The control plane VCN 716 can include a data plane mirrored app layer 740 that can include app subnet(s) 726. The app subnet(s) 726 included in the data plane mirrored app layer 740 can include a virtual network interface controller (VNIC) 742 on which a compute instance 744 can run. The compute instance 744 can communicatively couple the app subnet(s) 726 of the data plane mirrored app layer 740 to the app subnet(s) 726 that can be included in the data plane app layer 746.

[0082] The data plane VCN 718 can include a data plane app layer 746, a data plane DMZ layer 748, and a data plane data layer 750. The data plane DMZ layer 748 can include LB subnet(s) 722, which can be communicatively coupled to the app subnet(s) 726 of the data plane app layer 746 and an Internet gateway 734 of the data plane VCN 718. The app subnet(s) 726 can be communicatively coupled to a service gateway 736 of the data plane VCN 718 and a NAT gateway 738 of the data plane VCN 718. The data plane data layer 750 can also include DB subnet(s) 730, which can be communicatively coupled to the app subnet(s) 726 of the data plane app layer 746.

[0083] The internet gateways 734 of the control plane VCNs 716 and data plane VCNs 718 may be communicatively coupled to a metadata management service 752, which may be communicatively coupled to the public internet 754. The public internet 754 may be communicatively connected to NAT gateways 738 of the control plane VCNs 716 and data plane VCNs 718. The service gateways 736 of the control plane VCNs 716 and data plane VCNs 718 may be communicatively coupled to cloud services 756.

[0084] In some examples, a service gateway 736 in the control plane VCN 716 or data plane VCN 718 can make application programming interface (API) calls to cloud services 756 without traversing the public Internet 754. API calls from the service gateway 736 to the cloud services 756 can be one-way: the service gateway 736 can make an API call to the cloud services 756, and the cloud services 756 can send the requested data to the service gateway 736. However, the cloud services 756 may not be able to initiate the API call to the service gateway 736.

[0085] In some examples, the secure host tenant 704 can be directly connected to the service tenant 719 or can be otherwise separate. The secure host subnet 708 can communicate with the SSH subnet 714 through the LPG 710, which can allow bidirectional communication through otherwise separate systems. Connecting the secure host subnet 708 to the SSH subnet 714 can give the secure host subnet 708 access to other entities in the service tenant 719.

[0086] The control plane VCN 716 can enable users of a service tenant 719 to set up or provision desired resources. The desired resources provisioned in the control plane VCN 716 can be deployed or used in the data plane VCN 718. In some examples, the control plane VCN 716 can be separate from the data plane VCN 718, and the data plane mirror app layer 740 of the control plane VCN 716 can communicate with the data plane app layer 746 of the data plane VCN 718 via VNIC(s) 742, which can be included in the data plane mirror app layer 740 and the data plane app layer 746.

[0087] In some examples, a user or customer of the system may make a request, such as, for example, a create, read, update, or delete (CRUD) operation, via the public internet 754, which may communicate the request to a metadata management service 752. The metadata management service 752 may communicate the request to the control plane VCN 716 via an internet gateway 734. The request may be received by the LB subnet(s) 722 included in the control plane DMZ layer 720. The LB subnet(s) 722 may determine that the request is valid, and in response to this determination, the LB subnet(s) 722 may send the request to the app subnet(s) 726 included in the control plane app layer 724. If the request is validated and a call to the public internet 754 is required, the call to the public internet 754 may be sent to the NAT gateway 738, which may make the call to the public internet 754. Memory that may be desirable to be stored with the request may be stored in the DB subnet(s) 730.

[0088] In some examples, the data plane mirror app layer 740 can facilitate direct communication between the control plane VCN 716 and the data plane VCN 718. For example, it may be desirable to apply configuration changes, updates, or other suitable modifications to resources included in the data plane VCN 718. Through the VNIC 742, the control plane VCN 716 can communicate directly with the resources included in the data plane VCN 718, thereby performing configuration changes, updates, or other suitable modifications to the resources included in the data plane VCN 718.

[0089] In some embodiments, the control plane VCN 716 and the data plane VCN 718 can be included in the service tenant 719. In this case, a user or customer of the system may not own or operate either the control plane VCN 716 or the data plane VCN 718. Instead, an IaaS provider may own or operate the control plane VCN 716 and the data plane VCN 718, both of which may be included in the service tenant 719. This embodiment may allow for network isolation that may prevent a user or customer from interacting with the resources of other users or other customers. This embodiment also allows a user or customer of the system to store databases privately without having to rely on the public Internet 754, which may not have the desired level of threat protection for storage.

[0090] In other embodiments, the LB subnet(s) 722 included in the control plane VCN 716 may be configured to receive signals from the service gateway 736. In this embodiment, the control plane VCN 716 and the data plane VCN 718 may be configured to be called by the IaaS provider's customers without calling the public Internet 754. The IaaS provider's customers may desire this embodiment because the database(s) used by the customers may be stored in a service tenant 719 that is controlled by the IaaS provider and may be isolated from the public Internet 754.

[0091] 8 is a block diagram 800 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 802 (e.g., service operator 702 of FIG. 7 ) can be communicatively coupled to a secure host tenant 804 (e.g., secure host tenant 704 of FIG. 7 ), which can include a virtual cloud network (VCN) 806 (e.g., VCN 706 of FIG. 7 ) and a secure host subnet 808 (e.g., secure host subnet 708 of FIG. 7 ). The VCN 806 can include a local peering gateway (LPG) 810 (e.g., LPG 710 of FIG. 7 ), which can be communicatively coupled to a secure shell (SSH) VCN 812 (e.g., SSH VCN 712 of FIG. 7 ) via the LPG 710 included in the SSH VCN 812. SSH VCN 812 can include an SSH subnet 814 (e.g., SSH subnet 714 in FIG. 7), which can be communicatively coupled to a control plane VCN 816 (e.g., control plane VCN 716 in FIG. 7) via an LPG 810 included in the control plane VCN 816. The control plane VCN 816 can be included in a service tenant 819 (e.g., service tenant 719 in FIG. 7), and the data plane VCN 818 (e.g., data plane VCN 718 in FIG. 7) can be included in a customer tenant 821, which can be owned or operated by a user or customer of the system.

[0092] The control plane VCN 816 may include a control plane DMZ layer 820 (e.g., control plane DMZ layer 720 of FIG. 7 ) that may include LB subnet(s) 822 (e.g., LB subnet(s) 722 of FIG. 7 ), a control plane app layer 824 (e.g., control plane app layer 724 of FIG. 7 ) that may include app subnet(s) 826 (e.g., app subnet(s) 726 of FIG. 7 ), a control plane data layer 828 (e.g., control plane data layer 728 of FIG. 7 ) that may include database (DB) subnet(s) 830 (e.g., similar to DB subnet(s) 730 of FIG. 7 ). The LB subnet(s) 822 included in the control plane DMZ tier 820 can be communicatively coupled to app subnet(s) 826 included in the control plane app tier 824 and to an Internet gateway 834 (e.g., Internet gateway 734 in FIG. 7 ) that may be included in the control plane VCN 816, and the app subnet(s) 826 can be communicatively coupled to DB subnet(s) 830 included in the control plane data tier 828, as well as to a service gateway 836 (e.g., service gateway in FIG. 7 ) and a network address translation (NAT) gateway 838 (e.g., NAT gateway 738 in FIG. 7 ). The control plane VCN 816 can include the service gateway 836 and the NAT gateway 838.

[0093] The control plane VCN 816 can include a data plane mirror app layer 840 (e.g., data plane mirror app layer 740 of FIG. 7 ), which can include app subnet(s) 826. The app subnet(s) 826 included in the data plane mirror app layer 840 can include virtual network interface controllers (VNICs) 842 (e.g., VNICs 742) on which compute instances 844 (e.g., similar to compute instances 744 of FIG. 7 ) can run. The compute instances 844 can facilitate communication between the app subnet(s) 826 and the app subnet(s) 826 of the data plane mirror app layer 840, which can be included in the data plane app layer 846 (e.g., data plane app layer 746 of FIG. 7 ), via the VNICs 842 included in the data plane mirror app layer 840 and the VNICs 842 included in the data plane app layer 846.

[0094] An Internet gateway 834 included in the control plane VCN 816 can be communicatively coupled to a metadata management service 852 (e.g., metadata management service 752 in FIG. 7), which can be communicatively coupled to a public Internet 854 (e.g., public Internet 754 in FIG. 7). The public Internet 854 can be communicatively coupled to a NAT gateway 838 included in the control plane VCN 816. A service gateway 836 included in the control plane VCN 816 can be communicatively coupled to cloud services 856 (e.g., cloud services 756 in FIG. 7).

[0095] In some examples, the data plane VCN 818 can be included in a customer tenant 821. In this case, the IaaS provider can provide a control plane VCN 816 for each customer, and the IaaS provider can set up a unique compute instance 844 included in a service tenant 819 for each customer. Each compute instance 844 may enable communication between the control plane VCN 816 included in the service tenant 819 and the data plane VCN 818 included in the customer tenant 821. The compute instance 844 may enable resources provisioned in the control plane VCN 816 included in the service tenant 819 to be deployed or otherwise used in the data plane VCN 818 included in the customer tenant 821.

[0096] In another example, an IaaS provider customer may have a database that resides in a customer tenant 821. In this example, the control plane VCN 816 may include a data plane mirror app layer 840, which may include app subnet(s) 826. The data plane mirror app layer 840 may reside in the data plane VCN 818, but the data plane mirror app layer 840 may not reside in the data plane VCN 818. That is, the data plane mirror app layer 840 has access to the customer tenant 821, but the data plane mirror app layer 840 may not reside in the data plane VCN 818 or may be owned or operated by the IaaS provider customer. The data plane mirror app layer 840 may be configured to make calls to the data plane VCN 818, but may not be configured to make calls to any entities included in the control plane VCN 816. A customer may desire to deploy or otherwise use resources in the data plane VCN 818 that are provisioned in the control plane VCN 816, and the data plane mirror app layer 840 can facilitate the customer's desired deployment or other use of the resources.

[0097] In some embodiments, the IaaS provider's customer can apply filters to the data plane VCN 818. In this embodiment, the customer can determine what the data plane VCN 818 can access, and the customer can limit access from the data plane VCN 818 to the public Internet 854. The IaaS provider may not be able to apply filters or control the data plane VCN 818's access to external networks or databases. Applying customer filters and controls to the data plane VCN 818 contained in a customer tenant 821 can help isolate the data plane VCN 818 from other customers and the public Internet 854.

[0098] In some embodiments, cloud services 856 may be called by service gateway 836 to access services that may not be on the public Internet 854, on the control plane VCN 816, or on the data plane VCN 818. The connection between cloud services 856 and the control plane VCN 816 or the data plane VCN 818 may not be live or continuous. Cloud services 856 may be on another network owned or operated by the IaaS provider. Cloud services 856 may be configured to receive calls from service gateway 836 or may be configured not to receive calls from the public Internet 854. Some cloud services 856 may be isolated from other cloud services 856, and control plane VCN 816 may be isolated from cloud services 856 that may not be in the same region as control plane VCN 816. For example, control plane VCN 816 may be located in “Region 1” and cloud service “Deployment 7” may be located in Region 1 and Region 2. If a call to deployment 7 is made by a service gateway 836 included in control plane VCN 816 in region 1, the call may be sent to deployment 7 in region 1. In this example, control plane VCN 816, or deployment 7 in region 1, may not be communicatively coupled to or in communication with deployment 7 in region 2.

[0099] 9 is a block diagram 900 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 902 (e.g., service operator 702 of FIG. 7 ) can be communicatively coupled to a secure host tenant 904 (e.g., secure host tenant 704 of FIG. 7 ), which can include a virtual cloud network (VCN) 906 (e.g., VCN 706 of FIG. 7 ) and a secure host subnet 908 (e.g., secure host subnet 708 of FIG. 7 ). The VCN 906 can include an LPG 910 (e.g., LPG 710 of FIG. 7 ) that can be communicatively coupled to an SSH VCN 912 via an LPG 910 included in the SSH VCN 912 (e.g., SSH VCN 712 of FIG. 7 ). The SSH VCN 912 can include an SSH subnet 914 (e.g., SSH subnet 714 in FIG. 7), which can be communicatively coupled to a control plane VCN 916 via an LPG 910 included in the control plane VCN 916 (e.g., control plane VCN 716 in FIG. 7) and to a data plane VCN 918 via an LPG 910 included in the data plane VCN 918 (e.g., data plane 718 in FIG. 7). The control plane VCN 916 and the data plane VCN 918 can be included in a service tenant 919 (e.g., service tenant 719 in FIG. 7).

[0100] The control plane VCN 916 may include a control plane DMZ layer 920 (e.g., control plane DMZ layer 720 of FIG. 7 ) that may include load balancer (LB) subnet(s) 922 (e.g., LB subnet(s) 722 of FIG. 7 ), a control plane app layer 924 (e.g., control plane app layer 724 of FIG. 7 ) that may include app subnet(s) 926 (e.g., similar to app subnet(s) 726 of FIG. 7 ), a control plane data layer 928 (e.g., control plane data layer 728 of FIG. 7 ) that may include DB subnet(s) 930. The LB subnet(s) 922 included in the control plane DMZ tier 920 may be communicatively coupled to app subnet(s) 926 included in the control plane app tier 924 and to an Internet gateway 934 (e.g., Internet gateway 734 of FIG. 7 ) that may be included in the control plane VCN 916, and the app subnet(s) 926 can be communicatively coupled to DB subnet(s) 930, service gateway 936 (e.g., service gateway of FIG. 7 ), and network address translation (NAT) gateway 938 (e.g., NAT gateway 738 of FIG. 7 ) included in the control plane data tier 928. The control plane VCN 916 may include the service gateway 936 and the NAT gateway 938.

[0101] The data plane VCN 918 can include a data plane app layer 946 (e.g., data plane app layer 746 of FIG. 7), a data plane DMZ layer 948 (e.g., data plane DMZ layer 748 of FIG. 7), and a data plane data layer 950 (e.g., data plane data layer 750 of FIG. 7). The data plane DMZ layer 948 can include LB subnet(s) 922, which can be communicatively coupled to trusted app subnet(s) 960 and untrusted app subnet(s) 962 of the data plane app layer 946, and to an Internet gateway 934 included in the data plane VCN 918. The trusted app subnet(s) 960 can be communicatively coupled to a service gateway 936 included in the data plane VCN 918, a NAT gateway 938 included in the data plane VCN 918, and a DB subnet(s) 930 included in the data plane data layer 950. The untrusted app subnet(s) 962 can be communicatively coupled to a service gateway 936 included in the data plane VCN 918 and to a DB subnet(s) 930 included in the data plane data layer 950. The data plane data layer 950 can include the DB subnet(s) 930 that can be communicatively coupled to a service gateway 936 included in the data plane VCN 918.

[0102] The untrusted app subnet(s) 962 can include one or more primary VNIC(s) 964(1)-(N) that can be communicatively coupled to tenant virtual machines (VM(s)) 966(1)-(N). Each tenant VM 966(1)-(N) can be communicatively coupled to a respective app subnet 967(1)-(N), which can be included in a respective container egress VCN(s) 968(1)-(N), which can be included in a respective customer tenant 970(1)-(N). Each secondary VNIC(s) 972(1)-(N) can facilitate communication between the untrusted app subnet(s) 962 included in the data plane VCN 918 and the app subnet included in the container egress VCN(s) 968(1)-(N). Each container egress VCN(s) 968(1)-(N) can include a NAT gateway 938 that can be communicatively coupled to the public Internet 954 (e.g., the public Internet 754 in FIG. 7).

[0103] The internet gateway 934 included in the control plane VCN 916 and in the data plane VCN 918 can be communicatively coupled to a metadata management service 952 (e.g., metadata management system 752 of FIG. 7 ), which can be communicatively coupled to the public internet 954. The public internet 954 can be communicatively coupled to a NAT gateway 938 included in the control plane VCN 916 and in the data plane VCN 918. The service gateway 936 included in the control plane VCN 916 and in the data plane VCN 918 can be communicatively coupled to cloud services 956.

[0104] In some embodiments, the data plane VCN 918 can be integrated with customer tenants 970. This integration can be beneficial or desirable for the IaaS provider's customers, such as when support is needed when running code. The customer may provide code for execution that may be disruptive, communicate with other customer resources, or cause other undesirable effects. In response, the IaaS provider can decide whether to run code provided to the IaaS provider by the customer.

[0105] In some examples, an IaaS provider's customer may grant the IaaS provider temporary network access to request functionality to be added to data plane layer app 946. The code to execute the functionality may be executed in VM(s) 966(1)-(N), and the code may not be configured to execute elsewhere on data plane VCN 918. Each VM(s) 966(1)-(N) may be connected to one customer tenant 970. Each container 971(1)-(N) contained in VM 966(1)-(N) may be configured to execute code. In this case, double isolation may exist (e.g., code execution in container 971(1)-(N), container 971(1)-(N) may be contained in at least VM 966(1)-(N) contained in untrusted app subnet(s) 962), which may help prevent errant or unwanted code from damaging the IaaS provider's network or damaging another customer's network. Containers 971(1)-(N) may be communicatively coupled to customer tenants 970 and may be configured to send or receive data from customer tenants 970. Containers 971(1)-(N) may not be configured to send or receive data from any other entities in data plane VCN 918. Once code execution is complete, the IaaS provider may kill or otherwise destroy containers 971(1)-(N).

[0106] In some embodiments, trusted app subnet(s) 960 may execute code that may be owned or operated by the IaaS provider. In this embodiment, trusted app subnet(s) 960 may be communicatively coupled to DB subnet(s) 930 and configured to perform CRUD operations within DB subnet(s) 930. Untrusted app subnet(s) 962 may be communicatively coupled to DB subnet(s) 930, although in this embodiment, the untrusted app subnet(s) may be configured to perform read operations in DB subnet(s) 930. Containers 971(1)-(N) that may be included in each customer's VMs 966(1)-(N) and that may execute code from the customer may not be communicatively coupled to DB subnet(s) 930.

[0107] In other embodiments, the control plane VCN 916 and the data plane VCN 918 may not be directly communicatively coupled. In this embodiment, there may not be direct communication between the control plane VCN 916 and the data plane VCN 918. However, communication may occur indirectly through at least one method. The LPG 910 may be established by an IaaS provider that may facilitate communication between the control plane VCN 916 and the data plane VCN 918. In another example, the control plane VCN 916 or the data plane VCN 918 may make a call to a cloud service 956 through a service gateway 936. For example, a call from the control plane VCN 916 to the cloud service 956 may include a request for a service that may communicate with the data plane VCN 918.

[0108] 10 is a block diagram 1000 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1002 (e.g., service operator 702 of FIG. 7 ) can be communicatively coupled to a secure host tenant 1004 (e.g., secure host tenant 704 of FIG. 7 ), which can include a virtual cloud network (VCN) 1006 (e.g., VCN 706 of FIG. 7 ) and a secure host subnet 1008 (e.g., secure host subnet 708 of FIG. 7 ). The VCN 1006 can include an LPG 1010 (e.g., LPG 710 of FIG. 7 ) that can be communicatively coupled to an SSH VCN 1012 (e.g., SSH VCN 712 of FIG. 7 ) via an LPG 1010 included in the SSH VCN 1012. The SSH VCN 1012 can include an SSH subnet 1014 (e.g., SSH subnet 714 in FIG. 7 ), which can be communicatively coupled to a control plane VCN 1016 via an LPG 1010 included in the control plane VCN 1016 (e.g., control plane VCN 716 in FIG. 7 ) and to a data plane VCN 1018 via an LPG 1010 included in the data plane VCN 1018 (e.g., data plane 718 in FIG. 7 ). The control plane VCN 1016 and the data plane VCN 1018 can be included in a service tenant 1019 (e.g., service tenant 719 in FIG. 7 ).

[0109] The control plane VCN 1016 may include a control plane DMZ layer 1020 (e.g., control plane DMZ layer 720 of FIG. 7 ) that may include LB subnet(s) 1022 (e.g., LB subnet(s) 722 of FIG. 7 ), a control plane app layer 1024 (e.g., control plane app layer 724 of FIG. 7 ) that may include app subnet(s) 1026 (e.g., app subnet(s) 726 of FIG. 7 ), a control plane data layer 1028 (e.g., control plane data layer 728 of FIG. 7 ) that may include DB subnet(s) 1030 (e.g., DB subnet(s) 930 of FIG. 9 ). The LB subnet(s) 1022 included in the control plane DMZ tier 1020 may be communicatively coupled to app subnet(s) 1026 included in the control plane app tier 1024 and may be communicatively coupled to an Internet gateway 1034 (e.g., Internet gateway 734 of FIG. 7 ) included in the control plane VCN 1016, which may be communicatively coupled to a DB subnet(s) 1030 included in the control plane data tier 1028 and may be communicatively coupled to a service gateway 1036 (e.g., service gateway of FIG. 7 ) and a network address translation (NAT) gateway 1038 (e.g., NAT gateway 738 of FIG. 7 ). The control plane VCN 1016 may include the service gateway 1036 and the NAT gateway 1038.

[0110] The data plane VCN 1018 can include a data plane app layer 1046 (e.g., data plane app layer 746 of FIG. 7), a data plane DMZ layer 1048 (e.g., data plane DMZ layer 748 of FIG. 7), and a data plane data layer 1050 (e.g., data plane data layer 750 of FIG. 7). The data plane DMZ layer 1048 can include LB subnet(s) 1022, which can be communicatively coupled to trusted app subnet(s) 1060 (e.g., trusted app subnet(s) 960 of FIG. 9), and to untrusted app subnet(s) 1062 (e.g., untrusted app subnet(s) 962 of FIG. 9) of the data plane app layer 1046 and an Internet gateway 1034 included in the data plane VCN 1018. The trusted app subnet(s) 1060 may be communicatively coupled to a service gateway 1036 included in the data plane VCN 1018, and may be communicatively coupled to a NAT gateway 1038 included in the data plane VCN 1018, and to a DB subnet(s) 1030 included in the data plane data layer 1050. The untrusted app subnet(s) 1062 may be communicatively connected to a service gateway 1036 included in the data plane VCN 1018 and to a DB subnet(s) 1030 included in the data plane data layer 1050. The data plane data layer 1050 may include a DB subnet(s) 1030 that may be communicatively coupled to a service gateway 1036 included in the data plane VCN 1018.

[0111] The untrusted app subnet(s) 1062 may include primary VNIC(s) 1064(1)-(N), which may be communicatively coupled to tenant virtual machines (VM(s)) 1066(1)-(N) that reside in the untrusted app subnet(s) 1062. Each tenant VM 1066(1)-(N) may execute code within a respective container 1067(1)-(N), which may be communicatively coupled to an app subnet 1026 that may be included in a data plane app tier 1046 that may be included in a container egress VCN 1068. Each secondary VNIC(s) 1072(1)-(N) may facilitate communication between the untrusted app subnet(s) 1062 included in the data plane VCN 1018 and the app subnet included in the container egress VCN 1068. The container egress VCN may include a NAT gateway 1038 that may be communicatively coupled to the public Internet 1054 (e.g., public Internet 754 of FIG. 7).

[0112] An Internet gateway 1034 included in the control plane VCN 1016 and included in the data plane VCN 1018 can be communicatively coupled to a metadata management service 1052 (e.g., metadata management system 752 of FIG. 7 ), which can be communicatively coupled to the public Internet 1054. The public Internet 1054 can be communicatively coupled to a NAT gateway 1038 included in the control plane VCN 1016 and included in the data plane VCN 1018. A service gateway 1036 included in the control plane VCN 1016 and included in the data plane VCN 1018 can be communicatively coupled to cloud services 1056.

[0113] In some examples, the pattern illustrated by the architecture of block diagram 1000 of FIG. 10 may be an exception to the pattern illustrated by the architecture of block diagram 900 of FIG. 9 and may be desirable for an IaaS provider's customer when the IaaS provider cannot communicate directly with the customer (e.g., in a disconnected area). Each container 1067(1)-(N) contained in each customer's VM(s) 1066(1)-(N) is accessible in real time by the customer. The containers 1067(1)-(N) may be configured to make calls to a respective secondary VNIC(s) 1072(1)-(N) contained in the app subnet(s) 1026 of the data plane app tier 1046 that may be contained in a container egress VCN 1068. The secondary VNIC(s) 1072(1)-(N) may send the calls to a NAT gateway 1038 that may send the calls to the public Internet 1054. In this example, containers 1067(1)-(N) that a customer can access in real time can be isolated from the control plane VCN 1016 and can be isolated from other entities contained in the data plane VCN 1018. Containers 1067(1)-(N) may be isolated from resources from other customers.

[0114] In another example, a customer can use containers 1067(1)-(N) to invoke cloud service 1056. In this example, the customer can execute code within containers 1067(1)-(N) that requests a service from cloud service 1056. Containers 1067(1)-(N) can send the request to secondary VNIC(s) 1072(1)-(N), which can send the request to a NAT gateway that can send the request to public Internet 1054. Public Internet 1054 can send the request via Internet gateway 1034 to LB subnet(s) 1022 included in control plane VCN 1016. In response to determining that the request is valid, the LB subnet(s) can send the request to app subnet(s) 1026, which can send the request to cloud service 1056 via service gateway 1036.

[0115] It should be understood that the IaaS architectures 700, 800, 900, 1000 depicted in the figures may have components other than those depicted. Additionally, the embodiments depicted in the figures are only some examples of cloud infrastructure systems that may incorporate embodiments of the present disclosure. In other embodiments, the IaaS systems may have more or fewer components than depicted, may combine two or more components, or may have different configurations or arrangements of components.

[0116] In certain embodiments, the IaaS system described herein may include a suite of application, middleware, and database services products that are delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. One example of such an IaaS system is Oracle Cloud Infrastructure (OCI), offered by the present assignee.

[0117] 11 illustrates an exemplary computer system 1100 in which various embodiments may be implemented. The system 1100 may be used to implement any of the computer systems described above. As shown in the figure, the computer system 1100 includes a processing unit 1104 that communicates with a number of peripheral subsystems via a bus subsystem 1102. These peripheral subsystems may include a processing accelerator 1106, an I / O subsystem 1108, a storage subsystem 1118, and a communication subsystem 1124. The storage subsystem 1118 includes a tangible computer readable storage medium 1122 and a system memory 1110.

[0118] Bus subsystem 1102 provides a mechanism that allows the various components and subsystems of computer system 1100 to communicate with each other as intended. Although bus subsystem 1102 is shown diagrammatically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 1102 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus. It may be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.

[0119] The processing unit 1104 may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers) and controls the operation of the computer system 1100. One or more processors may be included in the processing unit 1104. These processors may include single-core processors or multi-core processors. In particular embodiments, the processing unit 1104 may be implemented as one or more independent processing units 1132 and / or 1134 with a single-core processor or a multi-core processor included in each processing unit. In other embodiments, the processing unit 1104 may be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.

[0120] In various embodiments, the processing unit 1104 may execute various programs in response to program code and may maintain multiple simultaneously executing programs or processes. At any time, some or all of the program code being executed may reside in the processor(s) 1104 and / or in the storage subsystem 1118. Through appropriate programming, the processor(s) 1104 may provide various functions as discussed above. The computer system 1100 may further include a processing accelerator 1106, which may include a digital signal processor (DSP), a special purpose processor, or the like.

[0121] The I / O subsystem 1108 can include user interface input devices and user interface output devices. User interface input devices can include keyboards, pointing devices such as mice or trackballs, touchpads or touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, audio input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices can include, for example, motion sensing and / or gesture recognition devices such as Microsoft Kinect® motion sensors, allowing users to control and interact with input devices such as Microsoft Xbox® 360 game controllers through a natural user interface using gestures and voice commands. User interface input devices can also include eye gesture recognition devices such as the Google Glass® blink detector that detects eye activity from a user (e.g., “blinking” while taking a picture and / or selecting a menu) and translates eye gestures as input to an input device (e.g., Google Glass®). Additionally, the user interface input devices may include voice recognition sensing devices that allow a user to interact with a voice recognition system (eg, the Siri® navigator) through voice commands.

[0122] User interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, game pads and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers 3D scanners, 3D printers, laser range finders, and eye-tracking devices. Additionally, user interface input devices may include medical imaging input devices such as, for example, computed tomography, magnetic resonance imaging, position emission tomography, and medical ultrasound machines. User interface input devices may also include audio input devices such as, for example, MIDI keyboards, digital musical instruments, and the like.

[0123] User interface output devices may include non-visual displays such as a display subsystem, indicator lights, or audio output devices. The display subsystem may be a flat panel device such as one using a cathode ray tube (CRT), liquid crystal display (LCD) or plasma display, a projection device, a touch screen, etc. In general, use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from computer system 1100 to a user or to another computer. For example, user interface output devices include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automobile navigation systems, plotters, voice output devices, and modems.

[0124] Computer system 1100 may include a storage subsystem 1118 that comprises software elements shown as currently located in system memory 1110. The system memory 1110 may store program instructions that are loadable and executable on the processing unit 1104, as well as data generated during the execution of these programs.

[0125] Depending on the configuration and type of computer system 1100, the system memory 1110 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that are immediately accessible to and / or currently being operated on and executed by the processing unit 1104. In some implementations, the system memory 1110 may include a number of different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS), containing the basic routines that help to transfer information between elements within the computer system 1100, such as during start-up, may typically be stored in ROM. By way of example and not limitation, the system memory 1110 also illustrates application programs 1112, program data 1114, and an operating system 1116, which may include client applications, a web browser, a mid-tier application, a relational database management system (RDBMS), and the like. By way of example, operating systems 1116 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® 11 OS, and Palm® OS operating systems.

[0126] The storage subsystem 1118 may also provide a tangible computer-readable storage medium for storing the basic programming and data structures that provide the functionality of some embodiments. Software (programs, code modules, instructions) that, when executed by the processor, provide the functionality described above may be stored in the storage subsystem 1118. These software modules or instructions may be executed by the processing unit 1104. The storage subsystem 1118 may also provide a repository for storing data used in accordance with the present disclosure.

[0127] The storage subsystem 1100 may also include a computer readable storage medium reader 1120 that may be further coupled to a computer readable storage medium 1122. Together, and optionally in combination with the system memory 1110, the computer readable storage medium 1122 may comprehensively represent remote, local, fixed, and / or removable storage devices, as well as storage media for containing, storing, transmitting, and retrieving computer readable information on a temporary and / or more permanent basis.

[0128] The computer readable storage medium 1122 containing the code or portions of code may include any suitable medium known or used in the art, including, but not limited to, storage media and communication media, such as volatile and non-volatile, removable and non-removable media implemented in any manner or technology for storing and / or transmitting information. This may include tangible computer readable storage media, such as RAM, ROM, Electronically Erasable Programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage, or other tangible computer readable medium. This may also include intangible computer readable media, such as a data signal, data transmission, or any other medium that can be used to transmit the desired information and that can be accessed by the computing system 1100.

[0129] As an example, the computer readable storage medium 1122 may include a hard disk drive that reads or writes to a non-removable non-volatile magnetic medium, a magnetic disk drive that reads or writes to a removable non-volatile magnetic disk, and an optical disk drive that reads or writes to a removable non-volatile optical disk, such as a CD-ROM, DVD, Blu-Ray® disk, or other optical medium. The computer readable storage medium 1122 may include, but is not limited to, a Zip® drive, a flash memory card, a Universal Serial Bus (USB) flash drive, a Secure Digital (SD) card, a DVD disk, a digital video tape, and the like. The computer readable storage medium 1122 may also include a solid state drive (SSD) based on non-volatile memory, such as a flash memory-based SSD, an enterprise flash drive, an SSD based on volatile memory, such as solid state RAM, dynamic RAM, static RAM, such as solid state ROM, a DRAM-based SSD, a magnetoresistive RAM (MRAM) SSD, and a hybrid SSD that uses a combination of DRAM and a flash memory-based SSD. The disk drives and their associated computer-readable media may provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for computer system 1100.

[0130] The communication subsystem 1124 provides an interface to other computer systems and networks. The communication subsystem 1124 serves as an interface for transmitting and receiving data from the computer system 1100 to and from other systems. For example, the communication subsystem 1124 may enable the computer system 1100 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1124 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular technology, advanced data network technologies such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution)), WiFi (IEEE 802.11 family standard, or other mobile communication technologies, or any combination thereof), global positioning system (GPS) receiver components, and / or other components. In some embodiments, the communication subsystem 1124 may provide a wired network connection (e.g., Ethernet) in addition to or instead of a wireless interface.

[0131] In some embodiments, the communications subsystem 1124 may also receive incoming communications in the form of structured and / or unstructured data feeds 1126, event streams 1128, event updates 1130, etc., on behalf of one or more users who may use the computer system 1100.

[0132] As an example, the communications subsystem 1124 may be configured to receive data feeds 1126 in real time from users of social networks and / or other communications services, such as web feeds, such as Twitter® feeds, Facebook® updates, Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third party information sources.

[0133] Additionally, the communications subsystem 1124 may be configured to receive data in the form of a continuous data stream, which may include an event stream 1128 of real-time events and / or event updates 1130, which may be continuous or essentially unlimited with no apparent end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, and the like.

[0134] The communications subsystem 1124 may also be configured to output structured and / or unstructured data feeds 1126, event streams 1128, event updates 1130, etc. to one or more databases that may be in communication with one or more streaming data source computers coupled to the computer system 1100.

[0135] The computer system 1100 may be one of a variety of types, including a handheld portable device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google Glass® head mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack, or other data processing system.

[0136] Due to the ever-changing nature of computers and networks, the description of the computer system 1100 shown in the figure is intended as a specific example only. Many other configurations are possible, with more or fewer components than the system shown in the figure. For example, customized hardware may also be used, or particular elements may be implemented in hardware, firmware, software (including applets), or a combination thereof. Additionally, connections to other computing devices, such as network input / output devices, may be used. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will appreciate other approaches and / or methods for implementing various embodiments.

[0137] Although specific embodiments have been described, various modifications, variations, alternative constructions, and equivalents are within the scope of the disclosure. The embodiments are not limited to operating in any particular data processing environment, but may freely operate in multiple data processing environments. Furthermore, while the embodiments have been described using a particular sequence of transactions and steps, it will be apparent to those skilled in the art that the scope of the disclosure is not limited to the sequence of transactions and steps described. Various features and aspects of the above-described embodiments may be used individually or in combination.

[0138] Furthermore, while embodiments have been described using a particular combination of hardware and software, it should be appreciated that other combinations of hardware and software are within the scope of the present disclosure. The embodiments may be implemented using only hardware, only software, or a combination thereof. The various processes described herein may be implemented on the same processor or any combination of different processors. Thus, where a component or module is described as being configured to perform a particular operation, such configuration may be achieved, for example, by designing electronic circuitry to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, or any combination thereof. Processes may communicate using a variety of techniques, including, but not limited to, conventional techniques for inter-process communication, with different pairs of processes using different techniques and the same pair of processes using different techniques at different times.

[0139] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will be apparent, however, that additions, subtractions, deletions, and other modifications and alterations may be made without departing from the broader spirit and scope of the appended claims. Accordingly, although certain disclosed embodiments have been described, they are not intended to be limiting. Various modifications and equivalents are intended to be within the scope of the following claims.

[0140] Use of the terms "a," "an," "the," and similar referents in the context of describing the disclosed embodiments (particularly in the context of the claims that follow) should be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "comprise," "have," "include," and "contain" should be construed as open-ended terms (i.e., meaning "including but not limited to"), unless otherwise noted. The term "connected" should be construed as being partially or wholly contained within, attached to, or connected to, even if there is something intervening. The recitation of ranges of values ​​herein is merely intended to serve as a shorthand method of individually referring to each individual value falling within the range, unless otherwise stated herein, and each individual value is incorporated into the specification as if it were individually set forth herein. All methods described herein can be performed in any suitable order, unless otherwise indicated herein or clearly contradicted by context. Any examples provided herein, or the use of exemplary language (e.g., "such as"), are intended merely to better illustrate the embodiments and do not impose limitations on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.

[0141] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is intended to be understood in context as generally used to indicate that an item, term, etc. can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless specifically stated otherwise. Thus, such disjunctive language is generally not intended to, and should not, imply that a particular embodiment requires that at least one of X, at least one of Y, or at least one of Z, respectively, be present.

[0142] Preferred embodiments of the present disclosure are described herein, including the best mode known for carrying out the present disclosure. Variations of these preferred embodiments will become apparent to those skilled in the art upon reading the foregoing description. Those skilled in the art should be able to adopt such variations as necessary, and the present disclosure may be carried out in ways other than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, unless otherwise indicated herein, this disclosure encompasses any combination of the above-described elements in all possible variations thereof.

[0143] All references cited in this specification, including publications, patent applications, and patents, are herein incorporated by reference to the same extent as if each reference was individually and specifically indicated to be incorporated by reference and was set forth in its entirety herein.

[0144] In the foregoing specification, aspects of the disclosure have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that the disclosure is not limited thereto. Various features and aspects of the above-described disclosure can be used individually or in combination. Moreover, the embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Thus, the specification and drawings should be regarded as illustrative rather than restrictive.

Claims

1. 1. A computer-implemented method comprising: A distributed computing system for providing analytical data processing services includes: implementing a cluster including a plurality of nodes; the distributed computing system maintaining a cluster state including a plurality of token boundaries uniformly associated with the plurality of nodes; a driver node of the plurality of nodes receiving a query for execution; identifying, by the driver node based at least in part on the query, a set of one or more data segments corresponding to the query; said driver node computing a set of tokens corresponding to said set of one or more data segments; launching, on a first execution unit executing on a first worker node of the plurality of nodes, a first task for processing a first data segment from the set of one or more data segments, wherein the first worker node selected is based at least in part on a first token of the set of tokens corresponding to a first pair of token boundaries of the plurality of token boundaries, the first pair of token boundaries being associated with the first worker node; The method comprises: The method further comprising the first worker node obtaining the first data segment.

2. The step of maintaining the state of the cluster comprises: storing a mapping of one or more execution units to the plurality of nodes; and assigning the plurality of token boundaries to the one or more execution units executing on the plurality of nodes based at least in part on the mapping, wherein the plurality of token boundaries are uniformly distributed to the one or more execution units, and wherein a token boundary of the first pair of the plurality of token boundaries is assigned to a first execution unit of the one or more execution units.

3. receiving, by the driver node, an indication that the cluster has changed; updating the plurality of token boundaries in response to the indication; 3. The computer-implemented method of claim 2, further comprising: assigning the updated token boundaries to the one or more execution units, wherein the updated token boundaries are uniformly distributed among the one or more execution units.

4. 2. The computer-implemented method of claim 1, wherein obtaining the first data segment comprises: the first execution unit determining that the first data segment resides in a first cache associated with the first worker node; and reading the first data segment from the first cache.

5. The step of obtaining the first data segment comprises: the first execution unit determining that the first data segment is not present in a first cache associated with the first worker node; sending a request to one or more neighboring nodes of the plurality of nodes; In response to the request, a second execution unit executing on a second worker node of the one or more neighboring nodes determines that the first data segment is present in a second cache associated with the second worker node; the first worker node copying the first data segment from the second cache to the first cache; and reading the first data segment from the first cache.

6. The computer-implemented method of claim 1 , wherein computing the set of tokens comprises computing a hash value that uniquely identifies the one or more data segments.

7. 7. The computer-implemented method of claim 6, wherein the token boundaries include integer values ​​that divide a range of integer token values, and the calculated hash value is within the range of integer token values.

8. the driver node sending a housekeeping request to the first worker node; in response to the housekeeping request, the first executor determines one or more outlier data segments present in a first cache associated with the first worker node, the one or more outlier data segments being determined at least in part based on one or more tokens of the set of tokens that correspond to the one or more outlier data segments that are outside of the first pair of token boundaries assigned to the first executor; the first execution unit transmitting identifiers of the one or more outlier data segments to the plurality of nodes; 2. The computer-implemented method of claim 1, further comprising: a second worker node of the plurality of nodes copying the one or more outlier data segments to a second cache associated with the second worker node based at least in part on the identifier.

9. the driver node sending a housekeeping request to the first worker node; determining, by the first execution unit, a set of valid data segments in response to the housekeeping request; determining, based at least in part on the set of valid data segments, one or more invalid data segments present in a first cache associated with the first worker node; removing the one or more invalid data segments from the first cache; the first worker node determining a current storage availability associated with the first worker node; If the current storage availability is less than a threshold: the first execution unit determining one or more target data segments present in the first cache; The computer-implemented method of claim 1 , further comprising: deleting the one or more target data segments.

10. The computer-implemented method of claim 9 , wherein the one or more target data segments are determined based at least in part on a set of segment temperatures.

11. 1. A distributed computing system for providing analytical data processing services, comprising: one or more processors; and one or more memories having computer executable instructions stored thereon, the instructions, when executed by one or more processors, providing the distributed computing system with: Running a cluster with multiple nodes, and causing the distributed computing system to maintain a state of the cluster, the state including a plurality of token boundaries uniformly associated with the plurality of nodes. causing a driver node of the plurality of nodes to receive a query for execution; identifying a set of one or more data segments corresponding to the query based at least in part on the query; causing the driver node to compute a set of tokens corresponding to the set of one or more data segments; launching, on a first execution unit executing on a first worker node of the plurality of nodes, a first task for processing a first data segment from the set of one or more data segments, the first worker node being selected based at least in part on a first token of the set of tokens corresponding to a first pair of token boundaries of the plurality of token boundaries, the first pair of token boundaries being associated with the first worker node; The first worker node retrieves the first data segment.

12. Maintaining the state of the cluster includes: storing a mapping of one or more execution units to the plurality of nodes; and assigning the plurality of token boundaries to the one or more execution units executing on the plurality of nodes based at least in part on the mapping, wherein the plurality of token boundaries are uniformly distributed among the one or more execution units, and the first pair of token boundaries of the plurality of token boundaries is assigned to a first execution unit of the one or more execution units.

13. When the computer-executable instructions are executed, the distributed computing system: the driver node receives an indication that the cluster has changed; updating the plurality of token boundaries in response to the indication; and 13. The distributed computing system of claim 12, further comprising: allocating the updated token boundaries to the one or more execution units, the updated token boundaries being uniformly distributed among the one or more execution units.

14. 14. The distributed computing system of claim 11, wherein obtaining the first data segment includes the first execution unit determining that the first data segment is present in a first cache associated with the first worker node, and reading the first data segment from the first cache.

15. Obtaining the first data segment comprises: the first execution unit determining that the first data segment is not present in a first cache associated with the first worker node; sending a request to one or more neighboring nodes of the plurality of nodes; In response to the request, a second execution unit executing on a second worker node of the one or more neighboring nodes determines that the first data segment is present in a second cache associated with the second worker node; and the first worker node copying the first data segment from the second cache to the first cache; and reading the first data segment from the first cache.

16. A computer program product configured to cause one or more processors to execute a method according to any one of claims 1 to 10.