Methods, systems, and computer program products for function scheduling on distributed systems
Patent Information
- Application Number
- CN202480086320.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2026-09-25
AI Technical Summary
然而,关于无服务器架构的设计,存在批判看法
Smart Images

Figure CN122826549A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to serverless computing, and in non-limiting embodiments or aspects, to methods, systems, and computer program products for efficiently scheduling serverless functions on distributed systems. Background Technology
[0002] Serverless computing is gaining popularity for applications seeking on-demand scalability, deployment abstraction, and / or unit billing. However, there are critical perspectives on the design of serverless architectures. In their 2018 paper, "Serverless Computing: One Step Forward, Two Steps Back," Hellerstein et al. argued that serverless architecture design inherently suffers from two problems: 1) Function as a Service (FaaS) is a data delivery architecture, and 2) FaaS hinders distributed computing (e.g., because serverless functions are not addressable, they cannot communicate directly with each other and instead use expensive (in terms of read / write throughput) storage operations to transfer data). Summary of the Invention
[0003] Therefore, improved methods, systems, and computer program products for function scheduling on distributed systems are provided.
[0004] According to a non-limiting embodiment or aspect, a method is provided, comprising: storing a plurality of data points and a mapping between a plurality of storage nodes on which the plurality of data points are stored using at least one processor; receiving a request to execute a function using the at least one processor, wherein the request to execute the function includes at least one identifier associated with at least one data point to be accessed by the function during the execution of the function; determining, using the at least one processor, at least one of the plurality of storage nodes that has stored or will store the at least one data point, based on the at least one identifier associated with the at least one data point and the mapping; determining, using the at least one processor, at least one of the plurality of computing nodes to schedule the execution of the function, based on distance data associated with a plurality of distances between the at least one storage node and the plurality of computing nodes; and scheduling the function using the at least one processor for execution on the at least one computing node.
[0005] In some non-limiting embodiments or aspects, the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
[0006] In some non-limiting embodiments or aspects, the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on adjustable parameters.
[0007] In some non-limiting embodiments or aspects, the distance data associated with the plurality of distances between the at least one storage node and the plurality of computing nodes includes a plurality of ping latencies between the at least one storage node and the plurality of computing nodes.
[0008] In some non-limiting embodiments or aspects, the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes multiple identifiers of multiple data points to be accessed by the function during the execution of the function, and wherein the multiple data points include one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
[0009] In some non-limiting embodiments or aspects, the method further includes: executing the function on the at least one computing node using the at least one processor, wherein executing the function on the at least one computing node causes the at least one data point to be accessed at the at least one storage node where the at least one data point has been or will be stored in the plurality of storage nodes.
[0010] In some non-limiting embodiments or aspects, the method further includes: updating the mapping between the plurality of data points and the plurality of storage nodes based on accessing the at least one data point at the at least one storage node among the plurality of storage nodes using the at least one processor.
[0011] According to some non-limiting embodiments or aspects, a system is provided, comprising: at least one processor coupled to a memory and configured to: store a plurality of data points and a mapping between a plurality of storage nodes on which the plurality of data points are stored; receive a request to execute a function, wherein the request to execute the function includes at least one identifier associated with at least one data point to be accessed by the function during the execution of the function; determine at least one storage node among the plurality of storage nodes that has stored or will store the at least one data point, based on the at least one identifier associated with the at least one data point and the mapping; determine at least one computing node among the plurality of computing nodes to be scheduled to execute the function, based on distance data associated with a plurality of distances between the at least one storage node and the plurality of computing nodes; and schedule the function for execution on the at least one computing node.
[0012] In some non-limiting embodiments or aspects, the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
[0013] In some non-limiting embodiments or aspects, the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on adjustable parameters.
[0014] In some non-limiting embodiments or aspects, the distance data associated with the plurality of distances between the at least one storage node and the plurality of computing nodes includes a plurality of ping latencies between the at least one storage node and the plurality of computing nodes.
[0015] In some non-limiting embodiments or aspects, the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes multiple identifiers of multiple data points to be accessed by the function during the execution of the function, and wherein the multiple data points include one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
[0016] In some non-limiting embodiments or aspects, the at least one processor is further configured to execute the function on the at least one computing node, wherein executing the function on the at least one computing node enables access to the at least one data point at the at least one storage node where the at least one data point has been or will be stored in the plurality of storage nodes.
[0017] In some non-limiting embodiments or aspects, the at least one processor is further configured to update the mapping between the plurality of data points and the plurality of storage nodes based on accessing the at least one data point at the at least one storage node among the plurality of storage nodes.
[0018] According to some non-limiting embodiments or aspects, a computer program product is provided comprising at least one non-transitory computer-readable medium, the at least one non-transitory computer-readable medium comprising program instructions that, when executed by at least one processor, cause the at least one processor to: store a plurality of data points and a mapping between a plurality of storage nodes on which the plurality of data points are stored; receive a request to execute a function, wherein the request to execute the function includes at least one identifier associated with at least one data point to be accessed by the function during the execution of the function; determine, based on the at least one identifier associated with the at least one data point and the mapping, at least one of the plurality of storage nodes that has stored or will store the at least one data point; determine, based on distance data associated with a plurality of distances between the at least one storage node and the plurality of computing nodes, at least one of the plurality of computing nodes to schedule the execution of the function; and schedule the function for execution on the at least one computing node.
[0019] In some non-limiting embodiments or aspects, the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
[0020] In some non-limiting embodiments or aspects, the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on adjustable parameters.
[0021] In some non-limiting embodiments or aspects, the distance data associated with the plurality of distances between the at least one storage node and the plurality of computing nodes includes a plurality of ping latencies between the at least one storage node and the plurality of computing nodes.
[0022] In some non-limiting embodiments or aspects, the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes multiple identifiers of multiple data points to be accessed by the function during the execution of the function, and wherein the multiple data points include one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
[0023] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, further cause the at least one processor to: execute the function on the at least one computing node, wherein executing the function on the at least one computing node causes access to the at least one data point at the at least one storage node where the at least one data point has been stored or will be stored in the plurality of storage nodes; and update the mapping between the plurality of data points and the plurality of storage nodes based on access to the at least one data point at the at least one storage node in the plurality of storage nodes.
[0024] Other non-limiting embodiments or aspects are set forth in the following numbered clauses: Clause 1. A method comprising: storing, using at least one processor, a mapping between a plurality of data points and a plurality of storage nodes on which the plurality of data points are stored; receiving, using the at least one processor, a request to execute a function, wherein the request to execute the function includes at least one identifier associated with at least one data point to be accessed by the function during the execution of the function; determining, using the at least one processor, at least one of the plurality of storage nodes that has stored or will store the at least one data point, based on the at least one identifier associated with the at least one data point and the mapping; determining, using the at least one processor, at least one of the plurality of computing nodes to schedule the execution of the function, based on distance data associated with a plurality of distances between the at least one storage node and the plurality of computing nodes; and scheduling, using the at least one processor, the function for execution on the at least one computing node.
[0025] Clause 2. The method according to Clause 1, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
[0026] Clause 3. The method according to Clause 1 or 2, wherein the at least one computing node to be scheduled to execute the function among the plurality of computing nodes is further determined based on adjustable parameters.
[0027] Clause 4. The method according to any one of Clauses 1 to 3, wherein the distance data associated with the plurality of distances between the at least one storage node and the plurality of compute nodes includes a plurality of ping delays between the at least one storage node and the plurality of compute nodes.
[0028] Clause 5. The method according to any one of Clauses 1 to 4, wherein the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes a plurality of identifiers of a plurality of data points to be accessed by the function during the execution of the function, and wherein the plurality of data points includes one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
[0029] Clause 6. The method according to any one of Clauses 1 to 5 further comprises: executing the function on the at least one computing node using the at least one processor, wherein executing the function on the at least one computing node causes the at least one data point to be accessed at the at least one storage node where the at least one data point has been or will be stored in the plurality of storage nodes.
[0030] Clause 7. The method according to any one of Clauses 1 to 6 further comprises: updating the mapping between the plurality of data points and the plurality of storage nodes based on accessing the at least one data point at the at least one storage node among the plurality of storage nodes using the at least one processor.
[0031] Clause 8. A system comprising: at least one processor coupled to a memory and configured to: store a plurality of data points and a mapping between a plurality of storage nodes on which the plurality of data points are stored; receive a request to execute a function, wherein the request to execute the function includes at least one identifier associated with at least one data point to be accessed by the function during the execution of the function; determine, based on the at least one identifier associated with the at least one data point and the mapping, at least one of the plurality of storage nodes that has stored or will store the at least one data point; determine, based on distance data associated with a plurality of distances between the at least one storage node and the plurality of computing nodes, at least one computing node among the plurality of computing nodes to be scheduled to execute the function; and schedule the function for execution on the at least one computing node.
[0032] Clause 9. In the system of Clause 8, the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
[0033] Clause 10. The system according to Clause 8 or 9, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on adjustable parameters.
[0034] Clause 11. The system according to any one of Clauses 8 to 10, wherein the distance data associated with the plurality of distances between the at least one storage node and the plurality of compute nodes includes a plurality of ping delays between the at least one storage node and the plurality of compute nodes.
[0035] Clause 12. The system according to any one of Clauses 8 to 11, wherein the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes a plurality of identifiers of a plurality of data points to be accessed by the function during the execution of the function, and wherein the plurality of data points includes one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
[0036] Clause 13. The system according to any one of Clauses 8 to 12, wherein the at least one processor is further configured to: execute the function on the at least one computing node, wherein executing the function on the at least one computing node causes the at least one data point to be accessed at the at least one storage node where the at least one data point has been or will be stored in the plurality of storage nodes.
[0037] Clause 14. The system according to any one of Clauses 8 to 13, wherein the at least one processor is further configured to: update the mapping between the plurality of data points and the plurality of storage nodes based on access to the at least one data point at the at least one storage node among the plurality of storage nodes.
[0038] Clause 15. A computer program product comprising at least one non-transitory computer-readable medium, said at least one non-transitory computer-readable medium comprising program instructions that, when executed by at least one processor, cause said at least one processor to: store a plurality of data points and a mapping between a plurality of storage nodes on which said plurality of data points are stored; receive a request to execute a function, wherein the request to execute said function includes at least one identifier associated with at least one data point to be accessed by said function during execution of said function; determine, based on said at least one identifier associated with said at least one data point and the mapping, at least one of the plurality of storage nodes that has stored or will store said at least one data point; determine, based on distance data associated with said at least one storage node and a plurality of computing nodes, at least one computing node to be scheduled to execute said function; and schedule said function for execution on said at least one computing node.
[0039] Clause 16. The computer program product according to Clause 15, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
[0040] Clause 17. The computer program product according to Clause 15 or 16, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on adjustable parameters.
[0041] Clause 18. A computer program product according to any one of Clauses 15 to 17, wherein the distance data associated with the plurality of distances between the at least one storage node and the plurality of computing nodes includes a plurality of ping delays between the at least one storage node and the plurality of computing nodes.
[0042] Clause 19. A computer program product according to any one of Clauses 15 to 18, wherein the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes a plurality of identifiers of a plurality of data points to be accessed by the function during the execution of the function, and wherein the plurality of data points includes one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
[0043] Clause 20. A computer program product according to any one of Clauses 15 to 19, wherein the program instructions, when executed by the at least one processor, further cause the at least one processor to: execute the function on the at least one computing node, wherein executing the function on the at least one computing node causes access to the at least one data point at the at least one storage node where the at least one data point has been stored or will be stored in the plurality of storage nodes; and update the mapping between the plurality of data points and the plurality of storage nodes based on access to the at least one data point at the at least one storage node in the plurality of storage nodes.
[0044] The operational methods and manufacturing economics of these and other features and characteristics of this disclosure, as well as combinations of related structural elements and parts, will become more apparent when considering the following description and appended claims with reference to the accompanying drawings, all of which form part of this specification, wherein similar reference numerals denote corresponding parts in the figures. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to be a definition of limitation on the disclosed subject matter. Attached Figure Description
[0045] Further advantages and details are explained in more detail below with reference to the non-limiting exemplary embodiments shown in the illustrative drawings, in which: Figure 1 It is a schematic diagram of a system for function scheduling on a distributed system according to some non-limiting embodiments or aspects; Figure 2 Based on some non-limiting embodiments or aspects Figure 1 A schematic diagram of example components of one or more devices; Figure 3 This is a flowchart of a method for function scheduling on a distributed system according to some non-limiting embodiments or aspects; Figure 4 This is a schematic diagram of an implementation scheme of a system for function scheduling on a distributed system, based on some non-limiting embodiments or aspects; Figure 5 These are schematic diagrams illustrating implementations of a system for function scheduling on a distributed system, based on some non-limiting embodiments or aspects; and Figure 6 This is a schematic diagram of an implementation scheme for a system for function scheduling on a distributed system, based on some non-limiting embodiments or aspects. Detailed Implementation
[0046] For the purposes of the following description, the terms “end,” “upper,” “lower,” “right,” “left,” “vertical,” “horizontal,” “top,” “bottom,” “lateral,” “longitudinal,” and their derivatives should be associated with the orientation of the embodiments in the accompanying drawings. However, it should be understood that various alternative variations and sequences of steps may be employed in this disclosure, except where explicitly specified otherwise. It should also be understood that the specific apparatus and processes shown in the accompanying drawings and described in the following specification are merely exemplary and non-limiting embodiments or aspects of the disclosed subject matter. Therefore, specific dimensions and other physical characteristics relating to the embodiments or aspects disclosed herein should not be considered limiting.
[0047] This document describes some non-limiting embodiments or aspects in conjunction with thresholds. As used herein, satisfying a threshold can refer to a value greater than a threshold, more than a threshold, higher than a threshold, greater than or equal to a threshold, less than a threshold, less than a threshold, lower than a threshold, less than or equal to a threshold, equal to a threshold, etc.
[0048] The terms "aspect," "component," "element," "structure," "action," "step," "function," and "instruction" used herein should not be construed as critical or essential unless explicitly stated otherwise. Furthermore, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more" and "at least one." Additionally, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.) and may be used interchangeably with "one or more" or "at least one." The term "a" or similar language is used where only one item is intended. Moreover, as used herein, the terms "has," "have," and "having" are intended to be open-ended terms. Additionally, unless explicitly stated otherwise, the phrase "based on" is intended to mean "at least partially based on." Furthermore, a reference to an action based on a condition may indicate that the action is "in response to" the condition. For example, in some non-limiting embodiments or aspects, the phrases “based on” and “in response to” may refer to conditions for automatically triggering actions (e.g., specific operations of electronic devices such as computing devices, processors, etc.).
[0049] As used herein, the term "communication" can refer to the receiving, accepting, sending, transmitting, providing, etc., of data (e.g., information, signals, messages, instructions, commands, etc.). For one unit (e.g., apparatus, system, component of an apparatus or system, combination thereof, etc.) to communicate with another unit means that the first unit is able to receive information directly or indirectly from and / or send information to the other unit. This can refer to a direct or indirect connection that is inherently wired and / or wireless (e.g., a direct communication connection, an indirect communication connection, etc.). Furthermore, the two units can communicate with each other even though the transmitted information may be modified, processed, relayed, and / or routed between them. For example, the first unit can communicate with the second unit even if it passively receives information and does not actively send information to the second unit. As another example, the first unit can communicate with the second unit if at least one intermediate unit processes information received from the first unit and transmits the processed information to the second unit. In some non-limiting embodiments or aspects, a message can refer to a network packet (e.g., a data packet, etc.) that includes data. It should be understood that many other arrangements are possible.
[0050] As used herein, the term "computing device" can refer to one or more electronic devices configured to process data. In some examples, a computing device may include essential components for receiving, processing, and outputting data, such as a processor, display, memory, input device, network interface, etc. A computing device can be a mobile device. As examples, a mobile device may include a cellular phone (e.g., a smartphone or standard cellular phone), a portable computer, a wearable device (e.g., a watch, glasses, lenses, clothing, etc.), a personal digital assistant (PDA), and / or other similar devices. A computing device may also be a desktop computer or other form of non-mobile computer.
[0051] As used herein, the term "server" may refer to or include one or more computing devices operated by or facilitating communication and processing among multiple parties in a network environment such as the Internet, but it should be understood that communication may be facilitated through one or more public or private network environments, and various other arrangements may be possible. Furthermore, multiple computing devices (e.g., servers, point-of-sale (POS) devices, mobile devices, etc.) communicating directly or indirectly in a network environment may constitute a "system".
[0052] As used herein, the term "system" may refer to one or more computing devices or a combination of computing devices (e.g., processor, server, client device, software application, components of such computing devices, etc.). As used herein, references to "device," "server," "processor," etc., may refer to a previously described device, server, or processor, different devices, servers, or processors, and / or combinations of devices, servers, and / or processors, described as performing a preceding step or function. For example, as used in the specification and claims, a first device, first server, or first processor described as performing a first step or a first function may refer to the same or different devices, servers, or processors described as performing a second step or a second function.
[0053] As used herein, the term "communication network" can refer to one or more wired and / or wireless networks. For example, a communication network can include cellular networks (e.g., Long Term Evolution (LTE®) networks, third-generation (3G) networks, fourth-generation (4G) networks, fifth-generation (5G) networks, Code Division Multiple Access (CDMA) networks, etc.), Public Land Mobile Networks (PLMNs), Local Area Networks (LANs), Wide Area Networks (WANs), Metropolitan Area Networks (MANs), Telephone Networks (e.g., Public Switched Telephone Networks (PSTN)), Private Networks, Self-organizing Networks, Intranets, the Internet, Fiber-based Networks, Cloud Computing Networks, etc., and / or combinations of these or other types of networks.
[0054] Non-limiting embodiments or aspects of this disclosure may provide a method, system, and / or computer program product for: storing a plurality of data points and a mapping between a plurality of storage nodes on which the plurality of data points are stored; receiving a request to execute a function, wherein the request to execute the function includes at least one identifier associated with at least one data point to be accessed by the function during the execution of the function; determining at least one storage node among the plurality of storage nodes that has stored or will store the at least one data point based on the at least one identifier associated with the at least one data point and the mapping; determining at least one computing node among the plurality of computing nodes to schedule the execution of the function based on distance data associated with a plurality of distances between the at least one storage node and the plurality of computing nodes; and scheduling the function for execution on the at least one computing node.
[0055] In this way, non-limiting embodiments or aspects of this disclosure can enable ensuring that functions are scheduled for computations as close to storage as possible. This helps address the first problem of Function as a Service (FaaS) being a data delivery architecture, and / or can enable improved or optimal scheduling of a group of functions that can be part of the same call stack (e.g., so-called sequential calls or calls in some predefined order). This helps address the second problem of FaaS hindering distributed computing and / or reducing network overhead, as functions can communicate indirectly, allowing functions to be scheduled closer to each other. Therefore, non-limiting embodiments or aspects of this disclosure can reinstate the fundamental principle of distributed computing, namely the co-location of computation and storage, while ensuring that the flexibility offered by serverless computing is not lost, and / or help developers define a way to ensure that related functions or function calls that are part of the same call stack can be located close to each other. This can reduce the cost of communicating via expensive database reads / writes by placing functions closer to shared storage nodes and reducing network latency.
[0056] Figure 1 This is a schematic diagram of a system 100 for function scheduling on a distributed system, based on some non-limiting embodiments or aspects. For example... Figure 1As shown, system 100 may include a function scheduler system 102, a storage information node 104, multiple computing nodes C1, C2, C3, C4, C5, ... CN, multiple storage nodes storage 1, storage 2, ... storage N, and / or multiple clients SR1, SR2, ... SRN. The function scheduler system 102, storage information node 104, the multiple computing nodes C1, C2, C3, C4, C5, ... CN, the multiple storage nodes storage 1, storage 2, ... storage N, and / or the multiple clients SR1, SR2, ... SRN may be interconnected via wired connections, wireless connections, or a combination of wired and wireless connections (e.g., establishing connections for communication).
[0057] The function scheduler system 102 may include one or more devices capable of receiving information and / or data (e.g., via a communication network, etc.) from storage information node 104, multiple computing nodes C1, C2, C3, C4, C5, ... CN, multiple storage nodes storage 1, storage 2, ..., storage N, and / or multiple client SR1, SR2, ... SRN, and / or transmitting information and / or data (e.g., via a communication network, etc.) to storage information node 104, multiple computing nodes C1, C2, C3, C4, C5, ... CN, multiple storage nodes storage 1, storage 2, ..., storage N, and / or multiple client SR1, SR2, ... SRN. For example, the function scheduler system 102 may include computing devices, such as servers, server groups, and / or other similar devices.
[0058] The function scheduler system 102 may include a highly available entity that acts as an interface between multiple client SR1, SR2, ..., SRN that can generate requests to execute functions and multiple compute nodes C1, C2, C3, C4, C5, ... CN or a distributed cluster on which the requested functions can be executed. The function scheduler system 102 may be responsible for scheduling functions for execution on the multiple compute nodes C1, C2, C3, C4, C5, ... CN based on input from the storage information node 104 and / or the function definition itself. The function scheduler system 102 may also use an adjustable parameter α to influence the strictness of the policy applied to attempt to schedule compute nodes as close as possible to the storage node to be accessed in order to schedule functions for execution on the multiple compute nodes C1, C2, C3, C4, C5, ... CN.
[0059] Storage information node 104 may include one or more devices capable of receiving information and / or data (e.g., via a communication network, etc.) from function scheduler system 102, multiple computing nodes C1, C2, C3, C4, C5, ... CN, multiple storage nodes storage 1, storage 2, ... storage N, and / or multiple clients SR1, SR2, ... SRN, and / or transmitting information and / or data (e.g., via a communication network, etc.) to function scheduler system 102, multiple computing nodes C1, C2, C3, C4, C5, ... CN, multiple storage nodes storage 1, storage 2, ... storage N, and / or multiple clients SR1, SR2, ... SRN. For example, storage information node 104 may include computing devices, such as servers, server groups, and / or other similar devices. In some non-limiting embodiments or aspects, storage information node 104 may be included in and / or implemented by function scheduler system 102.
[0060] Storage information node 104 can be configured to store storage information for multiple storage nodes storage 1, storage 2, ..., storage N. For example, storage information node 104 can maintain a mapping between multiple storage nodes storage 1, storage 2, ..., storage N and multiple data points, so that when storage information node 104 is queried (e.g., by function scheduler system 102, etc.), storage information node 104 can provide storage node information indicating which data point is stored on which storage node.
[0061] Multiple computing nodes C1, C2, C3, C4, C5, ... CN may include one or more devices capable of receiving information and / or data (e.g., via a communication network, etc.) from the function scheduler system 102, storage information node 104, multiple storage nodes storage 1, storage 2, ... storage N, and / or multiple client SR1, SR2, ... SRN, and / or transmitting information and / or data (e.g., via a communication network, etc.) to the function scheduler system 102, storage information node 104, multiple storage nodes storage 1, storage 2, ... storage N, and / or multiple client SR1, SR2, ... SRN. For example, multiple computing nodes C1, C2, C3, C4, C5, ... CN can be implemented in a distributed system, wherein each node of the multiple computing nodes C1, C2, C3, C4, C5, ... CN can be implemented within a single device and / or system, or distributed across multiple devices and / or systems in the distributed system. For example, nodes among the multiple computing nodes C1, C2, C3, C4, C5, ... CN may include computing devices and / or may be implemented by computing devices, such as servers, server clusters, and / or other similar devices. As an example, one or more nodes among the multiple computing nodes C1, C2, C3, C4, C5, ... CN (e.g., implementing the one or more nodes among the multiple computing nodes C1, C2, C3, C4, C5, ... CN, or including one or more computing devices in the one or more nodes) may be located in different physical locations from one or more other nodes among the multiple computing nodes C1, C2, C3, C4, C5, ... CN (e.g., implementing the one or more other nodes among the multiple computing nodes C1, C2, C3, C4, C5, ... CN, or including one or more other computing devices in the one or more other nodes).
[0062] In some non-limiting embodiments or aspects, the plurality of computing nodes C1, C2, C3, C4, C5...CN includes a plurality of heterogeneous nodes. For example, different nodes among the plurality of computing nodes C1, C2, C3, C4, C5,...CN may include different types of hardware, firmware, or combinations of hardware and software (e.g., random access memory (RAM) of different types and / or speeds, processors of different types and / or speeds, etc.).
[0063] Multiple computing nodes C1, C2, C3, C4, C5, ... CN can be configured to execute functions. For example, a computing node among the multiple computing nodes C1, C2, C3, C4, C5, ... CN can be configured to execute functions assigned to the nodes by the function scheduler system 102.
[0064] Multiple storage nodes (Storage 1, Storage 2, ..., Storage N) may include one or more devices capable of receiving information and / or data (e.g., via a communication network, etc.) from the function scheduler system 102, storage information node 104, multiple computing nodes C1, C2, C3, C4, C5, ..., CN, and / or multiple client SR1, SR2, ..., SRN, and / or transmitting information and / or data (e.g., via a communication network, etc.) to the function scheduler system 102, storage information node 104, multiple computing nodes C1, C2, C3, C4, C5, ..., CN, and / or multiple client SR1, SR2, ..., SRN. For example, multiple storage nodes (Storage 1, Storage 2, ..., Storage N) can be implemented in a distributed system, wherein each node in the multiple storage nodes (Storage 1, Storage 2, ..., Storage N) can be implemented within a single device and / or system or distributed across multiple devices and / or systems in the distributed system. For example, nodes in a plurality of storage nodes (Storage 1, Storage 2, ..., Storage N) may include computing devices and / or may be implemented by computing devices, such as servers, server clusters, and / or other similar devices. As an example, one or more nodes in a plurality of storage nodes (Storage 1, Storage 2, ..., Storage N) (e.g., implementing one or more nodes in a plurality of storage nodes (Storage 1, Storage 2, ..., Storage N) or including one or more computing devices in said one or more nodes) may be located in different physical locations from one or more other nodes in a plurality of storage nodes (e.g., implementing one or more other nodes in a plurality of storage nodes (Storage 1, Storage 2, ..., Storage N) or including one or more other computing devices in said one or more other nodes).
[0065] In some non-limiting embodiments or aspects, the multiple storage nodes storage 1, storage 2, ... storage N include multiple heterogeneous nodes. For example, different nodes in the multiple storage nodes storage 1, storage 2, ... storage N may include different types of hardware, firmware, or combinations of hardware and software (e.g., random access memory (RAM) of different types and / or speeds, processors of different types and / or speeds, etc.).
[0066] Multiple storage nodes (Store 1, Storage 2, ..., Storage N) can be configured to store multiple data points. For example, a function executed by a compute node among multiple compute nodes C1, C2, C3, C4, C5, ..., CN can access (e.g., read, write, etc.) one or more of the multiple data points at the multiple storage nodes (Store 1, Storage 2, ..., Storage N).
[0067] Data points can include discrete units of information, such as data blocks comprising sequences of bits or bytes, which may contain an integer number of records, have a maximum length (e.g., block size), and so on. For example, a data block can include a sequence of data in bit or byte form that can be transmitted as a whole. Multiple data points can include multiple different types of data. In some non-limiting embodiments or aspects, multiple data points can include transaction data associated with multiple transactions processed in an electronic payment network. For example, transaction data can include parameters associated with the transaction, such as account identifiers (e.g., PANs), transaction amounts, transaction dates and times, types of products and / or services associated with the transaction, currency exchange rates, currency types, merchant types, merchant names, merchant locations, transaction approval (and / or rejection) rates, etc.
[0068] Multiple client SR1, SR2, ... SRNs (e.g., multiple client systems, multiple client devices, etc.) may include one or more devices capable of receiving information and / or data (e.g., via a communication network, etc.) from the function scheduler system 102, storage information node 104, multiple computing nodes C1, C2, C3, C4, C5, ... CN, and / or multiple storage nodes storage 1, storage 2, ... storage N, and / or transmitting information and / or data (e.g., via a communication network, etc.) to the function scheduler system 102, storage information node 104, multiple computing nodes C1, C2, C3, C4, C5, ... CN, and / or multiple storage nodes storage 1, storage 2, ... storage N. For example, multiple client SR1, SR2, ... SRNs may include computing devices, such as servers, server groups, and / or other similar devices.
[0069] Each of the multiple clients SR1, SR2, ... SRN can be configured to provide a request to the function scheduler system 102 for execution of a function (e.g., a software function).
[0070] like Figure 1 The diagram further illustrates that the function scheduler system 102, multiple compute nodes C1, C2, C3, C4, C5, ... CN, and / or multiple storage nodes Storage 1, Storage 2, ... Storage N can provide or implement a background daemon service "Storage Update". This service is configured to monitor the reading and writing of data points at multiple storage nodes Storage 1, Storage 2, ... Storage N by functions executed through multiple compute nodes C1, C2, C3, C4, C5, ... CN, and to provide indications or logs of reading and writing to storage information node 104 via the function scheduler system 102, so as to update the mapping between multiple data points and multiple storage nodes Storage 1, Storage 2, ... Storage N based on reading and writing, for example, so that the function scheduler system 102 can take this information into account when making decisions for future function execution requests.
[0071] In some non-limiting embodiments or aspects, the background daemon service "Storage Update" can provide multiple load or capacity factors associated with multiple storage nodes Storage 1, Storage 2, ... Storage N. For example, the load or capacity factor of a storage node can include the ratio of the number of data blocks stored at the node to the number of addresses within the storage node (e.g., the number of storage blocks, the number of memory locations, etc.). As an example, multiple load or capacity factors can be used to instruct the function scheduler system 102 to take into account the load or capacity factors of each storage node when scheduling functions for execution on multiple compute nodes C1, C2, C3, C4, C5, ... CN.
[0072] like Figure 1 The diagram further illustrates that the function scheduler system 102 and / or multiple storage nodes, storage 1, storage 2, ..., storage N, can provide or implement a background daemon service called "status update". This service is configured to monitor the health status of multiple compute nodes C1, C2, C3, C4, C5, ..., CN and provide the function scheduler system 102 with associated node health data or parameters. This allows the function scheduler system 102 to take this information into account when making decisions about function scheduling. For example, unhealthy compute nodes or those not reaching acceptable or optimal capacity should not be assigned more function execution requests.
[0073] Node health data or parameters may include, and / or the function scheduler system 102 may determine read latency and / or write latency based on node health data, including the latency (e.g., amount of time, etc.) introduced or incurred by the compute node when servicing a request to access a data block at the storage node (e.g., the amount of time used by the compute node to servicing the request, the amount of time spent by the compute node processing, etc.). For example, node health data may include, and / or the function scheduler system 102 may determine average read and / or write latency based on node health data, including the average latency (e.g., average amount of time, etc.) introduced or incurred by the compute node when servicing one or more requests to access one or more data blocks at the storage node (e.g., the average amount of time used by the compute node to servicing multiple requests from one or more clients, the average amount of time spent by the compute node processing, etc.). As an example, for each of the multiple compute nodes C1, C2, C3, C4, C5, ... CN, the node health data associated with multiple read and / or write latencies associated with the multiple compute nodes C1, C2, C3, C4, C5, ... CN may include read and / or write latencies, including the average amount of time spent by the node service accessing one or more data blocks at one or more storage nodes.
[0074] Now for reference Figure 2 The diagram illustrates example components of a device 200 according to a non-limiting embodiment. As an example, device 200 may correspond to a function scheduler system 102, a storage information node 104, one or more computing nodes C1, C2, C3, C4, C5, ... CN, one or more storage nodes Storage 1, Storage 2, ... Storage N, and / or one or more clients SR1, SR2, ... SRN. In some non-limiting embodiments or aspects, such a system or device may include at least one device 200 and / or at least one component of device 200. The number and arrangement of components shown are provided as examples. In some non-limiting embodiments or aspects, device 200 may include additional components, fewer components, different components, or components arranged in a manner different from those illustrated. Additionally or alternatively, a set of components of device 200 (e.g., one or more components) may perform one or more functions described as being performed by another set of components of device 200.
[0075] like Figure 2 As shown, device 200 may include bus 202, processor 204, memory 206, storage unit 208, input unit 210, output unit 212, and communication interface 214. Bus 202 may include components that allow communication between components of device 200. In some non-limiting embodiments, processor 204 may be implemented in hardware, firmware, or a combination of hardware and software. For example, processor 204 may include processors (e.g., central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), etc.), microprocessors, digital signal processors (DSPs), and / or any processing unit that can be programmed to perform functions (e.g., field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), etc.). Memory 206 may include random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, optical memory, etc.) that stores information and / or instructions for use by processor 204.
[0076] Continue to refer to Figure 2Storage component 208 may store information and / or software related to the operation and use of device 200. For example, storage component 208 may include a hard disk (e.g., magnetic disk, optical disk, magneto-optical disk, solid-state disk, etc.) and / or another type of computer-readable medium. Input component 210 may include components that allow device 200 to receive information, such as via user input (e.g., touch screen display, keyboard, keypad, mouse, buttons, switches, microphone, etc.). Alternatively, input component 210 may include sensors for sensing information (e.g., Global Positioning System (GPS) components, accelerometers, gyroscopes, actuators, etc.). Output component 212 may include components that provide output information from device 200 (e.g., display screen, speaker, one or more light-emitting diodes (LEDs), etc.). Communication interface 214 may include transceiver components (e.g., transceiver, separate receiver and transmitter, etc.) that enable device 200 to communicate with other devices, for example, via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 214 may allow device 200 to receive information from another device and / or provide information to another device. For example, communication interface 214 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi® interface, a cellular network interface, etc.
[0077] Apparatus 200 can perform one or more processes described herein. Apparatus 200 can perform these processes based on processor 204 executing software instructions stored in a computer-readable medium, such as memory 206 and / or storage unit 208. The computer-readable medium may include any non-transient memory device. The memory device includes memory space located within a single physical storage device or memory space extended across multiple physical storage devices. Software instructions may be read into memory 206 and / or storage unit 208 via communication interface 214 from another computer-readable medium or from another device. When executed, the software instructions stored in memory 206 and / or storage unit 208 may cause processor 204 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in conjunction with the software instructions to perform one or more processes described herein. Therefore, the embodiments described herein are not limited to any particular combination of hardware circuitry and software. As used herein, the term “configured to” may refer to an arrangement of software, apparatus, and / or hardware for performing and / or implementing one or more functions (e.g., actions, processes, steps of processes, etc.). For example, "a processor configured to..." can refer to a processor that executes software instructions (such as program code) that cause the processor to execute one or more functions.
[0078] Now for reference Figure 3The diagram illustrates a flowchart of a method for function scheduling on a distributed system according to some non-limiting embodiments or aspects. Figure 3 The steps shown are for illustrative purposes only. It will be understood that additional, fewer, different, and / or different orders of steps may be used in some non-limiting embodiments or aspects. In some non-limiting embodiments or aspects, steps may be performed automatically in response to the execution and / or completion of previous steps.
[0079] like Figure 3 As shown, at step 302, method 300 includes storing a mapping between multiple data points and multiple storage nodes on which the multiple data points are stored. For example, storage information node 104 may store a mapping between multiple data points and multiple storage nodes (storage 1, storage 2, ..., storage N) on which the multiple data points are stored. As an example, function scheduler system 102 may store (and / or update) the mapping between multiple data points and multiple storage nodes (storage 1, storage 2, ..., storage N) on which the multiple data points are stored at storage information node 104. In such examples, storage information node 104 may act as a bookkeeping entity to maintain information about which storage node stores what data (e.g., mappings, etc.).
[0080] like Figure 3 As shown, at step 304, method 300 includes receiving a request to execute a function. For example, function scheduler system 102 may receive a request to execute a function (e.g., a software function, etc.). As an example, function scheduler system 102 may receive a request to execute a function (e.g., F1, etc.) from a client among multiple client SR1, SR2, ... SRNs. In such examples, the request to execute a function may include at least one identifier (e.g., P1, P2, etc.) associated with at least one data point that the function will access during function execution. For example, a request to run / execute a function on a distributed cluster may be received by function scheduler system 102, and / or a function requesting execution on a distributed cluster may declare external data connections (e.g., reads, writes, etc.) planned to be performed by the function throughout its execution lifecycle by informing function scheduler system 102 of the details of the precise data points that the function plans to access before execution. This information enables function scheduler system 102 to consider the type of data the function is expected to access and make scheduling decisions based on said information.
[0081] In some non-limiting embodiments or aspects, at least one identifier of at least one data point to be accessed by the function during function execution includes multiple identifiers of multiple data points to be accessed by the function during function execution, and / or multiple data points include one or more data points to be read by the function during function execution and one or more data points to be written by the function during function execution. For example, a request to execute a function may include an indication of at least one storage node about which at least one data point will be written by the function during function execution. For example, the sharing of data access information may not be limited to data blocks read by the function, but may also include information about writes planned by the function, such that any future function execution requests that depend on data previously written by the function can be intelligently scheduled to be close to nodes where data has been previously written by the function, thereby helping to offset the second problem mentioned in the background section, namely, "FaaS hinders distributed computing". Because functions can precisely identify the reads and writes planned by the function during its lifecycle, the function scheduler system 102 can take this information into account to ensure that computation is as close to storage as possible, which is a fundamental principle of distributed computing.
[0082] like Figure 3 As shown, at step 306, method 300 includes determining at least one storage node among a plurality of storage nodes that has stored or will store the at least one data point, based on the at least one identifier associated with the at least one data point and the mapping. For example, function scheduler system 102 may determine at least one storage node among a plurality of storage nodes (storage 1, storage 2, ..., storage N) that has stored or will store the at least one data point, based on the at least one identifier associated with the at least one data point and the mapping. As an example, function scheduler system 102 may query storage information node 104 based on the at least one identifier associated with the at least one data point, and in response to the query including the at least one identifier, storage information node 104 may return at least one storage node among a plurality of storage nodes (storage 1, storage 2, ..., storage N) that has stored or will store the at least one data point. In such examples, declared external data connections that a function plans to make throughout its execution lifetime are available to function scheduler system 102 for scheduling functions for execution on at least one compute node. For example, once the function scheduler system 102 shares details of the precise data points accessed by the function plan with the storage information node 104, the storage information node 104 can indicate the location or node identifier of the at least one node containing those data points.
[0083] like Figure 3As shown, at step 308, method 300 includes determining at least one compute node among the plurality of compute nodes to schedule the execution of a function based on distance data associated with multiple distances between the at least one storage node and the plurality of compute nodes. For example, function scheduler system 102 may determine at least one compute node among the plurality of compute nodes C1, C2, C3, C4, C5, ... CN to schedule the execution of a function based on distance data associated with multiple distances between the at least one storage node and the plurality of compute nodes C1, C2, C3, C4, C5, ... CN.
[0084] In some non-limiting embodiments or aspects, the distance data associated with multiple distances between the at least one storage node and the plurality of compute nodes C1, C2, C3, C4, C5, ... CN includes multiple ping latencies between the at least one storage node and the plurality of compute nodes C1, C2, C3, C4, C5, ... CN. For example, the distance between a compute node and a storage node can be calculated based on the ping latencies measured between the compute node and the storage node, and can be a relative calculation (e.g., the higher the ping latency between two nodes, the greater the distance between the two nodes, etc.).
[0085] In some non-limiting embodiments or aspects, the at least one computing node among the plurality of computing nodes to be scheduled to execute a function is further determined based on the node capacities associated with the plurality of computing nodes C1, C2, C3, C4, C5, ... CN, the node health parameters associated with the plurality of computing nodes C1, C2, C3, C4, C5, ... CN, one or more adjustable parameters, or any combination thereof. For example, the function scheduler system 102 may determine the at least one computing node among the plurality of computing nodes C1, C2, C3, C4, C5, ... CN to be scheduled to execute a function based on distance data associated with multiple distances between the at least one storage node and the plurality of computing nodes C1, C2, C3, C4, C5, ... CN, the node capacities associated with the plurality of computing nodes C1, C2, C3, C4, C5, ... CN, the node health parameters associated with the plurality of computing nodes C1, C2, C3, C4, C5, ... CN, one or more adjustable parameters, or any combination thereof. As an example, the function scheduler system 102 can determine the function to be scheduled for execution among multiple computing nodes C1, C2, C3, C4, C5, ... CN according to the following equation. At least one computing node:
[0086] in It is a computing node. It is a storage node. It is the first adjustable parameter. It is the second adjustable parameter. It is the third adjustable parameter.
[0087] Adjustable parameters You can configure or control how close compute nodes and storage should be to each other for the strictness of function execution. For example, also refer to... Figure 4 , Figure 4 This is a schematic diagram of an implementation 400 of a system for function scheduling on a distributed system, based on some non-limiting embodiments or aspects. If α is high (e.g., close to or equal to 1), then the function scheduler system 102 can schedule functions with a lower or least stringent degree. This could mean that the function scheduler system 102 may not assign much weight to the distance between the storage node the function will use and the potential computing node on which the function can run or execute. As an example, if α equals 1, such as Figure 4 As shown, the function scheduler system 102 can use more lenient or most lenient requirements to locate computation and storage together, where the function scheduler system 102 may not assign any weight to co-location when making scheduling decisions. In such examples, Running a distributed cluster with =1 can be a manifestation of the existing behavior of scheduling serverless functions, regardless of the stored information.
[0088] Now for reference Figure 5 , Figure 5 This is a schematic diagram of an implementation 500 of a system for function scheduling on a distributed system, based on some non-limiting embodiments or aspects. If α equals 0.5, then the function scheduler system 102 can use moderate requirements to locate computation and storage together, such that for a given storage node, the function scheduler system 102 can schedule... Figure 5 The function on any of the three computation nodes shown is compared to Figure 4 The implementation scheme 400 has a more lenient requirement, and this may help the function scheduler system 102 to be more flexible in making scheduling decisions, because the function scheduler system 102 can have more options to schedule functions and assign more weight to other parameters, such as the node health and node capacity of the expanded set of compute nodes, in making scheduling decisions.
[0089] Now for reference Figure 6 , Figure 6 This is a schematic diagram of an implementation scheme 600 for a system of function scheduling on a distributed system, based on some non-limiting embodiments or aspects. If α is low (e.g., close to or equal to 0), then the function scheduler system 102 can schedule functions with a higher or highest strictness. This could mean that the function scheduler system 102 can assign greater or maximum weight to the distance between the storage node to which the function will use and the potential computation node on which the function can run or execute (e.g., to locate storage and computation together and / or ensure that the function is scheduled to a node as close as possible to the node storing the information for the function to use). As an example, if α equals 0, such as Figure 6 As shown, the function scheduler system 102 can use strict requirements to locate computation and storage as close as possible, so that for a given storage node, the function scheduler system 102 can only schedule... Figure 4 The function is shown on either of the two compute nodes closest to the specific storage node. Therefore, in this most stringent scenario where α equals 0, the requirement may violate the server anonymity property of serverless architecture.
[0090] Second adjustable parameter The weights of node health information or parameters, and / or a third adjustable parameter, can be set or controlled during scheduling. The weight assigned to node capacity during scheduling functions can be set or controlled.
[0091] like Figure 3 As shown, at step 310, method 300 includes scheduling a function for execution on the at least one compute node. For example, function scheduler system 102 can schedule a function for execution on at least one compute node among a plurality of compute nodes C1, C2, C3, C4, C5, ... CN.
[0092] like Figure 3 As shown, at step 312, method 300 includes executing a function on the at least one compute node. For example, the at least one compute node among a plurality of compute nodes C1, C2, C3, C4, C5, ... CN can execute the function. As an example, function scheduler system 102 can cause the at least one compute node among a plurality of compute nodes C1, C2, C3, C4, C5, ... CN to execute functions in a scheduled order or relative to one or more other functions scheduled to be executed on the at least one compute node. In such examples, executing a function on the at least one compute node can enable access (e.g., reading, writing, etc.) of the at least one data point at the at least one storage node where the at least one data point has been stored or will be stored in a plurality of storage nodes.
[0093] like Figure 3As shown, at step 314, method 300 includes updating the mapping between multiple data points and multiple storage nodes based on access to the at least one data point at the at least one storage node among multiple storage nodes. For example, storage information node 104 may update the mapping between multiple data points and multiple storage nodes storage 1, storage 2, ... storage N based on access to the at least one data point at the at least one storage node among multiple storage nodes storage 1, storage 2, ... storage N. As an example, function scheduler system 102 may update the mapping between multiple data points and multiple storage nodes storage 1, storage 2, ... storage N at storage information node 104. In such examples, the background daemon service "Storage Update" may provide function scheduler system 102 with indications or logs of access to the at least one data point at the at least one storage node among multiple storage nodes storage 1, storage 2, ... storage N (e.g., where reads and writes are performed by executed functions, etc.). In some non-limiting embodiments or aspects, it may be assumed that functions utilize external memory to transfer processing results.
[0094] Although embodiments have been described in detail for illustrative purposes, it should be understood that such details are for the purposes described only, and this disclosure is not limited to the disclosed embodiments or aspects, but rather is intended to cover modifications and equivalent arrangements that fall within the spirit and scope of the appended claims. For example, it should be understood that this disclosure contemplates, as far as possible, that one or more features of any embodiment or aspect may be combined with one or more features of any other embodiment or aspect.
Claims
1. A method comprising: A mapping is used between storing multiple data points and multiple storage nodes storing the multiple data points on at least one processor; The at least one processor receives a request to execute a function, wherein the request to execute the function includes at least one identifier associated with at least one data point that the function wants to access during the execution of the function; Based on the at least one identifier associated with the at least one data point and the mapping, the at least one processor determines at least one storage node among the plurality of storage nodes that has stored or will store the at least one data point; Based on distance data associated with multiple distances between the at least one storage node and the plurality of computing nodes, the at least one processor determines at least one computing node among the plurality of computing nodes to be scheduled to execute the function; as well as The function is scheduled to be executed on the at least one computing node using the at least one processor.
2. The method of claim 1, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
3. The method of claim 2, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on adjustable parameters.
4. The method of claim 1, wherein the distance data associated with the plurality of distances between the at least one storage node and the plurality of computing nodes includes a plurality of ping delays between the at least one storage node and the plurality of computing nodes.
5. The method of claim 1, wherein the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes a plurality of identifiers of a plurality of data points to be accessed by the function during the execution of the function, and wherein the plurality of data points includes one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
6. The method of claim 1, further comprising: The function is executed on the at least one computing node using the at least one processor, wherein executing the function on the at least one computing node enables access to the at least one data point at the at least one storage node where the at least one data point has been or will be stored in the plurality of storage nodes.
7. The method of claim 6, further comprising: Based on accessing the at least one data point at the at least one storage node among the plurality of storage nodes, the at least one processor updates the mapping between the plurality of data points and the plurality of storage nodes.
8. A system comprising: At least one processor, said at least one processor being coupled to memory and configured to: A mapping between storing multiple data points and multiple storage nodes on which the multiple data points are stored; Receive a request to execute a function, wherein the request to execute the function includes at least one identifier associated with at least one data point that the function wants to access during the execution of the function; Based on the at least one identifier associated with the at least one data point and the mapping, determine at least one storage node among the plurality of storage nodes that has stored or will store the at least one data point; Based on distance data associated with multiple distances between the at least one storage node and the plurality of computing nodes, at least one computing node among the plurality of computing nodes is determined to be scheduled to execute the function; as well as The function is scheduled for execution on the at least one computing node.
9. The system of claim 8, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
10. The system of claim 9, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on adjustable parameters.
11. The system of claim 8, wherein the distance data associated with the plurality of distances between the at least one storage node and the plurality of computing nodes includes a plurality of ping delays between the at least one storage node and the plurality of computing nodes.
12. The system of claim 8, wherein the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes a plurality of identifiers of a plurality of data points to be accessed by the function during the execution of the function, and wherein the plurality of data points includes one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
13. The system of claim 8, wherein the at least one processor is further configured to: The function is executed on the at least one computing node, wherein executing the function on the at least one computing node enables access to the at least one data point at the at least one storage node where the at least one data point has been or will be stored in the plurality of storage nodes.
14. The system of claim 13, wherein the at least one processor is further configured to: Based on accessing the at least one data point at the at least one of the plurality of storage nodes, the mapping between the plurality of data points and the plurality of storage nodes is updated.
15. A computer program product comprising at least one non-transitory computer-readable medium, the at least one non-transitory computer-readable medium comprising program instructions that, when executed by at least one processor, cause the at least one processor to: A mapping between storing multiple data points and multiple storage nodes on which the multiple data points are stored; Receive a request to execute a function, wherein the request to execute the function includes at least one identifier associated with at least one data point that the function wants to access during the execution of the function; Based on the at least one identifier associated with the at least one data point and the mapping, determine at least one storage node among the plurality of storage nodes that has stored or will store the at least one data point; Based on distance data associated with multiple distances between the at least one storage node and the plurality of computing nodes, at least one computing node among the plurality of computing nodes is determined to be scheduled to execute the function; as well as The function is scheduled for execution on the at least one computing node.
16. The computer program product of claim 15, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on the plurality of node capacities associated with the plurality of computing nodes and the plurality of node health parameters associated with the plurality of computing nodes.
17. The computer program product of claim 16, wherein the at least one computing node among the plurality of computing nodes to be scheduled to execute the function is further determined based on adjustable parameters.
18. The computer program product of claim 15, wherein the distance data associated with the plurality of distances between the at least one storage node and the plurality of computing nodes includes a plurality of ping delays between the at least one storage node and the plurality of computing nodes.
19. The computer program product of claim 15, wherein the at least one identifier of at least one data point to be accessed by the function during the execution of the function includes a plurality of identifiers of a plurality of data points to be accessed by the function during the execution of the function, and wherein the plurality of data points includes one or more data points to be read by the function during the execution of the function and one or more data points to be written by the function during the execution of the function.
20. The computer program product of claim 15, wherein the program instructions, when executed by the at least one processor, further cause the at least one processor to: Executing the function on the at least one computing node, wherein executing the function on the at least one computing node causes the at least one data point to be accessed at the at least one storage node where the at least one data point has been or will be stored in the plurality of storage nodes; and Based on accessing the at least one data point at the at least one of the plurality of storage nodes, the mapping between the plurality of data points and the plurality of storage nodes is updated.