Dispatcher

The method and system efficiently manage tasks in a distributed computing environment by using a dispatcher to assign tasks based on node properties and task requirements, addressing inefficiencies and security vulnerabilities in decentralized systems.

GB2637665APending Publication Date: 2025-08-06GREEN CLOUD COMPUTING LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
GB2022008778
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-08-06

AI Technical Summary

Technical Problem

Existing distributed computing environments face challenges in efficiently assigning tasks to appropriate nodes, particularly in decentralized systems where nodes can join or leave freely, leading to inefficiencies and potential security vulnerabilities.

Method used

A method and system for managing tasks in a distributed computing environment using a dispatcher to assign tasks to channels based on node properties and task requirements, allowing nodes to join or leave without disrupting the system, and ensuring tasks are executed by nodes with the appropriate capabilities and availability.

Benefits of technology

This approach enables efficient task scheduling, reduces computational inefficiencies, and enhances security by ensuring tasks are assigned to suitable nodes, even as nodes dynamically join or leave the network, while promoting processing using renewable energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for managing a task in a distributed computing environment comprising: one or more nodes configured to execute a task, a plurality of channels, and a dispatcher configured to assign the task to a channel from the plurality of channels. The method comprising: determining at least one property of each node; assigning each node to a channel of the plurality of channels dependent on the at least one property of each node; monitoring each node as they join and / or leave the distributed computing environment; receiving a task at the dispatcher from an external device; assigning, at the dispatcher, the task to a channel selected from the plurality of channels, wherein the channel is selected dependent on at least one property of the task; monitoring an availability of a subset of the nodes, wherein the subset of the nodes are assigned to the selected channel; determining that a node from the subset of the nodes is available; and transmitting the task to the node determined as available for execution at the available node.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD The present disclosure relates to a system and method for managing a task in a distributed computing environment. In particular, but not exclusively, the system and method are used to assign tasks to nodes in the distributed computing environment running on renewable energy sources. BACKGROUND Existing solutions for remote code execution are usually centralised or decentralised. Centralised solutions are vulnerable since they have one single point of failure. If the central component goes down, the entire network is rendered inoperative. A centralised solution may be a data centre or fixed network of computing devices. De-centralised solutions introduce a less vulnerable approach by slicing the network in sections, limiting the damage to a smaller portion of the network in case of failure. De-centralised solution may be in the form of a cloud computing network with devices located at multiple locations. A problem associated with distributed computing is how to assign tasks to different nodes within the system. Specifically, how to ensure tasks are being performed by an appropriate node. As such there is a desire to be able to efficiently and effectively assign and manage tasks in a distributed computing environment. SUMMARY OF THE INVENTION Aspects and embodiments of the invention provide a system and method for managing a task in a distributed computing environment as claimed in the appended claims. An aspect of the invention introduces a way to create a distributed solution for remote code execution, where nodes can join or leave the network at will, without interfering with the regular operation of the system. The network is capable of both shrinking and expanding as nodes are constantly joining and leaving. According to an aspect of the invention there is provided a method for managing a task in a distributed computing environment, wherein the distributed computing environment comprises: one or more nodes configured to execute a task, a plurality of channels, and a dispatcher configured to assign the task to a channel from the plurality of channels. The method comprising: determining at least one property of each node of the one or more nodes; assigning each node of the one or more nodes to a channel of the plurality of channels dependent on the at least one property of each node of the one or more nodes; monitoring each node of the one or more nodes as they join and / or leave the distributed computing environment; receiving a task at the dispatcher from an external device; assigning, at the dispatcher, the task to a channel selected from the plurality of channels, wherein the channel is selected dependent on at least one property of the task; monitoring an availability of a subset of the one or more nodes, wherein the subset of the one or more nodes are assigned to the selected channel; determining that a node from the subset of the one or more nodes is available; and transmitting the task to the node determined as available for execution at the available node. Such a process allows for the more effective scheduling of tasks within a distributed computing environment. The method provides for the scheduling of computational processing utilising a disparate network of computers, optionally all powered via renewable sources. The process enables tasks to be assigned to an appropriate node with the correct capabilities thus resulting in a more efficient system. In such a distributed computing environment the availability of individual nodes may change over time and is not necessarily consistent. This represents a further challenge in the distribution and scheduling of tasks as the status of the nodes may vary over time. By monitoring the availability of nodes and sending tasks to available nodes, nodes are able to join and leave the system at will, without imposing significant overall impact on the process. Further, the security of the system is increased by communicating via the dispatcher. Optionally wherein the one or more nodes join and / or leave the distributed computing environment dependent on at least one criteria. Optionally wherein the at least one criteria comprises whether the one or more nodes are currently being used. Optionally wherein the at least one criteria comprises the time of day. Optionally wherein the at least one criteria comprises whether the one or more nodes are currently running on renewable energy. Such a process allows customers to process computational data with a focus on the processing occurring using renewable energy. Optionally the method further comprising assigning each node of the one or more nodes to a channel from the plurality of channels upon registering each node of the one or more nodes with the dispatcher. Optionally wherein the at least one property of the task comprises the length of time to complete the task. Optionally wherein the at least one property of the task comprises how computationally costly the task is. Optionally wherein the at least one property of the node comprises latency. Optionally wherein the at least one property of the node comprises memory. Optionally the method further comprising determining a geolocation of the task. Optionally the method further comprising assigning the task to the channel selected from the plurality of channels depending on the geolocation of the task such that a carbon footprint per byte is reduced. Optionally wherein the plurality of channels comprise a first channel for performing tasks with the least computational cost. Optionally wherein the plurality of channels comprise a second channel for performing tasks with the most computational cost. There is also provided a dispatcher configured to perform the method of: determining at least one property of each node of one or more nodes; assigning each node of the one or more nodes to a channel of a plurality of channels dependent on the at least one property of each node of the one or more nodes; monitoring each node of the one or more nodes as they join and / or leave the distributed computing environment; receiving a task from an external device; assigning the task to a channel selected from the plurality of channels, wherein the channel is selected dependent on at least one property of the task; monitoring an availability of a subset of the one or more nodes, wherein the subset of the one or more nodes are assigned to the selected channel; determining that a node from the subset of the one or more nodes is available; and transmitting the task to the node determined as available for execution at the available node. There is also provided a system for managing a task in a distributed computing environment, the system comprising a node configured to execute a task and a dispatcher as described above. There is also provided a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out any of the above recited method steps. Within the scope of this application it is expressly intended that the various aspects, embodiments, examples and alternatives set out in the preceding paragraphs, in the claims and / or in the following description and drawings, and in particular the individual features thereof, may be taken independently or in any combination. That is, all embodiments and / or features of any embodiment can be combined in any way and / or combination, unless such features are incompatible. The applicant reserves the right to change any originally filed claim or file any new claim accordingly, including the right to amend any originally filed claim to depend from and / or incorporate any feature of any other claim although not originally claimed in that manner. BRIEF DESCRIPTION OF THE DRAWINGS One or more embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which: Figure 1 is a schematic representation of the apparatus according to an aspect of the invention; Figure 2 is a process diagram of receiving and outputting a task according to an aspect of the invention; Figure 3 is a flow chart of the process when a node joins the distributed computing environment according to an aspect of the invention; Figure 4 is a flow chart of the process of receiving and outputting a task according to an aspect of the invention; Figure 5 is a flow chart of the process of receiving and outputting a task including monitoring node availability according to an aspect of the invention. DETAILED DESCRIPTION The present invention provides a system and method for assigning a task in a distributed computing environment. In particular, but not exclusively, the present invention may be used in systems having distributed computers of varying capabilities and computational power. The invention is able to analyse these capabilities and assign tasks appropriately. Advantageously, this results in a more efficient system as tasks are handled by the nodes or devices which have the appropriate capabilities and computational power, and which are available at the time of execution of the task. By ensuring the nodes which have the appropriate capabilities and computational power are selected, nodes which for example, are underpowered in terms of processing power, are not selected for computationally intensive tasks as this would be inefficient. Conversely, tasks which have relatively lower demands in computational power are preferentially assigned to devices with, say, lower computational power thus ensuring that devices with higher computational power remain available for computationally intensive tasks. During periods of low demand, tasks with lower demands may be assigned to nodes with a high computational power. Figure 1 is a schematic representation of the system for a distributed computing environment in accordance with an embodiment of the invention. In Figure 1 there is shown a system 100 for managing task assignments in a distributed computing environment. The system 100 comprises a user device 102 which is arranged to deliver a task 104 to a dispatcher 106. The dispatcher 106 is arranged to assign and deliver tasks via a first channel 108, a second channel 110 and / or a third channel 112 to nodes 114, 116, and 118. For the purpose of understanding, only three channels are shown though in further embodiments the number of channels may be greater than, or less than, three. Similarly for the purpose of understanding only a single node is shown to be assigned to each channel. In practice multiple nodes, for example thousands, may be assigned to each channel. Tasks are distributed and assigned to channels 108-112 for execution by their respective nodes 114-118 as described in further detail below. User device 102 may also be considered a node in the distributed computing environment capable of executing a task. Similarly, one or more of nodes 114-118 may be considered a user device capable of delivering a task to the dispatcher. As such, user device 102 and nodes 114-118 may be one and the same within the distributed computing environment. The user device 102 and the node 114-118 can be any suitable device such as tablet computers, laptop computers, desktop computers, smart phones, raspberry pi etc. The user device 102 and the nodes 114-118 need not be the same type of device. Indeed, the distributed computing environment can be formed from a variety of different types of devices. Whilst system 100 shows one node 114 associated with the first channel 108, one node 116 associated with the second channel 110, and one node 118 associated with the third channel 112, it is to be understood that a plurality of nodes may be associated with each channel. A single node per channel is depicted for ease of illustration. In one example, the system may have any number of channels suitable for the distributed computing environment. Figure 1 shows node 114 associated with the first channel 108, node 116 associated with the second channel 110, and node 118 associated with the third channel 112. Each node is assigned to a channel. This process is described in more detail below with reference to Figure 3. This process may be performed when the node first joins the distributed computing environment. Preferably, the assignment of a node to a channel is dependent on one or more properties of the node. For example, the system may assess the computational power of the node to determine strengths and weaknesses of the device. By way of example, the system may determine the latency and memory of the node. Advantageously, this assignment allows nodes to be grouped together, preferably in the same channel 108-112, dependent on computational power. This enables the system to distribute one or more tasks in the most efficient manner. Nodes may also be grouped together in the same channel depending on a geolocation of the node and / or a reliability score associated with the node, as discussed in more detail in relation to Figure 3. Figure 1 shows the first channel 108, the second channel 110, and the third channel 112. Any number of channels may be implemented in system 100. The system may include one or more channels for long-running, central processing unit (CPU) intensive, and / or memory eager tasks. The system may include one or more channels for short-term, high disk and / or low network latency tasks. As an example, the first channel 108 may be associated with one or more nodes 114 which have the lowest computational power of the nodes within the distributed computing environment. For example, node 114 may be a device such as a raspberry pi. The third channel 112 may be associated with one or more nodes 118 which have the most computational power of the nodes within the distributed computing environment. For example, node 118 may be a device such as a high performing desktop computer. System 100 may include the second channel 110 which may be associated with one or more nodes 116 which have a computational power higher than the nodes 114 with the lowest computational power and / or a computational power lower than the nodes 118 with the highest computational power. Therefore, nodes with the same and / or similar properties are grouped into the same channel. The dispatcher 106 is configured to assess task properties and distribute tasks to nodes based on the task properties. The dispatcher may be a piece of software that hosts a database and a queue which is split into channels. The queue may be split into channels depending on node capability as described directly above. The database may operate to keep metadata on the network and replicate amongst several nodes which may each run a copy of the dispatcher. In some examples, the dispatcher may rank each device as it joins the distributed computing environment. The ranking may be based on performance capabilities of the device such as computational power, latency, memory etc. The dispatcher operates as a de-centralised computing device which receives and collects data from external devices and nodes within the distributed computing environment. Some of the data collected by the dispatcher can be shared with the external devices and / or the nodes. For example, information received from an external device about a task is shared with one or more nodes via the dispatcher. Similarly, information collected from the nodes such as the task output is shared with an external device via the dispatcher. In this way the dispatcher operates as a de-centralised computing device capable of communicating with the other devices within the distributed computing system. Optionally, the distributed computing system may comprise more than one dispatcher each having the capabilities discussed above in relation to a single dispatcher. In one example, there may be a different dispatcher for different physical locations. For example, there may a first dispatcher operable within the UK and a second dispatcher operable within Australia. In further examples there may be ,multiple dispatchers in the same country or continent. Figure 2 is a worker diagram of the process of managing tasks in a distributed computing environment. At step S202, a node (also known as a worker) may transmit data to the dispatcher to register the node with the distributed computing environment. As an example, the node may transmit data which identifies the node. Preferably, the node transmits data relating to the computational power and / or performance capabilities of the node itself. This data may include the latency and / or memory properties of the node. Based on this data, the dispatcher can rank the nodes and determine which tasks are sent to which nodes. In an example, the dispatcher may use the transmitted data from the node (e.g., the memory and latency properties) to assign the node to a channel. For example, a subset of the nodes may be grouped as having a low computational power and assigned to a first channel, and a different subset may be grouped as having a high computational power and assigned to a second channel. This enables the dispatcher to assign tasks to a channel, and thus a node, having the most appropriate capabilities for that task. Nodes may also be grouped together in the same channel depending on a geolocation of the node and a reliability score associated with the node, as discussed in more detail in relation to Figure 3. At step S204, the dispatcher receives a task. The task may be received from an external device. As an example, the external device may be one of the nodes within the distributed computing environment. In other examples the external device is a device which does not form a node within the distributed computing environment. The tasks may be transmitted to the dispatcher as Hypertext Transfer Protocol Secure (HTTPS) requests which can be a click of a button, a timer or a schedule firing the HTTPS request to the dispatcher. The user requesting execution of a task does not communicate directly with the node that performs the execution. The user communicates with the dispatcher. This increases the security of the overall system. In further examples any other suitable form of secure protocol may be used. The dispatcher may analyse the task to determine one or more properties associated with the task. For example, the properties may be the length of time taken to complete the task and / or the computational cost of performing the task. The dispatcher may determine a geolocation of the task. Depending on the geolocation, the dispatcher can assign the task to a channel encompassing the geolocation of the task such that a carbon footprint per byte is reduced. For example, a task registered in Australia would preferably be performed by a node assigned to a channel encompassing Australia whereas a task registered in the UK would not be assigned to this channel. The dispatcher may then assign the task to a channel and dispatch the task to a node within that channel at step S206. In some examples, the dispatcher may monitor the availability of nodes within a channel to determine which nodes are available to execute the task as described in relation to Figure 5. At step S208, the node may use a known Function-as-a-Service (FaaS) runner to execute the task. At step S210 the output of the task is returned by the node and transmitted to the dispatcher at step S212. The nodes do not communicate directly with the user device which transmitted the task. The communication occurs indirectly via the dispatcher. Figure 3 is a flow chart of the process 300 followed when a node joins the distributed computing environment. Preferably, this process is completed when the node first joins the distributed computing environment although it is to be understood that such a process may additionally be performed at any time. For example, the process may be run periodically to capture any improvements or deteriorations of any node within the distributed computing environment. For example, after joining the distributed computing environment, the nodes may be assessed based on properties such as the length of time the node is available each day, the number of days the node has been running within the environment, and the number of incidents the node has had while operating in the distributed computing environment. Such properties allow the dispatcher to calculate a reliability score for each node. This reliability score may be updated with time. In a node has a score under a certain reliability threshold, a node may not handle a task by itself and will share the processing responsibility with a fallback node. In this way, if the assigned node fails to deliver a result, the dispatcher will assign the fallback node to provide the result. The fallback node may be the systems internal hardware or a reliable node member of the distributed computing environment. If a high reliability score node fails to deliver an expected result, the systems internal hardware will process the task after a grace period. This guarantees delivery of a result despite a minor delay. The length of the grace period depends on the task being run. For example, a long running task will have a greater delay in comparison to a critical task. For every incident a node has its reliability score will be adjusted accordingly. If a slow node comes back with a response before the fallback node finishes executing the task, the latter computation will be discarded, and no penalty will be recorded to the assigned node. That protocol ensures that a node is free to join or leave the network at will, as the network will adjust itself to deliver long running tasks to well established reliable nodes, critical tasks to powerful nodes and highly demanding memory processing tasks to big RAM nodes. At step S302, a node registers with the dispatcher. Upon registering with the dispatcher, the node provides information relating to its capabilities. For example, the node may transmit data relating to its computational power. Such information may include the type of microprocessor(s) and the chipset(s) in the device, or other suitable information relating to the performance and capabilities of the device. Preferably, such data may comprise the memory and / or latency of the node. At step S304, the dispatcher analyses the information provided by the node and determines the properties of the node. Preferably, the dispatcher ranks the nodes based on computational power and / or performance capabilities. At step S306, the dispatcher may use the determined properties to assign the node to a channel. This enables the dispatcher to group nodes together based on one or more properties of the node. For example, nodes having identical microprocessors, or microprocessors with similar performance metrics, are grouped in the same channel. In further examples one or more different properties such as device type, memory, latency etc., may be used. This allows tasks to be efficiently assigned to appropriate channels and thus nodes. Further, the nodes may be further grouped dependent on geolocation to form one or more channels encompassing a specific region. For example, a node located within the UK may be assigned to a channel wherein each node in that channel is also located in the UK. Similarly, a node located in Australia may be assigned to a channel which groups nodes located in Australia. These channels are optionally further divided into performance capabilities and / or reliability scores as described above. The nodes may also be grouped together in a channel dependent on properties relating to the processing reliability of the node within the environment as described in more detail above. Figure 4 is a flow chart of the process of managing a task in a distributed computing environment, such as system 100, according to an aspect of the invention. At step S402, a task is registered by an external device. The external device then transmits the task to the dispatcher. Each task may be transmitted as a HTTPS request, or via some other form of secure transfer. At step S404, the dispatcher receives the task from the external device. At step S406, the dispatcher may assess the properties of the task to determine the resource demands the task may require. For example, the dispatcher may estimate the length of time likely required to complete the task. Alternatively, or additionally, the dispatcher may determine the computational cost of the task and / or the geolocation of the task. At step S408, the dispatcher may assign the received task to a channel. This assignment can be based on the properties of the task and / or the classification of the channel. For example, a resource intensive task may be assigned to a channel which is associated with high performing nodes. A task registered in Australia may be assigned to an appropriate geographic channel. A critical task may be assigned to a channel associated with nodes having a high reliability score regardless of geography. At step S410, the dispatcher may transmit the task to a node associated with the assigned channel. In some examples, the node may return a confirmation message to the dispatcher to confirm that the node will run the task. At step S412, the node performs the task. For example, the node may perform the task via a FaaS as discussed in relation to Figure 2. At step S414, the node transmits the task response to the dispatcher which then may output the task response to the user device. Figure 5 is a flow chart of a process of monitoring node availability in accordance with an embodiment of the invention. Step S502 may be the same as step S402 described in relation to Figure 4. Step S504 may be the same as step S404 described in relation to Figure 4. Step S506 may be the same as step S406 described in relation to Figure 4. Step S508 may be the same as step S408 described in relation to Figure 4. At step S510, the dispatcher may monitor the nodes associated with the assigned channel to determine an available node for executing the task. In an example, the availability of a node may be dependent on one or more criteria. Preferably, the one or more criteria may comprise whether the node is running on renewable energy. The one or more criteria may comprise whether the node is currently being used. For example, the dispatcher may determine whether the node is currently being operated by a user. As another example, the criteria may comprise a particular time of day. For example, the node may be classed as available during a period at night such as between 11pm and 5am. Alternatively, the node may be classed as available between 8am and 3pm, for example. It is to be understood that the node may be classed as available for any suitable time period chosen by the owner / user of the node. The criteria may comprise more than one of whether the node is running on renewable energy, whether the node is currently being used, and / or the time of day. Advantageously, when the dispatcher assesses the length of time taken to complete a task (as described at step S406 of Figure 4), the dispatcher may consider which node to assign the task to dependent on the availability of the node. For example, a task with an estimated run time of 20 minutes would preferably not be assigned to a node which is classed as available for another 10 minutes (for example, the node is classed as available for certain hours of the day and is approaching the end of its available hours). This improves the efficiency of the system and ensures tasks are handled by nodes with the appropriate capacity. At step S512, the dispatcher may determine that a node associated with the assigned channel is not currently available. In this case, the dispatcher returns to step S510 to continue monitoring the nodes for availability. This process may be repeated until a node is determined as available. Once a node is determined as available, the dispatcher may transmit the task to the available node for execution. The task may then be returned as described at steps S412 and S414 of Figure 4. The approaches described herein allow for an effective method of managing tasks in a distributed computing environment. Such an approach allows customers to 5 process computational data, preferably, with a focus on the processing occurring using renewable energy. Tasks can be scheduled more effectively and efficiently by assigning tasks to distributed computers within the environment with the most suitable resources and availability classification to carry out each task. Further, by operating via the dispatcher the user devices and the nodes do not communicate 10 with one another thus improving the security of the system.

Claims

1. A method for managing a task in a distributed computing environment, wherein the distributed computing environment comprises:one or more nodes configured to execute a task,a plurality of channels, anda dispatcher configured to assign the task to a channel from the plurality of channels;the method comprising:determining at least one property of each node of the one or more nodes;assigning each node of the one or more nodes to a channel of the plurality of channels dependent on the at least one property of each node of the one or more nodes;monitoring each node of the one or more nodes as they join and / or leave the distributed computing environment;receiving a task at the dispatcher from an external device;assigning, at the dispatcher, the task to a channel selected from the plurality of channels, wherein the channel is selected dependent on at least one property of the task;monitoring an availability of a subset of the one or more nodes, wherein the subset of the one or more nodes are assigned to the channel selected from the plurality of channels;determining that a node from the subset of the one or more nodes is available; andtransmitting the task to the node determined as available for execution at the node determined as available.

2. The method of claim 1, wherein the one or more nodes join and / or leave the distributed computing environment dependent on at least one criteria.

3. The method of claim 2, wherein the at least one criteria comprises whether the one or more nodes are currently being used.

4. The method of claim 2, wherein the at least one criteria comprises a time of day.

5. The method of claim 2, wherein the at least one criteria comprises whether the one or more nodes are currently running on renewable energy.

6. The method of any preceding claim, further comprising assigning each node of the one or more nodes to a channel from the plurality of channels upon registering each node of the one or more nodes with the dispatcher.

7. The method of any preceding claim, wherein the at least one property of the task comprises a length of time to complete the task.

8. The method of any preceding claim, wherein the at least one property of the task comprises a computational cost of the task.

9. The method of any preceding claim, wherein the at least one property of each node of the one or more nodes comprises latency.10.The method of any preceding claim, wherein the at least one property of each node of the one or more nodes comprises memory.

11. The method of any preceding claim further comprising determining a geolocation of the task.

12. The method of claim 11 further comprising assigning the task to the channel selected from the plurality of channels depending on the geolocation of the task such that a carbon footprint per byte is reduced.13.The method of any preceding claim, wherein the plurality of channels comprise a first channel for performing tasks with a low computational cost.14.The method of any preceding claim, wherein the plurality of channels comprise a second channel for performing tasks with a high computational cost.

15. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the steps of the method of any of claims 1-14.

16. A dispatcher configured to perform a method comprising:determining at least one property of each node of one or more nodes;assigning each node of the one or more nodes to a channel of a plurality of channels dependent on the at least one property of each node of the one or more nodes;monitoring each node of the one or more nodes as they join and / or leave a distributed computing environment;receiving a task from an external device;assigning the task to a channel selected from the plurality of channels, wherein the channel is selected dependent on at least one property of the task;monitoring an availability of a subset of the one or more nodes, wherein the subset of the one or more nodes are assigned to the channel selected from the plurality of channels;determining that a node from the subset of the one or more nodes is available; andtransmitting the task to the node determined as available for execution at the node determined as available.

17. The dispatcher of claim 16, wherein the dispatcher is further configured to perform the method of any of claims 2-14.

18. A system for managing a task in a distributed computing environment, the system comprising:a node configured to execute a task; andthe dispatcher of claim 16.