Distributed task scheduling system oriented to edge environment
By designing a distributed task scheduling system for edge environments, the task interruption caused by resource management limitations and node exceptions in the existing technology is solved, more accurate resource allocation and task scheduling is achieved, and the stability of the system and task execution efficiency are improved.
Patent Information
- Application Number
- CN202510234602.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
The existing task scheduling tools have limitations in resource management, and fail to fully consider the actual needs of tasks and network conditions, resulting in uneven resource allocation and load imbalance. The lack of flexible task resource management strategies when node abnormalities may lead to interruption of task operation.
A distributed task scheduling system for edge environments is designed, including communication module, data collection module, task management module, task scheduling module and data cache module. The system monitors node status in real time, adjusts resource allocation dynamically, uses genetic algorithms to schedule tasks, and backs up task cache data in multiple nodes to ensure that tasks are efficiently executed on the most suitable nodes.
It realizes more accurate resource allocation and task scheduling, improves resource utilization and task execution efficiency, enhances system stability and fault tolerance, and ensures task continuity and data security.
Smart Images

Figure CN120066733A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of edge computing, and specifically to a distributed task scheduling system for edge environments. Background Art
[0002] Cloud computing is a network-based computing model that provides computing resources and services on demand over the Internet. Its core lies in the utilization of large data centers and virtualization technologies to achieve flexible resource allocation and efficient management. In contrast, edge computing emphasizes data processing and storage at the source of data generation or close to the data source, aiming to reduce the time for data to be transmitted to the data center, thereby reducing latency and improving response speed, especially suitable for scenarios with high latency requirements. In an edge cluster, the master-slave architecture, as a common design pattern, optimizes task offloading and resource allocation through centralized management and distributed execution, while using containerization technology to further reduce system resource overhead.
[0003] However, current task scheduling tools have limitations in resource management. Although these tools can effectively manage CPU and memory resources, they often overlook key factors such as node heterogeneity, network bandwidth, and disk capacity. Existing scheduling algorithms, such as the best fit algorithm and the minimum load method, etc., although they have improved resource utilization to a certain extent, often fail to fully consider the actual requirements of tasks and network conditions, resulting in frequent problems such as uneven resource allocation and load imbalance. In addition, when nodes in the cluster encounter abnormalities, such as insufficient resources or communication interruptions, existing tools lack flexible task resource management strategies and cannot perform dynamic scheduling according to the actual situation of the cluster, which may lead to task running interruptions and thus affect the performance of the entire system. In view of the deficiencies of the prior art, the present invention provides a distributed task scheduling system for edge environments to solve the above problems. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides a distributed task scheduling system for edge environments, which solves
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A distributed task scheduling system for edge environments, including a communication module, a data collection module, a task management module, a task scheduling module, and a data caching module;
[0006] The communication module establishes a connection between the cloud node and the edge node, where the cloud node acts as the client and the edge node acts as the server, and the cloud node regularly sends liveness detection information to the edge node;
[0007] The data collection module regularly sends data collection instructions to the nodes in the "alive" state through the communication module to obtain the resource usage of the nodes and the resource occupancy of the task containers;
[0008] The task management module is responsible for managing and monitoring the life cycle of tasks, pulling task images from the image repository and deploying tasks;
[0009] The task scheduling module adjusts the execution order of the task queue and formulates the task scheduling plan;
[0010] The data cache module is responsible for storing key task data and cache data generated during task execution, and stores the data in multiple nodes.
[0011] Preferably, the communication module uses AES symmetric encryption technology to encrypt information sent by the sending node, and transmits the encrypted message to the target node.
[0012] Preferably, the communication module assigns a unique identifier to each node, the identifier is used to manage the nodes in the network, and regularly sends survival detection information to the edge nodes to monitor the operating status of each node in real time;
[0013] The communication module monitors the running status of the node in real time. When a node stops running, it notifies the task management module to check the tasks being executed on the node and migrate the tasks to other nodes for continued execution according to task requirements.
[0014] Preferably, the data collection module can collect the load status of the node in real time. When the node resources are insufficient, it notifies the task management module to obtain the tasks running on the node, and migrates the lower priority tasks to the nodes with more sufficient resources according to the task priority to ensure the smooth execution of key tasks.
[0015] Preferably, the task management module is responsible for managing the task life cycle, monitoring the deployment and running status of the task, and when the task deployment fails or the running fails, the task scheduling strategy is re-formulated through the task scheduling module to reallocate and deploy the task.
[0016] Preferably, the task management module adapts to dynamic changes in task requirements and dynamically adjusts resource allocation of tasks when resources required for task execution are insufficient.
[0017] Preferably, the task scheduling module adjusts the execution order of tasks according to factors such as task priority and task execution time to ensure that important tasks are executed first.
[0018] Preferably, the task scheduling method adopted by the task scheduling module is divided into two stages, including:
[0019] In the pre-selection stage, nodes that do not meet the task requirements are filtered out according to the running conditions of different tasks;
[0020] In the optimization stage, the optimal node is selected for task deployment according to the resource requirements of the task and the resource occupancy of different nodes.
[0021] Preferably, the data cache module backs up task cache data among multiple nodes.
[0022] Preferably, the data cache module adopts a cache update strategy and uses a cache replacement algorithm to manage task cache data.
[0023] The present invention discloses a distributed task scheduling system for edge environments, and its beneficial effects are as follows:
[0024] 1. The distributed task scheduling system for edge environments realizes more accurate resource allocation and task scheduling by comprehensively considering key factors such as node heterogeneity, network bandwidth, and disk capacity. This not only improves the utilization rate of resources but also ensures that tasks can be efficiently executed on the most suitable nodes, thereby enhancing the performance and response speed of the entire system.
[0025] 2. The distributed task scheduling system for edge environments adopts advanced communication technologies and encryption means to ensure the security and reliability of information transmission between nodes. At the same time, by real-time monitoring the running status of nodes and timely performing task migration, it effectively avoids task interruption and system performance degradation caused by node anomalies, enhancing the stability and fault tolerance of the system.
[0026] 3. The distributed task scheduling system for edge environments realizes the consistency and redundant backup of task cache data through the effective management of the data cache module, preventing the loss of key task data. At the same time, the cache strategy is dynamically adjusted according to the remaining capacity of the cache space, optimizing the utilization efficiency of the cache space and further enhancing the reliability and data security of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0028] Figure 1 is a schematic structural diagram of the distributed task scheduling system for edge environments provided by the present invention;
[0029] Figure 2 is a schematic working diagram of communication between nodes in the distributed task scheduling system for edge environments provided by the present invention;
[0030] Figure 3 It is a schematic flowchart of the symmetric encryption operation in the distributed task scheduling system for the edge environment provided by the present invention;
[0031] Figure 4 It is a schematic flowchart of the working process of the communication module in the distributed task scheduling system for the edge environment provided by the present invention;
[0032] Figure 5 It is a schematic flowchart of the working process of the data collection module in the distributed task scheduling system for the edge environment provided by the present invention;
[0033] Figure 6 It is a schematic flowchart of the working process of the task management module in the distributed task scheduling system for the edge environment provided by the present invention;
[0034] Figure 7 It is a schematic flowchart of the genetic algorithm in the distributed task scheduling system for the edge environment provided by the present invention;
[0035] Figure 8 It is a schematic flowchart of the working process of the data cache module in the distributed task scheduling system for the edge environment provided by the present invention.
[0036] In the figure: 101, communication module; 102, data collection module; 103, task management module; 104, task scheduling module; 105, data cache module. Specific embodiments
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] By providing a distributed task scheduling system for the edge environment in the embodiments of the present application, the problems that the current task scheduling tools have limitations in resource management, fail to fully consider the actual requirements of tasks and network conditions, resulting in frequent problems such as uneven resource allocation and load imbalance, and when nodes in the cluster are abnormal, such as resource shortage or communication interruption, the existing tools lack flexible task resource management strategies and cannot perform dynamic scheduling according to the actual situation of the cluster, which may cause task operation interruption and thus affect the performance of the entire system are solved.
[0039] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0040] Example 1:
[0041] An embodiment of the present invention discloses a distributed task scheduling system for an edge environment. As shown in the attached Figure 1-8 figure, it includes a communication module 101, a data collection module 102, a task management module 103, a task scheduling module 104, and a data cache module 105;
[0042] The communication module 101 establishes a connection between the cloud node and the edge node through the GRPC communication protocol, where the cloud node acts as the client and the edge node acts as the server. The cloud node periodically sends liveness detection information to the edge node to ensure the normal operation of the edge node;
[0043] The data collection module 102 periodically sends data collection instructions to the nodes in the "alive" state through the communication module 101 to obtain the resource usage of the nodes and the resource occupancy of the task containers;
[0044] The task management module 103 is responsible for managing tasks and monitoring the life cycle of tasks. According to the allocation strategy formulated by the task scheduling module 104, it pulls task images from the image repository and deploys tasks; when the resources required for task execution are insufficient, the task management module 103 will dynamically adjust the resource allocation to ensure the smooth execution of tasks;
[0045] The task scheduling module 104 can adjust the execution order of the task queue according to factors such as the priority and execution time of the tasks, and uses a genetic algorithm to formulate a task scheduling plan to ensure the balanced allocation of cluster load resources;
[0046] The data cache module 105 is responsible for storing key task data and cache data generated during task execution. Adopting a backup storage strategy, it stores the data on multiple nodes respectively to prevent the loss of key task data.
[0047] As Figure 2 shown, the communication module 101 uses the GRPC communication protocol to establish a connection between nodes, taking the cloud node as the client and the edge node as the server; GRPC is an efficient and open-source remote procedure call protocol that supports cross-language communication and allows clients and servers to exchange data through predefined interfaces. It adopts a cross-platform communication mechanism, supports the interaction and data transmission between services in a large-scale distributed environment, and is widely used in scenarios such as cloud computing, real-time applications, and big data processing.
[0048] In an embodiment of the present invention, the communication module 101 uses the symmetric encryption algorithm AES algorithm to encrypt the message sent by the cloud node with a preset symmetric key, and sends the encrypted message to the edge node; the edge node receives the encrypted information, decrypts the message with the symmetric key, and performs corresponding operations according to the message type. The schematic flow diagram of the symmetric encryption operation is as Figure 3 shown;
[0049] The AES algorithm is a widely used symmetric encryption algorithm for ensuring data security. It uses a fixed-length key and realizes data encryption and decryption through multiple rounds of encryption processes. In a distributed task scheduling system for edge environments, AES can ensure secure message transmission between nodes while maintaining high processing speed, making it suitable for use in distributed environments.
[0050] As Figure 4 shown, the communication module 101 assigns a unique identifier and IP address to each node, and sends a liveness detection signal to the nodes in the cluster every 10 seconds to ensure the normal operation of the nodes. When an abnormality is detected in a certain node, the communication module 101 will notify the task management module 103 to check the task information running on that node, such as the resources required by the task, the task image, etc., and formulate an allocation strategy through the task scheduling module 104 based on the task information, and re-pull the image to complete the task deployment. To ensure the integrity of global monitoring, the communication module 101 traverses all edge nodes during each detection.
[0051] As Figure 5 shown, the data collection module 102 is responsible for collecting load information such as CPU, memory, network bandwidth, and disk capacity of the edge nodes, as well as the resource occupancy of tasks. When the CPU, memory, or disk usage rate of a certain node exceeds the preset threshold, an alarm message will be sent, and the task management module 103 will be notified to obtain the priority and resource occupancy of the tasks running on that node. Subsequently, some tasks with lower priorities will be scheduled to nodes with sufficient resources to ensure the smooth execution of important tasks.
[0052] In an embodiment of the present invention, the data collection module 102 has multi-dimensional monitoring capabilities. In addition to collecting the CPU and memory usage of edge nodes, it further collects load information such as network bandwidth and disk capacity, and obtains the resource occupancy data of tasks in real time. Through comprehensive monitoring and analysis of the node status, the system can accurately evaluate the running status of nodes based on multiple key indicators, provide data support for resource scheduling and task allocation, and ensure the optimization of resource allocation and the efficient execution of task scheduling.
[0053] As Figure 6As shown in the figure, the task management module 103 uniformly receives requests sent by the front-end program, parses the request information, sends the resource information required for the task to the task scheduling module 104, and formulates a reasonable allocation strategy using a scheduling algorithm; subsequently, it uses the communication module 101 to notify the task deployment node to pull the task image from the image repository and execute the task.
[0054] The task management module 103 is also responsible for managing the life cycle of the task and monitoring the deployment and running status of the task. When the task deployment or running fails, the task management module 103 will re-formulate the scheduling strategy through the task scheduling module 104 and redeploy and allocate the task. At the same time, the task management module 103 can adjust the resource allocation according to the dynamic changes in task requirements. When the resources required for the task are insufficient, the task management module 103 will re-optimize the resource allocation plan according to the resource requirements of the task and the resource status of the node.
[0055] In the embodiment of the present invention, the task management module 103 is combined with the communication module 101, the data collection module 102, the task scheduling module 104, and the data cache module 105, which can effectively cope with various complex scenarios, realize the efficient allocation of resources and the dynamic scheduling of tasks. The system can give priority to ensuring the normal operation of critical tasks when resources are insufficient, and at the same time migrate low-priority tasks to nodes with sufficient resources; when a certain node stops running, the task can be timely migrated to other available nodes to ensure that the task is not interrupted; in the case of task running failure, the scheduling strategy can be re-formulated and the task can be restarted; in addition, the system can adjust the resource allocation strategy in real time according to the task running status, further improving the flexibility of task scheduling and the stability of system operation.
[0056] In the embodiment of the present invention, Docker containerization technology is adopted to manage tasks, and the application and its dependencies are packaged into independent containers to ensure the consistency and efficiency of tasks in different edge environments. Through containerization, the system can achieve rapid startup, resource isolation, and efficient scheduling, enabling tasks to migrate seamlessly between various hardware and operating system platforms. Docker containers share the host operating system kernel but are completely isolated from each other, which provides strong support for the independent operation, automated deployment, and elastic scaling of tasks, thus improving the flexibility and efficiency of task management.
[0057] In the embodiment of the present invention, the task scheduling module 104 can dynamically adjust the execution order of tasks in the List collection according to factors such as the priority and execution time of the tasks, so as to ensure that important tasks can be deployed and executed first.
[0058] The task scheduling module 104 adopts a two-stage task scheduling strategy, which mainly includes:
[0059] In the pre-selection stage, the task scheduling module 104 filters out some nodes according to the execution requirements of the tasks and the rules preset by the user, thereby narrowing the range of optional nodes and reducing the computational complexity in the optimization stage.
[0060] In the optimization stage, the task scheduling module 104 uses a genetic algorithm to formulate a task allocation strategy. First, a population is initialized, where each individual represents a task allocation scheme. The load balancing degree of the cluster is used as the fitness function to evaluate the quality of the individuals. The load balancing degree is the standard deviation of the loads among the nodes. Then, the selection operation is used to select parental individuals from the current population according to the fitness values, and the uniform crossover operator is used for the crossover operation to generate new offspring individuals. To maintain the diversity of the population, the mutation operation is further performed to randomly modify some genes in the individuals. By repeating the selection, crossover, and mutation operations, the task allocation strategy is gradually optimized until the maximum number of iterations is reached. Finally, the optimal task allocation scheme is output and passed to the task management module 103 for task deployment operations. The execution process of the algorithm is as Figure 7 shown.
[0061] In the embodiment of the present invention, a genetic algorithm is used to formulate a task allocation strategy. Compared with the traditional best-fit algorithm and minimum load method, the genetic algorithm has obvious advantages in resource scheduling. Traditional algorithms are usually only applicable to a single resource, such as CPU or memory, and cannot effectively handle the comprehensive optimization problems of multiple resource types, such as CPU, memory, network bandwidth, disk, etc. The genetic algorithm represents multi-dimensional resource allocation through the population and explores the global optimal solution by using crossover and mutation operations, avoiding the dilemma of local optimality. This algorithm can perform global optimization of task scheduling while considering task priorities and resource constraints, ensuring the reasonable utilization of resources and load balancing.
[0062] In the embodiment of the present invention, the data caching module 105 uses the MariaDB database to store node and task load data, and uses the NFS protocol to ensure the consistency of task cache data among nodes and prevent task data loss.
[0063] Among them, the NFS protocol is a protocol that allows different hosts to share file systems through the network. In this system, the NFS protocol is used to ensure data consistency among multiple nodes. Through the NFS protocol, each node can access and update the shared task cache data, thereby realizing the synchronization and coordination of task data. This protocol guarantees data consistency among different nodes, prevents task loss or execution errors caused by data synchronization problems, and improves the reliability and stability of the system in a distributed environment. Among them, the cloud node acts as the server, and the edge node acts as the client.
[0064] such as Figure 8As shown, the data cache module 105 can effectively manage the task cache space. When the remaining capacity of the cache space is lower than the preset threshold, a warning will be triggered to update the cache space. If the currently used algorithm is the LRU algorithm, the data cache module 105 will delete the least frequently accessed task cache data. If the FIFO algorithm is used, it will delete the task data that entered the cache folder first.
[0065] The device embodiments described above are merely illustrative. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0066] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0067] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0068] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A distributed task scheduling system for edge environments, characterized by: include: The communication module establishes a connection between the cloud node and the edge node, where the cloud node acts as the client and the edge node acts as the server. The cloud node periodically sends survival detection information to the edge node. The data collection module periodically sends data collection instructions to nodes in the "survival" state through the communication module to obtain the resource usage of the nodes and the resource occupancy of the task container; The task management module is responsible for managing and monitoring the life cycle of tasks, pulling task images from the image repository and deploying tasks; Task scheduling module, adjusts the execution order of task queues and formulates task scheduling plans; The data cache module is responsible for storing key task data and cache data generated during task execution, and stores the data in multiple nodes.
2. The distributed task scheduling system for edge environments according to claim 1, characterized in that: The communication module uses AES symmetric encryption technology to encrypt the information sent by the sender node and transmits the encrypted message to the target node.
3. The distributed task scheduling system for edge environments according to claim 1, characterized in that: The communication module assigns a unique identifier to each node, which is used to manage nodes in the network, and regularly sends survival detection information to edge nodes to monitor the operating status of each node in real time; The communication module monitors the running status of the node in real time. When a node stops running, it notifies the task management module to check the tasks being executed on the node and migrate the tasks to other nodes for continued execution according to task requirements.
4. The distributed task scheduling system for edge environments according to claim 1, characterized in that: The data collection module can collect the load status of the node in real time. When the node resources are insufficient, it notifies the task management module to obtain the tasks running on the node. According to the task priority, the tasks with lower priority are migrated to the nodes with more sufficient resources to ensure the smooth execution of key tasks.
5. The distributed task scheduling system for edge environments according to claim 1, characterized in that: The task management module is responsible for managing the task life cycle, monitoring the deployment and running status of tasks, and when the task deployment fails or the running fails, the task scheduling module re-formulates the task scheduling strategy to reallocate and deploy tasks.
6. The distributed task scheduling system for edge environments according to claim 4, characterized in that: The task management module adapts to the dynamic changes in task requirements and dynamically adjusts the resource allocation of the task when the resources required for task execution are insufficient.
7. The distributed task scheduling system for edge environments according to claim 1, characterized in that: The task scheduling module adjusts the execution order of tasks according to factors such as task priority and task execution time to ensure that important tasks are executed first.
8. The distributed task scheduling system for edge environments according to claim 7, characterized in that: The task scheduling method adopted by the task scheduling module is divided into two stages, including: In the pre-selection stage, nodes that do not meet the task requirements are filtered out according to the running conditions of different tasks; In the optimization stage, the optimal node is selected to deploy the task based on the resource requirements of the task and the resource usage of different nodes.
9. The distributed task scheduling system for edge environments according to claim 1, characterized in that: The data cache module backs up task cache data in multiple nodes.
10. The distributed task scheduling system for edge environments according to claim 9, characterized in that: The data cache module adopts a cache update strategy and uses a cache replacement algorithm to manage task cache data.
Citation Information
Cited By
Data processing task scheduling method and device, electronic equipment and medium
CN120578483A