Cluster task scheduling method and system based on prior value

Through the prior value knowledge base analysis of the calculation node score and combining the task resource situation, the problem of unoptimized task allocation in the existing cluster scheduling method is solved, and the optimal task allocation and system performance improvement are achieved.

CN120045320APending Publication Date: 2025-05-27HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510109086.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing cluster scheduling method is difficult to evaluate the matching relationship between the current task and the cluster server, resulting in unoptimized task allocation.

Method used

The scoring of the computing node is analyzed through the prior value knowledge base, and combined with the resource requirements of the task and the resource usage of the computing node, the task is assigned to the computing node with the highest score using the scoring formula.

Benefits of technology

The optimal allocation of tasks is achieved, the accuracy of computing node scores is improved, and the system can provide sufficient computing resources at high loads, which improves performance and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045320A_ABST
    Figure CN120045320A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cluster scheduling, in particular to a cluster task scheduling method and system based on a priori value. The method comprises the following steps that: a cluster server monitors the computing resource use condition of each computing node through a central node; the central node is used for realizing resource monitoring, center registration and task scheduling of the cluster server, the computing node is used for collecting local resource data and processing tasks, and the resource data comprises a CPU core number, a CPU utilization rate, a total memory amount and a memory usage amount; automatically putting tasks created by a user into a task queue, analyzing resource conditions required by to-be-processed tasks in the task queue according to a time sequence on the basis of a priori value knowledge base, and scoring the computing nodes according to the resource conditions required by the to-be-processed tasks and the computing resource use conditions of the computing nodes; and scheduling the to-be-processed task to the computing node with the highest score for processing through the central node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cluster scheduling, and particularly to a cluster task scheduling method and system based on prior values. Background Art

[0002] With the rapid development of cloud computing and big data technologies, various systematic services emerge in an endless stream and become a very important part of people's daily lives. With the gradual expansion of the network scale, the access volume to the network background servers also shows an exponential growth trend. If the huge access volume cannot be processed and responded to in a timely manner, it will seriously affect the user experience and even lead to problems such as server congestion, downtime, and crashes. Therefore, in the face of scenarios with a large number of users and high concurrency, a single server cannot support a large number of user requests. In actual production scenarios, multiple servers need to be combined into a server cluster to jointly provide service support for users.

[0003] Monitoring systems are divided into open-source systems, commercial systems, and self-developed systems in the current environment. They each have a certain usage proportion on the Internet, different target audiences, and corresponding characteristic advantages.

[0004] Traditional cluster scheduling methods use the load status of cluster servers as load evaluation parameters, calculate corresponding weight values based on the change amount of each parameter within two seconds, and finally allocate tasks to the server with the smallest load. This method has the problem of being difficult to evaluate the matching relationship between the current task and the cluster server. Summary of the Invention

[0005] In order to solve the problem that the existing cluster scheduling methods are difficult to evaluate the matching relationship between the current task and the cluster server, the present invention provides a cluster task scheduling method and system based on prior values, which analyzes and calculates the scores of computing nodes through a prior value knowledge base to achieve optimal task allocation, and solves the problem of being difficult to evaluate the matching relationship between the current task and the cluster server.

[0006] In a first aspect, a cluster task scheduling method based on prior values provided by the present invention includes:

[0007] The cluster server monitors the computing resource usage of each computing node through a central node; the central node is used to implement resource monitoring, registration center, and task scheduling of the cluster server, the computing node is used for collecting local resource data and processing tasks, and the resource data includes the number of CPU cores, CPU usage rate, total memory, and memory usage.

[0008] Automatically put the tasks created by the user into the task queue, analyze the resource requirements of the tasks to be processed in the task queue according to the prior value knowledge base in chronological order, and score each computing node according to the resource requirements of the tasks to be processed and the computing resource usage of each computing node;

[0009] Dispatch the task to be processed to the computing node with the highest score for processing through the central node.

[0010] Furthermore, the method further includes: monitoring the computing resource usage occupied by the processing process of the task to be processed to update the prior value knowledge base.

[0011] Furthermore, the computing node deployment includes a windows-exporter collection module, an altermanager alarm module, and a pushgateway push module, and the central node deploys a prometheus service module;

[0012] The windows-exporter collection module is used to collect the resource data of the computing node and adaptively adjust the collection frequency according to the resource usage of the current computing node;

[0013] The altermanager alarm module is used to implement the monitoring and alarm function, and notify the administrator by email after receiving the alarm information from the central node;

[0014] The pushgateway push module is used to push the received data to the prometheus service module;

[0015] The peometheus service module is used to collect, store, and query time series data, and accept the resource data collected by the windows-exporter collection module.

[0016] Furthermore, the prior value knowledge base is constructed as follows:

[0017] Establish a mathematical model for the computing resources and time required to process tasks of different sizes to form the prior knowledge base.

[0018] Furthermore, the central node also deploys a registration center service, which is specifically used for: putting the scrape_configs.targets in the configuration file of the promethus service module into the cache of the registration center service, notifying the registration center service when any computing node is created or destroyed, and after the registration center service receives the notification, pulling the current computing node registration list and comparing it with the cache. If they are inconsistent, update the cache and restart the premethus service module.

[0019] Further, the method of scoring each computing node according to the resource requirements of the to-be-processed task and the computing resource usage of each computing node is as follows:

[0020]

[0021]

[0022] Among them, FinalScore i represents the final score of computing node i, α, β, and γ represent weight parameters, and Score i represents the score of computing node i for the task, Penalty i represents the penalty standard deviation of computing node i, represents the average remaining CPU resources of the server within time T, represents the average CPU usage rate of the task, represents the average remaining memory resources of the server within time T, represents the average memory usage rate of the task, CPU server represents the CPU usage rate of computing node i, Memory server represents the memory usage rate of computing node i, n represents the number of sampling points, and T represents the task execution time.

[0023] Further, the method further includes: dynamically reducing or increasing the number of computing nodes by analyzing the number of to-be-processed tasks in the current task queue and the resource usage of each computing node.

[0024] Further, the dynamically reducing or increasing the number of computing nodes by analyzing the number of to-be-processed tasks in the current task queue and the resource usage of each computing node specifically includes:

[0025] Setting a maximum queue length and a rejection policy;

[0026] When the task queue reaches the maximum queue length, obtain the amount of to-be-processed tasks in the task queue. If the total resource requirement of the current to-be-processed task amount is greater than the maximum acceptance capacity of a computing node, determine whether the number of computing nodes has reached the maximum number. If the number of computing nodes has not reached the maximum number, add virtual computing nodes through cloud resources;

[0027] When the task queue does not reach the maximum queue length and the current cluster server is not fully loaded, give priority to allocating tasks to non-virtual computing nodes; and set a timeout for newly added virtual computing nodes. When the virtual computing node does not receive tasks within the timeout range, destroy the virtual computing node.

[0028] In a second aspect, a cluster task scheduling system based on prior values includes:

[0029] A computing resource monitoring module for monitoring the usage of computing resources of each computing node through a central node; the central node is used to implement resource monitoring, a registration center, and task scheduling of a cluster server, and the computing node is used for collecting local resource data and processing tasks, and the resource data includes the number of CPU cores, CPU usage rate, total memory, and memory usage;

[0030] A to-be-processed task analysis module for automatically putting the tasks created by users into a task queue, analyzing the resource requirements of the to-be-processed tasks in the task queue in chronological order based on a prior value knowledge base, and scoring each computing node according to the resource requirements of the to-be-processed tasks and the usage of computing resources of each computing node;

[0031] A task allocation module for scheduling the to-be-processed tasks to the computing node with the highest score for processing through the central node.

[0032] In a third aspect, a computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described above.

[0033] Advantages of the present invention:

[0034] (1) By monitoring the changes in the computing resources of computing nodes in a cluster server, scoring each computing node according to the to-be-processed tasks, and allocating the to-be-processed tasks to the node with the highest score. When scoring the computing nodes, the standard deviation is introduced to measure the volatility and a penalty term is added to improve the accuracy of scoring.

[0035] (2) The monitoring and collection service used in the present invention is the Prometheus open-source monitoring component. Generally, open-source monitoring systems are built for units with spare capacity by combining open-source components, reducing the development volume, quickly providing monitoring capabilities, and laying the foundation for subsequent customized monitoring systems.

[0036] (3) The dynamic and scalable node method provided by the present invention can dynamically increase or decrease nodes according to the changes in the current tasks and the computing resources of computing nodes, ensuring that the system can provide sufficient computing resources to improve its performance under high load, and can save resources under low load. By dynamically adjusting the resource allocation of nodes, load balancing is achieved, and the overall performance and response speed of the system are improved. Description of the Drawings

[0037] Figure 1Schematic flowchart of a cluster task scheduling method based on prior values provided by an embodiment of the present invention;

[0038] Figure 2 Schematic flowchart of a dynamically scalable node provided by an embodiment of the present invention. Detailed implementation manners

[0039] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] As Figure 1 shown, a cluster task scheduling method based on prior values provided by an embodiment of the present invention includes:

[0041] S1: The cluster server monitors the usage of computing resources of each computing node through the central node; the central node is used to implement resource monitoring, registration center, and task scheduling of the cluster server, the computing node is used for collecting local resource data and processing tasks, and the resource data includes the number of CPU cores, CPU usage rate, total memory, and memory usage;

[0042] Specifically, the computing node includes a windows-exporter collection module, an altermanager alarm module, and a pushgateway push module, and the central node deploys a prometheus service module.

[0043] It can be understood that the user can select different exporter collectors or customize exporter collectors according to their own needs.

[0044] The windows-exporter collection module is used to collect the resource data of the computing node and adaptively adjust the collection frequency according to the current usage of resources of the computing node.

[0045] It can be understood that windows-exporter can collect various metrics of the windows system and expose these metrics to peometheus for monitoring and analysis. The user can flexibly enable or disable specific metric collectors through a configuration file according to their own needs to meet different monitoring requirements.

[0046] The altermanager alert module is used to implement the monitoring and alerting function. After receiving the alert information from the central node, it notifies the administrator via email. By customizing the rules for alerts, when the central node receives the resource data of the computing nodes, it determines whether the threshold is exceeded according to the rules. If the threshold is exceeded, an alert message will be sent to the alertmanager, and the alertmanager will notify the system administrator via email.

[0047] Specifically, define the notification receiving method and routing rules in alertmanager.yml, and add the alerting configuration in prometheus.yml. Moreover, the alertmanager can send the alert information to different recipients.

[0048] The pushgateway push module is used to push the received data to the prometheus service module. The pushgateway is actively sent by customizing the monitoring metric script. In the background Java program, it is necessary to record the task processing situation of the computing nodes, the current number of tasks being processed, and the total number of tasks being processed in real time. These custom information are statistically calculated by the background Java program and sent to the pushgateway push module via HTTP in two ways: regular push and completed task push. After receiving the data, the pushgateway push module pushes the data to the peometheus service module. In the background service, the collected resource data is actively pushed to the pushgateway push module via requests.

[0049] The peometheus service module is used to collect, store, and query time series data, and receive the resource data collected by the windows-exporter collection module.

[0050] Furthermore, the central node also deploys a registration center service, puts the scrape_configs.targets in the promethus configuration file into the cache of the registration center service. When a computing node is created or destroyed, it notifies the registration center service. After receiving the notification, the registration center service pulls the current computing node registration list and compares it with the cache. If they are inconsistent, it updates the cache and restarts the premethus service.

[0051] Specifically, the monitoring targets of Prometheus are based on the configuration item "targets" in prometheus.yml, and each startup can only detect according to the "targets" in the current prometheus.yml. In the cluster server, the startup or shutdown of the computing node will cause incomplete Prometheus monitoring. Deploy a registration center on the central node and deploy a registration center service on the central node. Put the "scrape_configs.targets" in the Prometheus configuration file into the cache of the registration center service. When a computing node is created or destroyed, it will notify the registration center service. After receiving the notification, the registration center service will pull the current computing node registration list and compare it with the cache. If they are inconsistent, it will update the cache and restart the Prometheus service to achieve the dynamic monitoring function of the cluster server.

[0052] Specifically, the prior knowledge base is constructed as follows:

[0053] Establish a mathematical model for the computing resources and time required to process tasks of different sizes to form a prior knowledge base.

[0054] Specifically, different algorithms are used to process different types of remote sensing images (tasks). For example, the GeoRectify_S algorithm is used for geometric types, the RAD_ZY303_TMS_S algorithm is used for radiation types, the LSR_ZY303_TMS_S algorithm is used for land surface types, the NDVI_ZY303_TMS_S algorithm is used for vegetation types, the AOD_ZY303_TMS_S algorithm is used for atmospheric types, and the SSC_ZY303_TMS_S algorithm is used for water body types. By establishing a prior knowledge base, a mathematical model of the computing resources required for algorithm processing of remote sensing image data of different sizes and the time curve of the required computing resources are constructed.

[0055] S2: Automatically put the tasks created by the user into the task queue, analyze the resource requirements of the tasks to be processed in the task queue in chronological order based on the prior knowledge base, and score each computing node according to the resource requirements of the tasks to be processed and the computing resource usage of each computing node.

[0056] Specifically, all the tasks to be inspected submitted by the user on the web page are put into the task queue, and a persistent operation is performed on the task queue to prevent task loss caused by system downtime or service crash. Define the CPU utilization rate of each server node as CPU server , and the memory resource utilization rate as Memory server , by analyzing the CPU utilization rate CPU task required by the task and the memory usage rate Memory taskScore each computing node according to the real-time load of the cluster server. Use the following formula to calculate the score of each node of the server Score i To avoid assigning tasks to servers with large resource fluctuations, the standard deviation is introduced to measure the volatility and a penalty term, Penalty i is the penalty standard deviation of each computing node. Finally, the computing node with the highest score is used to process the task:

[0057] Specifically, the scoring formula is as follows:

[0058]

[0059] Among them, FinalScore i represents the final score of computing node i, α, β, and γ represent weight parameters, Score i represents the score of computing node i for the task, Penalty i represents the penalty standard deviation of computing node i, represents the average remaining CPU resources of the server within time T, represents the average CPU usage rate of the task, represents the average remaining memory resources of the server within time T, represents the average memory usage rate of the task, CPU server represents the CPU usage rate of computing node i, Memory server represents the memory usage rate of computing node i, n represents the number of sampling points, and T represents the task execution time.

[0060] S3: Schedule the task to be processed to the computing node with the highest score through the central node for processing.

[0061] In the embodiment of the present invention, by monitoring the change of computing resources of computing nodes in the cluster server, scoring each computing node according to the task to be processed, and allocating the task to be processed to the node with the highest score. When scoring the computing node, the standard deviation is introduced to measure the volatility and a penalty term is added to improve the accuracy of the score.

[0062] Based on the above embodiment, the method provided in this embodiment further includes: monitoring the usage of computing resources occupied during the processing of the task to be processed to update the prior knowledge base.

[0063] Based on the above embodiment, the method provided in this embodiment further includes: as Figure 2 shown, dynamically reduce or increase the number of computing nodes by analyzing the number of tasks to be processed in the current task queue and the resource usage of each computing node.

[0064] Specifically, set the maximum queue length and rejection policy.

[0065] When the task queue reaches the maximum queue length, obtain the amount of tasks to be processed in the task queue. If the total resource requirements of the current tasks to be processed are greater than the maximum acceptance capacity of a computing node, determine whether the number of computing nodes has reached the maximum number. If the number of computing nodes has not reached the maximum number, add virtual computing nodes through cloud resources; where the maximum number of computing nodes is determined according to the size of the cloud resources, and the maximum number is the computing nodes + virtual computing nodes. Since the number of computing nodes is fixed, the maximum number of computing nodes is determined by the size of the cloud resources, that is, how many virtual computing nodes can be expanded at most.

[0066] Specifically, create virtual machines in the way of KVM, and the image file of the virtual machine is an image that has been prepared in advance and has the service capacity of computing nodes. The total resource requirements of the current tasks to be processed are greater than the maximum acceptance capacity of a computing node. Its design purpose is to ensure that tasks can be processed as soon as possible. If it is set as the maximum acceptance capacity of all computing nodes, more tasks will accumulate, affecting the final completion time of the tasks.

[0067] If the number of computing nodes reaches the maximum number, store the tasks in the database.

[0068] When the task queue has not reached the maximum queue length and the current cluster server is not fully loaded, preferentially allocate tasks to non-virtual computing nodes; and set the timeout for newly added virtual computing nodes. When a virtual computing node does not receive tasks within the timeout period, destroy the virtual computing node; if the current cluster server is fully loaded, return the tasks to the task queue.

[0069] It can be understood that if the virtual computing nodes exist for too long, it will cause waste of resources, and the processing ability of virtual computing nodes is not as good as that of physical computing nodes. Tasks need to be concentrated on physical computing nodes, that is, the original computing nodes in the cluster server.

[0070] The dynamic scalable node method provided by the embodiments of the present invention can dynamically increase or decrease nodes according to the changes of current tasks and the computing resources of computing nodes, ensuring that the system can provide sufficient computing resources to improve its performance under high load, while saving resources under low load. By dynamically adjusting the resource allocation of nodes, load balancing is achieved, and the overall performance and response speed of the system are improved.

[0071] The embodiments of the present invention also provide a cluster task scheduling system based on prior values, including:

[0072] A computing resource monitoring module is used to monitor the computing resource usage of each computing node through a central node; the central node is used to implement resource monitoring, a registration center, and task scheduling for a cluster server, and the computing node is used to collect local resource data and process tasks. The resource data includes the number of CPU cores, CPU usage rate, total memory, and memory usage.

[0073] A to-be-processed task analysis module is used to automatically put the tasks created by users into a task queue, analyze the resource requirements of the to-be-processed tasks in the task queue in chronological order based on a prior value knowledge base, and score each computing node according to the resource requirements of the to-be-processed tasks and the computing resource usage of each computing node.

[0074] A task allocation module is used to schedule the to-be-processed tasks to the computing node with the highest score for processing through the central node.

[0075] An embodiment of the present invention further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement any one of the above methods. For example:

[0076] The cluster server monitors the computing resource usage of each computing node through the central node; the central node is used to implement resource monitoring, a registration center, and task scheduling for the cluster server, and the computing node is used to collect local resource data and process tasks. The resource data includes the number of CPU cores, CPU usage rate, total memory, and memory usage.

[0077] Automatically put the tasks created by users into a task queue, analyze the resource requirements of the to-be-processed tasks in the task queue in chronological order based on a prior value knowledge base, and score each computing node according to the resource requirements of the to-be-processed tasks and the computing resource usage of each computing node.

[0078] Schedule the to-be-processed tasks to the computing node with the highest score for processing through the central node.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cluster task scheduling method based on a priori value, characterized in that: include: The cluster server monitors the computing resource usage of each computing node through the central node; the central node is used to implement resource monitoring, registration center and task scheduling of the cluster server, and the computing node is used to collect local resource data and process tasks. The resource data includes the number of CPU cores, CPU usage, total memory and memory usage; Automatically put the tasks created by the user into the task queue, analyze the resource requirements of the pending tasks in the task queue in chronological order based on the prior value knowledge base, and score each computing node according to the resource requirements of the pending tasks and the computing resource usage of each computing node; The central node dispatches the task to be processed to the computing node with the highest score for processing.

2. The cluster task scheduling method based on a priori value according to claim 1, characterized in that: The method further includes: monitoring the usage of computing resources occupied by the processing of the pending task to update the prior value knowledge base.

3. The cluster task scheduling method based on a priori value according to claim 1, characterized in that: The computing node deployment includes windows-exporter collection module, altermanager alarm module and pushgateway push module, and the central node deploys prometheus service module; The windows-exporter collection module is used to collect resource data of computing nodes and adaptively adjust the collection frequency according to the resource usage of the current computing nodes; The altermanager alarm module is used to implement the monitoring alarm function and notify the administrator by email after receiving the alarm information of the central node; The pushgateway push module is used to push the received data to the prometheus service module; The peometheus service module is used to collect, store and query time series data, and accept resource data collected by the windows-exporter collection module.

4. The cluster task scheduling method based on a priori value according to claim 1, characterized in that: The prior value knowledge base is constructed as follows: A mathematical model of the computing resources and time required to process tasks of different sizes is established to form the prior knowledge base.

5. The cluster task scheduling method based on a priori value according to claim 1, characterized in that: The central node also deploys a registration center service, which is specifically used to: put the scrape_configs.targets in the configuration file of the promethus service module into the cache of the registration center service, notify the registration center service when any computing node is created or destroyed, and after receiving the notification, the registration center service pulls the current computing node registration list and compares it with the cache. If there is inconsistency, the cache is updated and the premethus service module is restarted.

6. The cluster task scheduling method based on a priori value according to claim 1, characterized in that: The computing nodes are scored according to the resource requirements of the tasks to be processed and the computing resource usage of the computing nodes. The scoring formula is as follows: Among them, FinalScore i represents the final score of computing node i, α, β and γ represent weight parameters, Score i represents the score of the task given by computing node i, Penalty i represents the penalty standard deviation of computing node i, Indicates the average remaining CPU resources of the server within T time. Indicates the average CPU usage of the task. Indicates the average remaining memory resources of the server within T time. Indicates the average memory usage of the task, CPU server Indicates the CPU usage of computing node i, Memory server represents the memory usage of computing node i, n represents the number of sampling points, and T represents the task execution time.

7. The cluster task scheduling method based on a priori value according to claim 1, characterized in that: The method further includes: dynamically reducing or increasing the number of computing nodes by analyzing the number of pending tasks in the current task queue and the resource usage of each computing node.

8. The cluster task scheduling method based on a priori value according to claim 7, characterized in that: The dynamically reducing or increasing the number of computing nodes by analyzing the number of pending tasks in the current task queue and the resource usage of each computing node specifically includes: Set maximum queue length and rejection policy; When the task queue reaches the maximum queue length, the amount of pending tasks in the task queue is obtained, and if the sum of resource requirements of the current amount of pending tasks is greater than the maximum acceptance capacity of a computing node, it is determined whether the number of computing nodes has reached the maximum number, and if the number of computing nodes has not reached the maximum number, virtual computing nodes are added through cloud resources; When the task queue has not reached the maximum queue length and the current cluster server has not reached full load, tasks are preferentially assigned to non-virtual computing nodes; and a timeout period for adding new virtual computing nodes is set. When the virtual computing node does not receive a task within the timeout period, the virtual computing node is destroyed.

9. A cluster task scheduling system based on prior values, characterized in that: include: The computing resource monitoring module is used to monitor the computing resource usage of each computing node through the central node; the central node is used to implement resource monitoring, registration center and task scheduling of the cluster server, and the computing node is used to collect local resource data and process tasks. The resource data includes the number of CPU cores, CPU usage, total memory and memory usage; The pending task analysis module is used to automatically put the tasks created by the user into the task queue, analyze the resource requirements of the pending tasks in the task queue in chronological order based on the prior value knowledge base, and score each computing node according to the resource requirements of the pending tasks and the computing resource usage of each computing node; The task allocation module is used to dispatch the to-be-processed tasks to the computing nodes with the highest scores for processing through the central node.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.