Big Data-Based Task Scheduling Optimization Analysis Method and System
By collecting and analyzing task data in real time in distributed big data processing scenarios, and using the Hadoop framework for scheduling clustering and optimization, the problem of unreasonable resource allocation of traditional task scheduling algorithms in the big data era is solved, and task execution efficiency and system performance are improved.
Patent Information
- Application Number
- CN202510387541.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Traditional task scheduling algorithms are difficult to effectively handle the growth of task scale and complexity in computer systems in the big data era, and cannot fully consider the real-time requirements of tasks, resource usage volatility and dependencies, resulting in unreasonable resource allocation and inefficient task execution.
By collecting task-related data in real time in distributed big data processing scenarios, using the Hadoop distributed framework for big data feature mining and scheduling cluster analysis, a task scheduling cluster group is generated, and the scheduling priority is determined based on the task urgency and resource requirements, and a dynamic optimization scheduling scheme is generated.
Accurate task planning and progress tracking are realized, resource waste and delays are avoided, computing resource usage efficiency and task completion timeliness, and computing performance and task completion efficiency of distributed big data processing platform are improved.
Smart Images

Figure CN119883580B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and system for optimizing and analyzing task scheduling based on big data. Background Art
[0002] In a computer system, task scheduling is a key link for reasonably allocating system resources and ensuring efficient task execution. Traditional task scheduling algorithms mostly make scheduling decisions based on simple factors such as task priority and resource requirements, such as the First-Come, First-Served (FCFS), Shortest Job First (SJF) algorithms. At the same time, with the advent of the big data era, the scale and complexity of tasks faced by computer systems have increased exponentially, and these traditional algorithms have gradually exposed many deficiencies.
[0003] In scenarios such as cloud computing and data centers, a large number of tasks are submitted simultaneously, and the types of tasks are rich and diverse, including data processing, machine learning model training, real-time communication, etc. However, traditional scheduling algorithms are difficult to comprehensively consider the complex characteristics and dynamic changes of distributed tasks between computer devices, such as the real-time requirements of tasks, the volatility of resource usage, and the dependencies between different tasks, lacking in-depth analysis of task computing nodes and system resource usage, and unable to fully exploit data value, resulting in unreasonable resource allocation and low task execution efficiency. Summary of the Invention
[0004] Based on this, it is necessary for the present invention to provide a method and system for optimizing and analyzing task scheduling based on big data to solve at least one of the above technical problems.
[0005] To achieve the above object, a method for optimizing and analyzing task scheduling based on big data includes the following steps:
[0006] Step S1: Real-time collect task-related data in a corresponding distributed big data processing scenario among servers, storage devices, and clients, including task submission time, task deadline, task type, and task resource requirements; calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenario based on the task resource requirements to obtain task system resource usage data; measure the task status of the task computing nodes corresponding to the distributed big data processing scenario to obtain task execution status data, including task execution duration, actual task resource usage, and task completion status.
[0007] Step S2: Use the Hadoop distributed framework to perform big data feature mining and analysis on task-related data, task system resource usage data, and task execution status data to obtain a distributed task big data feature set; based on the distributed task big data feature set, perform task scheduling clustering analysis on the corresponding task computing nodes to generate a task scheduling cluster group corresponding to similar resource requirements;
[0008] Step S3: Calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the actual task resource usage to obtain the remaining execution calculation duration of the task; perform task urgency accounting on the corresponding task computing nodes based on the task deadline and the remaining execution calculation duration of the task to obtain the task node scheduling urgency;
[0009] Step S4: Determine the scheduling priority of the corresponding task computing nodes according to the task node scheduling urgency to obtain the task node scheduling priority; perform task scheduling optimization analysis on the task scheduling cluster group corresponding to similar resource requirements based on the task node scheduling priority to generate a task dynamic optimization scheduling plan.
[0010] Further, step S1 includes the following steps:
[0011] Step S11: Configure and deploy each distributed network device between the server, storage device, and client to ensure a stable network connection between the devices in the corresponding task access test environment of the server, storage device, and client, and generate a corresponding distributed big data processing scenario;
[0012] Step S12: Real-time collect task-related data of the corresponding distributed access operation tasks in the distributed big data processing scenario between the server, storage device, and client, including task submission time, task deadline, task type, and task resource requirements;
[0013] Step S13: Calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenario based on the task resource requirements to obtain the task system resource usage data;
[0014] Step S14: Deploy data collection interfaces at the system kernel, task management module, and resource monitoring tool locations of the task computing nodes corresponding to the distributed big data processing scenario, and use the data collection interfaces to measure the task status of the task computing nodes corresponding to the distributed big data processing scenario to obtain the task execution status data, including task execution duration, actual task resource usage, and task completion status.
[0015] Further, the task resource requirements described in step S12 specifically include the resource requirements corresponding to CPU, memory, storage, and network bandwidth.
[0016] Further, step S13 includes the following steps:
[0017] Step S131: Calculate the CPU utilization rate of the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the CPU, so as to obtain the CPU utilization rate corresponding to each computing node;
[0018] Step S132: Conduct memory free space accounting for the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the memory, so as to obtain the memory free space corresponding to each computing node;
[0019] Step S133: Conduct storage read and write statistical analysis on the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the storage, so as to obtain the storage read and write speed corresponding to each computing node;
[0020] Step S134: Conduct bandwidth occupancy accounting for the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the network bandwidth, so as to obtain the network bandwidth occupancy corresponding to each computing node;
[0021] Step S135: Merge the CPU utilization rate, memory free space, storage read and write speed, and network bandwidth occupancy corresponding to each computing node for resource usage to obtain the task system resource usage data.
[0022] Further, step S2 includes the following steps:
[0023] Step S21: Use the Hadoop distributed framework to perform big data multi-source heterogeneous storage on the task-related data, task system resource usage data, and task execution status data to obtain a distributed task big data storage set;
[0024] Step S22: Conduct big data cleaning on the distributed task big data storage set to remove noise, duplicate values, and outliers in the dataset to obtain a distributed task big data cleaning set;
[0025] Step S23: Conduct data normalization on the distributed task big data cleaning set to unify the task resources corresponding to different magnitudes to the same scale to obtain a distributed task big data normalization set;
[0026] Step S24: Conduct big data feature mining and analysis on the distributed task big data normalization set to analyze the distributed task description using text classification methods to determine the category to which the task belongs, and use data mining techniques to extract the resource requirement pattern features corresponding to the distributed task to obtain a distributed task big data feature set;
[0027] Step S25: Perform task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set to generate a task scheduling cluster group corresponding to similar resource requirements.
[0028] Further, step S3 includes the following steps:
[0029] Step S31: Calculate the resource usage fluctuation of the task computing nodes corresponding to the distributed big data processing scenario based on the actual resource usage of the tasks to obtain the task node resource usage volatility;
[0030] Step S32: Calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the task node resource usage volatility to obtain the remaining execution calculation duration of the task;
[0031] Step S33: Perform task urgency accounting on the corresponding task computing nodes based on the task deadline and the remaining execution calculation duration of the task using the task scheduling urgency calculation formula to obtain the task node scheduling urgency.
[0032] Further, step S32 includes the following steps:
[0033] Step S321: Statistically calculate the task node resource usage variance and the task node resource usage covariance corresponding to the task node resource usage volatility;
[0034] Step S322: Perform resource load balancing evaluation on the task computing nodes corresponding to the distributed big data processing scenario based on the task node resource usage variance and the task node resource usage covariance to obtain the task node resource load balancing degree;
[0035] Step S323: Predict the remaining resource capacity of the corresponding task computing nodes based on the task node resource load balancing degree to obtain the remaining resource capacity of the task nodes;
[0036] Step S324: Estimate the remaining execution duration of the corresponding task computing nodes based on the task execution duration and the remaining resource capacity of the task nodes to obtain the remaining execution calculation duration of the task.
[0037] Further, the task scheduling urgency calculation formula described in step S33 is specifically:
[0038] ;
[0039] In the formula, is the task node scheduling urgency, is the number of task computing nodes, is the The system resource usage of a task computing node is the maximum resource usage of the task computing node is the task deadline is the current time point is the remaining execution computing duration of the task
[0040] Further, step S4 includes the following steps:
[0041] Step S41: Determine the scheduling priority of the corresponding task computing node according to the scheduling urgency of the task node to obtain the task node scheduling priority;
[0042] Step S42: Arrange the execution order within the group for the task scheduling cluster groups corresponding to similar resource requirements based on the task node scheduling priority to obtain the task node intra-group scheduling arrangement sequence corresponding to each task scheduling cluster group;
[0043] Step S43: Optimize the inter-group task scheduling according to the task node intra-group scheduling arrangement sequence corresponding to each task scheduling cluster group, prioritize the inter-group task scheduling and execute the task node scheduling arrangement corresponding to the task scheduling cluster group to generate a task dynamic optimization scheduling plan.
[0044] Further, the present invention also provides a task scheduling optimization analysis system based on big data for executing the above-mentioned task scheduling optimization analysis method based on big data. The task scheduling optimization analysis system based on big data includes:
[0045] A distributed task big data collection module for real-time collecting task-related data in the corresponding distributed big data processing scenarios among servers, storage devices, and clients, including task submission time, task deadline, task type, and task resource requirements; calculating the system resource usage of the task computing nodes corresponding to the distributed big data processing scenarios based on the task resource requirements to obtain task system resource usage data; measuring the task status of the task computing nodes corresponding to the distributed big data processing scenarios to obtain task execution status data, including task execution duration, actual task resource usage, and task completion status;
[0046] A task resource requirement scheduling clustering module for performing big data feature mining and analysis on task-related data, task system resource usage data, and task execution status data by using the Hadoop distributed framework to obtain a distributed task big data feature set; performing task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set to generate task scheduling cluster groups corresponding to similar resource requirements;
[0047] The task node scheduling emergency evaluation module is used to calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the actual task resource usage, so as to obtain the remaining execution computing duration of the task; and conduct task emergency accounting for the corresponding task computing nodes based on the task deadline and the remaining execution computing duration of the task, thereby obtaining the task node scheduling emergency level.
[0048] The task scheduling optimization analysis module is used to determine the scheduling priority of the corresponding task computing nodes according to the task node scheduling emergency level, so as to obtain the task node scheduling priority; and conduct task scheduling optimization analysis on the task scheduling cluster groups with similar resource requirements based on the task node scheduling priority, so as to generate a task dynamic optimization scheduling scheme.
[0049] Advantages of the present invention:
[0050] 1. The task scheduling optimization analysis method based on big data proposed by the present invention, compared with the prior art, the beneficial effects of the present application are as follows: By collecting task-related data in the distributed big data processing scenario in real time, the collected data includes task submission time, task deadline, task type, and task resource requirements, etc., which ensures that during the task execution process, the start and end times of each task, the types of tasks, and the computing resources required by the tasks can be accurately grasped. The acquisition of these data not only helps the system to perform precise task planning in the early stage of task execution, but also can effectively track the task progress, avoiding task delays or resource waste caused by insufficient or excessive resource allocation. Further, through the combination of system resource usage calculation and task execution status measurement, the resource consumption of computing nodes can be evaluated in real time, providing detailed feedback on task execution, including task execution duration, actual resource usage, and task completion status. This provides accurate dynamic data for subsequent task scheduling, helping the system to make effective resource adjustments during task execution, thereby improving the usage efficiency of computing resources and the timeliness of task completion. Secondly, by applying the Hadoop distributed framework to perform big data feature mining and analysis on task-related data, task system resource usage data, and task execution status data, the hidden features and rules in the task execution process can be identified. This analysis process enables the system to extract the similarity features between tasks, providing a more accurate reference basis for subsequent task scheduling. Through scheduling clustering analysis of tasks, tasks with similar resource requirements can be grouped into the same scheduling cluster, thereby achieving efficient resource allocation. For example, some tasks require a large amount of computing resources or memory, while other tasks have lower requirements for computing resources, thus avoiding mixed scheduling of tasks with mismatched resource requirements and avoiding problems such as resource contention, low efficiency, or task delay, which can significantly improve the computing performance and task completion efficiency of the entire distributed big data processing platform. Then, by calculating the remaining execution duration of the task and the actual resource usage of the task, the progress and remaining time of the task can be accurately monitored dynamically. The calculation of the remaining execution duration helps to predict the completion time of each task, and thus provides a basis for task scheduling and resource allocation. By combining the task deadline and the remaining execution duration, the scheduling urgency of each task node can be evaluated. This analysis provides a basis for task priority sorting in the task scheduling system, ensuring that urgent tasks can obtain computing resources first, avoiding timeouts or uncompleted situations caused by improper resource scheduling of tasks, and thus improving the service quality of the entire distributed big data processing system.Finally, by calculating the scheduling priority of task nodes and dynamically optimizing the task scheduling cluster group, efficient scheduling of tasks can be achieved according to the urgency and resource requirements of the tasks. The core of this process lies in determining the scheduling priority of task nodes, enabling the system to allocate resources precisely according to the actual requirements of the tasks. For tasks with high urgency, the scheduling priority will be higher to ensure that these tasks can obtain the required computing resources in a timely manner. Tasks with fewer resources and lower urgency can be delayed for execution when resources permit. This priority scheduling based on task urgency can effectively avoid resource waste or task delays caused by unreasonable task scheduling. At the same time, the scheduling optimization based on priority can make a balance between different task groups, preventing a certain cluster group from occupying too many resources, resulting in resource shortages for other tasks. This dynamically optimized scheduling method can greatly improve the overall performance of the system, fully explore the value of data at task computing nodes, reduce waste of computing resources, and ensure that tasks are completed on time, thereby improving the execution efficiency of tasks.
[0051] 2. The task scheduling optimization analysis system based on big data proposed by the present invention is generally composed of a distributed task big data collection module, a task resource demand scheduling clustering module, a task node scheduling emergency evaluation module, and a task scheduling optimization analysis module, and can implement any task scheduling optimization analysis method based on big data of the present invention. It is used to realize the task scheduling optimization analysis method based on big data through the operations between computer programs running on each module. The internal structure of the system cooperates with each other, which can greatly reduce repetitive work and manpower investment, and can quickly and effectively provide a more accurate and efficient task scheduling optimization analysis process based on big data, thereby simplifying the operation process of the task scheduling optimization analysis system based on big data. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Other features, purposes, and advantages of the present invention will become more obvious by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0053] Figure 1 It is a schematic flowchart of the steps of the task scheduling optimization analysis method based on big data of the present invention;
[0054] Figure 2 is Figure 1 a detailed schematic flowchart of step S1 in
[0055] Figure 3 is Figure 2 a detailed schematic flowchart of step S13 in DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0057] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.
[0058] It should be understood that although terms such as "first" and "second" may be used here to describe each unit, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0059] To achieve the above object, please refer to Figures 1 to 2 , the present invention provides a method for optimizing and analyzing task scheduling based on big data, and the method includes the following steps:
[0060] Step S1: Real-time collect task-related data in the corresponding distributed big data processing scenario among the server, storage device, and client, including task submission time, task deadline, task type, and task resource requirements; calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenario based on the task resource requirements to obtain task system resource usage data; measure the task status of the task computing nodes corresponding to the distributed big data processing scenario to obtain task execution status data, including task execution duration, actual task resource usage, and task completion status;
[0061] Step S2: Use the Hadoop distributed framework to perform big data feature mining and analysis on the task-related data, task system resource usage data, and task execution status data to obtain a distributed task big data feature set; perform task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set to generate a task scheduling cluster group corresponding to similar resource requirements;
[0062] Step S3: Calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the actual usage of task resources, so as to obtain the remaining execution calculation duration of the task; perform task urgency accounting on the corresponding task computing nodes based on the task deadline and the remaining execution calculation duration of the task, and obtain the task node scheduling urgency level;
[0063] Step S4: Determine the scheduling priority of the corresponding task computing nodes according to the task node scheduling urgency level, so as to obtain the task node scheduling priority; perform task scheduling optimization analysis on the task scheduling cluster groups corresponding to similar resource requirements based on the task node scheduling priority, so as to generate a task dynamic optimization scheduling scheme.
[0064] In the embodiment of the present invention, please refer to Figure 1 As shown in the figure, it is a schematic flow chart of the steps of the task scheduling optimization analysis method based on big data of the present invention. In this example, the task scheduling optimization analysis method based on big data includes the following steps:
[0065] Step S1: Real-time collect task-related data in the corresponding distributed big data processing scenario among the server, storage device, and client, including task submission time, task deadline, task type, and task resource requirements; calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenario based on the task resource requirements, and obtain the task system resource usage data; measure the task status of the task computing nodes corresponding to the distributed big data processing scenario, and obtain the task execution status data, including task execution duration, actual task resource usage, and task completion status;
[0066] In the embodiments of the present invention, in the distributed big data processing scenario, the collection of task-related data is crucial. First, when a task is submitted, the system automatically records information such as the task submission time, task deadline, task type, and task resource requirements. The task resource requirements include the required amounts of resources such as CPU, memory, storage space, and bandwidth. This data is collected in real time among servers, storage devices, and clients through a distributed data collection system and stored in a distributed database in a standardized format. For the calculation of the resource usage of task computing nodes, through the interface with a cluster management tool (such as YARN or Mesos), the resource usage of each node is obtained, including the resource usage corresponding to CPU, memory, storage, and network bandwidth, and the resource allocation amount of each computing node is calculated to generate task system resource usage data. In addition, the measurement of the task execution status is monitored through system logs, the task execution duration is automatically recorded through timestamps, the actual task resource usage is tracked in real time by the system monitoring tool, the task completion status is classified by determining whether the task has successfully terminated or there are any exceptions, and the task status data is automatically updated, finally obtaining the task execution status data.
[0067] Step S2: Use the Hadoop distributed framework to perform big data feature mining and analysis on the task-related data, task system resource usage data, and task execution status data to obtain a distributed task big data feature set; perform task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set to generate a task scheduling cluster group corresponding to similar resource requirements;
[0068] In the embodiments of the present invention, by using the Hadoop distributed framework to perform big data feature mining and analysis on the collected task data, first, the task-related data, task system resource usage data, and task execution status data are cleaned and preprocessed to remove outliers and redundant information. Then, the MapReduce model or Apache Spark is applied for parallel data processing, combined with machine learning algorithms such as clustering analysis and regression analysis, to extract key features such as the time feature, resource consumption feature, and execution efficiency feature of the task, forming a distributed task big data feature set. When performing task scheduling clustering analysis, the tasks are clustered based on features such as task resource requirements and execution time. Through the K-means clustering algorithm or hierarchical clustering algorithm, task clusters with similar resource requirements and execution time are identified. Each task cluster corresponds to a specific task scheduling cluster group, and these cluster groups will help optimize the subsequent scheduling strategy, finally generating a task scheduling cluster group corresponding to similar resource requirements.
[0069] Step S3: Calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the actual usage of task resources, so as to obtain the remaining execution calculation duration of the task; conduct a task urgency accounting for the corresponding task computing nodes based on the task deadline and the remaining execution calculation duration of the task, and obtain the scheduling urgency degree of the task nodes.
[0070] In the embodiment of the present invention, the task execution duration and the actual usage of task resources are key indicators for calculating the remaining execution duration of the task. In this step, first, obtain the executed time and the consumed resources through the task execution status data, so as to calculate the current remaining execution duration of each task. Specifically, the system will establish a regression model based on historical data and real-time data in combination with the task type, task resource usage pattern and execution duration to predict the remaining execution time. The task deadline is known. Therefore, by subtracting the remaining execution duration of the task from the task deadline, the remaining execution calculation duration of the task can be obtained. Then, use the difference between the remaining execution calculation duration of the task and the task deadline to evaluate the urgency degree of the task node. The way of task urgency accounting is to calculate the weight of the remaining execution time of the task to evaluate whether the task needs to be scheduled preferentially. Tasks with a higher urgency degree will be marked as high-priority tasks, and finally the scheduling urgency degree of the task nodes is obtained.
[0071] Step S4: Determine the scheduling priority of the corresponding task computing nodes according to the scheduling urgency degree of the task nodes, so as to obtain the scheduling priority of the task nodes; conduct a task scheduling optimization analysis on the task scheduling cluster groups corresponding to similar resource requirements based on the scheduling priority of the task nodes, so as to generate a task dynamic optimization scheduling plan.
[0072] In the embodiment of the present invention, by according to the scheduling urgency degree of the task nodes, the task scheduling priority can be calculated for each task. The calculation of the priority takes into account the task deadline, the remaining execution duration, the task resource requirements and the urgency degree of the task. First, by setting a rule, the system can automatically assign a preliminary priority level, and tasks with a higher urgency degree are assigned as high priority. Then, the system will conduct an optimization analysis on the task scheduling cluster groups with similar resource requirements through optimization algorithms such as genetic algorithms or simulated annealing algorithms. This optimization analysis is based on factors such as the resource requirements of the task, the task execution duration, the remaining resources, etc., and combines the idle resource situation of the computing nodes to dynamically optimize the task scheduling plan. During the optimization process, tasks with a higher task scheduling priority will be preferentially scheduled to the computing nodes with sufficient resources for execution. The system will generate a task dynamic optimization scheduling plan, schedule the task to the most suitable computing node to ensure that the task can be completed within the specified time, and at the same time improve the utilization rate of computing resources, and finally generate a task dynamic optimization scheduling plan.
[0073] Further, step S1 includes the following steps:
[0074] Step S11: Configure and deploy each distributed network device among the server, the storage device, and the client to ensure a stable network connection among the devices in the corresponding task access test environment between the server, the storage device, and the client, and generate a corresponding distributed big data processing scenario;
[0075] Step S12: Real-time collect task-related data of the corresponding distributed access operation task in the corresponding distributed big data processing scenario among the server, the storage device, and the client, including the task submission time, the task deadline, the task type, and the task resource requirements;
[0076] Step S13: Calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenario based on the task resource requirements to obtain task system resource usage data;
[0077] Step S14: Deploy data collection interfaces at the system kernel, the task management module, and the resource monitoring tool positions of the task computing nodes corresponding to the distributed big data processing scenario, and use the data collection interfaces to measure the task status of the task computing nodes corresponding to the distributed big data processing scenario to obtain task execution status data, including the task execution duration, the actual task resource usage, and the task completion status.
[0078] As an embodiment of the present invention, refer to Figure 2 shown in Figure 1 is a detailed step flow diagram of step S1 in
[0079] Step S11: Configure and deploy each distributed network device among the server, the storage device, and the client to ensure a stable network connection among the devices in the corresponding task access test environment between the server, the storage device, and the client, and generate a corresponding distributed big data processing scenario;
[0080] In an embodiment of the present invention, during the creation process of a distributed big data processing scenario, the network connections between servers, storage devices, and clients are first configured and deployed to ensure stable and efficient data exchange between these devices. In a specific implementation, a physical network connection is established between different devices using high-speed network switches and routers. At the same time, logical isolation is achieved through technologies such as configuring virtual local area networks (VLANs) and virtual private networks (VPNs) to ensure the security and reliability of data transmission. To ensure the smooth transfer of data streams in the big data processing environment, it is necessary to reasonably allocate bandwidth according to the task loads and traffic requirements of each device. A network load balancer (such as a load balancing device from F5 or Cisco) is used to achieve optimized traffic allocation. An efficient storage network (such as Fibre Channel SAN or iSCSI) is deployed between the server and the storage device to improve the speed of data access and ensure stable and efficient data transmission between the client and the computing node. In this way, it is ensured that the task access of each device in the entire distributed environment can proceed smoothly, and finally, the corresponding distributed big data processing scenario is configured and generated.
[0081] Step S12: In the corresponding distributed big data processing scenario among the server, storage device, and client, task-related data of the corresponding distributed access operation task is collected in real time, including the task submission time, task deadline, task type, and task resource requirements.
[0082] In an embodiment of the present invention, after the deployment of the distributed big data processing environment is completed, the task operation data is monitored and recorded in real time through the configured data collection system. In a specific implementation, distributed log collection tools (such as ELK Stack, Logstash) are used to deploy log collectors on each node to record in detail the submission time, deadline, type, and resource requirements of the task. A unique task identifier is generated when each task is submitted. The task is uniformly scheduled and tracked through a task management system (such as Apache Oozie, Airflow). The submission time and deadline of the task are recorded through an automatic timestamp and transmitted to the data warehouse through the log collection tool. The type and resource requirements of the task are standardized through configured predefined templates or interfaces and are collected and uploaded to the data center for storage in real time. All the collected task-related data will be provided to the subsequent resource scheduling and optimization analysis module, and finally, the task-related data of the distributed access operation task is obtained, including the task submission time, task deadline, task type, and task resource requirements.
[0083] Step S13: Calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenario based on the task resource requirements to obtain the task system resource usage data.
[0084] In an embodiment of the present invention, after obtaining the resource requirement data of a task, for the resource requirements of each task (i.e., the resource requirements corresponding to CPU, memory, storage, and network bandwidth), the calculation of the system resource usage of the task computing node is performed. Specifically, first, by using a cluster resource management tool (such as Apache YARN or Kubernetes) to monitor the resource utilization of each task computing node, these tools can obtain the usage of resources such as CPU, memory, disk I / O, and network bandwidth of the node in real time, and through algorithms (such as resource scheduling algorithms, task priority sorting algorithms), predictive calculations are performed according to the resource requirements and priorities of the tasks. Through these tools, the amount of resources required by each task during execution can be obtained, and then the task resource scheduling can be optimized. In this process, combined with a machine learning model, different types of task resource requirements can be learned and modeled to further improve the accuracy and efficiency of resource calculation. The result of the resource usage calculation will be used as the basis for subsequent task scheduling and task execution monitoring, including the CPU usage rate, memory free amount, storage read and write speed, and network bandwidth occupancy of each computing node, and finally, the task system resource usage data is obtained.
[0085] Step S14: Deploy data collection interfaces at the system kernel, task management module, and resource monitoring tool positions of the task computing node corresponding to the distributed big data processing scenario, and use the data collection interfaces to measure the task status of the task computing node corresponding to the distributed big data processing scenario to obtain task execution status data, including task execution duration, actual task resource usage, and task completion status.
[0086] In the embodiments of the present invention, by deploying data collection interfaces on each task computing node, the specific operation is to deploy a resource monitoring module and a task status monitoring module on the computing node kernel using system monitoring tools (such as Prometheus or Zabbix). These modules collect key metric data during the task execution process, such as task execution duration, resource usage (such as real-time occupancy of CPU and memory), and task completion status. When the task is executed, the monitoring module will record this data in real time and transmit it to the central database or data warehouse through the data collection interface for storage and analysis. In addition, by combining a task management module (such as Oozie or Airflow), the system can trigger automatic collection operations at the start and end of the task, record the complete life cycle data of the task, and judge and feedback on the task status according to preset conditions. After the status judgment is completed, the system will update the execution result (success or failure) of the task to the database. The monitoring tool displays the task execution process and resource utilization through a visual dashboard (such as Grafana), further optimizing the resource scheduling of the task, and finally obtaining task execution status data, specifically including task execution duration, actual task resource usage, and task completion status.
[0087] Further, the task resource requirements described in step S12 specifically include resource requirements corresponding to CPU, memory, storage, and network bandwidth.
[0088] Further, step S13 includes the following steps:
[0089] Step S131: Calculate the CPU usage rate of the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the CPU to obtain the CPU usage rate corresponding to each computing node;
[0090] Step S132: Conduct memory free calculation on the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the memory to obtain the memory free amount corresponding to each computing node;
[0091] Step S133: Conduct storage read and write statistical analysis on the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the storage to obtain the storage read and write speed corresponding to each computing node;
[0092] Step S134: Conduct bandwidth occupancy calculation on the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the network bandwidth to obtain the network bandwidth occupancy amount corresponding to each computing node;
[0093] Step S135: Combine the CPU usage rate, memory free amount, storage read and write speed, and network bandwidth occupancy amount corresponding to each computing node for resource usage to obtain task system resource usage data.
[0094] As an embodiment of the present invention, with reference to Figure 3 shown in Figure 2 is a detailed step flow schematic diagram of step S13 in
[0095] Step S131: Calculate the CPU usage rate of the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the CPU, so as to obtain the CPU usage rate corresponding to each computing node;
[0096] In the embodiment of the present invention, by obtaining the CPU resource information of each computing node, including the total number of CPU cores, frequency, and the amount of used CPU resources of each computing node, monitoring tools such as top, htop, psutil, etc. can be used to obtain the current usage of each core on the computing node. Then, the CPU usage rate of the computing node can be obtained by real-time monitoring of the CPU load. For example, the CPU usage rate can be calculated by the following formula: CPU usage rate = used CPU time / total CPU time × 100%. In this step, by statistically summarizing the usage of each CPU core of each computing node, the CPU usage rate corresponding to each computing node is finally obtained.
[0097] Step S132: Conduct memory free accounting on the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the memory, so as to obtain the memory free amount corresponding to each computing node;
[0098] In the embodiment of the present invention, by obtaining the total capacity of the memory and the currently used memory capacity on each computing node, the total amount and remaining amount of the memory can be extracted by using commands such as free, vmstat, or the psutil library. Specifically, first determine the total amount of memory on each computing node and obtain the amount of used memory. Then, calculate the free memory amount through the following formula: free memory amount = total memory - used memory. The free memory amount will reflect whether each computing node has enough memory to process new tasks. This calculation can not only help evaluate the current load of the computing node, but also provide a basis for optimizing the utilization rate of memory resources. Finally, the memory free amount corresponding to each computing node is obtained.
[0099] Step S133: Conduct storage read and write statistical analysis on the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the storage, so as to obtain the storage read and write speed corresponding to each computing node;
[0100] In an embodiment of the present invention, by obtaining the read and write rate information of the storage device of the computing node, the I / O operation of the disk can be obtained through monitoring tools such as iostat and dstat. To obtain the accurate storage read and write speed, the following steps can be taken: First, read the total read and write bytes of the disk I / O of each computing node, and then calculate the storage read and write speed according to the set time window. For example: storage read and write speed = total read and write bytes / time window. This will provide the performance data of the storage device for each computing node and help evaluate whether there is resource consumption or delay caused by disk I / O bottleneck, and finally obtain the storage read and write speed corresponding to each computing node.
[0101] Step S134: Based on the resource requirements corresponding to the network bandwidth, conduct bandwidth occupancy accounting for the task computing nodes corresponding to the distributed big data processing scenario to obtain the network bandwidth occupancy corresponding to each computing node;
[0102] In an embodiment of the present invention, by monitoring the network bandwidth usage of each computing node, the network bandwidth occupancy is usually monitored and recorded in real time by tools such as iftop, netstat, and nload. By obtaining the upload and download traffic data of the network interface of each computing node, calculate the network bandwidth utilization rate of each node. The calculation method is: bandwidth occupancy = data transmission volume / time window. Through this method, it is possible to evaluate the bandwidth consumption of each computing node during data transmission and understand whether the network has become a potential bottleneck, and finally obtain the network bandwidth occupancy corresponding to each computing node.
[0103] Step S135: Combine the CPU utilization rate, memory free amount, storage read and write speed, and network bandwidth occupancy corresponding to each computing node for resource usage combination to obtain the task system resource usage data.
[0104] In an embodiment of the present invention, by integrating the various resource indicators (CPU utilization rate, memory free amount, storage read and write speed, network bandwidth occupancy) calculated in steps S131 to S134, the specific operation can integrate the corresponding CPU utilization rate, memory free amount, storage read and write speed, and network bandwidth occupancy into the same big data set, and finally obtain the task system resource usage data.
[0105] Further, step S2 includes the following steps:
[0106] Step S21: Use the Hadoop distributed framework to perform big data multi-source heterogeneous storage on the task-related data, task system resource usage data, and task execution status data to obtain a distributed task big data storage set;
[0107] In an embodiment of the present invention, the task-related data, task system resource usage data, and task execution status data are stored in a multi-source heterogeneous manner through the Hadoop distributed framework. The HDFS (Hadoop Distributed FileSystem) is used as the storage layer to ensure that all data can be distributed and stored among multiple nodes. The task-related data includes the task submission time, task deadline, task type, and task resource requirements. The task system resource usage data includes the usage requirements of resources such as CPU, memory, disk I / O (i.e., storage), and network bandwidth for each task. The task execution status data records the task execution duration, the actual usage amount of task resources, and the task completion status (success, failure, or suspension), etc. The Hadoop MapReduce framework is used for the preliminary processing and distribution of data. The task data is stored as different data files in the HDFS. By setting a reasonable sharding strategy, the large-scale task data is partitioned and stored to improve the subsequent processing efficiency and ensure the reliability and availability of the data, finally forming a distributed task big data storage set.
[0108] Step S22: Perform big data cleaning on the distributed task big data storage set to remove noise, duplicate values, and outliers in the data set, obtaining a distributed task big data cleaning set.
[0109] In an embodiment of the present invention, by performing data cleaning on the distributed task big data storage set, first, Apache Spark is used as the data processing engine for parallel computing. In the data cleaning stage, de-duplication algorithms such as hash de-duplication and timestamp-based de-duplication are adopted to clean duplicate task data. Then, statistical methods such as Z-Score or IQR (Interquartile Range) are used for outlier detection and processing to identify abnormal data in task execution time and resource usage, and corresponding corrections or deletions are made. For noise data, a multiple data filtering mechanism can be constructed to eliminate records that do not conform to the expected pattern. For example, the K-Means clustering method is used to identify noise points, and the data is smoothed by the weighted average method. After the cleaning is completed, a distributed task big data cleaning set is finally obtained.
[0110] Step S23: Perform data normalization on the distributed task big data cleaning set to unify the task resources corresponding to different magnitudes to the same scale, obtaining a distributed task big data normalization set.
[0111] In an embodiment of the present invention, after data cleaning of the distributed task big data cleaning set, data normalization is performed. The goal is to convert resource items with large magnitude differences in different task data (such as CPU usage, memory usage, disk I / O, network bandwidth, etc.) to the same scale. The Min-Max Normalization method is used to process all task resources, and the resource metrics of each task are scaled between 0 and 1 to eliminate the influence of different resource usage metric dimensions, ensuring that various metrics are compared on the same scale. Through the standardized data, the bias caused by different data units can be eliminated, making each task have a relatively fair weight in the resource scheduling process. In addition, the Z-Score standardization method can also be used to perform standardization processing with a mean of 0 and a variance of 1 on the data to adapt to different machine learning and data mining algorithms, and finally obtain the distributed task big data normalization set.
[0112] Step S24: Perform big data feature mining and analysis on the distributed task big data normalization set to analyze the distributed task description using a text classification method to determine the category to which the task belongs, and use data mining techniques to extract the resource demand pattern features corresponding to the distributed tasks to obtain the distributed task big data feature set;
[0113] In an embodiment of the present invention, after data normalization, feature mining and analysis are performed on the distributed task big data normalization set. By using a text classification method to analyze the task description, with the help of natural language processing (NLP) technology, the TF-IDF (term frequency-inverse document frequency) algorithm is used to extract the keywords in the task description, and the task category is identified through a Naive Bayes classifier or a support vector machine (SVM) model to determine the category to which each task belongs. Further applying data mining techniques, the resource demand pattern features corresponding to each task are extracted from the cleaned task execution data. Through clustering analysis algorithms such as K-Means or DBSCAN, combined with the resource consumption behavior during the task execution process, the demand patterns of tasks for computing resources are automatically identified. These pattern features will help understand the resource demand laws of tasks, thereby providing a basis for subsequent scheduling and resource allocation, and finally obtaining the distributed task big data feature set, including key information such as the category of tasks and resource demand characteristics.
[0114] Step S25: Perform task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set to generate a task scheduling cluster group corresponding to similar resource demands.
[0115] In the embodiments of the present invention, based on the obtained distributed task big data feature set, task scheduling clustering analysis is further performed on task computing nodes. First, the K-Means algorithm is used to cluster tasks. The goal of clustering is to classify tasks with similar resource demand patterns into one cluster. The tasks in each cluster have high similarity in computing resource requirements during execution and can be scheduled on the same or similar computing nodes to avoid resource waste and excessive competition. By calculating the resource demand similarity between tasks, the clustering algorithm can automatically generate multiple task scheduling cluster groups. The tasks in each cluster group can share resources, thereby improving the resource utilization rate of the system. In addition, clustering analysis can also help identify the dependency relationships between tasks and provide strong support for optimizing the execution order and scheduling strategy of tasks. Through task scheduling clustering analysis, task scheduling cluster groups with similar resource requirements are finally generated to promote the efficient scheduling and execution of tasks.
[0116] Further, step S3 includes the following steps:
[0117] Step S31: Calculate the resource usage fluctuation of the task computing nodes corresponding to the distributed big data processing scenario based on the actual resource usage of tasks to obtain the task node resource usage volatility;
[0118] In the embodiments of the present invention, by collecting indicators such as the CPU usage rate, memory consumption, storage, and network bandwidth usage of task nodes within a certain period of time, a resource usage model is established based on historical data. This model calculates the resource usage fluctuation of computing nodes at each moment and quantifies the volatility through statistical methods such as standard deviation and coefficient of variation. By real-time monitoring the resource consumption of each computing node, the resource usage volatility of each node is obtained, reflecting the amplitude of the resource consumption fluctuation of the node. To accurately calculate the volatility, the system needs to discretize the resource usage data of each node for each time period and further eliminate extreme fluctuations through methods such as sliding window or weighted average. Finally, the task node resource usage volatility is obtained.
[0119] Step S32: Calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the task node resource usage volatility to obtain the remaining execution calculation duration of the task;
[0120] In the embodiments of the present invention, after obtaining the resource usage volatility of the task computing node, the system combines the actual execution duration of the task and the volatility of the node to further calculate the remaining execution duration of the task. First, the system needs to obtain the current execution progress of the task and the time that has been consumed. This information is usually tracked in real time through the monitoring module in the task scheduling system. On this basis, the system estimates according to the resource volatility of the task node. If the volatility of a certain node is large, it will lead to the instability of the calculation progress, thereby prolonging the execution duration of the task. By comprehensively considering the impact of volatility, the system uses linear regression or a multivariable prediction model to correct the prediction of the remaining execution duration. In this process, factors such as the computing load of the node and network latency need to be considered, and the model is adjusted according to the characteristics of the task (such as compute-intensive or I / O-intensive), and finally the remaining execution calculation duration of the task is obtained.
[0121] Step S33: Based on the task deadline and the remaining execution calculation duration of the task, use the task scheduling emergency calculation formula to perform task emergency accounting on the corresponding task computing node to obtain the task node scheduling emergency level.
[0122] In the embodiments of the present invention, by combining the number of task computing nodes, the system resource usage, the maximum resource usage of the task computing node, the task deadline, the current time point, the remaining execution calculation duration of the task, and relevant parameters, a suitable task scheduling emergency calculation formula is formed to perform quantitative calculation of the task emergency level, so as to quantitatively calculate the remaining available time window of the task. The emergency coefficient formula can also be used: Emergency level = (Task deadline - Current time) / Remaining execution duration. In this formula, if the emergency level is close to 1, it indicates that the task is very urgent and must be processed immediately; if the emergency level is low, the task can be scheduled in a subsequent period, and finally the task node scheduling emergency level is obtained.
[0123] Further, step S32 includes the following steps:
[0124] Step S321: Calculate the corresponding task node resource usage variance and task node resource usage covariance through the task node resource usage volatility statistics.
[0125] In an embodiment of the present invention, by collecting and counting the resource usage of each task node, these resources include but are not limited to CPU utilization, memory occupancy, network bandwidth, etc. To evaluate the volatility of task node resource usage, it is necessary to statistically analyze the resource consumption data of each node within a certain period of time, calculate its volatility. The calculation of volatility is based on the standard deviation and average value of resource usage, and then the variance is obtained. Then, for the resource usage data between different task nodes, the covariance is calculated to understand the correlation between the resource usage of each task node. This process relies on time series analysis methods and uses statistical tools such as MATLAB, R language, or the pandas library in Python to complete the calculation of variance and covariance. Specifically, when implementing, first clean the dataset of each task node to remove outliers, then use the sum of squared deviations of the data to calculate the variance, and then obtain the covariance of the resource usage of the task node through the covariance formula. Finally, the variance of the resource usage of the task node and the covariance of the resource usage of the task node are obtained.
[0126] Step S322: Based on the variance of the resource usage of the task node and the covariance of the resource usage of the task node, evaluate the resource load balance of the task computing nodes corresponding to the distributed big data processing scenario, and obtain the resource load balance degree of the task node;
[0127] In an embodiment of the present invention, through the previously obtained variance and covariance of the resource usage of the task node, key data can be provided for load balance evaluation. Specifically, when operating, first judge the stability of the node resource usage according to the resource variance of the task node. A larger variance indicates that the resource usage of this node fluctuates greatly, and the load pressure may be greater; on the contrary, a smaller variance indicates that the resource usage is relatively stable. Then, use the covariance to evaluate the resource usage dependence relationship between each task node. If the covariance is high, it indicates that there is a strong correlation between the resource usage changes of the two nodes, and coordinated scheduling is required. The load balance degree can be quantified by comprehensively considering the variance and covariance between nodes. Generally speaking, the goal of the load balance degree is to make the resource usage of each task node more uniform by adjusting the task allocation. When evaluating, use the weighted average method or the linear programming method to comprehensively evaluate each node, so as to obtain the resource load balance degree. This balance degree measures the uniformity of the load of each computing node in the system. Finally, the resource load balance degree of the task node is obtained.
[0128] Step S323: Based on the resource load balance degree of the task node, predict the remaining resource capacity of the corresponding task computing node, and obtain the remaining resource capacity of the task node;
[0129] In an embodiment of the present invention, by evaluating the remaining resource capacity of each node according to the previously obtained task node resource load balancing degree, a node with a lower load balancing degree indicates that its resources are relatively overloaded. Therefore, it is necessary to predict its remaining resources through scheduling optimization. To calculate the remaining resource capacity of each task node, the current resource occupancy of the node can be predicted. Combining the historical resource usage data of the task node, a time series prediction method (such as the ARIMA model or the LSTM neural network) is used to predict the resource usage trend. The input of the prediction model is the historical resource usage data of the node, and the output is the predicted resource usage value of the node in the future time period. By comparing the difference between the total resource capacity of the node and the predicted resource demand, the remaining resource capacity of each node is obtained. For different types of tasks (such as compute-intensive tasks or IO-intensive tasks), different prediction algorithms can be used to ensure the accuracy of the prediction, and finally the remaining resource capacity of the task node is obtained.
[0130] Step S324: Estimate the remaining execution duration of the corresponding task computing node based on the task execution duration and the remaining resource capacity of the task node, so as to obtain the remaining execution computing duration of the task.
[0131] In an embodiment of the present invention, the estimation of the remaining execution duration needs to be performed based on the task execution duration and the remaining resource capacity of the task node. Specifically, first, collect the execution progress information of the current task, including the completed task volume and the estimated total execution duration. Combining the previously obtained remaining resource capacity of the task node, use the relationship model between task execution time and resource consumption to estimate the execution duration. Usually, by analyzing the resource consumption and progress of the task when running on different nodes, a resource-time mapping model can be established. This model can estimate the remaining execution duration based on the remaining resources of the node. By comparing the current progress of the task with the expected completion progress and combining the resource consumption situation of the current node, the remaining execution duration can be predicted. To improve the estimation accuracy, regression analysis, machine learning algorithms (such as decision trees or support vector machines) can be used to train and optimize the mapping relationship between task execution duration and resource capacity. This method considers task types, resource usage patterns, and historical execution data, thereby improving the estimation accuracy of the remaining execution duration of the task, and finally obtaining the remaining execution computing duration of the task.
[0132] Further, the specific formula for the task scheduling urgency in step S33 is:
[0133] ;
[0134] In the formula, is the scheduling urgency of the task node, is the number of task computing nodes, is the The system resource usage of a task computing node is the maximum resource usage of the task computing node is the task deadline is the current time point is the remaining execution computing duration of the task
[0135] The present invention obtains an emergency calculation formula for task scheduling through using a specific mathematical model and verification, which is used to conduct emergency accounting for corresponding task computing nodes. By considering the resource usage fluctuations of task computing nodes, this emergency calculation formula for task scheduling realizes dynamic prediction of task execution duration, which means that task scheduling can adapt to the current resource usage situation and avoid situations such as computing node overload or uneven resource allocation. This formula takes into account the resource usage of task nodes and the maximum resource usage, and accurately reflects the changes in the system resource requirements of task nodes at different time periods, thereby providing a real-time reference basis for task scheduling decisions. The fluctuations in the resource usage (such as CPU, memory, storage, etc.) of computing nodes during the execution process will affect the remaining execution duration of tasks. Secondly, the part combines the remaining time of the task with the deadline, reflecting the urgency of task completion. In this way, the scheduling system can better evaluate the matching degree between the remaining duration of the current task and the deadline, and thus reasonably arrange the task allocation of computing nodes to ensure that tasks are completed on time. To sum up, this formula fully considers the emergency degree of task node scheduling , the number of task computing nodes , the th system resource usage of the task computing node , the maximum resource usage of the task computing node , the task deadline , the current time point , the remaining execution computing duration of the task , according to the emergency degree of task node scheduling and the mutual correlation relationship between the above parameters constitutes a functional relationship , this formula can realize the emergency accounting process of corresponding task computing nodes, thereby improving the accuracy and applicability of the emergency calculation formula for task scheduling
[0136] Furthermore, step S4 includes the following steps:
[0137] Step S41: Determine the scheduling priority of corresponding task computing nodes according to the emergency degree of task node scheduling to obtain the task node scheduling priority
[0138] In an embodiment of the present invention, a numerical emergency level is assigned to each task node according to the scheduling urgency of the task node obtained through previous quantization calculations. Task nodes with a high scheduling urgency (such as tasks approaching the deadline or requiring a large amount of computing resources) will be assigned a higher priority. Then, the scheduling system sorts the task nodes according to their urgency levels, and through algorithms such as weighted sorting or priority queue method, the scheduling priorities of each task node are obtained. During this process, if the urgency levels of multiple task nodes are the same, the sorting is further refined according to the computational complexity or dependency relationship of the task nodes to ensure the efficiency and accuracy of task scheduling, and finally the scheduling priorities of the task nodes are obtained.
[0139] Step S42: Based on the scheduling priorities of the task nodes, arrange the execution order within the task scheduling cluster groups corresponding to similar resource requirements to obtain the in-group scheduling arrangement sequences of the task nodes corresponding to each task scheduling cluster group;
[0140] In an embodiment of the present invention, by analyzing the resource requirements of the task nodes, the task nodes with similar resource requirements are identified and assigned to different scheduling cluster groups. The task nodes within each task scheduling cluster group have similar resource requirement characteristics, such as the number of CPU cores, memory size, storage capacity, etc. When arranging the execution order of these task nodes within the same cluster group, the scheduling priority of the task nodes must be considered first. For the task nodes within the same cluster group, the scheduling system sorts them according to the previously determined task node priorities, and preferentially schedules the task nodes with a high urgency level. If there are multiple task nodes within the cluster group with the same priority, the scheduling system further uses other scheduling strategies, such as first-come-first-served (FCFS) or shortest job first (SJF), etc., to finely arrange the execution order of the task nodes to ensure the efficient use of resources and the balance of task execution, and finally obtain the in-group scheduling arrangement sequences of the task nodes corresponding to each task scheduling cluster group.
[0141] Step S43: Optimize the inter-group task scheduling according to the in-group scheduling arrangement sequences of the task nodes corresponding to each task scheduling cluster group, prioritize the inter-group task scheduling and execute the task node scheduling arrangement corresponding to the task scheduling cluster group to generate a task dynamic optimization scheduling plan.
[0142] In an embodiment of the present invention, by obtaining the in-group scheduling arrangement sequences of the task node groups of each task scheduling cluster group, the scheduling system performs inter-group task scheduling optimization based on these in-group sorting results. In the inter-group optimization, the scheduling system comprehensively schedules the task nodes of different cluster groups to ensure that when multiple tasks are executed concurrently, the task scheduling order between different cluster groups can balance the use of system resources. The cluster group with a higher scheduling priority will be preferentially supported in resource allocation, while the cluster group with a lower priority is scheduled according to the idle resources. When the task nodes of a certain cluster group are executed, the task nodes of the next cluster group are immediately scheduled according to the dynamic resource availability to avoid resource waste. The scheduling system generates a dynamic optimization scheduling plan based on the above optimization results. This plan can achieve optimal resource allocation, improve task execution efficiency, and ensure that high-priority tasks can be completed on time, and finally generate a task dynamic optimization scheduling plan.
[0143] Further, the present invention also provides a task scheduling optimization analysis system based on big data for performing the task scheduling optimization analysis method based on big data as described above. The task scheduling optimization analysis system based on big data includes:
[0144] A distributed task big data collection module, which is used to collect task-related data in real time in the corresponding distributed big data processing scenarios among servers, storage devices, and clients, including task submission time, task deadline, task type, and task resource requirements; calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenarios based on the task resource requirements to obtain task system resource usage data; measure the task status of the task computing nodes corresponding to the distributed big data processing scenarios to obtain task execution status data, including task execution duration, actual task resource usage, and task completion status;
[0145] A task resource requirement scheduling clustering module, which is used to perform big data feature mining and analysis on task-related data, task system resource usage data, and task execution status data by using the Hadoop distributed framework to obtain a distributed task big data feature set; perform task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set to generate task scheduling cluster groups corresponding to similar resource requirements;
[0146] A task node scheduling emergency evaluation module, which is used to calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenarios based on the task execution duration and the actual task resource usage to obtain the remaining execution calculation duration of the task; perform task emergency accounting on the corresponding task computing nodes based on the task deadline and the remaining execution calculation duration of the task to obtain the task node scheduling emergency level;
[0147] The task scheduling optimization analysis module is used to determine the scheduling priority of the corresponding task computing nodes according to the urgency of task node scheduling, so as to obtain the task node scheduling priority; based on the task node scheduling priority, perform task scheduling optimization analysis on the task scheduling cluster groups corresponding to similar resource requirements, so as to generate a task dynamic optimization scheduling plan.
[0148] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes that fall within the meaning and scope of the equivalent elements of the application documents within the present invention.
[0149] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A task scheduling optimization analysis method based on big data, characterized in that It includes the following steps: Step S1: Real-time collect task-related data in the corresponding distributed big data processing scenario among the server, storage device, and client, including task submission time, task deadline, task type, and task resource requirements; Calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenario based on the task resource requirements to obtain task system resource usage data; measure the task status of the task computing nodes corresponding to the distributed big data processing scenario to obtain task execution status data, including task execution duration, actual task resource usage, and task completion status; Step S2: Use the Hadoop distributed framework to perform big data feature mining and analysis on the task-related data, task system resource usage data, and task execution status data to obtain a distributed task big data feature set; perform task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set to generate a task scheduling cluster group corresponding to similar resource requirements; Step S3: Calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the actual task resource usage to obtain the remaining execution calculation duration of the task; Perform task urgency accounting on the corresponding task computing nodes based on the task deadline and the remaining execution calculation duration of the task to obtain the scheduling urgency degree of the task nodes; among them, Step S3 includes the following steps: Step S31: Calculate the resource usage fluctuation of the task computing nodes corresponding to the distributed big data processing scenario based on the actual task resource usage to obtain the resource usage volatility of the task nodes; Step S32: Calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the resource usage volatility of the task nodes to obtain the remaining execution calculation duration of the task; among them, Step S32 includes the following steps: Step S321: Statistically calculate the resource usage variance and covariance of the corresponding task nodes through the resource usage volatility of the task nodes; Step S322: Perform resource load balancing evaluation on the task computing nodes corresponding to the distributed big data processing scenario based on the resource usage variance and covariance of the task nodes to obtain the resource load balancing degree of the task nodes; Step S323: Predict the remaining resource capacity of the corresponding task computing nodes based on the resource load balancing degree of the task nodes to obtain the remaining resource capacity of the task nodes; Step S324: Estimate the remaining execution duration of the corresponding task computing nodes based on the task execution duration and the remaining resource capacity of the task nodes to obtain the remaining execution calculation duration of the task; Step S33: Perform task urgency accounting on the corresponding task computing nodes based on the task deadline and the remaining execution calculation duration using the task scheduling urgency calculation formula to obtain the scheduling urgency degree of the task nodes; among them, the specific task scheduling urgency calculation formula is: ; In the formula, is the urgency of task node scheduling, is the number of task computing nodes, is the system resource usage of the th task computing node, is the task deadline, is the current time point, is the remaining execution computing duration of the task; Step S4: Determine the scheduling priorities of the corresponding task computing nodes according to the urgency of the task node scheduling, so as to obtain the task node scheduling priorities; based on the task node scheduling priorities, conduct task scheduling optimization analysis on the task scheduling cluster groups corresponding to similar resource requirements, so as to generate a task dynamic optimization scheduling plan.
2. The method for optimizing and analyzing task scheduling based on big data according to claim 1, wherein Step S1 includes the following steps: Step S11: Configure and deploy each distributed network device among the server, storage device, and client to ensure a stable network connection among the devices in the corresponding task access test environment between the server, storage device, and client, and generate a corresponding distributed big data processing scenario; Step S12: Real-time collect the task-related data of the corresponding distributed access operation tasks in the distributed big data processing scenario among the server, storage device, and client, including the task submission time, task deadline, task type, and task resource requirements; Step S13: Calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenario based on the task resource requirements to obtain the task system resource usage data; Step S14: Deploy data acquisition interfaces at the system kernel, task management module, and resource monitoring tool locations of the task computing nodes corresponding to the distributed big data processing scenario, and use the data acquisition interfaces to measure the task status of the task computing nodes corresponding to the distributed big data processing scenario to obtain the task execution status data, including the task execution duration, the actual task resource usage, and the task completion status.
3. The method for optimizing and analyzing task scheduling based on big data according to claim 2, wherein, The task resource requirements described in Step S12 specifically include the resource requirements corresponding to the CPU, memory, storage, and network bandwidth.
4. The method for optimizing and analyzing task scheduling based on big data according to claim 3, wherein, Step S13 includes the following steps: Step S131: Calculate the CPU usage rate of the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the CPU to obtain the CPU usage rate of each computing node; Step S132: Conduct memory idle accounting for the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the memory to obtain the memory idle amount of each computing node; Step S133: Conduct storage read / write statistical analysis on the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the storage to obtain the storage read / write speed of each computing node; Step S134: Conduct bandwidth occupancy accounting for the task computing nodes corresponding to the distributed big data processing scenario based on the resource requirements corresponding to the network bandwidth to obtain the network bandwidth occupancy of each computing node; Step S135: Combine the CPU usage rate, memory idle amount, storage read / write speed, and network bandwidth occupancy of each computing node for resource usage to obtain the task system resource usage data.
5. The method for optimizing and analyzing task scheduling based on big data according to claim 1, wherein Step S2 includes the following steps: Step S21: Use the Hadoop distributed framework to perform big data multi-source heterogeneous storage on the task-related data, task system resource usage data, and task execution status data to obtain a distributed task big data storage set; Step S22: Perform big data cleaning on the distributed task big data storage set to remove noise, duplicate values, and outliers in the data set, obtaining a distributed task big data cleaning set; Step S23: Perform data normalization on the distributed task big data cleaning set to unify task resources corresponding to different magnitudes to the same scale, obtaining a distributed task big data normalization set; Step S24: Perform big data feature mining and analysis on the distributed task big data normalization set to analyze the distributed task description using text classification methods to determine the category to which the task belongs, and use data mining techniques to extract the resource requirement pattern features corresponding to the distributed task, obtaining a distributed task big data feature set; Step S25: Perform task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set to generate a task scheduling cluster group corresponding to similar resource requirements.
6. The method for optimizing and analyzing task scheduling based on big data according to claim 1, wherein Step S4 includes the following steps: Step S41: Determine the scheduling priority of the corresponding task computing nodes according to the scheduling urgency of the task nodes, obtaining the task node scheduling priority; Step S42: Arrange the intra-group order of scheduling for the task scheduling cluster group corresponding to similar resource requirements based on the task node scheduling priority, obtaining the intra-group scheduling arrangement sequence of task nodes corresponding to each task scheduling cluster group; Step S43: Perform inter-group task scheduling optimization according to the intra-group scheduling arrangement sequence of task nodes corresponding to each task scheduling cluster group, with inter-group task scheduling taking precedence and executing the task node scheduling arrangement corresponding to the task scheduling cluster group, to generate a task dynamic optimization scheduling plan.
7. A task scheduling optimization analysis system based on big data, characterized in that, A task scheduling optimization analysis system based on big data for executing the method as described in claim 1, the task scheduling optimization analysis system based on big data includes: A distributed task big data collection module, configured to collect task-related data in real time in the corresponding distributed big data processing scenarios among servers, storage devices, and clients, including task submission time, task deadline, task type, and task resource requirements; calculate the system resource usage of the task computing nodes corresponding to the distributed big data processing scenarios based on the task resource requirements, obtaining task system resource usage data; measure the task status of the task computing nodes corresponding to the distributed big data processing scenarios, thereby obtaining task execution status data, including task execution duration, actual task resource usage, and task completion status; A task resource requirement scheduling clustering module, configured to perform big data feature mining and analysis on the task-related data, task system resource usage data, and task execution status data using the Hadoop distributed framework to obtain a distributed task big data feature set; perform task scheduling clustering analysis on the corresponding task computing nodes based on the distributed task big data feature set, thereby generating a task scheduling cluster group corresponding to similar resource requirements; The task node scheduling emergency evaluation module is used to calculate the remaining execution duration of the task computing nodes corresponding to the distributed big data processing scenario based on the task execution duration and the actual usage of task resources, so as to obtain the remaining execution calculation duration of the task; and conduct task emergency accounting for the corresponding task computing nodes based on the task deadline and the remaining execution calculation duration of the task, thereby obtaining the emergency degree of task node scheduling. The task scheduling optimization analysis module is used to determine the scheduling priority of the corresponding task computing nodes according to the emergency degree of task node scheduling, so as to obtain the task node scheduling priority; and conduct task scheduling optimization analysis on the task scheduling cluster groups with similar resource requirements based on the task node scheduling priority, so as to generate a task dynamic optimization scheduling scheme.
Citation Information
Patent Citations
Resource service system and resource distribution method thereof
CN104168318A
Intelligent task scheduling system and method based on machine learning
CN116909712A