A server resource scheduling system integrating AI and edge computing
Through the server resource scheduling system integrating AI and edge computing, real-time monitoring and analysis of server status, dynamic task processing demand prediction and resource allocation, the problems of resource waste and load imbalance in traditional scheduling methods are solved, and efficient load balancing and adaptive optimization are achieved.
Patent Information
- Application Number
- CN202510221796.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-02-27
AI Technical Summary
When facing rapidly changing workloads and dynamic environments, traditional server resource scheduling methods are prone to waste of resources or unbalanced load allocation, and cannot meet the needs of efficient service response and user experience.
A server resource scheduling system that integrates AI and edge computing is adopted, including a computing node module, a resource call module, a task traversal module, a demand prediction module and a resource allocation module. By monitoring the server status in real time, analyzing resource call behavior and task status, dynamic task processing demand prediction and resource allocation are performed, and load scheduling is optimized.
It improves the stability and efficiency of the system, optimizes resource utilization, realizes load balancing and adaptability, and improves the overall performance and response speed of the system.
Smart Images

Figure CN119718682B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of server resource scheduling, and particularly to a server resource scheduling system integrating AI and edge computing. Background Art
[0002] With the rapid development of cloud computing and big data technologies, the scale and complexity of data centers and enterprise IT infrastructures have been continuously increasing. As the core component supporting modern information applications, servers undertake a large number of computing and storage tasks. The effective scheduling and management of server resources have become key factors affecting system performance and service quality. In practical applications, server resource scheduling faces many challenges. Firstly, with the wide application of large-scale virtualization and containerization technologies, resource requirements show highly dynamic characteristics. Secondly, the requirements for service quality and user experience are increasing day by day, and the system needs to adjust resource allocation according to the real-time load situation to ensure efficient service response.
[0003] Traditional server resource scheduling methods usually rely on centralized scheduling systems, mainly based on preset rules or static policies to allocate resources. This approach seems inadequate in the face of rapidly changing workloads and dynamic environments, and is prone to resource waste or uneven resource load distribution. Therefore, to address the limitations of traditional server resource scheduling methods, a more intelligent server resource scheduling system is needed. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a server resource scheduling system integrating AI and edge computing to solve at least one of the above technical problems.
[0005] To achieve the above object, the present invention provides a server resource integrating AI and edge computing. The server resource scheduling system integrating AI and edge computing includes a computing node module, a resource invocation module, a task traversal module, a demand prediction module, a resource allocation module, and a load scheduling module:
[0006] The computing node module is used to monitor the real-time operating status of the server computing nodes, and perform distributed node architecture fitting to construct a distributed server node network;
[0007] The resource invocation module is used to obtain the server operation logs; perform server resource invocation behavior analysis on the server operation logs, and conduct resource invocation time series feature evolution to generate server resource invocation time series features;
[0008] The task traversal module is used to traverse and identify unprocessed tasks in the server operation logs, and mine task status features to construct an unprocessed task status mapping sequence;
[0009] A demand prediction module, configured to perform dynamic task processing demand prediction on the unprocessed task status mapping sequence according to the timing characteristics of server resource calls, so as to obtain dynamic task processing demand prediction data;
[0010] A resource allocation module, configured to perform callable resource matching on the dynamic task processing demand prediction data according to the distributed server node network, and perform dynamic computing resource allocation, so as to obtain a dynamic computing resource allocation strategy;
[0011] A load scheduling module, configured to perform task resource allocation evolution based on the dynamic computing resource allocation strategy, and then perform adaptive load scheduling optimization to construct a load scheduling optimization model.
[0012] By monitoring the running status of the server computing nodes in real time, the present invention can help discover potential problems and ensure system stability. Constructing a distributed server node network helps improve the scalability and fault tolerance of the system, optimize resource utilization. Analyzing the resource call behavior in the server running log can reveal the system resource utilization situation, providing a basis for performance optimization. Generating resource call timing characteristics helps understand the evolution of the resource call pattern, providing data support for subsequent task processing demand prediction. Identifying unprocessed tasks and mining task status characteristics can help optimize task scheduling and resource allocation, improving system efficiency. Constructing an unprocessed task status mapping sequence helps understand the task processing flow, providing a basis for subsequent demand prediction. Predicting the dynamic task processing demand according to the resource call timing characteristics can allocate resources in advance, optimizing the system response speed. Obtaining the dynamic task processing demand prediction data helps reasonably plan the resource allocation strategy, improving the overall system efficiency. Performing resource matching and dynamic computing resource allocation according to the distributed server node network can achieve reasonable resource utilization, improving the overall system performance. Obtaining the dynamic computing resource allocation strategy helps adjust the resource allocation according to real-time demands, optimizing the system load balance. Performing task resource allocation evolution and load scheduling optimization based on the dynamic computing resource allocation strategy can achieve system load balance, improving system stability and performance. Constructing a load scheduling optimization model helps continuously improve the system load scheduling strategy, improving the system adaptability and efficiency.
[0013] Preferably, the computing node module is configured to monitor the real-time running status of the server computing nodes, and perform distributed node architecture fitting to construct a distributed server node network, specifically:
[0014] Identify all server computing nodes;
[0015] Identify the terminal device types of the server computing nodes, and mark the central device computing nodes and edge device computing nodes;
[0016] Monitor the real-time operating status of the central device computing node and the edge device computing node to obtain the central node operating status parameters and the edge node operating status parameters;
[0017] Conduct node load balancing feature analysis on the central node operating status parameters and the edge node operating status parameters to obtain the central node load balancing features and the edge node load balancing features;
[0018] Perform distributed node architecture fitting on the terminal device computing node and the edge device computing node according to the central node load balancing features and the edge node load balancing features, and construct a distributed server node network.
[0019] The present invention can establish an overall view by identifying all server computing nodes, which helps to comprehensively monitor and manage server resources. Marking the central device computing node and the edge device computing node helps to distinguish the functions of different nodes and locate them, providing a basis for subsequent load balancing. Real-time monitoring of the operating status of the central device and the edge device computing node can timely detect problems and take measures to maintain system stability. Obtaining the operating status parameters of the central node and the edge node helps to quantify the operating conditions of the nodes, providing data support for load balancing feature analysis. Analyzing the load balancing features of the central node and the edge node can understand the load distribution between nodes, providing a basis for resource scheduling. Extracting the load balancing features helps to evaluate the performance of the nodes, providing a reference for subsequent architecture fitting and resource allocation. Conducting distributed node architecture fitting according to the load balancing features of the central node and the edge node can achieve reasonable resource allocation and load balancing, improving the overall performance of the system. Constructing a distributed server node network helps to optimize the data processing process, improve the response speed and efficiency of the system, and thus achieve the optimization of resource load scheduling.
[0020] Preferably, the specific steps of conducting node load balancing feature analysis on the central node operating status parameters and the edge node operating status parameters to obtain the central node load balancing features and the edge node load balancing features are as follows:
[0021] Calculate the CPU utilization rate of each node one by one for the central node operating status parameters and the edge node operating status parameters to obtain the CPU utilization rate of each node;
[0022] Conduct utilization rate fluctuation analysis on the CPU utilization rate of each node to generate node utilization rate fluctuation features;
[0023] Conduct real-time memory occupancy calculation on the central node operating status parameters and the edge node operating status parameters to obtain the memory occupancy value of each node;
[0024] Analyze the operating state parameters of the central node and the network throughput of the edge nodes, and extract the network throughput of each node.
[0025] Conduct an analysis of the current load characteristics of the nodes based on the node utilization fluctuation characteristics, the memory occupancy value of each node, and the network throughput of each node, so as to obtain the current load characteristics of each node.
[0026] Identify the node types based on the current load characteristics of each node, and extract the load characteristics of each central node and each edge node.
[0027] Identify the load differences between each central node's load characteristics and each edge node's load characteristics, and extract the load difference data between the central node and the edge node.
[0028] Conduct an analysis of the node load balancing characteristics based on the load difference data between the central node and the edge node, so as to obtain the central node load balancing characteristics and the edge node load balancing characteristics.
[0029] By calculating the CPU utilization rate of each node, the present invention helps to understand the consumption of computing resources of the nodes, provides an important indicator for load balancing. Analyzing the fluctuation characteristics of the CPU utilization rate can reveal the stability and volatility of the node load, providing a reference for load scheduling optimization. Calculating the memory occupancy value of each node in real time helps to monitor the usage of memory resources of the nodes, providing a basis for resource allocation. Analyzing the network throughput of the nodes can evaluate the communication efficiency between the nodes, providing data support for network topology optimization and load balancing. Considering comprehensively the CPU utilization rate, memory occupancy, and network throughput characteristics of the nodes can comprehensively understand the current load status of the nodes, providing a basis for load balancing decisions. Identifying the node types can distinguish between central nodes and edge nodes, helping to formulate targeted load balancing strategies and optimize the resource scheduling effect. Extracting the load difference data between the central node and the edge node can help to discover the load imbalance between the nodes, providing clues for further optimizing the load distribution. Conducting an analysis of the load balancing characteristics based on the load difference data can adjust the resource allocation strategy, achieve load balancing between the nodes, and improve the overall system performance.
[0030] Preferably, the resource invocation module is used to obtain the server operation logs; conduct an analysis of the server resource invocation behavior on the server operation logs, and perform the evolution of the resource invocation time series characteristics to generate the server resource invocation time series characteristics, specifically used for:
[0031] Obtain the server operation logs; conduct an analysis of the server resource invocation behavior on the server operation logs, and extract the server resource invocation behavior data.
[0032] Perform a time-series window partitioning on the server resource call behavior data to obtain the resource call behavior data of multiple time windows;
[0033] Perform a scheduling count on the resource call behavior data of multiple time windows to obtain the scheduling count for each time window;
[0034] Calculate the scheduling frequency based on the scheduling count of each time window to generate the resource call frequency of each window;
[0035] Perform an evolution of the resource call time-series characteristics for the resource call frequency of each window to generate the server resource call time-series characteristics.
[0036] The present invention can record various operations and events during the server operation by obtaining the server operation log, providing a data basis for subsequent analysis. Extracting the server resource call behavior data helps to understand the specific usage of server resources, providing a basis for resource scheduling optimization. Partitioning the resource call behavior data by time-series window can serialize the resource call behavior in time, helping to discover the rules and trends of resource calls. Processing the resource call behavior data of multiple time windows can provide a data basis for subsequent frequency statistics and feature analysis. Counting the scheduling count of each time window can understand the frequency of resource scheduling, providing a basis for formulating resource scheduling strategies. Calculating the scheduling frequency can quantify the frequency change of resource calls, helping to discover the regularity and change trend of resource scheduling. Through the evolution of resource call time-series characteristics, the evolution rule of resource call behavior can be revealed, providing a reference for future resource scheduling optimization. Generating the server resource call time-series characteristics helps to comprehensively understand the scheduling situation of server resources, providing data support for load scheduling optimization.
[0037] Preferably, a task traversal module is used to identify unprocessed tasks in the server operation log and mine task status characteristics, constructing an unprocessed task status mapping sequence, specifically used for:
[0038] Identify unprocessed tasks in the server operation log and extract multiple unprocessed tasks of the server;
[0039] Calculate the remaining time for task processing of multiple unprocessed tasks of the server and extract the unprocessed task timestamps;
[0040] Calculate the task resource occupancy of the unprocessed tasks of the server and extract the task resource occupancy data;
[0041] Mine the task status characteristics from the unprocessed task timestamps and task resource occupancy data to generate the status characteristics of each unprocessed task;
[0042] Perform task sequence encoding on multiple unprocessed tasks of the server to obtain an unprocessed task encoding sequence;
[0043] According to each unprocessed task status feature, perform corresponding task status mapping on the unprocessed task coding sequence to construct an unprocessed task status mapping sequence.
[0044] The present invention can help understand the workload to be processed in the system by identifying the unprocessed tasks of the server, providing basic data for resource scheduling and optimization. Extracting the unprocessed tasks of multiple servers helps to comprehensively understand the task distribution in the system, providing data support for subsequent processing. Calculating the remaining processing time of tasks can evaluate the timeliness of task completion, which is helpful for reasonably arranging resources and task scheduling. Extracting the timestamps of unprocessed tasks can record the initiation time of tasks, providing a time clue for analyzing task status changes. Calculating the resource occupancy of tasks helps to evaluate the consumption of system resources by tasks, providing a basis for resource scheduling optimization. Through mining task status features, the state evolution law of tasks can be deeply understood, providing a reference basis for task scheduling and optimization. Performing sequence coding on unprocessed tasks can convert task status into a data form that can be processed by a computer, facilitating subsequent analysis and processing. Constructing an unprocessed task status mapping sequence can map task status features into corresponding coding sequences, which is helpful for visualizing and analyzing task status.
[0045] Preferably, a demand prediction module is used to perform dynamic task processing demand prediction on the unprocessed task status mapping sequence according to the server resource call timing characteristics to obtain dynamic task processing demand prediction data, specifically used for:
[0046] Perform multi-point resource call prediction on the server resource call timing characteristics to generate resource call prediction data at multiple time points;
[0047] Perform resource call timing change analysis on the resource call prediction data at multiple time points to obtain resource call timing change data;
[0048] Perform per-task resource requirement analysis on the unprocessed task status mapping sequence to obtain the resource requirement data for each task;
[0049] Based on the resource call timing change data, perform dynamic task processing demand prediction on the resource requirement data for each task to obtain dynamic task processing demand prediction data.
[0050] Through resource call prediction, the present invention can help the system make resource allocation decisions in advance, optimize resource utilization. Analyzing the temporal changes in resource calls helps to understand the trends and patterns of resource calls, providing a basis for subsequent resource scheduling. Analyzing the resource requirements of each task one by one can deeply understand the resource requirements of each task, providing guidance for task scheduling and resource allocation. Predicting the dynamic task processing requirements based on the temporal change data of resource calls can predict the changes in resource requirements of tasks according to the resource call situation, helping the system make timely resource scheduling and allocation.
[0051] Preferably, the specific steps of performing multi-point resource call prediction on the temporal characteristics of server resource calls to generate resource call prediction data at multiple time points are as follows:
[0052] Perform multi-point segmentation processing on the temporal characteristics of server resource calls to generate resource call temporal data at multiple time points;
[0053] Extract the timestamps of the resource call temporal data at multiple time points;
[0054] Perform multi-point temporal fitting according to the timestamps to obtain a multi-point resource call temporal curve;
[0055] Perform deep call feature learning on the multi-point resource call temporal curve to generate server resource call rules;
[0056] Perform short-term resource call simulation on the temporal characteristics of server resource calls based on the server resource call rules to generate resource call simulation data;
[0057] Define the prediction time node;
[0058] Perform resource call prediction on the resource call simulation data according to the prediction time node to generate resource call prediction data at multiple time points.
[0059] Through the processing of multi-point resource call temporal data, the present invention can divide resource call data by time point, which helps to analyze the resource usage in different time periods. Through temporal fitting, the trends and patterns of resource calls at multiple time points can be found, providing basic data for subsequent deep learning. Through deep call feature learning, hidden patterns and rules in resource call data can be mined, providing a deeper understanding for resource scheduling optimization. Generating server resource call rules helps the system understand the characteristics and change trends of resource calls, providing guidance for resource allocation and scheduling strategies. Performing short-term resource call simulation can simulate the actual situation of resource calls and evaluate the effectiveness of resource allocation strategies. Through resource call prediction, future resource requirements can be predicted, helping the system make reasonable resource allocation and scheduling decisions.
[0060] Preferably, it is used to perform callable resource matching on the dynamic task processing demand prediction data according to the distributed server node network, and perform dynamic computing resource allocation, so as to obtain a dynamic computing resource allocation strategy, specifically used for:
[0061] Identify callable computing nodes in the distributed server node network to mark callable computing nodes;
[0062] Calculate the idle resources of callable computing nodes, and extract the idle resource data of each computing node;
[0063] Perform callable resource matching on the dynamic task processing demand prediction data based on the idle resource data of each computing node to obtain task processing resource matching data;
[0064] Perform dynamic computing resource allocation on the distributed server node network based on the task processing resource matching data, so as to obtain a dynamic computing resource allocation strategy.
[0065] By identifying callable computing nodes, the present invention can determine which nodes can be used for task processing, improve resource utilization rate. Extracting the idle resource data of each computing node helps to understand the available resources in the system and provides a basis for task allocation. Performing dynamic task processing demand prediction based on the idle resource data of computing nodes can find the best computing nodes for task allocation according to the actual resource situation. Generating task processing resource matching data can ensure that tasks receive sufficient resource support, improve task processing efficiency and system performance. Performing dynamic computing resource allocation based on task processing resource matching data can reasonably allocate resources according to actual needs, improve the overall performance of the system. The dynamic computing resource allocation strategy can flexibly adjust resource allocation according to the current task situation and the resource status of computing nodes, optimize the server resource load scheduling, and improve the stability and efficiency of the system.
[0066] Preferably, the load scheduling module is used to perform task resource allocation evolution based on the dynamic computing resource allocation strategy, and then perform adaptive load scheduling optimization to construct a load scheduling optimization model, specifically used for:
[0067] Perform task resource allocation evolution on the distributed server node network based on the dynamic computing resource allocation strategy to generate task resource allocation evolution data; perform computing node load analysis on the task resource allocation simulation data to obtain the load value of each computing node;
[0068] Perform node load peak analysis on the load value of each computing node to generate node load peaks;
[0069] Perform adaptive load scheduling optimization on the dynamic computing resource allocation strategy based on the node load peaks to construct a load scheduling optimization model.
[0070] By analyzing the evolution data of task resource allocation, the present invention can understand the change trend of task allocation, provide historical data support for load scheduling. The load analysis of computing nodes helps to evaluate the load situation of each node and provides a reference for load scheduling optimization. By analyzing the node load peak, the highest point of node load can be determined, which helps the system predict the peak load period. The adaptive load scheduling optimization based on the node load peak can adjust the resource allocation strategy according to the node load situation, avoid the situation of too high or too low load, and improve the load balance and performance stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 FIG. is a schematic structural diagram of a server resource scheduling system integrating AI and edge computing according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0073] The embodiments of the present application provide a server resource scheduling system integrating AI and edge computing. The execution subjects of the server resource scheduling system integrating AI and edge computing include, but are not limited to, the following general computing nodes carrying the system: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. The data processing platform includes, but is not limited to, at least one of an audio and image management system, an information management system, and a cloud data management system.
[0074] Please refer to Figure 1 , the present invention provides a server resource scheduling system integrating AI and edge computing. The server resource scheduling system integrating AI and edge computing includes a computing node module, a resource invocation module, a task traversal module, a demand prediction module, a resource allocation module, and a load scheduling module:
[0075] The computing node module is used to monitor the real-time operating status of the server computing nodes, perform distributed node architecture fitting, and construct a distributed server node network;
[0076] The resource invocation module is used to obtain the server operation logs; perform server resource invocation behavior analysis on the server operation logs, and perform resource invocation time series feature evolution to generate server resource invocation time series features;
[0077] The task traversal module is used to traverse and identify unprocessed tasks in the server operation logs, and mine task status features to construct an unprocessed task status mapping sequence;
[0078] A demand prediction module, configured to perform dynamic task processing demand prediction on the unprocessed task status mapping sequence according to the time series characteristics of server resource calls, so as to obtain dynamic task processing demand prediction data;
[0079] A resource allocation module, configured to perform callable resource matching on the dynamic task processing demand prediction data according to the distributed server node network, and perform dynamic computing resource allocation, so as to obtain a dynamic computing resource allocation strategy;
[0080] A load scheduling module, configured to perform task resource allocation evolution based on the dynamic computing resource allocation strategy, and then perform adaptive load scheduling optimization to construct a load scheduling optimization model.
[0081] By monitoring the running status of the server computing nodes in real time, the present invention can help discover potential problems, ensure system stability. Constructing a distributed server node network helps improve the scalability and fault tolerance of the system, optimize resource utilization. Analyzing the resource call behavior in the server running logs can reveal the system resource utilization situation, providing a basis for performance optimization. Generating resource call time series characteristics helps understand the evolution of resource call patterns, providing data support for subsequent task processing demand prediction. Identifying unprocessed tasks and mining task status characteristics can help optimize task scheduling and resource allocation, improving system efficiency. Constructing an unprocessed task status mapping sequence helps understand the task processing process, providing a basis for subsequent demand prediction. Predicting the dynamic task processing demand according to the resource call time series characteristics can allocate resources in advance, optimizing the system response speed. Obtaining the dynamic task processing demand prediction data helps reasonably plan the resource allocation strategy, improving the overall system efficiency. Performing resource matching and dynamic computing resource allocation according to the distributed server node network can achieve reasonable resource utilization, improving the overall system performance. Obtaining the dynamic computing resource allocation strategy helps adjust the resource allocation according to real-time demands, optimizing the system load balance. Performing task resource allocation evolution and load scheduling optimization based on the dynamic computing resource allocation strategy can achieve system load balance, improving system stability and performance. Constructing a load scheduling optimization model helps continuously improve the system load scheduling strategy, improving the system's adaptability and efficiency.
[0082] In an embodiment of the present invention, refer to Figure 1 , which is a schematic diagram of the step flow of a server resource scheduling system integrating AI and edge computing according to the present invention. In this example, the server resource scheduling system integrating AI and edge computing includes a computing node module, a resource call module, a task traversal module, a demand prediction module, a resource allocation module, and a load scheduling module:
[0083] A computing node module, configured to perform real-time running status monitoring on the server computing nodes, and perform distributed node architecture fitting to construct a distributed server node network;
[0084] In this embodiment, key performance indicators (KPIs) to be monitored are determined, such as CPU usage, memory occupancy, network bandwidth, disk I / O, etc., and thresholds are set to trigger an alarm when resource usage exceeds a certain limit. Common monitoring tools include Prometheus, Nagios, Zabbix, and Grafana. These tools can provide real-time monitoring, data collection, and visualization functions. For example, Prometheus can collect and store monitoring data through its time series database, while Grafana can be used for data visualization. A monitoring agent (such as Node Exporter) is installed on each computing node to collect the running status data of the node. The monitoring agent is configured to ensure that it can periodically collect and send data to the monitoring server. The monitoring tool is configured to regularly collect the running status data of each node at a set time interval (such as every minute) to ensure that the data of all nodes can be effectively collected and stored in the monitoring system. The collected data is cleaned to remove duplicate values, missing values, and outliers to ensure the accuracy and integrity of the data. Necessary standardization processing is performed for different monitoring indicators for subsequent analysis. According to the characteristics of the monitoring data, a suitable model is selected for distributed node architecture fitting. Machine learning methods (such as clustering analysis, regression models) or statistical methods (such as time series analysis) can be used. For example, the k-means clustering algorithm can be used to classify nodes and identify nodes with similar performance characteristics. The processed monitoring data (such as CPU usage, memory occupancy, etc.) is used as input to train the fitting model. The performance of the model is evaluated by adjusting the model parameters and using the training set (such as historical data) to ensure that it can accurately reflect the running status of the node. Methods such as cross-validation are used to evaluate the accuracy and robustness of the model. If the model performance does not meet the requirements, the model parameters need to be optimized or a different fitting method needs to be selected. According to the fitting results, the structure of the distributed server node network is designed, including the connection relationship between nodes, data flow, and load balancing strategy. The role of each node (such as master node, slave node, load balancing node) is determined to optimize resource usage and task processing efficiency. The network architecture is deployed to ensure that all nodes can communicate effectively. Virtual networks (such as Docker network, Kubernetes service) or physical networks can be used. Network switches and routers are configured to ensure that data packets can flow smoothly between nodes. The monitoring tool is used to monitor the running status of the distributed server node network in real time to ensure its stability and reliability. The network performance data is analyzed regularly to evaluate the load situation and response time of the nodes, and the network configuration is adjusted in a timely manner to cope with changing workloads.
[0085] A resource invocation module, which is used to obtain the server operation logs; perform server resource invocation behavior analysis on the server operation logs, and conduct resource invocation time-series feature evolution to generate server resource invocation time-series features;
[0086] In this embodiment, determine the types of logs to be analyzed, such as system logs ( / var / log / syslog), application logs (such as Nginx and Apache logs), monitoring logs, etc. Ensure that the paths and formats of the log files are known, usually in text format or JSON format. Tools such as Filebeat, Fluentd, or Logstash can be used. These tools can collect logs from different sources and send them to a centralized log management system (such as Elasticsearch). Configure the log collection tool to collect log files regularly and store the collected logs in a queryable database, such as Elasticsearch, MongoDB, or a relational database (such as MySQL). Ensure that the log data can be retrieved quickly by timestamp for subsequent analysis. According to the format of the logs (such as JSON, CSV, Plain Text), use appropriate parsing methods to convert the log data into a structured format. The pandas library in Python or other data processing tools can be used for parsing. Clean the log data by removing unnecessary fields, duplicate entries, and null values to ensure the accuracy and integrity of the data. Uniformly format the timestamps to ensure that all recorded timestamps are in the same format (such as ISO 8601). Filter out the log entries related to resource calls according to the requirements, such as call records related to CPU, memory, I / O operations, etc. Use statistical analysis methods to identify the behavior patterns of server resource calls, which may include call frequency, call duration, resource consumption, etc. Cluster analysis (such as K-means) can be used to identify different call patterns. Use data visualization tools (such as Grafana, Tableau, Matplotlib) to graphically represent the resource call behavior to help quickly identify trends and anomalies. The visualized content includes the change trend of the call count, the time distribution of resource usage, etc. According to the cleaned log data, construct time series data of resource calls, which includes recording the resource usage at each time point (such as CPU usage rate, memory usage). The timestamps should be consistent to ensure that the data can be arranged in chronological order. Extract features from the time series data, including: Trend: Observe the long-term trend of resource usage (rising, falling, or stable). Seasonality: Identify the periodic changes in resource usage (such as peak and trough periods). Sudden behavior: Analyze the abnormal fluctuations and sudden situations in resource usage. Apply time series analysis models (such as ARIMA, LSTM) to perform evolutionary analysis on the time series characteristics of resource calls and predict the future resource usage trend. Record the key features and change points identified during the evolution process. Integrate the extracted time series features into a structured data set, which can be a table containing information such as timestamps, resource usage amounts, call frequencies, etc. Ensure that the integrated data can be used for subsequent analysis and modeling.Save the generated resource call timing characteristics in a suitable format (such as CSV, JSON) for subsequent use. The data can be stored in a data warehouse or data lake for further analysis.
[0087] The task traversal module is used to traverse and identify unprocessed tasks in the server operation log, mine task status characteristics, and construct an unprocessed task status mapping sequence;
[0088] In this embodiment, determine the types of logs to be analyzed, such as application logs, task scheduling logs, etc., to ensure that relevant information about unprocessed tasks can be obtained. Use tools such as Filebeat, Fluentd, or Logstash to centralize the log data to an analyzable location (such as Elasticsearch or a database). Store the collected logs in a queryable database to ensure quick access and retrieval of log records. According to the format of the logs (such as JSON, CSV, PlainText), use appropriate methods to convert the log data into a structured format for subsequent analysis. Clean the log data by removing unnecessary fields, duplicate entries, and null values to ensure the accuracy and integrity of the data. Filter out the log entries related to the task status, such as status information like "Task Created", "Task Running", "Task Failed", etc. Read and process the log data line by line to find relevant information about unprocessed tasks, which usually includes task ID, task creation time, task status, etc. Determine what kind of tasks are considered "unprocessed", for example, tasks that are not completed or not successful within a specific time. Identify unprocessed tasks through status tags (such as "Created", "Running", "Failed"). Create a data structure (such as a dictionary, list, or DataFrame) to record information such as the ID, status, and creation time of unprocessed tasks for subsequent analysis. Analyze the status changes of unprocessed tasks, extract status transition features, such as the transition from "Created" to "Running" or "Failed". Statistically analyze the frequency and duration of each status occurrence to identify patterns of status changes. Time series analysis methods can be used to observe the trend of status changes. Use tools (such as Matplotlib, Seaborn, or Grafana) to visualize the status features to help identify patterns and anomalies in status changes. Based on the status changes of unprocessed tasks, construct a status mapping sequence, where each sequence records the task ID and the timestamp of its status change, for example: [(task_id, timestamp, "Created"), (task_id, timestamp, "Running")...]. Organize the status mapping sequence into a structured format (such as a list or DataFrame) for subsequent analysis and storage.
[0089] A demand prediction module for dynamically predicting the demand for task processing of the unprocessed task status mapping sequence according to the time series characteristics of server resource calls to obtain dynamically predicted demand data for task processing;
[0090] In this embodiment, the server resource call timing characteristics (such as CPU usage rate, memory usage, etc.) are merged with the mapping sequence of unprocessed task states to form a comprehensive data set. The data set should include timestamps, resource usage characteristics, and the status information of unprocessed tasks. Deal with missing values, outliers, and duplicate data to ensure the quality of the data, ensure that the timestamp formats are consistent, and sort the data for subsequent analysis. Use feature selection methods (such as LASSO regression, importance analysis of random forests) to identify the features that have a greater impact on the prediction of task processing requirements, and ensure that the selected features can fully reflect the changes in task processing requirements. Select a suitable machine learning model according to the characteristics of the data and the prediction target. Common models include: Linear regression: suitable for simple predictions of linear relationships; Decision tree or random forest: suitable for dealing with complex non-linear relationships; Time series models: such as ARIMA or LSTM, suitable for dealing with data with time dependence. Divide the data set into a training set (for training the model) and a test set (for evaluating the model performance). Usually, 70% of the data is used for training and 30% of the data is used for testing. Use the training set to train the selected model and adjust the model parameters to improve the prediction accuracy. For example, for a random forest, parameters such as the number of trees and the maximum depth can be adjusted. Use the test set to evaluate the prediction performance of the model, and adopt metrics such as root mean square error (RMSE), mean absolute error (MAE), etc. Further verify the stability and generalization ability of the model through methods such as cross-validation. Save the predicted results of dynamic task processing requirements to a database or a file for subsequent analysis and scheduling decisions. Check the rationality and feasibility of the prediction results, analyze the accuracy, deviation degree, etc. of the prediction, and display the trend of the prediction results through visualization tools (such as Matplotlib, Seaborn) for easy understanding and decision-making. According to the predicted dynamic task processing requirements, adjust the server resource allocation strategy to ensure that the resources can meet the expected task requirements. Preventive measures can be formulated, such as allocating resources in advance, dynamically adjusting the task scheduling strategy, etc.
[0091] A resource allocation module, configured to perform callable resource matching on the dynamic task processing requirement prediction data according to the distributed server node network, and perform dynamic computing resource allocation, so as to obtain a dynamic computing resource allocation strategy;
[0092] In this embodiment, the prediction data is formatted into a table or a structured format, including fields such as task ID, resource requirements, and priority, to ensure that the data can be effectively read by subsequent processing tools. The status information of distributed server nodes, including the CPU usage rate, memory usage, network bandwidth, and load of the nodes, is obtained in real time through monitoring tools (such as Prometheus and Grafana). According to the collected node status, the available resource amount of each node is evaluated, and the available CPU, memory, and other information of each node are recorded for subsequent matching. The criteria for resource matching are determined, such as task priority and resource demand. Tasks with high demand are preferentially matched to nodes with sufficient available resources. Resource matching can be performed using algorithms such as the greedy algorithm, the minimum cost flow algorithm, or the Hungarian algorithm, which can help achieve efficient resource allocation. The predicted task demand data is traversed and matched with the available resources of distributed server nodes, and the corresponding relationship between each task and the matching node is recorded. If the resource requirements of a task cannot be matched to a suitable node within a time window, it should be recorded as a task to be assigned for subsequent processing. According to the resource matching results, a dynamic computing resource allocation strategy is formulated, which includes allocating specific computing resources (such as CPU cores and memory amounts) to each matching task, considering the task priority and the node load situation to ensure the rationality of resource allocation. The tasks are assigned to the corresponding computing nodes, and the status information of the nodes is updated, recording the resource usage after allocation. The real-time effect of resource allocation is monitored to ensure that the tasks can be executed smoothly. The generated dynamic computing resource allocation strategy is recorded in the database, including information such as task ID, allocated node, allocated resource amount, and allocation time. The effectiveness of the resource allocation strategy is evaluated regularly, and indicators such as the success rate of task execution and resource utilization efficiency are analyzed. Based on the monitoring data and historical allocation results, the resource allocation strategy is optimized to cope with future dynamic task demands.
[0093] The load scheduling module is used to perform the evolution of task resource allocation based on the dynamic computing resource allocation strategy, and then perform adaptive load scheduling optimization to construct a load scheduling optimization model.
[0094] In this embodiment, detailed information of tasks to be processed is collected, including task ID, required resource amounts (CPU, memory, I / O), priority, etc. This information can be obtained through a task scheduling system or a monitoring tool. According to the collected task requirements, the previously defined dynamic computing resource allocation policy is applied in real time to determine how to allocate resources to each task. Considering the priority of the task and the current load situation of the node, it is ensured that high-priority tasks can obtain resources first. After each resource allocation, detailed information about the resource allocation is recorded, including the allocation time, task ID, allocated node, resource amount, etc., to form historical data on the evolution of resource allocation, which is convenient for subsequent analysis and optimization. Monitoring tools (such as Prometheus, Grafana) are used to monitor the load situation of computing nodes in real time, including CPU, memory, network usage, etc. Thresholds are set, and when the node load exceeds a certain critical value, an adaptive scheduling mechanism is triggered. According to the monitored load situation, adaptive load scheduling rules are defined. For example, if the load of a certain node exceeds 80%, new tasks will be allocated to nodes with lower load. Load balancing algorithms (such as round-robin, least connections, weighted round-robin, etc.) can also be used to optimize resource allocation. According to the load monitoring results and the defined scheduling rules, resource allocation is dynamically adjusted. For example, tasks are migrated from high-load nodes to low-load nodes, or the allocation strategy for new tasks is adjusted. Scheduling decisions and results are recorded for subsequent analysis. Historical resource allocation data, load situations, task execution results, etc. are collected as the basic dataset for building an optimization model. Key features are extracted from the historical data, such as the number of tasks, resource usage, node load, scheduling delay, etc., to ensure that the features can accurately reflect the influencing factors of task scheduling. The selected model is trained using the prepared dataset, and the model parameters are optimized to improve the prediction accuracy. Cross-validation can be used to evaluate the performance of the model to ensure the generalization ability of the model. Metrics such as RMSE, MAE, R² are used to evaluate the prediction performance of the model to ensure that the model can accurately reflect the requirements of load scheduling. According to the evaluation results, the model is continuously optimized, and feature selection and model parameters are adjusted to improve the scheduling efficiency and accuracy. The trained model is incorporated into the load scheduling module to predict and optimize task resource allocation in real time, and the actual effect of the model is monitored to ensure its effectiveness in a dynamic environment.
[0095] In this embodiment, the computing node module is used to monitor the real-time operating status of the server computing nodes and perform distributed node architecture fitting to construct a distributed server node network. Specifically, it is used for:
[0096] Identify all server computing nodes;
[0097] Identify the types of terminal devices of the server computing nodes and mark the central device computing nodes and edge device computing nodes;
[0098] Monitor the real-time operating status of the central device computing node and the edge device computing node to obtain the central node operating status parameters and the edge node operating status parameters;
[0099] Conduct node load balancing feature analysis on the central node operating status parameters and the edge node operating status parameters to obtain the central node load balancing features and the edge node load balancing features;
[0100] According to the central node load balancing features and the edge node load balancing features, perform distributed node architecture fitting on the terminal device computing node and the edge device computing node, and construct a distributed server node network.
[0101] In this embodiment, use a network scanning tool (such as Nmap or Angry IP Scanner) to scan the network, identify all active server computing nodes, collect basic information such as the IP address, MAC address, and hostname of the nodes, determine the device type of the nodes by analyzing the open ports and services of the nodes (such as HTTP, SSH, FTP, etc.), use SNMP (Simple Network Management Protocol) to obtain the detailed information of the nodes, such as the device manufacturer, model, and its functions, mark the server nodes as central device computing nodes (such as data center servers) and edge device computing nodes (such as IoT devices, edge gateways, etc.) according to the identified device types, store the classification results in the database to form a list of central nodes and edge nodes, select a suitable monitoring tool (such as Zabbix, Prometheus, or Nagios) for real-time operating status monitoring, determine the operating status parameters to be monitored, such as CPU usage, memory usage, network bandwidth, disk I / O, etc., configure the monitoring tool to ensure that the operating status parameters of the central node and the edge node can be collected in real time, set the data storage policy to ensure the long-term preservation and access of the monitoring data, clean the collected operating status parameters to remove invalid data and outliers, use statistical analysis methods (such as mean, standard deviation) and machine learning algorithms (such as clustering analysis) to analyze the load balancing features, calculate the load balancing features of the central node and the edge node, such as load distribution, response time, throughput, etc., use a visualization tool (such as Matplotlib, Tableau) to generate a load balancing feature map to help understand the load situation of the nodes, design a suitable distributed node architecture according to the load balancing features to ensure the efficient utilization of computing resources, use optimization algorithms (such as genetic algorithms, particle swarm optimization) for distributed node architecture fitting to achieve the best resource allocation and load balancing, construct a distributed server node network according to the fitting results to ensure efficient communication and data flow between the central device and the edge device, and record the network topology structure, including information such as the connection relationship and bandwidth allocation of each node.
[0102] In this embodiment, the specific steps for performing node load balancing feature analysis on the central node operation state parameters and the edge node operation state parameters to obtain the central node load balancing features and the edge node load balancing features are as follows:
[0103] Calculate the CPU utilization rate of each node for the central node operation state parameters and the edge node operation state parameters to obtain the CPU utilization rate of each node;
[0104] Perform utilization rate fluctuation analysis on the CPU utilization rate of each node to generate node utilization rate fluctuation features;
[0105] Calculate the real-time memory occupancy for the central node operation state parameters and the edge node operation state parameters to obtain the memory occupancy value of each node;
[0106] Analyze the network throughput of the central node operation state parameters and the edge nodes, and extract the network throughput of each node;
[0107] Perform node current load feature analysis on the node utilization rate fluctuation features, the memory occupancy value of each node, and the network throughput of each node to obtain the current load features of each node;
[0108] Perform node type identification on the current load features of each node, and extract the load features of each central node and the load features of each edge node;
[0109] Perform load difference identification on the load features of each central node and the load features of each edge node, and extract the load difference data between the central node and the edge node;
[0110] Perform node load balancing feature analysis based on the load difference data between the central node and the edge node to obtain the central node load balancing features and the edge node load balancing features.
[0111] In this embodiment, a monitoring tool (such as Prometheus, Zabbix, or Grafana) is used to configure node monitoring to ensure that CPU usage data can be collected. The CPU utilization rate = (total time / total usage time) × 100%. For each node, the CPU usage time (user mode, kernel mode, etc.) within the monitoring period is collected, and the CPU utilization rate of each node is calculated. The fluctuation characteristics of the CPU utilization rate of each node are calculated, including the mean, standard deviation, maximum value, minimum value, etc. Through time series analysis, the utilization rate fluctuation pattern is identified, and the utilization rate fluctuation characteristics (such as fluctuation amplitude, frequency, etc.) of each node are recorded in the database to provide a basis for subsequent analysis. The monitoring tool is configured to set the monitoring parameters for memory occupancy (such as total memory, used memory, free memory), calculate the memory occupancy value for each node, and record the results. A network monitoring tool (such as Wireshark, ntop) is used to monitor network traffic, and necessary parameters (such as bandwidth, traffic) are configured. The network throughput can be calculated by the following formula: network throughput = time / amount of data transmitted (such as Mbps). The CPU utilization rate fluctuation characteristics, memory occupancy value, and network throughput of each node are integrated into the current load characteristics. Data analysis tools (such as Pandas, NumPy) are used to analyze the integrated data, extract the current load characteristics of each node, including CPU, memory, and network loads. Based on the load characteristics of the nodes (such as CPU, memory, network traffic patterns), node type identification is performed to determine whether it is a central node or an edge node. The load characteristics of each central node and edge node are extracted and recorded to ensure the integrity of the characteristics. Statistical methods such as t-tests or analysis of variance are used to compare the load characteristics of central nodes and edge nodes to identify load differences. The identified load difference data is recorded in the database for subsequent analysis. Methods such as clustering analysis and regression analysis are used to analyze the load balancing characteristics, and based on the load difference data between central nodes and edge nodes, the load balancing characteristics of central nodes and edge nodes are extracted.
[0112] In this embodiment, the resource call module is used to obtain the server operation logs; perform server resource call behavior analysis on the server operation logs, and perform the evolution of the resource call timing characteristics to generate the server resource call timing characteristics, specifically used for:
[0113] Obtain the server operation logs; perform server resource call behavior analysis on the server operation logs, and extract the server resource call behavior data;
[0114] Perform time series window partitioning on the server resource call behavior data to obtain the resource call behavior data of multiple time windows;
[0115] Count the number of scheduling times for the resource call behavior data of multiple time windows to obtain the scheduling times for each time window;
[0116] Calculate the scheduling frequency based on the scheduling times for each time window to generate the resource call frequency for each window;
[0117] Evolve the resource call timing characteristics for the resource call frequency of each window to generate the server resource call timing characteristics.
[0118] In this embodiment, determining the types of logs to be collected includes: operating system logs such as syslog and dmesg, which record system events and errors; application logs, which are the detailed information during the operation of applications, such as the logs of web servers (Apache, Nginx) or databases (MySQL, PostgreSQL); security logs, which record security-related events such as user logins and permission changes. Using log collection tools (such as rsyslog or Filebeat), configure the log source and the target storage location to ensure that the tool can read the log files in real-time or periodically and transfer them to a central storage system (such as the ELK Stack or Hadoop). Set up a log rotation mechanism to prevent the log files from becoming too large and affecting the server performance. Use log parsing tools (such as Logstash, Fluentd) to parse the running logs and extract key fields, including: timestamp, resource type (CPU, memory, disk, network, etc.), call duration or request size. Define parsing rules to ensure the accuracy and consistency of the data. Extract the server resource call behavior data from the parsed logs to form a structured data set, recording the detailed information of each call. According to the characteristics of server resource usage and the analysis requirements, select a suitable time window size (such as 1 minute, 5 minutes, or 10 minutes). The window size needs to consider the analysis precision and data processing capabilities. Divide the resource call behavior data within the entire time range into multiple time windows to ensure the data integrity within each time window. A sliding window (for example, updating the window at regular intervals) or a fixed window (each window has a fixed time period) can be used for division. Count the resource call behaviors within each time window and calculate the scheduling times (i.e., the total number of resource calls) for each window. This can be achieved through a simple counter. Record the scheduling times of each time window and its corresponding time window information in a database or a data frame for subsequent analysis and query. The scheduling frequency can be calculated using the following formula: scheduling frequency = window size (time) / scheduling times. Apply this formula to each time window to calculate the resource call frequency, which reflects the activity level of resource usage within each window. Conduct time series analysis on the resource call frequency of each window to extract time series features, including: average frequency: the average resource call frequency within the window; volatility: the degree of change in frequency, which may be calculated using the standard deviation or coefficient of variation; trend: analyze the change trend of frequency over time to identify increasing or decreasing patterns; periodicity: check for the existence of periodic patterns (such as peak times every day). Record the extracted time series features (such as frequency change trends, fluctuations) in the database for subsequent analysis and report generation.
[0119] In this embodiment, the task traversal module is used to traverse and identify unprocessed tasks in the server operation log, mine task status features, and construct an unprocessed task status mapping sequence. Specifically, it is used for:
[0120] Traverse and identify unprocessed tasks in the server operation log, and extract multiple unprocessed tasks of the server;
[0121] Calculate the remaining time for task processing of multiple unprocessed tasks of the server, and extract the unprocessed task timestamps;
[0122] Calculate the task resource occupancy of the unprocessed tasks of the server, and extract the task resource occupancy data;
[0123] Mine task status features from the unprocessed task timestamps and task resource occupancy data to generate the status features of each unprocessed task;
[0124] Encode the task sequences of multiple unprocessed tasks of the server to obtain an unprocessed task encoding sequence;
[0125] Perform corresponding task status mapping on the unprocessed task encoding sequence according to the status features of each unprocessed task, and construct an unprocessed task status mapping sequence.
[0126] In this embodiment, the types of logs to be collected are determined, mainly including: Operating system logs: such as syslog or journalctl, which record system events and error messages; Application logs: such as Web server logs (e.g., Apache, Nginx) and database logs (e.g., MySQL, PostgreSQL), which record the running status of applications and request processing; Task scheduling logs: monitor the logs generated by task schedulers (e.g., Cron, Celery), which record the start, completion, and failure of tasks. Use parsing tools (e.g., Logstash, Fluentd, or custom Python scripts) to parse the log files according to predefined formats. When parsing, pay attention to the following fields: Task ID: uniquely identifies each task; Task status: such as "not started", "in progress", "failed", "completed", etc.; Timestamp: records the creation time, start time, and end time of the task. Extract these fields through regular expressions or log analysis libraries (e.g., pandas). Store the unprocessed task information (including task ID, timestamp, and status) filtered out into a database (e.g., MySQL, PostgreSQL) or a data frame (e.g., Pandas DataFrame) to form a list of unprocessed tasks. For each unprocessed task, determine its estimated completion time. This time may be directly provided in the log or need to be calculated based on the task type and historical data. If there is no direct estimated completion time, the following methods can be used for estimation: Historical average time: Analyze the historical data of similar tasks and calculate their average processing time; Dynamic estimation: Dynamically adjust according to the current system load, resource usage, and task complexity. Record the remaining time (in seconds or minutes) of the unprocessed tasks calculated and the timestamp into the database to form a record containing the task ID, remaining time, and timestamp. Use monitoring tools (e.g., Prometheus, Grafana, or system monitoring commands such as top, htop) to obtain the resource occupancy of the current unprocessed tasks. The key resources include: CPU occupancy: the CPU time or percentage used by each task; Memory occupancy: the amount of memory used by each task; I / O occupancy: the number and speed of disk read and write operations. Associate the monitored resource usage with the corresponding task ID. You can use calculation scripts (e.g., Python scripts combined with the psutil library) to perform statistics on the resource occupancy of each unprocessed task and store the results in the database. Determine the key status characteristics of the unprocessed tasks, such as: Remaining time: the time difference from the current time to the task's estimated completion time; Current resource occupancy: the real-time values of CPU, memory, and I / O occupancy; Task priority: if the task priority is defined in the task management system, record the task priority; Historical status: the past success rate, number of failures, etc. of the task.Use statistical analysis tools (such as pandas) to extract features for each unprocessed task, generate a feature vector for each task. You can use clustering analysis techniques (such as K-means) to classify tasks and extract relevant features. To assign a unique code to each unprocessed task, the following methods can be adopted: directly use the task ID: simple and direct, facilitating subsequent reference; hash encoding: use a hash function (such as SHA256) to hash the task ID or task features to generate a fixed-length code. According to the status features of each unprocessed task, define a mapping rule to associate each task in the code sequence with the corresponding status feature. By combining the unprocessed task code sequence with the corresponding status features, generate an unprocessed task status mapping sequence. This sequence will be used for subsequent data analysis and decision support to help identify the priority of task processing and resource allocation.
[0127] In this embodiment, the demand prediction module is used to dynamically predict the demand for task processing of the unprocessed task status mapping sequence based on the timing characteristics of server resource calls to obtain dynamic task processing demand prediction data. Specifically, it is used for:
[0128] Perform multi-point resource call prediction on the timing characteristics of server resource calls to generate resource call prediction data at multiple time points;
[0129] Conduct resource call timing change analysis on the resource call prediction data at multiple time points to obtain resource call timing change data;
[0130] Conduct per-task resource requirement analysis on the unprocessed task status mapping sequence to obtain the resource requirement data for each task;
[0131] Based on the resource call timing change data, perform dynamic task processing demand prediction on the resource requirement data for each task to obtain dynamic task processing demand prediction data.
[0132] In this embodiment, historical resource call data is collected: historical resource call time series data is extracted from a monitoring system or a log collection tool. These data should include timestamps, CPU usage, memory occupancy, network bandwidth, etc. Missing values and outliers are cleaned to ensure data integrity, and normalization or standardization processing is performed as needed for model training. Additional features are extracted from the timestamps, such as hour, day, week, season, etc., to capture potential periodic behaviors. Sliding window features: The sliding window technique is used to create additional features, such as the average and maximum values of the past N time points, to help the model capture time dependence. Model selection: A suitable prediction model is selected. For example: ARIMA: suitable for stationary time series; LSTM: a deep learning model suitable for capturing long-term and short-term dependencies; Prophet: suitable for time series with strong seasonal and trend changes. The selected model is trained using the training set (usually the first 80%-90% of the historical data), and the model performance is evaluated using the test set (the remaining part). Evaluation is performed through metrics such as root mean square error (RMSE). Analyze the changes in predicted data at multiple time points. Use the percentage change between adjacent time points: Change rate = (Predicted value at the current time point - Predicted value at the previous time point) / Predicted value at the previous time point × 100%. By plotting a time series graph, observe the trend line of the predicted values and identify rising, falling, or stable patterns. The change rate, trend direction, and fluctuation amplitude are recorded as a new data set for subsequent analysis and decision-making. The generated unprocessed task status mapping sequence should contain the status information of each task, such as not started, in progress, paused, etc. Define resource requirement characteristics for each task, including: CPU requirement: the percentage of CPU required for each task; Memory requirement: the amount of memory (MB) required for task execution; I / O requirement: the amount of disk I / O operations required for the task. Traverse the unprocessed task status mapping sequence and extract the resource requirement data for each task based on historical data or the resource usage of similar tasks. This can be done through statistical analysis methods, such as calculating the average resource requirements of historical similar tasks. Integrate the resource call time series change data and the resource requirements data of unprocessed tasks to form a comprehensive data set. This data set should include time series features, resource change rates, and task requirement features. According to the characteristics of the integrated data, select a suitable dynamic demand prediction model. For example: Regression models: such as linear regression, random forest regression; Deep learning models: such as LSTM, suitable for capturing dynamic changes in time series data. Use the integrated data set to train the prediction model and evaluate the model performance to ensure that the model can effectively capture the relationship between resource requirements and call changes. Based on the trained model, predict the future task execution situation and generate dynamic task processing demand prediction data. These data should contain the resource requirement predictions for each task at different time points.
[0133] In this embodiment, the specific steps of performing multi-point resource call prediction on the server resource call timing characteristics to generate resource call prediction data at multiple time points are as follows:
[0134] Perform multi-point segmentation processing on the server resource call timing characteristics to generate resource call timing data at multiple time points;
[0135] Extract the timestamps of the resource call timing data at multiple time points;
[0136] Perform multi-point timing fitting according to the timestamps to obtain a multi-point resource call timing curve;
[0137] Perform deep call feature learning on the multi-point resource call timing curve to generate server resource call rules;
[0138] Based on the server resource call rules, perform short-term resource call simulation on the server resource call timing characteristics to generate resource call simulation data;
[0139] Define the prediction time node;
[0140] Perform resource call prediction on the resource call simulation data according to the prediction time node to generate resource call prediction data at multiple time points.
[0141] In this embodiment, according to requirements, the time series data is segmented into multiple time periods (such as every minute, every hour, or each fixed time window). The sliding window or fixed window method can be used. The data for each time period will be used as the basis for subsequent analysis. The resource call data for multiple time points after segmentation is stored in a database or data framework to form a structured time series data set. The timestamps for each time period are extracted from the segmented time series data to ensure one-to-one correspondence with the resource call data. Timestamps usually include date and time information for subsequent analysis. The extracted timestamps are stored together with the corresponding resource call data to form a complete time series data set containing timestamps. A suitable fitting model is selected according to the data characteristics, such as: linear regression model: suitable for relatively simple linear relationships; polynomial regression: suitable for more complex relationships; time series prediction models: such as ARIMA, SARIMA, used to handle trends and seasonality in time series data. The extracted timestamps and the corresponding resource call data are fitted, and the selected model is used to train the data to obtain the resource call time series curve. The fitting effect is evaluated. Common metrics include root mean square error (RMSE) and coefficient of determination (R²). Determine the features to be extracted, such as: peak moment of call frequency, call pattern (such as periodicity, trend). Deep learning models (such as LSTM, GRU) can be used for feature learning to capture complex patterns in time series data. Prepare the training set and validation set. Use historical resource call data for model training. Train the deep learning model, and use the backpropagation algorithm to optimize the weights and biases, and improve the learning effect of the model through multiple iterations. According to the extracted resource call rules, select a suitable simulation method. The following methods can be used: Monte Carlo simulation: generate multiple possible call scenarios through random sampling; model-based simulation: use the time series model established before to generate short-term resource call data. Generate short-term resource call simulation data for the future according to the extracted rules and historical data, and record the simulation results. Define the prediction time nodes according to business requirements and prediction period, such as specific time points every 5 minutes, every hour, or every day. Combine the generated resource call simulation data with the prediction time nodes for subsequent prediction analysis. According to the characteristics of the simulation data, select a suitable prediction model (such as SARIMA, LSTM) for short-term prediction. Use the trained model to predict the resource call simulation data according to the defined prediction time nodes to generate resource call prediction data for multiple time points.
[0142] In this embodiment, the resource allocation module is used to perform callable resource matching on the dynamic task processing demand prediction data according to the distributed server node network, and perform dynamic computing resource allocation, so as to obtain a dynamic computing resource allocation strategy. Specifically, it is used for:
[0143] Identify callable computing nodes in the distributed server node network to mark callable computing nodes;
[0144] Calculate the idle resources of callable computing nodes and extract the idle resource data of each computing node;
[0145] Perform callable resource matching on the dynamic task processing demand prediction data based on the idle resource data of each computing node to obtain task processing resource matching data;
[0146] Perform dynamic computing resource allocation on the distributed server node network based on the task processing resource matching data to obtain a dynamic computing resource allocation strategy.
[0147] In this embodiment, node status information is obtained from the monitoring platform of the distributed system (such as Kubernetes, Apache Mesos, or OpenStack). These information usually include the health status of the node, IP address, CPU and memory usage, etc. Use APIs or command-line tools to query the real-time status of the node to ensure that the latest node information can be obtained. Determine the criteria for callable computing nodes. For example, the node must be in the "online" state and the current CPU usage rate is lower than the set threshold (such as 70%). Additional criteria can be defined, such as memory usage rate, load balancing status, etc., to ensure that the node has sufficient resources to handle new tasks. Traverse the collected node information and mark the computing nodes that meet the callable criteria. This step usually records the ID, status, and current resource usage of the node for subsequent resource calculation and matching. Use system monitoring tools (such as Prometheus, Nagios, or Zabbix) to obtain the real-time resource usage data of callable nodes. These tools provide rich monitoring metrics that can help us evaluate the resource usage of each node. In particular, extract the CPU and memory usage to ensure the accuracy and timeliness of the data. The calculation of idle resources includes subtracting the current usage from the total resources. For example, the idle CPU can be obtained by subtracting the currently used cores from the total number of CPU cores, and the idle memory is calculated by subtracting the currently used memory from the total memory. Organize the idle resource data of each node into a structured format (such as a table or database) for subsequent analysis and decision-making. Obtain the resource requirement information of the tasks to be processed from the task management system or scheduler. These information usually include the number of CPU cores, memory size, and other resource requirements required by the tasks. Ensure that the requirement information of the tasks is the latest and classify them according to the priority of the tasks. Determine how to match the task requirements with the idle resources of the callable computing nodes. Simple conditional judgments (such as whether the idle resources are sufficient) or more complex algorithms (such as greedy algorithms, minimum cost flow algorithms) can be used for matching. Traverse the resource requirements of each task and find callable computing nodes that meet the requirements. Record the corresponding relationship between each task and its matching node for subsequent resource allocation. This step needs to ensure that each task is effectively matched to a suitable computing node. Based on the priority of the tasks and the load conditions of the callable nodes, design a dynamic resource allocation strategy. Tasks with higher priorities should be allocated to nodes with more idle resources as much as possible without affecting the overall performance of the system. Consider using load balancing algorithms to ensure the uniform distribution of resources and avoid overloading some nodes while other nodes are idle. According to the resource matching results, dynamically allocate tasks to suitable computing nodes. Monitor the allocation and execution status of tasks to ensure that each task can run efficiently on the most suitable node. Use monitoring tools to continuously observe the task execution situation and the resource usage of computing nodes.Monitoring metrics include CPU and memory usage, task execution time, etc. If it is found that the resource usage of some nodes is unbalanced or the task processing is delayed, the system should be able to quickly re-evaluate the resource allocation strategy and re-allocate tasks if necessary to ensure the efficient operation of the system.
[0148] In this embodiment, the load scheduling module is used to perform task resource allocation evolution based on the dynamic computing resource allocation strategy, and then perform adaptive load scheduling optimization to construct a load scheduling optimization model. Specifically, it is used for:
[0149] Perform task resource allocation evolution on the distributed server node network based on the dynamic computing resource allocation strategy to generate task resource allocation evolution data;
[0150] Perform computing node load analysis on the task resource allocation simulation data to obtain the load value of each computing node;
[0151] Perform node load peak analysis on the load value of each computing node to generate node load peaks;
[0152] Perform adaptive load scheduling optimization on the dynamic computing resource allocation strategy based on the node load peak to construct a load scheduling optimization model.
[0153] In this embodiment, the previously designed dynamic computing resource allocation strategy is used. By monitoring the current task requirements and the idle resources of computing nodes, real-time task allocation is performed. During the allocation process, the task priority, node load conditions, and resource availability are considered to flexibly adjust the resource allocation. After each resource allocation, the allocation situation of each task is recorded, including the task ID, the ID of the allocated computing node, the allocated resource amount (such as CPU and memory), etc. These information are stored in a database or a logging system to form a structured data format for subsequent analysis. Real-time load data of each computing node are collected from a monitoring system (such as Prometheus, Grafana, or a custom monitoring tool), which includes CPU usage rate, memory usage, I / O operations, etc. Ensure the accuracy of the timestamp information of the data for subsequent analysis. Calculate the load value of each node based on the collected data. The load value can be obtained through the following formula: Load value = CPU usage rate + Memory usage rate + I / O waiting time (weighted according to specific situations). Record the load value of each node to form a load analysis data set. Analyze the load value of each computing node within a certain time window to identify the load peak. Statistical methods (such as moving average, maximum value) can be used to identify the peak. Set a suitable time window (such as 5 minutes, 10 minutes), and calculate the maximum value of the load value within this window. Record the load peak of each node to form a node load peak data set. This data set will be used for subsequent load scheduling optimization. Based on the node load peak data, design an adaptive load scheduling optimization model. The model can adopt machine learning algorithms (such as regression model, decision tree, random forest, etc.) to predict load changes and resource requirements. The model input can include historical load data, task characteristics, node status, etc., and the output is the optimal resource allocation strategy for each node. Use the historical data set (including task resource allocation evolution data and node load data) to train the optimization model. Through cross-validation of the model, evaluate its accuracy and generalization ability, and adjust the model parameters to improve the prediction accuracy and scheduling efficiency. During the actual operation process, use the optimization model to adjust the resource allocation strategy in real time. When it is detected that the load of a certain node is close to the peak, the system should automatically adjust the task allocation, reduce the number of tasks on the node with too high load, or migrate the tasks to the node with lower load. Monitor the scheduling effect and continuously adjust and optimize the model according to the feedback.
[0154] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to encompass all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.
[0155] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A server resource scheduling system integrating AI and edge computing, characterized in that, The server resource scheduling system integrating AI and edge computing includes a computing node module, a resource invocation module, a task traversal module, a demand prediction module, a resource allocation module, and a load scheduling module: The computing node module is used for real-time operation status monitoring and performing distributed node architecture fitting to construct a distributed server node network; The resource invocation module is used for obtaining the server operation logs; Performing server resource invocation behavior analysis on the server operation logs and evolving the timing characteristics of resource invocation to generate server resource invocation timing characteristics; The task traversal module is used for traversing and identifying unprocessed tasks in the server operation logs and mining task status characteristics to construct an unprocessed task status mapping sequence; The demand prediction module is used for dynamically predicting the task processing demands of the unprocessed task status mapping sequence according to the server resource invocation timing characteristics to obtain dynamic task processing demand prediction data; The resource allocation module is used for matching the callable resources for the dynamic task processing demand prediction data according to the distributed server node network and performing dynamic computing resource allocation to obtain a dynamic computing resource allocation strategy; The load scheduling module is used for evolving the task resource allocation based on the dynamic computing resource allocation strategy, then performing adaptive load scheduling optimization to construct a load scheduling optimization model; Among them, the computing node module is used for real-time operation status monitoring and performing distributed node architecture fitting to construct a distributed server node network. Specifically, it is used for: Identifying all server computing nodes; Identifying the types of terminal devices of the server computing nodes and marking the central device computing nodes and edge device computing nodes; Performing real-time operation status monitoring on the central device computing nodes and edge device computing nodes to obtain central node operation status parameters and edge node operation status parameters; Performing node load balancing characteristic analysis on the central node operation status parameters and edge node operation status parameters to obtain central node load balancing characteristics and edge node load balancing characteristics; Performing distributed node architecture fitting on the terminal device computing nodes and edge device computing nodes according to the central node load balancing characteristics and edge node load balancing characteristics to construct a distributed server node network; Among them, the task traversal module is used for traversing and identifying unprocessed tasks in the server operation logs and mining task status characteristics to construct an unprocessed task status mapping sequence. Specifically, it is used for: Traversing and identifying unprocessed tasks in the server operation logs and extracting multiple unprocessed server tasks; Calculating the remaining time for processing the multiple unprocessed server tasks and extracting the unprocessed task timestamps; Calculating the task resource occupancy of the server unprocessed tasks and extracting the task resource occupancy data; Mining the task status characteristics of the unprocessed task timestamps and task resource occupancy data to generate the status characteristics of each unprocessed task; Encoding the task sequences of the multiple unprocessed server tasks to obtain an unprocessed task encoding sequence; Performing corresponding task status mapping on the unprocessed task encoding sequence according to the status characteristics of each unprocessed task to construct an unprocessed task status mapping sequence.
2. The server resource scheduling system integrating AI and edge computing according to claim 1, characterized in that, The specific steps for performing node load balancing feature analysis on the central node operation status parameters and edge node operation status parameters to obtain the central node load balancing features and edge node load balancing features are as follows: Calculate the CPU utilization rate of each node for the central node operation status parameters and edge node operation status parameters to obtain the CPU utilization rate of each node; Conduct utilization rate fluctuation analysis on the CPU utilization rate of each node to generate node utilization rate fluctuation features; Calculate the real-time memory occupancy for the central node operation status parameters and edge node operation status parameters to obtain the memory occupancy value of each node; Analyze the network throughput of the central node operation status parameters and edge nodes, and extract the network throughput of each node; Conduct node current load feature analysis on the node utilization rate fluctuation features, the memory occupancy value of each node, and the network throughput of each node to obtain the current load features of each node; Conduct node type identification on the current load features of each node to extract the load features of each central node and the load features of each edge node; Conduct load difference identification on the load features of each central node and the load features of each edge node to extract the load difference data between the central node and the edge node; Based on the load difference data between the central node and the edge node, perform node load balancing feature analysis to obtain the central node load balancing features and edge node load balancing features.
3. The server resource scheduling system integrating AI and edge computing according to claim 1, wherein The resource invocation module is used to obtain the server operation log; conduct server resource invocation behavior analysis on the server operation log, and perform resource invocation timing feature evolution to generate server resource invocation timing features, specifically used for: Obtain the server operation log; conduct server resource invocation behavior analysis on the server operation log, and extract server resource invocation behavior data; Perform time window partitioning on the server resource invocation behavior data to obtain resource invocation behavior data for multiple time windows; Conduct scheduling times statistics on the resource invocation behavior data for multiple time windows to obtain the scheduling times for each time window; Calculate the scheduling frequency based on the scheduling times for each time window to generate the resource invocation frequency for each window; Perform resource invocation timing feature evolution on the resource invocation frequency for each window to generate server resource invocation timing features.
4. The server resource scheduling system integrating AI and edge computing according to claim 1, wherein The demand prediction module is used to perform dynamic task processing demand prediction on the unprocessed task status mapping sequence according to the server resource invocation timing features to obtain dynamic task processing demand prediction data, specifically used for: Perform multi-point resource invocation prediction on the server resource invocation timing features to generate resource invocation prediction data for multiple time points; Conduct resource invocation timing change analysis on the resource invocation prediction data for multiple time points to obtain resource invocation timing change data; Conduct per-task resource requirement analysis on the unprocessed task status mapping sequence to obtain the resource requirement data for each task; Based on the resource invocation timing change data, perform dynamic task processing demand prediction on the resource requirement data for each task to obtain dynamic task processing demand prediction data.
5. The server resource scheduling system integrating AI and edge computing according to claim 4, characterized in that, The specific steps for predicting resource calls at multiple time points for the timing characteristics of server resource calls, thereby generating resource call prediction data at multiple time points are as follows: Perform multi-time point segmentation processing on the timing characteristics of server resource calls to generate resource call timing data at multiple time points; Extract the timestamps of the resource call timing data at multiple time points; Perform multi-time point timing fitting based on the timestamps to obtain a multi-time point resource call timing curve; Perform in-depth call feature learning on the multi-time point resource call timing curve to generate server resource call rules; Based on the server resource call rules, perform short-term resource call simulation on the timing characteristics of server resource calls to generate resource call simulation data; Define the prediction time node; Based on the prediction time node, perform resource call prediction on the resource call simulation data, thereby generating resource call prediction data at multiple time points.
6. The server resource scheduling system integrating AI and edge computing according to claim 1, characterized in that, The resource allocation module is used to perform callable resource matching on the dynamic task processing demand prediction data according to the distributed server node network, and perform dynamic computing resource allocation, thereby obtaining a dynamic computing resource allocation strategy, and is specifically used for: Identify callable computing nodes in the distributed server node network to mark callable computing nodes; Calculate the idle resources of the callable computing nodes and extract the idle resource data of each computing node; Perform callable resource matching on the dynamic task processing demand prediction data based on the idle resource data of each computing node to obtain task processing resource matching data; Based on the task processing resource matching data, perform dynamic computing resource allocation on the distributed server node network, thereby obtaining a dynamic computing resource allocation strategy.
7. The server resource scheduling system integrating AI and edge computing according to claim 1, wherein, The load scheduling module is used to perform task resource allocation evolution based on the dynamic computing resource allocation strategy, and then perform adaptive load scheduling optimization to construct a load scheduling optimization model, and is specifically used for: Perform task resource allocation evolution on the distributed server node network based on the dynamic computing resource allocation strategy to generate task resource allocation evolution data; Perform computing node load analysis on the task resource allocation simulation data to obtain the load value of each computing node; Perform node load peak analysis on the load value of each computing node to generate node load peaks; Based on the node load peaks, perform adaptive load scheduling optimization on the dynamic computing resource allocation strategy to construct a load scheduling optimization model.
Citation Information
Patent Citations
Model computing resource dynamic planning method
CN117762633A
Computing power data management system and method based on distributed computing
CN119025283A
Methods and Systems for Intelligent Distribution of Workloads to Multi-Access Edge Compute Nodes on a Communication Network
US20200351336A1