Resource scheduling method of server and electronic equipment

By combining load prediction models with temporal and spatial characteristics, dynamically adjusting the task allocation strategy, the problem of being unable to quickly optimize server resource scheduling in the existing technology is solved, and efficient resource utilization and task allocation of server clusters are realized.

CN120256114APending Publication Date: 2025-07-04INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510352887.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the server resource scheduling method cannot be optimized based on fluctuations in task load and resource consumption changes in a short time, resulting in excessive loading of the server or waste of resources.

Method used

By introducing spatial characteristics based on the temporal characteristics of historical load data and the load correlation of adjacent servers, the load prediction model is used to perform server load prediction, and dynamically adjust the resource allocation and scheduling strategies of tasks in combination with task type weights.

Benefits of technology

It realizes the coordinated optimization of server cluster resources and task allocation efficiency, avoids server overload and resource waste, and improves resource utilization and task execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256114A_ABST
    Figure CN120256114A_ABST
Patent Text Reader

Abstract

The invention discloses a resource scheduling method of a server and electronic equipment, and relates to the technical field of network optimization. Time features based on historical load data and spatial features based on load correlation of adjacent servers are introduced in a feature extraction stage; the dynamic change rule of the server load and the collaborative influence of the cluster environment are comprehensively captured, and the input information richness and context sensing ability of the prediction model are improved. The method comprises the following steps: predicting load data at a target moment through a load prediction model, carrying out weighted calculation on a server load state in combination with a task type weight, quantitatively evaluating actual occupation demands of different task types on server resources, and dynamically adjusting a task allocation strategy based on an inverse relation between the load state and an allocation priority. According to the method and the device, the problem that a load balancing strategy cannot be optimized in a short time according to fluctuation of task loads and changes of resource consumption can be solved, and refined scheduling of server cluster resources and collaborative optimization of task allocation efficiency are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network optimization technologies, and particularly to a method for resource scheduling of a server and an electronic device. Background Art

[0002] With the rapid development of technologies such as cloud computing, big data, and the Internet of Things, the scale of data centers has expanded rapidly, and the computing power of servers and the demand for resource usage have also increased accordingly.

[0003] Currently, the resource scheduling of servers mainly relies on static load balancing strategies. Although this scheduling method can, to a certain extent, adjust the server load, with the fluctuations of task loads and changes in resource consumption, this method cannot optimize the load balancing strategy in a short time, resulting in situations such as overloading of servers or waste of server resources. Summary of the Invention

[0004] This application provides a method for resource scheduling of a server and an electronic device to at least solve the problem in the related art that the load balancing strategy cannot be optimized in a short time, resulting in overloading of servers or waste of server resources.

[0005] This application provides a method for resource scheduling of a server, including:

[0006] Performing feature extraction on the monitoring data of the server to obtain time features and spatial features; wherein, the time features include statistical features of historical load data, and the spatial features include load correlation features of adjacent servers of the server;

[0007] Inputting the time features and spatial features into a load prediction model to perform server load prediction, and obtaining a load prediction result of the server at a target moment;

[0008] Dynamically adjusting the resource allocation and scheduling strategy of the task based on the load prediction result and the server task type.

[0009] This application also provides a resource scheduling device for a server, including:

[0010] A feature extraction unit, configured to perform feature extraction on the monitoring data of the server to obtain time features and spatial features; wherein, the time features include statistical features of historical load data, and the spatial features include load correlation features of adjacent servers of the server;

[0011] A prediction unit, configured to input the time features and spatial features into a load prediction model to perform server load prediction, and obtaining a load prediction result of the server at a target moment;

[0012] An adjustment unit for dynamically adjusting the resource allocation and scheduling strategy of tasks based on the load prediction result and the server task type.

[0013] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above server resource scheduling methods when executing the computer program.

[0014] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above server resource scheduling methods when executed by a processor.

[0015] This application also provides a computer program product including a computer program, which implements the steps of any of the above server resource scheduling methods when executed by a processor.

[0016] Through this application, since both the time features based on historical load data and the spatial features based on the load correlation of adjacent servers are introduced in the feature extraction stage, the dynamic change law of server load and the collaborative influence of the cluster environment can be captured more comprehensively, improving the richness of input information and the context awareness ability of the prediction model. Through the prediction of the load data at the target moment by the load prediction model, combined with the task type weights, the server load status is weighted and calculated to quantitatively evaluate the actual occupancy requirements of different task types for server resources, and then the task allocation strategy is dynamically adjusted based on the inverse relationship between the load status and the allocation priority. Therefore, this application can solve the problem of being unable to optimize the load balancing strategy in a short time according to the fluctuations of task loads and the changes in resource consumption, and realize the fine-grained scheduling of server cluster resources and the collaborative optimization of task allocation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0018] Figure 1 It is a schematic flowchart of a server resource scheduling method provided by an embodiment of this application;

[0019] Figure 2 It is a schematic diagram of a server resource scheduling method provided by an embodiment of this application;

[0020] Figure 3 It is a schematic flowchart of a server resource scheduling method provided by an embodiment of this application;

[0021] Figure 4 A structural schematic diagram of a resource scheduling device for a server provided by an embodiment of the present application;

[0022] Figure 5 A structural schematic diagram of a resource scheduling device for a server provided by an embodiment of the present application. Detailed implementation manners

[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.

[0024] The following describes a resource scheduling method and an electronic device for a server according to embodiments of the present disclosure with reference to the accompanying drawings.

[0025] Figure 1 A flowchart of a resource scheduling method for a server provided by an embodiment of the present application.

[0026] As Figure 1 shown, the method includes the following steps:

[0027] Step 101: Extract features from the monitoring data of the server to obtain time features and spatial features; wherein, the time features include statistical features of historical load data, and the spatial features include load correlation features of adjacent servers of the server.

[0028] In server resource management and task scheduling, feature extraction aims to capture the dynamic change law of server load. Time features are mainly based on statistical features of historical load data, such as mean, variance, maximum value, minimum value, trend, and periodicity, etc., which are used to reflect the change law of server load over time, such as average load level, fluctuation degree, extreme value range, and long-term change trend and periodic pattern. Spatial features are extracted by analyzing the load correlation between the server and its adjacent servers, including load correlation coefficient, load distribution of adjacent servers, and overall load balance state of the cluster, etc., which are used to reflect the collaborative effect and load distribution of the server in the cluster environment.

[0029] By combining time features and spatial features, the spatio-temporal dynamic changes of server load can be comprehensively captured, providing more accurate data support for load prediction and task allocation.

[0030] Step 102: Input the time feature and the space feature into the load prediction model for server load prediction, and obtain the load prediction result at the target moment of the server.

[0031] Time features (such as the mean, variance, trend, and periodicity of historical load data) provide the regular information of the server load changing over time, helping the model understand the historical behavior of the load; space features (such as the load correlation coefficient of adjacent servers and the cluster load distribution) reflect the collaborative effect and load distribution of the server in the cluster environment, enhancing the model's perception ability of the cluster dynamics. By jointly inputting these two types of features into the load prediction model, the load prediction model can comprehensively consider the variation law in the time dimension and the environmental impact in the space dimension, so as to more accurately predict the server load status at the target moment.

[0032] In some embodiments, the load prediction model can be a Long Short-Term Memory (LSTM) network. The LSTM model not only processes the information in the time dimension but also processes the space information through spatio-temporal feature fusion.

[0033] The prediction result of the load prediction model provides an important basis for subsequent task allocation and resource scheduling, helping to achieve refined management and efficient utilization of server resources.

[0034] Step 103: Dynamically adjust the resource allocation and scheduling strategy of the task based on the load prediction result and the server task type.

[0035] In some embodiments, according to the load prediction result, identify the servers with low load currently and in a period of time in the future, and combine with the demand characteristics of the task type (such as compute-intensive, storage-intensive, or network-intensive), and allocate the task to the most suitable server. For example, compute-intensive tasks are preferentially allocated to servers with sufficient CPU resources, while storage-intensive tasks are allocated to servers with rich storage resources.

[0036] By introducing the weight coefficient of the task type, quantitatively evaluate the resource occupancy of different task types on the server, ensure that high-priority tasks can obtain resource support first, and at the same time avoid low-priority tasks occupying too many resources. In some embodiments, dynamically monitor the change of the server load, and adjust the task allocation strategy in real time. For example, when the load of a certain server suddenly increases, migrate some tasks to the server with lower load, or when the overall load of the cluster is low, make full use of the idle resources to execute low-priority tasks. Through this dynamic adjustment mechanism, it can effectively avoid server overload or resource waste, improve resource utilization and task execution efficiency, and thus achieve intelligent scheduling and optimized management of the server cluster.

[0037] In some embodiments, before performing step 101, sensors and monitoring tools are deployed on each server node in a data center to collect the resource usage of the servers in real time. Please refer to Figure 2 , Figure 2 which is a schematic diagram of a resource scheduling method for a server provided by an embodiment of the present application. As shown in Figure 2 , start a data collection module to periodically collect real-time monitoring data and create a database table to store the collected data:

[0038] Server resource usage data: including CPU usage rate, memory occupancy rate, disk space usage, network bandwidth, etc.; please refer to Table 1, which is a server resource usage data table provided by an embodiment of the present application;

[0039] Table 1

[0040] server_id Server unique identifier timestamp Recording time cpu_usage CPU usage rate at the current moment memory_usage Memory usage rate at the current moment disk_usage Hard disk usage at the current moment network_in Network receive bandwidth usage network_out Network transmit bandwidth usage

[0041] Task execution status data: including task type, priority, execution duration, status (such as running, completed, etc.); please refer to Table 2, which is a task information table provided by an embodiment of the present application:

[0042] Table 2

[0043]

[0044] Historical load data: provides training data for the LSTM model for load prediction; please refer to Table 3, which is a historical load data table provided by an embodiment of the present application:

[0045] Table 3

[0046] server_id Server unique identifier timestamp Recording time cpu_usage CPU usage rate memory_usage Memory usage rate disk_usage Disk usage rate network_in Network receive bandwidth usage network_out Network transmit bandwidth usage

[0047] Spatial load data: stores the load relationship between different servers for modeling spatial features. Please refer to Table 4, which is a spatial load table provided by an embodiment of the present application:

[0048] Table 4

[0049]

[0050] In some embodiments, the feature extraction of the monitoring data of the server in step 101 further includes:

[0051] Obtain the monitoring data corresponding to a preset number of time steps in a preset database; wherein, the monitoring data includes the load data of the server and the network traffic data of the adjacent server;

[0052] Please continue to refer to Figure 2, start the load prediction module, extract historical data and multi-dimensional input parameters, which are used to capture the load change pattern of the server in the time dimension and the synergy effect with adjacent servers in the space dimension. Specifically, the load data of the server reflects its resource usage (such as the usage rates of CPU, memory, disk, and network), while the network traffic data of adjacent servers provides the communication load and dependency relationships between servers in the cluster environment. By obtaining these data, comprehensive input information can be provided for subsequent feature extraction and load prediction, so as to more accurately analyze the spatio-temporal dynamic changes of the server load and lay a data foundation for resource scheduling and task allocation.

[0053] A time step refers to dividing continuous time into discrete time units or intervals in time series analysis or prediction. Each time step represents a specific time point or time period and is used to capture the change pattern of data over time. In time series data, a time step is the basic unit for discretizing continuous time.

[0054] For example, if one hour is taken as a time step, then the time series data can be represented as data points per hour (such as the load data of the server per hour); in the embodiments of this application, the data of the past 5 minutes is extracted, and one minute is taken as a time step, resulting in 5 time steps. It should be noted that this description method is only an exemplary illustration and not a specific limitation on the specific number of time steps.

[0055] Extract the statistical features of the historical load data from the monitoring data of each of the time steps, and extract the network traffic and load dependencies of the adjacent servers to obtain the time features and space features of each of the time steps;

[0056] From the monitoring data of each time step, extract the statistical features of the historical load data (such as data of CPU, memory, network bandwidth, etc.) to capture the change pattern of the server load over time, and at the same time extract the network traffic data of adjacent servers and their load dependency relationships to reflect the synergy effect and spatial correlation of the server in the cluster environment.

[0057] This feature extraction method combining the time dimension and the space dimension can comprehensively reflect the spatio-temporal dynamic characteristics of the server load and provide more accurate and reliable data support for subsequent load prediction and resource scheduling.

[0058] The step of inputting the time features and the space features into a load prediction model to perform server load prediction and obtaining the load prediction result of the server at the target moment includes:

[0059] Input the preset number of time steps into the load prediction model.

[0060] Please continue to refer to Figure 2, input the multi-dimensional input parameters obtained daily into the LSTM model. Each time step contains the extracted temporal features and spatial features. Combine the temporal features and spatial features to form an input sequence. Assume that the input to the model is the load data of the past 5 time steps (5 is an example and can be multiple) and the network traffic data of adjacent servers. The input for each time step contains multiple features, in the following form:

[0061] X t = [CPU t , Memory t , NetworkSend t , NetworkErcv t ,

[0062] NetworkTraffic t-1 , NetworkTraffic t

[0063] Among them, cPU t is the CPU usage rate; Memory t is the server memory usage rate; NetworkSend t is the amount of network data sent by the server; NetworkRecv t is the amount of network data received by the server; NetworkTraffic t-1 is the network traffic data of adjacent servers at time step t-1; NetworkTraffic t is the network traffic data of adjacent servers at time step t; Substitute X t-4 , X t-3 , X t-2 , X t-1 , X t into the formula of the LSTM forward propagation. Specifically, it includes calculating the output of the gates: initializing model parameters such as weight matrices and bias terms, performing multi-gate operations, outputting the prediction results for the next moment, and updating the cell state Ct according to the LSTM formula by calculating the forget gate, input gate, output gate, and candidate cell state and combining temporal features and spatial features.

[0064] Forget gate f t :

[0065] f t = σ(W f [h t-1 , x t ) + b f

[0066] Among them, h t-1 is the hidden layer state of the previous moment, x t ​is the input feature at the current moment, W f , b f are the weights and bias terms of the forget gate, and σ is the sigmoid activation function, which is used to map the input to the range (0, 1), representing the degree of "forgetting". 0 means completely forgetting, and 1 means completely retaining. The output of the forget gate determines which information will be discarded at the current moment (i.e., which information will be forgotten from the cell state).

[0067] Input gate i t :

[0068] i t = σ(W i [h t-1 , x t ) + b i

[0069] where W i , b i are the weights and bias terms of the input gate. The output of the input gate determines which new information will be added to the cell state.

[0070] Candidate cell state

[0071]

[0072] where W C , b C are the weights and bias terms of the cell state, and tanh is the hyperbolic tangent function, which maps the output to the range (-1, 1) to ensure that the data does not exceed this range. The candidate cell state represents the "new information" that will be added to the current cell state.

[0073] Cell state C t :

[0074]

[0075] where C t-1 is the cell state at the previous moment. The cell state at the current moment represents the memory of the LSTM, which will contain important information for long-term storage.

[0076] Output gate o t :

[0077] o t = σ(W o [h t-1 , x t ) + b o

[0078] where W o , b oare the weights and bias terms of the output gate. The output of the output gate determines which information in the cell state at the current time will affect the final output (i.e., the hidden layer state).

[0079] Hidden layer state h t :

[0080] h t = o t · tanh(C t )

[0081] Calculate the final hidden layer state h t , through y pred = W y h t + b y Get the final predicted output, which contains the summary of all information at the current time.

[0082] Among them, y pred is the prediction result of the load prediction model; W y is the weight matrix of the output layer, and b y is the bias vector of the output layer.

[0083] In some embodiments, the input dimension of the load prediction model is the same as the hidden layer dimension. Since there are 6 features in the time step, the input dimension of the load prediction model is set to 6, and the hidden layer dimension is also set to 6.

[0084] By inputting these time step data into the load prediction model, the model can comprehensively analyze the load change trend of the time series and the cluster dynamics of the spatial dimension, so as to more accurately predict the server load status at the target time.

[0085] Figure 3 The following is a schematic flow chart of a server resource scheduling method provided by an embodiment of the present application. Please refer to Figure 3 , including:[[]]

[0086] Step 201, calculate the load status of the server through weighted summation based on the load prediction result and the weight of the task type of the server;

[0087] The load prediction value at the target time is the expected load level of the server at the target time. In some embodiments, if the time step is 1 minute, then the target time is the expected load level of the server 1 minute later; it should be noted that the target time varies with the time step, and this description method is not a specific limitation on the target time, and the embodiments of the present application do not limit this.

[0088] According to the demand characteristics of server resources for each task type (such as compute-intensive, storage-intensive, network-intensive), assign corresponding weights to each task type. For example: Compute-intensive tasks have a higher CPU weight; Storage-intensive tasks have higher memory and disk weights; Network-intensive tasks have a higher network bandwidth weight.

[0089] Multiply the predicted load value of each task type by its corresponding weight to obtain a weighted load value, and sum up the weighted load values of all task types to obtain the comprehensive load status of the server.

[0090] In some embodiments, calculating the load status of the server by weighted summation based on the predicted load data and the weights of the task types of the server includes:

[0091] Perform normalization processing on the predicted load data to eliminate the dimension difference;

[0092] Based on the normalized predicted load data and the weights of the task types, calculate the comprehensive load score of the server by weighted summation;

[0093] Based on a preset threshold and the comprehensive load score, determine the load status of the server.

[0094] Please continue to refer to Figure 2 , start the intelligent scheduling module, perform normalization processing on the prediction results, and perform normalization processing on the predicted load data obtained from the load prediction model. Specifically, define variables as follows. Please refer to Table 5. Table 5 is a schematic table of variable definitions provided in the embodiments of the present application:

[0095] Table 5

[0096]

[0097] Normalize the predicted server resource utilization rate of LSTM to eliminate the dimension difference:

[0098]

[0099] Where represents the predicted value of the jth resource of all servers, is the minimum resource prediction value, is the maximum resource prediction value.

[0100] According to the task type weights and the normalized resource values, calculate the comprehensive load score of the server:

[0101]

[0102] In some embodiments, since the predicted items of resources and the task types are different, the corresponding weights are also different. For example, the weight of a compute-intensive task: w task = [0.4, 0.3, 0.1, 0.1, 0.1, 0.0]; the weight of a storage-intensive task: w task = [0.1, 0.1, 0.3, 0.2, 0.2, 0.1]. It should be noted that this description is only an exemplary one and is a specific limitation on the weights of specific predicted items and task types.

[0103] Divide the server load status according to the score:

[0104]

[0105] In practical applications, the high-load threshold and the low-load threshold can be determined according to actual needs, and the embodiments of the present application do not limit this.

[0106] Step 202, dynamically adjust the resource allocation and scheduling strategy of the task according to the load status; wherein, the load status is inversely proportional to the allocation priority of the server.

[0107] Please continue to refer to Figure 2 , determine whether the task type is a compute task. When the task type is a compute task, calculate the load according to the weight of the compute-intensive task. When the task type is not a compute task, calculate the load according to the weight of the storage-intensive task. Specifically, when executing this step, it can be executed according to the following steps:

[0108] Based on the task type of the task to be allocated and the priority of the task to be allocated, allocate the task to be allocated to a server with the same processing task type and a low load status; wherein, the higher the priority of the task to be allocated, the lower the comprehensive load score of the allocated server; and / or

[0109] Based on the server performance requirements of the task to be allocated, allocate the task to be allocated to a server that meets the server performance requirements and has a low load status; wherein, the server performance requirements include computing power requirements and storage resource requirements; and / or

[0110] Based on the allocation priority of the server, allocate the task to be allocated according to the allocation priority; wherein, the allocation priority is calculated according to the remaining capacity of the server and the comprehensive load score.

[0111] According to the type of the task (compute-intensive, storage-intensive, etc.) and the priority of the task, combined with the load prediction result, select a server with a lower load for task allocation.

[0112] For compute-intensive tasks, servers with stronger computing power and lower load are preferentially selected; for storage-intensive tasks, servers with sufficient storage resources are preferentially selected.

[0113] Through the scheduling algorithm, the system ensures that the load of each server is as balanced as possible, thus avoiding the situation of overloading a certain server.

[0114] Based on the task type and the remaining capacity of the server, calculate the allocation priority:

[0115] P i,k = α·(1 - S i ) + β·C i

[0116] where α and β are weight coefficients (e.g., α = 0.7, β = 0.3), and α + β = 1 needs to be satisfied. The task is assigned to the server with the highest priority P i,k highest

[0117] In some embodiments, after calculating the load status of the server by weighted summation based on the predicted load data and the weights of the task types of the server, the following is further included:

[0118] When determining that the load status of the server is high load, in the server, determine the to-be-processed tasks whose migration cost is lower than the preset migration cost threshold; wherein, the migration cost is calculated based on the task resource occupancy and the migration time;

[0119] Please continue to refer to Figure 2 , start the real-time monitoring module. When the server load exceeds the preset threshold, obtain low-load servers that meet the conditions, and perform task migration to ensure load balance; when determining that the load status of the server is high load, that is, the load of the server exceeds the preset threshold, the system will automatically migrate some tasks to the servers with lower load. Screen out the to-be-processed tasks whose migration cost is lower than the preset migration cost threshold from this server. The specific method is as follows: First, calculate the migration cost of each task according to the task resource occupancy (such as CPU, memory, storage, etc.) and the migration time (such as data transfer time, task restart time). The calculation formula is as follows:

[0120] Migration task selection conditions:

[0121] Compare the calculated migration cost with the preset migration cost threshold, and screen out the tasks with migration cost lower than the threshold as the to-be-migrated tasks. Ensure that the selection of migration tasks is within an acceptable range in terms of both resource overhead and time cost.

[0122] Migrate the to-be-processed tasks to the adjacent servers with the low-load status.

[0123] Select a server with a low load status and sufficient resources from adjacent servers as the target server to ensure that the target server can accommodate the migration task without overloading. Through this process, the pressure on high-load servers can be effectively reduced, and at the same time, the idle resources of low-load servers can be fully utilized to achieve the balance of cluster load and the efficient utilization of resources.

[0124] In some embodiments, after inputting the time feature and the space feature into the load prediction model to predict the server load and obtaining the load prediction result at the target moment of the server, the method further includes:

[0125] Obtain the real load data at the target moment, and calculate the error between the predicted load data and the real load data based on the mean square error;

[0126] Please continue to refer to Figure 2 , calculate the loss of the prediction result according to the actual data, adjust the initial value. After obtaining the real load data at the target moment, calculate the error between the predicted load data and the real load data based on the mean square error (MSE). The specific method is as follows: First, obtain the real load data (such as CPU usage rate, memory usage rate, etc.) at the target moment from the server monitoring system as the benchmark for evaluating the prediction accuracy; then, compare the predicted load data with the real load data point by point and calculate the error of each sample:

[0127]

[0128] where sample error, y pred is the predicted value, y true is the real value; by calculating the mean square error, the deviation between the predicted load data and the real load data can be quantified, providing a clear error metric basis for subsequent model optimization.

[0129] Calculate the gradient information of each parameter in the load prediction model based on the error;

[0130] Update the parameters in the load prediction model based on the gradient information.

[0131] Starting from the output layer, calculate the partial derivative of the loss function with respect to each parameter such as the weight matrix and the bias term layer by layer, and use the chain rule to transfer the gradient information from the output layer to the input layer to ensure that the gradient information of each parameter accurately reflects its contribution to the error; then, update the parameters in the load prediction model based on the gradient information, use the gradient descent method or an optimizer (such as Adam) to adjust the parameter values, and control the step size of parameter update; continuously optimize the parameters of the load prediction model to gradually improve the prediction accuracy and generalization ability of the model.

[0132] In some embodiments, after obtaining the actual load data at the target moment and calculating the error between the predicted load data and the actual load data based on the mean square error, the method further includes:

[0133] When the error between the predicted load data and the actual load data is greater than a preset error threshold, re - execute the allocation of the task to be allocated based on the actual load data.

[0134] Compare the error between the predicted load data and the actual load data. If the error exceeds the preset threshold, it is determined that the prediction result is unreliable; then, re - evaluate the load status of the server based on the actual load data, calculate the comprehensive load score of the server by combining the task type and its weight, and determine the load status of the server according to the preset threshold; then, according to the re - evaluated load status, dynamically adjust the task allocation strategy, give priority to allocating tasks to servers with a low - load status, and at the same time avoid allocating tasks to high - load servers; finally, update the task allocation record and monitor the task execution situation to ensure the rationality of resource allocation and the efficiency of task execution. When the prediction error is large, timely correct the task allocation strategy to avoid resource waste or server overload problems caused by inaccurate prediction.

[0135] Please continue to refer to Figure 2 , evaluate the current performance and feedback for optimizing and adjusting the model. During the operation of the system, dynamic adjustment is the key to ensuring efficient resource utilization. The following is the process of dynamic adjustment:

[0136] Real - time monitoring: The system monitors the resource usage of each server in real - time, especially the load changes during peak load periods; compare the real - time load data with the prediction results of the LSTM model to check whether the load prediction is accurate.

[0137] Real - time feedback: If there is a large difference between the predicted load and the actual load, the system will trigger a feedback mechanism to adjust the resource allocation strategy in real - time; for example, if the load prediction of a certain server is low while the actual load suddenly increases, the system can timely adjust the task allocation strategy and migrate more tasks to other servers with lower loads.

[0138] Dynamic task scheduling:

[0139] During the task execution process, the system dynamically adjusts the task scheduling according to the real - time monitored load data. For example, if a certain server is about to reach the load limit, the system will migrate tasks to other servers in advance to avoid overload.

[0140] System optimization includes continuous optimization of aspects such as load prediction accuracy, task scheduling efficiency, and energy efficiency. The specific process is as follows:

[0141] Performance evaluation: Monitor the system performance by regularly evaluating indicators such as task execution time, resource utilization, and energy efficiency; Task completion time: Evaluate the execution duration of tasks to ensure that tasks can be completed on time. Resource utilization: Monitor the usage of resources such as CPU, memory, and network bandwidth of each server to ensure that resources are fully utilized; Energy efficiency evaluation: Evaluate the energy consumption during task execution to ensure energy efficiency optimization.

[0142] Prediction accuracy feedback: Compare real-time load data with the prediction results of the LSTM model to evaluate the prediction accuracy of the model. If there is a large deviation, the system will trigger a feedback mechanism to retrain the model to improve the accuracy.

[0143] Scheduling optimization: Optimize the task scheduling strategy according to the performance evaluation results. For example, adjust the load balancing algorithm, task priority strategy, etc. to further improve resource utilization and task execution efficiency.

[0144] Model optimization: Periodically retrain the LSTM model with new historical data to improve the accuracy of load prediction; Adjust the model input features (such as introducing new spatio-temporal features or enhancing existing features) to improve the prediction performance.

[0145] Corresponding to the above server resource scheduling method, the present invention also proposes a server resource scheduling device. Since the device embodiment of the present invention corresponds to the above method embodiment, for the details not disclosed in the device embodiment, reference may be made to the above method embodiment, and the present invention will not be elaborated herein.

[0146] Figure 4 FIG. is a schematic structural diagram of a server resource scheduling device provided by an embodiment of the present application, as Figure 4 shown, including:

[0147] A feature extraction unit 31, configured to extract features from the monitoring data of the server to obtain time features and spatial features; wherein, the time features include statistical features of historical load data, and the spatial features include load correlation features of adjacent servers of the server;

[0148] A prediction unit 32, configured to input the time features and the spatial features into a load prediction model to perform server load prediction, and obtain a load prediction result of the server at a target moment;

[0149] An adjustment unit 33, configured to dynamically adjust the resource allocation and scheduling strategy of the task based on the load prediction result and the server task type.

[0150] Further, in a possible implementation manner of an embodiment of the present disclosure, as Figure 5 shown, the feature extraction unit 31 is further configured to:

[0151] Obtain the monitoring data corresponding to a preset number of time steps in a preset database; wherein, the monitoring data includes the load data of the server and the network traffic data of the adjacent server;

[0152] Extract the statistical features of the historical load data from the monitoring data at each of the time steps, and extract the network traffic and load dependencies of the adjacent servers, to obtain the time features and space features at each of the time steps;

[0153] The inputting the time features and the space features into a load prediction model to perform server load prediction, and obtaining the load prediction result at the target moment of the server includes:

[0154] Input the preset number of time steps into the load prediction model.

[0155] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 5 shown, the apparatus further includes:

[0156] A first calculation unit 34, configured to, before an adjustment unit 33 dynamically adjusts the resource allocation and scheduling policy of a task based on the load prediction result and the server task type, calculate the load status of the server by weighted summation based on the load prediction result and the weight of the server task type;

[0157] The adjustment unit 33 is further configured to:

[0158] Dynamically adjust the resource allocation and scheduling policy of the task according to the load status; wherein, the load status is inversely proportional to the allocation priority of the server.

[0159] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 5 shown, the first calculation unit 34 is further configured to:

[0160] Perform normalization processing on the predicted load data to eliminate the dimension difference;

[0161] Based on the predicted load data after normalization processing and the weight of the task type, calculate the comprehensive load score of the server by weighted summation;

[0162] Based on a preset threshold and the comprehensive load score, determine the load status of the server.

[0163] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 5 shown, the adjustment unit 33 is further configured to:

[0164] Based on the task type of the task to be assigned and the priority of the task to be assigned, assign the task to be assigned to a server with the same processing task type and a low load status; wherein, the higher the priority of the task to be assigned, the lower the comprehensive load score of the assigned server; and / or

[0165] Based on the server performance requirements of the task to be assigned, assign the task to be assigned to a server that meets the server performance requirements and has a low load status; wherein, the server performance requirements include computing power requirements and storage resource requirements; and / or

[0166] Based on the allocation priority of the server, assign the task to be assigned according to the allocation priority; wherein, the allocation priority is calculated based on the remaining capacity of the server and the comprehensive load score.

[0167] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 5 shown, the device further includes:

[0168] A determination unit 35, configured to, after a first calculation unit 34 calculates the load status of the server by weighted summation based on the predicted load data and the weight of the task type of the server, when determining that the load status of the server is a high load, determine, in the server, a task to be processed whose migration cost is lower than a preset migration cost threshold; wherein, the migration cost is calculated based on task resource occupancy and migration time;

[0169] A migration unit 36, configured to migrate the task to be processed to an adjacent server with a low load status.

[0170] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 5 shown, the device further includes:

[0171] An acquisition unit 37, configured to, after a prediction unit 32 inputs the time feature and the space feature into a load prediction model for server load prediction to obtain a load prediction result of the server at a target moment, acquire the real load data at the target moment, and calculate the error between the predicted load data and the real load data based on the mean square error;

[0172] A second calculation unit 38, configured to calculate the gradient information of each parameter in the load prediction model based on the error;

[0173] An update unit 39, configured to update the parameters in the load prediction model based on the gradient information.

[0174] Further, in a possible implementation manner of the embodiments of the present disclosure, asFigure 5 As shown, the device further includes:

[0175] An adjustment unit 33 is further configured to, after the acquisition unit 37 acquires the actual load data at the target moment and calculates the error between the predicted load data and the actual load data based on the mean square error, when the error between the predicted load data and the actual load data is greater than a preset error threshold, re - execute the allocation of the task to be allocated based on the actual load data.

[0176] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 5 shown, the input dimension of the load prediction model is the same as the dimension of the hidden layer.

[0177] For the description of the features in the corresponding embodiments of the resource scheduling device of the server, reference can be made to the relevant description in the corresponding embodiments of the resource scheduling method of the server, which will not be elaborated here one by one.

[0178] The embodiments of the present application further provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above - mentioned embodiments of the resource scheduling method of the server.

[0179] The embodiments of the present application further provide a computer - readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above - mentioned embodiments of the resource scheduling method of the server when running.

[0180] In an exemplary embodiment, the above - mentioned computer - readable storage medium may include, but is not limited to: USB flash drive, read - only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk, or optical disc and other various media that can store computer programs.

[0181] The embodiments of the present application further provide a computer program product. The above - mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above - mentioned embodiments of the resource scheduling method of the server.

[0182] The embodiments of the present application further provide another computer program product, including a non - volatile computer - readable storage medium. The non - volatile computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above - mentioned embodiments of the resource scheduling method of the server.

[0183] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0184] The above has introduced in detail a resource scheduling method and an electronic device of a server provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A resource scheduling method for a server, characterized in that Including: Performing feature extraction on the monitoring data of the server to obtain time features and spatial features; wherein, the time features include statistical features of historical load data, and the spatial features include load correlation features of adjacent servers of the server; Inputting the time features and the spatial features into a load prediction model to perform server load prediction, and obtaining a load prediction result at a target time of the server; Dynamically adjusting the resource allocation and scheduling strategy of tasks based on the load prediction result and the server task type.

2. The resource scheduling method of the server according to claim 1, wherein The performing feature extraction on the monitoring data of the server to obtain time features and spatial features includes: Obtaining the monitoring data corresponding to a preset number of time steps in a preset database; wherein, the monitoring data includes load data of the server and network traffic data of adjacent servers; Extracting statistical features of the historical load data and extracting network traffic and load dependencies of adjacent servers from the monitoring data at each time step, to obtain time features and spatial features at each time step; The inputting the time features and the spatial features into a load prediction model to perform server load prediction, and obtaining a load prediction result at a target time of the server includes: Inputting the preset number of time steps into the load prediction model.

3. The resource scheduling method of the server according to claim 1, characterized in that, Before dynamically adjusting the resource allocation and scheduling strategy of tasks based on the load prediction result and the server task type, the method further includes: Calculating the load status of the server by weighted summation based on the load prediction result and the weight of the server task type; The dynamically adjusting the resource allocation and scheduling strategy of tasks based on the load prediction result and the server task type includes: Dynamically adjusting the resource allocation and scheduling strategy of tasks according to the load status; wherein, the load status is inversely proportional to the allocation priority of the server.

4. The resource scheduling method of the server according to claim 3, wherein The calculating the load status of the server by weighted summation based on the predicted load data and the weight of the server task type includes: Performing normalization processing on the predicted load data to eliminate the dimension difference; Based on the predicted load data after normalization processing and the weight of the task type, calculating the comprehensive load score of the server by weighted summation; Determining the load status of the server based on a preset threshold and the comprehensive load score.

5. The resource scheduling method of the server according to claim 4, wherein, The dynamically adjusting the resource allocation and scheduling strategy of tasks according to the load status further includes: Based on the task type of the task to be allocated and the priority of the task to be allocated, allocating the task to be allocated to a server with the same processing task type and a low load status; wherein, the higher the priority of the task to be allocated, the lower the comprehensive load score of the allocated server; and / or Based on the server performance requirements of the task to be allocated, allocating the task to be allocated to a server that meets the server performance requirements and has a low load status; wherein, the server performance requirements include computing power requirements and storage resource requirements; and / or Based on the allocation priority of the server, allocate the task to be allocated according to the allocation priority; wherein, the allocation priority is calculated based on the remaining capacity of the server and the comprehensive load score.

6. The resource scheduling method of the server according to claim 3, characterized in that, After calculating the load status of the server by weighted summation based on the predicted load data and the weight of the task type of the server, the method further includes: When determining that the load status of the server is high load, in the server, determine the task to be processed whose migration cost is lower than the preset migration cost threshold; wherein, the migration cost is calculated based on the task resource occupancy and the migration time; Migrate the task to be processed to an adjacent server with a low load status.

7. The resource scheduling method of the server according to claim 1, characterized in that After inputting the time feature and the space feature into the load prediction model for server load prediction to obtain the load prediction result of the server at the target moment, the method further includes: Obtain the real load data at the target moment, and calculate the error between the predicted load data and the real load data based on the mean square error; Calculate the gradient information of each parameter in the load prediction model based on the error; Update the parameters in the load prediction model based on the gradient information.

8. The resource scheduling method of the server according to claim 7, characterized in that, After obtaining the real load data at the target moment and calculating the error between the predicted load data and the real load data based on the mean square error, the method further includes: When the error between the predicted load data and the real load data is greater than the preset error threshold, re - execute the allocation of the task to be allocated based on the real load data.

9. The resource scheduling method of the server according to any one of claims 1-8, characterized in that The input dimension of the load prediction model is the same as the dimension of the hidden layer.

10. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the resource scheduling method of the server according to any one of claims 1 to 9 when executing the computer program.

Citation Information

Cited By

  • Task scheduling method and device, data processing unit, program product and medium

    CN120762868A

  • Task scheduling methods, devices, data processing units, program products, media

    CN120762868B

  • Server resource scheduling system for high-density computing environment

    CN120872610A

  • Software testing method

    CN120929386A

  • Load prediction-based server management method, program product and device

    CN120950342A