Computing Power Resource Scheduling Method, System and Storage Medium
By using neural network models in the computing power cluster for computing power prediction and pattern judgment, intelligent scheduling of computing power resources is realized, the problem of uneven resource allocation in the computing power cluster is solved, and the success rate and node utilization rate of computing tasks are improved.
Patent Information
- Application Number
- CN202510242550.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-03
AI Technical Summary
It is difficult to intelligently rationalize the computing power resource scheduling in existing computing power clusters, resulting in a low success rate of computing tasks.
By obtaining real-time data and historical data of computing power nodes, using neural network models to predict computing power, judging working modes and scheduling resources, ensuring the matching degree between nodes and tasks and the rationality of resource allocation.
It improves the success rate of computing tasks, avoids waste of computing resources, and improves the utilization rate and computing efficiency of nodes.
Smart Images

Figure CN119739535B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing power scheduling, and particularly to a computing power resource scheduling method, system and storage medium. Background Art
[0002] With the development of technologies such as big data, machine learning, and deep learning, technologies such as cloud computing and edge computing have also flourished. Cloud computing platforms are generally built in a cluster manner, connecting multiple computing power nodes together to form a huge computing power cluster. In a computing power cluster, the computing capabilities of each computing power node may vary. Therefore, deploying different computing tasks to different computing power nodes can ensure the computing efficiency of the tasks. In addition, due to the large number of computing power nodes in the computing power cluster and the continuous change of computing tasks, the computing power resources of the computing power nodes need to be reasonably configured and scheduled to fully utilize the computing power resources, improve the computing efficiency, and reduce the computing energy consumption at the same time. Summary of the Invention
[0003] The main purpose of the present invention is to provide a computing power resource scheduling method, system and storage medium, aiming to solve the technical problem that it is difficult to intelligently and reasonably schedule computing power resources in the case of a large number of nodes and diverse computing tasks in the existing computing power cluster, thereby reducing the success rate of computing tasks.
[0004] The first aspect of the present invention provides a computing power resource scheduling method, and the computing power resource scheduling method includes:
[0005] Obtain the real-time computing power data of each computing power node and the computing tasks to be allocated with computing power resources;
[0006] Based on the real-time computing power data of each computing power node and the computing tasks to be allocated with computing power resources, perform computing power node matching to select target computing power nodes;
[0007] Obtain the historical computing power data of each target computing power node, and input the real-time computing power data and historical computing power data of each target computing power node into a pre-set neural network model for computing power prediction in turn to obtain the predicted real-time computing power data of each target computing power node;
[0008] Based on the real-time computing power data and predicted real-time computing power data of each target computing power node, respectively judge the working modes of each computing power node, and the working modes include a stable working mode and / or a fluctuating working mode;
[0009] If the target computing power node is in the stable working mode, perform computing power resource scheduling according to the historical computing power data of the target computing power node;
[0010] If the target computing power node is in the fluctuating working mode, perform computing power resource scheduling according to the real-time computing power data and predicted real-time computing power data of the target computing power node.
[0011] In a second aspect of the present invention, a computing power resource scheduling system is provided. The computing power resource scheduling system includes:
[0012] An acquisition module, configured to acquire the real-time computing power data of each computing power node and the computing tasks of the computing power resources to be allocated;
[0013] A matching module, configured to perform matching of computing power nodes based on the real-time computing power data of each computing power node and the computing tasks of the computing power resources to be allocated, so as to select target computing power nodes;
[0014] A prediction module, configured to acquire the historical computing power data of each target computing power node, and sequentially input the real-time computing power data and the historical computing power data of each target computing power node into a preset neural network model for computing power prediction, so as to obtain the predicted real-time computing power data of each target computing power node;
[0015] A judgment module, configured to respectively judge the working modes of each computing power node based on the real-time computing power data and the predicted real-time computing power data of each target computing power node, where the working modes include a stable working mode and / or a fluctuating working mode;
[0016] A scheduling module, configured to, if the target computing power node is in the stable working mode, perform computing power resource scheduling according to the historical computing power data of the target computing power node; if the target computing power node is in the fluctuating working mode, perform computing power resource scheduling according to the real-time computing power data and the predicted real-time computing power data of the target computing power node.
[0017] In a third aspect of the present invention, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the computing power resource scheduling method described in any one of the above.
[0018] In the technical solution provided by the present invention, before performing computing power resource scheduling, first match the computing power characteristics of the computing power nodes with the task characteristics of the computing tasks for which the computing power resources are to be allocated, so as to screen out the target computing power nodes that meet the requirements of the computing tasks and ensure the adaptability between the computing power nodes and the computing tasks to be calculated; then further use the neural network model of deep learning to predict the computing power of the computing power nodes, ensure that the predicted computing power data is closer to the actual computing power data, and then improve the prediction accuracy and the accuracy of computing power node resource scheduling. Then, based on the real-time computing power data and the predicted real-time computing power data of each target computing power node, determine the working mode of each target computing power node and allocate resources based on different working modes. The present invention grasps the fluctuation state of the computing power of the nodes in real time according to the node computing power prediction, allocates resources according to the node computing power, avoids the waste of the computing power resources of the computing power nodes, makes the allocation of the computing power resources of the computing power nodes more reasonable, and avoids the load imbalance caused by uneven resource allocation. This is not only conducive to improving the success rate of computing tasks, but also further improves the utilization rate of node computing power resources while ensuring that the computing power nodes are fully utilized. Description of the Drawings
[0019] Figure 1 It is a schematic diagram of an embodiment of the computing power resource scheduling method in an embodiment of the present invention. Detailed Embodiments
[0020] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0021] For ease of understanding, the specific process of the embodiment of the present invention will be described below. Please refer to Figure 1 An embodiment of the computing power resource scheduling method in an embodiment of the present invention includes:
[0022] 101. Obtain the real-time computing power data of each computing power node and the computing tasks for which the computing power resources are to be allocated;
[0023] The scheduling system needs to first collect the real-time computing power data of all computing power nodes. The computing power data includes, but is not limited to, CPU usage rate, memory occupancy rate, disk I / O speed, network bandwidth, etc. At the same time, the scheduling system also needs to obtain the computing tasks for which computing power resources are to be allocated currently. These tasks usually include information such as the computing volume of the task, the required resource type, and the computing priority.
[0024] The collection of computing power node data can obtain the computing power data in real time through the monitoring software or API interface deployed on the computing power nodes. The monitoring software or API interface should be able to send the computing power status information of the nodes to the central scheduling system regularly (such as every second or every minute). The computing tasks are usually submitted to the scheduling system by users or application programs. The scheduling system should provide a user-friendly interface or API that allows users to submit computing tasks and includes the detailed description of the tasks and the required resources.
[0025] 102. Match the computing power nodes based on the real-time computing power data of each computing power node and the computing tasks for which computing power resources are to be allocated, so as to select the target computing power nodes;
[0026] In this embodiment, the system needs to analyze the specific computing power requirements of each computing task, including the required CPU, memory, disk, and network resources, etc. At the same time, a suitable matching algorithm is adopted to compare the task requirements with the real-time computing power data of the nodes. Such as the greedy algorithm, heuristic algorithm, or machine learning algorithm. Then, according to the matching results, one or more computing power nodes that are most suitable for executing the current task are selected as the target computing power nodes.
[0027] In an alternative embodiment, the above step 102 further includes:
[0028] 1021. According to the node running time in the real-time computing power data of each computing power node and the task execution time of the computing tasks for which computing power resources are to be allocated, determine whether the computing power node matching condition is satisfied;
[0029] Collect the running time data of each computing power node in real time, which usually includes the node start running time, the estimated end running time (or the current time if the node runs continuously), etc. At the same time, obtain the task execution time of the computing tasks to be allocated from the task queue, including the task start time and the task end time. For each computing power node, check whether its running time interval (start time to end time) contains the task execution time interval (start time to end time) of the task to be allocated. If the running time interval of the node completely covers or partially covers the time interval of the task (that is, the task start time is not earlier than the node start time, and the task end time is not later than the node end time), it is initially considered that the node is available in terms of time.
[0030] Based on the above comparison results, all available computing power nodes in terms of time are selected as the candidate node set for the subsequent steps. If the candidate node set is empty, it indicates that there are currently no computing power nodes that can meet the time requirements of the task, and the task execution time needs to be readjusted or wait for more nodes to become available. For continuously running nodes, their end time can be regarded as a very large value (such as a future time point), or simply not set an end time.
[0031] 1022. If the node running time of the computing power node is earlier than the task execution time of the computing task for which computing power resources are to be allocated, then determine whether the task execution time of the computing task for which computing power resources are to be allocated is between the low - valley running time and the peak running time of the current computing power node;
[0032] Obtain its low - valley running time and peak running time from the real - time data of the computing power node. These times are usually determined based on factors such as the node's historical running data, load conditions, user behavior patterns, etc. The low - valley running time refers to the time period when the node load is low and the computing power utilization rate is low; the peak running time is the time period when the node load is high and the computing power utilization rate is high. For each candidate node selected in the time matching step, check whether the task execution time of the task to be allocated is between its low - valley running time and peak running time.
[0033] If the time interval of the task intersects with the low - valley - to - peak time interval of the node, it is considered that the task execution time is within this interval. Based on the above comparison results, further select those computing power nodes that meet the requirements both in terms of time and load. If the task execution time of a certain node is not within its low - valley - to - peak time interval, remove it from the candidate node set.
[0034] Considering that the low - valley and peak times of the node may change over time (such as daily or weekly periodic changes), the system needs to be able to dynamically update this time information. Machine learning algorithms or statistical methods can be used to predict future low - valley and peak times and adjust the matching strategy accordingly.
[0035] In actual implementation, the determination of low - valley and peak times may be a complex process that needs to consider multiple factors. For some special types of tasks (such as emergency tasks, high - priority tasks), it may be necessary to allocate computing power resources even during peak times.
[0036] 1023. If the task execution time of the computing task for which computing power resources are to be allocated is between the low - valley running time and the peak running time of the current computing power node, then determine the current computing power node as a candidate computing power node;
[0037] Determine the final set of candidate nodes based on the screening results of the previous two steps. These nodes cover the task execution time in terms of time, and the task execution time is between its low - peak running time and peak running time. If the set of candidate nodes contains multiple nodes and there are differences in computing power characteristics among these nodes (such as the number of CPUs, memory size, network bandwidth, etc.), the candidate nodes can be sorted according to these characteristics. The basis for sorting can be factors such as the computing power utilization rate of the nodes, historical task completion status, user evaluations, etc.
[0038] 1024. Perform feature matching between the computing power characteristics of each candidate computing power node and the task characteristics of the computing task to be allocated, and select the candidate computing power node with the highest feature matching degree as the target computing power node.
[0039] For each candidate node, extract its computing power characteristics, such as CPU model, memory size, network bandwidth, storage capacity, computing power utilization rate, etc. For the computing task to be allocated, extract its task characteristics, such as the amount of computation, memory requirements, IO requirements, network requirements, task type (such as CPU - intensive, IO - intensive, memory - intensive, etc.).
[0040] Design a feature matching algorithm to calculate the matching degree between the candidate node characteristics and task characteristics. The algorithm can be implemented based on methods such as weighted summation, distance metrics (such as Euclidean distance, Manhattan distance), similarity metrics (such as cosine similarity, Jaccard similarity), etc. In the weighted summation method, each feature can be assigned a weight indicating its importance in the matching. The weights can be adjusted according to actual needs. For each candidate node, use the feature matching algorithm to calculate its matching degree with the task. Sort the candidate nodes according to the matching degree, and select the node with the highest matching degree as the target computing power node. Consider using machine - learning algorithms to optimize the feature matching process and improve the matching accuracy and efficiency.
[0041] 103. Obtain the historical computing power data of each target computing power node, and sequentially input the real - time computing power data and historical computing power data of each target computing power node into a pre - set neural network model for computing power prediction to obtain the predicted real - time computing power data of each target computing power node;
[0042] In this embodiment, first obtain the computing power data of the target computing power node from the historical database. The obtained computing power data should cover a long enough time period to reflect the computing power change trend of the node. At the same time, it is also necessary to select a suitable neural network model, such as a long - short - term memory network or a gated recurrent unit, and input the historical computing power data into the neural network model for training until the model reaches a satisfactory prediction accuracy. During the training process, methods such as cross - validation can be used to prevent overfitting.
[0043] After completing the model training, the real-time computing power data and historical computing power data of the target computing power nodes can be input into the trained neural network model to obtain the predicted future computing power data. The predicted data will be used for subsequent working mode judgment and resource scheduling decision-making.
[0044] In an optional embodiment, the neural network model is obtained by using the following training method:
[0045] Obtain the historical real-time computing power data and historical predicted computing power data of each computing power node. Among them, the historical computing power data of each period includes the error characteristic parameters of the actual computing power value and the predicted computing power value corresponding to the time stamp;
[0046] Construct an initial neural network model. The input layer of the initial neural network model inputs the historical real-time computing power data and historical predicted computing power data, and the output layer outputs the currently predicted real-time computing power data predicted by the computing power node;
[0047] According to the historical real-time computing power data and historical predicted computing power data of each computing power node, construct the loss function of the initial neural network model, and train the initial neural network model to obtain the trained neural network model;
[0048] Among them, the input of the hidden layer of the initial neural network model is as follows:
[0049]
[0050] Among them, represents the historical real-time computing power data vector of the kth input of the computing power node i in the hidden layer, represents the historical predicted computing power data vector of the kth input of the computing power node i in the hidden layer, is the historical real-time computing power data of the kth input of the computing power node i, is the input weight of the historical real-time computing power data of the kth input of the computing power node i, is the bias term of the computing power node i, is the input weight coefficient of the kth input, N is the number of neurons in the input layer, f is the activation function, is the historical predicted computing power data of the kth input of the computing power node i, is the input weight of the historical predicted computing power data of the kth input of the computing power node i, is the bias term of the computing power node i, is the input weight coefficient of the kth input, M is the number of neurons in the input layer;
[0051] Among them, the loss function of the initial neural network model is as follows:
[0052]
[0053] Among them, represents the loss function of the initial neural network model, is the computing power data distribution matrix between computing power node i and computing power node j in the historical time period, S is the number of computing power node samples, is the real-time computing power data of computing power node j, represents the predicted computing power data of computing power node j predicted by the model.
[0054] 104. Based on the real-time computing power data and predicted real-time computing power data of each target computing power node, respectively determine the working mode of each computing power node, and the working mode includes a stable working mode and / or a fluctuating working mode;
[0055] In this embodiment, the stable working mode means that the computing power data of the computing power node remains relatively stable within a period of time with a small fluctuation range; the fluctuating working mode means that the computing power data of the node changes frequently with a large fluctuation range. Specifically, statistical methods (such as variance, standard deviation, etc.) or machine learning algorithms (such as clustering algorithms) can be used to determine the working mode of the node, and finally the determination result is stored in the database for subsequent resource scheduling decisions.
[0056] 105. If the target computing power node is in the stable working mode, perform computing power resource scheduling according to the historical computing power data of the target computing power node;
[0057] When the target computing power node is in the stable working mode, it indicates that the computing power resources of the current target computing power node are relatively stable, and resource scheduling can be performed based on the historical computing power data.
[0058] In this embodiment, in the stable working mode, a relatively simple resource allocation strategy can be adopted, such as resource allocation based on historical average computing power. Specifically, the average computing power of the node can be calculated according to its historical computing power data, and then resources are allocated to those nodes whose average computing power meets the task requirements according to the task requirements. According to the resource allocation strategy, a scheduling instruction is sent to the target computing power node to instruct it to start executing the computing task. The scheduling instruction should include detailed information about the task, the required resource amount, and the execution time, etc. In addition, during the execution process, the system should continuously monitor the computing power data of the node. If it is found that the computing power of the target computing power node changes significantly, it is necessary to re-evaluate its working mode and adjust the resource allocation strategy.
[0059] 106. If the target computing power node is in the fluctuating working mode, perform computing power resource scheduling according to the real-time computing power data and predicted real-time computing power data of the target computing power node.
[0060] When the target computing power node is in the fluctuating working mode, it indicates that the computing power resources of the current target computing power node change frequently, and more flexible resource scheduling needs to be performed based on real-time computing power data and predicted computing power data.
[0061] In this embodiment, in the fluctuating working mode, a more dynamic resource allocation strategy needs to be adopted. This can be achieved by monitoring the computing power data of the node in real time and dynamically adjusting the resource allocation according to the prediction results. For example, a reinforcement learning algorithm can be used to optimize the resource allocation strategy to minimize resource waste while meeting the task requirements. Since the computing power of the node fluctuates greatly, it may be necessary to adjust the resource allocation frequently. Therefore, the system should support elastic scheduling, that is, it can dynamically increase or decrease the amount of resources allocated to the node according to real-time computing power data and prediction results.
[0062] In addition, in this embodiment, when the target computing power node is in the fluctuating working mode, task execution may fail or be delayed due to computing power fluctuations. Therefore, the system should have risk management and fault tolerance mechanisms, such as task retry, resource backup, etc., to ensure the smooth execution of tasks. At the same time, the system should send scheduling instructions to the target computing power node and execute computing tasks. The system should continuously monitor the computing power data of the node and the task execution situation to detect and handle potential problems in a timely manner.
[0063] In an alternative embodiment, step 106 above further includes:
[0064] 1061. If the target computing power node is in the fluctuating working mode, compare the numerical difference between the real-time computing power data of the target computing power node and the predicted real-time computing power data in real time;
[0065] In this embodiment, once it is confirmed that the node is in the fluctuating working mode, the system needs to obtain the computing power data of the node in real time. This can be achieved through API calls, log analysis, or directly collecting data from the node. The real-time computing power data may include indicators such as CPU usage rate, memory occupancy rate, and GPU computing power. At the same time, the system also needs to obtain the predicted real-time computing power data for the node, and then compare the numerical difference between the real-time computing power data and the predicted real-time computing power data in real time. This is completed by calculating the absolute difference or relative difference between the two. The magnitude of the difference reflects the deviation degree between the actual computing power and the predicted computing power.
[0066] 1062. If the numerical difference exceeds the preset difference range, determine that there is a computing power anomaly in the target computing power node currently, and increase the number of anomaly occurrences of the target computing power node by one;
[0067] The system needs to set a preset difference range for each computing power metric. This range can be determined based on the historical performance of the node, business requirements, or industry standards. The difference range can be a fixed value or a percentage based on the predicted value. The system compares the numerical difference between the real-time computing power data and the predicted computing power data to check if it exceeds the preset difference range. If it is found that the numerical difference exceeds the preset range, it is determined that there is a computing power anomaly in the target computing power node currently. The system needs to maintain a data structure (such as a database table or in-memory data structure) that records the number of anomalies and increment the number of anomalies of the target computing power node by one.
[0068] 1063. When the cumulative number of anomalies of the target computing power node is within the first preset range, it is determined that the target computing power node is in a regular fluctuation state. When the cumulative number of anomalies of the target computing power node is within the second preset range, it is determined that the target computing power node is in an irregular fluctuation mode.
[0069] The system needs to set the preset ranges of the cumulative number of anomalies for the regular fluctuation state and the irregular fluctuation mode. These ranges can be determined based on the historical performance of the node, business requirements, or industry standards. For example, the first preset range may be 0 - 5 times, and the second preset range may be greater than 6 times. The system checks the cumulative number of anomalies of the target computing power node to determine if it is within the preset range. This can be achieved by querying the data structure of the number of anomalies maintained in step 1062. According to the range where the cumulative number of anomalies is located, the system determines whether the target computing power node is in a regular fluctuation state or an irregular fluctuation mode.
[0070] 1064. If the target computing power node is in a regular fluctuation state, obtain the computing power fluctuation threshold of the target computing power node and perform computing power resource scheduling according to the computing power fluctuation threshold.
[0071] The system needs to collect the historical computing power data of the target computing power node, including the real-time computing power data, predicted computing power data, and the differences between them over a period of time. Based on the historical computing power data, the system uses statistical analysis methods (such as mean, standard deviation, percentile, etc.) to determine the computing power fluctuation threshold in advance before performing computing power resource scheduling. This threshold reflects the normal fluctuation range of the computing power of the node in the regular fluctuation state. Once the computing power fluctuation threshold is determined, the system can dynamically adjust the computing power resources according to this threshold. For example, when the real-time computing power data approaches or exceeds the threshold, the system can trigger resource expansion operations (such as increasing the number of CPU cores, allocating more memory, etc.) to ensure that the computing power requirements of the node are met.
[0072] 1065. If the target computing power node is in an irregular fluctuation state, obtain the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node.
[0073] The system needs to obtain the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node in real time. This can be achieved through API calls, log analysis, or directly collecting data from the node. Similar to step 1061, but here it is a specific acquisition operation for the non-regular fluctuation state. After obtaining the data, the system should perform necessary data verification to ensure the accuracy and integrity of the data. For example, it can check whether the data is missing, whether it exceeds a reasonable range, etc.
[0074] 1066. Calculate the difference between the latest real-time computing power data and the latest predicted real-time computing power data to obtain the computing power fluctuation range of the target computing power node;
[0075] First, preprocess the latest real-time computing power data and the latest predicted real-time computing power data to ensure that the data formats are consistent and the timestamps are aligned. Then, calculate the difference between the real-time computing power data and the predicted computing power data at each time point to obtain a computing power fluctuation difference sequence. The difference calculation formula is: computing power fluctuation difference = real-time computing power data - predicted computing power data.
[0076] Perform statistical analysis on the computing power fluctuation difference sequence, and calculate statistical quantities such as the average value, maximum value, and minimum value of its absolute value. These statistical quantities are used as measurement indicators for the computing power fluctuation range. Of course, more complex measurement methods can also be used, such as calculating the variance and standard deviation of the fluctuation difference, to more comprehensively reflect the amplitude and stability of the computing power fluctuation. Store the calculated computing power fluctuation range in the database for query and analysis in subsequent steps.
[0077] 1067. Based on the computing power fluctuation range of the target computing power node, determine whether the target computing power node has computing power loss and the computing power prediction accuracy of the target computing power node in the fluctuation working mode;
[0078] Preset a threshold for computing power loss in advance. This threshold can be set according to the actual needs of the system. For example, it can be that the real-time computing power data is lower than a certain preset value, or the difference between the real-time computing power data and the predicted computing power data exceeds a certain range. If the real-time computing power data of the target computing power node is lower than the computing power loss threshold, or the computing power fluctuation range exceeds the set fluctuation range, it is considered that the node has computing power loss.
[0079] For the target computing power node in the fluctuation working mode, it is necessary to evaluate the accuracy of its computing power prediction. Multiple evaluation indicators can be used to measure the prediction accuracy, such as mean square error, mean absolute error, root mean square error, etc. When calculating these evaluation indicators, the predicted computing power data needs to be compared with the actual computing power data. According to the values of the evaluation indicators, the prediction accuracy can be divided into different levels such as high, medium, and low. Record the results of the computing power loss judgment and the prediction accuracy judgment to provide a basis for decision-making in subsequent steps.
[0080] In an alternative embodiment, step 1067 further includes:
[0081] S101. If the value of the computing power fluctuation range of the target computing power node is negative, it is determined that the target computing power node has a computing power shortage. If the value of the computing power fluctuation range of the target computing power node is positive, it is determined that the target computing power node has no computing power shortage.
[0082] Computing power fluctuation range = (current computing power - expected computing power) / expected computing power * 100%, where the current computing power is the actual computing power of the target computing power node at a certain moment, and the expected computing power is the computing power that the target computing power node should reach at that moment.
[0083] Obtain the current computing power and expected computing power of the target computing power node through a monitoring or logging system, and calculate the computing power fluctuation range. If the computing power fluctuation range < 0, it is determined that the target computing power node has a computing power shortage. If the computing power fluctuation range >= 0, it is determined that the target computing power node has no computing power shortage.
[0084] S102. Obtain the maximum historical computing power fluctuation range among the historical computing power fluctuation ranges of the target computing power node in the fluctuating working mode.
[0085] S103. When the difference between the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node exceeds the deviation threshold range, calculate the difference between the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node to obtain the real-time computing power fluctuation range of the target computing power node.
[0086] In this embodiment, a deviation threshold is preset in advance to determine whether the difference between the latest real-time computing power data and the latest predicted real-time computing power data exceeds the acceptable range.
[0087] Obtain the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node through a monitoring or logging system. Calculate the difference between the latest real-time computing power data and the latest predicted real-time computing power data. If the difference exceeds the deviation threshold, proceed to the next step to calculate the real-time computing power fluctuation range; otherwise, do not perform the calculation of this step.
[0088] Real-time computing power fluctuation range = (latest real-time computing power data - latest predicted real-time computing power data) / latest predicted real-time computing power data * 100%
[0089] S104. Calculate the ratio of the real-time computing power fluctuation range of the target computing power node to the maximum historical computing power fluctuation range. If the ratio is greater than or equal to the preset ratio threshold, it is determined that the computing power prediction accuracy of the target computing power node is at a high level. If the ratio is less than the ratio threshold, it is determined that the computing power prediction accuracy of the target computing power node is at a low level.
[0090] Set a preset ratio threshold for judging the accuracy level of computing power prediction. Obtain the real-time computing power fluctuation range of the target computing power node through the previous steps. Obtain the maximum historical computing power fluctuation range of the target computing power node through the previous steps.
[0091] Ratio = Real-time computing power fluctuation range / Maximum historical computing power fluctuation range
[0092] Judge the accuracy level of computing power prediction:
[0093] If the ratio >= the preset ratio threshold, it is judged as a high level.
[0094] If the ratio < the preset ratio threshold, it is judged as a low level.
[0095] 1068. If the target computing power node has a computing power shortage and the computing power prediction accuracy is at a high level, mark the type of the target computing power node as a stable computing power shortage node;
[0096] 1069. If the target computing power node has a computing power shortage and the computing power prediction accuracy is at a low level, mark the type of the target computing power node as a sudden computing power shortage node;
[0097] According to the results of the previous steps, judge whether the target computing power node simultaneously meets the two conditions of "having a computing power shortage" and "high computing power prediction accuracy". If both conditions are met, proceed to the next step.
[0098] If the type of the target computing power node is marked as a "stable computing power shortage node", although there is a computing power shortage for this node, its computing power fluctuation has a certain regularity, and the prediction model can better predict its future computing power change trend. For stable computing power shortage nodes, targeted computing power resource scheduling strategies can be formulated. For example, a certain amount of computing power resources can be reserved in advance, or resources can be preferentially scheduled from other nodes to make up for the shortage when there is a computing power shortage.
[0099] If the type of the target computing power node is a "sudden computing power shortage node", the computing power shortage of this node is sudden, and the prediction model cannot accurately predict its future computing power change trend. For sudden computing power shortage nodes, more flexible computing power resource scheduling strategies need to be formulated. For example, the computing power status of the node can be monitored in real time. Once a computing power shortage is detected, resources are immediately scheduled from other nodes to handle it. At the same time, the monitoring and early warning of these nodes can be strengthened to detect and handle potential computing power problems in a timely manner.
[0100] 10610. Perform computing power resource scheduling based on the type of the target computing power node.
[0101] According to the type of target computing power nodes (stable missing computing power nodes or burst missing computing power nodes), select the corresponding computing power resource scheduling strategy. For stable missing computing power nodes, strategies such as reserved resources and priority scheduling can be adopted. For burst missing computing power nodes, a more flexible and rapid resource scheduling strategy is required.
[0102] In this embodiment, before resource scheduling, the available computing power resources in the system can be further evaluated, including calculating the remaining computing power of each node, evaluating factors such as communication latency and bandwidth between different nodes, etc. Then, based on the results of the resource evaluation, a specific computing power resource scheduling plan is formulated, including determining from which nodes to schedule resources, how much resources to schedule, when to schedule, etc. And the scheduling decision is transformed into specific execution operations, such as starting or stopping certain computing tasks, adjusting the load distribution of nodes, establishing or disconnecting communication connections between nodes, etc.
[0103] (1) The type of target computing power node is a stable missing computing power node
[0104] In this embodiment, when the type of target computing power node is a stable missing computing power node, the following method is used for computing power resource scheduling:
[0105] 1.1. Obtain the real-time computing power data and predicted real-time computing power data of the stable missing computing power node at different times and perform regression operations to obtain the stable predicted computing power of the stable missing computing power node within the deviation threshold range;
[0106] In this embodiment, since the computing power data may fluctuate greatly at different time points, in order to improve the stability and accuracy of the prediction model, it is necessary to perform normalization processing on the obtained real-time computing power data and predicted real-time computing power data at different times, so as to convert the data to the same scale. Among them, the predicted real-time computing power data needs to be predicted through a prediction model.
[0107] For example, pre-train the selected prediction algorithm (such as a recurrent neural network) using historical real-time computing power data, and optimize the model through methods such as cross-validation to improve the accuracy and robustness of the prediction. During the training process, the parameters of the model, such as the learning rate and the number of iterations, can be adjusted to achieve the best prediction effect. After the model training is completed, use the real-time computing power data as input and obtain the predicted real-time computing power data through the prediction model. Then, perform regression operations on the predicted real-time computing power data and the real-time computing power data to evaluate the accuracy of the prediction.
[0108] Based on the regression operation, set a deviation threshold (such as ±10%). Consider the part where the deviation between the predicted real-time computing power data and the real-time computing power data is within the threshold as stable predicted computing power. This can be achieved by calculating the error between the predicted value and the actual value and filtering out the data points where the error is within the threshold range.
[0109] 1.2. Add the stable predicted computing power and the predicted real-time computing power of the stable missing computing power node to obtain the predicted real-time computing power increase value of the stable missing computing power node. Subtract the stable predicted computing power from the predicted real-time computing power of the stable missing computing power node to obtain the predicted real-time computing power decrease value of the stable missing computing power node.
[0110] After obtaining the stable predicted computing power, perform an addition operation with the predicted real-time computing power data at the same time point. This will result in a value representing the computing power increase trend, that is, the predicted real-time computing power increase value. This value reflects the part where the node's computing power may increase in the future period. Similarly, perform a subtraction operation between the stable predicted computing power and the predicted real-time computing power data to obtain a value representing the computing power decrease trend, that is, the predicted real-time computing power decrease value. This value reflects the part where the node's computing power may decrease in the future period.
[0111] 1.3. Schedule the computing power resources corresponding to the predicted real-time computing power increase value of the stable missing computing power node to the stable missing computing power node, and schedule the computing power resources corresponding to the predicted real-time computing power decrease value of the stable missing computing power node to other computing power nodes.
[0112] Based on the magnitudes and directions of the predicted real-time computing power increase value and decrease value, formulate a computing power resource scheduling strategy, including determining the allocation ratio of computing power resources, scheduling timing, priority, etc. When formulating the strategy, factors such as the computing power requirements of the node, resource utilization rate, and task priority need to be comprehensively considered.
[0113] According to the formulated computing power resource scheduling strategy, allocate the computing power resources corresponding to the predicted real-time computing power increase value to the stable missing computing power node. This can be achieved by adjusting the computing power configuration of the node, optimizing task allocation, etc. At the same time, it is also necessary to pay attention to the real-time computing power changes of the node in order to adjust the resource allocation plan in a timely manner. For the computing power resources corresponding to the predicted real-time computing power decrease value, they can be recycled and reallocated to other computing power nodes. This can be achieved by releasing idle resources, optimizing task scheduling, etc. During the process of resource recycling and reallocation, it is necessary to ensure the effective utilization of resources and avoid waste.
[0114] (2)The type of the target computing power node is a burst missing computing power node
[0115] In this embodiment, when the type of the target computing power node is a burst missing computing power node, the following method is used for computing power resource scheduling:
[0116] 2.1. Classify the nodes with sudden lack of computing power into the first nodes with lack of computing power whose lack-of-computing-power threshold is lower than the preset computing-power threshold and the second nodes with lack of computing power whose lack-of-computing-power threshold is greater than or equal to the computing-power threshold;
[0117] The system needs to define a preset computing-power threshold. This threshold can be set based on historical data, business requirements, or system load conditions. For example, if the average computing power of the system is 100 units, a threshold of 80 units can be set, and nodes below this value will be considered as having insufficient computing power. For each node, calculate its current lack-of-computing-power threshold, which is obtained by comparing the node's current computing power with its maximum computing power or expected computing power. For example, if the maximum computing power of a node is 120 units and the current computing power is 60 units, then the lack-of-computing-power threshold is 60 units (i.e., half of the maximum computing power).
[0118] According to the preset computing-power threshold, classify the nodes into two categories:
[0119] The first nodes with lack of computing power: Nodes whose lack-of-computing-power threshold is lower than the preset computing-power threshold.
[0120] The second nodes with lack of computing power: Nodes whose lack-of-computing-power threshold is greater than or equal to the preset computing-power threshold.
[0121] 2.2. If the type of the node with sudden lack of computing power is the first node with lack of computing power, obtain the adjacent computing-power nodes of the first node with lack of computing power, predict the first computing-power threshold of the first node with lack of computing power according to the real-time computing-power data of the adjacent computing-power nodes, and schedule the computing-power resources in the first node with lack of computing power to the adjacent computing-power nodes and the first node with lack of computing power according to the preset ratio;
[0122] For each first node with lack of computing power, its adjacent computing-power nodes need to be found, which can be achieved through the network topology diagram or the connection relationship between nodes. For example, the adjacency list or adjacency matrix of a graph can be used to represent the connection relationship between nodes. Use machine learning algorithms or statistical models to predict the first computing-power threshold of the first node with lack of computing power, which can be achieved based on factors such as the real-time computing-power data of adjacent nodes, historical computing-power data, and the correlation between nodes. For example, algorithms such as linear regression, decision tree, or neural network can be used to establish a prediction model.
[0123] According to the predicted first computing-power threshold and the preset ratio, schedule the computing-power resources in the first node with lack of computing power to the adjacent computing-power nodes and the first node with lack of computing power, which can be achieved by adjusting the computing-power allocation strategy or task scheduling strategy of the node. For example, a part of the computing-power resources can be migrated from the first node with lack of computing power to the adjacent nodes to balance the computing-power load of the entire system. Finally, update the computing-power status of the first node with lack of computing power and the adjacent nodes, and record it in the system log or database.
[0124] 2.3. If the type of the suddenly missing computing power node is the second missing computing power node, monitor the change of the computing power data of the second missing computing power node within a preset time;
[0125] Within the preset time, the system needs to continuously collect the real-time computing power data of the second missing computing power node, including key indicators such as the CPU usage rate, memory occupancy rate, disk I / O speed, and network bandwidth of the node. Analyze the change of the collected computing power data to identify the computing power change trend of the node. For example, it can be achieved through time series analysis algorithms to capture the periodic, trend, and random components in the data.
[0126] 2.4. If the real-time computing power data of the second missing computing power node exceeds the first threshold range or the second threshold range, determine that the second missing computing power node is a suddenly changing computing power node; if the real-time computing power data of the second missing computing power node is within the first threshold range or the second threshold range, determine that the second missing computing power node is a continuously decreasing computing power node;
[0127] The first threshold range and the second threshold range are set according to the system's tolerance for computing power fluctuations. The first threshold range is usually narrower and is used to capture significant computing power mutations; the second threshold range is wider and is used to identify the trend of gradually decreasing computing power. These thresholds can be determined based on historical data, business requirements, and system stability requirements. During the real-time monitoring process, the system needs to continuously compare the real-time computing power data of the second missing computing power node with the set threshold range. According to the comparison result of the real-time computing power data and the threshold range, the system classifies the second missing computing power node into two categories:
[0128] Suddenly changing computing power node: A node whose real-time computing power data exceeds the first threshold range or the second threshold range. The computing power of such nodes changes significantly and unpredictably, and may be caused by hardware failures, software anomalies, or external attacks, etc.
[0129] Continuously decreasing computing power node: A node whose real-time computing power data is within the first threshold range or the second threshold range. The computing power of such nodes gradually decreases, and may be caused by resource overload, hardware aging, or software performance bottlenecks, etc.
[0130] 2.5. If the second missing computing power node is a suddenly changing computing power node, remove the second missing computing power node; if the second missing computing power node is a continuously decreasing computing power node, schedule the computing power resources in the second missing computing power node to other computing power nodes according to a preset ratio.
[0131] For the second missing computing power node determined to be a sudden change computing power node, the system needs to immediately remove it from the computing power resource pool. This can be achieved by stopping the node's services, releasing the resources it occupies, and updating the node status, etc. The purpose of removal is to prevent this node from further affecting the overall performance of the system. For the second missing computing power node determined to be a continuously decreasing computing power node, the system needs to schedule its remaining computing power resources to other computing power nodes according to a preset ratio. The scheduling ratio can be determined based on the real-time computing power requirements of the nodes, the priorities, and the overall computing power allocation strategy of the system. The purpose of scheduling is to maximize the utilization of system resources and ensure the continuity and stability of the business. After completing the removal or scheduling operation, the system needs to update the node status and resource configuration information, including marking the status of the removed node as "unavailable", updating the computing power resources of the scheduled nodes to the resource management system, and adjusting the load balancing strategy of the relevant nodes, etc.
[0132] In one embodiment, to further improve the efficiency and accuracy of computing power resource scheduling, this embodiment further provides a computing power node recommendation function. In this embodiment, the computing power resource scheduling method further includes:
[0133] S201. Obtain the average resource utilization rate and the maximum resource utilization rate corresponding to each target computing power node at each time period within the historical period;
[0134] Obtain the relevant data of each target computing power node, including the timestamp, computing power node ID, resource utilization rate (CPU, memory, disk, network, etc.). Then calculate the average resource utilization rate and the maximum resource utilization rate. The historical period can be divided into multiple time periods, and each time period can be an hour, a day, a week, etc., depending on the business requirements. For each time period, calculate the average value of the resource utilization rate of each target computing power node. This usually involves summing up the resource utilization rates of all sampling points within the time period and then dividing by the number of sampling points. Similarly, for each time period, find the maximum value of the resource utilization rate of each target computing power node. Finally, store the calculated average resource utilization rate and maximum resource utilization rate in the database for subsequent steps to use.
[0135] S202. Subtract the average resource utilization rate corresponding to each target computing power node at each time period within the historical period from the maximum resource utilization rate to obtain the resource utilization rate difference corresponding to each target computing power node at each time period within the historical period;
[0136] Read the average resource utilization rate and the maximum resource utilization rate calculated in step S201 from the database. For each target computing power node, at each time period, subtract the average resource utilization rate from the maximum resource utilization rate to obtain the resource utilization rate difference. This difference reflects the fluctuation of the resource utilization rate, that is, the gap between the peak value and the average value of the resource utilization rate.
[0137] S203. Multiply the computing power resources corresponding to each target computing power node at each time period within the historical period by the difference in resource utilization rate corresponding to the corresponding time period to obtain the computing power resource utilization rate corresponding to each target computing power node at each time period within the historical period;
[0138] Read from the database the difference in resource utilization rate calculated in step S202 and the computing power resources (such as the number of CPU cores, memory capacity, etc.) of the target computing power node at each time period. For each target computing power node, at each time period, multiply the computing power resources by the corresponding difference in resource utilization rate to obtain the computing power resource utilization rate. This value reflects the actual usage of computing power resources under the condition of fluctuating resource utilization rate.
[0139] S204. Compare the computing power resource utilization rate corresponding to each target computing power node at each time period within the historical period with the set resource utilization rate. If the computing power resource utilization rate corresponding to each target computing power node at each time period within the historical period is higher than or equal to the set resource utilization rate, then use the corresponding target computing power node as the historical recommended computing power node for the next computing power resource scheduling;
[0140] According to the business requirements, set a resource utilization rate threshold. This threshold can be fixed or dynamically adjusted, depending on the business requirements and resource management strategies. Read from the database the computing power resource utilization rate calculated in step S203. For each target computing power node, at each time period within the historical period, compare the computing power resource utilization rate with the set resource utilization rate threshold. If the computing power resource utilization rate in all time periods is higher than or equal to the threshold, then use the target computing power node as the historical recommended computing power node for the next computing power resource scheduling.
[0141] S205. Obtain the actual scheduling resources and the estimated scheduling resources of each historical recommended computing power node, where the estimated scheduling resources are obtained by predicting through a pre-set intelligent prediction model;
[0142] Read from the database the actual scheduling resource data of the historical recommended computing power node in the past period of time, including the allocation and usage of computing power resources, etc. According to the business requirements, extract the features affecting the computing power resource scheduling, such as time, the geographical location of the computing power node, hardware configuration, historical resource usage, etc. According to the characteristics of the data and the business requirements, select a suitable intelligent prediction model, such as a time series analysis model, a machine learning model (such as a regression model, decision tree, random forest, etc.) or a deep learning model (such as a neural network). Use the pre-processed data to train the intelligent prediction model to obtain the model parameters. Adjust the model parameters according to the evaluation results or select other models. Use the trained intelligent prediction model to predict the estimated scheduling resources of each historical recommended computing power node in the future period of time according to the current feature data.
[0143] S206. Determine the initial values of computing power resource scheduling corresponding to each historical recommended computing power node based on the actual scheduled resources and the predicted scheduled resources of each historical recommended computing power node.
[0144] Read the actual scheduled resources and the predicted scheduled resources obtained in step S205 from the database. For each historical recommended computing power node, calculate the initial value of computing power resource scheduling according to the actual demand and business strategy, in combination with the actual scheduled resources and the predicted scheduled resources. This initial value can be the average value, weighted average value or other statistical quantities of the actual scheduled resources and the predicted scheduled resources, depending on the business requirements and resource management strategy.
[0145] S207. Compare the initial values of resource scheduling corresponding to each historical recommended computing power node with the set initial value of resource scheduling. If the initial value of resource scheduling corresponding to each historical recommended computing power node is higher than or equal to the set initial value of resource scheduling, then use the corresponding historical recommended node as the historical recommended computing power node for the final value of demand resource scheduling; if the initial value of resource scheduling corresponding to each historical recommended computing power node is lower than the set initial value of resource scheduling, then use the corresponding historical recommended computing power node as the historical recommended computing power node for the final value of restricted resource scheduling.
[0146] Set a threshold for the initial value of resource scheduling according to the business requirements. This threshold can be fixed or dynamically adjusted, depending on the business requirements and the resource management strategy. Read the initial value of computing power resource scheduling calculated in step S206 from the database. For each historical recommended computing power node, compare its initial value of computing power resource scheduling with the set threshold for the initial value of resource scheduling. If the initial value of computing power resource scheduling is higher than or equal to the threshold, then use the historical recommended computing power node as the historical recommended computing power node for the final value of demand resource scheduling; if the initial value of computing power resource scheduling is lower than the threshold, then use it as the historical recommended computing power node for the final value of restricted resource scheduling.
[0147] S208. Input the final values of demand resource scheduling and the final values of restricted resource scheduling respectively into the intelligent prediction model for prediction to obtain the schedulable resources corresponding to each historical recommended computing power node.
[0148] Extract the features affecting the computing power resource scheduling according to the business requirements, such as time, geographical location of the computing power node, hardware configuration, historical resource usage, etc. For the historical recommended computing power nodes with the final values of demand resource scheduling and the final values of restricted resource scheduling, extract the corresponding feature data respectively. Use the intelligent prediction model trained in step S205 to predict the schedulable resources. Specifically, input the feature data of the historical recommended computing power nodes with the final values of demand resource scheduling and the final values of restricted resource scheduling into the intelligent prediction model respectively to predict the schedulable resources of each historical recommended computing power node in a future period of time.
[0149] S209. Calculate the difference between the schedulable resources corresponding to each historical recommended computing power node and the initial resource scheduling value corresponding to the corresponding historical recommended computing power node to obtain the final resource scheduling value of the corresponding historical recommended computing power node.
[0150] Read the schedulable resources corresponding to each historical recommended computing power node predicted in step S208 from the database. At the same time, read the initial resource scheduling values corresponding to each historical recommended computing power node calculated in step S206 from the database. Ensure that the schedulable resources of each historical recommended computing power node can be correctly matched with its corresponding initial resource scheduling value. This is usually achieved through the unique identifier (such as ID) of the computing power node. For each historical recommended computing power node, calculate the difference between its schedulable resources and its initial resource scheduling value. This difference reflects the adjustment amount from the initial value to the final value, that is, the adjustment that needs to be made to the initial value according to the prediction result of the intelligent prediction model.
[0151] The difference calculation formula can be expressed as: Final resource scheduling value = Schedulable resources - Initial resource scheduling value. The difference is directly used as the adjustment amount. Therefore, the final resource scheduling value is actually equal to the initial value plus this difference. It should be noted that "difference calculation" does not mean that the result obtained by direct subtraction is the final scheduling value, but this difference is used to represent the adjustment direction and amplitude from the initial value to the final value. The final scheduling value (i.e., the final resource scheduling value) should be the initial value plus (or minus, depending on the sign of the difference) this adjustment amount.
[0152] In addition, if the calculated difference is very large (positive or negative), it may be necessary to further check the accuracy of the data and the prediction performance of the model. If the difference exceeds the preset reasonable range, other strategies can be considered for adjustment, such as limiting the maximum adjustment amplitude, using a smoothing algorithm, etc.
[0153] The above describes the computing power resource scheduling method in the embodiments of the present invention. Next, the computing power resource scheduling system in the embodiments of the present invention will be described. An embodiment of the computing power resource scheduling system in the embodiments of the present invention includes:
[0154] An acquisition module for acquiring the real-time computing power data of each computing power node and the computing tasks of the computing power resources to be allocated;
[0155] A matching module for matching the computing power nodes based on the real-time computing power data of each computing power node and the computing tasks of the computing power resources to be allocated to select the target computing power nodes;
[0156] A prediction module for acquiring the historical computing power data of each target computing power node, and sequentially inputting the real-time computing power data and historical computing power data of each target computing power node into a preset neural network model for computing power prediction to obtain the predicted real-time computing power data of each target computing power node;
[0157] A judgment module, configured to judge the working modes of each computing power node respectively based on the real-time computing power data and predicted real-time computing power data of each target computing power node, where the working modes include a stable working mode and / or a fluctuating working mode;
[0158] A scheduling module, configured to, if the target computing power node is in the stable working mode, perform computing power resource scheduling according to the historical computing power data of the target computing power node; if the target computing power node is in the fluctuating working mode, perform computing power resource scheduling according to the real-time computing power data and predicted real-time computing power data of the target computing power node.
[0159] Optionally, in an embodiment, the matching module is specifically configured to:
[0160] Judge whether the computing power node matching condition is satisfied according to the node running time in the real-time computing power data of each computing power node and the task execution time of the computing task to which the computing power resource to be allocated;
[0161] If the node running time of the computing power node is earlier than the task execution time of the computing task to which the computing power resource to be allocated, judge whether the task execution time of the computing task to which the computing power resource to be allocated is between the low valley running time and the peak running time of the current computing power node;
[0162] If the task execution time of the computing task to which the computing power resource to be allocated is between the low valley running time and the peak running time of the current computing power node, determine the current computing power node as a candidate computing power node;
[0163] Perform feature matching according to the computing power characteristics of each candidate computing power node and the task characteristics of the computing task to which the computing power resource to be allocated, and select the candidate computing power node with the highest feature matching degree as the target computing power node.
[0164] Optionally, in an embodiment, the scheduling module is specifically configured to:
[0165] If the target computing power node is in the fluctuating working mode, compare the numerical difference between the real-time computing power data and the predicted real-time computing power data of the target computing power node in real time;
[0166] If the numerical difference exceeds the preset difference range, determine that there is a computing power anomaly in the target computing power node currently, and increase the number of anomaly times of the target computing power node by one;
[0167] When the cumulative number of anomaly times of the target computing power node is within the first preset range, determine that the target computing power node is in a regular fluctuation state; when the cumulative number of anomaly times of the target computing power node is within the second preset range, determine that the target computing power node is in an irregular fluctuation mode;
[0168] If the target computing power node is in a regular fluctuation state, obtain the computing power fluctuation threshold of the target computing power node, and perform computing power resource scheduling according to the computing power fluctuation threshold;
[0169] If the target computing power node is in an irregular fluctuation state, obtain the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node;
[0170] Calculate the difference between the latest real-time computing power data and the latest predicted real-time computing power data to obtain the computing power fluctuation range of the target computing power node;
[0171] Based on the computing power fluctuation range of the target computing power node, determine whether the target computing power node has a computing power shortage and determine the computing power prediction accuracy of the target computing power node in the fluctuation working mode;
[0172] If the target computing power node has a computing power shortage and the computing power prediction accuracy is at a high level, mark the type of the target computing power node as a stable computing power shortage node;
[0173] If the target computing power node has a computing power shortage and the computing power prediction accuracy is at a low level, mark the type of the target computing power node as a sudden computing power shortage node;
[0174] Perform computing power resource scheduling based on the type of the target computing power node.
[0175] Optionally, in one embodiment, the judging module is specifically configured to:
[0176] If the value of the computing power fluctuation range of the target computing power node is negative, it is determined that the target computing power node has a computing power shortage. If the value of the computing power fluctuation range of the target computing power node is positive, it is determined that the target computing power node does not have a computing power shortage;
[0177] Obtain the maximum historical computing power fluctuation range in the historical computing power fluctuation ranges of the target computing power node in the fluctuation working mode;
[0178] When the difference between the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node exceeds the deviation threshold range, calculate the difference between the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node to obtain the real-time computing power fluctuation range of the target computing power node;
[0179] Calculate the ratio of the real-time computing power fluctuation range of the target computing power node to the maximum historical computing power fluctuation range. If the ratio is greater than or equal to the preset ratio threshold, it is determined that the computing power prediction accuracy of the target computing power node is at a high level. If the ratio is less than the ratio threshold, it is determined that the computing power prediction accuracy of the target computing power node is at a low level.
[0180] Optionally, in one embodiment, the scheduling module is specifically further configured to:
[0181] If the type of the target computing power node is a stable missing computing power node, obtain the real-time computing power data and predicted real-time computing power data of the stable missing computing power node at different times and perform a regression operation to obtain the stable predicted computing power of the stable missing computing power node within the deviation threshold range;
[0182] Add the stable predicted computing power and the predicted real-time computing power of the stable missing computing power node to obtain the predicted real-time computing power increase value of the stable missing computing power node, and subtract the stable predicted computing power and the predicted real-time computing power of the stable missing computing power node to obtain the predicted real-time computing power decrease value of the stable missing computing power node;
[0183] Schedule the computing power resources corresponding to the predicted real-time computing power increase value of the stable missing computing power node to the stable missing computing power node, and schedule the computing power resources corresponding to the predicted real-time computing power decrease value of the stable missing computing power node to other computing power nodes.
[0184] Optionally, in an embodiment, the scheduling module is further specifically configured to:
[0185] If the type of the target computing power node is a sudden missing computing power node, classify the sudden missing computing power node into a first missing computing power node with a computing power missing threshold lower than a preset computing power threshold and a second missing computing power node with a computing power missing threshold greater than or equal to the computing power threshold;
[0186] If the type of the sudden missing computing power node is a first missing computing power node, obtain the adjacent computing power nodes of the first missing computing power node, predict the first computing power threshold of the first missing computing power node according to the real-time computing power data of the adjacent computing power nodes, and schedule the computing power resources in the first missing computing power node to the adjacent computing power nodes and the first missing computing power node according to a preset ratio;
[0187] If the type of the sudden missing computing power node is a second missing computing power node, monitor the change of the computing power data of the second missing computing power node within a preset time;
[0188] If the real-time computing power data of the second missing computing power node exceeds the first threshold range or the second threshold range, determine that the second missing computing power node is a suddenly changing computing power node, and if the real-time computing power data of the second missing computing power node is within the first threshold range or the second threshold range, determine that the second missing computing power node is a continuously decreasing computing power node;
[0189] If the second missing computing power node is a suddenly changing computing power node, remove the second missing computing power node, and if the second missing computing power node is a continuously decreasing computing power node, schedule the computing power resources in the second missing computing power node to other computing power nodes according to a preset ratio.
[0190] Optionally, in an embodiment, the neural network model is obtained by using the following training method:
[0191] Obtain the historical real-time computing power data and historical predicted computing power data of each computing power node, where the historical computing power data for each period includes the error characteristic parameters of the actual computing power value and the predicted computing power value corresponding to the time stamp;
[0192] Construct an initial neural network model, where the input layer of the initial neural network model inputs the historical real-time computing power data and the historical predicted computing power data, and the output layer outputs the currently predicted real-time computing power data predicted by the computing power node;
[0193] Construct the loss function of the initial neural network model according to the historical real-time computing power data and the historical predicted computing power data of each computing power node, and train the initial neural network model to obtain a trained neural network model;
[0194] Among them, the input of the hidden layer of the initial neural network model is as follows:
[0195]
[0196] Among them, represents the historical real-time computing power data vector of the kth input of the computing power node i in the hidden layer, represents the historical predicted computing power data vector of the kth input of the computing power node i in the hidden layer, is the historical real-time computing power data of the kth input of the computing power node i, is the input weight of the historical real-time computing power data of the kth input of the computing power node i, is the bias term of the computing power node i, is the input weight coefficient of the kth input, N is the number of neurons in the input layer, f is the activation function, is the historical predicted computing power data of the kth input of the computing power node i, is the input weight of the historical predicted computing power data of the kth input of the computing power node i, is the bias term of the computing power node i, is the input weight coefficient of the kth input, M is the number of neurons in the input layer;
[0197] Among them, the loss function of the initial neural network model is as follows:
[0198]
[0199] Among them, represents the loss function of the initial neural network model, is the computing power data distribution matrix between the computing power node i and the computing power node j in the historical time period, S is the number of computing power node samples, is the real-time computing power data of the computing power node j, It represents the predicted computing power data of computing power node j predicted by the model.
[0200] Optionally, in one embodiment, the computing power resource scheduling device further includes:
[0201] A utilization rate calculation module, configured to obtain the average resource utilization rate and the maximum resource utilization rate corresponding to each target computing power node at each time period within the historical period; subtract the average resource utilization rate corresponding to each target computing power node at each time period within the historical period from the maximum resource utilization rate, to obtain the resource utilization rate difference corresponding to each target computing power node at each time period within the historical period; multiply the computing power resources corresponding to each target computing power node at each time period within the historical period by the corresponding resource utilization rate difference within the corresponding time period, to obtain the computing power resource utilization rate corresponding to each target computing power node at each time period within the historical period.
[0202] A utilization rate comparison module, configured to compare the computing power resource utilization rate corresponding to each target computing power node at each time period within the historical period with the set resource utilization rate. If the computing power resource utilization rate corresponding to each target computing power node at each time period within the historical period is higher than or equal to the set resource utilization rate, then use the corresponding target computing power node as the historical recommended computing power node for the next computing power resource scheduling.
[0203] A node recommendation module, configured to obtain the actual scheduled resources and the predicted scheduled resources of each historical recommended computing power node, where the predicted scheduled resources are obtained by predicting through a preset intelligent prediction model; based on the actual scheduled resources and the predicted scheduled resources of each historical recommended computing power node, determine the initial value of the computing power resource scheduling corresponding to each historical recommended computing power node; compare the initial value of the resource scheduling corresponding to each historical recommended computing power node with the set initial value of the resource scheduling. If the initial value of the resource scheduling corresponding to each historical recommended computing power node is higher than or equal to the set initial value of the resource scheduling, then use the corresponding historical recommended node as the historical recommended computing power node for the final value of the required resource scheduling; if the initial value of the resource scheduling corresponding to each historical recommended computing power node is lower than the set initial value of the resource scheduling, then use the corresponding historical recommended computing power node as the historical recommended computing power node for the final value of the restricted resource scheduling.
[0204] A scheduled resource calculation module, configured to respectively input each final value of the required resource scheduling and each final value of the restricted resource scheduling into the intelligent prediction model for prediction, to obtain the schedulable resources corresponding to each historical recommended computing power node; perform a difference calculation on the schedulable resources corresponding to each historical recommended computing power node and the initial value of the resource scheduling corresponding to the corresponding historical recommended computing power node, to obtain the final value of the resource scheduling corresponding to the corresponding historical recommended computing power node.
[0205] Since the embodiments of the system part correspond to the embodiments of the above method, the description of the computing power resource scheduling system provided by the present invention can refer to the above method embodiments, and the present invention will not be repeated here. It has the same beneficial effects as the above computing power resource scheduling method.
[0206] The present invention also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the computing power resource scheduling method.
[0207] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A computing power resource scheduling method, characterized in that, The computing power resource scheduling method includes: Obtaining the real-time computing power data of each computing power node and the computing tasks for which computing power resources are to be allocated; Performing computing power node matching based on the real-time computing power data of each computing power node and the computing tasks for which computing power resources are to be allocated to select target computing power nodes; Obtaining the historical computing power data of each target computing power node, and sequentially inputting the real-time computing power data and historical computing power data of each target computing power node into a pre-set neural network model for computing power prediction to obtain the predicted real-time computing power data of each target computing power node; Based on the real-time computing power data and predicted real-time computing power data of each target computing power node, respectively determine the working mode of each computing power node, where the working mode includes a stable working mode and a fluctuating working mode; If the target computing power node is in the stable working mode, perform computing power resource scheduling according to the historical computing power data of the target computing power node; If the target computing power node is in the fluctuating working mode, perform computing power resource scheduling according to the real-time computing power data and predicted real-time computing power data of the target computing power node; Among them, the step of if the target computing power node is in the fluctuating working mode, perform computing power resource scheduling according to the real-time computing power data and predicted real-time computing power data of the target computing power node includes: If the target computing power node is in the fluctuating working mode, compare the numerical difference between the real-time computing power data and the predicted real-time computing power data of the target computing power node in real time; If the numerical difference exceeds the pre-set difference range, it is determined that there is a computing power anomaly in the target computing power node currently, and the number of anomalies of the target computing power node is increased by one; When the accumulated number of anomalies of the target computing power node is within the second preset range, it is determined that the target computing power node is in a non-regular fluctuation state; If the target computing power node is in a non-regular fluctuation state, obtain the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node; Calculate the difference between the latest real-time computing power data and the latest predicted real-time computing power data to obtain the computing power fluctuation amplitude of the target computing power node; Based on the computing power fluctuation amplitude of the target computing power node, determine whether the target computing power node has a computing power shortage and determine the computing power prediction accuracy of the target computing power node in the fluctuating working mode; If the target computing power node has a computing power shortage and the computing power prediction accuracy is at a high level, mark the type of the target computing power node as a stable computing power shortage node; If the target computing power node has a computing power shortage and the computing power prediction accuracy is at a low level, mark the type of the target computing power node as a sudden computing power shortage node; Perform computing power resource scheduling based on the type of the target computing power node.
2. The computing power resource scheduling method according to claim 1, wherein The step of performing computing power node matching based on the real-time computing power data of each computing power node and the computing tasks for which computing power resources are to be allocated to select target computing power nodes includes: According to the node running time in the real-time computing power data of each computing power node and the task execution time of the computing task for which computing power resources are to be allocated, determine whether the computing power node matching condition is satisfied; If the node running time of the computing power node is earlier than the task execution time of the computing task for which computing power resources are to be allocated, determine whether the task execution time of the computing task for which computing power resources are to be allocated is between the low valley running time and the peak running time of the current computing power node; If the task execution time of the computing task for which computing power resources are to be allocated is between the low - valley running time and the peak running time of the current computing power node, then determine the current computing power node as a candidate computing power node; Perform feature matching based on the computing power characteristics of each candidate computing power node and the task characteristics of the computing task for which computing power resources are to be allocated, and select the candidate computing power node with the highest feature matching degree as the target computing power node.
3. The computing power resource scheduling method according to claim 1, wherein The judging whether the target computing power node has computing power shortage and the judging of the computing power prediction accuracy of the target computing power node in the fluctuating working mode based on the computing power fluctuation range of the target computing power node include: If the value of the computing power fluctuation range of the target computing power node is negative, determine that the target computing power node has a computing power shortage; if the value of the computing power fluctuation range of the target computing power node is positive, determine that the target computing power node does not have a computing power shortage; Obtain the maximum historical computing power fluctuation range in the historical computing power fluctuation ranges of the target computing power node in the fluctuating working mode; When the difference between the latest real - time computing power data and the latest predicted real - time computing power data of the target computing power node exceeds the deviation threshold range, calculate the difference between the latest real - time computing power data and the latest predicted real - time computing power data of the target computing power node to obtain the real - time computing power fluctuation range of the target computing power node; Calculate the ratio of the real - time computing power fluctuation range of the target computing power node to the maximum historical computing power fluctuation range. If the ratio is greater than or equal to the preset ratio threshold, determine that the computing power prediction accuracy of the target computing power node is at a high level; if the ratio is less than the ratio threshold, determine that the computing power prediction accuracy of the target computing power node is at a low level.
4. The computing power resource scheduling method according to claim 1, characterized in that The computing power resource scheduling based on the type of the target computing power node includes: If the type of the target computing power node is a stable computing power shortage node, obtain the real - time computing power data and the predicted real - time computing power data of the stable computing power shortage node at different times and perform a regression operation to obtain the stable predicted computing power of the stable computing power shortage node within the deviation threshold range; Add the stable predicted computing power and the predicted real - time computing power of the stable computing power shortage node to obtain the predicted real - time computing power increase value of the stable computing power shortage node, and subtract the predicted real - time computing power from the stable predicted computing power of the stable computing power shortage node to obtain the predicted real - time computing power decrease value of the stable computing power shortage node; Schedule the computing power resources corresponding to the predicted real - time computing power increase value of the stable computing power shortage node to the stable computing power shortage node, and schedule the computing power resources corresponding to the predicted real - time computing power decrease value of the stable computing power shortage node to other computing power nodes.
5. The computing power resource scheduling method according to claim 1, wherein The computing power resource scheduling based on the type of the target computing power node includes: If the type of the target computing power node is a sudden computing power shortage node, classify the sudden computing power shortage node into a first computing power shortage node with a computing power shortage threshold lower than the preset computing power threshold and a second computing power shortage node with a computing power shortage threshold greater than or equal to the computing power threshold; If the type of the sudden computing power shortage node is a first computing power shortage node, obtain the adjacent computing power nodes of the first computing power shortage node, predict the first computing power threshold of the first computing power shortage node according to the real - time computing power data of the adjacent computing power nodes, and schedule the computing power resources in the first computing power shortage node to the adjacent computing power nodes and the first computing power shortage node according to a preset ratio; If the type of the suddenly missing computing power node is the second missing computing power node, the computing power data change of the second missing computing power node within a preset time is monitored in real time; If the real-time computing power data of the second missing computing power node exceeds the first threshold range or the second threshold range, it is determined that the second missing computing power node is a computing power node with sudden change. If the real-time computing power data of the second missing computing power node is within the first threshold range or the second threshold range, it is determined that the second missing computing power node is a computing power node with continuous reduction; If the second missing computing power node is a computing power node with sudden change, the second missing computing power node is removed. If the second missing computing power node is a computing power node with continuous reduction, the computing power resources in the second missing computing power node are scheduled to other computing power nodes according to a preset ratio; 6. The computing power resource scheduling method according to claim 1, wherein The neural network model is obtained by the following training method: Obtain the historical real-time computing power data and historical predicted computing power data of each computing power node. Among them, the historical computing power data of each period includes the error characteristic parameters of the actual computing power value and the predicted computing power value corresponding to the time stamp; Construct an initial neural network model. The input layer of the initial neural network model inputs the historical real-time computing power data and historical predicted computing power data, and the output layer outputs the currently predicted real-time computing power data predicted by the computing power node; According to the historical real-time computing power data and historical predicted computing power data of each computing power node, construct the loss function of the initial neural network model, and train the initial neural network model to obtain a trained neural network model; Among them, the input of the hidden layer of the initial neural network model is as follows: Among them, represents the historical real-time computing power data vector of the k-th input of the hidden layer computing power node i, represents the historical predicted computing power data vector of the k-th input of the hidden layer computing power node i, is the historical real-time computing power data of the k-th input of the computing power node i, w ik is the input weight of the historical real-time computing power data of the k-th input of the computing power node i, b i is the bias term of the computing power node i, α k is the input weight coefficient of the k-th input, N is the number of neurons in the input layer, f is the activation function, is the historical predicted computing power data of the k-th input of the computing power node i, v ik is the input weight of the historical predicted computing power data of the k-th input of the computing power node i, c i is the bias term of the computing power node i, β k is the input weight coefficient of the k-th input, M is the number of neurons in the input layer and M is equal to N.
7. The computing power resource scheduling method according to any one of claims 1-6, characterized in that, The computing power resource scheduling method further includes: Obtain the average resource utilization rate and the maximum resource utilization rate corresponding to each time period of each target computing power node in the historical period; Subtract the maximum resource utilization rate corresponding to each time period of each target computing power node in the historical period from the average resource utilization rate to obtain the resource utilization rate difference corresponding to each time period of each target computing power node in the historical period; Multiply the computing power resources corresponding to each time period of each target computing power node in the historical period by the corresponding resource utilization rate difference in the corresponding time period to obtain the computing power resource utilization rate corresponding to each time period of each target computing power node in the historical period; Compare the computing power resource utilization rate corresponding to each time period of each target computing power node in the historical period with the set resource utilization rate. If the computing power resource utilization rate corresponding to each time period of each target computing power node in the historical period is higher than or equal to the set resource utilization rate, the corresponding target computing power node is used as the historical recommended computing power node for the next computing power resource scheduling; Obtain the actual scheduled resources and the expected scheduled resources of each historical recommended computing power node, where the expected scheduled resources are obtained by predicting through a preset intelligent prediction model; Based on the actual scheduled resources and the expected scheduled resources of each historical recommended computing power node, determine the initial value of the computing power resource scheduling corresponding to each historical recommended computing power node; Compare the initial resource scheduling values corresponding to each historical recommended computing power node with the set initial resource scheduling value. If the initial resource scheduling value corresponding to each historical recommended computing power node is higher than or equal to the set initial resource scheduling value, then use the corresponding historical recommended node as the historical recommended computing power node for the final demand resource scheduling value. If the initial resource scheduling value corresponding to each historical recommended computing power node is lower than the set initial resource scheduling value, then use the corresponding historical recommended computing power node as the historical recommended computing power node for the final restricted resource scheduling value; Input each final demand resource scheduling value and each final restricted resource scheduling value into the intelligent prediction model for prediction respectively to obtain the schedulable resources corresponding to each historical recommended computing power node; Perform a difference calculation between the schedulable resources corresponding to each historical recommended computing power node and the initial resource scheduling value corresponding to the corresponding historical recommended computing power node to obtain the final resource scheduling value of the corresponding historical recommended computing power node.
8. A computing power resource scheduling system, characterized in that, The computing power resource scheduling system includes: An acquisition module for acquiring the real-time computing power data of each computing power node and the computing tasks of the computing power resources to be allocated; A matching module for matching the computing power nodes based on the real-time computing power data of each computing power node and the computing tasks of the computing power resources to be allocated to select the target computing power nodes; A prediction module for acquiring the historical computing power data of each target computing power node and sequentially inputting the real-time computing power data and historical computing power data of each target computing power node into a preset neural network model for computing power prediction to obtain the predicted real-time computing power data of each target computing power node; A judgment module for respectively judging the working modes of each computing power node based on the real-time computing power data and predicted real-time computing power data of each target computing power node, and the working modes include a stable working mode and a fluctuating working mode; A scheduling module for, if the target computing power node is in the stable working mode, performing computing power resource scheduling according to the historical computing power data of the target computing power node; if the target computing power node is in the fluctuating working mode, performing computing power resource scheduling according to the real-time computing power data and predicted real-time computing power data of the target computing power node; Among them, the scheduling module is specifically used for: If the target computing power node is in the fluctuating working mode, then compare the real-time computing power data of the target computing power node with the predicted real-time computing power data in real time; If the numerical difference exceeds the preset difference range, then determine that there is a computing power anomaly in the target computing power node currently and increase the number of anomaly times of the target computing power node by one; When the accumulated number of anomaly times of the target computing power node is within the second preset range, then determine that the target computing power node is in an irregular fluctuation state; If the target computing power node is in the irregular fluctuation state, then acquire the latest real-time computing power data and the latest predicted real-time computing power data of the target computing power node; Calculate the difference between the latest real-time computing power data and the latest predicted real-time computing power data to obtain the computing power fluctuation range of the target computing power node; Based on the computing power fluctuation range of the target computing power node, judge whether the target computing power node has a computing power shortage and judge the computing power prediction accuracy of the target computing power node in the fluctuating working mode; If the target computing power node has a computing power shortage and the computing power prediction accuracy is at a high level, then mark the type of the target computing power node as a stable computing power shortage node; If the target computing power node has a computing power shortage and the computing power prediction accuracy is at a low level, then mark the type of the target computing power node as a sudden computing power shortage node; Perform computing power resource scheduling based on the type of the target computing power node.
9. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is caused to execute the steps of the computing power resource scheduling method according to any one of claims 1-7.
Citation Information
Patent Citations
Computing power data management system and method based on distributed computing
CN119025283A
Computing resource scheduling optimization method for large model
CN119311395A