Super-computing resource intelligent allocation method and system based on AI load monitoring
By obtaining real-time load data of the computing nodes of the supercomputing system and performing feature extraction and prediction model analysis, the problem of unreasonable resource allocation in traditional supercomputing system resource management is solved, intelligent resource allocation is realized, and the efficiency and stability of the system are improved.
Patent Information
- Application Number
- CN202510869195.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The traditional supercomputing system resource management method fails to effectively utilize historical data and real-time information, resulting in unreasonable resource allocation and low task execution efficiency, affecting the overall performance and stability of the system.
By obtaining real-time load data of the computing node of the supercomputing system, extracting load characteristics and using the pre-built resource demand prediction model for analysis, generating resource allocation priority and migration strategies, and realizing intelligent resource allocation.
It improves the resource utilization rate of the supercomputing system, optimizes the task execution process, shortens the task completion time, and improves the overall energy efficiency ratio and stability of the system.
Smart Images

Figure CN120429121A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for intelligent allocation of supercomputing resources based on AI load monitoring. Background Art
[0002] In the field of high-performance computing, supercomputer systems (abbreviated as supercomputers) serve as core infrastructure for handling complex tasks such as large-scale scientific calculations, engineering simulations, and data analysis. Their operational efficiency and resource utilization are directly related to the speed and quality of computing task completion. Traditional supercomputer system resource management methods rely on static configuration or dynamic adjustments based on simple rules. These methods struggle to accurately capture the actual load changes of computing nodes within different time windows, leading to problems such as irrational resource allocation, inefficient task execution, and low overall system energy efficiency.
[0003] Specifically, existing supercomputing system resource management often ignores the dynamic nature of computing node loads and fails to effectively leverage historical data and real-time information to predict future resource needs, resulting in a lack of foresight and flexibility in resource allocation. This not only wastes computing resources but can also lead to task execution delays or system overloads due to insufficient or excessive resource allocation, impacting the overall performance and stability of the supercomputing system. Summary of the Invention
[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a method for intelligent allocation of supercomputing resources based on AI load monitoring, the method comprising: Acquire a real-time load data set of multiple computing nodes in a supercomputer system, wherein the real-time load data set includes a resource occupancy index and a task queue status parameter of each computing node in a continuous time window; Performing load feature extraction processing on the real-time load data set to generate a load feature set of each computing node, the load feature set including resource occupancy fluctuation features and task queue evolution features; Calling a pre-built resource demand prediction model to perform resource demand analysis on the load feature set to generate a resource demand prediction result for each computing node in a subsequent time window; Determining resource allocation priority parameters and resource migration strategy parameters for each computing node in the supercomputing system based on the resource demand prediction result; A resource dynamic allocation instruction is generated according to the resource allocation priority parameter and the resource migration policy parameter, and the resource dynamic allocation instruction is sent to the target computing node to trigger a resource reallocation operation.
[0005] On the other hand, an embodiment of the present invention also provides an intelligent allocation system for supercomputing resources based on AI load monitoring, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0006] Based on the above aspects, the embodiment of the present invention obtains a real-time load data set of multiple computing nodes in a supercomputing system, and conducts an in-depth analysis of the resource occupancy rate indicators and task queue status parameters therein, thereby extracting a load feature set that can reflect the load characteristics of the computing nodes. Furthermore, by using a pre-built resource demand prediction model, the resource demand of each computing node in the subsequent time window is accurately predicted based on these load feature sets, thereby realizing the transformation of resource management from passive response to active prediction. Based on the resource demand prediction results, the resource allocation priority parameters and resource migration strategy parameters of each computing node can be dynamically determined, and resource dynamic allocation instructions can be generated accordingly, thereby realizing the intelligent and efficient allocation of supercomputing system resources, which not only significantly improves the resource utilization of the supercomputing system and reduces resource waste, but also optimizes the task execution process, shortens the task completion time, and improves the overall energy efficiency and stability of the supercomputing system. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 This is a schematic diagram of the execution flow of the method for intelligent allocation of supercomputing resources based on AI load monitoring provided by an embodiment of the present invention.
[0008] Figure 2 This is a schematic diagram of exemplary hardware and software components of the supercomputing resource intelligent allocation system based on AI load monitoring provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0009] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a method for intelligent allocation of supercomputing resources based on AI load monitoring provided by an embodiment of the present invention. The following is a detailed introduction to the method for intelligent allocation of supercomputing resources based on AI load monitoring.
[0010] Step S110: obtaining a real-time load data set of multiple computing nodes in a supercomputer system, wherein the real-time load data set includes resource occupancy indicators and task queue status parameters of each computing node within a continuous time window.
[0011] Supercomputing systems are run collaboratively by numerous computing nodes, each of which performs different computing tasks and is the fundamental unit that enables the system's powerful computing capabilities. To achieve intelligent allocation of supercomputing resources, we first collect real-time load data from multiple computing nodes.
[0012] The computing nodes mentioned here, such as Node A, Node B, and Node C in a supercomputing system, all have independent computing and task processing capabilities. A continuous time window is a pre-defined, uninterrupted time range that limits the period of data collection.
[0013] Resource utilization measures the percentage of various resources a compute node utilizes within a continuous time window. These resources include, but are not limited to, central processing unit (CPU), memory, and storage resources. For example, for node A, its CPU utilization reflects the extent to which the node's CPU is being used within the continuous time window. Memory utilization reflects the percentage of memory used.
[0014] Task queue status parameters primarily reflect the status of a compute node's task queue. A task queue is a collection of tasks awaiting processing by a compute node. Task queue status parameters include the order in which tasks are queued, which allows us to understand the order in which tasks enter the queue. They also include information about the waiting time of tasks in the queue, which can be used to assess task processing delays. By collecting these resource utilization metrics and task queue status parameters for each compute node within a continuous time window, we can obtain a real-time load data set.
[0015] Step S120: performing load feature extraction processing on the real-time load data set to generate a load feature set of each computing node, wherein the load feature set includes resource occupancy fluctuation features and task queue evolution features.
[0016] After obtaining the real-time load data set, since the original data is relatively complex, it is necessary to extract and process the load features to generate a load feature set that includes resource occupancy fluctuation features and task queue evolution features.
[0017] Step S121: performing data cleaning processing on the real-time load data set, removing the load data units corresponding to the abnormal time windows, and performing interpolation filling processing on the load data units of the missing time windows to generate a standardized load data set.
[0018] During the data collection process, the real-time load data set may be affected by various factors, resulting in data anomalies or missing data.
[0019] The load data units corresponding to abnormal time windows need to be identified and eliminated. There are many ways to identify abnormal data units, such as by setting a reasonable data boundary range for judgment. Taking the resource utilization rate indicator as an example, under normal circumstances, the CPU resource utilization rate should be within a reasonable range. If the CPU resource utilization rate recorded in a certain time window exceeds the range, such as an unreasonable value that is too high or too low, then the load data unit corresponding to the time window can be judged as abnormal. By the same token, for the task queue status parameters, if the task waiting time has a value that does not conform to the actual logic, it will also be regarded as abnormal data. The load data units corresponding to these abnormal time windows are then removed from the real-time load data set.
[0020] For load data units in missing time windows, interpolation is performed to ensure data integrity. For example, using the linear interpolation method, suppose that within a continuous time window, for a compute node, data for memory resource utilization is available at time points t1 and t3, M1 and M3, respectively, but data for time point t2 is missing. Since the time points are continuous, it can be assumed that memory resource utilization exhibits an approximately linear change over this period. The missing data is then filled using the following calculation logic: First, the ratio of the time interval between time points t2 and t1 to the time interval between time points t3 and t1 is calculated. Then, based on this ratio, the memory resource utilization at the missing time point t2 is calculated using M1 and M3. Similar processing is performed on load data units in all missing time windows to generate a standardized load data set. This standardized load data set provides more accurate and complete data.
[0021] Step S122: performing time series division processing on the standardized load data set to generate a plurality of load data subsets of continuous time segments.
[0022] The standardized load data set needs to be further divided into time series in order to analyze the characteristics of the standardized load data more finely.
[0023] Step S1221: Detect the timestamp continuity parameters of the load data units in each time window in the standardized load data set, perform segment marking processing on the load data units with timestamp jumps, and remove the segment if the number of jump time windows exceeds a preset threshold.
[0024] First, the timestamps of the load data units in each time window in the standardized load data set must be checked. The timestamp records the specific time point of data collection. Under normal circumstances, the timestamps should be continuously increasing. If it is found that the timestamp of a certain time window is discontinuous with the timestamp of the previous time window and there is a jump, that is, the time sequence is interrupted, then these load data units with timestamp jumps need to be segmented and marked. For example, according to the normal order, the time windows should be t1, t2, t3..., but if t1, t2, t4 appear in the record and t3 is skipped, then the load data unit corresponding to t4 is a timestamp jump, and the segment it is in needs to be marked.
[0025] At the same time, a preset threshold is set to measure the severity of timestamp jumps. If the number of jump time windows in a marked segment exceeds the preset threshold, it indicates that the data continuity of the segment is severely damaged, which may significantly interfere with subsequent analysis. Therefore, the segment needs to be removed from the standardized load data set. After this processing, the temporal continuity of the data will be more reliable, which facilitates the subsequent accurate analysis of load characteristics.
[0026] Step S1222: performing sliding window segmentation processing on the normalized load data set after the segmentation marking processing based on a preset time segment length threshold, and generating multiple load data subsets with equal time lengths.
[0027] After dealing with the problem of timestamp jumps, the standardized load data set that has been processed by segmentation marking needs to be segmented into sliding windows based on the preset time segment length threshold. The preset time segment length threshold is a fixed time length value determined according to the actual analysis requirements. For example, the time segment length threshold is set to a set number of time windows. Using this length as the window size, starting from the starting position of the standardized load data set, the data is segmented by moving the distance of a time window each time. Each segmentation will obtain a subset containing a fixed number of continuous time window load data. In this way, multiple load data subsets with equal time length are generated, thereby dividing the overall data according to a fixed time length, which is convenient for subsequent load feature analysis of each relatively independent time segment.
[0028] Step S1223: performing time dimension alignment processing on the resource occupancy rate indicators in each load data subset, so that the time intervals of the load data units in each time window are evenly distributed.
[0029] After each load data subset is generated, the time intervals between load data units in each time window may not be completely uniform due to slight time errors during the data collection process. To ensure the accuracy of subsequent analysis, the resource utilization indicators within each load data subset must be aligned along the time dimension.
[0030] The specific operation method is to use the time point of the first time window in the subset as the benchmark, and adjust the time points of the subsequent time windows according to the actual time intervals between each time window. For example, if the time interval between the time point originally recorded in the second time window and the first time window is longer than the expected standard time interval by a certain amount, then the resource utilization index data of the second time window is allocated to the increased time interval according to the set ratio, so that the time interval of each time window after adjustment is consistent with the expected standard time interval, thereby achieving a uniform distribution of the time intervals of the load data units of each time window. After the above processing, when analyzing the changes in the resource utilization index over time, the data in the time dimension will be more consistent and comparable.
[0031] Step S1224: performing queue identifier matching processing on the task queue state parameters in each load data subset to generate load data subsets of multiple consecutive time segments.
[0032] After aligning the resource utilization metrics over time, we then perform queue identifier matching on the task queue status parameters within each load data subset. Each task is assigned a unique queue identifier upon entering the task queue. This identifier identifies the task's identity and position within the queue.
[0033] By matching the task queue status parameters in each load data subset with queue identifiers, we can ensure that the relevant parameters of the same task in different time windows can be accurately associated. For example, for a certain compute node's load data subset, the relevant information of tasks in the task queue is recorded in different time windows. By matching queue identifiers, we can accurately match the parameters of the same task such as the waiting time and task status changes in different time windows, thus avoiding the confusion of task information. After the above processing, multiple load data subsets of consecutive time segments are generated.
[0034] Step S123: performing resource occupancy fluctuation analysis on each load data subset, and extracting resource occupancy change gradient parameters and resource occupancy peak duration parameters of each computing node in the corresponding time segment.
[0035] After obtaining multiple load data subsets through time series partitioning, it is necessary to perform resource occupancy fluctuation analysis on each load data subset to extract key characteristic parameters.
[0036] Step S1231: performing smoothing filtering on the resource occupancy rate indicators in the load data subset to generate a filtered resource occupancy rate time series.
[0037] Since the original resource utilization index data may be interfered by various noises and show irregular fluctuations, this situation is not conducive to accurately analyzing its changing trend. The role of smoothing filtering is to remove these noise interferences and make the data smoother. For example, a moving average filtering method is used to set a filter window size, assuming it is a set number of time windows. For the resource utilization index of each time window, the average value of its own resource utilization index and the resource utilization index of the adjacent time windows is taken as the resource utilization value after filtering of the time window. In this way, each time window in the load data subset is processed in turn to generate a filtered resource utilization time series sequence, which can more clearly reflect the real changing trend of the resource utilization.
[0038] Step S1232: Calculate the absolute value of the difference of the filtered resource occupancy time series in adjacent time windows to generate a resource occupancy change gradient parameter.
[0039] The resource utilization rate change gradient of adjacent time windows can reflect how quickly resource utilization rates change over a short period of time. For every two adjacent time windows in the filtered resource utilization rate time series, subtract the resource utilization rate value of the previous time window from the resource utilization rate value of the latter time window, and then take the absolute value. The result is the resource utilization rate change gradient parameter within the two adjacent time windows. For example, in the filtered time series, the resource utilization rate value of time window t1 is R1, and the resource utilization rate value of time window t2 is R2. Then the resource utilization rate change gradient parameter of the two adjacent time windows is R2 minus the absolute value of R1. By calculating the change gradient parameter of each adjacent time window, we can fully understand the speed of change of resource utilization rates within that time segment.
[0040] Step S1233: Detecting peak points exceeding a preset threshold in the filtered resource occupancy time series, and recording the start timestamp and end timestamp corresponding to each peak point.
[0041] The preset threshold is a reference value set based on the historical data of the supercomputing system and actual operating experience. The preset threshold is used to determine whether the resource utilization rate has reached a peak. When the filtered resource utilization rate value exceeds the preset threshold, and within the set range before and after the point, for example, a time window before and after, the value of the point is the largest, then the point is determined to be the peak point. Record the starting timestamp of each peak point, that is, the time point when the peak point first exceeds the preset threshold, and the end timestamp, that is, the time point when the peak point last exceeds the preset threshold. In this way, the specific time period when the resource utilization rate peaks within the time segment can be accurately captured.
[0042] Step S1234: Calculate the resource occupancy peak duration parameter according to the difference between the start timestamp and the end timestamp, and perform cumulative summation processing on the duration parameters of consecutive peak points.
[0043] In this embodiment, the duration parameter of each peak point is obtained by subtracting the start timestamp from the end timestamp of each peak point. If there are multiple consecutive peak points within a time segment, that is, the end timestamp of the previous peak point is adjacent to or overlaps with the start timestamp of the next peak point, the duration parameters of these consecutive peak points are accumulated and summed to obtain a comprehensive parameter for the duration of the resource utilization peak within the time segment. This resource utilization peak duration parameter reflects the total duration of resource utilization at peak level and is important for evaluating the performance of computing nodes during peak resource utilization periods.
[0044] Step S124: performing task queue evolution analysis on each load data subset, extracting the trend parameters of the number of pending tasks, the cumulative task execution delay parameters, and the task queue mutation amplitude parameters of each computing node in the corresponding time segment.
[0045] After analyzing the resource utilization fluctuations, the next step is to perform task queue evolution analysis on each load data subset to obtain key characteristic parameters related to the task queue.
[0046] Step S1241 , parsing the task queue state parameters in the load data subset, and extracting the time series of the number of tasks to be processed and the time series of the average execution delay of the tasks of each computing node.
[0047] The task queue status parameters contain a wealth of task queue information. By parsing them, we can isolate time-varying sequences of the number of pending tasks and the average task execution delay. For example, we extract the number of pending tasks within each time window from the task queue status parameters in chronological order, forming a time series of the number of pending tasks. Simultaneously, we extract the average task execution delay within each time window to form a time series of the average task execution delay. These time series of average task execution delays reflect the dynamic changes in the task queue within that time window.
[0048] Step S1242 , performing linear fitting processing on the time series of the number of tasks to be processed, and calculating the slope parameter of the fitting curve as the trend parameter of the number of tasks to be processed.
[0049] Linear fitting is a process that uses mathematical methods to find a straight line that best represents the trend of data changes. For a time series of the number of pending tasks, assuming the series contains multiple time windows with corresponding values for the number of pending tasks, a linear fitting method such as the least squares method is used to find a straight line that minimizes the sum of the squared errors between the line and these data points. The slope of this fitted line represents the trend of the number of pending tasks over time. A positive slope indicates an upward trend in the number of pending tasks; a negative slope indicates a downward trend. This pending task trend parameter provides an intuitive understanding of the increase or decrease in the number of pending tasks in the task queue.
[0050] Step S1243 : performing an integration operation on the task average execution delay timing sequence to calculate a task execution delay accumulation parameter within a preset time window.
[0051] In mathematics, integration is used to calculate the area under a function curve. Here, it can be understood as accumulating the average task execution delays within each time window. For a time series of average task execution delays, the average task execution delays within each time window are accumulated and summed to obtain the cumulative task execution delay parameter. This cumulative task execution delay parameter reflects the overall task execution delay within the preset time window and is an important reference for evaluating task processing efficiency.
[0052] Step S1244 , detecting a mutation point in the time series of the number of tasks to be processed, and calculating the absolute value of the difference between the number of tasks before and after the mutation point as a task queue mutation amplitude parameter.
[0053] A mutation point is a point at which the number of pending tasks changes significantly at a specific moment. By observing the time series of the number of pending tasks, we can identify a mutation point when the number of pending tasks in a given time window differs significantly from the number in the preceding and following time windows. The task queue mutation amplitude parameter is obtained by subtracting the number of pending tasks in the time window after the mutation point from the number of pending tasks in the time window before the mutation point and taking the absolute value. This parameter reflects the degree of sudden change in the number of tasks in the task queue during that time period and helps to promptly detect abnormal changes in the task queue.
[0054] Step S125: Input the resource occupancy rate change gradient parameter and the resource occupancy rate peak duration parameter into the feature encoder for normalized feature splicing processing to generate the resource occupancy fluctuation feature vector, and input the pending task quantity change trend parameter, the task execution delay accumulation parameter and the task queue mutation amplitude parameter into the feature encoder for normalized feature splicing processing to generate the task queue evolution feature vector, and output the resource occupancy fluctuation feature and the task queue evolution feature as the load feature set.
[0055] After extracting parameters related to resource occupancy fluctuation and task queue evolution, these parameters need to be further processed to generate a load feature set.
[0056] For the resource occupancy fluctuation feature, the resource occupancy rate change gradient parameter and the resource occupancy rate peak duration parameter are input into the feature encoder. The function of the feature encoder is to normalize the input parameters and perform feature splicing. The purpose of normalization is to convert parameters of different dimensions and value ranges to the same scale, which facilitates subsequent analysis and model processing. For example, assuming that the value range of the resource occupancy rate change gradient parameter is one range, and the value range of the resource occupancy rate peak duration parameter is another range, through normalization methods such as minimum-maximum normalization, the resource occupancy rate change gradient parameter is mapped to between 0 and 1, and the resource occupancy rate peak duration parameter is also mapped to between 0 and 1. Then, the two normalized parameters are spliced to form the resource occupancy fluctuation feature.
[0057] For the task queue evolution feature, similar processing is performed on the parameters for the trend of the number of pending tasks, the cumulative task execution delay, and the task queue mutation amplitude. Similarly, these three parameters are first normalized to bring them into the same scale range. For example, the trend of the number of pending tasks may have a range of values, the cumulative task execution delay may have a range of values, and the task queue mutation amplitude may have a range of values. Through normalization, these three parameters are mapped to a range between 0 and 1. The three normalized parameters are then concatenated to generate the task queue evolution feature vector.
[0058] Finally, the resource occupancy fluctuation feature vector and the task queue evolution feature vector are combined to form a load feature set, which comprehensively reflects the load characteristics of the computing node.
[0059] Step S130: calling a pre-built resource demand prediction model to perform resource demand analysis on the load feature set, and generating a resource demand prediction result for each computing node in a subsequent time window.
[0060] After obtaining the load feature set of each computing node, it is necessary to use the pre-built resource demand prediction model to analyze it, so as to predict the resource demand of each computing node in the subsequent time window.
[0061] Step S131: inputting the load feature set into the feature coding layer of the resource demand prediction model, performing correlation modeling processing on the resource occupancy fluctuation feature and the task queue evolution feature, and generating a fusion feature vector.
[0062] The resource demand forecasting model first processes the input load feature set through the feature encoding layer. The resource utilization fluctuation characteristics and task queue evolution characteristics within the load feature set are input into the feature encoding layer. The feature encoding layer typically consists of multiple neural network layers designed to explore and characterize the potential connections between these two characteristics.
[0063] Taking the multi-layer perceptron architecture as an example, the starting layer of the feature encoding layer receives the resource utilization fluctuation feature vector and the task queue evolution feature vector as input. The resource utilization fluctuation feature vector is assumed to include information on dimensions such as the gradient of resource utilization change and the duration of resource utilization peaks, while the task queue evolution feature vector includes information on dimensions such as the changing trend of the number of pending tasks, the accumulated task execution delays, and the magnitude of the task queue mutation. The neurons in the starting layer establish connections with each dimension of these input vectors and assign a set weight to each connection.
[0064] Specifically, each dimension of the resource utilization fluctuation feature vector and each dimension of the task queue evolution feature vector are multiplied by their corresponding weights, and all the multiplication results are accumulated to obtain the input value of each neuron. These weights are not fixed but are continuously adjusted during model training based on data feedback to enable the model to better capture the correlation between features. For example, if a potential connection is found between the gradient of resource utilization changes and the trend of changes in the number of tasks to be processed, the model will adjust the corresponding weights to strengthen the reflection of this connection in the neuron output.
[0065] After receiving the neuron's input value, it is passed to the activation function. The activation function introduces nonlinear characteristics to the model, enabling it to learn more complex feature relationships. For example, the Sigmoid activation function maps the input value to a range between 0 and 1, and outputs a corresponding nonlinear transformation based on the input value, which serves as the neuron's output.
[0066] The outputs of all neurons in the initial layer form a new vector, which serves as the input to the next layer, repeating the above calculation process. After multiple layers of processing, a fused feature vector is ultimately generated. This fused feature vector integrates information related to resource usage fluctuations and task queue evolution, providing a more valuable feature representation for subsequent models to analyze resource requirements.
[0067] Step S132: performing temporal dependency extraction processing on the fused feature vector through the temporal convolution layer of the resource demand prediction model to generate a temporal feature representation with long-term and short-term dependencies.
[0068] The temporal convolutional layer plays a key role in the resource demand prediction model. It is responsible for extracting temporal dependencies from the fused feature vector to generate temporal feature representations with long-term and short-term dependencies. The temporal convolutional layer primarily accomplishes this task through the collaborative efforts of several modules: a dilated causal convolution module, a gated activation function module, a residual connection module, and a layer-by-layer normalization module.
[0069] Step S1321: Input the fused feature vector into the dilated causal convolution module of the temporal convolution layer, use convolution kernels with different dilation rates to perform multi-scale temporal feature extraction on the fused feature vector, and obtain the convolution output result corresponding to each dilation rate.
[0070] The fused feature vector is fed into a dilated causal convolution module. This module uses convolution kernels with different dilation rates to capture the temporal characteristics of the fused feature vector at different scales. The dilation rate determines the sampling interval of the convolution kernel in the temporal dimension.
[0071] For example, when using multiple dilation rates, convolution kernels with different dilation rates sample the fused feature vector differently in the time dimension during convolution operations. For example, with a smaller dilation rate, the convolution kernel samples feature values at relatively close intervals in the time dimension, which helps capture short-term temporal feature changes. However, with a larger dilation rate, the convolution kernel samples at wider intervals, enabling the capture of long-term temporal feature information.
[0072] For each position of the fused feature vector in the time dimension, the convolution kernel corresponding to the dilation rate is calculated by combining the eigenvalues of the current position and other positions separated by the dilation rate. The specific calculation process is to multiply the eigenvalues of these positions by the weights of the convolution kernel, and then accumulate the products to obtain the output value after convolution at that position. In this way, for each convolution kernel with a different dilation rate, the fused feature vector is convolved, resulting in the convolution output result corresponding to each dilation rate. These convolution output results at different dilation rates reflect the characteristic information of the fused feature vector in the time dimension from multiple scales.
[0073] Step S1322: performing channel normalization processing on the convolution output results corresponding to each expansion rate, and inputting the normalized multi-scale time series features into the gated activation function module to generate a time series feature map with a nonlinear relationship.
[0074] After obtaining the convolution output corresponding to each dilation rate, it is necessary to perform channel normalization on it. Channel normalization aims to make the data of different channels have similar distribution characteristics for subsequent processing.
[0075] For each channel in the convolutional output, the mean and variance of all elements within that channel are calculated. Then, each element within that channel is normalized by subtracting the channel mean from the value and then dividing it by the channel standard deviation. This process makes the data distribution across channels more consistent, helping the model better learn and utilize these features.
[0076] The normalized multi-scale time series features are input into the gated activation function module. This module uses a predefined mechanism to filter and transform the input features to generate a time series feature map with nonlinear relationships. The gated activation function module dynamically adjusts its output based on the input features, highlighting features that are more critical for subsequent analysis and suppressing less important features. For example, for features that vary significantly over time and may be important indicators for resource demand forecasting, the gated activation function will enhance their representation in the output; while for relatively stable features that have less impact on forecasting, their influence may be appropriately weakened.
[0077] Step S1323: performing residual connection processing on the temporal feature map, and performing element-by-element addition processing on the original fused feature vector and the temporal feature map after gate activation.
[0078] After the gated activation function is processed, the generated time series feature map is processed with residual connections. Residual connections are used to solve the gradient vanishing or gradient exploding problems that may occur during the training of deep neural networks, which helps the model learn and optimize better.
[0079] The specific operation involves element-by-element addition of the original fused feature vector and the gated activated temporal feature map. The original fused feature vector retains the comprehensive information of the original input, while the gated activated temporal feature map incorporates the results of the temporal convolutional layer's mining of temporal dependencies. This element-by-element addition preserves the key information in the original features while incorporating the newly extracted temporal feature variations, resulting in a richer and more comprehensive feature representation.
[0080] Step S1324: The time series feature map after residual processing is input into the hierarchical normalization module to adjust the feature distribution and generate a time series feature representation with long-term and short-term dependencies.
[0081] The residual-processed time series feature map is fed into the hierarchical normalization module. This module normalizes the entire feature vector at different levels, further optimizing the distribution of features and making the scale of features more consistent across different dimensions.
[0082] The hierarchical normalization module comprehensively considers the statistical information of feature vectors at different levels and adjusts them. For example, it calculates the mean and variance of feature vectors at different levels and then normalizes the elements in the feature vectors based on these statistics. After processing by the hierarchical normalization module, a time series feature representation with long-term and short-term dependencies is generated. This time series feature representation comprehensively and accurately reflects the dependencies between the fused feature vectors in the time series.
[0083] Step S133: Call the multi-head attention layer of the resource demand prediction model to perform key time node identification processing on the temporal feature representation, and generate an attention weight distribution diagram of different time nodes.
[0084] The multi-head attention layer plays an important role in identifying key time nodes in the resource demand prediction model. By processing the time series feature representation, it generates the attention weight distribution map of different time nodes.
[0085] Step S1331: dividing the time series feature representation into multiple sub-feature vectors, each sub-feature vector corresponding to the feature representation of a different time node.
[0086] The time series feature representation with long-term and short-term dependencies is segmented along the time dimension to form multiple sub-feature vectors. Each sub-feature vector corresponds to the feature information of a different time node. For example, assuming the time series feature representation covers feature changes over a period of time, it is segmented according to a set time interval. The feature information of each segment constitutes a sub-feature vector. Each sub-feature vector contains a comprehensive expression of relevant features such as resource usage fluctuations and task queue evolution around that time node.
[0087] Step S1332: Perform linear projection processing on each sub-feature vector to generate a query vector, a key vector, and a value vector.
[0088] Each sub-eigenvector is linearly projected. Linear projection is achieved by multiplying the sub-eigenvector with a set weight matrix. For each sub-eigenvector, a different weight matrix is used to generate the query vector, key vector, and value vector.
[0089] Specifically, for a sub-feature vector, it is multiplied by the query weight matrix to obtain the corresponding query vector; multiplied by the key weight matrix to obtain the key vector; and multiplied by the value weight matrix to obtain the value vector. These weight matrices are continuously optimized during the model training process so that the generated query vector, key vector, and value vector can be effectively used for subsequent key time node identification.
[0090] Step S1333: Calculate the similarity scores between the query vector and the key vector at different time nodes to generate an initial attention score matrix.
[0091] Calculate the similarity between the query vectors and key vectors corresponding to different time nodes. For each query vector, calculate the similarity score with the key vectors of all other time nodes. For example, use the dot product operation to calculate the similarity score between the query vector and the key vector. Arrange the similarity scores calculated for a query vector and all key vectors together to form a row of data. Perform this operation on all query vectors, and finally generate an initial attention score matrix. Each element in this initial attention score matrix represents the similarity measure between the query vector and the key vector at different time nodes.
[0092] Step S1334: Mask the initial attention score matrix to shield the attention impact of future time nodes on historical time nodes.
[0093] Since in practice, the information of future time nodes should not affect the analysis of historical time nodes, it is necessary to perform masking on the initial attention score matrix. Masking is achieved by setting a mask matrix that sets the positions of future time nodes to 0 at the positions corresponding to the historical time nodes, while keeping the other positions unchanged.
[0094] The specific operation is to multiply the initial attention score matrix with the corresponding elements of the mask matrix, thereby shielding the attention influence of future time nodes on historical time nodes. After the masking process, the attention score matrix only reflects the relationship between time nodes that conform to the chronological logic.
[0095] Step S1335: Input the masked attention score matrix into the Softmax function for normalization to generate an attention weight distribution diagram at different time nodes.
[0096] The masked attention score matrix is input into the Softmax function. The Softmax function converts the elements in the matrix into a probability distribution form so that the sum of all elements is 1.
[0097] For each row of data in the masked attention score matrix, the Softmax function converts it into a set of probability values based on the numerical value of the row. These probabilities represent the degree of attention the model pays to each time node, namely the attention weight. The probability values for all rows are arranged together to generate a distribution graph of attention weights at different time nodes. This distribution graph intuitively shows how the model allocates attention at different time nodes, highlighting time nodes that may be more critical for resource demand prediction.
[0098] Step S1336: Perform weighted summation processing on the value vector based on the attention weight distribution map to generate a context-aware temporal feature representation.
[0099] Based on the generated attention weight distribution map, a weighted summation process is performed on the value vector corresponding to each time node. The attention weight of each time node is multiplied by the value vector corresponding to that time node, and all the products are accumulated to obtain a context-aware time series feature representation. This weighted summation process comprehensively considers the importance of different time nodes, generating a more representative time series feature representation and providing more targeted feature information for subsequent resource demand forecasting.
[0100] Step S134: Perform dynamic weighted aggregation processing on the temporal feature representation based on the attention weight distribution map to generate an aggregated feature vector that strengthens the key time node information.
[0101] Based on the generated attention weight distribution map, the temporal feature representation is dynamically weighted and aggregated to enhance the information of key time nodes and generate aggregated feature vectors.
[0102] The attention weight distribution diagram reflects the relative importance of different time nodes in resource demand prediction. For time series feature representation, the sub-feature vector corresponding to each time node is weighted according to the weight of the corresponding time node in the attention weight distribution diagram. That is, for each sub-feature vector, the value of each dimension is multiplied by the attention weight of the corresponding time node.
[0103] All weighted sub-feature vectors are then aggregated. Aggregation involves concatenating these weighted sub-feature vectors in a set order to form an aggregated feature vector that enhances information about key time nodes. This dynamic weighted aggregation process enhances the feature information of key time nodes in the aggregated feature vector, allowing the model to focus more on the information contained in these key time nodes in subsequent forecasts, thereby improving the accuracy of resource demand forecasts.
[0104] Step S135: inputting the aggregated feature vector into the regression output layer of the resource demand prediction model to perform resource demand prediction processing, and generating resource demand prediction results for each computing node in a subsequent time window.
[0105] The aggregated feature vectors that reinforce the key time node information are input into the regression output layer of the resource demand prediction model. The regression output layer is designed to predict the resource demand of each computing node in the subsequent time window based on the input aggregated feature vectors.
[0106] The regression output layer typically consists of one or more fully connected layers. The aggregated feature vector is fed into the first fully connected layer as input. Neurons in this fully connected layer establish connections with each dimension of the input vector and assign a weight to each connection. The neuron multiplies each dimension of the input vector by its corresponding weight, then adds the products and adds a bias term to produce the neuron's input value.
[0107] This input value is then processed by an activation function, such as a linear activation function (which directly outputs the input value), to obtain the output value of the neuron. The outputs of all neurons in the first fully connected layer form a new vector, which serves as the input of the next fully connected layer, and the above calculation process is repeated.
[0108] After processing through multiple fully connected layers, the final output is a vector whose elements correspond to the predicted demand for different types of resources for each compute node in the subsequent time window. For example, some elements in the vector represent the predicted demand for CPU resources, while others represent the predicted demand for memory resources. This process completes the conversion from the aggregated feature vector to the predicted resource demand for each compute node in the subsequent time window.
[0109] Step S140: Determine resource allocation priority parameters and resource migration strategy parameters for each computing node in the supercomputing system based on the resource demand prediction result.
[0110] After obtaining the resource demand prediction results of each computing node in the subsequent time window, the next step is to determine the resource allocation priority parameters and resource migration strategy parameters of each computing node in the supercomputing system based on these results.
[0111] Step S141: parsing the predicted resource occupancy rate index and the predicted task queue length index of each computing node in the resource demand prediction result.
[0112] The resource demand forecast results are analyzed to extract the predicted resource utilization and task queue length indicators for each compute node. The predicted resource utilization indicator reflects the predicted utilization ratio of various resources by the compute node in the subsequent time window, such as the predicted utilization of resources such as CPU and memory. The predicted task queue length indicator reflects the number of tasks waiting to be processed in the compute node's task queue.
[0113] For example, for a computing node, the resource demand forecast result includes resource utilization indicators such as the predicted CPU resource utilization and memory resource utilization, as well as the predicted number of tasks to be processed in the task queue, that is, the predicted task queue length indicator.
[0114] Step S142: Calculating a resource scarcity parameter of each computing node according to the predicted resource occupancy rate indicator, wherein the resource scarcity parameter is positively correlated with the predicted resource occupancy rate indicator.
[0115] Based on the extracted predicted resource utilization index, the resource scarcity parameter for each computing node is calculated. Since resource scarcity is closely related to the predicted resource utilization, generally speaking, a higher predicted resource utilization indicates a more urgent resource demand for the computing node and a higher resource scarcity. This means that the resource scarcity parameter and the predicted resource utilization index are positively correlated.
[0116] When calculating the resource scarcity parameter, a mapping relationship may be used to convert the predicted resource utilization indicator into the resource scarcity parameter. For example, a monotonically increasing function may be used to map the predicted resource utilization indicator to a set numerical range representing the resource scarcity level. Specifically, if the predicted resource utilization indicator is within a small range, the corresponding resource scarcity parameter is also low. As the predicted resource utilization indicator increases, the resource scarcity parameter increases accordingly to reflect the changes in the computing node's resource scarcity.
[0117] Step S143: calculating a task backlog risk parameter of each computing node according to the predicted task queue length indicator, wherein the task backlog risk parameter is positively correlated with a growth rate of the predicted task queue length indicator.
[0118] Based on the predicted task queue length indicator, we calculate the task backlog risk parameter for each compute node. The task backlog risk is closely related to the growth of the predicted task queue length. A higher growth rate in the predicted task queue length indicator indicates a greater likelihood of task backlogs and a higher backlog risk. Therefore, the task backlog risk parameter is positively correlated with the growth rate of the predicted task queue length indicator.
[0119] When calculating the task backlog risk parameter, you first need to calculate the growth rate of the predicted task queue length indicator. You can calculate the growth rate by comparing the current predicted task queue length with the predicted task queue length at a previous point in time. For example, subtract the predicted task queue length at the previous point in time from the current predicted task queue length, and then divide the result by the predicted task queue length at the previous point in time to obtain the growth rate of the predicted task queue length.
[0120] The growth rate is then converted into a task backlog risk parameter using a function related to the growth rate. Similar to the calculation of the resource scarcity parameter, a monotonically increasing function is used to map the growth rate to a numerical range representing the task backlog risk. Higher growth rates correspond to larger task backlog risk parameters, reflecting the degree of task backlog risk at the compute node.
[0121] Step S144: normalize the resource scarcity parameter and the task backlog risk parameter, and perform linear weighted summation on the normalized resource scarcity parameter and the task backlog risk parameter according to a preset resource allocation strategy weight coefficient to generate a comprehensive priority score.
[0122] In order to comprehensively consider the impact of resource scarcity and task backlog risk on resource allocation priorities, we first normalize the resource scarcity parameter and task backlog risk parameter. Normalization converts these two parameters to the same numerical range to facilitate comparison and comprehensive calculation.
[0123] For example, the minimum-maximum normalization method is used to map both the resource scarcity parameter and the task backlog risk parameter to a range of 0 to 1. Specifically, for the resource scarcity parameter, subtract its minimum value across all computing nodes from it, and then divide it by the difference between the maximum and minimum values to obtain the normalized resource scarcity parameter. The same operation is performed on the task backlog risk parameter to obtain the normalized task backlog risk parameter.
[0124] Next, a linear weighted summation of the normalized resource scarcity and task backlog risk parameters is performed based on the preset resource allocation strategy weight coefficients to generate a comprehensive priority score. The preset resource allocation strategy weight coefficients are pre-set based on multiple factors, including the actual operational requirements of the supercomputing system, resource characteristics, and task type. They determine the weight of the resource scarcity and task backlog risk parameters in the comprehensive priority score.
[0125] Assume that the normalized resource scarcity parameter is represented as A, the normalized task backlog risk parameter is represented as B, and the resource allocation strategy weight coefficients are α and β, respectively, where α represents the weight of the resource scarcity parameter and β represents the weight of the task backlog risk parameter, and α + β = 1. The comprehensive priority score is then calculated as: Comprehensive priority score = α × A + β × B.
[0126] Through such linear weighted summation, the two parameters reflecting different aspects are integrated into a comprehensive priority score, which can comprehensively reflect the priority of each computing node in terms of resource requirements and task processing status.
[0127] Step S145: Detect computing nodes whose comprehensive priority scores exceed a preset threshold and mark them as key computing nodes.
[0128] After obtaining the comprehensive priority scores for each compute node, it is necessary to screen out those with more urgent resource requirements and task processing status so that resource allocation can be prioritized and corresponding strategies can be adopted. This requires detecting compute nodes whose comprehensive priority scores exceed a preset threshold and marking them as critical compute nodes.
[0129] The preset threshold is a standard value set based on factors such as the overall resource status of the supercomputing system, general task execution requirements, and past experience. This threshold is used to prioritize compute nodes. Exceeding this threshold indicates that the compute node has relatively more urgent resource needs and task processing requirements, requiring special attention and priority processing.
[0130] The comprehensive priority score of each compute node is compared against a preset threshold. If a compute node's comprehensive priority score exceeds the preset threshold, it is identified as a critical compute node and marked. For example, these critical compute nodes are assigned a set identifier so that they can be quickly identified and distinguished from other compute nodes in subsequent processing. This screening and marking process clearly identifies which compute nodes in the supercomputing system have higher priority in resource allocation and task processing, providing clear goals for subsequent resource migration and allocation strategy formulation.
[0131] Step S146: adjusting the resource allocation priority parameter according to the distribution density parameter of the key computing nodes, wherein the resource allocation priority parameter is negatively correlated with the distribution density of the key computing nodes.
[0132] After determining the key computing nodes, the resource allocation priority parameters are further adjusted based on their distribution. The density of key computing nodes reflects the degree of concentration or dispersion of these important nodes in the supercomputing system, and the resource allocation priority parameters need to be dynamically adjusted based on this distribution. The two are negatively correlated: the denser the distribution of key computing nodes, the lower their resource allocation priority parameters; the more dispersed the distribution, the higher their resource allocation priority parameters.
[0133] First, calculate the distribution density parameters of key computing nodes. This requires determining the location information of key computing nodes in the supercomputing system (for example, their position in the supercomputing system's topology) and then calculating their distribution density. For example, the distribution density can be quantified by calculating the average distance between key computing nodes or the number of key computing nodes in a certain area. Assuming the method of calculating the average distance between key computing nodes is used, first calculate the distance between every two key computing nodes, add all these distance values, and then divide them by the total number of key computing nodes in each pair to obtain the average distance. The smaller the average distance, the denser the distribution of key computing nodes; conversely, the larger the average distance, the more dispersed the distribution.
[0134] Based on the average distance, a distribution density parameter is calculated using a pre-set functional relationship. This functional relationship is usually monotonically decreasing, that is, the smaller the average distance, the larger the corresponding distribution density parameter; the larger the average distance, the smaller the corresponding distribution density parameter.
[0135] The resource allocation priority parameter is then adjusted based on the distribution density parameter of key compute nodes. Since the resource allocation priority parameter is negatively correlated with the distribution density of key compute nodes, adjustment can be achieved using another pre-defined function. For example, the resource allocation priority parameter is multiplied by a coefficient related to the distribution density parameter, which decreases as the distribution density parameter increases. Assuming the distribution density parameter is D, a pre-defined adjustment coefficient function f(D) is specified such that f(D) decreases as D increases. The adjusted resource allocation priority parameter is then calculated as: original resource allocation priority parameter × f(D). This adjustment allows the resource allocation priority parameter to be adjusted appropriately based on the distribution of key compute nodes, providing more scientific guidance for resource allocation.
[0136] Step S147: Generate a resource allocation order list for each key computing node based on the adjusted resource allocation priority parameters, and determine resource migration strategy parameters based on the resource allocation order list for each computing node, wherein the resource migration strategy parameters include a migration target node identifier and a migration resource quantity parameter.
[0137] For example, step S1471: the key computing nodes are arranged in descending order according to the resource allocation priority parameter to generate an initial resource allocation sequence list.
[0138] This step involves performing a preliminary sorting of key compute nodes based on the adjusted resource allocation priority parameters. Since higher resource allocation priority parameters indicate a more urgent resource need for a key compute node, a descending sorting order is used, placing the most urgent key compute nodes at the top of the list. For example, assuming there are multiple key compute nodes, each with its own adjusted resource allocation priority parameter, these key compute nodes are sorted in descending order based on their resource allocation priority parameters to form an initial linear list. This is the initial resource allocation order list.
[0139] Step S1472: extracting the topological position coordinate parameters of each key computing node in the initial resource allocation order list, and calculating the distance set between adjacent key computing nodes based on the topological position coordinate parameters.
[0140] After generating the initial resource allocation order list, to further optimize the resource allocation order, it is necessary to consider the location information of key computing nodes in the supercomputing system topology. The topological position coordinate parameters of each key computing node in the initial resource allocation order list are extracted. These topological position coordinate parameters represent the location of the key computing node in the supercomputing system topology space.
[0141] Based on these topological location coordinate parameters, the distances between adjacent key computing nodes are calculated. For each pair of adjacent key computing nodes in the initial resource allocation order list, their topological location coordinates are used to calculate the distance between them using a distance calculation formula (e.g., a method similar to the Euclidean distance formula in two or three-dimensional space). The distance values of all adjacent key computing nodes are collected to form a distance set, which reflects the adjacent distances between key computing nodes in the topological structure.
[0142] Step S1473: Calculate the regional distribution density parameter of each key computing node according to the interval set, where the regional distribution density parameter is negatively correlated with the mean value of the interval set.
[0143] Once we have a set of distances between adjacent key computation nodes, we calculate the mean of this set to measure the average distance between them. We sum all the distances in the set and divide it by the number of elements in the set to get the mean.
[0144] The regional distribution density parameter is negatively correlated with the mean of the distance set. Specifically, the smaller the average distance, the denser the distribution of key computing nodes within a region, and the larger the regional distribution density parameter. Conversely, the larger the average distance, the smaller the regional distribution density parameter. The regional distribution density parameter is calculated using a pre-defined function, which is typically monotonically decreasing. For example, assuming the mean of the distance set is M, the regional distribution density parameter is calculated using the function g(M), where g(M) decreases as M increases. This yields the regional distribution density parameter for each key computing node, which quantifies the density of the regional distribution of key computing nodes within the supercomputing system topology.
[0145] Step S1474: performing a normalized conversion process on the regional distribution density parameter and the resource allocation priority parameter, and then performing an inverse proportional weighting process to generate a revised resource allocation priority parameter.
[0146] To comprehensively consider the regional distribution density and resource allocation priority of key computing nodes, the regional distribution density parameter and the resource allocation priority parameter are normalized. Normalization converts these two parameters to the same numerical range, facilitating unified calculation and comparison. For example, a minimum-maximum normalization method is used to map both the regional distribution density parameter and the resource allocation priority parameter to a range of 0 to 1.
[0147] The specific operation is to subtract the minimum value of the regional distribution density parameter among all key computing nodes from the parameter, and then divide it by the difference between the maximum and minimum values to obtain the normalized regional distribution density parameter; perform the same operation on the resource allocation priority parameter to obtain the normalized resource allocation priority parameter.
[0148] Then, an inversely proportional weighting process is performed on the standardized regional distribution density parameter and resource allocation priority parameter. Because the regional distribution density parameter and the resource allocation priority parameter are inversely proportional (i.e., the denser the distribution, the lower the priority), an inversely proportional weighting method is used to generate the revised resource allocation priority parameter. For example, the normalized resource allocation priority parameter is divided by the normalized regional distribution density parameter and then multiplied by a pre-set weight coefficient (which adjusts the strength of the inverse relationship) to obtain the revised resource allocation priority parameter. This process enables the resource allocation priority parameter to more accurately reflect the urgency of resource requirements of key computing nodes, taking into account regional distribution.
[0149] Step S1475: performing secondary sorting processing on the key computing nodes based on the revised resource allocation priority parameters to generate a final resource allocation sequence list.
[0150] After obtaining the revised resource allocation priority parameters, the key computing nodes are re-sorted based on these parameters. The key computing nodes are rearranged from high to low according to the revised resource allocation priority parameters (or according to the actual set sorting rules). For example, if a higher revised resource allocation priority parameter indicates higher resource priority, the key computing nodes are sorted from high to low according to the revised resource allocation priority parameters to obtain the final resource allocation order list. This final resource allocation order list comprehensively considers the resource demand urgency of the key computing nodes and their regional distribution in the supercomputing system topology, and is more scientific and reasonable than the initial resource allocation order list.
[0151] Step S1476: traverse the final resource allocation order list, and generate a resource gap parameter according to the difference between the predicted resource occupancy index of each key computing node and the current resource availability parameter.
[0152] Traverse each key computing node in the final resource allocation order list and calculate its resource gap parameter for each key computing node. The resource gap parameter reflects the gap between the current resource demand of the key computing node and the available resources.
[0153] The calculation method is based on the difference between the predicted resource utilization indicator and the current resource availability parameter for each key computing node. The predicted resource utilization indicator reflects the proportion of resources that the key computing node is expected to occupy in the subsequent time window, and the current resource availability parameter represents the actual amount of available resources currently possessed by the key computing node. The predicted resource utilization indicator is converted into actual resource demand (for example, if the predicted resource utilization indicator is for CPU resources and the total CPU resources of the current key computing node are known, then the predicted resource utilization indicator is multiplied by the total CPU resources to obtain the predicted CPU resource demand). Then, the difference between the predicted resource demand and the current resource availability parameter is the resource gap parameter. By calculating the resource gap parameter for each key computing node, the specific resource shortage situation of each key computing node is clarified.
[0154] Step S1477: Filter non-critical computing nodes in the supercomputing system whose resource scarcity parameter is lower than a preset lower limit and whose current available resources are greater than the resource gap parameter as a migration source node set.
[0155] After determining the resource shortage parameters for each key computing node, it is necessary to find suitable source nodes for resource migration within the supercomputing system. Select non-key computing nodes in the supercomputing system whose resource scarcity parameters are below the preset lower limit and whose current available resources are greater than the resource shortage parameters as the migration source node set.
[0156] The preset lower limit is a standard value set based on the overall resource status and operational requirements of the supercomputing system. It is used to measure whether the resource scarcity of non-critical computing nodes is within an acceptable range, meaning that the node has sufficient resources for migration. If the resource scarcity parameter is below the preset lower limit, the non-critical computing node has relatively sufficient resources.
[0157] At the same time, the non-critical compute node's currently available resources must be greater than the critical compute node's resource gap parameter to meet some or all of the critical compute node's resource needs. By examining all non-critical compute nodes in the supercomputing system one by one, we select those that meet these two conditions and form a set of migration source nodes. The nodes in this migration source node set are potential sources of resource migration for the critical compute nodes.
[0158] Step S1478: Dynamically allocate migration resource quantity parameters according to the difference between the current available resource quantity parameters and the resource gap quantity parameters of each node in the migration source node set, and associate with the corresponding migration target node identifier.
[0159] For each node in the migration source node set, the migration resource quantity parameter is dynamically allocated according to the difference between its current available resource quantity parameter and the resource gap quantity parameter of the key computing node.
[0160] First, for each key compute node, a suitable migration source node is selected from the set of migration source nodes. This selection can be based on factors such as the distance between the migration source node and the key compute node in the supercomputing system topology, and the degree of match between the resource type of the migration source node and the resource requirements of the key compute node. For example, migration source nodes that are close to the key compute node and have a high degree of resource type match are prioritized.
[0161] Then, the difference between the current available resource quantity parameter of the migration source node and the resource gap quantity parameter of the key computing node is calculated. If the current available resource quantity of the migration source node is greater than the resource gap quantity of the key computing node, the resource quantity to be migrated from the migration source node to the key computing node is determined based on the resource allocation strategy of the supercomputing system (such as balancing resource allocation as much as possible, prioritizing emergency needs, etc.). For example, if the resource allocation strategy tends to meet the resource needs of the key computing node as much as possible, and the current available resource quantity of the migration source node is sufficient, then the resource gap quantity of the key computing node can be used as the migration resource quantity parameter; if the current available resource quantity of the migration source node is limited and cannot fully meet the resource gap quantity of the key computing node, then the migration resource quantity parameter is allocated according to a set ratio or rule (such as the ratio between the current available resource quantity of the migration source node and the resource gap quantity).
[0162] After determining the migration resource quantity parameter, it is associated with the identifier of the corresponding key compute node (i.e., the migration target node). This determines how much resources each key compute node should obtain from which migration source node, forming part of the specific resource migration policy parameters. This dynamic allocation of migration resource quantity parameters and association with the migration target node identifier provides detailed operational guidance for resource migration of key compute nodes in supercomputing systems.
[0163] Step S1479: combining the migration target node identifier and the migration resource amount parameter into a resource migration strategy parameter.
[0164] Combine the previously determined migration target node identifier and migration resource quantity parameter to form a complete resource migration strategy parameter. Each key computing node has a corresponding migration target node identifier and migration resource quantity parameter. These two parameters are combined and represented in a predefined data structure. For example, the migration target node identifier can be used as one field, and the migration resource quantity parameter as another field. Together, they form a record of resource migration strategy parameters.
[0165] These resource migration policy parameters define how much resources should be migrated from which source nodes to which key compute nodes (migration target nodes) to meet the resource needs of key compute nodes in the supercomputing system. This combination provides a specific, executable strategy for the dynamic allocation of supercomputing system resources, enabling the system to more efficiently allocate resources, meet the needs of each compute node, and improve overall operational efficiency.
[0166] Step S150: Generate a resource dynamic allocation instruction according to the resource allocation priority parameter and the resource migration policy parameter, and send the resource dynamic allocation instruction to the target computing node to trigger a resource reallocation operation.
[0167] After determining the resource allocation priority parameters and resource migration strategy parameters, the next step is to generate dynamic resource allocation instructions and send them to the target computing node, thereby triggering resource reallocation operations and realizing intelligent allocation of supercomputing resources.
[0168] First, based on the resource allocation priority parameters and resource migration strategy parameters, a resource dynamic allocation instruction is generated according to the set instruction format. The design of the instruction format needs to consider many factors such as the architecture of the supercomputing system, the resource management mechanism, and the communication protocol. For example, the resource dynamic allocation instruction may contain an identifier of the target computing node, which is used to accurately locate the computing node that needs to be reallocated resources. The resource allocation priority parameter is also incorporated into the instruction in a set encoding method, which determines the order of resource allocation and the focus of resource allocation. The resource migration strategy parameter is also encoded according to the set rules, where the migration target node identifier specifies the source node of the resource, and the migration resource quantity parameter accurately specifies the quantity information of each type of resource that needs to be migrated.
[0169] Assuming a supercomputing system employs a hierarchical resource management architecture, with different layers managing and allocating resources differently, the instruction format needs to be adaptable to this hierarchical architecture. For example, instructions may contain identification information for different layers, allowing resource management modules at each layer to accurately identify and process the dynamic resource allocation instructions during transmission. Furthermore, to ensure reliable transmission of dynamic resource allocation instructions within the supercomputing system, some checksum or error correction code information may be added to detect and correct errors that may occur during instruction transmission.
[0170] After generating dynamic resource allocation instructions, they are sent to the target compute nodes via the supercomputing system's internal communication network. This network is responsible for transmitting information between compute nodes. This network may utilize a variety of protocols, such as those based on the Message Passing Interface (MPI) or other high-speed communication protocols designed specifically for supercomputing systems.
[0171] When sending a command, the dynamic resource allocation instruction is first encapsulated into a data packet that conforms to the communication protocol requirements. This data packet may contain the instruction content, source node information, destination node information, and some control information. The data packet's transmission path is then determined based on the communication network topology and routing algorithm. For example, if the supercomputing system's communication network uses a tree topology, the data packet will be transmitted step by step from the source node to the destination node along the branches of the tree structure. During transmission, the data packet may be forwarded through multiple intermediate nodes. Each intermediate node determines the next node to forward the data packet based on information in its routing table.
[0172] When a dynamic resource allocation instruction arrives at the target compute node, the target node's resource management module parses the instruction. Responsible for handling various resource-related matters, the resource management module first verifies the instruction's legitimacy, checking information such as the checksum or error correction code contained in the instruction to ensure that no errors occurred during transmission. If the instruction is valid, the resource management module extracts key information from the instruction, including the resource allocation priority parameter, the migration target node identifier, and the migration resource amount parameter.
[0173] Based on the resource allocation priority parameter, the resource management module determines the order in which resource reallocations are executed. For example, if multiple resource allocation tasks arrive simultaneously, the task with the higher priority parameter will be processed first. Regarding the migration resource quantity parameter, the resource management module checks the current resource status of the target compute node to ensure that sufficient resources are available to allocate or receive the migrated resources. Furthermore, based on the migration target node identifier, the resource management module establishes a communication connection with the migration source node to coordinate the specific aspects of the resource migration.
[0174] After establishing a connection with the migration source node, both parties confirm each other's status and capabilities through a pre-defined handshake protocol to ensure a smooth resource migration. For example, the migration source node sends the target compute node information about its current resource status, including the types and quantities of resources available for migration. The target compute node then provides feedback to the migration source node regarding its resource needs and the types and quantities of resources it can accept. Based on this information, both parties further negotiate specific resource migration details, such as the time and method of migration.
[0175] Once negotiation is complete, the source node begins migrating resources to the target compute node. The resource migration process may involve the transfer of various types of resources, such as moving data from storage devices and reallocating computing resources. High-speed data transfer protocols may be used to ensure data integrity and transfer speed. For computing resources, such as CPU cores, the target compute node's resource management module adjusts its internal task scheduling mechanism to reserve appropriate processing space for the upcoming migration.
[0176] During resource migration, the supercomputing system monitors the migration progress and resource status of the target compute node in real time. For example, a resource monitoring program running on the target compute node collects information such as resource usage and the progress of receiving migrated resources. If an anomaly occurs during migration, such as a network failure that interrupts resource transfer, the supercomputing system activates appropriate error handling mechanisms. It may attempt to reestablish a connection, resume unfinished resource migration tasks, or adjust resource migration strategies based on actual conditions to ensure the resource allocation task is ultimately completed.
[0177] Once resource migration is complete, the target compute node's resource management module updates its own resource status information and reports the results of the resource reallocation to the supercomputing system's central management module. The central management module records this information for subsequent resource management and scheduling decisions. Simultaneously, the target compute node adjusts its internal task execution plan based on the new resource allocation, fully utilizing the newly acquired resources and improving computing efficiency. For example, tasks in the task queue may be rescheduled, allowing some tasks previously waiting due to insufficient resources to be processed using the newly acquired resources.
[0178] Figure 2 A schematic diagram illustrates exemplary hardware and software components of a system 100 for intelligently allocating supercomputing resources based on AI load monitoring, which can implement the concepts of the present application, as provided in some embodiments of the present application. For example, the processor 120 can be used in the system 100 for intelligently allocating supercomputing resources based on AI load monitoring, and can be used to perform the functions described in the present application.
[0179] The AI load monitoring-based supercomputing resource intelligent allocation system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the AI load monitoring-based supercomputing resource intelligent allocation method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0180] For example, the supercomputing resource intelligent allocation system 100 based on AI load monitoring may include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the supercomputing resource intelligent allocation system 100 based on AI load monitoring may also include program instructions stored in ROM, RAM, or other types of non-temporary storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The supercomputing resource intelligent allocation system 100 based on AI load monitoring also includes an I / O interface 150 between the computer and other input and output devices.
[0181] For ease of explanation, only one processor is described in the supercomputing resource intelligent allocation system 100 based on AI load monitoring. However, it should be noted that the supercomputing resource intelligent allocation system 100 based on AI load monitoring in this application may also include multiple processors, so the steps performed by one processor described in this application may also be performed jointly or individually by multiple processors. For example, if the processor of the supercomputing resource intelligent allocation system 100 based on AI load monitoring executes step A and step B, it should be understood that step A and step B may also be executed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.
[0182] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned supercomputing resource intelligent allocation method based on AI load monitoring is implemented.
[0183] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A supercomputing resource intelligent allocation method based on AI load monitoring, characterized in that: The method comprises: Acquire a real-time load data set of multiple computing nodes in a supercomputer system, wherein the real-time load data set includes a resource occupancy index and a task queue status parameter of each computing node in a continuous time window; Performing load feature extraction processing on the real-time load data set to generate a load feature set of each computing node, the load feature set including resource occupancy fluctuation features and task queue evolution features; Calling a pre-built resource demand prediction model to perform resource demand analysis on the load feature set to generate a resource demand prediction result for each computing node in a subsequent time window; Determining resource allocation priority parameters and resource migration strategy parameters for each computing node in the supercomputing system based on the resource demand prediction result; A resource dynamic allocation instruction is generated according to the resource allocation priority parameter and the resource migration policy parameter, and the resource dynamic allocation instruction is sent to the target computing node to trigger a resource reallocation operation.
2. The method for intelligent allocation of supercomputing resources based on AI load monitoring according to claim 1 is characterized in that: The performing load feature extraction processing on the real-time load data set to generate a load feature set of each computing node includes: Performing data cleaning on the real-time load data set, removing load data units corresponding to abnormal time windows, and performing interpolation and filling processing on load data units in missing time windows to generate a standardized load data set; Performing time series division processing on the standardized load data set to generate a plurality of load data subsets of continuous time segments; Perform resource occupancy fluctuation analysis on each load data subset, extracting the resource occupancy change gradient parameters and resource occupancy peak duration parameters of each computing node in the corresponding time segment; Perform task queue evolution analysis on each load data subset to extract the trend parameters of the number of pending tasks, the cumulative task execution delay parameters, and the task queue mutation amplitude parameters of each computing node in the corresponding time segment; The resource occupancy rate change gradient parameter and the resource occupancy rate peak duration parameter are input into the feature encoder for normalized feature splicing processing to generate the resource occupancy fluctuation feature vector, and the pending task quantity change trend parameter, the task execution delay accumulation parameter and the task queue mutation amplitude parameter are input into the feature encoder for normalized feature splicing processing to generate the task queue evolution feature vector, and the resource occupancy fluctuation feature and the task queue evolution feature are output as the load feature set.
3. The method for intelligent allocation of supercomputing resources based on AI load monitoring according to claim 2 is characterized in that: The step of performing time series division processing on the standardized load data set to generate a plurality of load data subsets of consecutive time segments includes: Detecting the timestamp continuity parameter of the load data unit in each time window in the standardized load data set, performing segment marking processing on the load data unit with time stamp jump, and removing the segment if the number of jump time windows exceeds a preset threshold; Performing a sliding window segmentation process on the standardized load data set after the segmentation marking process based on a preset time segment length threshold to generate multiple load data subsets with equal time lengths; Performing time dimension alignment processing on the resource utilization rate indicators in each load data subset so that the time intervals of the load data units in each time window are evenly distributed; A queue identifier matching process is performed on the task queue state parameters in each load data subset to generate a plurality of load data subsets of continuous time segments.
4. The method for intelligent allocation of supercomputing resources based on AI load monitoring according to claim 2 is characterized in that: The resource occupancy rate fluctuation analysis processing is performed on each load data subset to extract the resource occupancy rate change gradient parameter and resource occupancy rate peak duration parameter of each computing node in the corresponding time segment, including: Performing smoothing filtering on the resource occupancy rate indicators in the load data subset to generate a filtered resource occupancy rate time series; Calculating the absolute value of the difference of the filtered resource occupancy time series in adjacent time windows to generate a resource occupancy change gradient parameter; Detecting peak points exceeding a preset threshold in the filtered resource occupancy time series, and recording a start timestamp and an end timestamp corresponding to each peak point; The resource occupancy peak duration parameter is calculated according to the difference between the start timestamp and the end timestamp, and the duration parameters of consecutive peak points are accumulated and summed.
5. The method for intelligent allocation of supercomputing resources based on AI load monitoring according to claim 2 is characterized in that: The task queue evolution analysis and processing is performed on each load data subset to extract the trend parameters of the number of tasks to be processed, the cumulative amount parameters of task execution delays, and the task queue mutation amplitude parameters of each computing node in the corresponding time segment, including: Analyzing the task queue status parameters in the load data subset to extract the time series of the number of tasks to be processed and the time series of the average execution delay of tasks of each computing node; Performing linear fitting processing on the time series of the number of tasks to be processed, and calculating the slope parameter of the fitting curve as a parameter of the change trend of the number of tasks to be processed; Performing an integration operation on the task average execution delay timing sequence to calculate a task execution delay accumulation parameter within a preset time window; A mutation point in the time series of the number of tasks to be processed is detected, and the absolute value of the difference between the number of tasks before and after the mutation point is calculated as a task queue mutation amplitude parameter.
6. The method for intelligent allocation of supercomputing resources based on AI load monitoring according to claim 1 is characterized in that: The calling of the pre-built resource demand prediction model to perform resource demand analysis on the load feature set to generate a resource demand prediction result for each computing node in a subsequent time window includes: Inputting the load feature set into the feature coding layer of the resource demand prediction model, performing correlation modeling on the resource occupancy fluctuation feature and the task queue evolution feature, and generating a fusion feature vector; Performing temporal dependency extraction processing on the fused feature vector through the temporal convolution layer of the resource demand prediction model to generate a temporal feature representation with long-term and short-term dependencies; Calling the multi-head attention layer of the resource demand prediction model to perform key time node identification processing on the time series feature representation, and generating an attention weight distribution map for different time nodes; Performing dynamic weighted aggregation processing on the temporal feature representation based on the attention weight distribution map to generate an aggregated feature vector that enhances key time node information; The aggregated feature vector is input into the regression output layer of the resource demand prediction model to perform resource demand prediction processing, and a resource demand prediction result of each computing node in a subsequent time window is generated.
7. The method for intelligent allocation of supercomputing resources based on AI load monitoring according to claim 6 is characterized in that: The step of performing temporal dependency extraction processing on the fused feature vector through the temporal convolution layer of the resource demand prediction model to generate a temporal feature representation with long-term and short-term dependencies includes: Inputting the fused feature vector into the dilated causal convolution module of the temporal convolution layer, performing multi-scale temporal feature extraction processing on the fused feature vector using convolution kernels with different dilation rates, and obtaining a convolution output result corresponding to each dilation rate; The convolution output results corresponding to each expansion rate are channel-normalized, and the normalized multi-scale time series features are input into the gated activation function module to generate a time series feature map with a nonlinear relationship; Performing residual connection processing on the temporal feature map, and performing element-by-element addition processing on the original fused feature vector and the temporal feature map after gate activation; The residual processed temporal feature map is input into the hierarchical normalization module to adjust the feature distribution and generate a temporal feature representation with long-term and short-term dependencies.
8. The method for intelligent allocation of supercomputing resources based on AI load monitoring according to claim 6 is characterized in that: The calling of the multi-head attention layer of the resource demand prediction model performs key time node identification processing on the time series feature representation to generate an attention weight distribution diagram for different time nodes, including: Splitting the time series feature representation into multiple sub-feature vectors, each sub-feature vector corresponding to a feature representation of a different time node; Perform linear projection on each sub-feature vector to generate query vector, key vector and value vector; Calculate the similarity scores between the query vector and the key vector at different time nodes to generate the initial attention score matrix; Performing masking on the initial attention score matrix to shield the attention influence of future time nodes on historical time nodes; The masked attention score matrix is input into the Softmax function for normalization to generate the attention weight distribution graph at different time nodes; A weighted summation process is performed on the value vector based on the attention weight distribution map to generate a context-aware temporal feature representation.
9. The method for intelligent allocation of supercomputing resources based on AI load monitoring according to claim 1, characterized in that: The determining of resource allocation priority parameters and resource migration strategy parameters of each computing node in the supercomputing system based on the resource demand prediction result includes: Analyze the predicted resource occupancy rate index and the predicted task queue length index of each computing node in the resource demand prediction result; Calculating a resource scarcity parameter of each computing node according to the predicted resource occupancy rate indicator, wherein the resource scarcity parameter is positively correlated with the predicted resource occupancy rate indicator; Calculating a task backlog risk parameter for each computing node based on the predicted task queue length indicator, wherein the task backlog risk parameter is positively correlated with a growth rate of the predicted task queue length indicator; Normalizing the resource scarcity parameter and the task backlog risk parameter, and performing a linear weighted summation process on the normalized resource scarcity parameter and the task backlog risk parameter according to a preset resource allocation strategy weight coefficient to generate a comprehensive priority score; Detecting computing nodes whose comprehensive priority scores exceed a preset threshold and marking them as key computing nodes; Adjusting a resource allocation priority parameter according to a distribution density parameter of key computing nodes, wherein the resource allocation priority parameter is negatively correlated with the distribution density of key computing nodes; A resource allocation order list of each key computing node is generated based on the adjusted resource allocation priority parameter, and resource migration strategy parameters are determined according to the resource allocation order list of each computing node. The resource migration strategy parameters include a migration target node identifier and a migration resource amount parameter.
10. A supercomputing resource intelligent allocation system based on AI load monitoring, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the supercomputing resource intelligent allocation method based on AI load monitoring as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method and device for computing power resource modeling and scoring in supercomputing center
CN117806931A
Node computing resource allocation method, network switching subsystem and intelligent computing platform
CN117997906A
SLA demand-oriented computing network multi-element fusion arrangement method
CN119536932A
AI-based big data distributed computing task automatic optimization method and system
CN119576507A
PCFarm resource scheduling method and system based on dynamic load prediction
CN120104355A
Cited By
Server resource scheduling system for high-density computing environment
CN120872610A
Load resource allocation method and system applied to intelligent supervision data analysis
CN121070631A
Power consumption and performance balanced scheduling method and system for heterogeneous multi-core processor
CN121455661A
Multi-level data intelligent statistical analysis system based on machine learning
CN121502159A
Data center resource dynamic scheduling method based on artificial intelligence
CN121658186A