A business demand-oriented dynamic computing power routing calculation method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本发明针对现有技术中算力调度缺乏负载趋势感知、两级调度协同性差以及业务需求评估粗放的技术问题,提供一种面向业务需求的动态算力路由计算方法及系统
[0017]相较于现有技术,本发明首先通过提取业务请求中的算力类型、算力大小需求、内存需求、最大容忍时延及最小带宽需求,实现了对用户业务需求的精细化感知,为调度决策提供了精准约束。其次,本发明在中心控制器融合区域算力状态信息与预测信息,以最大化全局资源利用率和区域间负载均衡为目标进行区域适配评估,引入负载趋势预测,避免了因资源波动导致的调度滞后与区域失衡。再次,本发明在适配边缘处理器基于区域内节点状态信息,以最大化区域内负载均衡为目标进行节点匹配分析,实现了区域选择与节点分配的协同优化。最后,本发明通过两级联动的动态算力路由机制,将业务请求精准调度至最优算力节点,在满足业务多重约束的前提下提升了全网算力资源利用率,保障了系统在不同负载条件下的均衡性与稳定性。
Smart Images

Figure CN122093468B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computing power network technology, specifically to a dynamic computing power routing calculation method and system oriented towards business needs. Background Technology
[0002] In a computing network architecture, computing tasks are widely distributed across edge nodes and various heterogeneous computing resources. With the explosive growth of low-latency services such as artificial intelligence and the Internet of Things, user service requests have placed differentiated demands on computing power type, computing power size, memory capacity, transmission latency, and network bandwidth. How to accurately schedule computing tasks to the most suitable computing nodes has become crucial for improving resource utilization and ensuring service quality.
[0003] Existing computing power scheduling methods typically employ centralized resource views for node selection or rely on simple threshold rules for region allocation. On one hand, these methods often depend solely on the current resource idle rate, failing to adequately consider the temporal variations in regional load. This leads to scheduling decisions lagging behind resource fluctuations, easily resulting in imbalances where some regions are overloaded due to sudden traffic surges while others remain idle for extended periods. On the other hand, existing methods lack collaborative optimization mechanisms at both the region selection and node matching levels. The load balancing objectives between regions and the node balancing objectives within regions are disconnected, making it difficult to maximize global resource utilization. Furthermore, for the multi-dimensional demands of user business requests, existing technologies generally employ rigid filtering methods, lacking refined evaluation of multiple constraints such as latency, bandwidth, and computing power type. This prevents dynamic adaptation of optimal resources while meeting basic business requirements. Summary of the Invention
[0004] This invention addresses the technical problems in existing computing power scheduling, such as lack of load trend perception, poor coordination between two-level scheduling, and crude assessment of business needs. It provides a dynamic computing power routing calculation method and system oriented towards business needs.
[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0006] In a first aspect, the present invention provides a dynamic computing power routing calculation method oriented towards business needs, comprising:
[0007] Receive user service requests, extract computing power type requirements, computing power size requirements, memory requirements, maximum tolerable latency, and minimum bandwidth requirements from the service requests and transmit them to the central controller as service requirement information;
[0008] In the central controller, based on several computing power status information and several computing power prediction information uploaded by several edge processors to several computing power scheduling areas, and combined with the business demand information, a regional adaptation assessment is performed with the goal of maximizing global resource utilization and inter-regional load balancing, to determine the suitable computing power scheduling areas and suitable edge processors.
[0009] In the adapted edge processor, based on the node status information of multiple computing nodes in the adapted computing power scheduling area at the current time, a computing power node matching analysis is performed with the goal of maximizing load balancing within the area to determine the optimal computing power node.
[0010] The service request is scheduled to the optimal computing power node for processing.
[0011] Secondly, the present invention provides a dynamic computing power routing calculation system oriented towards business needs, comprising:
[0012] The business requirement extraction module is used to receive user business requests and extract computing power type requirements, computing power size requirements, memory requirements, maximum tolerable latency and minimum bandwidth requirements from the business requests as business requirement information and transmit them to the central controller.
[0013] The central controller module is used to perform a regional adaptation assessment on the central controller based on several computing power status information and several computing power prediction information uploaded by several edge processors to several computing power scheduling areas, combined with the business demand information, with the goal of maximizing global resource utilization and inter-regional load balancing, and to determine the suitable computing power scheduling areas and suitable edge processors.
[0014] An adaptive edge processor module is used to perform a computing node matching analysis based on the node status information of multiple computing nodes in the adaptive computing power scheduling area at the current time, with the goal of maximizing load balancing within the area, to determine the optimal computing node.
[0015] The scheduling and execution module is used to schedule the service request to the optimal computing power node for processing.
[0016] The beneficial effects of this invention are:
[0017] Compared to existing technologies, this invention first achieves refined perception of user business needs by extracting computing power type, computing power size requirements, memory requirements, maximum tolerable latency, and minimum bandwidth requirements from business requests, providing precise constraints for scheduling decisions. Secondly, this invention integrates regional computing power status information and predictive information in the central controller to perform regional adaptation evaluation with the goal of maximizing global resource utilization and inter-regional load balancing, introducing load trend prediction to avoid scheduling lag and regional imbalance caused by resource fluctuations. Thirdly, this invention adapts edge processors based on node status information within the region to perform node matching analysis with the goal of maximizing load balancing within the region, achieving coordinated optimization of region selection and node allocation. Finally, this invention uses a two-level linked dynamic computing power routing mechanism to accurately schedule business requests to the optimal computing power node, improving the overall network computing power resource utilization while meeting multiple business constraints, and ensuring the system's balance and stability under different load conditions. Attached Figure Description
[0018] Figure 1 A flowchart illustrating a dynamic computing power routing calculation method oriented towards business needs provided by the present invention;
[0019] Figure 2 This is a schematic diagram of the structure of a dynamic computing power routing calculation system oriented towards business needs, provided by the present invention.
[0020] In the attached diagram, the components represented by each number are as follows:
[0021] The module includes a business requirement extraction module 11, a central controller module 12, an edge processor adaptation module 13, and a scheduling and execution module 14. Detailed Implementation
[0022] Example 1, as Figure 1 As shown, this embodiment of the invention provides a dynamic computing power routing calculation method oriented towards business needs, including:
[0023] S10: Receive the user's service request, extract the computing power type requirement, computing power size requirement, memory requirement, maximum tolerable latency and minimum bandwidth requirement from the service request and transmit them to the central controller as service requirement information;
[0024] First, the system receives user service requests. A user's service request is a computational task execution instruction initiated by the user, carrying resource specification constraints and service quality requirements for task execution. From this service request, the system extracts the following requirements: computing power type, computing power size, memory requirements, maximum tolerable latency, and minimum bandwidth requirements. Specifically, computing power type refers to the service request's compatibility with processor architecture or acceleration hardware type, such as a central processing unit, graphics processing unit, or neural network processor; computing power size refers to the total amount of computing resources required by the service request, typically measured in the number of processor cores or floating-point operations per second; memory requirements refer to the memory capacity required during the service request's execution; maximum tolerable latency refers to the maximum allowed time delay from submission to completion of the service request, ensuring the real-time requirements of the service; and minimum bandwidth requirements refer to the minimum network bandwidth required for data transmission, preventing transmission congestion due to insufficient bandwidth.
[0025] The extracted multi-dimensional resource constraint indicators are used as business requirement information. This business requirement information is a structured expression of quantifiable resource requirements in user business requests. It represents the user's specific expectations for computing resources in terms of type, capacity, storage, latency, and bandwidth, and is used to provide a constraint basis for subsequent region selection and node matching.
[0026] Then, the business requirement information is transmitted to the central controller, so that the central controller can accurately match each computing power scheduling area in the global view according to the requirement, and ensure that subsequent scheduling decisions achieve efficient allocation of resources while meeting the user's service quality.
[0027] Specifically, because user service requests have highly diverse resource requirements, using only resource idle rate as the scheduling basis will not be able to adapt to the differentiated requirements of different types of services for computing power architecture and network performance. Therefore, by extracting and structuring the requirements of service requests from multiple dimensions, user intent can be transformed into computable constraint parameters, providing a data foundation for two-level collaborative scheduling.
[0028] S20: In the central controller, based on several computing power status information and several computing power prediction information uploaded by several edge processors for several computing power scheduling areas, and combined with the business demand information, a regional adaptation assessment is performed with the goal of maximizing global resource utilization and inter-regional load balancing, and the suitable computing power scheduling areas and suitable edge processors are determined.
[0029] Specifically, the central controller receives and processes the aforementioned transmitted service request information. The central controller is a centralized decision-making unit deployed at the core layer of the computing power network. It aggregates the resource status and load trends of all computing power scheduling regions across the network and executes globally optimal region selection decisions based on service requirements. As the top-level decision-making node in the two-tier scheduling architecture, the central controller undertakes the core functions of cross-regional resource coordination and task distribution.
[0030] Specifically, the central controller contains computing power status information and computing power prediction information uploaded by several edge processors to several computing power scheduling regions. A computing power scheduling region refers to a geographical or responsibility area managed by one edge processor, containing multiple computing power nodes. The edge processor is responsible for collecting the aggregated status of computing power resources within the region and reporting it to the central controller. The computing power status information reflects the real-time resource status of each computing power scheduling region at the current moment, including indicators such as region average processor utilization, region average memory utilization, region total task queue length, region entry latency, region available bandwidth, and region supported computing power types. This information characterizes the region's current available capacity and quality of service assurance capabilities. The computing power prediction information reflects the load change trend of each computing power scheduling region within a preset future time window. It is obtained by the edge processor performing computing power load prediction based on historical load sequences, including a region predicted load sequence and a prediction confidence sequence. This information compensates for the limitation of real-time status information in predicting resource fluctuations.
[0031] By integrating the two types of information mentioned above, the central controller can not only grasp the current resource availability of the region, but also predict the future load trend of the region. Thus, in the process of regional adaptation assessment, it can take into account both short-term resource availability and medium- and long-term load balance, and provide data support for subsequent decisions aimed at maximizing global resource utilization and inter-regional load balance.
[0032] First, the steps for obtaining several computing power status information and several computing power prediction information from several computing power scheduling areas uploaded by several edge processors include:
[0033] Collect the regional average processor utilization, regional average memory utilization, regional total task queue length, regional entry latency, regional available bandwidth, and regional supported computing power types of each edge processor in its respective computing power scheduling region to obtain several computing power status information.
[0034] Based on the historical load sequences of the computing power scheduling regions where several edge processors are located, computing power load prediction is performed within a preset time window to obtain several computing power prediction information. Each computing power prediction information includes a region prediction load sequence and a prediction confidence sequence.
[0035] First, the average processor utilization, average memory utilization, total task queue length, entry latency, available bandwidth, and supported computing power types of each edge processor in its respective computing power scheduling region are collected to obtain several computing power status information.
[0036] As resource aggregation nodes within their respective computing power scheduling regions, edge processors are responsible for periodically collecting various resource metrics within those regions. Specifically, the region's average processor utilization reflects the average utilization of central processing unit resources across all computing power nodes within the scheduling region during a statistical time window; the region's average memory utilization reflects the average utilization of memory resources across all computing power nodes within the region; the region's total task queue length reflects the total number of tasks currently awaiting processing in the region; the region's entry latency characterizes the transmission and queuing delay experienced by a service request from accessing the network to reaching the edge processor in that region; the region's available bandwidth indicates the remaining bandwidth resources between the region and the external network; and the region's supported computing power type indicates the types of computing power services that the region can provide, such as general-purpose computing, graphics processor-accelerated computing, or neural network processor-accelerated computing.
[0037] By collecting the above six types of indicators, a structured computing power status information is formed. This computing power status information describes the real-time resource status of each computing power scheduling area at the current moment, providing a quantitative basis for the current available capacity for subsequent area adaptation assessment.
[0038] Secondly, based on the historical load sequences of the computing power scheduling regions of several edge processors, computing power load prediction is performed within a preset time window to obtain several computing power prediction information. The historical load sequence refers to the time-series data composed of the comprehensive regional load values at various sampling moments within a historical period, recorded and stored by the edge processors. This historical load sequence reflects the pattern and trend of regional load changes over time. The historical time refers to the time range covered by the historical data selected for training the computing power load prediction model. This historical time is comprehensively set according to the periodic change pattern of computing power load and the input length requirements of the prediction model; for example, it can be set to the past 24 hours or the past 7 days to cover the complete business load cycle. The sampling moment refers to the time point at which the edge processor collects the comprehensive regional load value. The time interval between two adjacent sampling moments can be set according to the system monitoring granularity and computational overhead; for example, it can be set to 30 seconds or 1 minute. The sampling moments are arranged sequentially according to this fixed time interval to form a continuous historical load sequence.
[0039] Furthermore, based on historical load sequences, computing load prediction is performed within a preset time window. The preset time window refers to the future period to be predicted, the length of which can be set according to the business scheduling cycle and system response requirements, for example, 5 minutes or 10 minutes. Specifically, computing load prediction is the process of inferring the comprehensive regional load value at each prediction moment within the preset time window by combining the regional load change patterns, periodic characteristics, and trend information contained in the historical load sequence with a time series prediction model or machine learning algorithm, thus forming a regional predicted load sequence.
[0040] Specifically, based on the historical load sequences of the computing power scheduling regions where several edge processors are located, computing power load prediction is performed within a preset time window to obtain several computing power prediction information, including:
[0041] Randomly select any edge processor from the plurality of edge processors as the first edge processor, and obtain the first historical load sequence of the first edge processor in the historical time zone. The historical load sequence includes the regional comprehensive load value at multiple historical moments. The regional comprehensive load value is obtained by weighted evaluation of the regional average CPU utilization, the regional average memory utilization, and the regional total task queue length.
[0042] The first load variation coefficient is obtained by calculating the load volatility of the first historical load sequence.
[0043] The first computing power load predictor is activated according to the first load variation coefficient, and the first computing power prediction information within the preset time window is predicted based on the first historical load sequence and added to the plurality of computing power prediction information.
[0044] First, any edge processor is randomly selected from several edge processors as the first edge processor, and the first historical load sequence of the first edge processor within the historical time zone is obtained. Each edge processor is responsible for maintaining the historical load data of its computing power scheduling area. The historical time zone refers to a pre-defined time range for storing historical data, such as the past 24 hours or the past 7 days.
[0045] The first historical load sequence contains regional comprehensive load values at multiple historical moments. Each historical moment corresponds to a sampling time point, and the time interval between adjacent sampling moments can be set according to the monitoring granularity, such as 30 seconds or 1 minute. The regional comprehensive load value is a comprehensive quantitative indicator of the overall resource utilization of the computing power scheduling region. Specifically, it can be obtained by weighting and evaluating the regional average CPU utilization, regional average memory utilization, and regional total task queue length. Among them, the regional average CPU utilization reflects the average CPU resource utilization of all computing power nodes in the region, the regional average memory utilization reflects the average memory resource utilization, and the regional total task queue length reflects the backlog of tasks to be processed in the region. The three indicators are weighted and summed using preset weight coefficients to obtain the regional comprehensive load value that can comprehensively characterize the regional load status.
[0046] The weighting coefficients are set comprehensively based on the core business type of the computing power scheduling region and the contribution of various resources to business processing capacity. For example, for compute-intensive business regions, the weighting coefficient for the region's average CPU utilization can be set to 0.6, the weighting coefficient for the region's average memory utilization can be set to 0.2, and the weighting coefficient for the region's total task queue length can be set to 0.2. For memory-intensive business regions, the weighting coefficient for the region's average CPU utilization can be set to 0.2, the weighting coefficient for the region's average memory utilization can be set to 0.6, and the weighting coefficient for the region's total task queue length can be set to 0.2. For regions with high task concurrency, the weighting coefficient for the region's total task queue length can be appropriately increased, for example, by setting the three weighting coefficients to 0.3, 0.3, and 0.4 respectively. Through the differentiated settings of the above weighting coefficients, the regional comprehensive load value can accurately reflect the actual resource occupancy under different business types.
[0047] Secondly, load volatility is calculated on the acquired first historical load sequence to obtain the first load variation coefficient. The load variation coefficient is a statistic that measures the dispersion of the historical load sequence and is used to characterize the severity of load fluctuations in the computing power scheduling area. Specifically, the average and standard deviation of the comprehensive load values for all areas in the first historical load sequence are first calculated. Then, the standard deviation is divided by the average to obtain the first load variation coefficient. The larger the first load variation coefficient, the more severe the historical load fluctuations in the area and the more complex the load change pattern; the smaller the first load variation coefficient, the more stable the historical load in the area and the simpler the load change pattern.
[0048] Then, the first computing power load predictor is activated according to the first load variation coefficient, and the first computing power prediction information within a preset time window is obtained based on the first historical load sequence, and added to several computing power prediction information sets. The first computing power load predictor is an integrated prediction model built on a Long Short-Term Memory (LSTM) network, used to predict the load change trend within a preset time window based on historical load sequences. Since the load fluctuation characteristics differ in different regions, regions with drastic load fluctuations require more complex prediction models to capture their changing patterns, while regions with stable loads can use relatively simple prediction models. Therefore, the number of model branches participating in the prediction in the first computing power load predictor needs to be dynamically determined according to the magnitude of the first load variation coefficient: the larger the load variation coefficient, the more model branches are activated, improving prediction accuracy through the integration of multi-branch prediction results; the smaller the load variation coefficient, the fewer model branches are activated to reduce computational overhead.
[0049] The first computing power prediction information includes a first region predicted load sequence and a first prediction confidence sequence. The first region predicted load sequence is the comprehensive load prediction value of the region at each prediction time within a preset time window. The first prediction confidence sequence is the confidence index corresponding to each prediction time, used to reflect the reliability of each prediction value. Through this method, each edge processor can adaptively adjust the complexity of the prediction model according to the load fluctuation characteristics of its own region, achieving efficient utilization of computing resources while ensuring prediction accuracy.
[0050] Specifically, the first computing power load predictor is activated according to the first load variation coefficient, and the first computing power prediction information is obtained based on the first historical load sequence, including:
[0051] Based on the historical operation records of the first edge processor, and constrained by the time span of the historical time zone and the preset time window, a sample historical load sequence set and a sample region predicted load sequence set are collected and used as sample training data.
[0052] The training data is subjected to K-fold cross-multiplication with replacement to obtain a training set of K samples;
[0053] Using the historical load sequence of the sample as input data and the predicted load sequence of the sample region as supervision label, the long short-term memory network is trained to converge using the K sample training sets to obtain K first computing load prediction branches, which are then combined to obtain the first computing load predictor.
[0054] The ratio of the first load variation coefficient to the preset baseline load variation coefficient of the first edge processor is used as the first prediction difficulty coefficient. The ratio of the first prediction difficulty coefficient to the preset baseline branch number is rounded down to obtain the number of adaptation branches P, where the preset baseline branch number is half K, and P is greater than or equal to 3 and less than or equal to K.
[0055] P branches are randomly selected from the K first computing load prediction branches of the first computing load predictor, and predictions are made according to the first historical load sequence to obtain P initial regional predicted load sequences. The first regional predicted load sequence is obtained by calculating the average value of the P initial regional predicted load sequences at the same time.
[0056] Calculate the reciprocal of the variance of the P initial region predicted load sequences at the same moment, and perform normalization processing to obtain the first prediction confidence sequence. Use the first region predicted load sequence and the first prediction confidence sequence as the first computing power prediction information.
[0057] First, based on the historical operation records of the first edge processor, and constrained by the time span of the historical time zone and a preset time window, a set of historical load sequences and a set of predicted load sequences for sample regions are collected and used as training data. The historical operation records are load data accumulated by the first edge processor during long-term operation, containing the comprehensive load values of regions within different time periods. Constrained by the time span of the historical time zone, the time length covered by each historical load sequence is determined, which is equal to the length of the historical time zone. Constrained by the time span of the preset time window, the future time length covered by each predicted load sequence for sample regions is determined, which is equal to the length of the preset time window.
[0058] Multiple sets of samples are extracted from historical operation records using a sliding window approach. Each set of samples contains a historical load sequence of samples and a corresponding sample region prediction load sequence. All samples constitute the sample training data for subsequent training of the prediction model.
[0059] Secondly, the training data is partitioned with replacement using a K-fold cross-splitting method to obtain K training sets. K-fold cross-splitting with replacement refers to the process of randomly drawing K samples with replacement from the original training data, each time drawing the same number of samples as the original samples, forming K different training sets. Here, K is a hyperparameter used to control the number of branches in the ensemble model. Its value is set comprehensively based on the total amount of training data and the system's requirements for balancing prediction accuracy and computational cost; for example, K can be set to 10 or 20. This partitioning method ensures that each training set has an independent sample distribution, with overlap and differences between the training sets, thereby enhancing the diversity between branches in subsequent model training and improving the generalization ability of the ensemble model.
[0060] Furthermore, using the historical load sequence of the samples as input data and the predicted load sequence of the sample region as supervision labels, the Long Short-Term Memory (LSTM) network is trained to convergence using K sample training sets, resulting in K first computing load prediction branches. These branches are then combined to obtain the first computing load predictor. The LSM network is a variant of a recurrent neural network suitable for time series prediction, capable of effectively capturing long-term dependencies in the load sequence.
[0061] For example, due to the complex characteristics of regional load sequences, such as time dependence, periodic fluctuations, and sudden changes, and the significant advantages of Long Short-Term Memory (LSTM) networks in capturing temporal dependencies and long-term memory capabilities, LTM networks are chosen to construct the computational load prediction model. Specifically, optionally, each first computational load prediction branch adopts a stacked LTM network structure, mainly including an input layer, an LTM network layer, an attention mechanism layer, and an output layer. The input layer receives the time-step input vector formed by the standardization of the first historical load sequence, with each time step corresponding to the regional comprehensive load value at a historical moment. The LTM network layer adopts a two-layer stacked structure, with the first layer containing 128 hidden units and the second layer containing 64 hidden units. Each LTM network layer uses the ReLU activation function to introduce nonlinear transformation capabilities, and a Dropout layer is embedded between network layers, with the dropout rate set between 0.2 and 0.3 to effectively suppress model overfitting and improve its generalization performance. The attention mechanism layer is used to weight and aggregate the hidden states of each time step output by the Long Short-Term Memory network layer, enabling the model to adaptively focus on historical moment information that contributes more to the prediction. The output layer adopts a fully connected structure, mapping the feature vector output by the attention mechanism layer to the regional comprehensive load value of each prediction moment within a preset time window.
[0062] During training, key hyperparameters included a learning rate of 0.001, a training epoch count of 200, a batch size of 32, and a time step length set to the number of sampling points within the historical time zone. The learning rate was set to balance training stability and convergence speed; the number of training epochs ensured the model fully learned the temporal dependency patterns in the load sequence; and the batch size balanced training efficiency with memory resource consumption.
[0063] Specifically, a supervised learning training method is adopted. Historical load sequences of samples from the pre-defined K sample training sets are read as input, and the predicted load sequences of the corresponding sample regions are read as supervision labels. The network weight parameters are iteratively optimized using the time backpropagation algorithm in conjunction with the Adam optimizer. The mean squared error loss function is used to measure the deviation between the predicted load sequence and the true load sequence. The training process is monitored using a validation set. Training is terminated when the validation set loss function value no longer decreases for several consecutive rounds and the model prediction accuracy reaches a predetermined threshold, resulting in a converged Long Short-Term Memory (LSTM) network model.
[0064] The convergence condition is set based on the trend of the loss function on the validation set and the required prediction accuracy. For example, training can be terminated when the loss function value on the validation set no longer decreases after 10 consecutive training iterations and the average absolute percentage error of the model on the validation set is less than 5%, thus obtaining the converged Long Short-Term Memory (LSTM) network model. By setting the above convergence condition, it is ensured that the model fully learns the temporal dependency patterns in the load sequence while avoiding overfitting caused by overtraining.
[0065] A Long Short-Term Memory (LSTM) network model is trained independently for each training set until the loss function of the model converges on the validation set, resulting in K trained first computing load prediction branches. These K branches are then combined to form the first computing load predictor, which is essentially an ensemble prediction model containing K heterogeneous model branches. Each branch has different prediction preferences and feature extraction capabilities due to differences in the training data.
[0066] Furthermore, the ratio of the first load variation coefficient to the preset baseline load variation coefficient of the first edge processor is used as the first prediction difficulty coefficient. The ratio of the first prediction difficulty coefficient to the preset baseline number of branches is rounded down to obtain the number of adapted branches P. The preset baseline load variation coefficient is a reference value obtained by statistically averaging the variation coefficients of the load sequences during typical periods of relatively stable load in the computing power scheduling region's historical operation. This preset baseline load variation coefficient is set comprehensively based on the business type and historical operating patterns of the region where the edge processor is located. For example, for business regions with relatively stable loads, the preset baseline load variation coefficient can be set to 0.2; for business regions with large load fluctuations, the preset baseline load variation coefficient can be set to 0.4.
[0067] The ratio of the first load variation coefficient to the preset baseline load variation coefficient is used as the first prediction difficulty coefficient. This first prediction difficulty coefficient is a dimensionless index that quantifies the difficulty of load prediction in the region, representing the degree of deviation of the current regional load fluctuation from the stable baseline state. Specifically, when the first prediction difficulty coefficient is greater than 1, it indicates that the regional load fluctuates drastically and the prediction difficulty is high; when the first prediction difficulty coefficient is less than 1, it indicates that the regional load is stable and the prediction difficulty is low.
[0068] Furthermore, the ratio of the first prediction difficulty coefficient to the preset baseline branch number is rounded down to obtain the number of adaptive branches, P. The preset baseline branch number is the number of branches that should be activated under the baseline state, and can be set to half of K. For example, when K equals 10, the preset baseline branch number is 5. This calculation method ensures that in regions with more drastic load fluctuations, the number of adaptive branches P is larger, requiring more model branch integration to improve prediction accuracy; conversely, in regions with more stable load fluctuations, the number of adaptive branches P is smaller, requiring reduced computational overhead. Simultaneously, a lower limit of 3 and an upper limit of K are set for P to ensure that the predictor activates at least 3 branches to guarantee prediction stability, and the total number of branches is not exceeded.
[0069] Furthermore, P branches are randomly selected from the K first computing load prediction branches of the first computing load predictor, and predictions are made according to the first historical load sequence to obtain P initial regional predicted load sequences. The first regional predicted load sequence is obtained by calculating the average of the P initial regional predicted load sequences at the same time.
[0070] Specifically, randomly selecting the first computing load prediction branch further enhances the diversity of prediction results and avoids systematic bias caused by branch selection preferences. Each selected first computing load prediction branch uses the first historical load sequence as input to independently predict the regional comprehensive load value at each moment within a preset time window, forming an initial regional predicted load sequence. Since the prediction results of different branches may differ, the predicted values of P initial regional predicted load sequences at the same prediction moment need to be arithmetically averaged to obtain the comprehensive predicted value at that moment. The comprehensive predicted values at all moments constitute the first regional predicted load sequence. By using the mean fusion method, the impact of single branch prediction errors on the final result can be effectively reduced. The final obtained first regional predicted load sequence is a sequence composed of the regional comprehensive load prediction values at each prediction moment within the preset time window. It represents the comprehensive prediction result of the resource occupancy of the computing power scheduling region in the future, and is used to provide a forward-looking load trend basis for subsequent regional adaptation assessment.
[0071] Then, the reciprocal of the variance of the P initial regional predicted load sequences at any given moment is calculated and normalized to obtain the first prediction confidence sequence. Specifically, for each prediction moment within a preset time window, the variance of the predicted values of the P initial regional predicted load sequences at that moment is calculated. The smaller the variance, the more consistent the prediction results of the P branches at that moment, and the higher the confidence of the prediction results; the larger the variance, the greater the prediction divergence between branches, and the lower the confidence of the prediction results. The reciprocal of the variance is used as the base confidence value for that moment; the larger the variance, the smaller the reciprocal, and the lower the confidence. Subsequently, the base confidence values for all prediction moments are normalized so that the sum of the confidence values at each moment is 1 or the value ranges between 0 and 1, thus obtaining the first prediction confidence sequence. This first prediction confidence sequence reflects the reliability of each predicted value in the first regional predicted load sequence at different moments, and is used to provide a quantitative basis for the reasonable use of prediction information in subsequent regional adaptation assessments.
[0072] Finally, the first region predicted load sequence and the first predicted confidence sequence are used as the first computing power prediction information. Through the above steps, the first edge processor adaptively activates the number of prediction branches according to the historical load fluctuation characteristics of its computing power scheduling region, generates the region predicted load sequence using the integrated prediction method, and synchronously outputs the confidence information at each prediction time. This provides complete prediction information for region adaptation evaluation, including both load change trends and prediction reliability.
[0073] Furthermore, based on several computing power status information and several computing power prediction information from several computing power scheduling regions, combined with the aforementioned business demand information, a region adaptation assessment is conducted with the goal of maximizing global resource utilization and inter-regional load balancing to determine suitable computing power scheduling regions and suitable edge processors, including:
[0074] In the central controller, the computing power type is matched for each computing power scheduling region according to the computing power type requirements in the business demand information, and a set of candidate regions that support the computing power type requirements is selected.
[0075] Based on the maximum tolerable latency and minimum bandwidth requirement in the business requirement information, the candidate region set is filtered by network constraints, and regions whose ingress latency exceeds the maximum tolerable latency or whose available bandwidth is less than the minimum bandwidth requirement are eliminated.
[0076] Obtain the average processor utilization, average memory utilization, and total task queue length of each region in the candidate region set, and calculate the current available capacity of each candidate region.
[0077] Obtain the predicted load sequence of each region in the candidate region set, and calculate the predicted available capacity of each candidate region;
[0078] Obtain the prediction confidence sequence of each candidate region, determine the fusion weight of the current available capacity and the predicted available capacity based on the prediction confidence sequence, and add the weighted current available capacity and the weighted predicted available capacity to obtain the comprehensive available capacity of each candidate region.
[0079] The prediction trend slope is calculated based on the current available capacity of each candidate region and the predicted load sequence of the region. Based on the prediction trend slope, each candidate region is determined to be a deteriorating region or an improving region, and the priority coefficient is adjusted accordingly.
[0080] The overall available capacity of each candidate region is multiplied by the priority coefficient, and load balancing is adjusted in combination with the number of tasks allocated to each candidate region in the current scheduling cycle to obtain the overall score of each candidate region.
[0081] The region with the highest overall score is selected as the adaptive computing power scheduling region, and the corresponding edge processor is determined as the adaptive edge processor.
[0082] First, at the central controller, the computing power type is matched to each computing power scheduling region based on the computing power type requirements in the business demand information, thus filtering out a set of candidate regions that support the computing power type requirements. Specifically, the computing power type requirements carried in the business demand information specify the type of computing power service required by the business, such as general-purpose computing, graphics processor accelerated computing, or neural network processor accelerated computing. Each computing power scheduling region indicates the type of computing power service it can provide by indicating the computing power type it supports. The central controller compares the computing power type of the business demand with the supported computing power types of each region, retaining only regions that support that computing power type, forming a set of candidate regions, thereby eliminating regions that do not have the corresponding computing power capabilities in the initial stage of region adaptation.
[0083] Secondly, network constraints are applied to the candidate region set based on the maximum tolerable latency and minimum bandwidth requirements in the business demand information, eliminating regions where the ingress latency exceeds the maximum tolerable latency or where the available bandwidth is less than the minimum bandwidth requirement. The maximum tolerable latency is the maximum end-to-end delay allowed for a business request; exceeding this latency will fail to meet the service quality requirements. The minimum bandwidth requirement is the minimum network bandwidth required during business processing; falling below this bandwidth will cause data transmission bottlenecks.
[0084] The central controller compares the regional ingress latency of each candidate region with the maximum tolerable latency, and compares the available bandwidth of the region with the minimum bandwidth requirement. Regions that do not meet the network constraints are eliminated, further narrowing down the candidate region range and ensuring that the selected regions have the capability to meet the business network requirements.
[0085] Then, the average processor utilization, average memory utilization, and total task queue length of each region in the candidate region set are obtained to calculate the current available capacity of each candidate region. The current available capacity reflects the region's ability to handle additional services at the current moment. Specifically, the average processor utilization and average memory utilization represent the occupancy of CPU and memory resources, respectively, while the total task queue length reflects the backlog of tasks to be processed.
[0086] Specifically, the current available capacity is obtained by weighted summing of the region's average processor idle rate, average memory idle rate, and queue idle rate. The region's average processor idle rate equals 1 minus processor utilization, the average memory idle rate equals 1 minus memory utilization, and the average queue idle rate is calculated based on the ratio of task queue length to the maximum queue capacity. The weighting coefficients for the weighted summation are set comprehensively based on the contribution of each resource to the region's overall service capacity and the type of core business in the region. For example, for compute-intensive service regions, the weighting coefficient for processor idle rate can be set to 0.6, the weighting coefficient for memory idle rate to 0.2, and the weighting coefficient for queue idle rate to 0.2. A higher calculated current available capacity indicates more abundant resource margins in the region at the current moment, making it more suitable for handling new service requests.
[0087] Furthermore, the predicted load sequence for each region in the candidate region set is obtained, and the predicted available capacity for each candidate region is calculated. The predicted load sequence reflects the comprehensive load value of the region at each predicted moment within a preset time window, representing the resource occupancy trend over a future period. By taking the arithmetic mean of the predicted values at each moment in the predicted load sequence, the average predicted load of the region over the future period can be obtained, and its complementary value is used as the predicted available capacity. Specifically, the predicted available capacity equals 1 minus the arithmetic mean of the predicted load sequence. The higher the predicted available capacity, the more abundant the resource margin of the region in the future period, and the more suitable it is to handle business requests that may last for a certain period of time.
[0088] Secondly, the prediction confidence sequence for each candidate region is obtained. Based on the prediction confidence sequence, the fusion weight between the current available capacity and the predicted available capacity is determined. The weighted current available capacity is then added to the weighted predicted available capacity to obtain the comprehensive available capacity for each candidate region. The prediction confidence sequence reflects the reliability of each predicted value in the regional predicted load sequence.
[0089] Specifically, the steps for calculating the current available capacity, predicted available capacity, and determining the fusion weights for each candidate region include:
[0090] The current available capacity of each candidate region is obtained by weighted summing of the region's average processor idle rate, region's average memory idle rate, and region's queue idle rate.
[0091] Calculate the arithmetic mean of the regional predicted load sequences for each candidate region, and use the complementary value of the arithmetic mean as the predicted available capacity for each candidate region.
[0092] Calculate the arithmetic mean of the prediction confidence sequence of each candidate region, multiply the arithmetic mean by the preset maximum prediction weight ratio to obtain the prediction weight value, and use the prediction weight value as the fusion weight of the prediction available capacity.
[0093] The difference between 1 and the predicted weight value is used as the fusion weight of the current available capacity.
[0094] First, the average processor idle rate, average memory idle rate, and average queue idle rate of each candidate region are weighted and summed to obtain the current available capacity of each candidate region. Similarly, the average processor idle rate equals 1 minus the average processor utilization rate, reflecting the remaining CPU resources in that region; the average memory idle rate equals 1 minus the average memory utilization rate, reflecting the remaining memory resources; and the average queue idle rate is calculated based on the ratio of the total task queue length to the preset queue capacity limit, reflecting the remaining task processing capacity of that region. These three idle rate indicators are then weighted and summed according to preset weighting coefficients to obtain the current available capacity. The weighting coefficients are set based on the contribution of various resources to the overall service carrying capacity of the region and the core service types of the region. This current available capacity is used to characterize the immediate resource margin of the candidate region for handling additional services at the current moment.
[0095] Secondly, the arithmetic mean of the regional predicted load sequences for each candidate region is calculated, and the complementary value of the arithmetic mean is used as the predicted available capacity for each candidate region. The regional predicted load sequence contains the regional composite load predicted values at multiple prediction times within a preset time window. The arithmetic mean of all predicted values in this sequence is calculated to obtain the average predicted load for the region over the future time period. Since the regional composite load value ranges from 0 to 1, its complementary value, 1 minus the average predicted load, reflects the average resource surplus of the region over the future time period, and this is used as the predicted available capacity. This predicted available capacity is used to characterize the expected resource margin for the candidate region to handle business in the future.
[0096] Then, the arithmetic mean of the prediction confidence sequence of each candidate region is calculated, and the arithmetic mean is multiplied by the preset maximum prediction weight ratio to obtain the prediction weight value. The prediction weight value is used as the fusion weight of the predicted available capacity, and the difference between 1 and the prediction weight value is used as the fusion weight of the current available capacity.
[0097] Specifically, the prediction confidence sequence reflects the reliability of each predicted value in the regional prediction load sequence. The overall confidence of the regional prediction results is obtained by calculating the arithmetic mean of this sequence. The preset maximum prediction weight percentage is a pre-defined upper limit of the weight that the predicted available capacity will occupy during the fusion process; for example, it can be set to 0.2, used to control the maximum impact of prediction information on the overall available capacity. The prediction weight value is obtained by multiplying the arithmetic mean of the prediction confidence scores by the preset maximum prediction weight percentage. This prediction weight value increases with increasing prediction confidence and decreases with decreasing prediction confidence. The prediction weight value is directly used as the fusion weight for the predicted available capacity, while the fusion weight for the current available capacity is equal to 1 minus the prediction weight value.
[0098] By determining the fusion weights as described above, when the confidence level of the prediction result is high, the system trusts the prediction information more, and the predicted available capacity has a higher weight in the overall available capacity; when the confidence level of the prediction result is low, the system relies more on the current measured information, and the current available capacity has a higher weight in the overall available capacity, thereby achieving adaptive fusion of prediction information and real-time information.
[0099] Furthermore, the predicted trend slope is calculated based on the current available capacity and the predicted load sequence of each candidate region. Based on the predicted trend slope, each candidate region is determined to be either a deteriorating or improving region, and the priority coefficient is adjusted accordingly. The predicted trend slope quantifies the changing trend of regional load over a future period; a positive slope indicates an upward trend in load and a deteriorating regional resource situation, while a negative slope indicates a downward trend in load and an improving regional resource situation.
[0100] Specifically, based on the predicted trend slope, each candidate region is determined to be either a deteriorating region or an improving region, and the priority coefficient is adjusted accordingly, including:
[0101] The load forecast values for the first and last time windows are extracted from the regional predicted load sequences of each candidate region. The slope of the prediction trend is calculated based on the ratio of the difference between the two values to the load forecast value of the first time window.
[0102] If the current available capacity is less than a preset low load threshold and the predicted trend slope is greater than a preset slope threshold, then the priority coefficient of the candidate region is set to the first priority coefficient value, wherein the first priority coefficient value is less than 1.
[0103] If the current available capacity is greater than the preset high load threshold and the predicted trend slope is less than the negative preset slope threshold, then the priority coefficient of the candidate region is set to the second priority coefficient value, wherein the second priority coefficient value is greater than 1.
[0104] If the current available capacity is greater than or equal to the preset low load threshold and less than or equal to the preset high load threshold, then the priority coefficient of the candidate region is set to 1.
[0105] In the specific calculation, firstly, the load forecast values for the first and last time windows are extracted from the regional predicted load sequences of each candidate region. The slope of the prediction trend is then calculated based on the ratio of the difference between the two values to the load forecast value for the first time window. The regional predicted load sequence covers multiple prediction moments within a preset time window. This sequence is then divided into multiple consecutive time windows, each containing several prediction moments. Optionally, the number of time windows is determined based on the total length of the preset time window and the granularity required for load trend analysis. For example, the preset time window can be divided into five consecutive time windows, each containing the same number of prediction moments. This ensures that the extraction of load forecast values from the first and last time windows effectively reflects the overall trend of regional load changes throughout the entire prediction period.
[0106] Specifically, the first time window corresponds to the initial stage of forecasting, and its load forecast value reflects the region's load level at the beginning of the forecast; the last time window corresponds to the final stage of forecasting, and its load forecast value reflects the region's load level at the end of the forecast. The formula for calculating the forecast trend slope is: the forecast trend slope equals the difference between the load forecast value of the last time window and the load forecast value of the first time window, divided by the load forecast value of the first time window. When the forecast trend slope is positive, it indicates that the regional load is trending upward and the resource situation is deteriorating; when it is negative, it indicates that the regional load is trending downward and the resource situation is improving; the larger the absolute value, the more significant the change trend.
[0107] Secondly, if the current available capacity is less than the preset low load threshold and the predicted trend slope is greater than the preset slope threshold, then the priority coefficient of the candidate region is set to the first priority coefficient value, where the first priority coefficient value is less than 1. The preset low load threshold is a critical value used to determine whether the region's current resources are strained; for example, it can be set to 0.2, meaning that resources are considered strained when the current available capacity is less than 20%. The preset slope threshold is a critical value used to determine whether the load change trend is significant; for example, it can be set to 0.1.
[0108] When a region is currently experiencing resource shortages and its load is showing a significant upward trend, it indicates that the region not only has limited current capacity but its resource situation will also deteriorate further in the future, classifying it as a deteriorating region. In this case, its priority coefficient should be set to a value less than 1, such as 0.5 or 0.6, to reduce the region's priority in subsequent comprehensive scoring and avoid scheduling new services to regions that are about to become overloaded.
[0109] Furthermore, if the current available capacity is greater than a preset high load threshold and the predicted trend slope is less than a negative preset slope threshold, then the priority coefficient of the candidate region is set to the second priority coefficient value, where the second priority coefficient value is greater than 1. The preset high load threshold is a critical value used to determine whether the current resources of a region are sufficient; for example, it can be set to 0.6, meaning that resources are considered sufficient when the current available capacity is higher than 60%.
[0110] When a region currently has ample resources and its load shows a significant downward trend, it indicates that the region not only has strong current capacity but also that its resource situation will further improve in the future, classifying it as an improving region. In this case, its priority coefficient is set to a value greater than 1, such as 1.5 or 2.0, to increase the region's priority in subsequent comprehensive scoring and guide business requests towards regions with improving resource conditions.
[0111] Furthermore, if the current available capacity is greater than or equal to the preset low load threshold and less than or equal to the preset high load threshold, the priority coefficient of the candidate region is set to 1. Specifically, when the current available capacity of a region is within the normal range, neither resource-scarce nor resource-sufficient, it indicates that the current load of the region is moderate and does not need to be adjusted through the priority coefficient. Therefore, the priority coefficient is set to 1 to keep its overall score unaffected by priority.
[0112] By setting the priority coefficients differently, the system can dynamically adjust the weight of each candidate region in the comprehensive score based on the current resource status and future load change trends of the region, so as to achieve forward-looking load balancing scheduling, avoid business requests being scheduled to regions that are about to deteriorate, and give priority to regions whose resource status is improving.
[0113] Furthermore, the overall available capacity of each candidate region is multiplied by a priority coefficient, and load balancing is adjusted based on the number of tasks already allocated to each candidate region within the current scheduling period to obtain a comprehensive score for each candidate region. Specifically, within the current scheduling period, each candidate region may have already handled several service requests. To avoid localized overload caused by centrally scheduling multiple service requests to the same region within the same scheduling period, a load balancing factor needs to be introduced.
[0114] Specifically, load balancing is adjusted based on the number of tasks already allocated to each candidate region within the current scheduling period to obtain a comprehensive score for each candidate region, including:
[0115] Obtain the number of tasks assigned to each candidate region within the current scheduling period, calculate the total number of tasks assigned to all candidate regions, and use the ratio of the number of tasks assigned to each candidate region to the total as the task allocation ratio.
[0116] The difference between 1 and the task allocation ratio is used as the load balancing factor for each candidate region. The region with more allocated tasks has a smaller load balancing factor.
[0117] The overall available capacity of each candidate region is multiplied by the priority coefficient and then by the load balancing factor to obtain the overall score of each candidate region.
[0118] First, obtain the number of tasks assigned to each candidate region within the current scheduling period, calculate the sum of the number of tasks assigned to all candidate regions, and use the ratio of the number of tasks assigned to each candidate region to the total as the task allocation percentage. The current scheduling period refers to the time window during which the central controller performs region adaptation evaluation. Within this period, the central controller may have assigned different service requests to multiple candidate regions. The number of assigned tasks reflects the service load that each candidate region has already undertaken within this scheduling period. By calculating the proportion of the number of tasks assigned to each candidate region to the total number of tasks assigned to all candidate regions, the task allocation percentage is obtained. The higher this task allocation percentage, the more tasks have been assigned to that region within this scheduling period, and the heavier the load.
[0119] Secondly, the difference between 1 and the task allocation ratio is used as the load balancing factor for each candidate region. The task allocation ratio reflects the concentration of task allocation in a region within the current scheduling cycle; a higher ratio indicates that the region is already handling a relatively large number of tasks. The load balancing factor is obtained by calculating 1 minus the task allocation ratio. When the task allocation ratio approaches 0, the load balancing factor approaches 1; when the task allocation ratio approaches 1, the load balancing factor approaches 0. Therefore, the region with the more tasks allocated, the smaller its load balancing factor, and the lower its weight in the subsequent comprehensive score. The introduction of this load balancing factor aims to avoid continuously scheduling multiple business requests to the same region within the same scheduling cycle, preventing some regions from becoming overloaded while others remain idle due to uneven allocation, thereby achieving load balancing between regions.
[0120] Then, the overall available capacity of each candidate region is multiplied by a priority coefficient and then by a load balancing factor to obtain the overall score for each candidate region. The overall available capacity reflects the region's resource carrying capacity at the current moment and over a future period; the priority coefficient reflects the impact of changes in the region's resource status on the selection priority; and the load balancing factor reflects the balance requirement for task allocation within the current scheduling cycle. The overall score obtained by multiplying these three factors takes into account the region's immediate resource margin, future resource trends, and cross-regional load balancing constraints.
[0121] Specifically, candidate regions with higher comprehensive scores have a greater advantage in maximizing global resource utilization and load balancing between regions. The central controller can then select the region with the highest comprehensive score as the appropriate computing power scheduling region. Through the calculation method of the comprehensive score described above, precise scheduling of business requests can be achieved from a global perspective, prioritizing regions with abundant resources and good trends while avoiding local overload within the same scheduling cycle.
[0122] Ultimately, the region with the highest comprehensive score is selected as the suitable computing power scheduling region, and the corresponding edge processor in that region is determined as the suitable edge processor. The comprehensive score comprehensively reflects the current resource status, future resource trends, prediction confidence, regional change trends, and load balancing status within the current scheduling cycle of the candidate region. The region with the highest score is the optimal choice under the goal of maximizing global resource utilization and inter-regional load balancing. Through the above multi-dimensional and multi-level region adaptation evaluation, the central controller can accurately schedule business requests to the most suitable computing power scheduling region from a global perspective, providing a foundation for subsequent matching of computing power nodes within the region.
[0123] S30: In the adaptive edge processor, based on the node status information of multiple computing nodes in the adaptive computing power scheduling area at the current time, a computing power node matching analysis is performed with the goal of maximizing load balancing within the area to determine the optimal computing power node.
[0124] The central controller determines the most suitable computing power scheduling region to handle current business requests through regional adaptation assessment. However, this region typically contains multiple computing nodes with computing resources, and the resource status, load level, and network conditions of each node vary. If scheduling is based solely on the regional assessment results, ignoring the differences in resource distribution at the node level, it may lead to uneven load distribution among nodes within the region, with some nodes overloaded while others are idle, affecting overall resource utilization efficiency. Therefore, after determining the suitable computing power scheduling region, further fine-grained scheduling within the region is required.
[0125] Specifically, based on the node status information of multiple computing nodes in the adapted computing power scheduling region at the current time, a computing power node matching analysis is performed with the goal of maximizing load balancing within the region to determine the optimal computing power node, including:
[0126] Collect node status information of each computing node in the adaptive computing power scheduling area, wherein the node status information includes node processor utilization, node memory utilization, node task queue length, node access latency, node available bandwidth, node computing power type and node available computing power size;
[0127] Based on the computing power type requirement, computing power size requirement, memory requirement, maximum tolerable latency and minimum bandwidth requirement in the business requirement information, each computing power node is filtered layer by layer to select a set of candidate nodes that meet all requirements.
[0128] Obtain the optimized weight vector transmitted by the central controller, wherein the optimized weight vector includes node processor utilization weight, node memory utilization weight, node task queue length weight, and node access latency weight.
[0129] The node available capacity of each candidate node is obtained by weighting and summing the node processor idle rate, node memory idle rate, node queue idle rate and node latency satisfaction of each candidate node according to the optimized weight vector.
[0130] Obtain the number of tasks assigned to each candidate node in the current scheduling period, calculate the sum of the number of tasks assigned to all candidate nodes, take the ratio of the number of tasks assigned to each candidate node to the sum as the node task allocation ratio, and take the difference between 1 and the node task allocation ratio as the node load balancing factor of each candidate node.
[0131] Multiply the available capacity of each candidate node by the node load balancing factor to obtain the fitness of each candidate node, and select the candidate node with the highest fitness as the optimal computing power node.
[0132] First, node status information for each computing node within the adaptive computing power scheduling area is collected. Node status information is key data reflecting the current resource status and network conditions of each computing node, specifically including node processor utilization, node memory utilization, node task queue length, node access latency, node available bandwidth, node computing power type, and node available computing power size. Specifically, node processor utilization and node memory utilization represent the occupancy of central processing unit (CPU) resources and memory resources, respectively; node task queue length reflects the backlog of tasks currently pending on the node; node access latency indicates the time required for a service request to be transmitted from the edge processor to the node; node available bandwidth indicates the remaining network bandwidth between the node and the edge processor; node computing power type indicates the type of computing power service the node can provide; and node available computing power size reflects the total remaining computing power resources of the node. Collecting this node status information provides the foundation for subsequent node screening and matching analysis.
[0133] Secondly, based on the computing power type requirement, computing power size requirement, memory requirement, maximum tolerable latency, and minimum bandwidth requirement in the business requirements information, each computing power node is filtered layer by layer to select a set of candidate nodes that meet all requirements. Specifically, this filtering process includes:
[0134] Based on the computing power type requirements in the business needs information, each computing power node is matched for computing power type, and nodes that support the computing power type requirements are selected. The computing power type requirements specify the category of computing power services required by the business, such as general computing, graphics processor accelerated, or neural network processor accelerated. Each computing power node indicates the category of computing power services it can provide through its node computing power type. The node computing power type is compared with the business computing power type requirements, and only nodes that support that computing power type are retained, forming a candidate node set.
[0135] The candidate node set is filtered based on the computing power requirements in the business needs information, removing nodes whose available computing power is less than the required computing power. A node's available computing power reflects its current total remaining computing resources, while the computing power requirement carried in the business request specifies the minimum amount of computing power resources required for business processing. For each node in the candidate node set, its available computing power is compared with the required computing power. If a node's available computing power is lower than the required value, the node is removed from the candidate node set, ensuring that the remaining nodes have the computing power capabilities to meet the business's computing power requirements.
[0136] The candidate node set is filtered based on the memory requirements in the business needs information, removing nodes whose remaining memory is less than the memory requirement. The remaining memory of a node reflects the currently available memory resource capacity of that node, while the memory requirement carried in the business request specifies the minimum memory capacity required for business processing. For each node in the candidate node set, its remaining memory is compared with the memory requirement. If the remaining memory of a node is lower than the requirement value, the node is removed from the candidate node set, ensuring that the retained nodes have the storage capacity to meet the business memory resource requirements.
[0137] The candidate node set is filtered based on the maximum tolerable latency and minimum bandwidth requirement in the business requirements information. Nodes with access latency exceeding the maximum tolerable latency and nodes with available bandwidth less than the minimum bandwidth requirement are removed. Node access latency represents the time required for a business request to travel from the edge processor to the node; the maximum tolerable latency carried in the business request indicates the maximum allowable delay for the business; available bandwidth indicates the remaining network bandwidth between the node and the edge processor; and the minimum bandwidth requirement carried in the business request indicates the minimum bandwidth guarantee required for business processing. For each node in the candidate node set, its access latency is compared with the maximum tolerable latency; if the access latency exceeds the maximum tolerable latency, the node is removed; its available bandwidth is compared with the minimum bandwidth requirement; if the available bandwidth is less than the minimum bandwidth requirement, the node is removed. Through this network constraint filtering, it is ensured that the retained nodes have the network transmission capabilities to meet the business service quality requirements.
[0138] In summary, through the above-mentioned layer-by-layer filtering, nodes that do not meet any of the requirements are eliminated, and nodes that simultaneously meet all hard constraints of computing power type, computing power size, memory capacity, access latency, and available bandwidth are retained to form a candidate node set, ensuring that the nodes in the subsequent matching analysis all have the basic ability to undertake the business request.
[0139] Next, the optimized weight vector transmitted by the central controller is obtained. This optimized weight vector is a weight configuration dynamically generated by the central controller based on the global resource optimization objective, including node processor utilization weight, node memory utilization weight, node task queue length weight, and node access latency weight. This optimized weight vector reflects the central controller's emphasis on different resource dimensions within the current scheduling cycle. For example, when processor resources are scarce globally, the central controller can increase the processor utilization weight, guiding edge processors to prioritize nodes with lower processor loads during node matching. By transmitting the optimized weight vector to the edge processors, the central controller indirectly controls node matching within the region, ensuring that node selection within the region aligns with the global optimization objective.
[0140] Secondly, the available capacity of each candidate node is obtained by weighting and summing the node processor idle rate, node memory idle rate, node queue idle rate, and node latency satisfaction based on the optimized weight vector. Specifically, the node processor idle rate equals 1 minus the node processor utilization rate, reflecting the remaining CPU resources of the node; the node memory idle rate equals 1 minus the node memory utilization rate, reflecting the remaining memory resources; the node queue idle rate is calculated based on the ratio of the node task queue length to the preset queue capacity limit, reflecting the remaining task processing capacity of the node; and the node latency satisfaction is calculated based on the ratio of the node access latency to the maximum tolerable latency of the service, reflecting the degree to which the node meets service requirements in terms of network latency. The above four indicators are then weighted and summed according to the corresponding weight coefficients in the optimized weight vector to obtain the node's available capacity. A higher node available capacity indicates that the node has more comprehensive resource margin at the current moment and is more suitable for handling new service requests.
[0141] Furthermore, obtain the number of tasks assigned to each candidate node within the current scheduling period, calculate the sum of the number of tasks assigned to all candidate nodes, take the ratio of the number of tasks assigned to each candidate node to the sum as the node task allocation ratio, and take the difference between 1 and the node task allocation ratio as the node load balancing factor of each candidate node.
[0142] Within the current scheduling cycle, the edge processor may have allocated different service requests to multiple candidate nodes. The number of allocated tasks reflects the service load that each node has already handled within this scheduling cycle. The higher the task allocation ratio of a node, the more tasks that node has been allocated within this scheduling cycle, and the heavier the load. By subtracting the node's task allocation ratio from 1, the node load balancing factor is obtained. The more tasks a node has been allocated, the smaller its load balancing factor. The introduction of this load balancing factor aims to avoid continuously scheduling multiple service requests to the same node within the same scheduling cycle, preventing node overload due to uneven allocation, thereby achieving load balancing within the region.
[0143] Finally, the available capacity of each candidate node is multiplied by the node load balancing factor to obtain the fitness of each candidate node. The candidate node with the highest fitness is selected as the optimal computing power node. Fitness comprehensively reflects the candidate node's real-time resource margin and the load balancing status within the current scheduling cycle. The node with the highest fitness is the optimal choice under the goal of maximizing load balancing within the region.
[0144] Through the above node matching analysis, the adapted edge processor can accurately select the most suitable computing power node to undertake the current business request from multiple candidate nodes within the region. This satisfies the various resource and network constraints of the business and achieves load balancing among nodes within the region, thereby completing the fine-grained scheduling of computing power resources based on regional adaptation.
[0145] S40: Schedule the service request to the optimal computing power node for processing.
[0146] Finally, the service request is scheduled to the optimal computing power node determined through the aforementioned steps for processing. After the central controller completes the regional adaptation assessment, determines the suitable computing power scheduling region and the suitable edge processor, and the adapted edge processor completes the node matching analysis and determines the optimal computing power node, the computing task corresponding to the service request is sent to the optimal computing power node, and the node actually executes the computing task required by the service.
[0147] Through the above scheduling process, service requests are accurately allocated to the computing power scheduling area with the most abundant resources and the most balanced load in the global scope, and further allocated to the computing power node with the most suitable resources and the most balanced load in this area, realizing two-level collaborative scheduling from the global area to the local node.
[0148] In summary, the embodiments of this application have at least the following technical effects:
[0149] This invention first extracts multi-dimensional information from business requests, including computing power type, computing power size requirements, memory requirements, maximum tolerable latency, and minimum bandwidth requirements, achieving a refined perception of user business needs and providing precise constraints for subsequent scheduling decisions. Second, at the central controller level, this invention integrates computing power status information and computing power prediction information uploaded by multiple edge processors. It performs region adaptation evaluation with the goal of maximizing global resource utilization and inter-regional load balancing. This not only considers the current resource idle status but also incorporates the predicted results of regional load trends, avoiding scheduling lag and inter-regional load imbalance caused by resource fluctuations. Third, at the edge processor adaptation level, this invention performs node matching analysis based on the status information of multiple computing power nodes within a region, aiming to maximize regional load balancing. This achieves coordinated optimization of region selection and node allocation, overcoming the shortcomings of the fragmented two-level scheduling in traditional methods. Finally, this invention uses a two-level linkage dynamic computing power routing mechanism to accurately schedule business requests to the optimal computing power node for processing. Under the premise of meeting multiple constraints such as business latency and bandwidth, it improves the overall utilization rate of the network's computing power resources and ensures the balance and stability of the system under different load conditions.
[0150] Example 2, as Figure 2As shown, based on the same inventive concept as the dynamic computing power routing calculation method for business needs provided in Embodiment 1, this embodiment of the invention also provides a dynamic computing power routing calculation system for business needs, including:
[0151] The business requirement extraction module 11 is used to receive user business requests and extract computing power type requirements, computing power size requirements, memory requirements, maximum tolerable latency and minimum bandwidth requirements from the business requests as business requirement information and transmit them to the central controller.
[0152] The central controller module 12 is used to perform a regional adaptation assessment on the central controller based on several computing power status information and several computing power prediction information of several computing power scheduling areas uploaded by several edge processors, combined with the business demand information, with the goal of maximizing global resource utilization and inter-regional load balancing, and to determine the suitable computing power scheduling areas and suitable edge processors.
[0153] The adaptive edge processor module 13 is used to perform a computing node matching analysis based on the node status information of multiple computing nodes in the adaptive computing power scheduling area at the current time, with the goal of maximizing load balancing within the area, and to determine the optimal computing node.
[0154] The scheduling and execution module 14 is used to schedule the service request to the optimal computing power node for processing.
[0155] Specifically, the business requirement extraction module 11 is used for:
[0156] The system receives user service requests and extracts the computing power type requirement, computing power size requirement, memory requirement, maximum tolerable latency, and minimum bandwidth requirement from the service requests as service requirement information, which is then transmitted to the central controller.
[0157] The central controller module 12 is specifically used for:
[0158] First, the steps for obtaining several computing power status information and several computing power prediction information from several computing power scheduling areas uploaded by several edge processors include:
[0159] Collect the regional average processor utilization, regional average memory utilization, regional total task queue length, regional entry latency, regional available bandwidth, and regional supported computing power types of each edge processor in its respective computing power scheduling region to obtain several computing power status information.
[0160] Based on the historical load sequences of the computing power scheduling regions where several edge processors are located, computing power load prediction is performed within a preset time window to obtain several computing power prediction information. Each computing power prediction information includes a region prediction load sequence and a prediction confidence sequence.
[0161] Specifically, based on the historical load sequences of the computing power scheduling regions where several edge processors are located, computing power load prediction is performed within a preset time window to obtain several computing power prediction information, including:
[0162] Randomly select any edge processor from the plurality of edge processors as the first edge processor, and obtain the first historical load sequence of the first edge processor in the historical time zone. The historical load sequence includes the regional comprehensive load value at multiple historical moments. The regional comprehensive load value is obtained by weighted evaluation of the regional average CPU utilization, the regional average memory utilization, and the regional total task queue length.
[0163] The first load variation coefficient is obtained by calculating the load volatility of the first historical load sequence.
[0164] The first computing power load predictor is activated according to the first load variation coefficient, and the first computing power prediction information within the preset time window is predicted based on the first historical load sequence and added to the plurality of computing power prediction information.
[0165] Specifically, the first computing power load predictor is activated according to the first load variation coefficient, and the first computing power prediction information is obtained based on the first historical load sequence, including:
[0166] Based on the historical operation records of the first edge processor, and constrained by the time span of the historical time zone and the preset time window, a sample historical load sequence set and a sample region predicted load sequence set are collected and used as sample training data.
[0167] The training data is subjected to K-fold cross-multiplication with replacement to obtain a training set of K samples;
[0168] Using the historical load sequence of the sample as input data and the predicted load sequence of the sample region as supervision label, the long short-term memory network is trained to converge using the K sample training sets to obtain K first computing load prediction branches, which are then combined to obtain the first computing load predictor.
[0169] The ratio of the first load variation coefficient to the preset baseline load variation coefficient of the first edge processor is used as the first prediction difficulty coefficient. The ratio of the first prediction difficulty coefficient to the preset baseline branch number is rounded down to obtain the number of adaptation branches P, where the preset baseline branch number is half K, and P is greater than or equal to 3 and less than or equal to K.
[0170] P branches are randomly selected from the K first computing load prediction branches of the first computing load predictor, and predictions are made according to the first historical load sequence to obtain P initial regional predicted load sequences. The first regional predicted load sequence is obtained by calculating the average value of the P initial regional predicted load sequences at the same time.
[0171] Calculate the reciprocal of the variance of the P initial region predicted load sequences at the same moment, and perform normalization processing to obtain the first prediction confidence sequence. Use the first region predicted load sequence and the first prediction confidence sequence as the first computing power prediction information.
[0172] Furthermore, based on several computing power status information and several computing power prediction information from several computing power scheduling regions, combined with the aforementioned business demand information, a region adaptation assessment is conducted with the goal of maximizing global resource utilization and inter-regional load balancing to determine suitable computing power scheduling regions and suitable edge processors, including:
[0173] In the central controller, the computing power type is matched for each computing power scheduling region according to the computing power type requirements in the business demand information, and a set of candidate regions that support the computing power type requirements is selected.
[0174] Based on the maximum tolerable latency and minimum bandwidth requirement in the business requirement information, the candidate region set is filtered by network constraints, and regions whose ingress latency exceeds the maximum tolerable latency or whose available bandwidth is less than the minimum bandwidth requirement are eliminated.
[0175] Obtain the average processor utilization, average memory utilization, and total task queue length of each region in the candidate region set, and calculate the current available capacity of each candidate region.
[0176] Obtain the predicted load sequence of each region in the candidate region set, and calculate the predicted available capacity of each candidate region;
[0177] Obtain the prediction confidence sequence of each candidate region, determine the fusion weight of the current available capacity and the predicted available capacity based on the prediction confidence sequence, and add the weighted current available capacity and the weighted predicted available capacity to obtain the comprehensive available capacity of each candidate region.
[0178] The prediction trend slope is calculated based on the current available capacity of each candidate region and the predicted load sequence of the region. Based on the prediction trend slope, each candidate region is determined to be a deteriorating region or an improving region, and the priority coefficient is adjusted accordingly.
[0179] The overall available capacity of each candidate region is multiplied by the priority coefficient, and load balancing is adjusted in combination with the number of tasks allocated to each candidate region in the current scheduling cycle to obtain the overall score of each candidate region.
[0180] The region with the highest overall score is selected as the adaptive computing power scheduling region, and the corresponding edge processor is determined as the adaptive edge processor.
[0181] The specific steps for calculating the current available capacity, predicted available capacity, and determining the fusion weights for each candidate region include:
[0182] The current available capacity of each candidate region is obtained by weighted summing of the region's average processor idle rate, region's average memory idle rate, and region's queue idle rate.
[0183] Calculate the arithmetic mean of the regional predicted load sequences for each candidate region, and use the complementary value of the arithmetic mean as the predicted available capacity for each candidate region.
[0184] Calculate the arithmetic mean of the prediction confidence sequence of each candidate region, multiply the arithmetic mean by the preset maximum prediction weight ratio to obtain the prediction weight value, and use the prediction weight value as the fusion weight of the prediction available capacity.
[0185] The difference between 1 and the predicted weight value is used as the fusion weight of the current available capacity.
[0186] Based on the predicted trend slope, each candidate region is determined to be either a deteriorating region or an improving region, and the priority coefficient is adjusted accordingly, including:
[0187] The load forecast values for the first and last time windows are extracted from the regional predicted load sequences of each candidate region. The slope of the prediction trend is calculated based on the ratio of the difference between the two values to the load forecast value of the first time window.
[0188] If the current available capacity is less than a preset low load threshold and the predicted trend slope is greater than a preset slope threshold, then the priority coefficient of the candidate region is set to the first priority coefficient value, wherein the first priority coefficient value is less than 1.
[0189] If the current available capacity is greater than the preset high load threshold and the predicted trend slope is less than the negative preset slope threshold, then the priority coefficient of the candidate region is set to the second priority coefficient value, wherein the second priority coefficient value is greater than 1.
[0190] If the current available capacity is greater than or equal to the preset low load threshold and less than or equal to the preset high load threshold, then the priority coefficient of the candidate region is set to 1.
[0191] Load balancing is adjusted based on the number of tasks already allocated to each candidate region within the current scheduling period, resulting in a comprehensive score for each candidate region, including:
[0192] Obtain the number of tasks assigned to each candidate region within the current scheduling period, calculate the total number of tasks assigned to all candidate regions, and use the ratio of the number of tasks assigned to each candidate region to the total as the task allocation ratio.
[0193] The difference between 1 and the task allocation ratio is used as the load balancing factor for each candidate region. The region with more allocated tasks has a smaller load balancing factor.
[0194] The overall available capacity of each candidate region is multiplied by the priority coefficient and then by the load balancing factor to obtain the overall score of each candidate region.
[0195] Specifically, the adaptive edge processor module 13 is used for:
[0196] Based on the node status information of multiple computing nodes in the adapted computing power scheduling region at the current time, a computing power node matching analysis is performed with the goal of maximizing load balancing within the region to determine the optimal computing power node, including:
[0197] Collect node status information of each computing node in the adaptive computing power scheduling area, wherein the node status information includes node processor utilization, node memory utilization, node task queue length, node access latency, node available bandwidth, node computing power type and node available computing power size;
[0198] Based on the computing power type requirement, computing power size requirement, memory requirement, maximum tolerable latency and minimum bandwidth requirement in the business requirement information, each computing power node is filtered layer by layer to select a set of candidate nodes that meet all requirements.
[0199] Obtain the optimized weight vector transmitted by the central controller, wherein the optimized weight vector includes node processor utilization weight, node memory utilization weight, node task queue length weight, and node access latency weight.
[0200] The node available capacity of each candidate node is obtained by weighting and summing the node processor idle rate, node memory idle rate, node queue idle rate and node latency satisfaction of each candidate node according to the optimized weight vector.
[0201] Obtain the number of tasks assigned to each candidate node in the current scheduling period, calculate the sum of the number of tasks assigned to all candidate nodes, take the ratio of the number of tasks assigned to each candidate node to the sum as the node task allocation ratio, and take the difference between 1 and the node task allocation ratio as the node load balancing factor of each candidate node.
[0202] Multiply the available capacity of each candidate node by the node load balancing factor to obtain the fitness of each candidate node, and select the candidate node with the highest fitness as the optimal computing power node.
[0203] Specifically, the scheduling execution module 14 is used for:
[0204] The service request is scheduled to the optimal computing power node for processing.
[0205] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0206] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0207] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0208] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A dynamic computing power routing calculation method oriented towards business needs, characterized in that, The method is applied to a dynamic computing power routing system, the system comprising a central controller and several edge processors, and the method includes: Receive user service requests, extract computing power type requirements, computing power size requirements, memory requirements, maximum tolerable latency, and minimum bandwidth requirements from the service requests and transmit them to the central controller as service requirement information; In the central controller, based on several computing power status information and several computing power prediction information uploaded by several edge processors to several computing power scheduling areas, and combined with the business demand information, a region adaptation assessment is performed with the goal of maximizing global resource utilization and inter-region load balancing. This assessment determines suitable computing power scheduling areas and suitable edge processors, including: In the central controller, the computing power type is matched for each computing power scheduling region according to the computing power type requirements in the business demand information, and a set of candidate regions that support the computing power type requirements is selected. Based on the maximum tolerable latency and minimum bandwidth requirement in the business requirement information, the candidate region set is filtered by network constraints, and regions whose ingress latency exceeds the maximum tolerable latency or whose available bandwidth is less than the minimum bandwidth requirement are eliminated. Obtain the average processor utilization, average memory utilization, and total task queue length of each region in the candidate region set, and calculate the current available capacity of each candidate region. Obtain the predicted load sequence of each region in the candidate region set, and calculate the predicted available capacity of each candidate region; Obtain the prediction confidence sequence of each candidate region, determine the fusion weight of the current available capacity and the predicted available capacity based on the prediction confidence sequence, and add the weighted current available capacity and the weighted predicted available capacity to obtain the comprehensive available capacity of each candidate region. The predicted trend slope is calculated based on the current available capacity of each candidate region and the predicted load sequence of the region. Based on the predicted trend slope, each candidate region is determined to be a deteriorating or improving region, and the priority coefficient is adjusted accordingly, including: The load forecast values for the first and last time windows are extracted from the regional predicted load sequences of each candidate region. The slope of the prediction trend is calculated based on the ratio of the difference between the two values to the load forecast value of the first time window. If the current available capacity is less than a preset low load threshold and the predicted trend slope is greater than a preset slope threshold, then the priority coefficient of the candidate region is set to the first priority coefficient value, wherein the first priority coefficient value is less than 1. If the current available capacity is greater than the preset high load threshold and the predicted trend slope is less than the negative preset slope threshold, then the priority coefficient of the candidate region is set to the second priority coefficient value, wherein the second priority coefficient value is greater than 1. If the current available capacity is greater than or equal to the preset low load threshold and less than or equal to the preset high load threshold, then the priority coefficient of the candidate region is set to 1; The overall available capacity of each candidate region is multiplied by the priority coefficient, and load balancing is adjusted based on the number of tasks already allocated to each candidate region in the current scheduling period to obtain a comprehensive score for each candidate region, including: Obtain the number of tasks assigned to each candidate region within the current scheduling period, calculate the total number of tasks assigned to all candidate regions, and use the ratio of the number of tasks assigned to each candidate region to the total as the task allocation ratio. The difference between 1 and the task allocation ratio is used as the load balancing factor for each candidate region. The region with more allocated tasks has a smaller load balancing factor. The overall available capacity of each candidate region is multiplied by the priority coefficient and then by the load balancing factor to obtain the overall score of each candidate region. Select the region with the highest comprehensive score as the adaptive computing power scheduling region, and determine the edge processor corresponding to the region as the adaptive edge processor; In the adapted edge processor, based on the node status information of multiple computing nodes in the adapted computing power scheduling area at the current time, a computing power node matching analysis is performed with the goal of maximizing load balancing within the area to determine the optimal computing power node. The service request is scheduled to the optimal computing power node for processing.
2. The dynamic computing power routing calculation method oriented towards business needs according to claim 1, characterized in that, The steps for obtaining several computing power status information and several computing power prediction information from several computing power scheduling areas uploaded by several edge processors include: Collect the regional average processor utilization, regional average memory utilization, regional total task queue length, regional entry latency, regional available bandwidth, and regional supported computing power types of each edge processor in its respective computing power scheduling region to obtain several computing power status information. Based on the historical load sequences of the computing power scheduling regions where several edge processors are located, computing power load prediction is performed within a preset time window to obtain several computing power prediction information. Each computing power prediction information includes a region prediction load sequence and a prediction confidence sequence.
3. The dynamic computing power routing calculation method oriented towards business needs according to claim 2, characterized in that, Based on the historical load sequences of the computing power scheduling regions where several edge processors are located, computing power load prediction is performed within a preset time window to obtain several computing power prediction information, including: Randomly select any edge processor from the plurality of edge processors as the first edge processor, and obtain the first historical load sequence of the first edge processor in the historical time zone. The historical load sequence includes the regional comprehensive load value at multiple historical moments. The regional comprehensive load value is obtained by weighted evaluation of the regional average CPU utilization, the regional average memory utilization, and the regional total task queue length. The first load variation coefficient is obtained by calculating the load volatility of the first historical load sequence. The first computing power load predictor is activated according to the first load variation coefficient, and the first computing power prediction information within the preset time window is predicted based on the first historical load sequence and added to the plurality of computing power prediction information.
4. The dynamic computing power routing calculation method oriented towards business needs according to claim 3, characterized in that, The first computing power load predictor is activated according to the first load variation coefficient, and the first computing power prediction information is obtained based on the first historical load sequence, including: Based on the historical operation records of the first edge processor, and constrained by the time span of the historical time zone and the preset time window, a sample historical load sequence set and a sample region predicted load sequence set are collected and used as sample training data. The training data is subjected to K-fold cross-multiplication with replacement to obtain a training set of K samples; Using the historical load sequence of the sample as input data and the predicted load sequence of the sample region as supervision label, the long short-term memory network is trained to converge using the K sample training sets to obtain K first computing load prediction branches, which are then combined to obtain the first computing load predictor. The ratio of the first load variation coefficient to the preset baseline load variation coefficient of the first edge processor is used as the first prediction difficulty coefficient. The ratio of the first prediction difficulty coefficient to the preset baseline branch number is rounded down to obtain the number of adaptation branches P, where the preset baseline branch number is half K, and P is greater than or equal to 3 and less than or equal to K. P branches are randomly selected from the K first computing load prediction branches of the first computing load predictor, and predictions are made according to the first historical load sequence to obtain P initial regional predicted load sequences. The first regional predicted load sequence is obtained by calculating the average value of the P initial regional predicted load sequences at the same time. Calculate the reciprocal of the variance of the P initial region predicted load sequences at the same moment, and perform normalization processing to obtain the first prediction confidence sequence. Use the first region predicted load sequence and the first prediction confidence sequence as the first computing power prediction information.
5. The dynamic computing power routing calculation method oriented towards business needs according to claim 1, characterized in that, The specific steps for calculating the current available capacity, predicted available capacity, and determining the fusion weights for each candidate region include: The current available capacity of each candidate region is obtained by weighted summing of the region's average processor idle rate, region's average memory idle rate, and region's queue idle rate. Calculate the arithmetic mean of the regional predicted load sequences for each candidate region, and use the complementary value of the arithmetic mean as the predicted available capacity for each candidate region. Calculate the arithmetic mean of the prediction confidence sequence of each candidate region, multiply the arithmetic mean by the preset maximum prediction weight ratio to obtain the prediction weight value, and use the prediction weight value as the fusion weight of the prediction available capacity. The difference between 1 and the predicted weight value is used as the fusion weight of the current available capacity.
6. The dynamic computing power routing calculation method oriented towards business needs according to claim 1, characterized in that, Based on the node status information of multiple computing nodes in the adapted computing power scheduling region at the current time, a computing power node matching analysis is performed with the goal of maximizing load balancing within the region to determine the optimal computing power node, including: Collect node status information of each computing node in the adaptive computing power scheduling area, wherein the node status information includes node processor utilization, node memory utilization, node task queue length, node access latency, node available bandwidth, node computing power type and node available computing power size; Based on the computing power type requirement, computing power size requirement, memory requirement, maximum tolerable latency and minimum bandwidth requirement in the business requirement information, each computing power node is filtered layer by layer to select a set of candidate nodes that meet all requirements. Obtain the optimized weight vector transmitted by the central controller, wherein the optimized weight vector includes node processor utilization weight, node memory utilization weight, node task queue length weight, and node access latency weight. The node available capacity of each candidate node is obtained by weighting and summing the node processor idle rate, node memory idle rate, node queue idle rate and node latency satisfaction of each candidate node according to the optimized weight vector. Obtain the number of tasks assigned to each candidate node in the current scheduling period, calculate the sum of the number of tasks assigned to all candidate nodes, take the ratio of the number of tasks assigned to each candidate node to the sum as the node task allocation ratio, and take the difference between 1 and the node task allocation ratio as the node load balancing factor of each candidate node. Multiply the available capacity of each candidate node by the node load balancing factor to obtain the fitness of each candidate node, and select the candidate node with the highest fitness as the optimal computing power node.
7. A dynamic computing power routing calculation system oriented towards business needs, characterized in that, A method for performing dynamic computing power routing calculation based on business needs as described in any one of claims 1-6, comprising: The business requirement extraction module is used to receive user business requests and extract computing power type requirements, computing power size requirements, memory requirements, maximum tolerable latency and minimum bandwidth requirements from the business requests as business requirement information and transmit them to the central controller. The central controller module is used to, in the central controller, perform a region adaptation assessment based on several computing power status information and several computing power prediction information uploaded by several edge processors to several computing power scheduling regions, and in conjunction with the business demand information, with the goal of maximizing global resource utilization and inter-region load balancing, to determine the suitable computing power scheduling regions and suitable edge processors, including: In the central controller, the computing power type is matched for each computing power scheduling region according to the computing power type requirements in the business demand information, and a set of candidate regions that support the computing power type requirements is selected. Based on the maximum tolerable latency and minimum bandwidth requirement in the business requirement information, the candidate region set is filtered by network constraints, and regions whose ingress latency exceeds the maximum tolerable latency or whose available bandwidth is less than the minimum bandwidth requirement are eliminated. Obtain the average processor utilization, average memory utilization, and total task queue length of each region in the candidate region set, and calculate the current available capacity of each candidate region. Obtain the predicted load sequence of each region in the candidate region set, and calculate the predicted available capacity of each candidate region; Obtain the prediction confidence sequence of each candidate region, determine the fusion weight of the current available capacity and the predicted available capacity based on the prediction confidence sequence, and add the weighted current available capacity and the weighted predicted available capacity to obtain the comprehensive available capacity of each candidate region. The predicted trend slope is calculated based on the current available capacity of each candidate region and the predicted load sequence of the region. Based on the predicted trend slope, each candidate region is determined to be a deteriorating or improving region, and the priority coefficient is adjusted accordingly, including: The load forecast values for the first and last time windows are extracted from the regional predicted load sequences of each candidate region. The slope of the prediction trend is calculated based on the ratio of the difference between the two values to the load forecast value of the first time window. If the current available capacity is less than a preset low load threshold and the predicted trend slope is greater than a preset slope threshold, then the priority coefficient of the candidate region is set to the first priority coefficient value, wherein the first priority coefficient value is less than 1. If the current available capacity is greater than the preset high load threshold and the predicted trend slope is less than the negative preset slope threshold, then the priority coefficient of the candidate region is set to the second priority coefficient value, wherein the second priority coefficient value is greater than 1. If the current available capacity is greater than or equal to the preset low load threshold and less than or equal to the preset high load threshold, then the priority coefficient of the candidate region is set to 1; The overall available capacity of each candidate region is multiplied by the priority coefficient, and load balancing is adjusted based on the number of tasks already allocated to each candidate region in the current scheduling period to obtain a comprehensive score for each candidate region, including: Obtain the number of tasks assigned to each candidate region within the current scheduling period, calculate the total number of tasks assigned to all candidate regions, and use the ratio of the number of tasks assigned to each candidate region to the total as the task allocation ratio. The difference between 1 and the task allocation ratio is used as the load balancing factor for each candidate region. The region with more allocated tasks has a smaller load balancing factor. The overall available capacity of each candidate region is multiplied by the priority coefficient and then by the load balancing factor to obtain the overall score of each candidate region. Select the region with the highest comprehensive score as the adaptive computing power scheduling region, and determine the edge processor corresponding to the region as the adaptive edge processor; An adaptive edge processor module is used to perform a computing node matching analysis based on the node status information of multiple computing nodes in the adaptive computing power scheduling area at the current time, with the goal of maximizing load balancing within the area, to determine the optimal computing node. The scheduling and execution module is used to schedule the service request to the optimal computing power node for processing.
Citation Information
Patent Citations
Efficient computing power routing method and system based on computing power network
CN117155847A
Cross-regional computing power resource collaborative allocation method based on computing power center
CN121326559A