Deep learning-based computing power pool intelligent integration and elastic scheduling method
By constructing an intelligent integration and elastic scheduling method for computing power pools of multi-level graph convolutional networks and graph reinforcement learning, the problems of unbalanced resource allocation and performance bottlenecks in deep learning scenarios are solved, efficient resource utilization and task isolation are achieved, and system stability and the collaborative efficiency of heterogeneous computing resources are improved.
Patent Information
- Application Number
- CN202510911765.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing computing resource scheduling solutions cannot effectively cope with the dynamic characteristics of deep learning workloads, resulting in unbalanced resource allocation and low utilization. They also lack an in-depth understanding of the dynamic interaction patterns between services, making it difficult to predict resource competition and performance bottlenecks. Especially in a multi-tenant shared environment, it is difficult to ensure resource isolation and performance for tasks of different priorities.
It adopts a deep learning-based intelligent integration and elastic scheduling method for computing power pools, builds a multi-level graph convolutional network model, dynamically adjusts the granularity of microservices, combines graph reinforcement learning and multi-objective optimization algorithms, realizes adaptive allocation and isolation of resources, predicts service interaction patterns and changes in resource demand, and dynamically adjusts service deployment and expansion.
It significantly improves computing resource utilization, enhances system stability and reliability, reduces resource fragmentation and performance interference, achieves a dynamic balance between latency, throughput, and resource efficiency, and ensures task quality and collaborative efficiency of heterogeneous computing resources in a multi-tenant environment.
Smart Images

Figure CN120803715A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of distributed computing and deep learning, more particularly, it relates to a deep learning-based computing power pool intelligent integration and elastic scheduling method. BACKGROUND
[0002] With the rapid development of artificial intelligence and deep learning technology, the demand for computing power resources has shown explosive growth. Large-scale distributed training and online inference services present high concurrency, large fluctuations, and strong heterogeneity in the use of computing resources. In order to improve resource utilization efficiency, microservice architecture is widely used in the management and scheduling of computing power resources. By splitting complex systems into loosely coupled microservice components, more flexible deployment and expansion can be achieved. However, traditional microservice architecture still faces many challenges in the face of deep learning scenarios.
[0003] The current mainstream computing power resource scheduling scheme is mainly based on static service dependency relationships and simple load balancing strategies. This approach cannot effectively deal with the dynamic characteristics of deep learning workloads, resulting in uneven resource allocation, low utilization, and other problems. At the same time, due to the lack of in-depth understanding of the dynamic interaction patterns between services, existing scheduling systems cannot accurately predict resource competition and performance bottlenecks, often resulting in resource fragmentation and performance interference.
[0004] In a multi-tenant shared environment, resource isolation and performance protection between tasks of different priorities also face great challenges. Existing technologies mainly rely on static resource quotas and simple priority mechanisms, which cannot dynamically adjust resource allocation strategies based on actual load conditions.
[0005] In addition, on a heterogeneous computing platform, how to fully utilize the computing characteristics of different types of hardware such as CPU and GPU to achieve optimal matching of computing tasks is also a pressing problem. SUMMARY
[0006] The present application provides a deep learning-based computing power pool intelligent integration and elastic scheduling method, which solves the technical problems of resource fragmentation and performance interference caused by fixed service granularity, static routing decision, lack of prediction ability, etc. in related technologies.
[0007] The present application provides a deep learning-based computing power pool intelligent integration and elastic scheduling method, comprising the following steps:
[0008] A multi-level graph convolution network is constructed to model the dynamic interaction relationship between microservices. The multi-level graph convolution network includes function-level, service-level, and cluster-level graph convolution layers, and outputs a service interaction pattern representation.
[0009] Based on the service interaction mode representation and the computing state indicators and network state indicators, a multi-dimensional scoring model is constructed to select the optimal execution path for the service request.
[0010] According to the service interaction mode representation and the scoring model results, the task characteristics and resource states are analyzed, and the microservice granularity is dynamically adjusted. The computing-intensive tasks are combined, and the tasks with high parallelism are split.
[0011] Based on the adjusted microservice structure, a graph reinforcement learning method is used to predict service interaction patterns and resource demand changes, deploy and expand key services in advance, and implement dynamic resource isolation strategies.
[0012] Based on the prediction results and resource isolation strategies, a multi-objective optimization algorithm is used to find the optimal service deployment scheme that balances delay, throughput, and resource efficiency, and a feedback adjustment mechanism is used to continuously optimize the orchestration strategy.
[0013] In a preferred embodiment, the step of constructing a multi-level graph convolutional network to model the dynamic interaction relationship between microservices includes:
[0014] Collecting interaction data and resource usage between microservices from the service mesh layer to construct an initial service interaction graph;
[0015] Constructing a graph convolutional network model containing three levels of function level, service level and cluster level to capture service interaction patterns at different granularities;
[0016] Integrating the outputs of the three levels of graph convolutional network through attention mechanism to generate a comprehensive service interaction representation.
[0017] In a preferred embodiment, the step of constructing a multi-dimensional scoring model based on the service interaction mode representation and the computing state indicators and network state indicators includes:
[0018] Collecting node computing state data through a lightweight monitoring agent, including CPU-related indicators, memory-related indicators, GPU-related indicators and node overall indicators;
[0019] Based on the collected computing state data and network state data, a multi-dimensional scoring model is constructed to calculate the comprehensive score of the service request to each candidate node;
[0020] A deep reinforcement learning method is used to dynamically optimize the dimension weights in the multi-dimensional scoring model, so that the system can adapt to routing requirements in different scenarios.
[0021] In a preferred embodiment, the step of dynamically adjusting the microservice granularity includes:
[0022] computing property analysis of microservice components, including computation intensity analysis, data dependency analysis, and parallelism analysis;
[0023] based on the computing property analysis results, determining the optimal service granularity and boundaries using a graph partitioning-based method;
[0024] based on the determined optimal service boundaries, dynamically performing service reorganization and resource allocation.
[0025] In a preferred embodiment, in the step of determining the optimal service granularity and boundaries using a graph partitioning-based method, the graph partitioning objective function considers the following objectives simultaneously:
[0026] minimizing inter-partition communication cost;
[0027] balancing the computing load of each partition;
[0028] limiting the resource demand of a single partition to not exceed the node capacity;
[0029] considering affinity and mutual exclusion constraints between services.
[0030] In a preferred embodiment, the step of predicting service interaction patterns and resource demand changes using a graph reinforcement learning method includes:
[0031] based on a multi-level graph convolutional network model, increasing the time series feature learning capability to build a service interaction prediction model;
[0032] based on the service interaction prediction results, implementing predictive service deployment and scaling;
[0033] based on service priority and predicted resource contention, implementing a dynamic resource isolation strategy.
[0034] In a preferred embodiment, in the step of implementing a dynamic resource isolation strategy, the resource isolation strategy includes the following levels:
[0035] priority scheduling isolation, providing basic isolation through scheduler priority mechanisms;
[0036] resource quota isolation, setting a resource usage upper limit through container technology;
[0037] resource reservation isolation, reserving dedicated resources for high-priority services;
[0038] physical isolation, deploying different services on physically isolated nodes.
[0039] In a preferred embodiment, the step of finding an optimal service deployment scheme that balances latency, throughput, and resource efficiency using a multi-objective optimization algorithm includes:
[0040] Formalize the service orchestration problem as a multi-objective optimization problem, considering multiple conflicting optimization objectives;
[0041] Solve the multi-objective orchestration optimization problem using a non-dominated sorting genetic algorithm;
[0042] Select a balanced solution from the obtained Pareto optimal solution set as the final orchestration scheme.
[0043] In a preferred embodiment, the step of continuously optimizing the orchestration strategy through a feedback adjustment mechanism comprises:
[0044] Continuously monitor system running status and collect key performance indicators;
[0045] Compare actual performance with expected targets and calculate deviations;
[0046] Adaptively adjust system parameters and optimization strategies according to deviation analysis results;
[0047] Use the updated model to re-solve the orchestration optimization problem and generate adjusted deployment and scheduling schemes.
[0048] In a preferred embodiment, a computer readable storage medium for storing computer readable instructions capable of running a deep learning-based computing power pool intelligent integration and elastic scheduling method when read by a computer.
[0049] The beneficial effects of the present application are:
[0050] Through dynamic service granularity adjustment and power-aware routing decision, the utilization rate of computing power resources is significantly improved, and the resource fragmentation problem is effectively solved. The system can adaptively combine computing-intensive services and split high-parallelism tasks according to task characteristics, achieving more optimal resource allocation efficiency.
[0051] Based on the service interaction prediction and resource demand prediction ability of graph reinforcement learning, the system can perceive potential performance bottlenecks and resource competition in advance, and timely perform preventive resource isolation and service expansion, greatly improving the stability and reliability of the system.
[0052] Using a multi-objective optimization method, a dynamic balance of delay, throughput and resource efficiency is achieved. Through a continuous feedback optimization mechanism, the system can find the best service orchestration strategy in different scenarios, maintaining a high overall performance level.
[0053] The multi-level resource isolation strategy and predictive service orchestration mechanism effectively reduce the performance interference in a multi-tenant environment, ensuring the service quality of tasks with different priorities, while improving the collaborative efficiency of heterogeneous computing resources. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 is a flowchart of a deep learning-based computing power pool intelligent integration and elastic scheduling method of the present application;
[0055] Figure 2 is a bar chart of resource utilization rate improvement comparison of the present application;
[0056] Figure 3 is a line chart of service stability index change of the present application;
[0057] Figure 4 is a radar chart of multi-objective optimization trade-off effect of the present application. DETAILED DESCRIPTION
[0058] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is merely meant to provide a better understanding of the subject matter described herein and can be changed in function and arrangement without departing from the scope of the present description. Various examples can omit, substitute, or add various procedures or components as appropriate or desired. Also, some features described with respect to one example can be combined with features of other examples.
[0059] In at least one embodiment of the present application, a deep learning-based computing power pool intelligent integration and elastic scheduling method is disclosed, as shown in Figure 1 includes the following steps:
[0060] Step 1, a multi-level graph convolutional network is constructed to model the dynamic interaction relationship between microservices, which contains function-level, service-level and cluster-level graph convolutional layers, and outputs a service interaction pattern representation;
[0061] Specifically, the following steps are included:
[0062] Step 1.1, collect service interaction data and construct an initial interaction graph;
[0063] Interaction data between microservices is collected from the service mesh layer, including call frequency, data transmission volume, response time and other indicators, as well as resource usage of each service node, including CPU utilization, memory occupancy, GPU computing load, etc.
[0064] Based on these data, an initial service interaction graph is constructed, in which nodes represent microservices and edges represent the calling relationship between services, and edge weights quantify the interaction intensity between services.
[0065] The data collection process needs to ensure low overhead and avoid significant impact on service performance. In addition, the data sampling frequency should be dynamically adjusted according to the system load change rate to ensure that load fluctuations are captured without generating too much redundant data.
[0066] Step 1.2, building a multi-level graph convolutional network model;
[0067] A graph convolutional network model containing three levels of function level, service level and cluster level is built to capture different granularity service interaction patterns. The graph convolutional network at each level extracts and converts features from the input data, realizing the mapping from the original interaction data to the high-level feature representation.
[0068] Function-level graph convolutional layer: capture the calling relationship and dependency between functions in microservices, learn function-level feature representation. This layer receives the function call graph as input, and generates the feature representation of each function by aggregating the information of adjacent function nodes.
[0069] Service-level graph convolutional layer: capture the direct calling relationship and data exchange pattern between microservices, learn service-level feature representation. This layer receives the service call graph as input, and generates the feature representation of each service by aggregating the information of adjacent service nodes.
[0070] Cluster-level graph convolutional layer: capture the resource competition and performance interference relationship between service clusters, learn cluster-level feature representation. This layer receives the cluster relationship graph as input, and generates the feature representation of each cluster by aggregating the information of adjacent cluster nodes.
[0071] The core of the graph convolution operation is to realize information aggregation through message passing, and the new feature of each node is composed of the weighted combination of its own feature and the feature of adjacent nodes, and the weight depends on the connection strength between nodes.
[0072]
[0073] where H (l) represents the node feature matrix of the l-th layer, which contains the feature vectors of all nodes in this layer; represents the adjacency matrix with self-connection, which describes the connection relationship between nodes in the graph, and self-connection represents the connection of a node with itself; represents the degree matrix, which is a diagonal matrix, and the elements on the diagonal represent the degree (the number of edges connected) of the corresponding node; W (l) represents the trainable weight matrix of the l-th layer, which is used to learn feature conversion; σ GCN represents a nonlinear activation function such as ReLU, which is used to introduce nonlinear transformation; H (l+1) represents the output feature matrix of the (l+1)-th layer; is a normalization term to prevent the feature scale from changing dramatically.
[0074] The multi-level graph convolutional network in this application is an innovative extension of the traditional graph convolutional network, and the specific implementation is as follows:
[0075] Network architecture: Each level of graph convolutional network consists of 2-3 graph convolutional layers, followed by a batch normalization layer and a ReLU activation function. The last layer of function-level, service-level, and cluster-level networks outputs 64, 128, and 256-dimensional feature vectors, respectively.
[0076] In different application scenarios, the network depth and width can be adjusted according to the service scale and computing resources. For example, for lightweight deployment in edge computing environments, the number of graph convolutional layers can be reduced to 1-2 layers, and the feature dimension can be reduced to 32, 64, and 128 dimensions; while for large-scale data center environments, the number of graph convolutional layers can be increased to 4-5 layers, and the feature dimension can be increased to 128, 256, and 512 dimensions to obtain richer feature representations.
[0077] Feature encoding:
[0078] Function-level node features include: function execution time distribution, call frequency, parameter type and size, resource consumption pattern, etc.
[0079] Service-level node features include: service response time, throughput, error rate, resource usage rate, etc.
[0080] Cluster-level node features include: cluster load distribution, resource utilization, network topology information, etc.
[0081] In some embodiments, different feature combinations can be selected according to specific business needs. For example, for delay-sensitive applications, the weight of time-related features can be increased; for throughput-sensitive applications, the weight of resource efficiency-related features can be increased. In addition, a dynamic feature selection mechanism can be introduced to dynamically adjust the weights of different features based on feature importance evaluation.
[0082] Learning process: The model is trained by minimizing the loss function, which includes the prediction task loss, the structure preservation loss, and the regularization term. The prediction task loss measures the difference between the model's predicted value and the true value; the structure preservation loss ensures that the learned features retain the original graph structure information; and the regularization term prevents the model from overfitting.
[0083] The loss function can be adjusted according to specific optimization goals. For example, an adversarial loss term can be added to improve the model's generalization ability and robustness. Another optional implementation is to use a hierarchical training strategy, first training each level of graph convolutional network separately, and then fine-tuning the whole model, which can speed up the training process and improve the model performance.
[0084] Step 1.3, integrate multi-level features and generate service interaction representation;
[0085] The attention mechanism integrates the outputs of three levels of graph convolutional networks to generate a comprehensive service interaction representation. The attention mechanism adaptively adjusts the importance weights of each level of features according to the characteristics of the current task, enabling the system to focus on the most relevant feature levels.
[0086] The attention weights are calculated by a trainable attention network that receives the current state and task features as input and outputs the weight distribution of each level of features.
[0087] This adaptive feature fusion method enables the system to flexibly handle different types of scheduling scenarios, such as compute-intensive tasks that may focus more on service-level features, and communication-intensive tasks that may focus more on function-level features.
[0088] In the context of a high-concurrency AI inference service system, the specific application of this multi-level graph convolutional network is as follows:
[0089] The function-level network captures the call relationships and data dependencies between components within the inference service, such as preprocessing, feature extraction, and classifiers;
[0090] The service-level network models the interaction patterns between different AI model services, such as image recognition, speech recognition, and natural language processing;
[0091] The cluster-level network analyzes the resource competition relationships between different resource domains, such as CPU computing clusters, GPU inference clusters, and storage clusters;
[0092] Through this multi-level modeling approach, the system can more accurately understand the characteristics and dependencies of different computing tasks, providing more comprehensive information support for subsequent routing decisions and resource allocation.
[0093] Step 2, based on the service interaction pattern representation and the computing power state indicators and network state indicators, build a multi-dimensional scoring model to select the optimal execution path for service requests;
[0094] Specifically, the following steps are included:
[0095] Step 2.1, collect and process computing power state data;
[0096] Deploy lightweight monitoring agents on each computing node to collect real-time node computing power state data, including CPU-related indicators (core utilization, load balancing, frequency status), memory-related indicators (usage, available amount, access delay), GPU / special accelerator-related indicators (compute unit utilization, video memory occupancy, temperature), and overall node indicators (energy consumption, temperature, resource reservation).
[0097] The collected raw data is standardized to eliminate the dimensional differences between different indicators, facilitating subsequent calculations. The standardization process uses the method of subtracting the mean value and dividing by the standard deviation to convert each indicator into the same numerical range.
[0098] The data collection process should be as lightweight as possible, using a combination of periodic sampling and event triggering to reduce monitoring overhead. In addition, the data storage uses a sliding window mechanism, which only retains data from the recent period, ensuring that trends are captured while avoiding excessive storage overhead.
[0099] Step 2.2, build a multi-dimensional scoring model;
[0100] Based on the collected computing power state data and network state data, a multi-dimensional scoring model is built to calculate the comprehensive score of service requests to each candidate node. The comprehensive score is obtained by weighted summation of scores from multiple dimensions, and the sum of weights of each dimension is 1. The scoring dimensions include but are not limited to node computing power load score, network state score, task affinity score, historical performance score, and resource efficiency score.
[0101] For different types of service requests, the importance of each scoring dimension is different. For example, for compute-intensive tasks, the weight of computing power load score should be higher; for communication-intensive tasks, the weight of network state score should be higher; for delay-sensitive tasks, the weight of historical performance score should be higher.
[0102] Step 2.3, apply reinforcement learning to optimize routing strategy;
[0103] The deep reinforcement learning method is used to dynamically optimize the dimension weights in the multi-dimensional scoring model, so that the system can adapt to routing requirements in different scenarios.
[0104] Specifically, the state space is defined to include information such as the resource state, load distribution, and request queue of the current computing power pool;
[0105] The action space is defined as adjusting the weights of each scoring dimension;
[0106] The reward function considers indicators such as response time, resource utilization, and load balancing degree;
[0107] The deep Q network (DeepQNetwork, DQN) is used to learn the optimal routing strategy.
[0108] The deep reinforcement learning routing strategy model in this application uses a double deep Q network (DoubleDeepQNetwork, DoubleDQN) architecture, and the specific implementation is as follows:
[0109] Network structure: The main network (Q-network) consists of 3 fully connected layers with 256, 128, and 64 neurons respectively, using ReLU as the activation function. The target network has the same structure as the main network and is used to stabilize the training process. The input layer receives the state vector and the output layer outputs the Q-values of each possible action.
[0110] In some implementations, different network architectures can be used to adapt to scheduling scenarios of different complexity. For example, for large-scale computing pools with a large state space, a deep network structure can be used, increasing the number of hidden layers to 5-6 layers. For scenarios with strong dynamics, recurrent neural network units such as LSTM or GRU can be introduced to capture the temporal dependencies of state sequences. For scenarios with spatial structure in state representation, convolutional neural networks can be used for feature extraction.
[0111] State representation: The state vector contains information such as computing node load indicators, network connection status, request queue information, and historical performance indicators.
[0112] To more effectively represent the high-dimensional state space, autoencoders or variational autoencoders can be used to reduce the dimensionality of the original state and extract key features. In addition, an attention mechanism can be introduced to dynamically adjust the importance weights of different state components based on the current task type.
[0113] Reward function design: The reward function considers multiple performance indicators, including request response time, resource utilization, load balancing, and energy cost, each weighted by a trade-off parameter.
[0114] For example, in different application scenarios, the weights of the reward function can be adjusted: for scenarios with high real-time requirements, the weight of the response time indicator can be increased; for resource-constrained environments, the weight of the resource utilization indicator can be increased; for high-concurrency systems, the weight of the load balancing indicator can be increased; for green computing requirements, the weight of the energy cost indicator can be increased.
[0115] Training process: Experience replay mechanism is used to maintain a fixed-size experience pool; every fixed number of steps, the main network parameters are copied to the target network; priority experience replay is used to preferentially sample high TD error experiences; ε-greedy strategy is used to balance exploration and exploitation.
[0116] Various improved training strategies can be used: for example, using the DuelingDQN architecture to separate state value and action advantage estimation; using distributed reinforcement learning methods to accelerate the training process; introducing human expert knowledge for imitation learning to initialize the policy network; using model prediction to assist training to improve sample efficiency.
[0117] In addition, a phased course learning can be set up to gradually transition from simple scenarios to complex scenarios, improving training efficiency and final performance.
[0118] In the edge-cloud collaborative computing environment scenario, the specific application of the deep reinforcement learning routing strategy model is as follows: for real-time AI inference tasks with strict delay requirements (such as real-time video analysis), the system increases the weights of the network state score and the task affinity score, and preferentially selects edge nodes with low delay connection and dedicated hardware accelerators;
[0119] For large-scale training tasks that are computationally intensive but not sensitive to delay, the system increases the weights of the computing power load score and the resource efficiency score, and preferentially selects cloud nodes with high-performance computing resources and low current load;
[0120] In the scenario of network condition fluctuation, the system can automatically adjust the weight of the network state score, dynamically balancing the influence of network delay and node computing capacity.
[0121] Through continuous learning and optimization, the system can continuously adjust the routing strategy as the environment changes and the load pattern evolves, improving the overall service quality and resource utilization efficiency.
[0122] Step 3, according to the service interaction mode representation and the score model result, analyze the task characteristics and resource state, dynamically adjust the microservice granularity, combine the computationally intensive tasks, and split the tasks with high parallelism;
[0123] Specifically, the following steps are included:
[0124] Step 3.1, analyze the service computing characteristics and resource requirements;
[0125] First, analyze the computing characteristics of the microservice components, including computational intensity analysis, data dependency analysis, and parallelism analysis.
[0126] Computational intensity analysis evaluates the ratio of computation to communication of the service, represented by the ratio of computation to communication, and services with high ratio are suitable for merging to reduce communication overhead;
[0127] Data dependency analysis evaluates the degree of data sharing between services, represented by the ratio of shared data to total data, and services with strong dependency are suitable for deployment on the same node;
[0128] Parallelism analysis evaluates the parallel execution potential of the service, represented by the ratio of sequential execution time to parallel execution time, and services with high parallelism are suitable for splitting to improve resource utilization.
[0129] The computing characteristic analysis process needs to balance accuracy and real-time performance. The present application adopts a method combining online lightweight analysis and offline deep analysis: online analysis is based on real-time monitoring data, providing fast but rough characteristic evaluation; offline analysis is based on historical execution records, providing more detailed and accurate characteristic modeling.
[0130] In addition, the present application also considers the following factors for more comprehensive service characteristic analysis:
[0131] Resource demand stability: evaluate the fluctuation degree of service resource demand, and services with large fluctuations are suitable for adopting more flexible granularity strategies to adapt to different load situations;
[0132] Performance sensitivity: analyze the sensitivity of the service to different resource types (such as CPU, memory, I / O, etc.), providing a basis for subsequent resource allocation;
[0133] State management complexity: evaluate the complexity of service state management, and services with complex states need to consider state synchronization overhead when splitting;
[0134] According to an optional embodiment of the present application, a machine learning method can be used to automatically extract and classify service features. Specifically, by monitoring the resource consumption mode, call relationship and performance indicators during service runtime, a classification model is trained to classify services into different categories such as compute-intensive, communication-intensive and memory-intensive, providing more accurate decision basis for granularity optimization.
[0135] Step 3.2, determine the optimal service granularity and boundary;
[0136] Based on the computing characteristic analysis results, a graph partitioning-based method is used to determine the optimal service granularity and boundary.
[0137] Construct a service dependency graph, where nodes represent functional modules and edges represent inter-module dependencies, and edge weights represent dependency strength.
[0138] Define a graph partitioning objective function that minimizes inter-partition communication cost and balances intra-partition computing load.
[0139] Use the normalized cut algorithm to solve the graph partitioning problem and obtain the optimal service boundary.
[0140] The graph partitioning algorithm used in the present application considers the following objectives:
[0141] Minimize cross-partition communication overhead;
[0142] Balance the computing load of each partition;
[0143] Limit the resource demand of a single partition to not exceed the node capacity;
[0144] Consider the affinity and mutual exclusion constraints between services.
[0145] When solving the graph partitioning problem, a multi-level partitioning strategy is adopted to improve algorithm efficiency:
[0146] Coarsen the original dependency graph and merge closely connected nodes;
[0147] Perform initial partitioning on the coarsened graph;
[0148] Restore the details of the original graph step by step through the refinement process, while optimizing the partitioning results.
[0149] In addition, the application also considers the cost factor of service granularity adjustment. For deployed services, re-adjusting the granularity will bring migration cost and temporary service interruption.
[0150] Therefore, only when the performance improvement brought by granularity adjustment significantly exceeds the adjustment cost, the system will perform actual granularity adjustment operation. Specifically, the system evaluates the adjustment benefit through the following formula:
[0151] Gain service =P imp -C adj ;
[0152] Where Gain service represents the net benefit of service granularity adjustment; P imp represents the performance improvement benefit after adjustment; C adj represents the cost generated by the adjustment process.
[0153] Where performance improvement is estimated by a prediction model, and adjustment cost includes migration time, resource consumption and service interruption impact, etc.
[0154] The application also provides a service granularity optimization method based on reinforcement learning. This method models the service granularity adjustment process as a Markov decision process, and continuously learns the optimal granularity adjustment strategy by interacting with the environment.
[0155] The state includes the current service distribution and system load, the action includes service merging, splitting and migration operation, and the reward is defined based on system performance indicators. This method can adapt to dynamically changing workloads and automatically explore better service granularity configuration.
[0156] Step 3.3, dynamically perform service reorganization and resource allocation;
[0157] According to the determined optimal service boundary, dynamically perform service reorganization and resource allocation:
[0158] For compute-intensive service clusters, perform service merging to reduce cross-service communication overhead;
[0159] For service clusters with high parallelism, service splitting is performed to improve resource utilization;
[0160] According to the service reorganization result, the resource allocation strategy is adjusted to ensure that each service obtains an appropriate amount of resources.
[0161] The service reorganization process adopts a gradual migration strategy to ensure that the system can still provide normal services during the adjustment process. Specifically, first, service instances are created according to the new boundaries, then the traffic is gradually migrated from the old instances to the new instances, and finally, after confirming that the new instances are running stably, the old instance resources are released.
[0162] The resource allocation process in this application fully considers the service characteristics and mutual influence, mainly including the following steps:
[0163] Preliminary resource demand estimation: based on the service historical resource usage mode and the current workload, the basic resource demand of each service is estimated;
[0164] Resource interference analysis: considering the possible resource competition and performance interference between services sharing the same physical node, the resource allocation is adjusted to reduce the negative impact;
[0165] Priority-aware allocation: according to the service priority difference, ensure that high-priority services get enough resources, while avoiding resource starvation of low-priority services;
[0166] Resource elasticity configuration: reserve appropriate resource elasticity space for services with large fluctuations to cope with sudden loads;
[0167] In some embodiments, resource oversubscription technology can be used to improve resource utilization. The system monitors the actual resource usage of each service, and when the peak resource demand of multiple services occurs at different times, it can moderately over-allocate resources to improve overall utilization efficiency. At the same time, through real-time monitoring and rapid response mechanism, when resource competition occurs, it is adjusted in time to ensure the stability of service performance.
[0168] In addition, the application also provides a resource pre-allocation method based on historical patterns. The system analyzes the historical patterns of service resource usage, identifies periodic changes and trends, and reserves resources in advance for services that will soon need more resources, reducing the impact of resource allocation delay on performance. This method is particularly suitable for workloads with obvious time patterns, such as daily business peaks, weekend traffic changes, etc.
[0169] Step 4, based on the adjusted microservice structure, use graph reinforcement learning method to predict service interaction mode and resource demand change, deploy and expand key services in advance, and implement dynamic resource isolation strategy;
[0170] Specifically, the following steps are included:
[0171] Step 4.1, constructing a service interaction prediction model;
[0172] Based on the multi-level graph convolutional network model constructed in step 1, the time series feature learning ability is increased, and the service interaction prediction model is constructed.
[0173] The service interaction graph sequence is represented as a time series graph, where each graph represents the service interaction state at a time point.
[0174] The spatio-temporal graph convolutional network is applied to learn the time series interaction mode, which combines the feature extraction capabilities of time and space dimensions.
[0175] Based on the learned spatio-temporal features, the service interaction mode at the future time is predicted.
[0176] The time series graph sequence construction process uses a sliding window mechanism to sample the service interaction data in the past period (such as the last 30 minutes), forming a graph sequence with time continuity. Each time point of the service interaction graph contains node features (such as service load, response time, etc.) and edge features (such as call frequency, data transmission volume, etc.), which dynamically change over time and reflect the evolution of system running state.
[0177] The spatio-temporal graph convolutional network in this application is an innovative model that combines time convolutional network and graph convolutional network, and its main components include:
[0178] Spatio-temporal convolution block: the basic building unit of the network, composed of three layers of "time convolution-graph convolution-time convolution", responsible for capturing time series patterns, graph structure features and their combined relationships. Multiple spatio-temporal convolution blocks are stacked to form a complete network, extracting higher-level spatio-temporal features layer by layer.
[0179] Time convolution layer: a one-dimensional convolution structure with a gating mechanism, effectively capturing time series dependencies. The gating mechanism is realized through two parallel convolution branches, one branch extracts features and the other branch controls information flow, similar to the gating unit in long short-term memory network, but with higher computational efficiency.
[0180] Graph convolution layer: a specialized convolution operation for processing graph structure data, aggregating information of nodes and their neighbors through message passing mechanism, learning the context representation of nodes. Unlike traditional convolution, graph convolution can handle irregularly structured data and is suitable for service call networks with such topologies.
[0181] Prediction layer: receives the final spatio-temporal feature representation and outputs the service interaction prediction result at the future time. Depending on the prediction task, it can be a regression layer (predicting continuous values such as load level) or a classification layer (predicting discrete states such as overload risk).
[0182] The model structure can be adjusted in different application scenarios:
[0183] For scenarios that require capturing long-term dependencies, dilated convolution can be used to expand the receptive field in the time dimension;
[0184] For complex systems with high feature dimension, a dimension reduction layer can be added before the spatio-temporal convolution block to reduce computational complexity;
[0185] For anomaly detection requirements, a reconstruction loss can be added to enable the model to have both prediction and anomaly detection capabilities;
[0186] In addition, the training of the prediction model uses a combination of online learning and transfer learning to enable the system to adapt to changing workload characteristics. Online learning updates model parameters incrementally, incorporating newly collected data into the training process; transfer learning allows knowledge learned in one environment to be transferred to a new environment, speeding up model adaptation to new scenarios.
[0187] Step 4.2, predictive service deployment and scaling;
[0188] Based on the predicted service interaction results, predictive service deployment and scaling are achieved. First, analyze the predicted service interaction graph to identify potential hot services and bottleneck nodes in the future.
[0189] Specifically, calculate the heat of each service node, which is determined by the sum of all interaction intensities received by the node.
[0190] Scale up services with heat exceeding the preset threshold in advance, pre-allocate computing resources to avoid performance bottlenecks during peak load periods.
[0191] Pre-allocate additional resources to hot services, with the allocation amount determined dynamically based on the deviation of service heat from the overall distribution.
[0192] According to an embodiment of the present application, predictive service deployment uses a multi-stage strategy:
[0193] Short-term prediction and immediate response: based on the prediction results of recent data (such as the past few minutes), quickly respond to upcoming load changes and adjust the resource allocation of existing service instances;
[0194] Medium-term prediction and scaling preparation: based on the prediction results of a longer time window (such as several hours), start preparing new service instances in advance, including image pulling, environment initialization, etc., to reduce scaling delay;
[0195] Long-term prediction and capacity planning: long-term trends (e.g., days or weeks) based on historical data analysis to guide system capacity planning and resource reservation strategies;
[0196] In addition, this application also considers the impact of prediction uncertainty on decision-making. The service interaction prediction model output not only includes point prediction (such as expected load value), but also includes the confidence interval of the prediction.
[0197] The system adjusts the response strategy according to the degree of uncertainty in the prediction:
[0198] For high-confidence predictions, take more aggressive pre-adjustment measures;
[0199] For low-confidence predictions, maintain more resource elasticity to cope with possible prediction deviations.
[0200] Step 4.3, implement dynamic resource isolation strategy;
[0201] Based on service priority and predicted resource contention, implement dynamic resource isolation strategy.
[0202] Calculate the resource contention degree between services, which is achieved by analyzing the usage patterns and competition degree of different resource types (such as CPU, memory, I / O, etc.) of each service.
[0203] According to the service priority and the resource contention degree, determine the isolation level. When the resource contention degree between high-priority services and low-priority services is high, a more stringent isolation mechanism is adopted.
[0204] Apply the corresponding resource isolation mechanism, including but not limited to CPU core binding, memory bandwidth limitation, I / O priority adjustment, and cache partitioning, etc.
[0205] It should be understood that the resource isolation level is closely related to the isolation effect. This application defines multiple isolation levels, from low to high:
[0206] Priority scheduling isolation: provide basic isolation through scheduler priority mechanism, achieve small overhead but limited isolation effect;
[0207] Resource quota isolation: set resource usage upper limit through container technology, prevent overuse while maintaining resource sharing;
[0208] Resource reservation isolation: reserve dedicated resources for high-priority services to ensure resource availability but may reduce overall utilization;
[0209] Physical isolation: deploy different services on physically isolated nodes to provide the strongest isolation effect but lower flexibility;
[0210] The system automatically selects appropriate isolation levels according to resource contention degree and service priority difference. For example, when there is moderate resource contention between two services with similar priorities, resource quota isolation may be selected; when a high-priority critical service coexists with multiple low-priority services and resource contention is severe, physical isolation may be selected.
[0211] In addition, the present application also proposes a dynamic fine-grained resource isolation method. This method not only considers service-level isolation, but also goes deep into the internal component level of microservices, identifies core components on the critical path, and provides stronger resource guarantees for these components. This fine-grained isolation method can maintain high resource utilization while ensuring the performance stability of critical functions.
[0212] In some embodiments, the isolation strategy can work in conjunction with the service granularity optimization in step 3. When it is found that there is severe resource competition between two services and it is difficult to solve by general isolation means, the system can consider adjusting the service granularity to separate the components of the competing resources into different services, and fundamentally solve the resource competition problem.
[0213] Resource isolation decisions are not static, but are continuously adjusted as system load and service characteristics change. The present application uses a closed-loop control mechanism to continuously monitor the effectiveness of isolation measures and automatically adjust isolation strategies when actual performance deviates from expected targets. This adaptive method ensures that the system can achieve a dynamic balance between resource utilization and service quality.
[0214] Step 5, based on the prediction results and resource isolation strategy, use multi-objective optimization algorithm to find the optimal service deployment scheme balancing delay, throughput and resource efficiency, and continuously optimize the orchestration strategy through feedback adjustment mechanism.
[0215] Specifically includes the following steps:
[0216] Step 5.1, define a multi-objective optimization problem;
[0217] Formalize the service orchestration problem as a multi-objective optimization problem, considering multiple usually conflicting optimization objectives. Optimization objectives include but are not limited to average response time, resource utilization (converted to negative to minimize the problem), energy consumption, and service waiting time variance (load balancing degree). The solution space of the problem is all feasible service deployment and scheduling schemes, subject to resource capacity, service dependency, and other constraint conditions.
[0218] Unlike traditional single-objective optimization, multi-objective optimization usually has no unique "optimal solution", but a set of Pareto-optimal solutions, which form a Pareto front in the objective space, any improvement in one objective will inevitably lead to the performance degradation of at least one other objective. Therefore, the core task of multi-objective optimization is to find a set of solutions on the Pareto front, providing system operators with different performance trade-off schemes.
[0219] The present application uses vector representation of multi-objective optimization problem, each solution corresponds to a target vector, and each component of the target vector represents the value of an optimization objective.
[0220] The target vectors are compared by the Pareto dominance relationship: if solution A is not worse than solution B in all objectives, and at least one objective is better than solution B, then solution A dominates solution B. The Pareto-optimal solution set consists of all solutions that are not dominated by any other solution.
[0221] According to an optional embodiment of the present application, multi-objective optimization can be converted into a scalar problem by introducing preference information, for example, using linear weighting method, combining multiple objective functions into a single objective function through weight parameters. This method is convenient for solving but needs to determine the weight of each objective in advance, and is sensitive to weight setting. Another optional method is to use constraint method, taking one objective as the optimization objective and setting other objectives as constraint conditions, which can optimize other indicators under the premise of guaranteeing certain key performance indicators.
[0222] Step 5.2, solve the scheduling optimization problem by applying a multi-objective evolutionary algorithm;
[0223] The non-dominated sorting genetic algorithm II (NSGA-II) is used to solve the multi-objective scheduling optimization problem. NSGA-II is an efficient multi-objective evolutionary algorithm that can obtain multiple Pareto-optimal solutions in one run. The main process of the algorithm includes:
[0224] Solution coding and initialization: encode the service deployment and resource allocation scheme into a chromosome to generate an initial population. The chromosome contains two parts of information: service deployment location and resource allocation, using hybrid coding method, integer part represents deployment location, and real part represents resource allocation proportion.
[0225] Non-dominated sorting: based on the Pareto dominance relationship, the solutions in the population are divided into different levels of non-dominated front. The first front contains the non-dominated solutions in the population, after removing these solutions, find the non-dominated solutions in the remaining solutions to form the second front, and so on.
[0226] Crowding degree calculation: within the same non-dominated front, the crowding degree is calculated according to the distribution density of the solution in the objective space, the solution with higher crowding degree is preferentially retained to maintain the diversity of the solution set.
[0227] Selection, crossover and mutation: Selection is based on non-dominated rank and crowding distance, and new generation is generated by crossover and mutation operation. Crossover operation exchanges fragments of parent chromosomes to generate offspring, and mutation operation randomly changes some genes in the chromosome to increase population diversity.
[0228] Elite strategy and iterative optimization: Elite strategy is adopted to retain the best non-dominated solution in each generation. After multiple iterations, the population gradually converges to the Pareto front, and a set of approximate Pareto optimal solutions is finally output.
[0229] The following improvements are made based on the standard NSGA-II algorithm:
[0230] Constraint handling mechanism: A dedicated constraint handling method is designed, including resource capacity constraints, service affinity constraints and service dependency constraints. The degree of constraint violation is converted into a penalty term of the objective function through a penalty function, so that the algorithm can effectively handle complex constraint conditions.
[0231] Adaptive evolutionary operation: According to the population diversity and optimization process, the crossover probability and mutation probability are dynamically adjusted to maintain a high exploration ability in the early search stage and enhance the fine search of local areas in the later search stage.
[0232] Hybrid search strategy: Local search method is combined to locally optimize the solutions on the Pareto front and accelerate convergence. The local search method adaptively selects different search strategies for different types of decision variables based on the characteristics of the solution.
[0233] Problem-specific initialization: Partial high-quality initial solutions are generated using professional knowledge, such as heuristic solutions based on historical deployment schemes and resource utilization patterns, to improve the starting point of the algorithm and accelerate the convergence process.
[0234] After obtaining the Pareto optimal solution set, a balanced solution needs to be selected as the actual deployment scheme. This application adopts a decision-making method based on fuzzy theory:
[0235] Define a membership function for each objective to represent the degree of solution satisfaction for that objective;
[0236] Calculate the total membership of each solution and select the solution with the highest total membership as the final scheme.
[0237] This method can find a balanced solution for each objective without introducing explicit weights.
[0238] Step 5.3, implement the feedback adjustment mechanism;
[0239] Construct a feedback adjustment mechanism to continuously optimize the scheduling strategy based on actual operation data. The feedback adjustment mechanism follows the principles of cybernetics and forms a closed-loop control through the three links of measurement, comparison and adjustment:
[0240] Performance Monitoring and Data Collection: Continuously monitor system operational status, collect key performance indicators, including response time distribution, resource utilization, energy consumption, etc. Data collection process focuses on the balance between sampling frequency and data representativeness, capturing system dynamic changes while avoiding excessive monitoring overhead.
[0241] Deviation Analysis and Root Cause Diagnosis: Compare actual performance with expected targets, calculate deviations. Deviation analysis not only focuses on deviation size, but also considers deviation trend and volatility. In addition, through correlation analysis and causal inference, diagnose the root cause of performance deviation, provide basis for subsequent adjustment.
[0242] Parameter Adjustment and Strategy Optimization: According to the results of deviation analysis, adaptively adjust system parameters and optimize strategies. Adjustment content includes:
[0243] Update the weights of each optimization target, increase the weights of large deviation targets;
[0244] Adjust the constraints, relax the constraints that cause performance bottlenecks;
[0245] Correct the prediction model, update the service interaction prediction model based on new data;
[0246] Optimize the arrangement decision, regenerate the deployment and scheduling scheme more suitable for the current state;
[0247] Gradual deployment and effect verification: Adopt gradual strategy to deploy the adjusted scheme, first verify the effect in small-scale environment, confirm the benefit before expanding the application range. This cautious deployment strategy reduces the risk of adjustment, ensures system stability.
[0248] Feedback adjustment should not be too frequent to avoid system shock.
[0249] This application adopts a gradual adjustment strategy with threshold:
[0250] Only when the performance deviation exceeds the preset threshold and lasts for a certain time, the adjustment will be triggered;
[0251] The adjustment amplitude is proportional to the deviation, but there is an upper limit to prevent excessive adjustment;
[0252] After adjustment, set a cooling period, during which no new adjustment is triggered, giving the system enough time to stabilize to the new state.
[0253] In addition, this application also constructs a multi-time scale feedback adjustment framework, which adopts different frequency adjustment strategies for different types of deviations:
[0254] For short-term fluctuations (such as temporary delay increase caused by sudden traffic), adopt fast response mechanism, relieve through temporary resource reallocation, etc.
[0255] For medium-term trends (such as changes in load patterns within a day), adapt by periodically re-optimizing the scheduling scheme;
[0256] For long-term evolution (such as system scale expansion or workload characteristics change), make fundamental adjustments through model retraining and optimization strategy updates.
[0257] This multi-time scale feedback framework enables the system to cope with changes of different cycles simultaneously, maintaining stability while having sufficient adaptability.
[0258] Application examples of this embodiment:
[0259] To verify the actual effect of the deep learning-based computing power pool intelligent integration and elastic scheduling method proposed in this application, the method is applied to the following practical scenarios for testing and evaluation.
[0260] Application scenarios:
[0261] The method proposed in this application has been deployed and tested in a distributed deep learning training environment of a large cloud computing platform. This environment has the following characteristics:
[0262] Computing resource pool: a heterogeneous computing cluster composed of 150 computing nodes, including standard CPU servers, GPU-accelerated servers (equipped with NVIDIA A100 GPUs), and dedicated AI acceleration card servers;
[0263] Workload characteristics: multiple deep learning training tasks are running simultaneously, including computer vision model training (such as ResNet, YOLO, etc.), natural language processing model training (such as Transformer, BERT, etc.), and recommendation system model training, which have different resource demand characteristics and priorities;
[0264] Business requirements: need to ensure the performance stability of high-priority tasks under limited resources, while improving overall resource utilization, reducing energy consumption and operating costs;
[0265] The main challenges faced by this application scenario include: large workload fluctuations, diverse resource demands, significant differences in task priorities, and severe resource competition. Traditional scheduling methods perform poorly in this scenario, with the main problems being low resource utilization (average only 35%), large fluctuations in high-priority task performance (maximum fluctuation up to 80%), and poor adaptability to sudden loads (expansion time up to several minutes).
[0266] Implementation process instance:
[0267] Multi-level graph convolutional network deployment and training:
[0268] In actual deployment, a three-layer graph convolutional network model is constructed, with the following specific configurations:
[0269] Function-level graph convolutional network: contains 2 layers of graph convolutional layers, each with 128 and 64 neurons respectively, capturing the call relationship and dependency within the operator level of the deep learning framework;
[0270] Service-level graph convolutional network: contains 3 layers of graph convolutional layers, each with 256, 192 and 128 neurons respectively, modeling the interaction relationship between different training tasks;
[0271] Cluster-level graph convolutional network: contains 2 layers of graph convolutional layers, each with 384 and 256 neurons respectively, analyzing the resource competition relationship between different resource domains;
[0272] The model training adopts a two-stage method: first, offline pre-training is performed using historical running data, and then the model parameters are continuously fine-tuned through online learning during system operation. The pre-training data set contains 30 days of system operation records, about 5 million service interaction data and 10 million resource usage records.
[0273] In order to adapt to different scale deployment environments, lightweight variants of the model are also defined, which reduce the computational complexity by reducing the number of graph convolutional layers and neurons, suitable for resource-constrained scenarios such as edge computing.
[0274] Implementation of power-aware routing decision:
[0275] In actual application, the power-aware routing decision module collects power state information through the sidecar proxy of the service mesh, including:
[0276] CPU utilization (core level), memory usage, GPU computing unit usage, GPU memory occupancy, I / O bandwidth usage, and energy consumption data of each computing node;
[0277] Network state information, including inter-node link delay, bandwidth occupancy, and packet loss rate;
[0278] Historical performance indicators, such as average execution time and resource consumption patterns of various tasks on different nodes;
[0279] The system defines dedicated routing strategies for different types of deep learning tasks:
[0280] For large model training tasks, prefer GPU memory capacity and high-bandwidth connections between nodes;
[0281] For small batch inference tasks, prefer low latency and stability;
[0282] For distributed training tasks, inter-node communication efficiency and resource consistency are prioritized.
[0283] The routing decision model is continuously optimized during system operation, learning the optimal dimension weight distribution strategy through deep reinforcement learning methods. In a 60-day test cycle, the model completed approximately 2 million decision optimization iterations, and the routing accuracy (defined as the proportion of selecting the optimal execution path) improved from 65% to 92%.
[0284] Adaptive microservice granularity optimization example:
[0285] In practical applications, the system performs service granularity dynamic adjustment for typical deep learning training tasks. Here is a specific case:
[0286] For a large Transformer model training task, the system initially divides it into data preprocessing, embedding layer calculation, multi-head attention calculation, feedforward network calculation, and backpropagation services.
[0287] Through characteristic analysis, the system identifies that there is frequent data exchange and close computational dependence between multi-head attention calculation and feedforward network calculation, and both components are computationally intensive.
[0288] Based on the results of the graph partitioning algorithm, the system merges these two components into one service, reducing cross-service communication overhead.
[0289] At the same time, for the data preprocessing service, the system finds that its internal parallelism is high and resource demand fluctuates greatly, so it further splits it into data loading, data transformation, and data batch processing three independent services, achieving more flexible resource allocation and load balancing.
[0290] Through such adaptive granularity adjustment, the end-to-end execution time of this training task is reduced by 27%, and the resource utilization is improved by 42%. The system dynamically adjusts the service granularity according to different stages of the training process (such as the warm-up period, stable training period, and convergence period), further optimizing performance.
[0291] Predictive service orchestration and resource isolation application:
[0292] In actual deployment, the system uses spatiotemporal graph convolutional networks to predict future workload changes. Here is a specific case of predictive orchestration:
[0293] The system found through analyzing historical data that 10:00-12:00 and 15:00-17:00 are peak periods for computer vision model training tasks. Based on this pattern, the system automatically starts resource preparation 30 minutes before the peak, including preheating GPUs, preloading common deep learning frameworks, and caching datasets. At the same time, the system predicts that there is significant resource contention between recommendation system model training and computer vision model training, especially in terms of GPU memory and PCIe bandwidth.
[0294] Based on the prediction results and task priorities, the system implements the following resource isolation strategies:
[0295] For high-priority computer vision model training tasks, allocate dedicated GPU cards and reserve a fixed proportion of memory and I / O bandwidth;
[0296] For medium-priority natural language processing model training tasks, use GPU time slice isolation and dynamic resource quotas;
[0297] For low-priority batch tasks, use opportunity scheduling strategy, only allocate execution resources when resources are sufficient;
[0298] Through this prediction-based proactive orchestration and multi-level isolation strategy, the performance fluctuation of high-priority tasks is reduced from 45% to 9%, while the overall resource utilization of the system remains at a high level (average 78%).
[0299] Multi-objective optimization and feedback adjustment implementation:
[0300] In practical applications, the system deploys a multi-objective optimization framework for distributed training environments, considering the following four optimization objectives:
[0301] Minimize training task completion time;
[0302] Maximize resource utilization;
[0303] Minimize energy consumption;
[0304] Maximize system throughput (number of training batches completed per hour);
[0305] An improved NSGA-II algorithm is used to solve the multi-objective optimization problem, with a population size of 200, 100 evolution generations, a crossover probability of 0.8, and a mutation probability of 0.1.
[0306] The optimization algorithm runs every 30 minutes to generate a new service deployment plan. To maintain system stability, an incremental deployment strategy is used, adjusting up to 20% of service instances each time.
[0307] The feedback adjustment mechanism implements three time-scale optimization closed loops:
[0308] Short-term adjustment (1-minute cycle): Mainly for sudden load and urgent resource competition, quickly respond through temporary resource reallocation and task priority adjustment;
[0309] Medium-term adjustment (30-minute cycle): Based on the latest collected performance data and load trends, re-optimize service deployment scheme;
[0310] Long-term adjustment (once a week): Comprehensive update of prediction models and optimization strategies to adapt to long-term workload evolution;
[0311] Through this multi-time scale feedback adjustment, the system can effectively adapt to environmental changes while maintaining stability, continuously optimizing performance.
[0312] Technical effect verification:
[0313] The method proposed in this application has achieved significant technical effects in practical application. The following are performance data collected in a 90-day actual operation, focusing on verifying two key technical effects: resource utilization efficiency improvement and service stability enhancement.
[0314] Resource utilization efficiency improvement:
[0315] This method significantly improves the resource utilization efficiency of the computing power pool through computing power-aware routing and adaptive service granularity optimization. Statistical data shows:
[0316] CPU resource utilization rate increased from an average of 35% to 82%, with a peak of 92%;
[0317] GPU computing unit utilization rate increased from an average of 42% to 87%, with a peak of 95%;
[0318] Memory utilization rate increased from an average of 45% to 76%;
[0319] Storage I / O bandwidth utilization rate increased from an average of 30% to 68%;
[0320] Compared with traditional fixed-grain microservice architecture, under the same workload, this method reduces the number of computing nodes required by 48%, while supporting an increase in concurrent task number by 125%. Especially in processing computationally intensive deep learning tasks, resource utilization efficiency is significantly improved, saving enterprises about 42% of infrastructure costs.
[0321] Through adaptive microservice granularity optimization, the system significantly reduces resource fragmentation.
[0322] Before using this method, there were 28% of fragmented resources (scattered resources that cannot be allocated to any task) on average per node;
[0323] After using this method, the resource fragmentation rate is reduced to 9%, and the effective capacity of the resource pool is improved.
[0324] Service stability is enhanced:
[0325] This method significantly enhances service stability through predictive service orchestration and dynamic resource isolation, especially the performance consistency of high-priority tasks. Statistics show:
[0326] The performance fluctuation of high-priority tasks (expressed as the standard deviation of completion time divided by the average value) is reduced from 45% to 9%;
[0327] The predictability of task execution time (the degree of agreement between predicted and actual values) is improved from 62% to 93%;
[0328] The service level agreement achievement rate is improved from 78% to 99.5%;
[0329] The number of performance interference events (significant performance decline due to resource competition) is reduced by 87%;
[0330] In handling sudden loads, this method performs particularly outstandingly. Under the traditional method, the average time for the system to respond to sudden loads is 8 minutes, and it often leads to a significant decline in the performance of running tasks;
[0331] After using this method, the system response time is shortened to 45 seconds (increased by 91%), and the performance of high-priority tasks is almost unaffected.
[0332] Through multi-level resource isolation strategies, the system can still maintain differentiated service quality in the case of intense resource competition.
[0333] In the simulated resource competition scenario, the performance difference between high-priority tasks and low-priority tasks under the traditional method is only 15% (there should be a significant difference in ideal conditions);
[0334] After using this method, the system can maintain a performance difference of 68%, ensuring that resource allocation meets business priority requirements.
[0335] These technical effect verification results fully prove the significant advantages of the proposed deep learning-based computing power pool intelligent integration and elastic scheduling method in practical applications, especially in resource utilization efficiency and service stability.
[0336] As shown in Figures 2 to 4 , respectively, are the comparison of resource utilization rate improvement; service stability index changes; and multi-objective optimization trade-off effects.
[0337] The above describes the embodiments of the present application, but the embodiments are not limited to the above specific embodiments, and the above specific embodiments are only illustrative but not restrictive, and the ordinary skilled in the art can make more equivalent embodiments under the inspiration of the embodiments, which are all within the protection scope of the embodiments.
Claims
1. A method for intelligent integration and flexible scheduling of computing power pools based on deep learning, characterized in that: The following steps are involved: Construct a multi-layer graph convolutional network to model the dynamic interaction relationship between microservices. The multi-layer graph convolutional network includes function-level, service-level, and cluster-level graph convolutional layers, and outputs a representation of the service interaction pattern. Build a multi-dimensional scoring model based on service interaction pattern representation, computing power status indicators, and network status indicators to select the optimal execution path for service requests; Based on the service interaction pattern representation and scoring model results, we analyze task characteristics and resource status, dynamically adjust the microservice granularity, merge computationally intensive tasks, and split tasks with high parallelism. Based on the adjusted microservice structure, we use graph reinforcement learning methods to predict service interaction patterns and changes in resource requirements, deploy and scale key services in advance, and implement dynamic resource isolation strategies. Based on the prediction results and resource isolation strategy, a multi-objective optimization algorithm is used to find the optimal service deployment solution that balances latency, throughput, and resource efficiency, and the orchestration strategy is continuously optimized through a feedback adjustment mechanism.
2. The method for intelligent integration and flexible scheduling of computing power pools based on deep learning according to claim 1 is characterized in that: The steps of constructing a multi-level graph convolutional network to model the dynamic interaction relationship between microservices include: Collect interaction data and resource usage between microservices from the service mesh layer to build an initial service interaction graph; Build a graph convolutional network model at three levels: function level, service level, and cluster level, to capture service interaction patterns at different granularities. The outputs of three levels of graph convolutional networks are integrated through the attention mechanism to generate a comprehensive service interaction representation.
3. The method for intelligent integration and flexible scheduling of computing power pools based on deep learning according to claim 1 is characterized in that: The steps of constructing a multi-dimensional scoring model based on the service interaction pattern representation and the computing power status indicators and network status indicators include: Collect node computing power status data through a lightweight monitoring agent, including CPU-related indicators, memory-related indicators, GPU-related indicators, and overall node indicators; Based on the collected computing power status data and network status data, a multi-dimensional scoring model is constructed to calculate the comprehensive score of service requests to each candidate node; A deep reinforcement learning method is used to dynamically optimize the dimension weights in the multi-dimensional scoring model, enabling the system to adapt to routing requirements in different scenarios.
4. The method for intelligent integration and flexible scheduling of computing power pools based on deep learning according to claim 1 is characterized in that: The steps of dynamically adjusting the microservice granularity include: Analyze the computational characteristics of microservice components, including computational intensity analysis, data dependency analysis, and parallelism analysis; Based on the results of computational characteristics analysis, the optimal service granularity and boundaries are determined using a graph partitioning-based approach; Based on the determined optimal service boundaries, service reconfiguration and resource allocation are performed dynamically.
5. The method for intelligent integration and flexible scheduling of computing power pools based on deep learning according to claim 4 is characterized in that: In the step of determining the optimal service granularity and boundary using the graph partitioning method, the graph partitioning objective function simultaneously considers the following objectives: Minimize inter-partition communication costs; Balance the computing load of each partition; Limit the resource requirements of a single partition to not exceed the node capacity; Consider affinity and mutual exclusion constraints between services.
6. The method for intelligent integration and flexible scheduling of computing power pools based on deep learning according to claim 1 is characterized in that: The step of using the graph reinforcement learning method to predict service interaction patterns and resource demand changes includes: Based on a multi-layer graph convolutional network model, we increase the ability to learn temporal features and build a service interaction prediction model. Based on the service interaction prediction results, predictive service deployment and scaling are achieved; Implement dynamic resource isolation policies based on service priorities and predicted resource contention.
7. The method for intelligent integration and flexible scheduling of computing power pools based on deep learning according to claim 6 is characterized in that: In the step of implementing the dynamic resource isolation strategy, the resource isolation strategy includes the following levels: Priority scheduling isolation, which provides basic isolation through the scheduler priority mechanism; Resource quota isolation, setting resource usage limits through container technology; Resource reservation isolation, reserving dedicated resources for high-priority services; Physical isolation: deploy different services on physically isolated nodes.
8. The method for intelligent integration and flexible scheduling of computing power pools based on deep learning according to claim 1 is characterized in that: The steps of using a multi-objective optimization algorithm to find an optimal service deployment solution that balances latency, throughput, and resource efficiency include: The service orchestration problem is formalized as a multi-objective optimization problem, considering multiple conflicting optimization objectives at the same time; A non-dominated sorting genetic algorithm is used to solve the multi-objective scheduling optimization problem; A balanced solution is selected from the obtained Pareto optimal solution set as the final arrangement plan.
9. The method for intelligent integration and flexible scheduling of computing power pools based on deep learning according to claim 1 is characterized in that: The steps of continuously optimizing the orchestration strategy through the feedback adjustment mechanism include: Continuously monitor system operating status and collect key performance indicators; Compare actual performance to expected targets and calculate deviations; Adaptively adjust system parameters and optimization strategies based on deviation analysis results; Use the updated model to re-solve the orchestration optimization problem and generate adjusted deployment and scheduling solutions.
10. A computer-readable storage medium, characterized in that It is used to store computer-readable instructions, which, when read by a computer, can run a deep learning-based computing power pool intelligent integration and flexible scheduling method as described in any one of claims 1-9.
Citation Information
Cited By
Multi-mode-based computing power scheduling method and device, server and storage medium
CN121116650A
Method, system and device for resource adaptation of an entertainment platform
CN122431908A