Artificial intelligence practical training resource elastic scheduling method for cloud edge hybrid deployment
By constructing a composite profile of tasks and nodes, resource requirements are accurately predicted and cloud-edge hybrid deployment is optimized. This solves the problems of low resource utilization and poor task continuity in AI training tasks in cloud-edge hybrid deployment, achieving efficient resource utilization and stable operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG PANLONG INFORMATION TECH CO LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies in cloud-edge hybrid deployments for AI training tasks suffer from low resource utilization, insufficient interactive response capabilities, and poor task continuity. In particular, the uneven demand for latency, computing power, storage, and network resources makes it impossible for single-edge-focused scheduling schemes to meet the requirements of continuous operation and overall elastic coordination.
By constructing composite task profiles and composite node profiles, resource requirements are accurately predicted and task urgency is calculated. Combined with multi-dimensional scoring, a comprehensive adaptation relationship is generated to optimize the cloud-edge hybrid deployment scheme. Through runtime monitoring, elastic scaling and fault migration mechanisms, the continuous and stable operation of training tasks is ensured.
It enables efficient resource utilization of AI training tasks under cloud-edge hybrid deployment, reduces node failure risk and service default probability, and improves task operation efficiency and resource utilization efficiency.
Smart Images

Figure CN122064461A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a method for elastic scheduling of artificial intelligence training resources deployed in a hybrid cloud-edge environment. Background Technology
[0002] As AI training platforms are increasingly used in university teaching, vocational training, and enterprise R&D training, these platforms typically need to continuously provide users with containerized development environments, CPU / GPU computing power, dataset access, model image loading, interactive sessions, and training task execution environments. Since training tasks include interactive coding, debugging, and inference verification, as well as longer training, fine-tuning, and batch processing jobs, their requirements for latency, computing power, storage, network, and operational continuity vary.
[0003] While traditional technologies rely solely on centralized cloud deployments for convenient centralized management and high-density computing power, they suffer from limitations in terms of interaction latency, data access based on proximity, and local continuity. Conversely, relying solely on edge node deployments is susceptible to factors such as small node resource scale, significant load fluctuations, and limited operational stability. Therefore, building a cloud-edge collaborative resource scheduling mechanism for AI training has become an important technological direction for improving resource utilization, interactive responsiveness, and task continuity.
[0004] Existing technologies have studied cloud-edge collaborative scheduling, task allocation, and elastic scaling from different perspectives. For example, Chinese patent CN113452566A discloses a cloud-edge-device collaborative resource management method, which obtains the status information of available cloud and edge resources and establishes an optimization model that considers reliability, latency, and resource capacity constraints to complete the matching of tasks and resource nodes; this solution demonstrates that existing technologies have already addressed the multiple resource constraints and reliability factors in cloud-edge collaborative scenarios. Chinese patent CN114301924A discloses an application task scheduling method for a cloud-edge collaborative environment, which selects target nodes based on latency-related parameters between tasks and candidate nodes, reflecting a low-latency task scheduling approach for cloud-edge collaboration. Chinese patent CN109117248A discloses a deep learning task elastic scaling system and method based on the Kubernetes platform, which improves deep learning training efficiency by collecting container resource utilization and increasing the number of containers according to preset rules, demonstrating that existing technologies have also addressed the elastic scaling problem of training tasks.
[0005] In terms of research papers, the Serverledge framework provides a function offloading and migration mechanism for the edge-cloud continuum, employing both horizontal and vertical function offloading. The LATM algorithm addresses the load distribution problem in edge networks, optimizing resource utilization through load-aware task migration. Thus, existing technologies offer feasible solutions for cloud-edge resource management, low-latency node selection, training task scaling, and task migration load balancing.
[0006] However, most of the aforementioned existing technologies focus on a single aspect: some emphasize node matching under latency constraints, some emphasize container replica scaling, and some emphasize the migration or unloading process itself. Based on the disclosed patents and papers, while existing solutions address resource management, node selection, elastic scaling, and migration issues, they are still insufficient for the continuous operation and overall elastic coordination of AI training tasks in a cloud-edge hybrid deployment. Summary of the Invention
[0007] In order to solve the above-mentioned technical problems, this application proposes the following technical solution: In a first aspect, embodiments of this application provide a method for elastic scheduling of artificial intelligence training resources in a hybrid cloud-edge deployment, including: Acquire task attribute data of the AI training tasks to be scheduled, as well as resource status data of cloud nodes and edge nodes, and construct composite task profiles and composite node profiles; Based on the aforementioned task composite profile and node composite profile, the resource requirements of each training task are predicted, and the overall urgency of the task is calculated. Based on the predicted demand results, urgency results, and overall node supply capacity, the data locality, resource matching degree, network advantage value, and node reliability score between each training task and each edge node and each cloud node are calculated, and the comprehensive adaptation relationship between task nodes is generated. Based on the comprehensive adaptation relationship and the comprehensive urgency of the tasks, the edge bearing weight, cloud-edge collaboration benefits and collaborative deployment utility of each training task are calculated, and the cloud-edge hybrid deployment scheme corresponding to each training task is determined under resource constraints. The system monitors the operational status of deployed training tasks, performs elastic scaling based on node pressure prediction, and triggers checkpoint migration and recovery when the source node experiences abnormal pressure, increased fault risk, or service default. Simultaneously, it updates scheduling parameters based on scheduling reports.
[0008] In one possible implementation, the step of acquiring task attribute data of the AI training task to be scheduled, as well as resource status data of cloud nodes and edge nodes, and constructing a composite task profile and a composite node profile includes: Collect resource request information, dataset information, latency constraint information, and interaction information of the AI training tasks to be scheduled, and simultaneously collect heterogeneous indicator information of each cloud node and edge node; The collected heterogeneous index information is normalized so that computing power indexes, link indexes and task indexes of different dimensions can enter the same computing space. After normalization, task composite profiles and node composite profiles are constructed respectively, and the comprehensive demand intensity of tasks, the volatility of task demand, and the comprehensive supply capacity of nodes are calculated to obtain the demand representation results on the task side and the supply representation results on the node side.
[0009] In one possible implementation, the construction of a task composite profile and a node composite profile, and the calculation of the overall task demand intensity, task demand volatility, and node overall supply capacity, to obtain task-side demand representation results and node-side supply representation results, include: For the Each training task constructs a composite profile vector. : in, , , , and They represent the first Each training task is at any time Requirements for CPU, GPU, memory, storage, and bandwidth. Indicates the scale of data associated with the task. This indicates the tolerable waiting time for the task. Indicates task priority. Indicates task latency sensitivity. Indicates the intensity of task interaction. Indicates the volatility of task requirements. This indicates the historical default risk value of the task; Regarding the The multidimensional requirements are coupled and modeled to obtain the comprehensive task requirement intensity. : in: Indicates the task is in the Normalized demand values in the resource category dimension Indicates the weight of a single-dimensional resource. This represents the coupling weight between different resource dimensions. , and These represent the weights for data size, latency sensitivity, and interaction intensity, respectively. The overall resource requirements of a task within a sliding window are statistically analyzed to obtain the task requirement fluctuation. : in: The length of the sliding window. This represents the average total resource requirements of the task within the window. To prevent tiny positive numbers with a denominator of zero; At the same time, for the first Each resource node builds a composite supply capacity. : in: For node number Normalized availability of class resources For node number Resource utilization rate For node health, To alleviate queuing pressure, For the risk of failure, To leverage cache hit advantage, , , , , , For the corresponding weights.
[0010] In one possible implementation, the prediction of resource requirements for each training task based on the task composite profile and node composite profile, and the calculation of the overall task urgency, includes: For the The task in the first Next-moment demand prediction in resource-like dimensions The following fusion prediction model was used to obtain the results: in: This is the historical smoothing coefficient. The coefficient for the trend term. For the coefficient of the periodic term, This is the error correction factor. This refers to the periodic demand component of the task in the corresponding resource dimension. This represents the current predicted residual; After obtaining the multidimensional resource forecasts, the forecast results of each dimension are integrated with the historical average demand, volatility, and historical default risk of the task to calculate the task's sudden risk coefficient. : in: This indicates the task is the first one in the history window. Average demand for such resources Indicates the degree of service breach in the past. Further calculate the overall urgency of the task. : in: to The fusion coefficient is... Let be the time normalization constant. This is a coupling term between latency sensitivity and interaction strength. The threshold for determining urgent tasks. This is an indicator function.
[0011] In one possible implementation, the process combines the predicted demand results, urgency results, and the overall supply capacity of nodes to calculate the data locality, resource matching degree, network advantage value, and node reliability score between each training task and each edge node and each cloud node, and generates a comprehensive adaptation relationship between task nodes, including: Calculate the task based on the degree of overlap between the data blocks required by the task and the data blocks already cached by the node. Relative to node Data locality : in: Indicates task Required data block set Represents a node The set of cached data blocks For data blocks Importance weight, To address the discrepancy between the cached data version on the node and the version expected by the task, This is the version deviation normalization constant; Then, based on the correspondence between the predicted task requirements and the available resources of the nodes, the resource matching degree is calculated. : in: For nodes In the Available capacity on class resources For the task In the Demand for similar resources To match the sensitivity index, The resource shortage penalty coefficient, To find the minimum value function, This is a function to find the maximum value. Further calculate the network dominance value : in: For task access side to node The time delay, For available bandwidth, As the reference bandwidth, For packet loss rate, For link jitter, to These are the corresponding weighting coefficients; At the same time, the reliability of the nodes themselves is scored, resulting in... : in: For node health, For the risk of failure, For queue occupancy rate, Mean time to recovery (MTBF) To recover the time normalization constant, These are the weighting coefficients; After completing the above scoring, it is fused with the projection similarity of the task profile and node profile to generate the task. With nodes Overall compatibility : in: and These are the projection matrices for the task profile and the node profile, respectively. The cosine similarity function is used. to This represents the corresponding fusion coefficient.
[0012] In one possible implementation, the step of calculating the edge bearing weight, cloud-edge collaboration benefits, and collaborative deployment utility of each training task based on the comprehensive adaptation relationship and the overall task urgency, and determining the cloud-edge hybrid deployment scheme corresponding to each training task under resource constraints, includes: Computational tasks Edge bearing weight : in: Indicates task Data locality on candidate edge sides To normalize GPU requirements, To normalize the data size, to This corresponds to the adjustment coefficient; Then, calculate the edge nodes. With cloud nodes Synergistic benefits between : in: For cloud-edge interconnect latency, For cloud-edge bandwidth, This represents the maximum observation bandwidth of the cloud-edge link. For cloud-edge link packet loss rate, and Leveraging the caching advantages of both edge and cloud sides, The cost of cloud-edge state synchronization These are the corresponding weight coefficients; Based on edge carrying weight, comprehensive adaptability of tasks and nodes, and cloud-edge collaboration benefits, combined with resource occupation cost, cold start cost, energy consumption cost, and operational risk cost, the collaborative deployment utility of each training task is calculated. Under resource capacity constraints and task unique deployment constraints, optimization solutions are performed to determine the cloud-edge hybrid deployment results of each training task.
[0013] In one possible implementation, the collaborative deployment utility of each training task is calculated based on edge bearer weight, the comprehensive adaptability of tasks and nodes, and cloud-edge collaboration benefits, combined with resource consumption costs, cold start costs, energy consumption costs, and operational risk costs. This is then optimized under resource capacity constraints and unique task deployment constraints to determine the cloud-edge hybrid deployment result for each training task, including: For the task Select edge nodes and cloud nodes Hybrid deployment schemes calculate the collaborative deployment utility : in: and These refer to the overall compatibility between tasks and edge nodes, and between tasks and cloud nodes. , , , These represent the costs of resource consumption, cold start, energy consumption, and operational risks, respectively. This corresponds to the penalty coefficient; Resource consumption cost and cold start cost The preferred method is to calculate as follows: in: This refers to the size of the task image. The available bandwidth from the source location to the target deployment location. , , This corresponds to the penalty coefficient; Under the premise of satisfying resource capacity and scheduling uniqueness constraints, the deployment decision variables are determined by the following objective function. : in: Represents the set of edge nodes. Represents a set of nodes in the cloud. This represents the node pressure value. For global average pressure, This is the load balancing penalty coefficient.
[0014] In one possible implementation, the step of monitoring the deployed training tasks in runtime, performing elastic scaling based on node pressure prediction, and triggering checkpoint migration and recovery when the source node experiences abnormal pressure, increased failure risk, or service default, while simultaneously updating scheduling parameters based on scheduling reports, includes: After the mixed deployment of the practical training tasks was completed, the first... Calculate the overall pressure value at each node. : in: , , , These are CPU, GPU, memory, and bandwidth utilization, respectively. Queue length The maximum allowed queue length, The average response time of the node. The rate of change of pressure, These are the corresponding weighting coefficients; Based on current and predicted pressure, nodes are elastically scaled up or down, with instance adjustments made. Obtained by the following formula: in: To predict the pressure in the next moment, The target pressure threshold, For the historical cumulative window, , , These represent the proportional, integral, and derivative control coefficients, respectively. When the source node where the task is located is detected to meet the preset abnormal conditions, the migration is triggered based on the task checkpoint status. The total migration cost and recovery utility are calculated for each candidate target node, and the target node with the largest recovery utility is selected to perform task migration recovery. After completing scaling up / down and migration recovery, a scheduling report is constructed and the scheduling weight parameters are adaptively updated based on the scheduling report to form an elastic closed-loop scheduling of artificial intelligence training resources in a cloud-edge hybrid deployment scenario.
[0015] In one possible implementation, when the source node where the task is located is detected to meet a preset abnormal condition, the migration is triggered based on the task checkpoint state. The total migration cost and recovery utility are calculated for each candidate target node, and the target node with the highest recovery utility is selected to perform task migration recovery. This includes: When the task execution process encounters increased pressure on the source node, increased risk of failure, increased network jitter, or service failure, the computation task... migration trigger value : in: The current pressure on the source node of the task. For the risk of source node failure, For the degree of network fluctuation related to the task, and These represent the number of task breaches and the number of detections, respectively. To check the freshness, These are the corresponding weighting coefficients; After triggering the migration, for each candidate target node Calculate the total migration cost : in: For the mirror volume, For the volume of the checkpoint, This represents the bandwidth from the source node to the target node. This refers to runtime memory state variables. For the target node I / O recovery rate, Assess the reliability of the target node. For heterogeneous architecture differences, To reduce the complexity of the reconstruction, This corresponds to the penalty coefficient; Subsequently, the recovery utility is calculated by combining the target node's adaptability and migration cost. : in: For the overall fit between the task and the target node, To restore the completion time, To recover the time normalization constant, The probability of future overload for the target node. These are the total migration cost penalty coefficient, the fast recovery reward coefficient, and the target node future overload probability penalty coefficient, respectively.
[0016] In one possible implementation, after completing scaling up / down and migration recovery, the process of constructing a scheduling report and adaptively updating the scheduling weight parameters based on the report to form an elastic closed-loop scheduling of AI training resources in a cloud-edge hybrid deployment scenario includes: After completing scaling up / down and migration recovery, the effectiveness of the current scheduling round is evaluated, and the scheduling reward function is calculated. : in: To improve the success rate of task initiation, For resource utilization, To improve service completion rate, For the benefit of local data, For the overall cost of resources, For migration cost or migration impact indicators, For system energy consumption, These are the corresponding weighting coefficients; Furthermore, the scheduling weights are adaptively updated based on the contribution of each evaluation factor to the current return: in: For the first The weight of each scheduling indicator in the current round. Based on current contribution level, To contribute to the change, This refers to historical mismatch values or regret values. , , To learn the control parameters.
[0017] In this embodiment, by constructing a composite profile of tasks and nodes, resource requirements are accurately predicted and task urgency is calculated, achieving precise matching between tasks and cloud-edge nodes. This balances the low-latency requirements of interactive tasks with the high-computing-power requirements of training tasks, overcoming the limitations of purely cloud or edge deployments. A comprehensive adaptation relationship is generated by combining multi-dimensional scoring, optimizing the cloud-edge hybrid deployment scheme and maximizing synergistic benefits and deployment utility under resource constraints. Through runtime monitoring, elastic scaling, and fault migration mechanisms, the continuous and stable operation of training tasks is ensured, reducing node failure risks and service default probabilities. Simultaneously, scheduling parameters are dynamically updated, and scheduling strategies are continuously optimized to adapt to fluctuations in training task load. This effectively solves the problems of existing technologies in cloud-edge deployment of AI training tasks, which focus on a single dimension, lack continuous operation, and are insufficient in overall elastic coordination, significantly improving the operational efficiency and resource utilization of training tasks. Attached Figure Description
[0018] Figure 1 A flowchart illustrating a method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment, as provided in this application embodiment; Figure 2 A bar chart showing resource utilization as provided in the embodiments of this application; Figure 3 The task startup latency curve provided in the embodiments of this application; Figure 4 Service achievement rate curve provided for embodiments of this application; Figure 5 The expansion and contraction histogram provided for embodiments of this application; Figure 6 The migration recovery curve provided in the embodiments of this application. Detailed Implementation
[0019] The present solution will now be described in conjunction with the accompanying drawings and specific embodiments.
[0020] The method provided in this implementation is designed for AI training scenarios where cloud and edge nodes collaborate. It provides a unified model for the resource requirements, latency constraints, data locality, node health status, link status, and operational risks of training tasks. Based on this model, it enables task profiling, demand prediction, node adaptation, hybrid cloud-edge deployment, elastic scaling, checkpoint migration and recovery, and closed-loop updates of scheduling parameters. The method is preferably executed by a processor deployed in the cloud control plane, edge access nodes, or a cloud-edge collaborative scheduling platform.
[0021] In this embodiment, all raw metrics are derived from the actual data that can be collected during system operation, preferably including: task submission information, container orchestration logs, GPU and CPU monitoring metrics, link telemetry metrics, image cache status, data cache status, node heartbeat information, anomaly alarm records, task execution logs, task timeout records, task restart records, and service completion logs.
[0022] The various weights, coefficients, thresholds, and matrix parameters in the formula are preferably determined by the method of "first offline calibration, then online correction". That is, initial values are first given based on historical sample statistics, expert experience, or capacity planning results, and then adaptive adjustments are made based on scheduling reports during operation.
[0023] See Figure 1 The cloud-edge hybrid deployment method for elastic scheduling of artificial intelligence training resources provided in this embodiment includes: S101 acquires task attribute data of the AI training tasks to be scheduled, as well as resource status data of cloud nodes and edge nodes, and constructs composite task profiles and composite node profiles.
[0024] In this step, resource request information, dataset-related information, latency constraint information, and interaction behavior information of the AI training tasks to be scheduled are first collected. Specifically, the task-side data includes at least: CPU requirements, GPU requirements, memory requirements, storage requirements, bandwidth requirements, task-related data scale, tolerable waiting time, task priority, task latency sensitivity, and task interaction intensity. Simultaneously, node-side status data of each candidate cloud node and edge node is collected, including at least: available CPU resources, available GPU resources, available memory resources, available storage resources, available bandwidth resources, cache status, resource utilization, queuing status, health status, and fault risk status.
[0025] Because different indicators have different units of measurement, it is preferable to perform normalization on each indicator before constructing the profile. For benefit-type indicators, i.e., indicators where larger values indicate better node or task status, such as available CPU, available GPU, cache hit rate, etc., the following linear range normalization function is preferred: in: This is the original value of the current indicator. and These are the minimum and maximum values of the indicator within the preset statistical window, respectively. To prevent the use of tiny positive numbers with a denominator of zero, it is preferable to select... .
[0026] For cost-related metrics, i.e., metrics where smaller values indicate better performance, such as latency, packet loss rate, queue length, and average recovery time, the reverse normalization method is preferred. In the presence of outliers, it is preferable to perform truncation first: Again Perform normalization, where and These are the lower and upper cutoff thresholds for the indicator, preferably the 5th and 95th percentiles of its historical distribution. This method helps avoid distortion of the subsequent profile caused by extreme monitoring values.
[0027] After normalization, for the first... Each training task constructs a composite profile vector. : in, , , , and They represent the first Each training task is at any time Requirements for CPU, GPU, memory, storage, and bandwidth. Indicates the scale of data associated with the task. This indicates the tolerable waiting time for the task. Indicates task priority. Indicates task latency sensitivity. Indicates the intensity of task interaction. Indicates the volatility of task requirements. This indicates the historical default risk value of the task.
[0028] To ensure the task profile reflects the joint resource allocation relationships between different resources, the overall task demand intensity is further calculated. The multidimensional requirements are coupled and modeled to obtain the comprehensive task requirement intensity. : in: Indicates the task is in the Normalized demand values in the resource category dimension Indicates the weight of a single-dimensional resource. This represents the coupling weight between different resource dimensions. , and These represent the weights for data size, latency sensitivity, and interaction intensity, respectively. Coupling term. This is used to characterize the joint consumption relationship between two types of resources, such as the coupling relationship between GPU and video memory, bandwidth and data scale, and CPU and interaction intensity. This coupling term takes the form of a bilinear product; in other implementations, a sum-of-squares coupling function, a geometric mean coupling function, or a kernel function coupling function can also be used, as long as the monotonicity of the resource coupling relationship remains unchanged.
[0029] In this embodiment, This is used to measure the basic contribution of resource dimensions such as CPU, GPU, memory, storage, and bandwidth to the resource pressure of a task. The larger the value, the more critical the corresponding resource dimension. All All are non-negative real numbers. In training-type practical tasks, the GPU and video memory corresponding to... The values are relatively large. In interactive training tasks, the corresponding CPU, memory, and bandwidth values... The value is relatively large. Used to characterize the joint occupancy relationship between two types of resources, such as the coupling effect between GPU and video memory, bandwidth and data scale, and CPU and interaction intensity. The larger the value, the stronger the linkage between the two types of resources. The values are non-negative real numbers, and their sum is not greater than the sum of the weights of single-dimensional resources, in order to avoid the coupling terms from excessively amplifying the intensity of comprehensive demand. This indicates the weight of the task data size in the overall demand intensity, reflecting the degree of impact of the total data volume on storage usage, cache warm-up, and data retrieval pressure. This indicates the weight of task latency sensitivity in the overall demand intensity. This indicates the weight of the task interaction intensity in the overall demand intensity. For training tasks that primarily involve low-latency interaction, and The value is relatively large; for training tasks on large datasets, The value is relatively large.
[0030] To characterize the degree of fluctuation in task requirements over time, it is preferable to calculate the dispersion of the overall task requirements within a sliding window to obtain the task requirement volatility. : in: The sliding window length is preferably set according to the scheduling cycle, such as 5 scheduling cycles, 10 scheduling cycles, or 30 scheduling cycles. This represents the average overall resource requirement of the task within the window. The above volatility formula is essentially in the form of a coefficient of variation. The larger the value, the more pronounced the fluctuation and the stronger the suddenness of the task's demand, requiring more elastic resources to be reserved in subsequent scheduling.
[0031] On the node side, simultaneously for the first Each resource node builds a composite supply capacity. : in: For node number Normalized availability of class resources For node number Resource utilization rate For node health, To alleviate queuing pressure, For the risk of failure, To leverage cache hit advantage, , , , , , For the corresponding weights.
[0032] Specifically, This represents the positive contribution weight of the k-th type of available resources at node in the overall supply capacity. The larger the value, the more important the corresponding available resource. This represents the negative penalty weight of the utilization rate of the k-th type of resource in the overall supply capacity. The larger the value, the more sensitive the system is to the current occupied status of that type of resource. For scarce resources such as GPUs and video memory, the corresponding... and The value is relatively large. The positive weights representing node health are used to characterize the degree of importance the system attaches to the health status of nodes. The penalty weight representing queuing pressure, The larger the value, the more the system tends to avoid congested nodes in the queue. The penalty weight represents the risk of failure. The larger the value, the better the system avoids high-failure nodes. The positive weight representing the advantages of caching The larger the value, the more the system values the deployment benefits derived from local data, images, and model file hits.
[0033] Specifically, Calculated in the following way: in, For the data cache hit rate of the node, For container image cache hit rate, For the local readiness rate of the model file, , and This refers to the weight of the cached components.
[0034] In this embodiment, The weights representing the data cache hit rate The weight representing the container image cache hit rate The weights represent the local readiness rate of the model file. , and It is a non-negative real number, and the sum of the three is 1. When the task's main time consumption is during data retrieval... A larger value indicates a longer image startup time. When the value is large, it indicates that the task strongly depends on the local model file. The values are relatively large. Through the above processing, we obtain the task composite profile, the overall demand intensity of the task, the volatility of the task demand, the historical default risk of the task, and the overall supply capacity of the nodes, which serve as inputs for subsequent prediction and scheduling.
[0035] Based on the constructed task composite profile and node comprehensive supply capabilities, various artificial intelligence training tasks are scheduled. After scheduling, the CPU, GPU, memory, and bandwidth utilization of cloud nodes and edge nodes are statistically analyzed, and the results are as follows: Figure 2 As shown. By Figure 2 As can be seen, based on task profiling, resource demand prediction, and task-node comprehensive adaptability calculation, this invention can more rationally allocate computing power and orchestrate instances between cloud nodes and edge nodes according to the heterogeneous resource requirements of training tasks, keeping the utilization rate of various resources within a high and relatively balanced range. This figure reflects that this invention does not merely improve the utilization rate of a single resource dimension, but rather achieves coordinated utilization of multiple resources such as CPU, GPU, memory, and bandwidth by comprehensively considering data locality, node status, and deployment costs. This reduces the phenomenon of some nodes being idle while others are overloaded, thereby improving overall resource scheduling efficiency.
[0036] Simultaneously, the entire time elapsed from task submission to completion and startup is recorded, and the resulting startup latency changes are as follows: Figure 3As shown in the figure, the change in startup latency can be used to reflect the effect of the present invention on dynamically optimizing the task deployment location under the constraints of edge bearer weight calculation, cloud-edge collaboration benefit evaluation, and cold start cost. For AI training tasks with high latency sensitivity and high interaction intensity, this embodiment prioritizes scheduling them to edge nodes or cloud-edge collaboration node combinations with high data locality and optimal link status, thereby reducing the additional waiting time caused by image pulling, data preheating, and remote access. Figure 2 This invention can effectively shorten the average waiting time during the task startup phase, and improve the response speed and user experience of interactive training tasks.
[0037] Based on the task composite profile and node composite profile, S102 predicts the resource requirements of each training task and calculates the overall urgency of the task.
[0038] In this step, based on the task composite profile and historical task execution trajectory constructed in step S101, the multi-dimensional resource requirements of the task in the next scheduling cycle are predicted. For the first... The task in the first Next-moment demand prediction in resource-like dimensions The following fusion prediction model was used to obtain the results: in: This is the historical smoothing coefficient. The coefficient for the trend term. For the coefficient of the periodic term, This is the error correction factor. This refers to the periodic demand component of the task in the corresponding resource dimension. This represents the current prediction residual. This prediction model comprehensively utilizes current observations, historical predictions, the difference trend between adjacent time points, periodic components, and prediction error feedback, enabling it to better adapt to the scheduled peaks and sudden loads of practical training operations.
[0039] In this embodiment, Used to balance the influence of current actual observations and historical predictions. The larger the value, the more the model relies on current observations; the smaller the value, the more the model relies on historical prediction trends. Used to measure the strength of the impact of the most recent change in resource demand on the forecast value at the next time point. The larger the value, the more sensitive the model is to short-term growth or decline trends. It is used to characterize the extent to which cyclical resource demand patterns play a role in forecast results, and is especially suitable for practical training operations with timed peaks, class-based visits, or obvious intraday patterns. Used to compensate for the current prediction deviation in the next prediction. The larger the value, the faster the model corrects prediction errors. All four coefficients are non-negative. Located in the interval [0,1], , , It can be calibrated based on the historical prediction error distribution and the intensity of sudden load bursts.
[0040] To avoid the source of periodic terms being unclear, in this embodiment, the periodic demand component... The preferred approach is to use the historical average over the same period: Where: P is the cycle length, preferably set according to the business scenario as a daily cycle, class hour cycle, or cycle; M is the number of review cycles. In addition to the above forms, sine-cosine expansion forms or timetable correction forms can also be used, as long as they can represent the repetitive fluctuation pattern of the task in the time dimension.
[0041] After obtaining the multidimensional resource forecasts, the forecast results of each dimension are integrated with the historical average demand, volatility, and historical default risk of the task to calculate the task's sudden risk coefficient. : in: This indicates the task is the first one in the history window. Average demand for such resources This indicates the degree of service breach in the past.
[0042] In this embodiment, the risk is approximately... The preferred method is to obtain the following: in: This represents the number of times the task times out within the statistics window. The number of times the task deployment failed. This refers to the number of times the task was restarted due to abnormalities. This represents the total number of task scheduling attempts. , and This represents the risk component weight. Therefore, historical default risk can be translated into a quantity directly obtained from the log system statistics. This indicates the timeout risk weight, which measures the contribution of task timeout behavior to historical default risk. This indicates the risk weight of deployment failure, which measures the contribution of task initiation failure or scheduling failure to the risk of historical default. This indicates the risk weight of abnormal restarts, which measures the degree to which interruptions and restarts during task execution contribute to the risk of historical defaults. , and All are non-negative real numbers, and their sum is 1. When the system is more concerned about a task failing to start, Greater than and When the system is more concerned about excessively long task queuing times, The value is relatively large. The logarithmic function in the above formula... This is used to compress extreme growth values, preventing a single resource spike from excessively amplifying the risk outcome. In other implementations, square root functions, hyperbolic tangent functions, or piecewise linear functions can also be used to achieve risk compression, as long as the risk increases monotonically with demand.
[0043] Based on the task's sudden risk coefficient, the overall task urgency is calculated by further considering task waiting time, task priority, latency sensitivity, and interaction intensity. : in: to The fusion coefficient is... Let be the time normalization constant. This is a coupling term between latency sensitivity and interaction strength. The threshold for determining urgent tasks. This is an indicator function. It represents the tolerable waiting time for the task. Less than or equal to the threshold When the threshold is set, the indicator function term is 1; otherwise, it is 0, thereby achieving additional priority for ultra-short time-limited tasks. It can be set based on the 10th percentile of the waiting tolerance time for all tasks, a fixed percentage of the average, or the upper limit of the system service level.
[0044] In this embodiment, This indicates the weight of the unexpected risk factor in the overall urgency of the task. The larger the value, the more the system prioritizes the risk of resource explosion. This indicates the weight of the waiting time limit item. The larger the value, the more the system values the tolerance time for task waiting. This indicates the task priority weight, which reflects the impact of business level or course level on the scheduling order. The weights of the coupling term between latency sensitivity and interaction strength are represented. The larger the value, the more the system tends to prioritize scheduling tasks that are both latency-sensitive and have strong interactive characteristics. This indicates the weight of the emergency waiting threshold bonus item, which is applied when the tolerable waiting time for a task is less than or equal to a preset threshold. At that time, this option can be used to further increase the urgency of the task. to All are non-negative real numbers and satisfy normalization constraints. For interactive debugging, online inference, and remote desktop tasks, , and The value is relatively large. For resource-explosive training tasks, The value is relatively large.
[0045] Through the above processing, we obtain the resource prediction results for the next cycle, the sudden risk coefficient of the task, and the overall urgency of the task for each task, providing input for the subsequent generation of task-node adaptation relationships.
[0046] S103 combines the predicted demand results, urgency results, and the comprehensive supply capacity of nodes to calculate the data locality, resource matching degree, network advantage value, and node reliability score between each training task and each edge node and each cloud node, and generates a comprehensive adaptation relationship between task nodes.
[0047] In this step, for each training task, the data locality, resource matching degree, network advantage value and node reliability score between it and each edge node and each cloud node are calculated. Combined with the projection similarity between the task profile and the node profile, the task-node comprehensive fit is generated.
[0048] First, calculate the task based on the degree of overlap between the data blocks required by the task and the data blocks already cached by the node. Relative to node Data locality : in: Indicates task Required data block set Represents a node The set of cached data blocks For data blocks Importance weight, To address the discrepancy between the cached data version on the node and the version expected by the task, This is the version deviation normalization constant.
[0049] In this embodiment, It can be determined as follows: in, This is the normalized value for the data block size. This is the normalized value of the historical access frequency of the data block. This is a value used to identify critical data blocks. When there's no need to differentiate data block importance, a uniform value can be used. . Indicates the weight of data block size. Indicates the weight of data block access frequency. Indicates the weight of the key data block identifier. , , All are non-negative real numbers, and their sum is 1. When the system places greater emphasis on warming up large files, The value is relatively large; when the system places more emphasis on hitting hot data, The value is relatively large; when the system places more emphasis on label files, vocabulary files, or core training samples, The value is relatively large.
[0050] To clarify the specific source of the version deviation, preferably, the version deviation... It can take any of the following forms: absolute version number difference, hash inconsistency indicator value, or file-level difference ratio. The preferred form is the simplest one: or This allows us to determine whether the cached data of a node can be directly used for the current task, rather than simply determining whether a hit has occurred.
[0051] Then, based on the correspondence between the predicted task requirements and the available resources of the nodes, the resource matching degree is calculated. : in: For nodes In the Available capacity on class resources For the task In the Demand for similar resources To match the sensitivity index, The resource shortage penalty coefficient, To find the minimum value function, This is the function for finding the maximum value.
[0052] In this embodiment, the sensitivity index Used to adjust the shape of the satisfaction curve for various resource dimensions; for critical resources, when... When the value is greater than 1, the matching score can only be significantly improved when the satisfaction value is large. For general resources, It can be set to 1 or slightly less than 1. The larger the value, the greater the impact on matching performance if that type of resource becomes insufficient. For example, for critical resources such as GPU and video memory, a larger value can be set. and To increase the penalty for insufficient critical resources; for general-purpose resources such as CPU and general-purpose memory, a smaller penalty can be set. and This allows us to differentiate between slightly insufficient resources and severely insufficient resources.
[0053] Considering the significant impact of node link status on low-latency training services, the network dominance value is further calculated. : in: For task access side to node The time delay, For available bandwidth, As the reference bandwidth, For packet loss rate, For link jitter, to These are the corresponding weighting coefficients. Specifically, Indicates the weight of the delay term. Indicates the weight of the bandwidth term. This indicates the weight of the inverse term of the packet loss rate. This indicates the weight of the jitter item. to Let be non-negative real numbers, and satisfy the condition that their sum is 1. For interactive tasks, and The value is relatively large; for large file transfer tasks, The value is relatively large; for remote training sessions or long-connection sessions, The value is relatively large. This formula is used to characterize the advantages of a node in carrying interactive AI training tasks in the current link environment.
[0054] Furthermore, the operational stability of the nodes themselves is evaluated to obtain a node reliability score. : in: For node health, For the risk of failure, For queue occupancy rate, Mean time to recovery (MTBF) To recover the time normalization constant, These are the weighting coefficients. Specifically, This indicates the weight of the node health item. Indicates the weight of the inverse term of failure risk. This indicates the weight of the reverse term in the queuing pressure. This indicates the weight of the fault recovery capability item. All are non-negative real numbers and satisfy normalization constraints. When the system prioritizes long-term stable operation, and The value is relatively large; when the system prioritizes immediate availability, The value is relatively large.
[0055] In this embodiment, It can be obtained from underlying monitoring indicators such as heart rate normality, container availability, disk health, temperature anomaly rate, and driver alarm rate. It can be estimated by combining historical anomaly frequency, current alarm status, and equipment anomaly records.
[0056] After completing the above scoring, it is fused with the projection similarity of the task profile and node profile to generate the task. With nodes Overall compatibility : in: and These are the projection matrices for the task profile and the node profile, respectively. The cosine similarity function is used. to This corresponds to the fusion coefficient. Specifically, Represents the locality weight of the data. Indicates the weight of resource matching degree. Represents the network advantage value weight. Indicates the node reliability score weight. This represents the weight of the high-dimensional similarity term between the task profile and the node profile. Let be non-negative real numbers, and satisfy the condition that their sum is 1. For tasks with large datasets and significant cold-start costs, The value is relatively large; for high-computing-power tasks, The value is relatively large; for remote interactive tasks, The value is relatively large; for long-term training tasks, The value is relatively large.
[0057] To avoid the use of projection matrices of unknown origin, in this embodiment, and It adopts a fixed weighted diagonal matrix form, that is: In this model, each diagonal element corresponds to the importance weight of each dimension of the task profile or node profile. This maps the task profile and node profile to a unified feature space, and cosine similarity is used to characterize their matching degree in the high-dimensional space. In other implementations, a principal component projection matrix or a mapping matrix obtained offline based on historical successful scheduling samples can also be used, as long as it enables the matching and comparison of task features and node features in a unified space. To ensure simplicity and transparency, a fixed weighted diagonal matrix is preferred.
[0058] Through the above processing, a comprehensive adaptation result for each task relative to each edge node and each cloud node is generated, which serves as a direct input for subsequent cloud-edge hybrid deployment decisions.
[0059] S104 calculates the edge bearing weight, cloud-edge collaboration benefits, and collaborative deployment utility of each training task based on the comprehensive adaptation relationship and the overall urgency of the task, and determines the cloud-edge hybrid deployment scheme corresponding to each training task under resource constraints.
[0060] In this step, instead of simply assigning tasks to either the cloud or the edge, we first calculate the edge-bearing weight of the task based on its preference for low-latency edge support. : in: Indicates task Data locality on candidate edge sides To normalize GPU requirements, To normalize the data size, to This is the corresponding adjustment coefficient.
[0061] Specifically, This indicates the weight of the latency sensitivity item, reflecting the degree to which the task depends on a low-latency response; This indicates the overall urgency weight of the task, reflecting the necessity of prioritizing the scheduling of the task to the low-latency side; This indicates the interaction intensity weight, reflecting the tendency to prioritize deploying interactive tasks closer to the user; This represents the locality weight of edge-side data, reflecting the facilitating effect of task data being locally ready on edge nodes on edge deployment; This indicates the GPU demand penalty weight, reflecting that high GPU-intensive tasks should not be overly burdened on the edge side; This represents the data size penalty weight, reflecting the inhibitory effect of ultra-large dataset tasks on edge storage and caching capabilities. to are non-negative real numbers, where to Constitutes an edge-bearing promotion item. and This constitutes an edge-bearing suppression term. For low-latency interactive tasks, , and The value is relatively large; for GPU-intensive tasks with large data volumes, and The value is relatively large.
[0062] The above embodiment uses the Sigmoid mapping function to compress the task's preference for edge-carrying to the [0,1] interval, thereby ensuring the comparability of edge preferences between different tasks. Typically, when a task has high latency sensitivity, high interaction intensity, and high edge-side data readiness rate, The value is relatively large; when the task's GPU load value is large and the data size is too large, The value is relatively small.
[0063] To avoid deployment based solely on the score of a single node, further evaluation of candidate edge nodes is needed. With candidate cloud nodes The collaborative capabilities between the two sides are evaluated to obtain the benefits of cloud-edge collaboration. : in: For cloud-edge interconnect latency, For cloud-edge bandwidth, This represents the maximum observation bandwidth of the cloud-edge link. For cloud-edge link packet loss rate, and Leveraging the caching advantages of both edge and cloud sides, The cost of cloud-edge state synchronization These are the corresponding weighting coefficients.
[0064] Furthermore, Indicates the weight of cloud-edge latency advantage. Indicates the weight of cloud-edge bandwidth advantage. This indicates the weight of the inverse term of the cloud edge packet loss rate. Indicates the weight of the collaborative advantages between cloud and edge caching. This represents the penalty weight for cloud-edge state synchronization costs. All are non-negative real numbers, among which As a positive return weight, This is a negative penalty weight. If the system is sensitive to the synchronization cost of the cloud-edge control plane, then... It is worth teaching.
[0065] In this embodiment, Obtain it as follows: in: For the total amount of task metadata and control metadata, This refers to the number of times the state is synchronized per unit of time. As a normalization benchmark for the number of synchronizations, This refers to the clock offset at the cloud edge. This indicates the penalty weight for the size of synchronized metadata. Indicates the synchronization frequency penalty weight. This indicates the clock skew penalty weight. If cloud-edge states are frequently synchronized, then... The value is relatively large; if the consistency of remote cloud-edge clocks has a significant impact on checkpoint recovery, then... The value is relatively large. Therefore, the synchronization cost can be explicitly reflected in the synergistic benefits.
[0066] After calculating the edge-bearing weight and cloud-edge collaborative benefits, the task... Select edge nodes and cloud nodes Hybrid deployment schemes calculate the collaborative deployment utility : in: and These refer to the overall compatibility between tasks and edge nodes, and between tasks and cloud nodes. , , , These represent the costs of resource consumption, cold start, energy consumption, and operational risks, respectively. This represents the corresponding penalty coefficient.
[0067] Furthermore, This represents the penalty coefficient for resource consumption. This represents the cold start cost penalty coefficient. This represents the energy consumption penalty coefficient. This represents the penalty coefficient for operational risk. All four are preferably non-negative real numbers, and the larger their values, the stronger the inhibitory effect of the corresponding cost term on the total utility. When the system emphasizes rapid startup, The value is relatively large. When system stability is emphasized, The value is relatively large; when the system places more emphasis on green energy conservation, The value is relatively large.
[0068] Among them, the cost of resource occupation and cold start cost The preferred method is to calculate as follows: in: This refers to the size of the task image. The available bandwidth from the source location to the target deployment location. , , This corresponds to the penalty coefficient. Specifically, This represents the penalty coefficient for the pressure of resource occupancy in the k-th category. For scarce resources such as GPUs, video memory, and NVMe storage, The preference is higher. In the cold start cost formula, This indicates the cold start penalty coefficient for mirror pull. This represents the cold start penalty coefficient for data preheating. When the image size is large and the container startup time is long, The value is relatively large; when the dataset is large and the locality of the target node data is low, The value is relatively large.
[0069] To compensate for the specific source of energy consumption costs, preferably, the energy consumption costs... Obtain it as follows: in, This refers to the computational energy consumption when the task is executed on the edge. This refers to the computational energy consumption when the task is executed on the cloud side. Energy consumption for task data transmission , and This represents the weighting of the energy consumption component. Specifically, This indicates the energy consumption weight calculated on the edge side. This indicates the weight of cloud-side computing energy consumption. This indicates the weight of network transmission energy consumption. When the power supply to edge nodes is limited... The value is relatively large; when the cost of cross-domain data transmission is high, The value is relatively large. The energy consumption can be estimated by multiplying the node's average power by the execution time.
[0070] To supplement the specific sources of operational risk costs, preferably, the operational risk costs... Obtain it as follows: in, and These represent the fault risks on the edge side and the cloud side, respectively. and These represent the future overload probabilities on the edge side and the cloud side, respectively. to For risk component weights. Specifically, Indicates the risk weight of edge node failure. Indicates the cloud node failure risk weight. This represents the weight of the probability of future overload for edge nodes. This represents the weight of the probability of future overload for cloud nodes. If the stability of edge nodes is relatively weak, then... and The value is relatively large; if the central cloud is more prone to congestion during peak hours, then... The value is relatively large.
[0071] After calculating the deployment utility, further optimization is performed under resource constraints and the constraint of unique task deployment. Deployment decision variables. satisfy: Under the premise of satisfying resource capacity and scheduling uniqueness constraints, the deployment decision variables are determined by the following objective function. : in: Represents the set of edge nodes. Represents a set of nodes in the cloud. This represents the node pressure value. For global average pressure, This is the load balancing penalty coefficient. The larger the value, the more the system tends to distribute tasks to reduce the probability of hotspot nodes forming. The smaller the value, the more the system tends to prioritize locally efficient deployment utility. By introducing this penalty term into the deployment utility maximization goal, a trade-off between "locally optimal deployment benefits" and "global load balancing" can be achieved. This determines the final cloud-edge hybrid deployment result and completes the allocation of container instances, data caching, and compute resource quotas.
[0072] In this embodiment ,in: This refers to the total set of nodes, comprising all cloud and edge nodes. The load balancing penalty is used to prevent tasks from being overly concentrated on a few high-performance nodes.
[0073] Through the above optimization solution, the cloud-edge hybrid deployment results corresponding to each training task are obtained, and container instances, image preheating, data caching and computing resource quota allocation are completed accordingly.
[0074] After generating the comprehensive adaptability of task nodes in step S103, step S104 determines the final hybrid deployment result by combining edge bearer weight, cloud-edge collaboration benefits, and deployment costs. To verify the impact of this deployment result on service stability, the proportion of tasks meeting the preset service requirements in each scheduling cycle is statistically analyzed, and the results are as follows: Figure 4 As shown. By Figure 4 As can be seen, the service achievement rate mainly reflects whether a task can complete its startup, operation, and result output under preset latency, preset resource guarantees, and preset availability conditions. Because this embodiment considers not only task resource requirements and node available resources during scheduling, but also node health status, failure risk, link quality, and migration recovery capabilities, it can maintain a high service achievement rate even in high-concurrency or load-fluctuation scenarios. This demonstrates that by combining resource demand prediction, node comprehensive adaptability, operational risk costs, and closed-loop parameter updates, this embodiment can improve the stability and continuity of task scheduling and reduce service breaches caused by node congestion, link anomalies, or improper deployment.
[0075] S105 performs operational monitoring of deployed training tasks, performs elastic scaling based on node pressure prediction, and triggers checkpoint migration and recovery when the source node pressure is abnormal, the risk of failure increases, or the service is in default. At the same time, it updates the scheduling parameters based on the scheduling report.
[0076] In this step, the training task that has completed the cloud-edge hybrid deployment is continuously monitored in its runtime state. The CPU utilization, GPU utilization, memory utilization, bandwidth utilization, queue length, average response time, and stress change rate of each node are collected, and the overall stress value of the node is calculated. : in: , , , These are CPU, GPU, memory, and bandwidth utilization, respectively. Queue length The maximum allowed queue length, The average response time of the node. The rate of change of pressure, These are the corresponding weighting coefficients. Specifically, Indicates CPU utilization weight. Indicates GPU utilization weight. Indicates memory utilization weight. Indicates the bandwidth utilization weight. Indicates the weight of queue length. Indicates the average response time weight. This indicates the weight of the rate of change of pressure. It is a non-negative real number and satisfies the normalization constraint. For the training cluster, The value is relatively large; for interactive clusters, and The value is relatively large; for scenarios with significant load fluctuations, The value is relatively large.
[0077] To avoid the predicted pressure being left unresolved during expansion and contraction, preferably, the node predicted pressure... Obtain it as follows: in, For smoothing coefficients, This is the trend compensation coefficient. and The value is in the interval [0,1].
[0078] Based on current and predicted pressure, nodes are elastically scaled up or down, with instance adjustments made. Obtained by the following formula: in: To predict the pressure in the next moment, The target pressure threshold, For the historical cumulative window, , , These represent the proportional, integral, and derivative control coefficients, respectively. Used to reflect the direct effect of current predicted pressure deviation on expansion and contraction capacity; Used to reflect historical cumulative deviation; This is used to reflect the advance response of the system to pressure change trends during expansion and contraction. If the system is prone to prolonged pressure buildup, then... The value is relatively large; if the system load fluctuates rapidly, then The value is relatively large.
[0079] Based on the instance adjustment amount, container instances, interactive session instances, training replicas, or GPU sandbox instances can be scaled up or down to achieve elastic resource control.
[0080] When the task execution process encounters increased pressure on the source node, increased risk of failure, increased network jitter, or service failure, the computation task... migration trigger value : in: The current pressure on the source node of the task. For the risk of source node failure, For the degree of network fluctuation related to the task, and These represent the number of task breaches and the number of detections, respectively. To check the freshness, These are the corresponding weighting coefficients.
[0081] Specifically, Indicates the pressure weight of the source node. Indicates the source node failure risk weight. Indicates network fluctuation weights. Indicates the weight of the degree of service breach. This indicates the penalty weight for insufficient freshness at checkpoints. It is a non-negative real number. If the system is more concerned with the impending instability of a node, then... and The value is relatively large; if the system is more concerned with whether checkpoints were saved in a timely manner before migration, then The value is relatively large.
[0082] In this embodiment, The following time decay function is used: in, Save the time of the most recent checkpoint for the task. This is the checkpoint time decay constant. When the checkpoint was last saved, Approximately 1. When not stored for a long time, This reduces the migration trigger value, thereby increasing it.
[0083] Calculate as follows: in, Represents the coefficient of variation. and These represent the latency and bandwidth sequences related to the task, respectively. This represents the current packet loss rate. This allows for the explicit mapping of link instability to migration triggering factors. Indicates the weight of latency fluctuations. Indicates bandwidth fluctuation weight. Indicates the packet loss rate weight; for interactive tasks, and The value is relatively large, which is beneficial for large file transfer tasks. The value is relatively large.
[0084] The system performs elastic scaling up and down based on the node's overall stress value and its prediction results, and statistically analyzes the number of scaling up and down operations, as well as the scale of instance adjustments. The results are as follows: Figure 5 As shown in the figure, this diagram illustrates the elastic regulation effect implemented by the present invention based on the node-based comprehensive pressure value, predicted pressure value, and proportional-integral-derivative control concept. Figure 5As can be seen, when the number of task requests increases, the node queuing length increases, or the response time becomes longer, the system can proactively expand the corresponding instances based on the predicted pressure; when the load decreases, it can promptly reduce redundant instances, thereby avoiding long-term resource idleness. This figure illustrates that, in AI training scenarios, this invention can not only achieve static deployment but also dynamically adjust the number of container instances, interactive session instances, or GPU sandbox instances based on changes in runtime status, thereby improving the system's adaptability to sudden business surges and periodic peaks. After the migration is triggered, if the migration trigger value exceeds a preset threshold, for each candidate target node... Calculate the total migration cost : in: For the mirror volume, For the volume of the checkpoint, This represents the bandwidth from the source node to the target node. This refers to runtime memory state variables. For the target node I / O recovery rate, Assess the reliability of the target node. For heterogeneous architecture differences, To reduce the complexity of the reconstruction, This represents the corresponding penalty coefficient.
[0085] Specifically, This represents the penalty coefficient for insufficient reliability of the target node. This represents the architectural difference penalty coefficient. This represents the penalty coefficient for dependency reconstruction complexity. If the cluster is highly heterogeneous, then... The value is relatively large; if the operating environment has complex dependencies, then... The value is relatively large.
[0086] In this embodiment, It can be obtained in the following way: in: This indicates the weight of the penalty for differences in the underlying architecture. Indicates the driving difference penalty weight, This indicates the runtime version difference penalty weight. If GPU driver incompatibility has the greatest impact on migration, then... The value is relatively large.
[0087] It can be obtained in the following way: Here, Arch, Driver, and CUDA represent the underlying architecture, driver version, and runtime version, respectively. The number of missing dependencies for the target node. This represents the total number of task dependencies.
[0088] Subsequently, the recovery utility is calculated by combining the target node's adaptability and migration cost. : in: For the overall fit between the task and the target node, To restore the completion time, To recover the time normalization constant, The probability of future overload for the target node. These are the total migration cost penalty coefficient, the rapid recovery reward coefficient, and the future overload probability penalty coefficient for the target node, respectively. If the system prioritizes the rapid recovery of task execution, then... The value is relatively large; if the system is more concerned about congestion recurring after task migration, then... The value is relatively large.
[0089] In this embodiment, This can be represented using the Logistic probability mapping function: in, to For regression coefficients, For bias terms, This represents the predicted pressure impact coefficient. This represents the influence coefficient of the pressure change rate. This represents the queue length influence coefficient. This set of coefficients is preferably obtained through offline fitting of historical overload samples, and is used to map the current state of a node to the probability of future overload. This allows pressure, pressure change trends, and queue length to be mapped to future overload risk.
[0090] By selecting the target node with the highest recovery utility, checkpoint-based task migration recovery and training state continuation are achieved. After scaling up / down and migration recovery are completed, the effectiveness of the current scheduling round is evaluated, and the scheduling reward function is calculated. : in: To improve the success rate of task initiation, For resource utilization, To improve service completion rate, For the benefit of local data, For the overall cost of resources, For migration cost or migration impact indicators, For system energy consumption, These are the corresponding weighting coefficients.
[0091] Specifically, Indicates the weight of the task startup success rate. Indicates the weight of resource utilization rate. Indicates the weight of service completion rate. Indicates the weight of locality of data benefits. Indicates the system cost penalty weight. This indicates that migration affects the penalty weight. This indicates the system's energy consumption penalty weight. If the system prioritizes successful task startup, then... The value is relatively large; if the system focuses more on overall resource utilization, then The value is relatively large; if the system is more focused on consistently achieving the service level, then The value is relatively large; if the system places greater emphasis on reducing migration interference and lowering energy consumption, then and The value is relatively large.
[0092] To clarify the specific source of each return component, the above indicators were obtained as follows: in, For the task to start successfully, For the number of task submissions, To serve the number of defaulted tasks, The total number of tasks. Indicates the actual deployment location of the task.
[0093] When a source node is detected to meet preset abnormal conditions, migration recovery is triggered based on the task checkpoint status. Statistics are then compiled on migration preparation time, status transmission time, recovery completion time, and post-migration stable operation. The results are as follows: Figure 6 As shown. Figure 6 This shows the relevant performance changes after the system performs checkpoint migration and recovery when the source node pressure is abnormally high, the risk of failure is increased, network fluctuations are aggravated, or the degree of service breach is increased. These changes include recovery time, recovery success rate, and the effect of stable operation after migration. Figure 6 The results reflect the migration trigger value calculation, total migration cost evaluation, and recovery utility selection mechanisms in this invention. By comparing the comprehensive adaptability, migration cost, and future overload risk of candidate target nodes, this invention can select the target node with the highest recovery utility to perform task migration recovery, thereby reducing task interruption time and maintaining the continuity of training or session states. This figure illustrates that this invention possesses good fault recovery capabilities and operational continuity in cloud-edge hybrid deployment scenarios, and can maintain the stable execution of AI training tasks under conditions of node anomalies or link fluctuations.
[0094] Finally, the scheduling weights are adaptively updated based on the contribution of each evaluation factor to the current return: in: For the first The weight of each scheduling indicator in the current round. Based on current contribution level, To contribute to the change, This refers to historical mismatch values or regret values. , , To learn the control parameters. Specifically, This represents the current contribution amplification factor. Indicates the amplification factor for the trend of contribution change. This represents the historical mismatch penalty coefficient. , , The value is in the range [0,1]. If the system places more emphasis on the performance in the current round, then... The value is relatively large; if the system places more emphasis on whether there has been continuous improvement in recent rounds, then... The value is relatively large; if the system aims to suppress indicators that have consistently performed poorly, then... The values are relatively large. This allows the system to adaptively adjust various weights based on actual operational results, ultimately forming a flexible closed-loop scheduling mechanism for AI training resources in hybrid cloud-edge deployment scenarios.
[0095] In this embodiment, It can be obtained by using partial derivatives of the scheduling reward or by finite difference approximation; It can be calculated based on the cumulative difference between this indicator and the optimal indicator in historical rounds. This allows for the elastic closed-loop scheduling of AI training resources in a hybrid cloud-edge deployment scenario.
[0096] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0097] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for elastic scheduling of artificial intelligence training resources in a hybrid cloud-edge deployment, characterized in that, include: Acquire task attribute data of the AI training tasks to be scheduled, as well as resource status data of cloud nodes and edge nodes, and construct composite task profiles and composite node profiles; Based on the aforementioned task composite profile and node composite profile, the resource requirements of each training task are predicted, and the overall urgency of the task is calculated. Based on the predicted demand results, urgency results, and overall node supply capacity, the data locality, resource matching degree, network advantage value, and node reliability score between each training task and each edge node and each cloud node are calculated, and the comprehensive adaptation relationship between task nodes is generated. Based on the comprehensive adaptation relationship and the comprehensive urgency of the tasks, the edge bearing weight, cloud-edge collaboration benefits and collaborative deployment utility of each training task are calculated, and the cloud-edge hybrid deployment scheme corresponding to each training task is determined under resource constraints. The system monitors the operational status of deployed training tasks, performs elastic scaling based on node pressure prediction, and triggers checkpoint migration and recovery when the source node experiences abnormal pressure, increased fault risk, or service default. Simultaneously, it updates scheduling parameters based on scheduling reports.
2. The method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment according to claim 1, characterized in that, The process of acquiring task attribute data of the AI training task to be scheduled, as well as resource status data of cloud nodes and edge nodes, and constructing composite task profiles and composite node profiles includes: Collect resource request information, dataset information, latency constraint information, and interaction information of the AI training tasks to be scheduled, and simultaneously collect heterogeneous indicator information of each cloud node and edge node; The collected heterogeneous index information is normalized so that computing power indexes, link indexes and task indexes of different dimensions can enter the same computing space. After normalization, task composite profiles and node composite profiles are constructed respectively, and the comprehensive demand intensity of tasks, the volatility of task demand, and the comprehensive supply capacity of nodes are calculated to obtain the demand representation results on the task side and the supply representation results on the node side.
3. The method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment according to claim 2, characterized in that, The process involves constructing composite task profiles and composite node profiles, and calculating the overall demand intensity, demand volatility, and overall node supply capacity to obtain task-side demand representation results and node-side supply representation results, including: For the Each training task constructs a composite profile vector. : in, , , , and They represent the first Each training task is at any time Requirements for CPU, GPU, memory, storage, and bandwidth. Indicates the scale of data associated with the task. This indicates the tolerable waiting time for the task. Indicates task priority. Indicates task latency sensitivity. Indicates the intensity of task interaction. Indicates the volatility of task requirements. This indicates the historical default risk value of the task; Regarding the The multidimensional requirements are coupled and modeled to obtain the comprehensive task requirement intensity. : in: Indicates the task is in the Normalized demand values in the resource category dimension Indicates the weight of a single-dimensional resource. This represents the coupling weight between different resource dimensions. , and These represent the weights for data size, latency sensitivity, and interaction intensity, respectively. The overall resource requirements of a task within a sliding window are statistically analyzed to obtain the task requirement fluctuation. : in: The length of the sliding window. This represents the average total resource requirements of the task within the window. To prevent tiny positive numbers with a denominator of zero; At the same time, for the first Each resource node builds a composite supply capacity. : in: For node number Normalized availability of class resources For node number Resource utilization rate For node health, To alleviate queuing pressure, For the risk of failure, To leverage cache hit advantage, , , , , , For the corresponding weights.
4. The method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment according to claim 3, characterized in that, Based on the task composite profile and node composite profile, the resource requirements of each training task are predicted, and the overall urgency of the task is calculated, including: For the The task in the first Next-moment demand prediction in resource-like dimensions The following fusion prediction model was used to obtain the results: in: This is the historical smoothing coefficient. The coefficient for the trend term. For the coefficient of the periodic term, This is the error correction factor. This refers to the periodic demand component of the task in the corresponding resource dimension. This represents the current predicted residual; After obtaining the multidimensional resource forecasts, the forecast results of each dimension are integrated with the historical average demand, volatility, and historical default risk of the task to calculate the task's sudden risk coefficient. : in: This indicates the task is the first one in the history window. Average demand for such resources Indicates the degree of service breach in the past. Further calculate the overall urgency of the task. : in: to The fusion coefficient is... Let be the time normalization constant. This is a coupling term between latency sensitivity and interaction strength. The threshold for determining urgent tasks. This is an indicator function.
5. The method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment according to claim 4, characterized in that, Combining the predicted demand results, urgency results, and overall node supply capacity, the data locality, resource matching degree, network advantage value, and node reliability score between each training task and each edge node and each cloud node are calculated, and a comprehensive adaptation relationship between task nodes is generated, including: Calculate the task based on the degree of overlap between the data blocks required by the task and the data blocks already cached by the node. Relative to node Data locality : in: Indicates task Required data block set Represents a node The set of cached data blocks For data blocks Importance weight, To address the discrepancy between the cached data version on the node and the version expected by the task, This is the version deviation normalization constant; Then, based on the correspondence between the predicted task requirements and the available resources of the nodes, the resource matching degree is calculated. : in: For nodes In the Available capacity on class resources For the task In the Demand for similar resources To match the sensitivity index, The resource shortage penalty coefficient, To find the minimum value function, This is a function to find the maximum value. Further calculate the network dominance value : in: For task access side to node The time delay, For available bandwidth, As the reference bandwidth, For packet loss rate, For link jitter, to These are the corresponding weighting coefficients; At the same time, the reliability of the nodes themselves is scored, resulting in... : in: For node health, For the risk of failure, For queue occupancy rate, Mean time to recovery (MTBF) To recover the time normalization constant, These are the weighting coefficients; After completing the above scoring, it is fused with the projection similarity of the task profile and node profile to generate the task. With nodes Overall compatibility : in: and These are the projection matrices for the task profile and the node profile, respectively. The cosine similarity function is used. to This represents the corresponding fusion coefficient.
6. The method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment according to claim 5, characterized in that, Based on the comprehensive adaptation relationship and the overall urgency of the tasks, the edge carrying weight, cloud-edge collaboration benefits, and collaborative deployment utility of each training task are calculated. Under resource constraints, a cloud-edge hybrid deployment scheme is determined for each training task, including: Computational tasks Edge bearing weight : in: Indicates task Data locality on candidate edge sides To normalize GPU requirements, To normalize the data size, to This corresponds to the adjustment coefficient; Then, calculate the edge nodes. With cloud nodes Synergistic benefits between : in: For cloud-edge interconnect latency, For cloud-edge bandwidth, This represents the maximum observation bandwidth of the cloud-edge link. For cloud-edge link packet loss rate, and Leveraging the caching advantages of both edge and cloud sides, The cost of cloud-edge state synchronization These are the corresponding weight coefficients; Based on edge carrying weight, comprehensive adaptability of tasks and nodes, and cloud-edge collaboration benefits, combined with resource occupation cost, cold start cost, energy consumption cost, and operational risk cost, the collaborative deployment utility of each training task is calculated. Under resource capacity constraints and task unique deployment constraints, optimization solutions are performed to determine the cloud-edge hybrid deployment results of each training task.
7. The method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment according to claim 6, characterized in that, The method, based on edge bearer weight, comprehensive adaptability of tasks and nodes, and cloud-edge collaboration benefits, combined with resource consumption costs, cold start costs, energy consumption costs, and operational risk costs, calculates the collaborative deployment utility of each training task. It then optimizes the solution under resource capacity constraints and unique task deployment constraints to determine the cloud-edge hybrid deployment results for each training task, including: For the task Select edge nodes and cloud nodes Hybrid deployment schemes calculate the collaborative deployment utility : in: and These refer to the overall compatibility between tasks and edge nodes, and between tasks and cloud nodes. , , , These represent the costs of resource consumption, cold start, energy consumption, and operational risks, respectively. This corresponds to the penalty coefficient; Resource consumption cost and cold start cost The preferred method is to calculate as follows: in: This refers to the size of the task image. The available bandwidth from the source location to the target deployment location. , , This corresponds to the penalty coefficient; Under the premise of satisfying resource capacity and scheduling uniqueness constraints, the deployment decision variables are determined by the following objective function. : in: Represents the set of edge nodes. Represents a set of nodes in the cloud. This represents the node pressure value. For global average pressure, This is the load balancing penalty coefficient.
8. The method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment according to claim 1, characterized in that, The process involves monitoring the operational status of deployed training tasks, performing elastic scaling based on node pressure prediction, and triggering checkpoint migration and recovery when source node pressure is abnormal, fault risk increases, or service defaults occur. Simultaneously, scheduling parameters are updated based on scheduling reports, including: After the mixed deployment of the practical training tasks was completed, the first... Calculate the overall pressure value at each node. : in: , , , These are CPU, GPU, memory, and bandwidth utilization, respectively. Queue length The maximum allowed queue length, The average response time of the node. The rate of change of pressure, These are the corresponding weighting coefficients; Based on current and predicted pressure, nodes are elastically scaled up or down, with instance adjustments made. Obtained by the following formula: in: To predict the pressure in the next moment, The target pressure threshold, For the historical cumulative window, , , These represent the proportional, integral, and derivative control coefficients, respectively. When the source node where the task is located is detected to meet the preset abnormal conditions, the migration is triggered based on the task checkpoint status. The total migration cost and recovery utility are calculated for each candidate target node, and the target node with the largest recovery utility is selected to perform task migration recovery. After completing scaling up / down and migration recovery, a scheduling report is constructed and the scheduling weight parameters are adaptively updated based on the scheduling report to form an elastic closed-loop scheduling of artificial intelligence training resources in a cloud-edge hybrid deployment scenario.
9. The method for elastic scheduling of artificial intelligence training resources in a cloud-edge hybrid deployment according to claim 8, characterized in that, When the source node where the task is located is detected to meet a preset abnormal condition, migration is triggered based on the task checkpoint status. The total migration cost and recovery utility are calculated for each candidate target node, and the target node with the highest recovery utility is selected to perform task migration recovery, including: When the task execution process encounters increased pressure on the source node, increased risk of failure, increased network jitter, or service failure, the computation task... migration trigger value : in: The current pressure on the source node of the task. For the risk of source node failure, For the degree of network fluctuation related to the task, and These represent the number of task breaches and the number of detections, respectively. To check the freshness, These are the corresponding weighting coefficients; After triggering the migration, for each candidate target node Calculate the total migration cost : in: For the mirror volume, For the volume of the checkpoint, This represents the bandwidth from the source node to the target node. This refers to runtime memory state variables. For the target node I / O recovery rate, Assess the reliability of the target node. For heterogeneous architecture differences, To reduce the complexity of the reconstruction, This corresponds to the penalty coefficient; Subsequently, the recovery utility is calculated by combining the target node's adaptability and migration cost. : in: For the overall fit between the task and the target node, To restore the completion time, To recover the time normalization constant, The probability of future overload for the target node. These are the total migration cost penalty coefficient, the fast recovery reward coefficient, and the target node future overload probability penalty coefficient, respectively.
10. The method for elastic scheduling of AI training resources in a cloud-edge hybrid deployment according to claim 9, characterized in that, After completing scaling up / down and migration recovery, the process involves constructing a scheduling report and adaptively updating the scheduling weight parameters based on the report to form an elastic closed-loop scheduling of AI training resources in a cloud-edge hybrid deployment scenario. This includes: After completing scaling up / down and migration recovery, the effectiveness of the current scheduling round is evaluated, and the scheduling reward function is calculated. : in: To improve the success rate of task initiation, For resource utilization, To improve service completion rate, For the benefit of local data, For the overall cost of resources, For migration cost or migration impact indicators, For system energy consumption, These are the corresponding weighting coefficients; Furthermore, the scheduling weights are adaptively updated based on the contribution of each evaluation factor to the current return: in: For the first The weight of each scheduling indicator in the current round. Based on current contribution level, To contribute to the change, This refers to historical mismatch values or regret values. , , To learn the control parameters.