Multi-edge node computing resource adaptive scheduling method and system
By constructing a multi-dimensional perception system and machine learning model, the problems of collaborative perception, accurate mapping, dynamic adaptation and cross-node collaborative scheduling in multi-edge node resource scheduling are solved, improving resource utilization and business determinism experience, and reducing system operation and maintenance complexity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TOWER CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-04-28
AI Technical Summary
Existing computing resource scheduling methods for multiple edge nodes have several drawbacks when adapting to deterministic network environments. These include difficulties in coordinating resource heterogeneity among multiple edge nodes with deterministic network assurance capabilities, lack of precise mapping between task-level resource requirements and network deterministic assurance, inability of static scheduling strategies to adapt to dynamic edge environments and business needs, and insufficient cross-node collaborative scheduling and model generalization capabilities.
By constructing a multi-dimensional perception system, a collaborative perception model of node computing power supply, network guarantee capability, and task resource requirements is established. A task-level resource requirement prediction model is established using machine learning methods, and a dynamic adaptive scheduling mechanism is designed to realize cross-node collaborative scheduling and knowledge transfer, thereby improving the generalization ability of the model.
It achieves collaborative perception and accurate mapping of resources across multiple edge nodes, possesses dynamic adaptive capabilities, improves resource utilization and business determinism, and reduces system operation and maintenance complexity and model training costs.
Smart Images

Figure CN121597428B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computing power networks and edge computing technology, specifically to an adaptive scheduling method and system for computing power resources of multiple edge nodes. Background Technology
[0002] With the trend of computing power networks extending from traditional cloud computing centers to edge computing nodes, deterministic networks, with their unique resource reservation and traffic shaping mechanisms, provide bounded latency, low jitter, and reliable bandwidth transmission guarantees for mission-critical services. However, existing computing resource scheduling methods for multiple edge nodes still face the following prominent challenges when adapting to deterministic network environments:
[0003] 1. The heterogeneity of resources among multiple edge nodes and the deterministic network guarantee capabilities are difficult to coordinate.
[0004] Existing scheduling methods are typically designed for homogeneous cloud data center environments and fail to effectively address the significant differences in computing power, storage performance, and network interfaces among edge nodes. In deterministic network environments, the guarantee capabilities (such as latency upper bound and available bandwidth lower bound) of network slices accessed by different edge nodes vary, and traditional schedulers lack awareness of the joint state of "node computing power characteristics - network guarantee capabilities," making it difficult to fully leverage the guarantee performance of deterministic networks.
[0005] 2. There is a lack of precise mapping between task-level resource requirements and network deterministic guarantees.
[0006] While deterministic networks offer reliable performance parameters, existing methods remain at the node or flow level with coarse-grained resource allocation. Their monitoring systems cannot establish a dynamic correlation model between computational resource consumption profiles at the "task cluster" granularity and deterministic network performance parameters. This often leads to overly conservative resource allocation, reserving far more resources than actual needs, resulting in low resource utilization; or overly aggressive allocation, failing to fully deliver on network performance promises and impacting the deterministic experience of services.
[0007] 3. Static scheduling strategies are difficult to adapt to dynamic edge environments and business needs.
[0008] Deterministic network assurance capabilities may dynamically adjust due to changes in network load or slice reconfiguration, while the available resource status of edge nodes and the load of connected services are also highly time-varying. Most existing scheduling schemes rely on preset static strategies, lacking the ability to adaptively adjust online based on changes in network performance, fluctuations in node resources, and shifts in service demands. This rigid scheduling mechanism cannot meet the dual dynamic optimization requirements of resource efficiency and deterministic assurance in edge scenarios, thus limiting overall system performance.
[0009] 4. Insufficient cross-node collaborative scheduling and model generalization capabilities.
[0010] In multi-edge node scenarios, tasks may require collaborative processing across nodes. Traditional single-node scheduling models cannot effectively coordinate the resource status and network assurance capabilities of multiple heterogeneous edge nodes to make globally optimal decisions. Furthermore, scheduling models trained for specific nodes or task types perform poorly on other nodes or task types, lacking the necessary generalization ability, increasing system operation and maintenance complexity, and limiting the large-scale application of scheduling solutions.
[0011] In summary, existing multi-edge node computing resource scheduling methods for deterministic networks have significant shortcomings in resource collaborative awareness, accurate mapping, dynamic adaptation, and cross-node generalization. Therefore, there is an urgent need for an adaptive scheduling method that can deeply integrate the characteristics of deterministic networks, support multi-edge node resource collaboration, and possess online learning capabilities, in order to maximize resource utilization efficiency while ensuring the quality of business services. Summary of the Invention
[0012] To address the aforementioned issues, this application provides an adaptive scheduling method and system for computing resources across multiple edge nodes. This overcomes the inherent limitations of traditional scheduling methods in environmental awareness, decision optimization, and system adaptation. Furthermore, it overcomes the following four key technical obstacles:
[0013] 1. Solve the challenge of collaborative perception between heterogeneous resources of multiple edge nodes and deterministic network guarantee capabilities.
[0014] Existing scheduling methods lack a unified framework for understanding the heterogeneous characteristics of edge computing nodes and the guarantee capabilities of deterministic networks. Different edge nodes exhibit significant differences in computing architecture (e.g., CPU / GPU), storage hierarchy, and network interfaces, while the performance parameters provided by deterministic networks, such as latency upper bounds and guaranteed bandwidth, also vary for each node. Traditional methods cannot establish a collaborative awareness model between "node computing power supply, network guarantee capabilities, and task resource requirements," resulting in an inability to adaptively generate appropriate computing resource allocation schemes based on task requirements and network performance. This application constructs a multi-dimensional awareness system capable of simultaneously perceiving the heterogeneous computing power characteristics of nodes and the deterministic guarantee capabilities of the network, providing a joint optimization foundation for precise scheduling.
[0015] 2. Break through the bottleneck of precise mapping between task-level resource requirements and network deterministic assurance.
[0016] While deterministic networks provide reliable performance guarantees, existing methods lack a precise mapping mechanism to translate this network-level guarantee capability into task-level resource allocation strategies. Especially in multi-edge node environments, when the same task is executed on different nodes, its resource requirements dynamically change due to differences in node characteristics and network guarantee capabilities. Traditional resource estimation methods based on fixed empirical values cannot adapt to this complex mapping relationship, leading to resource allocation that is either too conservative, resulting in waste, or too aggressive, affecting the achievement of expected task performance. This application establishes a task-level resource requirement prediction model that can dynamically adapt to node characteristics and network guarantees, achieving precise resource allocation.
[0017] 3. Overcome the adaptability gap between static scheduling strategies and dynamic edge environments.
[0018] The reliability of deterministic networks dynamically adjusts with network load and slice configuration, while the resource status and service load of edge nodes are also highly time-varying. Existing static scheduling strategies cannot adapt to this dual dynamic change and often fail rapidly after environmental changes, requiring manual intervention and reconfiguration. This lagging response mechanism makes it difficult to guarantee service continuity and stability. This application designs a dynamic scheduling mechanism that can adapt online based on different network performance, node resource fluctuations, and changes in service requirements, ensuring the system's continuous optimization capability in dynamic environments.
[0019] 4. Establish a global optimization mechanism for cross-node collaborative scheduling and knowledge transfer.
[0020] In multi-edge node scenarios, tasks may need to be executed collaboratively across multiple nodes, but the resource status and network assurance capabilities of different nodes vary. Traditional single-node scheduling models lack cross-node collaborative decision-making capabilities and cannot achieve global resource optimization. Furthermore, scheduling models trained for specific nodes are difficult to directly transfer to other nodes, resulting in poor model reusability and high training costs. This application constructs a global optimization framework that supports cross-node collaborative scheduling and knowledge transfer, enabling unified and efficient utilization of resources across multiple edge nodes.
[0021] By addressing the four key technical challenges mentioned above, this application aims to establish an adaptive scheduling method for computing resources across multiple edge nodes. This method achieves a technological leap from single-node local optimization to multi-node collaborative scheduling, from static configuration to dynamic adaptation, and from experience-driven to model-driven approaches. Ultimately, it significantly improves the overall utilization efficiency of computing resources while ensuring the quality of business services. The technical solution adopted in this application is as follows:
[0022] Firstly, this application provides an adaptive scheduling method for computing resources across multiple edge nodes, including:
[0023] Based on business objectives, each task is divided into multiple task clusters, and the basic data of each task cluster for each complete execution of business is obtained. Among them, the tasks included in a task cluster have the same business objectives, and the basic data includes: execution performance indicator input labels, deterministic network performance indicator input labels, and computing resource indicator output labels.
[0024] Based on the basic data, a hybrid training set is constructed, which includes the basic data training set of the target task cluster and the basic data training set of other task clusters. The weight of the basic data training set of the target task cluster is set higher than the weight of the basic data training set of other task clusters. The weighted mean square error is used as the loss function to train the computing resource demand prediction model.
[0025] Obtain the resource task clusters to be allocated, set the execution performance indicators of the resource task clusters to be allocated, and obtain the deterministic network performance indicators of the resource task clusters to be allocated; wherein, the resource task clusters to be allocated and the target task clusters have the same business objectives.
[0026] A prediction input vector is constructed based on execution performance metrics and deterministic network performance metrics. This prediction input vector is then fed into a computing resource demand prediction model. Based on this model, a prediction output vector of computing resource metric allocation values is output.
[0027] Secondly, this application also provides a multi-edge node computing resource adaptive scheduling system, including:
[0028] The historical data acquisition unit is used to divide each task into multiple task clusters according to business objectives and acquire the basic data of each task cluster for each complete execution of business. Among them, the tasks included in a task cluster have the same business objectives, and the basic data includes: execution performance indicator input labels, deterministic network performance indicator input labels, and computing resource indicator output labels.
[0029] The model training unit is used to construct a mixed training set based on basic data, including the basic data training set of the target task cluster and the basic data training set of other task clusters. The weight of the basic data training set of the target task cluster is set higher than the weight of the basic data training set of other task clusters. The weighted mean square error is used as the loss function to train the computing resource demand prediction model.
[0030] The incremental data acquisition unit is used to acquire resource task clusters to be allocated, set the execution performance indicators of the resource task clusters to be allocated, and acquire the deterministic network performance indicators of the resource task clusters to be allocated; wherein the resource task clusters to be allocated and the target task clusters have the same business objectives.
[0031] The resource scheduling unit is used to construct a prediction input vector based on execution performance indicators and deterministic network performance indicators, input the prediction input vector into the computing resource demand prediction model, and output a prediction output vector of computing resource indicator allocation values based on the computing resource demand prediction model.
[0032] Thirdly, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described adaptive scheduling method for multi-edge node computing resources.
[0033] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when instructed by a processor, implements the steps of the above-described adaptive scheduling method for multi-edge node computing resources.
[0034] The above-mentioned technical solution adopted in this application can achieve the following beneficial effects:
[0035] This application achieves collaborative perception of heterogeneous resources of multiple edge nodes and deterministic network assurance capabilities: By collecting deterministic network performance data, execution performance data and computing resource data of task clusters, this application establishes a collaborative perception model between "node computing power supply - network assurance capabilities - task resource requirements", which effectively solves the problem of the lack of a unified cognitive framework in existing methods and can give full play to the assurance effectiveness of deterministic networks.
[0036] A precise mapping relationship between task-level resource requirements and network deterministic assurance has been established: This application constructs a computing resource requirement prediction model based on machine learning methods. By using a hybrid training set and a weighted training strategy, it accurately captures the complex mapping relationship between task performance indicators, network performance indicators and resource requirements, thereby achieving precise allocation of task-level resources. This avoids the problem of overly conservative or aggressive resource allocation in traditional methods, and significantly improves resource utilization and business deterministic experience.
[0037] Possessing dynamic adaptive capabilities to adapt to dynamic changes in edge environments and business needs: The model continuous learning mechanism designed in this application can continuously optimize the computing resource demand prediction model based on the execution data of new task clusters, enabling the scheduling method to adapt to the dynamic adjustment of deterministic network assurance capabilities, fluctuations in edge node resource status, and changes in business needs. This solves the problem of rigidity in traditional static scheduling strategies and ensures the system's continuous optimization capabilities and business continuity in dynamic environments.
[0038] Improved cross-node collaborative scheduling and model generalization capabilities: The hybrid training set design of this application integrates historical data from the target task cluster and other task clusters, enabling the trained model to be applicable not only to specific nodes and task types, but also to other nodes and task types, thereby improving the model's generalization capabilities and reducing system operation and maintenance complexity and model training costs. Attached Figure Description
[0039] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0040] Figure 1 A flowchart illustrating an adaptive scheduling method for multi-edge node computing resources according to an embodiment of this application is shown.
[0041] Figure 2 A schematic diagram of the structure of a multi-edge node computing resource adaptive scheduling system according to an embodiment of this application is shown;
[0042] Figure 3 A schematic diagram of the resulting electronic device according to an embodiment of this application is shown. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0044] Figure 1 A flowchart illustrating an adaptive scheduling method for multi-edge node computing resources according to an embodiment of this application is shown. (Refer to...) Figure 1 As shown, this embodiment includes steps S110 to S140:
[0045] Step S110: Divide each task into multiple task clusters according to the business objectives, and obtain the basic data of each task cluster for each complete execution of the business; wherein, each task in a task cluster has the same business objective, and the basic data includes: execution performance indicator input labels, deterministic network performance indicator input labels, and computing resource indicator output labels.
[0046] This step establishes a task-level monitoring system to enable real-time collection and calculation of various indicators for each task cluster during each business execution.
[0047] A "task cluster" is a logical grouping of tasks with a common business objective, defined to achieve precise resource scheduling. It is the basic unit for multi-dimensional perception and demand forecasting in this application. Specifically, it involves determining the business objective corresponding to each task and grouping tasks with the same business objective into the same task cluster.
[0048] The criteria for determining business objectives can be that the core functions of the task are consistent with the service targets.
[0049] For example, in an industrial internet scenario, "real-time industrial data analysis" is a business objective, and the corresponding tasks include "equipment temperature data acquisition," "equipment vibration data filtering," "fault feature extraction," and "anomaly warning generation." These tasks all revolve around the core function of "real-time monitoring of industrial equipment operating status" and serve factory maintenance personnel, thus they are grouped into the same task cluster.
[0050] For example, in a smart city scenario, "dynamic traffic flow control" is a business objective, and the corresponding tasks include "intersection camera data collection," "traffic flow statistics," "congestion assessment," and "traffic light duration adjustment instruction generation." These tasks all revolve around the core function of "optimizing road traffic efficiency" and serve the traffic management platform, thus they are grouped into the same task cluster.
[0051] After determining the business objectives corresponding to each task, a business objective label can be preset for each task (e.g., "real-time analysis of industrial data", "dynamic control of traffic flow", etc.). Tasks with the same label will be automatically classified into the same task cluster, and a unique index will be assigned to each task cluster. .
[0052] For each task cluster, we monitor and calculate its basic data throughout the entire lifecycle of each complete execution, from start to finish, providing data input for training subsequent computing resource demand prediction models. The basic data includes: execution performance indicator input labels, deterministic network performance indicator input labels, and computing resource indicator output labels.
[0053] First, obtain the performance metrics input labels for each task cluster's complete execution of the service; among them, the performance metrics input labels include: average throughput. Average completion time Timeout rate and success rate The performance metrics input labels are obtained by parsing the task scheduler's operation logs and real-time monitoring calculations.
[0054] By combining task scheduler log analysis with real-time monitoring, a complete execution trajectory of each task cluster's full execution is constructed, forming input labels for execution performance metrics.
[0055] A combined approach of task scheduler log analysis and real-time status monitoring is employed to write business execution-related metrics into the task logs. Specifically, on one hand, key events in the business lifecycle are extracted by parsing the operation logs of the task scheduler (such as Kubernetes Job Controller or YARN Resource Manager) and written to the task logs; on the other hand, business status changes are obtained in real time through the scheduler API to monitor business execution progress and write it to the task logs; finally, the expected deadline, submission records, completion time, and completion status records of the tasks are correlated to construct a complete execution trajectory.
[0056] Using the above method, the task log records will include at least the expected deadline, start time, end time, and whether the execution was successful for each task in a complete execution of the business in the task cluster, as well as the total number of tasks in the task cluster, the earliest start time of the task cluster, and the latest end time of the task cluster.
[0057] That is, for any task cluster The task cluster is obtained by parsing the task scheduler's operation logs and real-time monitoring. Total number of tasks in a complete business execution Earliest start time for all tasks (Right now Latest end time for all tasks (Right now ), start time of each task The end time of each task Expected deadline for each task Successful execution flag for each task ;in, Indicates the task index and .
[0058] For any task cluster The following calculations yield four categories of indicators.
[0059] Task Cluster Average throughput of a complete business execution The average number of tasks successfully completed per unit of time (unit: tasks / minute) reflects the system's processing capacity.
[0060] Based on the total number of tasks Earliest start time for all tasks and the latest end time of all tasks The average throughput is calculated based on the following formula (1). :
[0061] , formula (1).
[0062] Task Cluster Average completion time of a complete business operation : The average actual completion time of the task (unit: seconds), reflecting the system processing speed.
[0063] Based on the total number of tasks The start time of each task and the end time of each task The average completion time is calculated based on the following formula (2). :
[0064] , formula (2).
[0065] Task Cluster Timeout rate for a complete business execution The percentage of tasks completed beyond the scheduled deadline (in %) reflects the system's ability to ensure timeliness.
[0066] Based on the total number of tasks The start time of each task The end time of each task and the expected deadline for each task The timeout rate is calculated based on the following formula (3). :
[0067] ,in, , formula (3).
[0068] Task Cluster Success rate of a complete business execution The percentage of successfully completed tasks out of the total number of submitted tasks (in %) reflects the system's reliability.
[0069] Based on the total number of tasks and a success flag for each task. The success rate is calculated based on the following formula (4). :
[0070] ,in, (The value is 1 if the execution is successful and 0 if the execution fails), formula (4).
[0071] Each complete execution of a task cluster obtains performance metric input tags using the same method.
[0072] Second, obtain the deterministic network performance indicator input labels for each task cluster's complete execution of the service; among them, the deterministic network performance indicator input labels include: lower limit of network bandwidth. Network latency limit upper limit of packet loss rate and network jitter limit The input labels for deterministic network performance metrics are obtained by querying the northbound API of the deterministic network controller.
[0073] Deterministic networking is a novel network architecture that provides bounded latency, extremely low jitter, and bandwidth guarantees for specific data flows through resource reservation, traffic shaping, and global scheduling techniques. Its core feature lies in transforming the traditional "best-effort" transmission mode based on statistical multiplexing into a "plan-first" deterministic quality of service (QoS) guarantee. With deterministic networking, network performance does not require real-time monitoring and statistical methods because the architecture pre-establishes strict transmission guarantee contracts for each service flow through the control plane. Network performance parameters (such as latency upper bounds, jitter range, and guaranteed bandwidth) are determined through contracts during the flow establishment phase. These metrics are expressed as pre-defined SLA (Service Level Agreement) commitment values rather than dynamic fluctuations. Therefore, these guaranteed values can be directly obtained by querying the network controller's northbound API (Application Programming Interface), which is more accurate and efficient than obtaining statistical averages through passive measurement, and completely avoids the overhead and uncertainty inherent in measurement itself.
[0074] In any task cluster At the start of a complete business execution, deterministic network performance metrics input tags during the execution process can be obtained by querying the network controller's northbound API. These specifically include the following four categories.
[0075] Task Cluster Minimum network bandwidth during a complete business process The minimum data transmission rate (in Mbps) supported by a network interface in a deterministic network reflects the lower limit of the network's actual data transmission capacity.
[0076] Task Cluster Network latency limit during a complete business process The maximum time (in milliseconds) required for a data packet to travel from the source address to the destination address and back in a deterministic network directly affects user experience and business real-time performance.
[0077] Task Cluster Maximum packet loss rate during a complete business process The maximum percentage (in %) of data packets lost during deterministic network transmission represents the reliability of the network.
[0078] Task Cluster Network jitter limit during a complete business execution The maximum value of the fluctuation in the transmission delay of consecutive data packets in a deterministic network (in milliseconds), which characterizes the determinism of the network.
[0079] Each complete execution of a task cluster obtains deterministic network performance metric input labels using the same method.
[0080] Third, obtain the computing resource metric output labels for each task cluster's complete execution of the business; among them, the computing resource metric output labels include: average CPU core utilization. Average memory usage Average GPU memory usage and average GPU core usage The output label of the resource index is calculated by collecting instantaneous indexes at fixed sampling intervals and then combining the number of samplings and the attenuation factor.
[0081] The system monitors and statistically analyzes the computing resource usage of each task cluster during each complete business execution process. It collects and calculates metrics through a task-level monitoring agent, generating computing resource metric output labels.
[0082] For any task cluster In the task cluster During a complete execution of a business process, at a fixed sampling interval (e.g., 100ms) for this task cluster The following four types of instantaneous values are collected during a complete execution of a business process.
[0083] Task Cluster CPU instantaneous usage during a complete business execution : For task clusters An independent monitoring container monitors resources using Cgroups (control groups, the Linux kernel's resource management mechanism). For the first Task clusters collected at each sampling interval The instantaneous CPU usage during a single complete execution of a business function.
[0084] Task Cluster Instantaneous memory usage during a complete business execution To monitor task-level memory usage using Linux's perf tool, For the first Task clusters collected at each sampling interval The instantaneous memory usage during a single complete business operation.
[0085] Task Cluster Instantaneous GPU memory usage during a complete business execution and GPU core instantaneous usage Monitor task clusters using GPU management interfaces (including but not limited to NVIDIA-SMI, etc.). GPU usage, then , The first Task clusters collected at each sampling interval Instantaneous GPU memory usage and GPU core usage during a single complete business execution.
[0086] Furthermore, obtain the task cluster. Number of data collections in a single complete business execution .
[0087] Based on instantaneous CPU usage Instantaneous memory usage GPU memory instantaneous usage GPU core instantaneous usage Number of times and attenuation factor ( (Pre-set), the average CPU core usage is calculated using the exponential moving average through the following formulas (5) to (8). Average memory usage Average GPU memory usage and average GPU core usage :
[0088] , formula (5);
[0089] , formula (6);
[0090] , formula (7);
[0091] , formula (8).
[0092] Task Cluster Average CPU core usage during a complete business execution Task cluster Average number of CPU cores used during execution (unit: cores).
[0093] Task Cluster Average memory usage per complete business execution ( ): Task Cluster Average memory usage during execution (in MB).
[0094] Task Cluster Average GPU memory usage during a complete business execution ( ): Task Cluster Average GPU memory usage during execution (in MB).
[0095] Task Cluster Average GPU core usage during a complete business execution ( ): Task Cluster Average GPU core usage during execution (in units).
[0096] Each complete execution of a task cluster obtains computing resource indicator output tags using the same method.
[0097] Step S120: Construct a hybrid training set based on the basic data, which includes the basic data training set of the target task cluster and the basic data training sets of other task clusters. Set the weight of the basic data training set of the target task cluster to be higher than the weight of the basic data training sets of other task clusters. Use the weighted mean square error as the loss function to train the computing power resource demand prediction model.
[0098] A computing resource demand prediction model is trained based on fundamental data. Machine learning methods are used to establish a mapping relationship between the input labels of task cluster execution performance metrics and deterministic network performance metrics, and the output labels of computing resource metrics. The computing resource demand model employs a weighted training strategy, prioritizing the learning of fundamental data from the current edge nodes and the target task cluster, while also considering fundamental data from other edge nodes or other task clusters, thus achieving accurate resource demand prediction.
[0099] First, construct a hybrid training set.
[0100] For any task cluster, construct an input feature vector based on the execution performance index input label and the deterministic network performance index input label of the task cluster, and construct an output feature vector based on the computational resource index output label of the task cluster.
[0101] Specifically, for any task cluster A complete execution of the business process constructs an 8-dimensional input feature vector. Construct 4 output feature vectors .
[0102] The input and output feature vectors corresponding to the target task cluster are used to form the basic data training set for the target task cluster; the input and output feature vectors corresponding to other task clusters are used to form the basic data training set for other task clusters.
[0103] Mixed training set It includes two parts:
[0104] The first part is the target task cluster. The basic data training set of the target task cluster, consisting of the corresponding input feature vectors and output feature vectors. In the basic training set of the target task cluster, the number of samples is the total number of times each complete execution of the business process of the target task cluster is performed.
[0105] The second part is other task clusters. The corresponding input and output feature vectors form the basic data training set for other task clusters. In the training set of the basic data for other task clusters, the number of samples is the total number of times each complete execution of the business process is performed in the other task clusters.
[0106] In other words, mixed training sets .
[0107] Then, a computing resource demand prediction model is trained.
[0108] The input-output relationship of the computing resource demand forecasting model can be expressed as follows: ;in, The input-output mapping relationship of the computing resource demand prediction model. These are the parameter values for the computing resource demand forecasting model. To predict the input vector; Given parameter values and a prediction input vector, this is the prediction output vector of the computing power resource demand prediction model.
[0109] To train the computing resource demand prediction model, the sample weights of the basic data training set for the target task cluster are set as follows: The sample weights of the basic training set for other task clusters are: ;in, The weights are pre-set. By differentiating the weights, the model pays more attention to the basic data of the target task cluster during training, thereby ensuring the accuracy of the model's prediction of the resource requirements of the target task cluster. At the same time, the generalization ability of the model is improved by using auxiliary training data from other task clusters.
[0110] Obtain the base model. The base model can be a deep learning model, including but not limited to GRU (Gated Recurrent Unit) and TCN (Temporal Convolutional Network), or a traditional machine learning algorithm, such as XGBoost and Random Forest. The specific model selection should be determined based on the data characteristics and real-time requirements of the actual business scenario.
[0111] Construct a weighted mean squared error loss function; where the expression of the weighted mean squared error loss function is as follows (9):
[0112] , formula (9);
[0113] in, Indicates the model parameter values. This represents the number of samples in the basic training set for the target task cluster. , The first data in the training set representing the basic data of the target task cluster. The input feature vector of each sample, Given parameter values and the base data training set of the target task cluster, the first... The predicted output of the computing resource demand prediction model given the input feature vector of a sample. The first data in the training set representing the basic data of the target task cluster. The output feature vector of each sample This represents the number of samples in the training set of the basic data for other task clusters. , The first data in the training set representing the basic data of other task clusters. The input feature vector of each sample, Given parameter values and the training set of the basic data for other task clusters, the first... The predicted output of the computing resource demand prediction model given the input feature vector of a sample. The first data in the training set representing the basic data of other task clusters. The output feature vector of each sample.
[0114] The basic model is trained iteratively using the gradient descent method; the iterative formula is as follows (10):
[0115] , formula (10);
[0116] in, , This represents the gradient of the weighted mean squared error loss function;
[0117] When the number of iterations reaches the preset maximum value Or the absolute value of the weighted mean squared error loss function is less than a preset threshold. The iteration stops at a certain point, resulting in a fully trained computing resource demand prediction model. (positive integers) and (Positive numbers) Preset. The parameter values of the trained computing resource demand prediction model are... .
[0118] Step S130: Obtain the resource task cluster to be allocated, set the execution performance index of the resource task cluster to be allocated, and obtain the deterministic network performance index of the resource task cluster to be allocated.
[0119] The trained computing resource demand prediction model is applied to actual resource scheduling scenarios to predict the resource demand of newly submitted task clusters to be allocated.
[0120] Obtain resource task clusters that have the same business objectives as the target task cluster.
[0121] The cluster of resource tasks to be allocated can be represented as Cluster of tasks awaiting allocation of resources With target task cluster They share the same business objectives.
[0122] Define the execution performance metrics for a single complete execution of the resource cluster to be allocated; these performance metrics include: target average throughput. Average time to complete the target Timeout rate and success rate The above four types of execution performance metrics support users in managing clusters of tasks awaiting resource allocation. The process will be dynamically adjusted during execution.
[0123] The deterministic network performance metrics for a single complete execution of a resource cluster to be allocated are obtained by querying the northbound API of the deterministic network controller. These deterministic network performance metrics include: the target network bandwidth lower limit. Target network latency limit Target packet loss rate upper limit and target network jitter limit .
[0124] In the cluster of tasks awaiting allocation of resources At the start of a complete business operation, the cluster of resource tasks to be allocated can be obtained by querying the northbound API of the network controller. Deterministic network performance metrics during execution. Specifically, these include: target network bandwidth minimum. Target network latency limit Target packet loss rate upper limit and target network jitter limit .
[0125] Step S140: Construct a prediction input vector based on execution performance indicators and deterministic network performance indicators, input the prediction input vector into the computing resource demand prediction model, and output the prediction output vector of computing resource indicator allocation values based on the computing resource demand prediction model.
[0126] Construct a prediction input vector based on the execution performance metrics and deterministic network performance metrics of a complete execution of a task cluster to be allocated resources.
[0127] For the cluster of tasks to be allocated resources A complete execution of the business logic constructs an 8-dimensional prediction input vector. .
[0128] The predicted input vector is fed into the computing resource demand prediction model. Based on the computing resource demand prediction model, the predicted output vector of the computing resource index allocation value of a complete execution of a business by a cluster of tasks to be allocated is output.
[0129] By inputting the predicted input vector into the computing resource demand prediction model, one can obtain the following based on the computing resource demand prediction model: .
[0130] The allocated values for computing resource metrics include: average utilization of target CPU cores. Target average memory usage Average usage of target GPU memory and average usage of target GPU cores .Right now .
[0131] Users can then use the predicted output vector from the computing resource demand prediction model to allocate resources to task clusters. Allocate computing resources.
[0132] In addition, the above method also includes: after the resource task cluster to be allocated is completed, obtaining the basic data of the completed resource task cluster; incrementally updating the hybrid training set based on the basic data of the completed resource task cluster, and using the updated hybrid training set to incrementally update the computing power resource demand prediction model.
[0133] After the task clusters to be allocated resources are completed, the computing resource demand prediction model is continuously optimized based on the actual execution effect of the task clusters to be allocated resources.
[0134] Due to the unassigned resource task cluster With target task cluster Because they share the same business objectives, resource task clusters to be allocated are recorded. Based on the performance of the computing resource allocation values predicted by the computing resource demand forecasting model, the clusters of tasks to be allocated can be obtained. The basic data from a complete business execution is used to build incremental training data. .
[0135] Incremental data With target task cluster Basic training data set Merging, i.e. This allows for the updating of the basic training data set. training sets of basic data from other task clusters Constructing an updated hybrid training set .
[0136] Utilizing updated mixed training sets Incrementally update the computing resource demand forecasting model and continuously optimize its parameter values. And continue to be used for subsequent tasks related to the target cluster. Resource requirement forecasting for task clusters with the same business objectives.
[0137] Figure 2 An adaptive scheduling system for multi-edge node computing resources according to an embodiment of this application is illustrated. (Refer to...) Figure 2 As shown, the multi-edge node computing resource adaptive scheduling system 200 includes:
[0138] The historical data acquisition unit 210 is used to divide each task into multiple task clusters according to the business objectives and acquire the basic data of each task cluster for each complete execution of the business; wherein, each task in a task cluster has the same business objective, and the basic data includes: execution performance indicator input labels, deterministic network performance indicator input labels, and computing resource indicator output labels;
[0139] The model training unit 220 is used to construct a mixed training set based on basic data, including the basic data training set of the target task cluster and the basic data training set of other task clusters. The weight of the basic data training set of the target task cluster is set higher than the weight of the basic data training set of other task clusters. The weighted mean square error is used as the loss function to train the computing power resource demand prediction model.
[0140] The incremental data acquisition unit 230 is used to acquire the resource task cluster to be allocated, set the execution performance index of the resource task cluster to be allocated, and acquire the deterministic network performance index of the resource task cluster to be allocated; wherein the resource task cluster to be allocated and the target task cluster have the same business objectives.
[0141] The resource scheduling unit 240 is used to construct a prediction input vector based on execution performance indicators and deterministic network performance indicators, input the prediction input vector into the computing resource demand prediction model, and output a prediction output vector of computing resource indicator allocation values based on the computing resource demand prediction model.
[0142] In some optional implementations, in the above system, the historical data acquisition unit 210 is used to: determine the business objectives corresponding to each task, and group tasks with the same business objectives into the same task cluster; acquire the execution performance indicator input labels for each task cluster's complete execution of the business each time; wherein, the execution performance indicator input labels include: average throughput Average completion time Timeout rate and success rate The performance indicator input labels are obtained by parsing the task scheduler's operation logs and calculating real-time monitoring results; the deterministic network performance indicator input labels for each task cluster's complete execution of the service are obtained; among them, the deterministic network performance indicator input labels include: lower limit of network bandwidth. Network latency limit upper limit of packet loss rate and network jitter limit The input labels for deterministic network performance metrics are obtained by querying the northbound API of the deterministic network controller; the output labels for computational resource metrics for each task cluster after each complete execution of the service are obtained; among them, the output labels for computational resource metrics include: average CPU core utilization. Average memory usage Average GPU memory usage and average GPU core usage The output label of the resource index is calculated by collecting instantaneous indicators at fixed sampling intervals and combining the number of samplings and the attenuation factor; among them, This represents the task cluster index.
[0143] In some optional implementations, in the above system, the historical data acquisition unit 210 is further configured to: target any task cluster The task cluster is obtained by parsing the task scheduler's operation logs and real-time monitoring. Total number of tasks in a complete business execution Earliest start time for all tasks Latest end time for all tasks Start time of each task The end time of each task Expected deadline for each task Successful execution flag for each task ;in, Indicates the task index and Based on the total number of tasks Earliest start time for all tasks and the latest end time of all tasks The average throughput is calculated based on the following formula. : Based on the total number of tasks The start time of each task and the end time of each task The average completion time is calculated based on the following formula. : Based on the total number of tasks The start time of each task The end time of each task and the expected deadline for each task The timeout rate is calculated based on the following formula. : ,in, Based on the total number of tasks and a success flag for each task. The success rate is calculated based on the following formula. : ,in, .
[0144] In some optional implementations, in the above system, the historical data acquisition unit 210 is further configured to: target any task cluster The task cluster is collected at a fixed sampling interval. CPU instantaneous usage during a complete business execution Instantaneous memory usage GPU memory instantaneous usage and GPU core instantaneous usage Obtain the task cluster Number of collections ;in, Indicates the index of the collection point; based on the instantaneous CPU usage. Instantaneous memory usage GPU memory instantaneous usage GPU core instantaneous usage Number of times and attenuation factor The average CPU core usage is calculated using the following formula based on the exponential moving average. Average memory usage Average GPU memory usage and average GPU core usage : , , , ,in, .
[0145] In some optional implementations, in the above system, the model training unit 220 is used to: construct an input feature vector for any task cluster based on the input labels of the execution performance indicators and the input labels of the deterministic network performance indicators for each execution of a complete service by the task cluster; construct an output feature vector based on the output labels of the computational resource indicators for each execution of a complete service by the task cluster; combine the input feature vectors and output feature vectors corresponding to the target task cluster into a basic data training set for the target task cluster; and combine the input feature vectors and output feature vectors corresponding to other task clusters into a basic data training set for other task clusters.
[0146] In some optional implementations, in the above system, the model training unit 220 is further configured to: set the sample weights of the basic data training set of the target task cluster as... The sample weights of the basic training set for other task clusters are: ;in, Obtain the basic model and construct the weighted mean squared error loss function; where the expression for the weighted mean squared error loss function is: ;in, Indicates the model parameter values. Indicates a mixed training set. This represents the number of samples in the basic training set for the target task cluster. Indicates the target task cluster, , The first data in the training set representing the basic data of the target task cluster. The input feature vector of each sample, Given parameter values and the base data training set of the target task cluster, the first... The predicted output of the computing resource demand prediction model given the input feature vector of a sample. This represents the input-output mapping relationship of the computing resource demand forecasting model. The first data in the training set representing the basic data of the target task cluster. The output feature vector of each sample This represents the number of samples in the training set of the basic data for other task clusters. Indicates other task clusters, , The first data in the training set representing the basic data of other task clusters. The input feature vector of each sample, Given parameter values and the training set of the basic data for other task clusters, the first... The predicted output of the computing resource demand prediction model given the input feature vector of a sample. The first data in the training set representing the basic data of other task clusters. The output feature vector of each sample; the base model is trained iteratively using gradient descent; the iterative formula is: ;in, , This represents the gradient of the weighted mean squared error loss function. When the number of iterations reaches the preset maximum value or the absolute value of the weighted mean squared error loss function is less than the preset threshold, the trained computing resource demand prediction model is obtained.
[0147] In some optional implementations, in the above system, the incremental data acquisition unit 230 is used to: acquire resource task clusters to be allocated that have the same business objectives as the target task cluster; and set the execution performance indicators for a single complete execution of the business by the resource task clusters to be allocated; wherein the execution performance indicators include: target average throughput. Average time to complete the target Timeout rate and success rate The deterministic network performance metrics for a single complete execution of a resource cluster to be allocated are obtained by querying the northbound API of the deterministic network controller. These deterministic network performance metrics include: the target network bandwidth lower limit. Target network latency limit Target packet loss rate upper limit and target network jitter limit ;in, This indicates a cluster of tasks awaiting resource allocation.
[0148] In some optional implementations, in the above system, the resource scheduling unit 240 is configured to: construct a prediction input vector based on the execution performance indicators and deterministic network performance indicators of a complete execution of a resource task cluster to be allocated; input the prediction input vector into a computing resource demand prediction model; and, based on the computing resource demand prediction model, output a prediction output vector of the computing resource indicator allocation values for a complete execution of a resource task cluster to be allocated; wherein the computing resource indicator allocation values include: the average utilization of the target CPU cores. Target average memory usage Average usage of target GPU memory and average usage of target GPU cores .
[0149] In some optional implementations, the system further includes a model update unit, configured to: after the resource task cluster to be allocated is completed, obtain the basic data of the completed resource task cluster; incrementally update the hybrid training set based on the basic data of the completed resource task cluster, and incrementally update the computing power resource demand prediction model using the updated hybrid training set.
[0150] It should be noted that the aforementioned multi-edge node computing resource adaptive scheduling system 200 can implement the aforementioned multi-edge node computing resource adaptive scheduling method one by one, which will not be elaborated further.
[0151] Figure 3 This invention illustrates a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 3 As shown, the electronic device includes a processor, internal memory, a network interface, and a non-volatile storage medium connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external devices via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the aforementioned multi-edge node computing resource adaptive scheduling method.
[0152] In one embodiment, the electronic device provided in this application includes a memory and a processor. The memory stores a database and a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the aforementioned adaptive scheduling method for multi-edge node computing resources.
[0153] The above is as stated in this application. Figure 2The method for implementing the multi-edge node computing resource adaptive scheduling system disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by software instructions. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The steps of the method disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0154] In one embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the aforementioned multi-edge node computing resource adaptive scheduling method.
[0155] It should be noted that the functions or steps that the above-mentioned electronic devices or computer-readable storage media can achieve can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0156] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0157] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0158] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for adaptive scheduling of computing resources across multiple edge nodes, characterized in that, include: Based on business objectives, each task is divided into multiple task clusters, and the basic data of each task cluster for each complete execution of business is obtained. Among them, the tasks included in a task cluster have the same business objectives, and the basic data includes: execution performance indicator input labels, deterministic network performance indicator input labels, and computing resource indicator output labels. Based on the basic data, a hybrid training set is constructed, which includes the basic data training set of the target task cluster and the basic data training set of other task clusters. The weight of the basic data training set of the target task cluster is set higher than the weight of the basic data training set of other task clusters. The weighted mean square error is used as the loss function to train the computing resource demand prediction model. Obtain the resource task clusters to be allocated, set the execution performance indicators of the resource task clusters to be allocated, and obtain the deterministic network performance indicators of the resource task clusters to be allocated; wherein, the resource task clusters to be allocated and the target task clusters have the same business objectives. A prediction input vector is constructed based on execution performance indicators and deterministic network performance indicators. The prediction input vector is then input into a computing resource demand prediction model. Based on the computing resource demand prediction model, a prediction output vector of computing resource indicator allocation values is output. Wherein, the weight of the basic data training set of the target task cluster is higher than the weight of the basic data training set of other task clusters, and the weighted mean square error is used as the loss function to train the computing resource demand prediction model, including: The sample weights of the basic training set for the target task cluster are set as follows: The sample weights of the basic training set for other task clusters are: ;in, ; Obtain the base model and construct the weighted mean squared error loss function; the expression for the weighted mean squared error loss function is: ; in, Indicates the model parameter values. Indicates a mixed training set. This represents the number of samples in the basic training set for the target task cluster. Indicates the target task cluster, , The first data in the training set representing the basic data of the target task cluster. The input feature vector of each sample, Given parameter values and the base data training set of the target task cluster, the first... The predicted output of the computing resource demand prediction model given the input feature vector of a sample. This represents the input-output mapping relationship of the computing resource demand forecasting model. The first data in the training set representing the basic data of the target task cluster. The output feature vector of each sample This represents the number of samples in the training set of the basic data for other task clusters. Indicates other task clusters, , The first data in the training set representing the basic data of other task clusters. The input feature vector of each sample, Given parameter values and the training set of the basic data for other task clusters, the first... The predicted output of the computing resource demand prediction model given the input feature vector of a sample. The first data in the training set representing the basic data of other task clusters. The output feature vector of each sample; The base model is trained iteratively using gradient descent; the iterative formula is as follows: ; in, , This represents the gradient of the weighted mean squared error loss function; When the number of iterations reaches the preset maximum value or the absolute value of the weighted mean square error loss function is less than the preset threshold, the trained computing resource demand prediction model is obtained.
2. The method according to claim 1, characterized in that, The process involves dividing each task into multiple task clusters based on business objectives, and obtaining the basic data for each task cluster's complete execution of the business process, including: Determine the business objectives for each task, and group tasks with the same business objectives into the same task cluster; Obtain the performance metrics input labels for each complete execution of the service in each task cluster; among which, the performance metrics input labels include: average throughput. Average completion time Timeout rate and success rate The performance metrics input labels are obtained by parsing the task scheduler's operation logs and real-time monitoring calculations; Obtain the deterministic network performance indicator input labels for each complete execution of the service in each task cluster; among which, the deterministic network performance indicator input labels include: lower limit of network bandwidth. Network latency limit upper limit of packet loss rate and network jitter limit The input labels for deterministic network performance metrics are obtained by querying the northbound API of the deterministic network controller. Obtain the output labels of computing resource metrics for each complete execution of the business process for each task cluster; among which, the output labels of computing resource metrics include: average CPU core utilization. Average memory usage Average GPU memory usage and average GPU core usage The output label of the resource index is calculated by collecting instantaneous indexes at fixed sampling intervals and then combining the number of samplings and the attenuation factor. Among them, superscript This represents the task cluster index.
3. The method according to claim 2, characterized in that, The performance metrics input labels are obtained by parsing the task scheduler's operation logs and calculating real-time monitoring results, including: For any task cluster The task cluster is obtained by parsing the task scheduler's operation logs and real-time monitoring. Total number of tasks in a complete business execution Earliest start time for all tasks Latest end time for all tasks Start time of each task The end time of each task Expected deadline for each task Successful execution flag for each task ;in, Indicates the task index and ; Based on the total number of tasks Earliest start time for all tasks and the latest end time of all tasks The average throughput is calculated based on the following formula. : ; Based on the total number of tasks The start time of each task and the end time of each task The average completion time is calculated based on the following formula. : ; Based on the total number of tasks The start time of each task The end time of each task and the expected deadline for each task The timeout rate is calculated based on the following formula. : ,in, ; Based on the total number of tasks and a success flag for each task. The success rate is calculated based on the following formula. : ,in, .
4. The method according to claim 2, characterized in that, The output label of the computing resource index is obtained by collecting instantaneous indexes at fixed sampling intervals and then combining the number of samplings and the attenuation factor, including: For any task cluster The task cluster is collected at a fixed sampling interval. CPU instantaneous usage during a complete business execution Instantaneous memory usage GPU memory instantaneous usage and GPU core instantaneous usage Obtain the task cluster Number of collections ;in, Indicates the index of the collection point; Based on instantaneous CPU usage Instantaneous memory usage GPU memory instantaneous usage GPU core instantaneous usage Number of times and attenuation factor The average CPU core usage is calculated using the following formula based on the exponential moving average. Average memory usage Average GPU memory usage and average GPU core usage : , , , ,in, .
5. The method according to claim 1, characterized in that, The construction of a hybrid training set based on basic data, including the basic data training set of the target task cluster and the basic data training sets of other task clusters, includes: For any task cluster, construct an input feature vector based on the input labels of the execution performance indicators and the input labels of the deterministic network performance indicators for each execution of the complete service of the task cluster, and construct an output feature vector based on the output labels of the computational resource indicators for each execution of the complete service of the task cluster. The input and output feature vectors corresponding to the target task clusters are used to form the basic data training set for the target task clusters. The input and output feature vectors corresponding to other task clusters are used to form the basic data training set for other task clusters.
6. The method according to claim 1, characterized in that, The steps of obtaining the task clusters to be allocated resources, setting the execution performance metrics of the task clusters to be allocated resources, and obtaining the deterministic network performance metrics of the task clusters to be allocated resources include: Obtain the resource task cluster to be allocated that has the same business objective as the target task cluster; Define the execution performance metrics for a single complete execution of the resource cluster to be allocated; these performance metrics include: target average throughput. Average time to complete the target Timeout rate and success rate ; The deterministic network performance metrics for a single complete execution of a resource cluster to be allocated are obtained by querying the northbound API of the deterministic network controller. These deterministic network performance metrics include: the target network bandwidth lower limit. Target network latency limit Target packet loss rate upper limit and target network jitter limit ; in, This indicates a cluster of tasks awaiting resource allocation.
7. The method according to claim 6, characterized in that, The process of constructing a prediction input vector based on execution performance metrics and deterministic network performance metrics, inputting the prediction input vector into a computing resource demand prediction model, and outputting a prediction output vector of computing resource metric allocation values based on the computing resource demand prediction model includes: Construct a prediction input vector based on the execution performance metrics and deterministic network performance metrics of a complete execution of a task cluster to be allocated resources; Input the predicted input vector into the computing power resource demand prediction model; Based on the computing resource demand prediction model, a predicted output vector of computing resource indicator allocation values for a complete execution of a task cluster to be allocated resources is output. These computing resource indicator allocation values include: the average utilization of the target CPU cores. Target average memory usage Average usage of target GPU memory and average usage of target GPU cores .
8. The method according to claim 1, characterized in that, The method further includes: After the cluster of resource tasks to be allocated is completed, obtain the basic data of the completed cluster of resource tasks to be allocated; The hybrid training set is incrementally updated based on the basic data of the completed clusters of tasks awaiting resource allocation, and the computing resource demand prediction model is updated incrementally using the updated hybrid training set.
9. A multi-edge node computing resource adaptive scheduling system, characterized in that, include: The historical data acquisition unit is used to divide each task into multiple task clusters according to business objectives and acquire the basic data of each task cluster for each complete execution of business. Among them, the tasks included in a task cluster have the same business objectives, and the basic data includes: execution performance indicator input labels, deterministic network performance indicator input labels, and computing resource indicator output labels. The model training unit is used to construct a mixed training set based on basic data, including the basic data training set of the target task cluster and the basic data training set of other task clusters. The weight of the basic data training set of the target task cluster is set higher than the weight of the basic data training set of other task clusters. The weighted mean square error is used as the loss function to train the computing resource demand prediction model. The incremental data acquisition unit is used to acquire resource task clusters to be allocated, set the execution performance indicators of the resource task clusters to be allocated, and acquire the deterministic network performance indicators of the resource task clusters to be allocated; wherein the resource task clusters to be allocated and the target task clusters have the same business objectives. The resource scheduling unit is used to construct a prediction input vector based on execution performance indicators and deterministic network performance indicators, input the prediction input vector into the computing resource demand prediction model, and output the prediction output vector of computing resource indicator allocation values based on the computing resource demand prediction model. Specifically, the model training unit is used for: The sample weights of the basic training set for the target task cluster are set as follows: The sample weights of the basic training set for other task clusters are: ;in, ; Obtain the base model and construct the weighted mean squared error loss function; the expression for the weighted mean squared error loss function is: ; in, Indicates the model parameter values. Indicates a mixed training set. This represents the number of samples in the basic training set for the target task cluster. Indicates the target task cluster, , The first data in the training set representing the basic data of the target task cluster. The input feature vector of each sample, Given parameter values and the base data training set of the target task cluster, the first... The predicted output of the computing resource demand prediction model given the input feature vector of a sample. This represents the input-output mapping relationship of the computing resource demand forecasting model. The first data in the training set representing the basic data of the target task cluster. The output feature vector of each sample This represents the number of samples in the training set of the basic data for other task clusters. Indicates other task clusters, , The first data in the training set representing the basic data of other task clusters. The input feature vector of each sample, Given parameter values and the training set of the basic data for other task clusters, the first... The predicted output of the computing resource demand prediction model given the input feature vector of a sample. The first data in the training set representing the basic data of other task clusters. The output feature vector of each sample; The base model is trained iteratively using gradient descent; the iterative formula is as follows: ; in, , This represents the gradient of the weighted mean squared error loss function; When the number of iterations reaches the preset maximum value or the absolute value of the weighted mean square error loss function is less than the preset threshold, the trained computing resource demand prediction model is obtained.
Citation Information
Patent Citations
Distributed heterogeneous node optimization method and system
CN120455463A
Distributed energy intelligent configuration method and device based on Internet of Things big data
CN121052961A