Federal mechanism-based computing power resource dynamic migration scheduling method and system
Through the dynamic migration and scheduling method of computing power resources based on the federated mechanism, the Karmada architecture and optimization algorithm are used to solve the problem of cross-cluster scheduling of computing power resources, and the efficient utilization of computing power resources and the improvement of user task processing efficiency are achieved.
Patent Information
- Application Number
- CN202510256174.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology cannot realize cross-cluster scheduling of computing power resources, resulting in insufficient or oversupply of resources and inability to fully and effectively utilize them.
The dynamic migration and scheduling method of computing power resources based on the federated mechanism is adopted, and the computing power resources connected and resource unified in the computing power nodes of the satellite subcluster are realized through the Karmada architecture, and the computing power optimization model is established, and the computing power scheduling algorithm based on the earliest task overall completion time and resource load balancing degree is solved to determine the task offload plan.
It realizes online dynamic migration and cross-cluster scheduling of computing power resources, greatly improving the utilization rate of computing power resources and improving the processing efficiency of user computing task requests.
Smart Images

Figure CN120335982A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent computing technology, and particularly relates to a method and system for dynamically migrating and scheduling computing power resources based on a federation mechanism. Background Art
[0002] Introducing cloud computing into the satellite Internet can effectively support data processing-intensive services such as image preprocessing, water area extraction, and target recognition for space-based and air-based users. Whether it is the on-board computing power resources in the satellite segment or the central cloud computing power resources in the ground segment, relevant application services can be deployed. During the concurrent stage of the service, there may be situations of insufficient resources and idle resources between computing power nodes in different segments.
[0003] Currently, computing power resources based on containers can be migrated and scaled within the same cluster, and the resources of all nodes within the container cluster can be fully utilized. However, cross-cluster migration of computing power resources cannot be achieved between multiple data center clusters. If most of the services deployed in the same cluster follow the cosine law or the sine law, then at the same moment of the peak, the use of computing power resources exceeds the scope that the cluster can bear, which will cause service failures. However, at this moment, the utilization rate of computing power resources in other clusters may be relatively low, and a large amount of computing power resources are in an idle state.
[0004] However, related technologies can only meet the scheduling of computing power resources within a single cluster and the dynamic migration of computing power resources, and cannot perform cross-cluster (or cross-domain) scheduling. When the computing power resources in a single cluster are insufficient, dynamic migration and horizontal expansion cannot be achieved, which will cause insufficient computing power resources or an overuse phenomenon in the use of computing power resources in other clusters. It can be seen that related technologies cannot achieve cross-cluster scheduling of computing power resources according to computing requirements, resulting in the inability to fully and effectively utilize computing power resources. Summary of the Invention
[0005] The technical problem to be solved by the present invention is the problem that cross-cluster scheduling of computing power resources cannot be achieved according to computing requirements.
[0006] To solve the above technical problem, the present invention provides a method and system for dynamically migrating and scheduling computing power resources based on a federation mechanism, and specifically adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for dynamically migrating and scheduling computing power resources based on a federated mechanism, including: First, obtain the computing power resource information of a federated computing power cluster, where the federated computing power cluster includes multiple satellite sub-cluster computing power nodes, and the computing power resource information is used to characterize the computing resources, storage resources, energy consumption, network topology, and network status of the satellite sub-cluster computing power nodes. Then, receive the computing task requests of users within a single time slot, and determine the task computing power request information according to the computing task requests. The task computing power request information is used to characterize the service location of the computing task requests, the size of the computing power resource requirements, the task deadline, and the service quality requirements. Next, establish a computing power optimization model based on the computing power resource information and the task computing power request information, and determine a task offloading scheme according to the optimization model. The task offloading scheme includes target computing power nodes, and the target computing power nodes are one or more of the multiple satellite sub-cluster computing power nodes. Finally, schedule the computing task requests to the target computing power nodes for computing according to the task offloading scheme.
[0008] This method can be applied to a multi-cloud management platform based on the Karmada architecture. Using satellite sub-cluster computing power nodes as sub-cluster nodes, the computing power resources of the satellite sub-cluster computing power nodes are interconnected and unified through the Karmada architecture, and are opened to users in the form of services. The offloading problem of users' computing task requests is modeled as an optimization problem, and a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance is used for solution to obtain an optimal task offloading scheme. Finally, based on the determined task offloading scheme, the computing task requests of users are distributed to the corresponding satellite sub-cluster computing power nodes for computing processing. In this way, according to the computing task requests of users, online dynamic migration of computing power resources can be achieved. And it can be cross-cluster scheduled among multiple satellite sub-cluster computing power nodes, greatly improving the utilization rate of computing power resources, making full and effective use of computing power resources, and improving the processing efficiency of users' computing task requests.
[0009] In combination with the first aspect, in an alternative implementation manner, the above-mentioned establishment of a computing power optimization model based on the computing power resource information and the task computing power request information, and determination of a task offloading scheme according to the optimization model, includes: First, perform information modeling on the computing power resource information and the task computing power request information to obtain the modeled computing power resource information and the modeled task computing power request information. Then, establish a computing power optimization model based on the modeled computing power resource information and the modeled task computing power request information. Finally, use a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance to solve the computing power optimization model to obtain a task offloading scheme.
[0010] In combination with the first aspect, in an alternative implementation manner, the expression of the above-mentioned modeled computing power resource information is:
[0011] V = {V1, V2, …, V N}, V n ∈V;
[0012] V n ' = {VC n , VR n , VF n , VE n};
[0013] Among them, V represents the federal computing power cluster, N represents the number of satellite sub-cluster computing power nodes, V n represents the nth satellite sub-cluster computing power node, V n ' represents the computing power resource information after modeling of the satellite sub-cluster computing power node V n , VC n represents the available computing resource size of the satellite sub-cluster computing power node V n , VR n represents the available storage resource size of the satellite sub-cluster computing power node V n , VF n represents the computing power provided by the satellite sub-cluster computing power node V n , VE n represents the upper limit of available energy consumption of the satellite sub-cluster computing power node V n . The expression of the above-mentioned task computing power request information after modeling is:
[0014] U = {U1, U2, …, U M}, U m ∈U;
[0015] U' m = {UV m , UC m , UR m , UF m , UE m , UD m};
[0016] Among them, U represents the set of computing task requests, M represents the number of computing task requests, U m represents the mth computing task request, U' m represents the task computing power request information after modeling of the computing task request U m , UV m represents the access satellite of the computing task request U m , UC m represents the computing resource demand of the computing task request U m , UR m represents the storage resource demand of the computing task request U m , UF mIndicates the computing task volume of computing task request U m UP m Indicates the transmission task volume of computing task request U m UE m Indicates the energy consumption of computing task request U m UD m Indicates the latest allowable completion time of computing task request U m of.
[0017] Combined with the first aspect, in an alternative implementation, the above computing power optimization model is an integer programming model. Specifically, the objective function of the computing power optimization model is:
[0018]
[0019] Among them,
[0020]
[0021] The constraint conditions of the computing power optimization model are:
[0022]
[0023] Among them, ω1 represents the proportion weight of resource load balance in the optimization objective, ω2 represents the proportion weight of the overall task completion time in the optimization objective, the sum of ω1 and ω2 is 1, x m,n Indicates that the computing task request U m is scheduled to the satellite sub-cluster computing power node V n The decision-making size on, x m,n = 1 indicates that the computing task request U m is scheduled to the satellite sub-cluster computing power node V n on, B m,n Indicates that the computing task request U m is scheduled to the satellite sub-cluster computing power node V n The load level of, NB m,n Indicates that the computing task request U m is scheduled to the satellite sub-cluster computing power node V n The normalized load level of, T m,n Indicates that the computing task request U m is scheduled to the satellite sub-cluster computing power node V n The task completion time of, NT m,n Indicates that the computing task request U m is scheduled to the satellite sub-cluster computing power node V n The normalized task completion time of, E n,UVm Indicates the access satellite of computing task request U m and the satellite sub-cluster computing power node Vn The network transmission capacity between, maxB represents B m,n The maximum value in, minB represents B m,n The minimum value in, maxT represents T m,n The maximum value in, minT represents T m,n The minimum value in.
[0024] Combined with the first aspect, in an alternative implementation, the above-mentioned method of solving the computing power optimization model by using the computing power scheduling algorithm based on the earliest overall task completion time and resource load balance includes: using a 1.5-approximate algorithm to solve the computing power optimization model.
[0025] In this implementation, the time complexity of the computing power scheduling algorithm based on the earliest overall task completion time and resource load balance is O(Mlog2N). Compared with the exact algorithms such as branch and bound with exponential time complexity O(2 MN ), this 1.5-approximate algorithm can obtain an optimized solution in a relatively short time. At the same time, compared with greedy algorithms such as first-come-first-served (also known as greedy algorithms), this 1.5-approximate algorithm can effectively and quickly solve problems such as unbalanced resource load and late overall task completion time. And it can combine the algorithms of cross-cluster scheduling and intelligent online scheduling to solve the problem of computing power resource integration and coordination.
[0026] Combined with the first aspect, in an alternative implementation, after scheduling the computing task request to the target computing power node for calculation according to the task offloading scheme, the method further includes: after the calculation task request is calculated and processed at the target computing power node, returning the calculation result corresponding to the calculation task request.
[0027] Combined with the first aspect, in an alternative implementation, after scheduling the computing task request to the target computing power node for calculation according to the task offloading scheme, the method further includes: synchronously updating the computing power resource information based on the federated mechanism of the Karmda architecture.
[0028] Second aspect, the present invention provides a dynamic migration scheduling system for computing power resources based on a federated mechanism. This system is constructed based on the Karmada architecture. Specifically, this system includes: a Monitoring module, an AI scheduling module, and an Operation Management (OperationManager) module. Among them, the Monitoring module is used to obtain the computing power resource information of the federated computing power cluster. The federated computing power cluster includes multiple satellite sub-cluster computing power nodes. The computing power resource information is used to characterize the computing resources, storage resources, energy consumption, network topology, and network status of the satellite sub-cluster computing power nodes. The Monitoring module is also used to receive the computing task requests of users within a single time slot and determine the task computing power request information according to the computing task requests. The task computing power request information is used to characterize the service location, the size of the computing power resource requirements, the task deadline, and the service quality requirements of the computing task requests. The AI scheduling module is used to establish a computing power optimization model according to the computing power resource information and the task computing power request information, and determine the task offloading plan according to the optimization model. The task offloading plan includes target computing power nodes, and the target computing power nodes are one or more of the multiple satellite sub-cluster computing power nodes. The OperationManager module is used to schedule the computing task requests to the target computing power nodes for computing according to the task offloading plan.
[0029] In combination with the second aspect, in an alternative implementation, the AI scheduling module is specifically used for: First, perform information modeling on the computing power resource information and the task computing power request information to obtain the modeled computing power resource information and the modeled task computing power request information. Then, establish a computing power optimization model according to the modeled computing power resource information and the modeled task computing power request information. Finally, use a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance to solve the computing power optimization model to obtain the task offloading plan.
[0030] In combination with the second aspect, in an alternative implementation, this system further includes: a Policy Management (PolicyManager) module. The Policy Management module is used to define scheduling policies, and / or create scheduling policies, and / or update scheduling policies, and / or delete scheduling policies, and / or view scheduling policies. The scheduling policies are used to instruct the AI scheduling module to determine the task offloading plan.
[0031] Third aspect, the present invention provides an electronic device, including: a memory, one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the method provided in the first aspect as described above.
[0032] Fourthly, the present invention provides a computer-readable storage medium, including computer instructions, which, when running on an electronic device, cause the electronic device to execute the method provided in the first aspect as described above.
[0033] It can be understood that the beneficial effects that can be achieved by the dynamic migration scheduling system of computing power resources based on the federation mechanism provided in the second aspect, the electronic device in the third aspect, and the computer-readable storage medium in the fourth aspect can refer to the beneficial effects in the first aspect and any of its possible design manners, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the method for dynamically migrating and scheduling computing power resources based on the federation mechanism provided by an embodiment of the present application;
[0035] Figure 2 It is a flowchart of the method for determining a task offloading scheme provided by an embodiment of the present application;
[0036] Figure 3 It is a schematic structural diagram of the dynamic migration scheduling system of computing power resources based on the federation mechanism provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0038] Introducing cloud computing into satellite Internet can effectively support data processing-intensive services such as image preprocessing, water area extraction, and target recognition for space-based and air-based users. Whether it is the on-board computing power resources in the satellite segment or the central cloud computing power resources in the ground segment, relevant application services can be deployed. During the concurrent stage of the services, there may be situations of insufficient resources and idle resources among the computing power nodes in different segments.
[0039] At present, container-based computing power resources can be migrated and scaled within the same cluster, and the resources of all nodes within the container cluster can be fully utilized. However, cross-cluster migration of computing power resources cannot be achieved between multiple data center clusters. Their businesses are often not static, and thus business access changes over time (e.g., within a day, a month, or even a year). Such changes can be assumed to follow a sine law. Correspondingly, some businesses may exhibit a cosine law. Specifically, when the business access volume is at its peak, the demand for computing power resources is relatively large, and when it is at its trough, the demand for computing power resources is relatively small. If most of the businesses deployed in the same cluster follow the cosine law or the sine law, then at the same moment of the peak, the use of computing power resources exceeds the capacity that the cluster can bear, which will cause business failures. However, at this moment, the utilization rate of computing power resources in other clusters may be relatively low, and a large amount of computing power resources are in an idle state.
[0040] However, related technologies can only meet the scheduling of computing power resources within a single cluster and the dynamic migration of computing power resources, and cannot perform cross-cluster (or cross-domain) scheduling. When the computing power resources of a single cluster are insufficient, dynamic migration and horizontal expansion cannot be achieved, which will result in insufficient computing power resources or an overuse of computing power resources in other clusters, and it is difficult to meet the situation where multiple services have peaks and troughs concurrently. At the same time, computing power fragmentation is relatively serious, resulting in a waste of a large number of small-scale computing powers that cannot be fully utilized. It can be seen that related technologies cannot achieve cross-cluster scheduling of computing power resources according to computing requirements, resulting in the inability to fully and effectively utilize computing power resources.
[0041] To solve the above problems, the embodiments of the present application provide a method and system for dynamically migrating and scheduling computing power resources based on a federation mechanism. The method and system can be applied to a multi-cloud management platform based on the Karmada architecture. Using satellite sub-cluster computing power nodes as sub-cluster nodes, the grid connection and resource unification of the computing power resources of the satellite sub-cluster computing power nodes are realized through the Karmada architecture and are opened to users in the form of services. The offloading problem of the user's computing task request is modeled as an optimization problem, and a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance is used for solving to obtain an optimal task offloading plan. Finally, based on the determined task offloading plan, the user's computing task request is distributed to the corresponding satellite sub-cluster computing power nodes for computing and processing. In this way, according to the user's computing task request, online dynamic migration of computing power resources can be achieved. And cross-cluster scheduling can be performed among multiple satellite sub-cluster computing power nodes, greatly improving the utilization rate of computing power resources and fully and effectively utilizing computing power resources.
[0042] The following introduces the solution provided by the embodiments of the present application in conjunction with the accompanying drawings.
[0043] Specifically,Figure 1 This is a flowchart of the dynamic migration scheduling method for computing power resources based on the federation mechanism provided by the embodiments of the present application. As Figure 1 shown, the dynamic migration scheduling method for computing power resources based on the federation mechanism provided by the embodiments of the present application includes the following steps S101 - S104:
[0044] S101. Obtain the computing power resource information of the federated computing power cluster.
[0045] In the embodiments of the present application, the federated computing power cluster includes multiple satellite sub - cluster computing power nodes, that is, multiple satellite sub - cluster computing power nodes can form a federated computing power cluster according to the cluster federation mechanism. In this way, it is convenient for the integration and full utilization of computing power resources. Specifically, the federated computing power cluster can be constructed based on the Karmada architecture.
[0046] Among them, the computing power resource information can be used to characterize the computing resources, storage resources, energy consumption, network topology, and network status of the satellite sub - cluster computing power nodes in the federated computing power cluster.
[0047] Exemplarily, the computing power resource information may include: the available computing resource size, available storage resource size, provided computing power, and available energy consumption limit of the satellite sub - cluster computing power nodes.
[0048] S102. Receive the computing task requests of users within a single time slot, and determine the task computing power request information according to the computing task requests.
[0049] In the embodiments of the present application, within a single time slot, one or more computing task requests of one user can be received, and multiple computing task requests of multiple users can also be received. The computing task requests may include the content to be calculated, processing requirements for the calculation, and requirements for the computing environment, etc.
[0050] Furthermore, perform task splitting according to the computing task requests to determine the task computing power request information. Among them, the task computing power request information is used to characterize the service location of the computing task requests, the size of the computing power resource requirements, the task deadline, and the service quality requirements.
[0051] Exemplarily, the task computing power request information may include: the access satellite corresponding to the computing task request, the computing resource demand, the storage resource demand, the computing task volume, the transmission task volume, the energy consumption, the allowed latest completion time, etc.
[0052] S103. Establish a computing power optimization model according to the computing power resource information and the task computing power request information, and determine the task offloading scheme according to the optimization model.
[0053] Next, the problem of matching and scheduling the computing power request of the computing task with the computing power resources of the federated computing power cluster can be modeled as an optimization problem based on the computing power resource information and the task computing power request information, that is, a computing power optimization model is established, and the established computing power optimization model is solved to determine the optimal offloading scheme, that is, the task offloading scheme. Among them, the task offloading scheme includes target computing power nodes, and the target computing power nodes are one or more of the computing power nodes of multiple satellite sub-clusters.
[0054] In some embodiments, Figure 2 is a flowchart of a method for determining a task offloading scheme provided by an embodiment of the present application, as Figure 2 shown, S103 may specifically include the following steps S1031-S1032:
[0055] S1031. Perform information modeling on the computing power resource information and the task computing power request information to obtain the modeled computing power resource information and the modeled task computing power request information.
[0056] Specifically, information modeling may be performed according to the information content included in the computing power resource information and the task computing power request information, corresponding expressions are constructed, and the modeled computing power resource information and the modeled task computing power request information are obtained. In this way, it is convenient to establish a computing power optimization model.
[0057] In some embodiments, the expression of the modeled computing power resource information is:
[0058] V = {V1, V2,..., V N}, V n ∈V;
[0059] V n ′ = {VC n , VR n , VF n , VE n};
[0060] Among them, V represents the federated computing power cluster, N represents the number of computing power nodes of the satellite sub-cluster, V n represents the nth computing power node of the satellite sub-cluster, V n ′ represents the modeled computing power resource information of the satellite sub-cluster computing power node V n , VC n represents the available computing resource size of the satellite sub-cluster computing power node V n , VR n represents the available storage resource size of the satellite sub-cluster computing power node V n , VF n represents the computing power provided by the satellite sub-cluster computing power node V n , VE n represents the satellite sub-cluster computing power node Vn The upper limit of available energy consumption.
[0061] The expression of the task computing power request information after modeling is:
[0062] U = {U1, U2,..., U M}, U m ∈U;
[0063] U' m = {UV m , UC m , UR m , UF m , UE m , UD m}.
[0064] Among them, U represents the set of computing task requests, M represents the number of computing task requests, U m represents the mth computing task request, U' m represents the task computing power request information of the computing task request U m after modeling, UV m represents the access satellite of the computing task request U m , UC m represents the computing resource demand of the computing task request U m , UR m represents the storage resource demand of the computing task request U m , UF m represents the computing task volume of the computing task request U m , UP m represents the transmission task volume of the computing task request U m , UE m represents the energy consumption of the computing task request U m , UD m represents the allowable latest completion time of the computing task request U m .
[0065] S1032. Establish a computing power optimization model according to the modeled computing power resource information and the modeled task computing power request information.
[0066] Then, with the system load balancing degree and the overall task completion time as the optimization objectives, establish a computing power optimization model according to the modeled computing power resource information and the modeled task computing power request information. The constraint conditions of this computing power optimization model include: computing resource size, storage resource size, energy consumption upper limit, network traffic balance, task deadline, and workflow task dependency, etc.
[0067] Specifically, schedule the computing task request U m to the satellite sub-cluster computing power node Vn The load level B m,n The expression is:
[0068]
[0069] Let maxB represent all B m,n the maximum value among them, and let minB represent all B m,n the minimum value among them. Then normalize B m,n to obtain. For the computing task request U m schedule it to the satellite sub-cluster computing power node V n The normalized load level NB m,n The expression is:
[0070]
[0071] For the computing task request U m schedule it to the satellite sub-cluster computing power node V n The task completion time T m,n The expression is:
[0072]
[0073] Among them, E n,UVm represents the network transmission capacity between the access satellite of the computing task request U m and the satellite sub-cluster computing power node V n
[0074] Let maxT represent all T m,n the maximum value among them, and minT represent all T m,n the minimum value among them. Then normalize T m,n to obtain. For the computing task request U m schedule it to the satellite sub-cluster computing power node V n The normalized task completion time NT m,n The expression is:
[0075]
[0076] Then the computing power optimization model can be established as an integer programming model. Specifically, the objective function of the computing power optimization model is:
[0077]
[0078] Among them, ω1 represents the proportion weight of the resource load balance degree in the optimization objective, ω2 represents the proportion weight of the overall task completion time in the optimization objective, the sum of ω1 and ω2 is 1, and x m,n represents scheduling the computing task request U m Scheduling to the computing power node V of the satellite sub-cluster n The decision-making magnitude on it. x m,n The larger x is, the more likely (probability or certainty) it is to schedule the computing task request U m to the computing power node V of the satellite sub-cluster n For example, when x m,n = 1, it means to schedule the computing task request U m to the computing power node V of the satellite sub-cluster n on it.
[0079] Furthermore, the constraint conditions of the above computing power optimization model are as follows:
[0080] Constraint condition 1:
[0081] Constraint condition 2:
[0082] Constraint condition 3:
[0083] Constraint condition 4:
[0084] Constraint condition 5:
[0085] Constraint condition 6:
[0086] Among them, constraint condition 1 means that the total computing resource requirements of all computing tasks (i.e., computing task requests) scheduled to the computing power node of the satellite sub-cluster do not exceed the computing resources (i.e., the available computing resource magnitude) that the computing power node of the satellite sub-cluster can provide. Constraint condition 2 means that the total storage resource requirements of all computing tasks scheduled to the computing power node of the satellite sub-cluster do not exceed the storage resources (i.e., the available storage resource magnitude) that the computing power node of the satellite sub-cluster can provide. Constraint condition 3 means that the total energy consumption of all computing tasks scheduled to the computing power node of the satellite sub-cluster does not exceed the upper limit of energy consumption (i.e., the available energy consumption upper limit) that the computing power node of the satellite sub-cluster can provide. Constraint condition 4 means that the processing time of the computing power node of the satellite sub-cluster allocated for the computing task cannot exceed the deadline for completing the computing task (i.e., the latest allowed completion time). Constraint condition 5 means that a computing task can only be scheduled to one computing power node of the satellite sub-cluster.
[0087] S1033. Solve the computing power optimization model by using a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance degree to obtain a task offloading solution.
[0088] Finally, based on the earliest overall task completion time and the resource load balancing degree, a computing power scheduling algorithm is used to solve the computing power optimization model under the above constraints to obtain a task offloading scheme.
[0089] In some embodiments, S1033 may specifically include: solving the computing power optimization model using a 1.5-approximation algorithm.
[0090] Since the computing power resources of all satellite sub-cluster computing power nodes in the federated computing power cluster and all computing task requests within the time slot interval are globally considered, the number of variables is too large. If an exact algorithm is used to solve it, it takes a long time. Therefore, in this embodiment, the solution idea of the 1.5-approximation algorithm based on the multi-cluster scheduling problem can be used to quickly solve the task offloading scheme based on the earliest overall task completion time and the resource load balancing degree.
[0091] Specifically, using the 1.5-approximation algorithm to solve the computing power optimization model may include the following steps S201-S208:
[0092] S201. Initialize the computing power resource information of the federated computing power cluster at the current moment and the computing task requests of the user.
[0093] S202. Sort all computing task requests in descending order of the computing task volume, and for computing task requests with the same computing task volume, sort them in descending order of the resource demand.
[0094] S203. Select the unscheduled computing tasks from all computing tasks as the to-be-scheduled computing tasks, and select the satellite sub-cluster computing power nodes in the federated computing power cluster that meet the computing task requirements (for example: computing resource demand, storage resource demand, etc.).
[0095] S204. Sort the satellite sub-cluster computing power nodes that meet the to-be-scheduled computing task requirements in ascending order of the earliest idle time, and for satellite sub-cluster computing power nodes with the same idle time, sort them in ascending order of the load degree.
[0096] S205. Schedule the to-be-scheduled computing tasks to the satellite sub-cluster computing power node with the earliest idle time and the lowest load degree.
[0097] S206. Update the computing power resource information, idle time, and load degree of the satellite sub-cluster computing power node that processes the to-be-scheduled computing task in S205.
[0098] S207. If there are still unscheduled computing tasks among all computing tasks, execute S203. Otherwise, execute S208.
[0099] S208. Output the scheduling result of the computing tasks at the current moment, which is used as the task offloading scheme.
[0100] In this way, through the above S201 - S208, a task offloading solution for solving the computing power optimization model can be achieved using the 1.5 - approximation algorithm. The time complexity of the computing power scheduling algorithm based on the earliest overall task completion time and resource load balance is O(Mlog2N). Compared with the exact algorithms such as the branch - and - bound method with exponential time complexity O(2 MN ), the 1.5 - approximation algorithm can obtain an optimized solution in a relatively short time. At the same time, compared with greedy algorithms such as first - come - first - served (also known as greedy algorithms), the 1.5 - approximation algorithm can effectively and quickly solve problems such as unbalanced resource load and late overall task completion time.
[0101] S104. Schedule the computing task request to the target computing power node for calculation according to the task offloading solution.
[0102] Finally, according to the target computing power node included in the task offloading solution, the computing task request can be scheduled to the corresponding target computing power node for calculation.
[0103] In one implementation, the establishment of mutual trust and interconnection between service nodes (i.e., between satellite sub - cluster computing power nodes) can be created. Specifically, based on the scheduling of the user's computing task request, a task offloading solution is determined (which can specifically include: offloading strategies and demands of offloaded tasks). Based on policy - based scheduling, services can be created and network traffic can be connected.
[0104] In some embodiments, as Figure 1 shown, after S104. Schedule the computing task request to the target computing power node for calculation, the method further includes:
[0105] S105. After the calculation task request is calculated and processed at the target computing power node, return the calculation result corresponding to the calculation task request.
[0106] Specifically, after the user's computing task request has gone through various scheduling and offloading, the computing task request is transferred to a certain satellite sub - cluster computing power node or multiple satellite sub - cluster computing power nodes for operation. After the calculation and processing are completed, the operation result of the task (i.e., the calculation result corresponding to the calculation task request) will be returned in sequence.
[0107] In some embodiments, as Figure 1 shown, after S104. Schedule the computing task request to the target computing power node for calculation, the method may further include:
[0108] S106. Synchronously update the computing power resource information based on the federation mechanism of the Karmda architecture.
[0109] Specifically, the calculation task request scheduling runs on the computing power nodes of different satellite sub-clusters, occupying different amounts of resources. The amount of resources changes continuously with the concurrent pressure of the task. The system will maintain a relatively good experience standard. The information fed back by each satellite sub-cluster computing power node will be summarized and synchronized based on the mechanism of federated scheduling, and the computing power resource information of the federated computing power cluster will be updated, and finally presented to the user.
[0110] The method for dynamically migrating and scheduling computing power resources based on the federated mechanism provided by the embodiments of the present application can be applied to a multi-cloud management platform based on the Karmada architecture. Using the computing power nodes of satellite sub-clusters as sub-cluster nodes, the grid connection and resource unification of the computing power resources of the satellite sub-cluster computing power nodes are realized through the Karmada architecture and opened to users in the form of services. The problem of offloading the computing task requests of users is modeled as an optimization problem, and the computing power scheduling algorithm based on the earliest overall task completion time and resource load balance is used to solve it, and the optimal task offloading plan is obtained. Finally, based on the determined task offloading plan, the computing task requests of users are distributed to the corresponding satellite sub-cluster computing power nodes for computing and processing.
[0111] This method makes scheduling decisions by globally considering all the computing task requests within a single time slot and the computing power resource information of the federated computing power cluster, thereby optimizing the overall load balance of the entire system and the overall completion time of all computing task requests. Specifically, the optimization problem modeling method is used to fully schedule the computing power resources of multiple federated member nodes (i.e., the computing power nodes of satellite sub-clusters) to achieve the optimal scheduling of the computing power resources of multiple member groups. The established optimization model jointly considers from the perspectives of computing power resources and tasks and includes multiple quality of service (QoS) constraint conditions. The time complexity of the computing power scheduling algorithm based on the earliest overall task completion time and resource load balance is O(Mlog2N). Compared with the exact algorithms such as branch and bound with exponential time complexity O(2 MN )), this algorithm can obtain an optimized solution in a relatively short time. At the same time, compared with greedy algorithms such as first-come-first-served, this algorithm can solve problems such as unbalanced resource load and late overall task completion time. And it can combine the algorithms of cross-cluster scheduling and intelligent online scheduling to solve the problem of fusion and coordination of computing power resources.
[0112] In this way, the method for dynamically migrating and scheduling computing power resources based on the federated mechanism provided by the embodiments of the present application can realize the online dynamic migration of computing power resources according to the computing task requests of users. And it can perform cross-cluster scheduling among multiple satellite sub-cluster computing power nodes, greatly improving the utilization rate of computing power resources, making full and effective use of computing power resources, and improving the processing efficiency of users' computing task requests.
[0113] The embodiment of the present application also provides a dynamic migration scheduling system for computing power resources based on the federation mechanism. Figure 3 FIG. is a schematic structural diagram of the dynamic migration scheduling system for computing power resources based on the federation mechanism provided by the embodiment of the present application. As Figure 3 shown, the dynamic migration scheduling system 300 for computing power resources based on the federation mechanism is constructed based on the Karmada architecture, and specifically includes: a Monitoring module 301, an AI scheduling module 302, and an Operation Management module 303.
[0114] Among them, the Karmada architecture is a computing power scheduling framework, mainly developed based on Cluster Federation V2, which supports optimizing the computing power scheduling ability and framework details with insufficient support. The multi-cloud and hybrid cloud multi-cluster environment it manages includes the following two types of clusters:
[0115] lHost cluster: A cluster composed of the karmada control plane, which receives the application deployment requirements submitted by users, synchronizes the application deployment requirements to the member clusters, and synchronizes the subsequent running status of the applications from the member clusters.
[0116] lMember cluster: Composed of one or more k8s clusters, responsible for running the applications submitted by users.
[0117] The dynamic migration scheduling system 300 for computing power resources based on the federation mechanism provided by the embodiment of the present application is developed and integrated with capabilities based on the Karmada architecture, and applies the scheduling model algorithm to the Karmada architecture to achieve Karmada+.
[0118] Among them, the Karmada framework mainly realizes the aggregation of multiple sub-clusters (i.e., satellite sub-cluster computing power nodes) into a unified cluster (i.e., federated computing power cluster), realizes the aggregation and networking of multiple satellite sub-cluster computing power nodes, and the ETCD (distributed scheduling storage system) completes the aggregation of the computing power resources and status of all satellite sub-cluster computing power nodes, realizing the unified management of multi-domain computing power and relevant policies of the traffic Ingress-controller.
[0119] The Monitoring module 301 can be used to complete the perception and discovery of computing power resource information, and complete the perception of computing power resource information through load, status monitoring, fault monitoring, etc.
[0120] Specifically, the Monitoring module 301 can be used to obtain the computing power resource information of the federated computing power cluster. The federated computing power cluster includes multiple satellite sub-cluster computing power nodes (such as Figure 3The satellite sub-cluster computing power nodes A, satellite sub-cluster computing power nodes B... satellite sub-cluster computing power nodes X shown in the figure, and the computing power resource information is used to characterize the computing resources, storage resources, energy consumption, network topology and network status of the satellite sub-cluster computing power nodes.
[0121] The OperationManager module 303 can be used to receive the computing task requests of users within a single time slot, and determine the task computing power request information according to the computing task requests. The task computing power request information is used to characterize the service location, computing power resource requirement size, task deadline and service quality requirement situation of the computing task requests.
[0122] The AI scheduling module 302 can be used to complete the processing, learning, training of algorithms, output of related capabilities such as model prediction, etc. of the perception data (for example: computing power resource information and task computing power request information).
[0123] Specifically, the AI scheduling module 302 can be used to establish an optimal computing power model according to the computing power resource information and the task computing power request information, and determine the task offloading plan according to the optimal model. The task offloading plan includes target computing power nodes, and the target computing power nodes are one or more of the multiple satellite sub-cluster computing power nodes.
[0124] In some embodiments, the AI scheduling module 302 is specifically used for: First, perform information modeling on the computing power resource information and the task computing power request information to obtain the modeled computing power resource information and the modeled task computing power request information. Then, establish an optimal computing power model according to the modeled computing power resource information and the modeled task computing power request information. Finally, use a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance to solve the optimal computing power model to obtain the task offloading plan.
[0125] The OperationManager module 303 can be used to complete the deployment of platform applications and schedule computing task requests.
[0126] Specifically, the OperationManager module 303 can also be used to schedule the computing task requests to the target computing power nodes for computing according to the task offloading plan.
[0127] In some embodiments, as Figure 3 shown, the computing power resource dynamic migration scheduling system 300 based on the federation mechanism further includes: the PolicyManager module 304.
[0128] The PolicyManager module 304 can be used to complete the definition, creation, update, deletion and viewing of scheduling policies.
[0129] Specifically, the PolicyManager module 304 can be used to define a scheduling policy, and / or create a scheduling policy, and / or update a scheduling policy, and / or delete a scheduling policy, and / or view a scheduling policy. The scheduling policy is used to instruct the AI scheduling module 302 to determine a task offloading solution.
[0130] In some embodiments, as Figure 3 shown, the dynamic migration scheduling system 300 of computing power resources based on the federation mechanism further includes: a deployment control Devops module 305.
[0131] Specifically, the deployment control Devops module 305 can be used to complete the unified management of code versions and the process standardization of releasing versions, and realize the management of the entire life cycle of an application from deployment to extinction. Among them, the deployment control Devops module 305 can be included in the Developer Center DevopersCenter.
[0132] In some embodiments, as Figure 3 shown, the Karmada architecture in the dynamic migration scheduling system 300 of computing power resources based on the federation mechanism provided by the embodiments of the present application can specifically include: a federated scheduler (Karmada Scheduler), a federated interface service (Karmada API Server), a distributed scheduling storage system (ETCD), and a federated controller (KarmadaController). Among them, the federated controller can specifically include: a cluster controller (Cluster Contolller), a policy controller (Policy Controller), a domain name service controller (Bindig Controller), and an execution controller (Execution Controller).
[0133] By using the dynamic migration scheduling system of computing power resources based on the federation mechanism provided by the embodiments of the present application, the system can realize the grid connection and resource unification of the computing power resources of the satellite sub-cluster computing power nodes through the Karmada architecture, and open them to users in the form of services. The offloading problem of the user's computing task request is modeled as an optimization problem, and it is solved based on a computing power scheduling algorithm of the earliest overall task completion time and resource load balance degree to obtain an optimal task offloading solution. Finally, based on the determined task offloading solution, the user's computing task request is distributed to the corresponding satellite sub-cluster computing power nodes for computing processing. In this way, the dynamic migration scheduling system of computing power resources based on the federation mechanism provided by the embodiments of the present application can realize the online dynamic migration of computing power resources according to the user's computing task request. And it can perform cross-cluster scheduling among multiple satellite sub-cluster computing power nodes, greatly improving the utilization rate of computing power resources, making full and effective use of computing power resources, and improving the processing efficiency of user computing task requests.
[0134] An embodiment of the present invention further provides an electronic device, which may include: a display screen, a memory, and one or more processors. The display screen, the memory, and the processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can execute each method or step executed in the above-described embodiment of the dynamic migration scheduling method of computing power resources based on the federation mechanism. Of course, the electronic device includes, but is not limited to, the above-mentioned display screen, memory, and one or more processors.
[0135] An embodiment of the present invention further provides a computer-readable storage medium for storing computer instructions for running the above-described dynamic migration scheduling method of computing power resources based on the federation mechanism.
[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0137] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0138] In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0139] For the similar parts between the embodiments provided in this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts belong to the protection scope of this application.
Claims
1. A dynamic migration scheduling method for computing power resources based on a federated mechanism, characterized in that Including: Obtain the computing power resource information of the federated computing power cluster, where the federated computing power cluster includes multiple satellite sub-cluster computing power nodes, and the computing power resource information is used to characterize the computing resources, storage resources, energy consumption, network topology, and network status of the satellite sub-cluster computing power nodes; Receive the computing task requests of users within a single time slot, and determine the task computing power request information according to the computing task requests. The task computing power request information is used to characterize the service location, computing power resource requirement size, task deadline, and service quality requirement of the computing task requests; Establish a computing power optimization model based on the computing power resource information and the task computing power request information, and determine a task offloading scheme according to the optimization model. The task offloading scheme includes target computing power nodes, and the target computing power nodes are one or more of the multiple satellite sub-cluster computing power nodes; Schedule the computing task requests to the target computing power nodes for computing according to the task offloading scheme.
2. The method according to claim 1, wherein The step of establishing a computing power optimization model based on the computing power resource information and the task computing power request information, and determining a task offloading scheme according to the optimization model includes: Perform information modeling on the computing power resource information and the task computing power request information to obtain the modeled computing power resource information and the modeled task computing power request information; Establish the computing power optimization model based on the modeled computing power resource information and the modeled task computing power request information; Use a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance to solve the computing power optimization model to obtain the task offloading scheme.
3. The method according to claim 2, wherein The expression of the modeled computing power resource information is: V = {V1, V2, …, V N}, V n ∈ V; V n ′ = {VC n , VR n , VF n , VE n}; Among them, V represents the federal computing power cluster, N represents the number of computing power nodes in the satellite sub-cluster, V n represents the nth computing power node in the satellite sub-cluster, V n ′ represents the computing power node in the satellite sub-cluster V n The computing power resource information after modeling, VC n represents the available computing resource size of the computing power node in the satellite sub-cluster V n represents the available storage resource size of the computing power node in the satellite sub-cluster V, VR n represents the available storage resource size of the computing power node in the satellite sub-cluster V n represents the available storage resource size of the computing power node in the satellite sub-cluster V, VF n represents the computing power provided by the computing power node in the satellite sub-cluster V n represents the computing power provided by the computing power node in the satellite sub-cluster V, VE n represents the available energy consumption upper limit of the computing power node in the satellite sub-cluster V n ; The expression of the modeled task computing power request information is: U = {U1, U2, …, U M}, U m ∈ U; U ′ m = {UV m , UC m , UR m , UF m , UE m , UD m}; Among them, U represents the set of computing task requests, M represents the number of the computing task requests, U m represents the m-th computing task request, U ′ m represents the computing task request U m the task computing power request information after modeling, UV m represents the computing task request U m the access satellite of the computing task request U, UC m represents the computing task request U m the computing resource requirement of the computing task request U, UR m represents the computing task request U m the storage resource requirement of the computing task request U, UF m represents the computing task request U m the computing task volume of the computing task request U, UP m represents the computing task request U m the transmission task volume of the computing task request U, UE m represents the computing task request U m the energy consumption of the computing task request U, UD m represents the computing task request U m the allowed latest completion time.
4. The method according to claim 3, wherein The computing power optimization model is an integer programming model; the objective function of the computing power optimization model is: Wherein, The constraint conditions of the computing power optimization model are: Among them, ω1 represents the proportion weight of the resource load balance degree in the optimization objective, ω2 represents the proportion weight of the overall task completion time in the optimization objective, and the sum of ω1 and ω2 is 1. x m,n represents the decision-making magnitude of scheduling the computing task request U m to the satellite sub-cluster computing power node V n . x m,n x = 1 means scheduling the computing task request U m to the satellite sub-cluster computing power node V n . B m,n represents the load degree of scheduling the computing task request U m to the satellite sub-cluster computing power node V n . NB m,n represents the normalized load degree of scheduling the computing task request U m to the satellite sub-cluster computing power node V n . T m,n represents the task completion time of scheduling the computing task request U m to the satellite sub-cluster computing power node V n . NT m,n represents the normalized task completion time of scheduling the computing task request U m to the satellite sub-cluster computing power node V n . E n,UVm represents the network transmission capacity between the access satellite of the computing task request U m and the satellite sub-cluster computing power node V n . maxB represents the maximum value in the B m,n , minB represents the minimum value in the B m,n , maxT represents the maximum value in the T m,n , minT represents the minimum value in the T m,n .
5. The method according to claim 4, characterized in that, The step of using a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance to solve the computing power optimization model includes: Use a 1.5-approximation algorithm to solve the computing power optimization model.
6. The method according to claim 1, characterized in that, After scheduling the computing task requests to the target computing power nodes for computing according to the task offloading scheme, the method further includes: After the target computing power nodes complete the computing processing of the computing task requests, return the computing results corresponding to the computing task requests.
7. The method according to claim 1, characterized in that, After scheduling the computing task requests to the target computing power nodes for computing according to the task offloading scheme, the method further includes: Synchronously update the computing power resource information based on the federated mechanism of the Karmada architecture.
8. A computing power resource dynamic migration and scheduling system based on a federated mechanism, characterized in that, The system is constructed based on the Karmada architecture, and the system includes: a Monitoring module, an AI Scheduling module, and an Operation Manager module; wherein, The Monition module is used to obtain the computing power resource information of the federated computing power cluster, where the federated computing power cluster includes multiple satellite sub-cluster computing power nodes, and the computing power resource information is used to characterize the computing resources, storage resources, energy consumption, network topology, and network status of the satellite sub-cluster computing power nodes; The OperationManager module is used to receive the computing task requests of users within a single time slot, and determine the task computing power request information according to the computing task requests. The task computing power request information is used to characterize the service location, the size of the computing power resource requirements, the task deadline, and the quality of service requirements of the computing task requests; The AI scheduling module is used to establish an optimal computing power model according to the computing power resource information and the task computing power request information, and determine the task offloading scheme according to the optimal model. The task offloading scheme includes target computing power nodes, and the target computing power nodes are one or more of the multiple satellite sub-cluster computing power nodes; The OperationManager module is further used to schedule the computing task requests to the target computing power nodes for computing according to the task offloading scheme.
9. The system according to claim 8, wherein The AI scheduling module is specifically used for: Performing information modeling on the computing power resource information and the task computing power request information to obtain the modeled computing power resource information and the modeled task computing power request information; Establishing the optimal computing power model according to the modeled computing power resource information and the modeled task computing power request information; Solving the optimal computing power model by using a computing power scheduling algorithm based on the earliest overall task completion time and resource load balance to obtain the task offloading scheme.
10. The system according to claim 8 or 9, characterized in that, The system further includes: a PolicyManager module for policy management; The policy management module is used to define, and / or create, and / or update, and / or delete, and / or view the scheduling policy, and the scheduling policy is used to instruct the AI scheduling module to determine the task offloading scheme.