Computing power scheduling method and system for AI server cluster
By conducting detailed analysis and optimization of the computing power resources of the AI server cluster, the problem of failure to fully consider data processing and collaborative processing capabilities in the existing technology is solved, and better computing power scheduling results are achieved.
Patent Information
- Application Number
- CN202410732152.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-06-06
AI Technical Summary
The computing power scheduling methods of existing AI server clusters fail to fully consider the data processing capabilities and collaborative processing capabilities of computing power resources, resulting in insufficient optimization of scheduling results.
By selecting the computing power analysis indicators of the AI server cluster, performing computing power measurements, calculating transmission and processing delays and energy consumption, generating computing power scheduling status and actions, constructing action value functions and state value functions, and optimizing these functions to determine the optimal computing power scheduling results.
The analysis of computing power resource data processing capabilities and collaborative processing capabilities has been improved, and the computing power scheduling results have been optimized to make them more in line with actual needs.
Smart Images

Figure CN118626224B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computing power scheduling technology, and in particular to a computing power scheduling method and system for an AI server cluster. Background Art
[0002] An AI server cluster refers to a collection of high-performance heterogeneous accelerated computing nodes equipped with the Lingjun optimization suite. Computing power scheduling refers to the process of scheduling computing power resources to users and processing data for them.
[0003] At present, there has been extensive research on the computing power scheduling of AI server clusters at home and abroad. The general computing power scheduling method is as follows: analyze the user task volume, evaluate the task processing efficiency of computing power resources, and construct an objective function for computing power resources to process the user's task volume. By finding the minimum value of the objective function, determine which computing power resources have the lowest cost to process the user's task volume, thereby determining the computing power resources that ultimately need to be scheduled.
[0004] In the above computing power scheduling process, the task processing efficiency of computing power resources is evaluated by calculating the energy consumption of computing power resources, that is, only the number, time and other costs of computing power resources for task processing are evaluated, but whether computing power resources have resources that can solve tasks, rather than the amount of resources, is not considered. Secondly, the objective function mainly includes the task processing cost of computing power resources. When the objective function is the minimum, the selected computing power resources are greatly affected by the task processing efficiency of computing power resources, and the data processing capacity of computing power resources is not considered. Finally, when selecting the action of computing power resource scheduling, only one computing power node is often selected for scheduling, and the problem of scheduling multiple computing power resources is not considered. Therefore, the computing power scheduling of AI server clusters does not adequately analyze the data processing capacity and collaborative processing capacity of computing power resources. Summary of the invention
[0005] In order to solve the above problems, the present invention provides a computing power scheduling method and system for an AI server cluster, which can increase the analysis of the data processing capability and collaborative processing capability of computing power resources.
[0006] In a first aspect, the present invention provides a computing power scheduling method for an AI server cluster, comprising:
[0007] Selecting a computing power analysis indicator of an AI server cluster, and using the computing power analysis indicator to measure the computing power resources in the AI server cluster to obtain a computing power measurement value;
[0008] Calculate the transmission delay and transmission energy consumption of the user's pending task from the user to the computing power resource, calculate the processing delay and processing energy consumption of the computing power resource processing the pending task, and calculate the first computing power scheduling cost of scheduling the computing power resource for the pending task based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption;
[0009] Query the collaborative computing resources when the computing resource processes the task to be processed, and calculate the second computing resource scheduling cost of the computing resource for scheduling the collaborative computing resource;
[0010] Generate a computing power scheduling state of the task to be processed by using the computing power measurement value, the first computing power scheduling cost, and the second computing power scheduling cost, generate a computing power scheduling action of the task to be processed, and construct an action value function and a state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action;
[0011] After optimizing the action value function and the state value function to obtain the optimized action value function and the optimized state value function, the optimized action value function and the optimized state value function are used to determine the computing power scheduling result of the task to be processed on the AI server cluster.
[0012] In a possible implementation of the first aspect, the selecting a computing power analysis indicator of the AI server cluster includes:
[0013] Dividing the primary computing power indicators of the AI server cluster;
[0014] The first-level computing power index includes logical operation index, parallel computing index and neural network computing index;
[0015] Divide the secondary computing power indicators of the primary computing power indicators;
[0016] The secondary computing power index includes the number of logical operations, the duration of logical operations and the speed of logical operations of the logical operation index, the number of parallel calculations, the duration of parallel calculations and the speed of parallel calculations of the parallel calculation index, and the number of neural network calculations, the duration of neural network calculations and the speed of neural network calculations of the neural network calculation index;
[0017] The number of logic operations, the speed of logic operations, the number of parallel calculations, the speed of parallel calculations, the number of neural network calculations, and the speed of neural network calculations are used as positive indicators of the AI server cluster, and the duration of logic operations, the duration of parallel calculations, and the duration of neural network calculations are used as negative indicators of the AI server cluster;
[0018] The positive indicator and the negative indicator are used as computing power analysis indicators of the AI server cluster.
[0019] In a possible implementation of the first aspect, the using the computing power analysis indicator to measure the computing power resources in the AI server cluster to obtain a computing power measurement value includes:
[0020] The indicator weight corresponding to the computing power analysis indicator is calculated using the following formula:
[0021]
[0022]
[0023] Among them, W u represents the indicator weight of the u-th indicator, ω uv represents the vth sampling value of the uth computing power analysis indicator, V represents the total number of samples of the uth computing power analysis indicator, e u The information entropy of the weight of the u-th indicator, U represents the total number of computing power analysis indicators;
[0024] Based on the indicator weight, the computing power measurement value corresponding to the computing power analysis indicator is calculated using the following formula:
[0025]
[0026] Among them, H represents the computing power measurement value, W u represents the indicator weight of the u-th indicator, ω uv Indicates the vth sample value of the uth computing power analysis indicator.
[0027] In a possible implementation manner of the first aspect, the calculating a transmission delay and a transmission energy consumption of a to-be-processed task of a user from the user to the computing power resource includes:
[0028] The following formula is used to calculate the transmission delay of the task to be processed by the user from the user to the computing resource:
[0029]
[0030] Wherein, Δt1 represents the transmission delay, L represents the distance of the task to be processed from the user to the computing resource, v1 represents the data propagation rate on the L distance, B represents the data frame size of the task to be processed, and v2 represents the data transmission rate on the L distance;
[0031] The following formula is used to calculate the transmission energy consumption of the task to be processed by the user from the user to the computing resource:
[0032] E1=α+βv2ε
[0033] Among them, E1 represents the transmission energy consumption, v2 represents the data transmission rate on the L distance, α represents the energy consumption of the task to be processed when it is idle on the distance from the user to the computing resource, and β and ε both represent positive integers used to fit the functional relationship between v2 and E1.
[0034] In a possible implementation manner of the first aspect, the calculating the processing delay and processing energy consumption of the computing resource for processing the task to be processed includes:
[0035] The processing delay of the computing resource for processing the task to be processed is calculated using the following formula:
[0036]
[0037] Among them, Δt2 represents the processing delay, t1 represents the start time of the computing resource processing the pending task, t2 represents the time when the pending task arrives at the computing resource, and CPI i represents the number of clock cycles required for the ith instruction of the task to be processed, n represents the number of instructions of the task to be processed, T CPU The clock cycle length corresponding to the number of clock cycles required for the i-th instruction;
[0038] The processing energy consumption of the computing resource for processing the task to be processed is calculated using the following formula:
[0039]
[0040] Wherein, E2 represents the processing energy consumption, t1 represents the start time of the computing resource processing the task to be processed, represents the end time of the computing resource processing the task to be processed, s(t) represents the electrical signal at time t when the computing resource processes the task to be processed, and α' represents the energy consumption when the computing resource is idle.
[0041] In a possible implementation of the first aspect, querying the collaborative computing resources when the computing resources process the task to be processed includes:
[0042] Obtain other computing resources in the AI server cluster except the computing resources;
[0043] Identifying response time limits of remaining tasks to be processed among the tasks to be processed;
[0044] The waiting response time of the remaining tasks to be processed waiting for the other computing resources is calculated using the following formula:
[0045] Δt3=t 1 +t2
[0046] Among them, Δt3 represents the waiting response time, t 1 Indicates the transmission time of the remaining pending tasks from the computing resource to the waiting queue of the other computing resources, t 2 Indicates the time taken by the other computing resources to process the tasks in the waiting queue that are ahead of the remaining tasks to be processed;
[0047] Determining whether the waiting response time exceeds the response time limit;
[0048] When the waiting response time does not exceed the response time limit, the other computing power resources are used as the collaborative computing power resources.
[0049] In a possible implementation manner of the first aspect, calculating the second computing power scheduling cost of the computing power resource scheduling the collaborative computing power resource includes:
[0050] Calculate the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource;
[0051] Identifying remaining tasks to be processed among the tasks to be processed;
[0052] Calculating the processing delay and processing energy consumption of the collaborative computing resources in processing the remaining tasks to be processed;
[0053] The second computing power scheduling cost is calculated by using the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource, and the processing delay and processing energy consumption of the collaborative computing power resource in processing the remaining tasks to be processed.
[0054] In a possible implementation manner of the first aspect, constructing the action value function and the state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action includes:
[0055] The action value function of the task to be processed under the computing power scheduling state and the computing power scheduling action is constructed using the following formula:
[0056]
[0057] Q π (s, a) = r (s, a) + γ∑ s′∈S P(s′|s,a)V π (s′)
[0058] Q π (s, a) = r t +γQ π (S t+1 , A t+1 )
[0059] Among them, Q π (s, a) = r t +γQ π S t+1 , A t+1 represents the action value function, S t , s represents the state selected from the computing power scheduling state at time t, A t , a indicates that the pending task has S t The action selected from the computing power scheduling action under the premise of the state, G t Represents the state s from time t t The sum of all reward decays from the start to the terminal state, represents the expected function, Q π (s, a) indicates that the task to be processed has S t Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t Action, and finally get reward G t The expectation of γ is the discount factor, R t+k represents the reward obtained at time t+k, k represents the constant used to limit the upper limit of t, R t represents the reward obtained at time t, R t+2 represents the reward obtained at time t+2, G t+1 Represents the state s from time t+1 t+1 The sum of the decays of all rewards from the start to the end state, r(s, a) represents the reward of the pending task selecting action a from the computing power scheduling action under the premise of having state s, S represents the set of computing power scheduling states, S t+1 , s′ represents the state selected from the computing power scheduling state at time t+1, V π (s′) indicates that at time t+1, t+1 Under the premise of state, follow the strategy π to select the state value function of any action from the computing power scheduling action, P(s′|s, a) represents the state transition function from state s to state s′ after executing action a, r(s, a) = r t , Q π (S t+1 , A t+1 )=∑ s′∈S P(s′|s,a)V π (s′), Q π (S t+1 , A t+1 ) indicates that the pending task has S t+1 Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t+1 Action, and finally get reward Gt+1 expectations;
[0060] The state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action is constructed using the following formula:
[0061]
[0062] V π (s)=∑ a∈A π(a|s)Q π (s, a)
[0063] Among them, V π (s)=∑ a∈A π(a|s)Q π (s, a) represents the state value function, π(a|s) represents the probability of executing action a under the premise of state s, Q π (s, a) indicates that the task to be processed has S t Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t Action, and finally get reward G t , A represents the set of computing power scheduling actions, S t , s represents the state selected from the computing power scheduling state at time t, A t , a indicates that the pending task has S t The action selected from the computing power scheduling action under the premise of the state, G t Represents the state s from time t t The sum of all reward decays from the start to the terminal state, Represents the expected function.
[0064] In a possible implementation manner of the first aspect, optimizing the action-value function and the state-value function to obtain an optimized action-value function and an optimized state-value function includes:
[0065] The weight parameters in the action value function are optimized using the following formula to obtain the optimized weight parameters:
[0066] Q π (s, a) = r t +γQ π (S t+1 , A t+1 )
[0067] Q π (S t+1 , A t+1 )=Q π (S t+1 , A t+1 , w)
[0068] δ t =Q π (S t , A t ,w)-Q π (S t+1 , A t+1 , w)
[0069]
[0070] Among them, w' represents the optimized weight parameter, w represents the weight parameter, α represents the learning rate, δ t represents Q at time t π (S t , A t , w) and Q π (S t+1 , A t+1 , w), S t , s represents the state selected from the computing power scheduling state at time t, A t , a indicates that the pending task has S t The action selected from the computing power scheduling action under the premise of having the state s, r(s,a) represents the reward of selecting action a from the computing power scheduling action under the premise of the pending task having the state s, r(s,a)=r t , Q π (s,a)=Q π S t ,A t ,w;
[0071] The action value function and the state value function are optimized using the optimized weight parameters to obtain an optimized action value function and an optimized state value function.
[0072] In a second aspect, the present invention provides a computing power scheduling system for an AI server cluster, the system comprising:
[0073] A computing power measurement module, used to select a computing power analysis indicator of the AI server cluster, and use the computing power analysis indicator to measure the computing power resources in the AI server cluster to obtain a computing power measurement value;
[0074] A first calculation module is used to calculate the transmission delay and transmission energy consumption of the pending task of the user from the user to the computing power resource, calculate the processing delay and processing energy consumption of the computing power resource processing the pending task, and calculate the first computing power scheduling cost of the computing power resource for scheduling the pending task based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption;
[0075] A second calculation module is used to query the collaborative computing resources when the computing resources process the task to be processed, and calculate the second computing power scheduling cost of the computing resources for scheduling the collaborative computing resources;
[0076] a function construction module, used to generate the computing power scheduling state of the task to be processed by using the computing power measurement value, the first computing power scheduling cost and the second computing power scheduling cost, and generate the computing power scheduling action of the task to be processed, and construct an action value function and a state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action;
[0077] The result determination module is used to optimize the action value function and the state value function, and after obtaining the optimized action value function and the optimized state value function, use the optimized action value function and the optimized state value function to determine the computing power scheduling result of the task to be processed on the AI server cluster.
[0078] Compared with the prior art, the technical principle and beneficial effects of this solution are:
[0079] By selecting the computing power analysis index of the AI server cluster, the embodiment of the present invention can consider whether the computing power resources have resource categories such as logical operation resources, parallel computing resources and neural network computing resources that can solve the task, rather than the number of resources. The embodiment of the present invention queries the collaborative computing power resources when the computing power resources process the pending tasks to consider the problem of scheduling multiple computing power resources to collaboratively process user tasks. Furthermore, the embodiment of the present invention generates the computing power scheduling state of the pending tasks by using the computing power measurement value, the first computing power scheduling cost and the second computing power scheduling cost, so as to add the computing power measurement value to the computing power scheduling state, and increase the consideration of data processing capability through the computing power measurement value. Secondly, in the computing power scheduling state, the second computing power scheduling cost is used to increase the consideration of collaborative processing capability through the second computing power scheduling cost. Therefore, the computing power scheduling method and system of an AI server cluster proposed in the embodiment of the present invention can increase the analysis of the data processing capability and collaborative processing capability of computing power resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0081] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0082] Figure 1 A flowchart of a computing power scheduling method for an AI server cluster provided in one embodiment of the present invention;
[0083] Figure 2 A schematic diagram of a module of a computing power scheduling system for an AI server cluster provided in accordance with an embodiment of the present invention. DETAILED DESCRIPTION
[0084] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0085] The embodiment of the present invention provides a computing power scheduling method for an AI server cluster, and the execution subject of the computing power scheduling method for the AI server cluster includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiment of the present invention. In other words, the computing power scheduling method for the AI server cluster can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0086] See also Figure 1 The figure shows the process of the computing power scheduling method of the AI server cluster provided by one embodiment of the present invention. Figure 1 The computing power scheduling method of the AI server cluster described in the article includes:
[0087] S1. Select a computing power analysis indicator of an AI server cluster, and use the computing power analysis indicator to measure the computing power resources in the AI server cluster to obtain a computing power measurement value.
[0088] In an embodiment of the present invention, the AI server cluster refers to a collection of high-performance heterogeneous accelerated computing nodes with a Lingjun optimization kit.
[0089] Furthermore, by selecting computing power analysis indicators of the AI server cluster, the embodiment of the present invention can consider whether the computing power resources include resource categories such as logical operation resources, parallel computing resources, and neural network computing resources that can solve the task, rather than the quantity of resources.
[0090] In one embodiment of the present invention, the selecting of computing power analysis indicators of the AI server cluster includes: dividing the first-level computing power indicators of the AI server cluster; wherein the first-level computing power indicators include logic operation indicators, parallel computing indicators and neural network computing indicators; dividing the second-level computing power indicators of the first-level computing power indicators; wherein the second-level computing power indicators include the number of logic operations, the duration of logic operations and the speed of logic operations of the logic operation indicators, the number of parallel calculations, the duration of parallel calculations and the speed of parallel calculations of the parallel calculation indicators, and the number of neural network calculations, the duration of neural network calculations and the speed of neural network calculations of the neural network calculation indicators; using the number of logic operations, the speed of logic operations, the number of parallel calculations, the parallel calculation speed, the number of neural network calculations and the speed of neural network calculations as positive indicators of the AI server cluster, and using the logic operation duration, the parallel calculation duration and the neural network calculation duration as negative indicators of the AI server cluster; using the positive indicators and the negative indicators as computing power analysis indicators of the AI server cluster.
[0091] Among them, the logical operation index refers to the index of the logical operation capability of the AI server cluster, the parallel computing index refers to the index of the parallel computing capability of the AI server cluster, the neural network computing index refers to the index of whether the AI server cluster can perform neural network computing, the number of logical operations, the number of parallel computing and the number of neural network computing all refer to the maximum amount of user tasks that the computing power resources can handle, the logical operation duration, the parallel computing duration and the neural network computing duration all refer to the duration required for the computing power resources to handle the maximum amount of user tasks. It should be noted that the number of logical operations and the logical operation duration collect data from the logical operation components of the AI server cluster, the number of parallel computing and the parallel computing duration collect data from the parallel computing components of the AI server cluster, and the number of neural network computing and the neural network computing duration collect data from the neural network computing components of the AI server cluster.
[0092] In one embodiment of the present invention, the step of measuring the computing power resources in the AI server cluster by using the computing power analysis indicator to obtain a computing power measurement value includes: calculating the indicator weight corresponding to the computing power analysis indicator by using the following formula:
[0093]
[0094]
[0095] Among them, W u represents the indicator weight of the u-th indicator, ω uv represents the vth sampling value of the uth computing power analysis indicator, V represents the total number of samples of the uth computing power analysis indicator, eu The information entropy of the weight of the u-th indicator, U represents the total number of computing power analysis indicators;
[0096] Based on the indicator weight, the computing power measurement value corresponding to the computing power analysis indicator is calculated using the following formula:
[0097]
[0098] Among them, H represents the computing power measurement value, W u represents the indicator weight of the u-th indicator, ω uv Indicates the vth sample value of the uth computing power analysis indicator.
[0099] S2. Calculate the transmission delay and transmission energy consumption of the user's pending tasks from the user to the computing power resources, calculate the processing delay and processing energy consumption of the computing power resources processing the pending tasks, and calculate the first computing power scheduling cost of scheduling the computing power resources for the pending tasks based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption.
[0100] In the embodiment of the present invention, the transmission delay includes the delay of data transmission to the communication path, that is, b / s, and also includes the propagation delay of data on the communication path, that is, distance / s.
[0101] In one embodiment of the present invention, the calculating the transmission delay and transmission energy consumption of the to-be-processed task of the user side from the user side to the computing power resource includes: calculating the transmission delay of the to-be-processed task of the user side from the user side to the computing power resource using the following formula:
[0102]
[0103] Wherein, Δt1 represents the transmission delay, L represents the distance of the task to be processed from the user to the computing resource, v1 represents the data propagation rate on the L distance, B represents the data frame size of the task to be processed, and v2 represents the data transmission rate on the L distance;
[0104] The following formula is used to calculate the transmission energy consumption of the task to be processed by the user from the user to the computing resource:
[0105] E1=α+βv2 ε
[0106] Among them, E1 represents the transmission energy consumption, v2 represents the data transmission rate on the L distance, α represents the energy consumption of the task to be processed when it is idle on the distance from the user to the computing resource, and β and ε both represent positive integers used to fit the functional relationship between v2 and E1.
[0107] In one embodiment of the present invention, the calculating of the processing delay and processing energy consumption of the computing resource for processing the task to be processed includes: calculating the processing delay of the computing resource for processing the task to be processed using the following formula:
[0108]
[0109] Among them, Δt2 represents the processing delay, t1 represents the start time of the computing resource processing the pending task, t2 represents the time when the pending task arrives at the computing resource, and CPI i represents the number of clock cycles required for the ith instruction of the task to be processed, n represents the number of instructions of the task to be processed, T CPU The clock cycle length corresponding to the number of clock cycles required for the i-th instruction;
[0110] The processing energy consumption of the computing resource for processing the task to be processed is calculated using the following formula:
[0111]
[0112] Wherein, E2 represents the processing energy consumption, t1 represents the start time of the computing resource processing the task to be processed, represents the end time of the computing resource processing the task to be processed, s(t) represents the electrical signal at time t when the computing resource processes the task to be processed, and α' represents the energy consumption when the computing resource is idle.
[0113] In one embodiment of the present invention, the first computing power scheduling cost of scheduling the computing power resource for the task to be processed based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption is calculated, including: based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption, the first computing power scheduling cost of scheduling the computing power resource for the task to be processed is calculated using the following formula:
[0114] C1=w 1 Δt1+w 2 E1+w 3 Δt2+w 4 E2
[0115] Among them, C1 represents the first computing power scheduling cost, w 1 、w 2 、w 3 、w 4 represents the weight, Δt1 represents the transmission delay, E1 represents the transmission energy consumption, Δt2 represents the processing delay, and E2 represents the processing energy consumption.
[0116] S3. Query the collaborative computing resources when the computing resources process the pending tasks, and calculate the second computing power scheduling cost of the computing resources for scheduling the collaborative computing resources.
[0117] The embodiment of the present invention is used to consider the problem of scheduling multiple computing resources to collaboratively process user tasks by querying the collaborative computing resources when the computing resources process the task to be processed.
[0118] In one embodiment of the present invention, the querying of the collaborative computing resources when the computing resources process the pending tasks includes: obtaining other computing resources in the AI server cluster except the computing resources; identifying the response time limit of the remaining pending tasks in the pending tasks; and calculating the waiting response time of the remaining pending tasks waiting for the other computing resources using the following formula:
[0119] Δt3=t 1 +t 2
[0120] Among them, Δt3 represents the waiting response time, t 1 Indicates the transmission time of the remaining pending tasks from the computing resource to the waiting queue of the other computing resources, t 2 Indicates the time taken by the other computing resources to process the tasks in the waiting queue that are ahead of the remaining tasks to be processed;
[0121] Determine whether the waiting response time exceeds the response time limit; when the waiting response time does not exceed the response time limit, use the other computing power resources as the collaborative computing power resources.
[0122] The response time limit refers to the period between the current time of the remaining tasks to be processed and the start time when the remaining tasks to be processed require the other computing resources to respond to the user's task processing request.
[0123] In one embodiment of the present invention, the calculation of the second computing power scheduling cost of the computing power resource for scheduling the collaborative computing power resource includes: calculating the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource; identifying the remaining tasks to be processed among the tasks to be processed; calculating the processing delay and processing energy consumption of the collaborative computing power resource for processing the remaining tasks to be processed; and calculating the second computing power scheduling cost using the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource and the processing delay and processing energy consumption of the collaborative computing power resource for processing the remaining tasks to be processed.
[0124] It should be noted that the process of calculating the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource; calculating the processing delay and processing energy consumption of the collaborative computing power resource for processing the remaining tasks to be processed; and calculating the second computing power scheduling cost using the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource, and the processing delay and processing energy consumption of the collaborative computing power resource for processing the remaining tasks to be processed is similar to the above-mentioned process of calculating the transmission delay and transmission energy consumption of the user's pending tasks from the user side to the computing power resource, calculating the processing delay and processing energy consumption of the computing power resource for processing the pending tasks, and calculating the first computing power scheduling cost of scheduling the computing power resource for scheduling the pending tasks based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption. The difference is that the above-mentioned process calculates the data of the user computing power resources, while the embodiment of the present invention calculates the data between the computing power resources and the collaborative computing power resources.
[0125] S4. Utilize the computing power measurement value, the first computing power scheduling cost, and the second computing power scheduling cost to generate the computing power scheduling state of the task to be processed, and generate the computing power scheduling action of the task to be processed, and construct the action value function and state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action.
[0126] Furthermore, the embodiment of the present invention generates the computing power scheduling state of the task to be processed by utilizing the computing power measurement value, the first computing power scheduling cost and the second computing power scheduling cost, so as to add the computing power measurement value to the computing power scheduling state, and increase the consideration for data processing capability through the computing power measurement value, and secondly, the second computing power scheduling cost in the computing power scheduling state, and increase the consideration for collaborative processing capability through the second computing power scheduling cost.
[0127] Optionally, the process of generating the computing power scheduling state of the task to be processed by using the computing power measurement value, the first computing power scheduling cost and the second computing power scheduling cost refers to the process of using the computing power measurement value, the first computing power scheduling cost and the second computing power scheduling cost as the computing power scheduling state.
[0128] Among them, the computing power scheduling action refers to the action of selecting computing power resources and the action of selecting collaborative computing power resources. The action of selecting computing power resources refers to the action of selecting which computing power resource among the computing power resources. The action of selecting collaborative computing power resources refers to the action of selecting which collaborative computing power resource among the collaborative computing power resources. For example, a computing power scheduling action is to simultaneously select the 5th computing power resource and the 7th collaborative computing power resource. It should be noted that the computing power scheduling action includes a set of single actions and a set of double actions. A single action refers to the action of executing the selection of computing power resources, and a double action refers to the action of executing the selection of computing power resources and the action of selecting collaborative computing power resources.
[0129] In one embodiment of the present invention, constructing the action value function and the state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action includes: constructing the action value function of the task to be processed under the computing power scheduling state and the computing power scheduling action using the following formula:
[0130]
[0131] Q π (s,a)=r(s,a)+γΣ s′∈S P(s′|s,a)V π (s′)
[0132] Q π (s,a)=r t +γQ π (S t+1 ,A t+1 )
[0133] Among them, Q π (s,a)=r t +γQ π S t+1 ,A t+1 ) represents the action value function, S t , s represents the state selected from the computing power scheduling state at time t, A t , a indicates that the pending task has S t The action selected from the computing power scheduling action under the premise of the state, G t Represents the state s from time t t The sum of all reward decays from the start to the terminal state, represents the expected function, Q π (s,a) indicates that the task to be processed has S t Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t Action, and finally get reward G t The expectation of γ is the discount factor, Rt+k represents the reward obtained at time t+k, k represents the constant used to limit the upper limit of t, R t represents the reward obtained at time t, R t+2 represents the reward obtained at time t+2, G t+1 Represents the state s from time t+1 t+1 The sum of the decays of all rewards from the beginning to the end state, r(s,a) represents the reward of the pending task selecting action a from the computing power scheduling action under the premise of having state s, S represents the set of computing power scheduling states, S t+1 , s′ represents the state selected from the computing power scheduling state at time t+1, V π (s′) indicates that at time t+1, t+1 Under the premise of state, follow the strategy π to select the state value function of any action from the computing power scheduling action, P(s′|s,a) represents the state transition function from state s to state s′ after executing action a, r(s,a)=r t , Q π S t+1 ,A t+1 )=∑ s′∈S P(s′|s,a)V π (s′), Q π (S t+1 ,A t+1 ) indicates that the pending task has S t+1 Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t+1 Action, and finally get reward G t+1 expectations;
[0134] The state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action is constructed using the following formula:
[0135]
[0136] V π (s)=Σ a∈A π(a|s)Q π (s,a)
[0137] Among them, V π (s)=∑ a∈A π(a|s)Q π (s,a) represents the state value function, π(a|s) represents the probability of executing action a under the premise of state s, Q π (s,a) indicates that the task to be processed has S t Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions tAction, and finally get reward G t , A represents the set of computing power scheduling actions, S t , s represents the state selected from the computing power scheduling state at time t, A t , a indicates that the pending task has S t The action selected from the computing power scheduling action under the premise of the state, G t Represents the state s from time t t The sum of all reward decays from the start to the terminal state, Represents the expected function.
[0138] S5. Optimize the action value function and the state value function. After obtaining the optimized action value function and the optimized state value function, use the optimized action value function and the optimized state value function to determine the computing power scheduling result of the task to be processed on the AI server cluster.
[0139] In one embodiment of the present invention, the optimizing the action value function and the state value function to obtain the optimized action value function and the optimized state value function includes: optimizing the weight parameters in the action value function using the following formula to obtain the optimized weight parameters:
[0140] Q π (s, a) = r t +γQ π (S t+1 , A t+1 )
[0141] Q π (S t+1 A t+1 )=Q π (S t+1 , A t+1 , w)
[0142] δ t =Qπ(S t , A t ,w)-Q π (S t+1 , A t+1 , w)
[0143]
[0144] Among them, w' represents the optimized weight parameter, w represents the weight parameter, α represents the learning rate, δ t represents Q at time t π (S t , A t , w) and Q π (S t+1 , At+1 , w), S t , s represents the state selected from the computing power scheduling state at time t, A t , a indicates that the pending task has S t The action selected from the computing power scheduling action under the premise of having the state s, r(s, a) represents the reward of selecting action a from the computing power scheduling action under the premise of having the state s, r(s, a) = r t , Q π (s, a) = Q π (S t , A t , w);
[0145] The action value function and the state value function are optimized using the optimized weight parameters to obtain an optimized action value function and an optimized state value function.
[0146] Optionally, the process of optimizing the action value function and the state value function by using the optimized weight parameter to obtain the optimized action value function and the optimized state value function refers to bringing the optimized weight parameter into the above-mentioned action value function to update the corresponding Q π (s, a), and using the updated Q π (s, a) Update V π (s).
[0147] Optionally, the process of determining the computing power scheduling result of the task to be processed on the AI server cluster by using the optimized action value function and the optimized state value function refers to determining the updated Q by using the optimized action value function and the optimized state value function. π (s, a) and V π (s)After that, you can use Q π (s, a) and V π (s) is used to calculate the reward. When the final return is calculated through the reward, when the return is the highest, the action of scheduling the computing power resources corresponding to the return is used as the computing power scheduling result of the AI server cluster.
[0148] It can be seen that the embodiment of the present invention can consider whether the computing power resources have resource categories such as logical operation resources, parallel computing resources and neural network computing resources that can solve the task, rather than the number of resources, by selecting the computing power analysis indicators of the AI server cluster. The embodiment of the present invention queries the collaborative computing power resources when the computing power resources process the pending tasks to consider the problem of scheduling multiple computing power resources to collaboratively process user tasks. Furthermore, the embodiment of the present invention generates the computing power scheduling state of the pending tasks by using the computing power measurement value, the first computing power scheduling cost and the second computing power scheduling cost, so as to add the computing power measurement value to the computing power scheduling state, and increase the consideration of data processing capability through the computing power measurement value. Secondly, the second computing power scheduling cost in the computing power scheduling state increases the consideration of collaborative processing capability through the second computing power scheduling cost. Therefore, the computing power scheduling method and system for an AI server cluster proposed in the embodiment of the present invention can increase the analysis of the data processing capability and collaborative processing capability of computing power resources.
[0149] like Figure 2 The figure shows a functional module diagram of the computing power scheduling system of the AI server cluster of the present invention.
[0150] The computing power scheduling system 200 of the AI server cluster of the present invention can be installed in an electronic device. According to the functions implemented, the computing power scheduling system of the AI server cluster can include a computing power measurement module 201, a first calculation module 202, a second calculation module 203, a function construction module 204 and a result determination module 205. The module of the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0151] In the embodiment of the present invention, the functions of each module / unit are as follows:
[0152] The computing power measurement module 201 is used to select a computing power analysis indicator of the AI server cluster, and use the computing power analysis indicator to measure the computing power resources in the AI server cluster to obtain a computing power measurement value;
[0153] The first calculation module 202 is used to calculate the transmission delay and transmission energy consumption of the user's pending task from the user to the computing resource, calculate the processing delay and processing energy consumption of the computing resource processing the pending task, and calculate the first computing power scheduling cost of the computing resource for scheduling the pending task based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption;
[0154] The second calculation module 203 is used to query the collaborative computing resources when the computing resources process the task to be processed, and calculate the second computing power scheduling cost of the computing resources for scheduling the collaborative computing resources;
[0155] The function construction module 204 is used to generate the computing power scheduling state of the task to be processed by using the computing power measurement value, the first computing power scheduling cost and the second computing power scheduling cost, and generate the computing power scheduling action of the task to be processed, and construct the action value function and state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action;
[0156] The result determination module 205 is used to optimize the action value function and the state value function. After obtaining the optimized action value function and the optimized state value function, the optimized action value function and the optimized state value function are used to determine the computing power scheduling result of the task to be processed for the AI server cluster.
[0157] In detail, each module in the computing power scheduling system 200 of the AI server cluster in the embodiment of the present invention is used in the same manner as described above. Figure 1 The computing power scheduling method of the AI server cluster described in the previous section is the same technical means and can produce the same technical effects, so I will not go into details here.
[0158] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0159] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0160] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0161] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0162] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A computing power scheduling method for an AI server cluster, characterized in that: The method comprises: Selecting a computing power analysis indicator of an AI server cluster, and using the computing power analysis indicator to measure the computing power resources in the AI server cluster to obtain a computing power measurement value; Calculate the transmission delay and transmission energy consumption of the user's pending task from the user to the computing power resource, calculate the processing delay and processing energy consumption of the computing power resource processing the pending task, and calculate the first computing power scheduling cost of scheduling the computing power resource for the pending task based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption; Querying the collaborative computing resources when the computing resource processes the task to be processed, and calculating the second computing resource scheduling cost of the computing resource for scheduling the collaborative computing resource, wherein the calculating the second computing resource scheduling cost of the computing resource for scheduling the collaborative computing resource includes: Calculate the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource; Identifying remaining tasks to be processed among the tasks to be processed; Calculating the processing delay and processing energy consumption of the collaborative computing resources in processing the remaining tasks to be processed; The second computing power scheduling cost is calculated by using the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource, and the processing delay and processing energy consumption of the collaborative computing power resource in processing the remaining tasks to be processed; Generate a computing power scheduling state of the task to be processed by using the computing power measurement value, the first computing power scheduling cost, and the second computing power scheduling cost, generate a computing power scheduling action of the task to be processed, and construct an action value function and a state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action; After optimizing the action value function and the state value function to obtain the optimized action value function and the optimized state value function, the optimized action value function and the optimized state value function are used to determine the computing power scheduling result of the task to be processed on the AI server cluster.
2. The method according to claim 1, characterized in that The computing power analysis indicators of the selected AI server cluster include: Dividing the primary computing power indicators of the AI server cluster; Among them, the first-level computing power indicators include logical operation indicators, parallel computing indicators and neural network computing indicators; Divide the secondary computing power indicators of the primary computing power indicators; The secondary computing power index includes the number of logical operations, the duration of logical operations and the speed of logical operations of the logical operation index, the number of parallel calculations, the duration of parallel calculations and the speed of parallel calculations of the parallel calculation index, and the number of neural network calculations, the duration of neural network calculations and the speed of neural network calculations of the neural network calculation index; The number of logic operations, the speed of logic operations, the number of parallel calculations, the speed of parallel calculations, the number of neural network calculations, and the speed of neural network calculations are used as positive indicators of the AI server cluster, and the duration of logic operations, the duration of parallel calculations, and the duration of neural network calculations are used as negative indicators of the AI server cluster; The positive indicator and the negative indicator are used as computing power analysis indicators of the AI server cluster.
3. The method according to claim 1, characterized in that The using the computing power analysis indicator to measure the computing power resources in the AI server cluster to obtain a computing power measurement value includes: The indicator weight corresponding to the computing power analysis indicator is calculated using the following formula: Among them, W u represents the indicator weight of the u-th indicator, ω uv represents the vth sampling value of the uth computing power analysis indicator, V represents the total number of samples of the uth computing power analysis indicator, e u The information entropy of the weight of the u-th indicator, U represents the total number of computing power analysis indicators; Based on the indicator weight, the computing power measurement value corresponding to the computing power analysis indicator is calculated using the following formula: Among them, H represents the computing power measurement value.
4. The method according to claim 1, characterized in that: The calculation of the transmission delay and transmission energy consumption of the task to be processed by the user from the user to the computing power resource includes: The following formula is used to calculate the transmission delay of the task to be processed by the user from the user to the computing resource: Wherein, Δt1 represents the transmission delay, L represents the distance of the task to be processed from the user to the computing resource, v1 represents the data propagation rate on the L distance, B represents the data frame size of the task to be processed, and v2 represents the data transmission rate on the L distance; The following formula is used to calculate the transmission energy consumption of the task to be processed by the user from the user to the computing resource: E1=α+βv2 ε Among them, E1 represents the transmission energy consumption, α represents the energy consumption of the task to be processed when it is idle during the transmission from the user to the computing resource, and β and ε both represent positive integers used to fit the functional relationship between v2 and E1.
5. The method according to claim 1, characterized in that The calculating the processing delay and processing energy consumption of the computing resource for processing the task to be processed includes: The processing delay of the computing resource for processing the task to be processed is calculated using the following formula: Among them, Δt2 represents the processing delay, t1 represents the start time of the computing resource processing the pending task, t2 represents the time when the pending task arrives at the computing resource, and CPI i represents the number of clock cycles required for the ith instruction of the task to be processed, n represents the number of instructions of the task to be processed, T CPU The clock cycle length corresponding to the number of clock cycles required for the i-th instruction; The processing energy consumption of the computing resource for processing the task to be processed is calculated using the following formula: Among them, E2 represents the processing energy consumption, represents the end time of the computing resource processing the task to be processed, s(t) represents the electrical signal at time t when the computing resource processes the task to be processed, and α' represents the energy consumption when the computing resource is idle.
6. The method according to claim 1, characterized in that The querying of the collaborative computing resources when the computing resources process the pending tasks includes: Obtain other computing resources in the AI server cluster except the computing resources; Identifying response time limits of remaining tasks to be processed among the tasks to be processed; The waiting response time of the remaining tasks to be processed waiting for the other computing resources is calculated using the following formula: Δt3=t 1 +t 2 Among them, Δt3 represents the waiting response time, t 1 Indicates the transmission time of the remaining pending tasks from the computing resource to the waiting queue of the other computing resources, t 2 Indicates the time taken by the other computing resources to process the tasks in the waiting queue that are ahead of the remaining tasks to be processed; Determining whether the waiting response time exceeds the response time limit; When the waiting response time does not exceed the response time limit, the other computing power resources are used as the collaborative computing power resources.
7. The method according to claim 1, characterized in that The step of constructing the action value function and the state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action includes: The action value function of the task to be processed under the computing power scheduling state and the computing power scheduling action is constructed using the following formula: Q π (s,a)=r(s,a)+γ∑ s′∈S P(s′∣s,a)V π (s′) Q π (s,a)=r t +γQ π (S t+1 ,A t+1 ) Among them, Q π (s,a)=r t +γQ π (S t+1 ,A t+1 ) represents the action value function, s represents the state selected from the computing power scheduling state at time t, S t represents any state in the state set at time t, S t =s means that s and S at time t t Consistent, a means that the task to be processed has S t The action selected from the computing power scheduling action under the premise of the state, A t represents any action in the action set, A t =a means that a and A at time t t Consistent, G. t Represents the state s from time t t The sum of all reward decays from the start to the terminal state, represents the expected function, Q π (s,a) indicates that the task to be processed has S t Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t Action, and finally get reward G t The expectation of γ represents the discount factor, R t+1 represents the reward obtained at time t+1, R t+k represents the reward obtained at time t+k, k represents the constant used to limit the upper limit of t, R t represents the reward obtained at time t, R t+2 represents the reward obtained at time t+2, G t+1 Represents the state s from time t+1 t+1 The sum of the decays of all rewards from the start to the end state, r(s,a) represents the reward of the pending task selecting action a from the computing power scheduling action under the premise of having state s, S represents the set of computing power scheduling states, s′ represents the state selected from the computing power scheduling state at time t+1, S t+1 represents any state in the state set at time t+1, S t+1 =s′ means that s′ and S at time t+1 t+1 Consistent, V π (s′) indicates that at time t+1, t+1 Under the premise of state, follow the strategy π to select the state value function of any action from the computing power scheduling action, P(s′|s,a) represents the state transition function from state s to state s′ after executing action a, r(s,a)=r t , Q π (S t+1 ,A t+1 )=∑ s′∈S P(s′|s,a)V π (s′), Q π (S t+1 ,A t+1 ) indicates that the pending task has S t+1 Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t+1 Action, and finally get reward G t+1 expectations; The state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action is constructed using the following formula: V π (s)=∑ a∈A π(a∣s)Q π (s,a) Among them, V π (s)=∑ a∈A π(a|s)Q π (s,a) represents the state value function, π(a|s) represents the probability of executing action a under the premise of state s, and A represents the set of computing power scheduling actions.
8. The method according to claim 1, characterized in that The step of optimizing the action value function and the state value function to obtain an optimized action value function and an optimized state value function includes: The weight parameters in the action value function are optimized using the following formula to obtain the optimized weight parameters: Q π (s,a)=r t +γQ π (S t+1 ,A t+1 ) Q π (S t+1 ,A t+1 )=Q π (S t+1 ,A t+1 ,w) δ t =Q π (S t ,A t ,w)-Q π (S t+1 ,A t+1 ,w) Among them, w' represents the optimized weight parameter, w represents the weight parameter, α represents the learning rate, δ t represents Q at time t π (S t ,A t ,w) and Q π (S t+1 ,A t+1 ,w), r(s,a) represents the reward of the pending task selecting action a from the computing power scheduling action under the premise of having state s, r(s,a)=r t , Q π (s,a)=Q π (S t ,A t ,w),Q π (s,a) indicates that the task to be processed has S t Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t Action, and finally get reward G t The expectation of , γ represents the discount factor, Q π (S t+1 ,A t+1 ) indicates that the pending task has S t+1 Under the premise of the state, follow the strategy π to select A from the computing power scheduling actions t+1 action, the expectation of eventually getting a reward, Q π (S t ,A t ,w) indicates that the task to be processed has S t Under the premise of the state, follow the strategy π and select A from the computing power scheduling actions according to the weight parameter w. t action, the expectation of eventually getting a reward, Q π (S t+1 ,A t+1 ,w) indicates that the pending task has S t+1 Under the premise of the state, follow the strategy π and select A from the computing power scheduling actions according to the weight parameter w. t+1 action, with the expectation of eventual reward; The action value function and the state value function are optimized using the optimized weight parameters to obtain an optimized action value function and an optimized state value function.
9. A computing power scheduling system for an AI server cluster, characterized in that: The system comprises: A computing power measurement module, used to select a computing power analysis indicator of the AI server cluster, and use the computing power analysis indicator to measure the computing power resources in the AI server cluster to obtain a computing power measurement value; A first calculation module is used to calculate the transmission delay and transmission energy consumption of the pending task of the user from the user to the computing power resource, calculate the processing delay and processing energy consumption of the computing power resource processing the pending task, and calculate the first computing power scheduling cost of the computing power resource for scheduling the pending task based on the transmission delay, the transmission energy consumption, the processing delay and the processing energy consumption; The second calculation module is used to query the collaborative computing resources when the computing resource processes the task to be processed, and calculate the second computing resource scheduling cost of the computing resource scheduling the collaborative computing resource, wherein the calculation of the second computing resource scheduling cost of the computing resource scheduling the collaborative computing resource includes: Calculate the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource; Identifying remaining tasks to be processed among the tasks to be processed; Calculating the processing delay and processing energy consumption of the collaborative computing resources in processing the remaining tasks to be processed; The second computing power scheduling cost is calculated by using the transmission delay and transmission energy consumption from the computing power resource to the collaborative computing power resource, and the processing delay and processing energy consumption of the collaborative computing power resource in processing the remaining tasks to be processed; a function construction module, used to generate the computing power scheduling state of the task to be processed by using the computing power measurement value, the first computing power scheduling cost and the second computing power scheduling cost, and generate the computing power scheduling action of the task to be processed, and construct an action value function and a state value function of the task to be processed under the computing power scheduling state and the computing power scheduling action; The result determination module is used to optimize the action value function and the state value function, and after obtaining the optimized action value function and the optimized state value function, use the optimized action value function and the optimized state value function to determine the computing power scheduling result of the task to be processed on the AI server cluster.
Citation Information
Patent Citations
O-RAN-oriented multilevel heterogeneous resource scheduling method
CN118535304A
Cited By
AI large model-based computing power resource optimization method
CN120803718A