Scheduling method and device for server cluster, equipment, storage medium and product
By receiving service requests and counting the number of service requests within the set time, combining the number of underutilized nodes in the server cluster, the target scheduling strategy is determined, and a scheduling strategy based on service type or Q learning network is adopted, the resource waste caused by traditional scheduling algorithms is solved, and more efficient resource utilization and service performance is achieved.
Patent Information
- Application Number
- CN202510073134.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-16
AI Technical Summary
When allocating server resources, traditional dynamic scheduling algorithms are prone to inappropriate allocation due to the habits and resource consumption differences of different user service types, resulting in waste of resources.
By receiving service requests and counting the number of service requests within the set time, the number of the latest underutilized nodes in the server cluster is determined, and the target scheduling policy is determined based on the comparison results of the first number and the second number, and the service nodes are assigned to the service request based on the target scheduling policy. If the first number is greater than the second number, a scheduling strategy based on the service type is adopted; if the first number is less than or equal to the second number, a scheduling strategy based on the Q learning network is adopted.
By comprehensively considering service types and node utilization, we optimize the allocation of service nodes, reduce resource waste, and improve the service performance of the server cluster.
Smart Images

Figure CN119922192A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing, and in particular to a scheduling method, apparatus, device, storage medium and product for a server cluster. Background Art
[0002] Cloud computing is a way of providing computing resources and services through the Internet. Its core is to distribute computing tasks on a large number of distributed computers, such as cloud server clusters.
[0003] Cloud server cluster is a method to improve system performance, reliability and scalability through multiple servers working together. It is suitable for various application scenarios, including Web applications, databases, big data processing and high-performance computing. Through load balancing, high availability and elastic scaling, cloud server cluster can provide better user experience and higher system stability.
[0004] In the related art, the load balancing technology of the Nginx server cluster distributes high-concurrency requests to each node in the server cluster reasonably through a scheduling algorithm. The Nginx server cluster usually refers to multiple Nginx server instances forming a whole, and distributes client requests through a load balancing device or the reverse proxy function of Nginx itself. Many dynamic scheduling algorithms have appeared in the field of load balancing technology. The dynamic load balancing algorithm collects the performance indicators of the server and dynamically distributes user requests through a certain mathematical model calculation, such as the minimum connection number algorithm, the weighted minimum connection number algorithm, the fastest response speed algorithm, the dynamic performance allocation algorithm, the fastest mode algorithm, the service type algorithm, the prediction mode algorithm, etc.
[0005] However, different users have different habits in using service types, and different service types have large differences in resource consumption. Traditional dynamic scheduling algorithms are prone to inappropriate allocation, resulting in partial waste of resources. Summary of the invention
[0006] In view of this, embodiments of the present application provide a scheduling method, apparatus, device, storage medium and product for a server cluster, aiming to reduce waste of resources.
[0007] The technical solution of the embodiment of the present application is implemented as follows:
[0008] In a first aspect, an embodiment of the present application provides a scheduling method for a server cluster, comprising:
[0009] Receive service requests and count a first number of service requests within a set time period; wherein each of the service requests at least includes first information indicating a service type;
[0010] determining a second number of most recent underutilized nodes in the server cluster;
[0011] Determining a target scheduling strategy for the server cluster based on a comparison result of the first quantity and the second quantity; wherein the target scheduling strategy includes an evaluation parameter corresponding to the service type;
[0012] A service node is allocated to each of the service requests based on the target scheduling policy.
[0013] In the above solution, determining the target scheduling strategy of the server cluster based on the comparison result of the first number and the second number includes:
[0014] If it is determined that the first number is greater than the second number, determining that the target scheduling strategy of the server cluster is a first target scheduling strategy for allocating service nodes based on service type;
[0015] If it is determined that the first number is less than or equal to the second number, the target scheduling strategy of the server cluster is determined to be a second target scheduling strategy for allocating service nodes based on a Q learning network.
[0016] In the above solution, if the target scheduling strategy is the first target scheduling strategy, allocating a service node to each service request based on the target scheduling strategy includes:
[0017] Determine a first service request set and a second service request set based on the first information of each service request, wherein the first service request set includes service requests corresponding to a first service type with a reception processing delay, and the second service request set includes service requests corresponding to a second service type that needs to be processed in a timely manner;
[0018] allocating service nodes to the service requests in the first service request set based on a first node allocation strategy corresponding to the first service type;
[0019] allocating service nodes to the service requests in the second service request set based on a second node allocation strategy corresponding to the second service type;
[0020] Among them, the first node allocation strategy is used to calculate the profit corresponding to the allocation method based on the discount coefficient corresponding to the processing delay, and the second node allocation strategy is used to calculate the profit corresponding to the allocation method based on the revenue of processed service requests and the penalty of unprocessed service requests.
[0021] In the above solution, if the target scheduling strategy is the second target scheduling strategy, allocating a service node to each service request based on the target scheduling strategy includes:
[0022] The service node corresponding to each service request is determined based on the trained Q learning network, and the service request is responded to based on the corresponding service node.
[0023] In the above scheme, the method further comprises:
[0024] The Q learning network is trained based on the state set and action set of the configured agent.
[0025] In the above scheme, the state set and action set of the configured agent are used to train the Q learning network, including:
[0026] Set the initialization parameters of the Q learning network;
[0027] Based on the configured state set and action set, a roulette wheel method is used to select an action, and a selection probability is calculated according to a probability optimization algorithm, and the Q learning network is trained based on the selection probability until a trained Q learning network is obtained.
[0028] In the above solution, determining the latest second number of underutilized nodes in the server cluster includes:
[0029] Based on the current load rate of each node in the service cluster and a set threshold, the latest underutilized node is determined, and the second number is obtained based on the latest underutilized node.
[0030] In a second aspect, an embodiment of the present application provides a scheduling device for a server cluster, including:
[0031] A first processing module, configured to receive service requests and count a first number of service requests within a set time period; wherein each of the service requests at least includes first information indicating a service type;
[0032] A second processing module, configured to determine a second number of latest underutilized nodes in the server cluster;
[0033] A determination module, configured to determine a target scheduling strategy for the server cluster based on a comparison result of the first quantity and the second quantity; wherein the target scheduling strategy includes an evaluation parameter corresponding to the service type;
[0034] The scheduling module is used to allocate a service node to each service request based on the target scheduling strategy.
[0035] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor, when used to run the computer program, executes the steps of the method described in the first aspect of the embodiment of the present application.
[0036] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect of the embodiment of the present application are implemented.
[0037] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method described in the first aspect of the embodiment of the present application.
[0038] The technical solution provided by the embodiment of the present application receives a service request and counts a first number of service requests within a set time period; wherein each service request includes at least first information indicating a service type; determines a second number of the latest underutilized nodes in the server cluster; determines a target scheduling strategy for the server cluster based on a comparison result of the first number and the second number; wherein the target scheduling strategy includes an evaluation parameter corresponding to the service type; and allocates a service node to each service request based on the target scheduling strategy. Since the embodiment of the present application takes into account the service type in the service request and the comparison result of the first number and the second number, a service node can be allocated to the service request in the server cluster on the basis of comprehensively considering the service type and the comparison result of the first number and the second number, thereby achieving optimal allocation of service nodes based on considering differences in resource consumption by different service types, thereby reducing resource waste and improving the service performance of the server cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A flow chart of a scheduling method for a server cluster according to an embodiment of the present application;
[0040] Figure 2 This is a schematic diagram of the structure of a scheduling device for a server cluster according to an embodiment of the present application;
[0041] Figure 3 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0042] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0044] The present application embodiment provides a scheduling method for a server cluster, which can be applied to a scheduling device of a server cluster, for example, a load balancing device or a reverse proxy device of a server cluster, such as Figure 1 As shown, the scheduling method includes:
[0045] Step 101, receiving service requests and counting a first number of service requests within a set time period; wherein each of the service requests at least includes first information indicating a service type.
[0046] Here, the scheduling device can receive service requests sent by multiple clients and count the first number of service requests within a set time period. The set time period can be reasonably set based on the processing delay requirements. For example, for service requests with higher real-time requirements, the set time period can be set shorter. For scenarios with higher resource balance requirements, the set time period can be set slightly longer. The embodiments of the present application do not make specific limitations on this.
[0047] Here, the service types include at least a first service type with a reception processing delay and a second service type that needs to be processed in a timely manner. For service requests of the first service type, the scheduling device may delay processing of this type of service requests when the number of service nodes is less than the service requests; for service requests of the second service type, the scheduling device needs to process this type of service requests in a timely manner after receiving the service requests.
[0048] Step 102: Determine the latest second number of underutilized nodes in the server cluster.
[0049] Here, the scheduling device may determine the latest second number of underutilized nodes in the server cluster before scheduling the service request.
[0050] Exemplarily, determining the latest second number of underutilized nodes in the server cluster includes:
[0051] Based on the current load rate of each node in the service cluster and a set threshold, the latest underutilized node is determined, and the second number is obtained based on the latest underutilized node.
[0052] It should be noted that in the related art, the scheduling device often only identifies the nodes that are initially underutilized in the server cluster, and subsequently allocates user requests to these nodes. However, this may eventually cause the initial underutilized nodes to be overloaded, because over time, more user requests will be allocated to these initial underutilized nodes, thereby affecting the performance and resource utilization of the server cluster.
[0053] In an embodiment of the present application, the second number of the latest underutilized nodes in the service cluster can be continuously identified based on the set frequency, and the second number can be used as a basis for subsequent scheduling to improve the overall performance of the server cluster.
[0054] Step 103: determining a target scheduling strategy for the server cluster based on a comparison result of the first quantity and the second quantity; wherein the target scheduling strategy includes an evaluation parameter corresponding to the service type.
[0055] It should be noted that, since the target scheduling strategy of the embodiment of the present application includes evaluation parameters corresponding to the service type, it can better balance the resource consumption of different service types during specific service scheduling, thereby reducing resource waste.
[0056] Step 104: Allocate a service node for each of the service requests based on the target scheduling policy.
[0057] It can be understood that since the embodiments of the present application take into account the service type in the service request and the comparison result of the first quantity and the second quantity, service nodes can be allocated to the service requests in the server cluster based on comprehensive consideration of the service type and the comparison result of the first quantity and the second quantity. Therefore, it is possible to achieve optimal allocation of service nodes based on differences in resource consumption due to different service types, thereby reducing resource waste and improving the service performance of the server cluster.
[0058] Exemplarily, determining the target scheduling strategy of the server cluster based on the comparison result of the first quantity and the second quantity includes:
[0059] If it is determined that the first number is greater than the second number, determining that the target scheduling strategy of the server cluster is a first target scheduling strategy for allocating service nodes based on service type;
[0060] If it is determined that the first number is less than or equal to the second number, the target scheduling strategy of the server cluster is determined to be a second target scheduling strategy for allocating service nodes based on a Q learning network.
[0061] Here, the scheduling device selects different target scheduling strategies based on the comparison result of the first number and the second number. Specifically, if the first number is greater than the second number, it indicates that some nodes need to process multiple service requests, and the first target scheduling strategy for allocating service nodes based on service type is adopted; if the first number is less than or equal to the second number, the second target scheduling strategy for allocating service nodes based on Q learning network is adopted; in this way, based on the classification of target scheduling strategies, the dynamic allocation of service nodes can be better realized, reducing the waste of resources of the server cluster.
[0062] Exemplarily, if the target scheduling policy is the first target scheduling policy, allocating a service node to each of the service requests based on the target scheduling policy includes:
[0063] Determine a first service request set and a second service request set based on the first information of each service request, wherein the first service request set includes service requests corresponding to a first service type with a reception processing delay, and the second service request set includes service requests corresponding to a second service type that needs to be processed in a timely manner;
[0064] allocating service nodes to the service requests in the first service request set based on a first node allocation strategy corresponding to the first service type;
[0065] allocating service nodes to the service requests in the second service request set based on a second node allocation strategy corresponding to the second service type;
[0066] Among them, the first node allocation strategy is used to calculate the profit corresponding to the allocation method based on the discount coefficient corresponding to the processing delay, and the second node allocation strategy is used to calculate the profit corresponding to the allocation method based on the revenue of processed service requests and the penalty of unprocessed service requests.
[0067] Here, the scheduling device can identify the service type corresponding to each service request based on the first information, and classify each service request based on the corresponding service type to obtain a first service request set and a second service request set, and then schedule the service nodes for the service requests in the first service request set based on the first node allocation strategy, and schedule the service nodes for the service requests in the second service request set based on the second node allocation strategy, thereby achieving optimal scheduling of service nodes while taking into account both full resource utilization and profit maximization.
[0068] Exemplarily, if the target scheduling policy is the second target scheduling policy, allocating a service node to each of the service requests based on the target scheduling policy includes:
[0069] The service node corresponding to each service request is determined based on the trained Q learning network, and the service request is responded to based on the corresponding service node.
[0070] Here, the Q-learning network algorithm is a type of reinforcement learning, or more precisely, a way of selecting strategies. The core and training goal of reinforcement learning is to select a suitable strategy (Policy) so that the sum of rewards (incentive values) obtained at the end of each epoch (training round) is maximized. The idea of the Q-learning network is: Q(S, A) = the sum of the reward values that will be obtained in the future after taking action A in state S. By maintaining a Q-table, it records what Q value can be obtained when taking what action in each state. In this way, as long as the program continuously updates this table during operation so that it can eventually converge, the algorithm can get the state by looking up the table to determine what behavior it should choose in order to obtain the maximum Q value, thereby realizing the intelligent selection of actions.
[0071] Here, since the service type in the service request is used as a dimension in the state set, and the user's service request is classified according to the service type, the Q value in the corresponding Q-table can be updated after the service request occurs, that is, the corresponding incentive value under the state and action is updated. In this way, the algorithm of the Q learning network can remember the user's habit of using service types, use the accumulated Q values to predict the user's future behavior, and improve resource utilization.
[0072] Exemplarily, the method further includes:
[0073] The Q learning network is trained based on the state set and action set of the configured agent.
[0074] Exemplarily, the training of the Q learning network based on the state set and action set of the configured agent includes:
[0075] Set the initialization parameters of the Q learning network;
[0076] Based on the configured state set and action set, a roulette wheel method is used to select an action, and a selection probability is calculated according to a probability optimization algorithm, and the Q learning network is trained based on the selection probability until a trained Q learning network is obtained.
[0077] In the embodiment of the present application, the random selection of actions in the greedy strategy is improved to selecting actions according to the roulette method, and the selection probability is calculated according to the probability optimization algorithm, which ensures the accuracy of the selection and is conducive to improving the selection performance of the Q learning network.
[0078] The present application is further described in detail below in conjunction with an application example.
[0079] This application embodiment provides an adaptive dynamic scheduling method applicable to an Nginx server cluster, which includes the following steps:
[0080] Step 1: Receive a user request (ie, a service request), wherein the user request includes the user's service type information (ie, first information).
[0081] Step 2: Determine the latest underutilized nodes and their number in the Nginx server cluster.
[0082] Here, the latest underutilized nodes and their number in the service cluster may be identified based on a set frequency, wherein the underutilized nodes may be nodes whose current load rates are less than a set threshold.
[0083] Step 3: Determine whether the number of nodes of the latest underutilized nodes is greater than the number requested by the user. If so, execute step 4; if not, execute step 5.
[0084] Here, according to the comparison result between the number of nodes that are not fully utilized and the number of user requests, the corresponding target scheduling strategy is selected, which can be divided into the following two cases:
[0085] Case 1: If the number of nodes of the latest underutilized nodes is less than the number of user requests, then the following step 4 is executed to determine the allocation method of each user request.
[0086] Case 2: If the number of nodes of the latest underutilized nodes is greater than or equal to the number of user requests, then the following step 5 is executed to determine the allocation method of each user request.
[0087] Step 4: Determine the processing method for each user request respectively, and then allocate the corresponding service node for processing according to the processing method and the preset allocation method.
[0088] If the current latest number of underutilized nodes is less than the number of user requests, it means that some nodes need to process multiple user requests. When the same node processes multiple requests, it is bound to cause processing delays for some user requests. Therefore, in this solution, the user requests that can receive processing delays (corresponding to the aforementioned first service type) and the user requests that cannot receive processing delays (corresponding to the aforementioned second service type) are first determined from the received user requests. It should be noted that the inability to receive processing delays here does not mean that no delays can be received, but only means that the user requests that need to be promptly assigned to the corresponding node for processing when the user request is received. Then, by controlling the processing amount of the user request, the processing delay of the user request is compensated, and finally the user request is assigned to the best service node for processing.
[0089] In this application embodiment, it is assumed that the service node in the Nginx server cluster needs the user to pay the processing amount when processing the user request, which is specifically divided into the following two situations:
[0090] 1) When the service node processes the user request immediately upon receiving it (i.e., the service request of the second service type), the corresponding processing amount is:
[0091]
[0092] in, It represents the processing amount, f0 represents the basic processing amount, ft represents the billing rate corresponding to the time consumption, tp represents the time consumed to process the user request, fd represents the billing rate corresponding to the transmission distance when the service node returns the processing result of the user request, and dp represents the transmission distance when the service node returns the processing result of the user request.
[0093] Correspondingly, the optimal service node allocation method for user requests in the above situation 1) is as follows:
[0094]
[0095] xp,v∈{0,1}
[0096] Where v represents any service node, p represents any user request, and Vp represents the service node that can process user request p; represents the cost of processing user request p by service node v. xp,v represents a binary variable with a value of {0,1}, which is used to indicate whether the service node is processing user requests. For example, if xp,v=1, it is considered that service node v is currently assigned to process user request p. |P| represents the number of user requests waiting to be processed. δ represents the profit parameter of the preset service node; γ represents the fine amount when the preset service node does not process the user request.
[0097] 2) When a service node receives a user request, it needs to process other user requests before processing this user request. The corresponding processing amount is:
[0098]
[0099] in, It represents the processing amount required when the service node processes the user request according to the aforementioned situation 1); φp represents the discount coefficient of the processing amount when the service node processes the user request according to situation 2), and the discount coefficient φp is strictly less than 1.
[0100] Exemplarily, the calculation formula of the discount coefficient φp is as follows:
[0101]
[0102] Among them, α represents the preset discount coefficient; b represents the basic rate of the service node for processing user requests, which is used to ensure that the service node can obtain processing revenue when processing user requests under discount conditions. Indicates the latency index for processing user requests; Indicates the concurrency coefficient of user requests processed by the service node.
[0103] For example, the delay index The calculation formula is as follows:
[0104]
[0105] in, Indicates the processing delay of the service node processing the user request after processing other user requests; Indicates the processing delay when the service node directly processes the user request without processing other user requests.
[0106] For example, the concurrency coefficient The calculation formula is as follows:
[0107]
[0108] in, kλ represents the total number of user requests processed, λ represents the concurrent processing node, and yp,λ represents a binary parameter, which is equal to 1 if the node concurrently processes request p and 0 otherwise.
[0109] Correspondingly, the optimal service node allocation method for user requests in the above situation 2) is as follows:
[0110]
[0111] xp,v∈{0,1}
[0112] Among them, ζ represents the penalty amount when the preset service node fails to process the user request, π P,v The calculation formula is as follows:
[0113]
[0114] Wherein, P represents the first service request set (i.e., the set of user requests with delayed reception and processing), represents the second service request set (i.e., the set of user requests that do not receive processing delays), Indicates the total amount of the first service request, Indicates the total amount of the second service request.
[0115] Step 5: Based on the improved Q-Learning adaptive dynamic scheduling algorithm, each user request is assigned a corresponding service node for processing.
[0116] The specific steps are as follows:
[0117] S1: Initialization parameters, such as Q-table, α, γ, ε0, t, m, n.
[0118] In step S1, the Q table is set to a zero matrix, and the learning rate α, the decay coefficient γ, the initial greedy coefficient ε0, the current step number t, the target step number m, and the number of nodes n are initialized.
[0119] S2: Set the State set (state set) and Action set (action set) of the agent.
[0120] For example, in step S2, the State set is defined as a combination of segmented real-time load and service type. The dimension is Where Dt(Si) is the real-time load of each node, Rk is the service type, and n is the number of nodes. The Action set is defined as the set of nodes, at = {S1, S2, ... Sn}, which means that the action is to select one of the nodes.
[0121] S3: At time t, initialize the agent’s state st.
[0122] S4: The agent selects action at according to the action selection strategy defined by the formula and executes the action.
[0123] For example, in step S4, the load-adaptive ε-greedy strategy is used as the selection strategy. In the greedy strategy of this solution, a 0-1 random number r is generated, compared with the greedy coefficient ε, and then action a is selected according to the formula, which is defined as follows:
[0124]
[0125] Among them, π(st,at) represents the action selection strategy, Roulette(at) represents the action selection according to the roulette method, and max Q(st,at) represents the action selection according to the maximum Q value principle. In order to enable the agent to explore more actions in the early stage and tend to choose better actions in the later stage, the decay t greedy method is used to improve the greedy coefficient ε. At the beginning, the value of ε is large, which means that the probability of taking random actions according to the roulette is large. At the same time, as the value of ε decays during the learning process, the action is more and more inclined to take actions according to the memorized Q value. The formula definition is as follows:
[0126]
[0127] Furthermore, the random selection action random(at) is improved to the roulette selection action Roulette(at), where the probability of each node being selected in the roulette method is Pt(Si). The probability optimization algorithm is introduced here, where Dt(Si) is the real-time load evaluation index of node Si. The formula definition of the probability optimization algorithm is as follows:
[0128]
[0129] S5: The agent obtains the environmental reward function Rt(Si) corresponding to the current action according to the formula, and calculates the Q-value function iterative update formula according to the value of the environmental reward function.
[0130] Exemplarily, in step S5, the environment reward function is defined as Rt(Si), and an environment reward function based on the minimum entropy increase principle is introduced, where D0(Si) is the initial comprehensive load evaluation index of each node in the server cluster, and the formula definition of the environment reward function is as follows:
[0131]
[0132] Furthermore, the formula definition of the initial comprehensive load evaluation index D0(Si) is as follows:
[0133]
[0134] Among them, the performance evaluation weight of each server node is Where j∈{C,M,I,N}, Wj(Si) is an initial performance value of node Si in the server cluster. The load evaluation indicators of each server node are used for calculation: CPU (C), memory (M), disk I / O (I) and network bandwidth (N), σ C is the weight coefficient corresponding to the CPU, σ M is the weight coefficient corresponding to the memory, σ I is the weight coefficient corresponding to disk I / O, σ N is the weight coefficient corresponding to the network bandwidth.
[0135] S6: The agent updates the Q-value function corresponding to the current state and action according to the formula;
[0136] Further, in step S6, the Q value function iterative update formula corresponding to the current state and action is updated according to the iterative update formula, where α is the learning rate and γ is the attenuation coefficient. The formula definition is as follows:
[0137]
[0138] S7: Determine the next state The dimension is The segmented real-time load Dt(Si) of each node is determined according to the real-time load of each node, and the service type Rk is determined according to the next request of the user.
[0139] Further, in step S7, the load evaluation index change rate △Di is used to determine whether the load information of a node Si is uploaded within the time period T. When △Di does not exceed the threshold △d, Dt +1 (Si) = Dt(Si), when △Di exceeds the threshold △d, Dt +1 (Si) Recalculate based on real-time comprehensive load evaluation indicators.
[0140] Furthermore, the real-time comprehensive load evaluation index is calculated using the load evaluation index of each server node: CPU (C), memory (M), disk I / O (I) and network bandwidth (N). When there are n nodes in the cluster, for each node Si∈{S1, S2, ... Sn}, n>1, the server comprehensive performance evaluation index is D(Si)∈{D(S1), D(S2), ...D(Sn)}, n>1. According to the performance of each item, the evaluation index can be set as D(Si)=σ C ×C C (Si)+σ M ×C M (Si)+σ I ×C I (Si)+σ N ×C N (Si), where the performance evaluation weight of each server node is Where j∈{C,M,I,N}, Wj(Si) is an initial performance value of node Si in the server cluster.
[0141] S8: until the target number of steps is reached, otherwise jump to step S4;
[0142] S9: t←t+1, jump to step S3.
[0143] In order to implement the method of the embodiment of the present application, the embodiment of the present application also provides a scheduling device for a server cluster, which corresponds to the above-mentioned scheduling method for a server cluster, and each step in the above-mentioned scheduling method embodiment for a server cluster is also fully applicable to the embodiment of the scheduling device for a server cluster.
[0144] like Figure 2As shown, the scheduling device for the server cluster includes: a first processing module 201, a second processing module 202, a determination module 203 and a scheduling module 204. The first processing module 201 is used to receive service requests and count the first number of service requests within a set time period; wherein each service request at least includes first information indicating a service type; the second processing module 202 is used to determine the second number of the latest underutilized nodes in the server cluster; the determination module 203 is used to determine the target scheduling strategy of the server cluster based on the comparison result of the first number and the second number; wherein the target scheduling strategy includes an evaluation parameter corresponding to the service type; the scheduling module 204 is used to allocate a service node to each service request based on the target scheduling strategy.
[0145] Exemplarily, the determination module 203 is specifically configured to:
[0146] If it is determined that the first number is greater than the second number, determining that the target scheduling strategy of the server cluster is a first target scheduling strategy for allocating service nodes based on service type;
[0147] If it is determined that the first number is less than or equal to the second number, the target scheduling strategy of the server cluster is determined to be a second target scheduling strategy for allocating service nodes based on a Q learning network.
[0148] Exemplarily, if the target scheduling strategy is the first target scheduling strategy, the scheduling module 204 is specifically configured to:
[0149] Determine a first service request set and a second service request set based on the first information of each service request, wherein the first service request set includes service requests corresponding to a first service type with a reception processing delay, and the second service request set includes service requests corresponding to a second service type that needs to be processed in a timely manner;
[0150] allocating service nodes to the service requests in the first service request set based on a first node allocation strategy corresponding to the first service type;
[0151] allocating service nodes to the service requests in the second service request set based on a second node allocation strategy corresponding to the second service type;
[0152] Among them, the first node allocation strategy is used to calculate the profit corresponding to the allocation method based on the discount coefficient corresponding to the processing delay, and the second node allocation strategy is used to calculate the profit corresponding to the allocation method based on the revenue of processed service requests and the penalty of unprocessed service requests.
[0153] Exemplarily, if the target scheduling strategy is the second target scheduling strategy, the scheduling module 204 is specifically configured to:
[0154] The service node corresponding to each service request is determined based on the trained Q learning network, and the service request is responded to based on the corresponding service node.
[0155] Exemplarily, the scheduling device further includes: a training module 205, which is used to train the Q learning network based on the state set and action set of the configured intelligent agent.
[0156] Exemplarily, the training module 205 is specifically used for:
[0157] Set the initialization parameters of the Q learning network;
[0158] Based on the configured state set and action set, a roulette wheel method is used to select an action, and a selection probability is calculated according to a probability optimization algorithm, and the Q learning network is trained based on the selection probability until a trained Q learning network is obtained.
[0159] Exemplarily, the second processing module 202 is specifically configured to:
[0160] Based on the current load rate of each node in the service cluster and a set threshold, the latest underutilized node is determined, and the second number is obtained based on the latest underutilized node.
[0161] In actual application, the first processing module 201, the second processing module 202, the determination module 203, the scheduling module 204 and the training module 205 can be implemented by a processor in the scheduling device. Of course, the processor needs to run the computer program in the memory to implement its functions.
[0162] It should be noted that: the scheduling device provided in the above embodiment only uses the division of the above program modules as an example when performing scheduling. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above. In addition, the scheduling device provided in the above embodiment and the scheduling method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0163] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiment of the present application, the embodiment of the present application also provides an electronic device, namely the aforementioned scheduling device. Figure 3 Only an exemplary structure of the electronic device is shown, not all structures, and it can be implemented as needed. Figure 3 Partial or complete structure shown.
[0164] like Figure 3As shown, the electronic device 300 provided in the embodiment of the present application includes: at least one processor 301, a memory 302, a user interface 303 and at least one network interface 304. The various components in the electronic device 300 are coupled together through a bus system 305. It can be understood that the bus system 305 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 305 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, in Figure 3 Various buses are labeled as bus system 305 .
[0165] The user interface 303 may include a display, a keyboard, a mouse, a trackball, a click wheel, keys, buttons, a touch pad or a touch screen.
[0166] The memory 302 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer program used to operate on the electronic device.
[0167] The scheduling method disclosed in the embodiment of the present application can be applied to the processor 301, or implemented by the processor 301. The processor 301 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the scheduling method can be completed by the hardware integrated logic circuit or software instructions in the processor 301. The above-mentioned processor 301 can be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 301 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiment of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 302. The processor 301 reads the information in the memory 302 and completes the steps of the scheduling method provided in the embodiment of the present application in combination with its hardware.
[0168] In an exemplary embodiment, the electronic device may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.
[0169] It can be understood that the memory 302 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.
[0170] In an exemplary embodiment, the present application also provides a computer storage medium, which can be a computer-readable storage medium, for example, a memory 302 storing a computer program, and the computer program can be executed by a processor 301 of an electronic device to complete the steps described in the method of the present application embodiment. The computer-readable storage medium can be a memory such as a ROM, a PROM, an EPROM, an EEPROM, a Flash Memory, a magnetic surface memory, an optical disk, or a CD-ROM.
[0171] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program, which can be executed by the processor 301 of the electronic device 300 to complete the steps described in the method of the embodiment of the present application.
[0172] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0173] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0174] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A scheduling method for a server cluster, characterized in that: include: Receive service requests and count a first number of service requests within a set time period; wherein each of the service requests at least includes first information indicating a service type; determining a second number of most recent underutilized nodes in the server cluster; Determining a target scheduling strategy for the server cluster based on a comparison result of the first quantity and the second quantity; wherein the target scheduling strategy includes an evaluation parameter corresponding to the service type; A service node is allocated to each of the service requests based on the target scheduling policy.
2. The method according to claim 1, characterized in that The determining the target scheduling strategy of the server cluster based on the comparison result of the first quantity and the second quantity includes: If it is determined that the first number is greater than the second number, determining that the target scheduling strategy of the server cluster is a first target scheduling strategy for allocating service nodes based on service type; If it is determined that the first number is less than or equal to the second number, the target scheduling strategy of the server cluster is determined to be a second target scheduling strategy for allocating service nodes based on a Q learning network.
3. The method according to claim 2, characterized in that If the target scheduling policy is the first target scheduling policy, allocating a service node for each service request based on the target scheduling policy includes: Determine a first service request set and a second service request set based on the first information of each service request, wherein the first service request set includes service requests corresponding to a first service type with a reception processing delay, and the second service request set includes service requests corresponding to a second service type that needs to be processed in a timely manner; allocating service nodes to the service requests in the first service request set based on a first node allocation strategy corresponding to the first service type; allocating service nodes to the service requests in the second service request set based on a second node allocation strategy corresponding to the second service type; Among them, the first node allocation strategy is used to calculate the profit corresponding to the allocation method based on the discount coefficient corresponding to the processing delay, and the second node allocation strategy is used to calculate the profit corresponding to the allocation method based on the revenue of processed service requests and the penalty of unprocessed service requests.
4. The method according to claim 2, characterized in that: If the target scheduling policy is the second target scheduling policy, allocating a service node to each of the service requests based on the target scheduling policy includes: The service node corresponding to each service request is determined based on the trained Q learning network, and the service request is responded to based on the corresponding service node.
5. The method according to claim 4, characterized in that The method further comprises: The Q learning network is trained based on the state set and action set of the configured agent.
6. The method according to claim 5, characterized in that The state set and action set of the configured agent are used to train the Q learning network, including: Set the initialization parameters of the Q learning network; Based on the configured state set and action set, a roulette wheel method is used to select an action, and a selection probability is calculated according to a probability optimization algorithm, and the Q learning network is trained based on the selection probability until a trained Q learning network is obtained.
7. The method according to claim 1, characterized in that The determining of the latest second number of underutilized nodes in the server cluster comprises: Based on the current load rate of each node in the service cluster and a set threshold, the latest underutilized node is determined, and the second number is obtained based on the latest underutilized node.
8. A scheduling device for a server cluster, characterized in that: include: A first processing module, configured to receive service requests and count a first number of service requests within a set time period; wherein each of the service requests at least includes first information indicating a service type; A second processing module, configured to determine a second number of latest underutilized nodes in the server cluster; A determination module, configured to determine a target scheduling strategy for the server cluster based on a comparison result of the first quantity and the second quantity; wherein the target scheduling strategy includes an evaluation parameter corresponding to the service type; The scheduling module is used to allocate a service node to each service request based on the target scheduling strategy.
9. An electronic device, characterized in that: include: A processor and a memory for storing a computer program that can be executed on the processor, wherein: The processor is used to execute the steps of the method according to any one of claims 1 to 7 when running a computer program.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method, device and equipment for dispatching GPU resource and computer readable storage medium
CN108363623A
Kubernetes resource scheduling method and device and electronic equipment
CN116991589A
TR-DQN-based high-performance computing cluster resource scheduling method and system
CN117591273A
Cloud traffic scheduling method, device, equipment, medium and computer program product
CN118869698A
Method and device for processing service access
US20160301765A1
Cited By
Server scheduling method, device, equipment, medium and computer program product
CN120371542A