A business server and a dedicated server scheduling method

Through the deep reinforcement learning framework and traffic prediction model, combined with multi-data centers and security factors, the problem of not being able to cope with dynamic business loads and lack of foresight in the existing technology is solved, intelligent server scheduling and resource planning are realized, and system performance and business security are improved.

CN119961002BActive Publication Date: 2025-07-01SHENZHEN ZRT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510436266.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-01
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing technology cannot effectively deal with dynamically changing business loads, lacks comprehensive perception of multi-dimensional system state, is difficult to balance multiple business goals, and lacks predictiveness of burst traffic.

Method used

Adopting the deep reinforcement learning framework, the deep reinforcement learning agent is initialized, and through the interaction and learning between the agent and the environment, we adaptively understand the dynamic changes in business characteristics and server status, and realize intelligent scheduling decisions. At the same time, traffic prediction models are introduced to build geographic information, network latency and energy cost information of multi-data centers, and integrate business security level information and server security threat information.

Benefits of technology

Intelligent scheduling decisions have been realized, and multiple performance indicators such as processing efficiency, resource utilization, service delay and load balancing have been comprehensively considered. Business peaks are predicted in advance and resource planning is carried out to reduce operating costs and ensure the safe and stable operation of the business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961002B_ABST
    Figure CN119961002B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of computer data processing, and discloses a service server and a dedicated server scheduling method, including: initializing a deep reinforcement learning agent; interacting the agent with the environment to obtain environmental feedback rewards and new states; storing the transfer data of the interaction between the agent and the environment in an experience pool and sampling; calculating the target Q value and loss function based on the sampled data, and updating the network parameters; performing policy optimization and model deployment according to preset conditions; The present invention adopts a deep reinforcement learning framework. Through the continuous interaction and learning between the agent and the environment, it can adaptively understand the dynamic changes of business characteristics and server states, realize intelligent scheduling decisions, and at the same time sets a multi-dimensional state space and reward function, comprehensively considering multiple performance indicators such as processing efficiency, resource utilization rate, service latency, and load balancing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data processing, and more specifically, to a business server and a dedicated server scheduling method. Background Art

[0002] With the continuous improvement of the enterprise informatization level, the complexity of business systems and the scale of servers are both growing rapidly; in the actual operation and maintenance process, administrators need to handle the resource scheduling problems of business servers and dedicated servers (such as database servers, cache servers, etc.) at the same time. Currently, the mainstream server scheduling solutions mainly adopt rule-based static scheduling strategies, such as scheduling based on server load thresholds, scheduling based on business priorities, etc. However, these traditional methods have the following technical problems:

[0003] Unable to effectively cope with dynamic business loads. During the business peak period, static scheduling rules often cannot adjust resource allocation in a timely manner, resulting in some servers being overloaded while other servers' resources are idle, affecting the overall system performance;

[0004] Lack of comprehensive perception of multi-dimensional system states. Traditional scheduling methods usually only consider a single dimension (such as CPU usage rate), ignoring other key indicators such as memory occupancy, network latency, and disk I / O, and cannot achieve intelligent scheduling in the true sense;

[0005] Difficult to balance multiple business goals. In actual scenarios, multiple goals such as processing efficiency, resource utilization rate, and service quality need to be considered simultaneously, and it is difficult for traditional rule-based scheduling strategies to find the optimal balance among these goals;

[0006] Lack of predictability for burst traffic. Traditional scheduling methods are often passive and reactive, unable to predict business peaks in advance and make resource preparations, which easily causes system response delays and service quality degradation. Summary of the Invention

[0007] The purpose of the present invention is to provide a business server and a dedicated server scheduling method to solve the above problems.

[0008] The present invention provides a business server and a dedicated server scheduling method, including:

[0009] Initializing a deep reinforcement learning agent, including constructing a state space, defining an action space, constructing a deep neural network, and initializing an exploration strategy;

[0010] Interacting the agent with the environment to obtain environmental feedback rewards and new states, and the environmental feedback reward formula is:

[0011] ;

[0012] Among them: is the environmental feedback reward value, , , , are weight coefficients, and each coefficient satisfies , , are negative numbers, represents the processing efficiency, and its calculation method is:

[0013] ;

[0014] Among them, is the number of successfully deployed services, is the total number of service requests;

[0015] represents the resource utilization rate, and its calculation method is:

[0016] ;

[0017] Among them, is the number of server cluster nodes, represents the th server node's used resource amount, represents the th server node's total resource amount;

[0018] represents the latency penalty, and its calculation method is:

[0019] ;

[0020] Among them, is the number of services, represents the th service's processing time, represents the th service's service level agreement specified time;

[0021] represents the overload penalty, and its calculation method is:

[0022] ;

[0023] is the number of server cluster nodes, is the set resource usage threshold, represents the th server node's used resource amount, represents the th server node's total resource amount;

[0024] Store the transfer data of the interaction between the agent and the environment in the experience pool and sample it;

[0025] Calculate the target Q-value and loss function based on the sampled data, and update the network parameters;

[0026] Perform policy optimization and model deployment according to the preset conditions.

[0027] Furthermore, the action space includes:

[0028] Deployment action, used to determine the allocation relationship of services to the server;

[0029] Scheduling action, used to perform resource reservation and priority adjustment.

[0030] Furthermore, the state space includes:

[0031] Service feature vector, which is collected in real time by the service monitoring system and normalized;

[0032] Server state matrix, which contains the key index data of the server cluster nodes.

[0033] Furthermore, the agent observes the current service feature vector and server state matrix, combines them into the current state; selects and executes an action according to the current policy, and obtains the environmental feedback reward and the new state.

[0034] Furthermore, the calculation method of the target Q-value is:

[0035] ;

[0036] Among them, represents the calculated th target Q-value, represents the environmental feedback reward value in the th transfer data sampled from the experience replay pool, represents the discount factor, represents the operation of taking the maximum value for all actions in the action space , represents the Q-value estimation of the target network for the action under the state , is the target network parameter;

[0037] The calculation method of the loss function is:

[0038] ;

[0039] Among them, represents the loss function, used to measure the current network parameters Under this condition, the difference between the estimated Q value and the target Q value represents the number of transfer data samples obtained from the experience replay pool, represents the main Q network in the state under which, for the action Q value estimation.

[0040] Furthermore, it also includes:

[0041] Construct a traffic prediction model to predict the business traffic peak and resource requirements;

[0042] Expand the state space and action space, and add information related to resource scaling;

[0043] Execute dynamic resource allocation based on the prediction results.

[0044] Furthermore, it also includes:

[0045] Introduce the geographical information, network latency information, and energy cost information of multiple data centers;

[0046] Reconstruct the state space and action space to achieve collaborative scheduling across data centers;

[0047] Optimize the deployment strategy across data centers to reduce the overall operating cost.

[0048] Furthermore, it also includes:

[0049] Introduce business security level information and server security threat information;

[0050] Expand the state space and action space, and add security-related scheduling strategies;

[0051] Conduct comprehensive optimization based on security level matching degree, processing efficiency, and threat level.

[0052] The present invention provides a computer storage medium for storing computer-readable instructions, which can execute the foregoing scheduling method for a business server and a dedicated server when the computer-readable instructions are read.

[0053] The beneficial effects of the present invention are as follows: The present invention adopts a deep reinforcement learning framework. Through continuous interaction and learning between the agent and the environment, it can adaptively understand the dynamic changes of business characteristics and server states, realize intelligent scheduling decisions, and at the same time set up a multi-dimensional state space and a reward function, comprehensively considering multiple performance indicators such as processing efficiency, resource utilization rate, service latency, and load balancing;

[0054] The present invention introduces a traffic prediction model, which can predict in advance the duration of business peaks and resource requirements, and achieve advance planning of resources and dynamic scaling. This forward-looking scheduling strategy effectively avoids the impact of sudden traffic on system performance;

[0055] The present invention supports collaborative scheduling of multiple data centers. By considering factors such as geographical distribution, network latency, and energy costs, it optimizes the resource allocation strategy across data centers, significantly reducing the overall operating cost;

[0056] The present invention incorporates security considerations into the scheduling decision. According to the security level of the business and the security threats faced by the server, it dynamically adjusts the deployment strategy, effectively preventing potential security risks and ensuring the secure and stable operation of the business. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a flowchart of a scheduling method for a business server and a dedicated server according to the present invention;

[0058] Figure 2 is an example diagram of business feature data according to the present invention;

[0059] Figure 3 is an example diagram of server status data according to the present invention;

[0060] Figure 4 is an example comparison diagram of scheduling optimization effects according to the present invention;

[0061] Figure 5 is an example diagram of scheduling actions executed by an agent according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] Now, the subject matter described herein will be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.

[0063] In at least one embodiment of the present invention, a scheduling method for a business server and a dedicated server is disclosed, as Figure 1 shown, including the following steps:

[0064] Step 100, initialize the deep reinforcement learning agent. The specific implementation is as follows:

[0065] (1) State space construction:

[0066] Construct a business feature vector , real-time collection is performed through the business monitoring system dimensional business feature data, for each dimension perform normalization processing, that is , where is the historical mean of this feature, is the historical standard deviation of this feature.

[0067] Construct the server status matrix , from the nodes of the server cluster collect key metrics (CPU usage rate, memory occupancy rate, disk IO, etc.) to form a dimensional status matrix, and at the same time standardize each element of the status matrix to the

[0068] Joint state space, form a state vector through feature concatenation , , represents the feature concatenation operation, and the dimension of the state vector is , ensuring dimension alignment between features.

[0069] (2) Definition of action space:

[0070] Deployment action : Adopt discrete action coding, each action corresponds to a specific deployment relationship of the business to the target server, such as represents the business flow is assigned to the server group .

[0071] Scheduling action : Includes operations such as resource reservation and priority adjustment, such as represents reserving 20% of the CPU resources for high-priority services.

[0072] Action space Adopt one-hot coding, and the dimension is determined according to the actual scale of the server cluster (50 - 200 discrete actions), and an action validity verification mechanism is established.

[0073] (3) Construction of the deep neural network:

[0074] The network structure adopts a fully connected DNN, including 3 hidden layers (the number of neurons is 256, 128, 64 respectively), and the number of input layer nodes matches the dimension of the state space.

[0075] The number of output layer nodes is , the activation function adopts ReLU (linear activation of the output layer), and the parameter initialization adopts the Xavier method.

[0076] Synchronously construct the target network and adopt a soft update strategy , where , is the current network parameter, is the target network parameter. The soft update strategy means that the target network parameter will be slowly updated to the current network parameter at a ratio of . .

[0077] (4)Initialization of exploration strategy:

[0078] Set -greedy strategy parameters: initial exploration rate , exponential decay factor , minimum exploration rate .

[0079] The experience replay pool adopts a prioritized experience replay mechanism, with the capacity adjusted to 100,000 experiences, and the sampling strategy is priority sampling based on TD error.

[0080] Step 200, the agent interacts with the environment. The agent observes the current business characteristics and the server status , combines them into the current state . Selects an action according to the current policy and executes it, obtaining the environmental feedback reward and the new state . The environmental feedback reward formula is:

[0081]

[0082] where: is the environmental feedback reward value, , , , are weight coefficients, and each coefficient satisfies , , are negative numbers, and the optimal weight combination is determined through Bayesian optimization; represents the processing efficiency, and its calculation method is:

[0083] ;

[0084] where, is the number of successfully deployed services, is the total number of service requests;

[0085] represents the resource utilization rate, and its calculation method is:

[0086]

[0087] Among them, is the number of server cluster nodes, represents the amount of resources used by the th server node, represents the total amount of resources of the th server node;

[0088] represents the delay penalty, and its calculation method is:

[0089]

[0090] Among them, is the number of services, represents the processing time of the th service, represents the time specified by the service level agreement of the th service;

[0091] represents the overload penalty, and its calculation method is:

[0092]

[0093] is the number of server cluster nodes, is the set resource usage threshold, represents the amount of resources used by the th server node, represents the total amount of resources of the th server node.

[0094] Step 300, experience storage and sampling. Store the transition in the experience pool (add the termination state flag ), and adopt a hierarchical sampling strategy: 70% of the samples come from the most recent experience, and 30% come from historical experience.

[0095] Step 400, network training and update:

[0096] (1) Calculate the target Q value, and its calculation method is:

[0097]

[0098] Among them, represents the th calculated target Q value, represents the from sampling in the experience replay poolThe environmental feedback reward value in the transferred data, represents the discount factor, represents all actions in the action space for the maximum operation, represents the target network in the state under the action Q-value estimation, is the target network parameter;

[0099] (2) The mean squared error loss function is adopted, and its calculation method is:

[0100]

[0101] where, represents the loss function, which is used to measure the difference between the estimated Q-value and the target Q-value under the current network parameter , represents the number of transferred data sampled from the experience replay pool, represents the Q-value estimation of the main Q network in the state under the action ;

[0102] (3) Use the Adam optimizer to update the parameters, and the learning rate ;

[0103] (4) Synchronously update the target network parameters every 10,000 steps .

[0104] Step 500, Policy Optimization and Model Deployment:

[0105] (1) Adopt periodic policy evaluation: Test the performance of the current policy every 5000 steps;

[0106] (2) The value linearly decays to ;

[0107] (3) The convergence condition is set as the average environmental feedback reward change rate is less than 2% and the Q-value fluctuation is less than 10%;

[0108] (4) Save the model parameters with the best performance on the validation set;

[0109] The risk control evaluation method based on multi-dimensional perception and enterprise data quantification realizes the intelligent scheduling of server resources through a deep reinforcement learning framework, effectively controls the server load while ensuring the business processing efficiency, and constructs a scalable intelligent operation and maintenance infrastructure.

[0110] In one embodiment of the present invention, to address the technical problem that when facing a sudden peak in business traffic, it is necessary not only to achieve reasonable deployment and scheduling of servers, but also to be able to quickly predict the duration of the peak and the specific resource requirements, so as to dynamically allocate resources and scale up or down in advance to avoid a decline in business processing performance, the following solutions are provided:

[0111] Step 601, construct a traffic prediction model. Utilize historical business traffic data , where represents the business traffic values in different past time periods, represents the total number of historical data, and use a time series analysis algorithm (such as the ARIMA model) or a recurrent neural network in deep learning (such as LSTM) to construct a traffic prediction model . The input of the model is historical business traffic data , and the output is the predicted future business traffic value as well as the peak duration and resource requirements .

[0112] Step 602, expand the state space and action space. On the basis of step 100, expand the state space to :

[0113]

[0114] where represents the Cartesian product operation, and the action space adds actions related to resource scaling up or down to form a new action space , such as operations like applying for new servers and releasing servers, where represents the original number of actions, represents the number of newly added actions. Initialize a deep neural network (such as the Q-network in the improved DQN), whose input is the state space , and the output is the Q value of each action , and initialize the network parameters .

[0115] Step 603, after the agent observes the current state , select an action according to the current policy and execute it. Among them, represents the current business characteristics, represents the current server state, represents the currently predicted traffic value, represents the current peak duration, represents the current resource requirements. After executing the action, the traffic reward and the new business characteristics , server state , predicted traffic value , peak duration and resource requirements , thus obtaining a new state . Traffic reward is set considering the rationality of resource scaling and business processing performance, and its calculation method is as follows:

[0116]

[0117] Among them, , , are weight coefficients and satisfy , represents the business processing performance reward, and its calculation method is as follows:

[0118]

[0119] Among them is the processing time of the th business request, is the processing time threshold; represents the resource scaling efficiency reward, and its calculation method is as follows:

[0120]

[0121] Among them, is the predicted resource demand, is the actually used resource amount; represents the resource cost reward, and its calculation method is as follows:

[0122]

[0123] Among them, is the used resource cost, is the budget cost ceiling.

[0124] Step 604, store the transfer of each step into the experience replay pool , randomly sample a batch of transfer data from the experience replay pool , where represents the sampled state, represents the sampled action, represents the sampled traffic reward, represents the sampled next state, from 1 to , represents the number of sampled data.

[0125] Step 605: Update the Q-network parameters according to the objective function of the deep reinforcement learning algorithm (such as the improved DQN). The calculation method is as follows:

[0126]

[0127] where represents the sampled traffic reward, represents the next state obtained by sampling, is the discount factor, which is used to measure the importance of future traffic rewards, are the target network parameters, represents selecting the action with the maximum Q-value among all possible actions . By minimizing the loss function, the calculation method is as follows:

[0128]

[0129] where represents the state obtained by sampling, represents the action obtained by sampling, is the number of sampled data, is the Q-value estimation of the main Q-network for the action in the state , represents the main Q-network parameters, and update the Q-network parameters using the gradient descent algorithm .

[0130] Step 606: According to the policy learned by the agent, when predicting a business traffic peak, based on the resource requirements and the peak duration , combined with the resource scaling-related information provided by the cloud resource provider , where represents the th resource configuration information, represents the total number of available resource configurations, such as scalable server types, quantities, costs, and scaling times, etc., perform resource scaling actions, such as applying for new server resources or releasing resources that are no longer needed.

[0131] By constructing a traffic prediction model, expanding the state space and action space, the agent can learn resource scaling-related policies. After learning and training, the agent can reasonably perform dynamic resource allocation and scaling in advance when facing sudden business traffic peaks, effectively avoiding the decline of business processing performance, improving the stability and reliability of business processing, and at the same time optimizing the resource usage cost.

[0132] In one embodiment of the present invention, for the technical problem of optimizing server deployment and scheduling strategies according to factors such as the geographical distribution, network latency, and energy cost of different data centers in a multi-data center environment, achieving collaborative operation and maintenance management across data centers, and reducing the overall operation cost, the following solutions are provided:

[0133] Step 701, reconstruct the state space and action space. Based on step 100, introduce the geographical information of each data center , where represents the geographical location coordinates of the data center; the network latency information between data centers , where 、 from 1 to , represents the data center to the data center network latency; the energy cost information of each data center , where represents the data center energy cost per unit time. Reconstruct the state space into :

[0134]

[0135] The action space is redefined as a set of server deployment and scheduling actions across data centers , for example, migrating a certain service from data center to data center and other operations. Initialize a deep neural network (such as the Actor-Critic network in the A2C algorithm based on policy gradients). The input of the Actor network is the state space , and the output is the probability of each action . The input of the Critic network is the state space , and the output is the state value . Initialize the Actor network parameters and the Critic network parameters .

[0136] Step 702, the agent observes the current business characteristics , server status , geographical information , network latency information and energy cost information , and combines them into the current state:

[0137]

[0138] The action probability output by the Actor network Sample and select an action And execute it. After executing the action, the cost reward And the new business characteristics 、Server status 、Geographical information 、Network latency information And energy cost information To obtain a new state :

[0139]

[0140] Cost reward The setting of the cost reward is based on the business objective and the operating cost, and the calculation method is as follows:

[0141]

[0142] Among them, 、 、 Are weight coefficients and satisfy , Represents the network latency reward, and the calculation method is as follows:

[0143]

[0144] Among them Is the network latency from the data center To , Is the data center To The business traffic of; Represents the energy cost reward, and the calculation method is as follows:

[0145]

[0146] Is the unit energy cost of the data center , Is the data center The energy consumption of; Represents the load balancing reward, and the calculation method is as follows:

[0147]

[0148] Is the data center The load rate of, Is the average load rate of all data centers.

[0149] Step 703, the transfer of each step Stored in the experience pool inside

[0150] Step 704, update the Actor-Critic network parameters according to the algorithm based on policy gradient (such as A2C). Calculate the advantage function , The calculation method is as follows

[0151]

[0152] where is the discount factor, used to measure the importance of future cost rewards represents the value evaluation of the next state by the Critic network of represents the value evaluation of the current state by the Critic network ; The update formula for the Actor network parameters is

[0153]

[0154] where is the learning rate of the Actor network represents taking the gradient of for updating the Actor network parameters represents the probability that the Actor network takes action in state ; The update formula for the Critic network parameters is

[0155]

[0156] where is the learning rate of the Critic network represents taking the gradient of for updating the Critic network parameters

[0157] Step 705, continuously repeat Step 702 - Step 704, so that the agent can continuously optimize the policy through learning to adapt to different business characteristics, server states and various data center-related factors in the multi-data center environment, and find the optimal cross-data center server deployment and scheduling strategy

[0158] By reconstructing the state space and action space, introducing multi-data center related information, and using a policy gradient-based algorithm for learning and optimization, the intelligent agent can learn cross-data center server deployment and scheduling strategies that consider factors such as geographical distribution, network latency, and energy costs. After implementing this solution, collaborative operation and maintenance management of multiple data centers can be achieved, effectively reducing the overall operating cost, improving the processing efficiency and resource utilization rate of the business in a multi-data center environment, and ensuring the service quality of the business at the same time.

[0159] In one embodiment of the present invention, to address the technical problem of dynamically adjusting server deployment and scheduling strategies according to the security level of the service and the security threats faced by the server while taking into account service security, the following solution is provided:

[0160] Step 801, expand the state space and action space. On the basis of step 100, introduce the security level information of the service , where represents the security level of the service (such as high, medium, low); the security threat information faced by the server , where represents the security threat indicators faced by the server node (such as DDoS attack risk score, number of vulnerabilities, etc.). Expand the state space to , and adjust the action space to a set of server deployment and scheduling actions considering security factors , such as deploying high-security level services to servers with low security threats and other operations. Initialize a deep neural network (such as a Dueling DQN network based on the double Q network), whose input is the state space , and the output is the Q value of each action , and initialize the network parameters .

[0161] Step 802, the intelligent agent observes the current service characteristics , server status , security level information and security threat information , and combines them into the current state :

[0162]

[0163] Select an action (such as the greedy policy) according to the current policy and execute it. After executing the action, the security reward and the new service characteristics , server status , security level information and security threat information to obtain a new state :

[0164]

[0165] Security rewards should be set considering both business security and processing efficiency, and the calculation method is as follows:

[0166]

[0167] where , , are weight coefficients and satisfy , represents the security level matching reward, and the calculation method is:

[0168]

[0169] where is the security level of the business , is the security protection level of the server , represents whether the business is deployed on the server as a 0-1 variable, yes is 1, no is 0;

[0170] represents the business processing efficiency reward, and the calculation method is:

[0171]

[0172] where is the processing delay of the business ;

[0173] represents the security threat reward, and the calculation method is:

[0174]

[0175] where is the security threat index of the server ;

[0176] Step 803, store the transfer of each step into the experience replay pool , and randomly sample a batch of transfer data from the experience replay pool , .

[0177] Step 804: Update the network parameters according to the Dueling DQN algorithm based on the double Q network. Two Q networks (the main Q network and the target Q network) are used respectively, where represents the Q-value estimation of the main Q network for the state and the action . represents the Q-value estimation of the target Q network for the state and the action . The calculation method of the target Q value is:

[0178]

[0179] where is the discount factor, are the target network parameters (synchronized with the main Q network parameters regularly). By minimizing the loss function:

[0180]

[0181] the main Q network parameters are updated using the gradient descent algorithm .

[0182] Step 805: Continuously repeat Step 802 - Step 804, enabling the agent to continuously optimize the strategy through learning to adapt to different business security levels and server security threat situations, and find the optimal server deployment and scheduling strategy that balances security and efficiency.

[0183] By expanding the state space and adjusting the action space, integrating business security level and server security threat information, and using the Dueling DQN algorithm based on the double Q network for learning, the agent can learn a server deployment and scheduling strategy that takes security into account. After the implementation of this solution, it is possible to optimize server deployment and scheduling on the premise of ensuring the secure operation of the business, avoid problems such as business interruption or data leakage caused by security threats, and at the same time maintain a high business processing efficiency and resource utilization rate.

[0184] Taking the order processing system of an e-commerce platform as an example, the platform has 10 business servers, and the scheduling optimization process during the "Double 11" event is as Figures 2 - 5 shown;

[0185] Through the scheduling optimization of the deep reinforcement learning model, the system shows good adaptability during the peak period:

[0186] The average server load is reduced from 85% to 65%, while maintaining a high resource utilization rate;

[0187] The order processing success rate is increased to 99.2%, and the user experience is significantly improved;

[0188] The system response time is reduced by 67.9%, from 2.8 seconds to 0.9 seconds;

[0189] The resource utilization balance is improved by 26.4%, effectively avoiding the overload of a single server;

[0190] In at least one embodiment of the present invention, a computer storage medium is provided, which is used to store computer-readable instructions that can execute the foregoing scheduling method for a service server and a dedicated server when the computer-readable instructions are read.

[0191] The embodiments of the present invention have been described above, but the embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of the present embodiment.

Claims

1. A method for scheduling a business server and a dedicated server, characterized in that: include: Initialize the deep reinforcement learning agent, including constructing the state space, defining the action space, building a deep neural network, and initializing the exploration strategy; The agent interacts with the environment to obtain environmental feedback rewards and new states. The environmental feedback reward formula is: ; in: is the reward value for the environment feedback, , , , is the weight coefficient, and each coefficient satisfies , , is a negative number, It represents the processing efficiency, which is calculated as follows: ; in, is the number of successfully deployed services, is the total number of business requests; Indicates resource utilization, which is calculated as follows: ; in, is the number of server cluster nodes, Indicates The amount of resources used by server nodes, Indicates Total resources of server nodes; represents the delay penalty, which is calculated as: ; in, is the number of businesses, Indicates The processing time of each business, Indicates The service level agreement of each business stipulates the time; Represents the overload penalty, which is calculated as: ; is the number of server cluster nodes, is the resource usage threshold that is set. Indicates The amount of resources used by server nodes, Indicates Total resources of server nodes; The transfer data of the interaction between the agent and the environment is stored in the experience pool and sampled; Calculate the target Q value and loss function based on the sampled data and update the network parameters; Optimize strategies and deploy models based on preset conditions.

2. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: The action space includes: Deployment action, used to determine the distribution relationship between services and servers; Scheduling actions, used to make resource reservations and priority adjustments; Deployment Actions : Using discrete action coding, each action corresponds to the deployment relationship from a specific business to the target server; Scheduling Actions : Includes resource reservation and priority adjustment operations; Action Space One-hot encoding is used, dimension The value is determined based on the actual server cluster size.

3. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: The state space consists of: Business feature vectors, which are collected and normalized in real time by a business monitoring system; A server status matrix, wherein the server status matrix includes key indicator data of server cluster nodes; The steps to generate the state space include: Constructing business feature vector , collected in real time through the business monitoring system Dimension business characteristic data, for each dimension Normalization is performed, that is, ,in is the historical mean of the feature, is the historical standard deviation of the feature; Building a server status matrix , from the server cluster Node collection key indicators, forming dimensional state matrix, and each element of the state matrix is ​​standardized to interval; Joint state space, forming a state vector by feature concatenation , , Represents the feature concatenation operation, and the state vector dimension is , ensuring that the dimensions between features are aligned.

4. A method for scheduling business servers and dedicated servers according to claim 3, characterized in that: The agent observes the current business feature vector and server state matrix, combines them into the current state; selects and executes actions according to the current strategy, and obtains environmental feedback rewards and new states.

5. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: The target Q value is calculated as follows: ; in, Indicates the calculated The target Q value, represents the first sampled from the experience replay pool The environmental feedback reward value in the transferred data, represents the discount factor, Represents all actions in the action space Take the maximum value operation, Indicates that the target network is in state Next, for action The Q value of is the target network parameter; The loss function is calculated as: ; in, Represents the loss function, which is used to measure the current network parameters The difference between the estimated Q value and the target Q value is Represents the number of transfer data sampled from the experience replay pool, Indicates that the main Q network is in state Next, for action The Q value of .

6. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: Also includes: Build a traffic prediction model to predict business traffic peaks and resource requirements; Expand the state space and action space, and add information related to resource expansion and contraction; Dynamic resource allocation is performed based on the prediction results.

7. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: Also includes: Introducing geographical information, network latency information, and energy cost information of multiple data centers; Reconstruct state space and action space to achieve cross-data center collaborative scheduling; Optimize deployment strategies across data centers and reduce overall operating costs.

8. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: Also includes: Introducing business security level information and server security threat information; Expand the state space and action space, and add safety-related scheduling strategies; Comprehensive optimization is performed based on security level matching, processing efficiency and threat level.

9. A computer storage medium, characterized in that It is used to store computer-readable instructions, and when the computer-readable instructions are read, a business server and dedicated server scheduling method as described in any one of claims 1-8 can be executed.

Citation Information

Patent Citations

  • Edge computing unloading and resource allocation method based on multi-agent reinforcement learning

    CN116321293A

  • Time-varying task scheduling method and system based on constraint near-end strategy optimization

    CN117851056A