Service server and dedicated server scheduling method
Through the deep reinforcement learning framework and traffic prediction model, combined with multi-data centers and security factors, the scheduling strategies of business servers and dedicated servers are optimized, and the problems of insufficient scheduling, lack of foresight and difficulty in balancing multiple goals in the existing technology are solved, and efficient and intelligent resource scheduling and business security guarantees are achieved.
Patent Information
- Application Number
- CN202510436266.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-09
AI Technical Summary
When handling resource scheduling of business servers and dedicated servers, the prior art cannot effectively deal with dynamically changing business loads, lacks comprehensive perception of multi-dimensional system status, is difficult to balance multiple business goals, and lacks predictiveness of burst traffic.
Adopt the deep reinforcement learning framework to initialize the deep reinforcement learning agent. Through the continuous interaction and learning between the agent and the environment, we adaptively understand the dynamic changes in business characteristics and server status, and realize intelligent scheduling decisions. At the same time, a traffic prediction model is introduced to construct geographic information, network latency information and energy cost information of multi-data centers, and integrate business security level information and server security threat information to optimize scheduling strategies.
It realizes intelligent scheduling decisions, can adaptively respond to dynamically changing business loads, comprehensively consider multiple performance indicators, predict business peaks in advance, optimize resource allocation, reduce operational costs, and improve system stability and reliability while ensuring business security.
Smart Images

Figure CN119961002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer data processing, and more specifically, to a method for scheduling a business server and a dedicated server. Background Art
[0002] As the degree of enterprise informatization continues to improve, the complexity of business systems and the scale of servers are growing rapidly. In the actual operation and maintenance process, administrators need to deal with the resource scheduling issues of business servers and dedicated servers (such as database servers, cache servers, etc.) at the same time. The current mainstream server scheduling solutions mainly adopt rule-based static scheduling strategies, such as scheduling based on server load thresholds, scheduling based on business priorities, etc. However, these traditional methods have the following technical problems: Unable to effectively cope with dynamically changing business loads. During business peaks, static scheduling rules often cannot adjust resource allocation in a timely manner, causing some servers to be overloaded while other server resources are idle, affecting overall system performance; Lack of comprehensive perception of multi-dimensional system status. Traditional scheduling methods usually only consider a single dimension (such as CPU usage), ignoring other key indicators such as memory usage, network latency, disk IO, etc., and cannot achieve truly intelligent scheduling; It is difficult to balance multiple business goals. In actual scenarios, multiple goals such as processing efficiency, resource utilization, and service quality need to be considered at the same time. Traditional rule-based scheduling strategies are difficult to find the optimal balance between these goals. Lack of predictability for sudden traffic. Traditional scheduling methods are often reactive and cannot predict business peaks in advance and prepare resources, which can easily cause system response delays and reduced service quality. Summary of the invention
[0003] The purpose of the present invention is to provide a business server and a dedicated server scheduling method in order to solve the above problems.
[0004] The present invention provides a method for scheduling a business server and a dedicated server, comprising: Initialize the deep reinforcement learning agent, including constructing the state space, defining the action space, building a deep neural network, and initializing the exploration strategy; Interact the agent with the environment to obtain environmental feedback rewards and new states; The transfer data of the interaction between the agent and the environment is stored in the experience pool and sampled; Calculate the target Q value and loss function based on the sampled data and update the network parameters; Optimize strategies and deploy models based on preset conditions.
[0005] Furthermore, the action space includes: Deployment action, used to determine the distribution relationship between services and servers; Scheduling actions are used to reserve resources and adjust priorities.
[0006] Furthermore, the state space includes: Business feature vectors, which are collected and normalized in real time by a business monitoring system; A server status matrix, wherein the server status matrix includes key indicator data of server cluster nodes.
[0007] Furthermore, the agent observes the current business feature vector and server state matrix, combines them into the current state; selects and executes actions according to the current strategy, and obtains environmental feedback rewards and new states.
[0008] Furthermore, the environmental feedback reward formula is: ; in: is the reward value for the environment feedback, , , , is the weight coefficient, and each coefficient satisfies , , is a negative number, It represents the processing efficiency, which is calculated as follows: ; in, is the number of successfully deployed services, is the total number of business requests; Indicates resource utilization, which is calculated as follows: ; in, is the number of server cluster nodes, Indicates The amount of resources used by server nodes, Indicates Total resources of server nodes; represents the delay penalty, which is calculated as: ; in, is the number of businesses, Indicates The processing time of each business, Indicates The service level agreement of each business stipulates the time; Represents the overload penalty, which is calculated as: ; is the number of server cluster nodes, is the resource usage threshold that is set. Indicates The amount of resources used by server nodes, Indicates The total resources of server nodes.
[0009] Furthermore, the target Q value is calculated as follows: ; in, Indicates the calculated The target Q value, represents the first sampled from the experience replay pool The environmental feedback reward value in the transferred data, represents the discount factor, Represents all actions in the action space Take the maximum value operation, Indicates that the target network is in status Next, for action The Q value of is the target network parameter; The loss function is calculated as: ; in, Represents the loss function, which is used to measure the current network parameters The difference between the estimated Q value and the target Q value is Represents the number of transfer data sampled from the experience replay pool, Indicates that the main Q network is in state Next, for action The Q value of .
[0010] Furthermore, it also includes: Build a traffic prediction model to predict business traffic peaks and resource requirements; Expand the state space and action space, and add information related to resource expansion and contraction; Dynamic resource allocation is performed based on the prediction results.
[0011] Furthermore, it also includes: Introducing geographical information, network latency information, and energy cost information of multiple data centers; Reconstruct state space and action space to achieve cross-data center collaborative scheduling; Optimize deployment strategies across data centers and reduce overall operating costs.
[0012] Furthermore, it also includes: Introducing business security level information and server security threat information; Expand the state space and action space, and add safety-related scheduling strategies; Comprehensive optimization is performed based on security level matching, processing efficiency and threat level.
[0013] The present invention provides a computer storage medium for storing computer-readable instructions, which can execute the aforementioned business server and dedicated server scheduling method when the computer-readable instructions are read.
[0014] The beneficial effects of the present invention are as follows: the present invention adopts a deep reinforcement learning framework, and through continuous interaction and learning between the intelligent agent and the environment, it can adaptively understand the dynamic changes of business characteristics and server status, realize intelligent scheduling decisions, and set up a multi-dimensional state space and reward function, comprehensively considering multiple performance indicators such as processing efficiency, resource utilization, service delay and load balancing; The present invention introduces a traffic prediction model that can predict the duration of business peaks and resource requirements in advance, and realize advance resource planning and dynamic expansion and contraction. This forward-looking scheduling strategy effectively avoids the impact of burst traffic on system performance; The present invention supports the coordinated scheduling of multiple data centers, optimizes the resource allocation strategy across data centers by considering factors such as geographical distribution, network latency and energy cost, and significantly reduces the overall operating cost; The present invention incorporates security considerations into scheduling decisions, dynamically adjusts deployment strategies according to the security level of the business and the security threats faced by the server, effectively prevents potential security risks, and ensures the safe and stable operation of the business. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flow chart of a method for scheduling a business server and a dedicated server of the present invention; Figure 2 is an example diagram of the service characteristic data of the present invention; Figure 3 is an example diagram of server status data of the present invention; Figure 4 This is a comparative example diagram of the scheduling optimization effect of the present invention; Figure 5 It is an example diagram of the scheduling action performed by the intelligent agent of the present invention. DETAILED DESCRIPTION
[0016] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that the discussion of these embodiments is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the contents of this specification. Each example may omit, replace or add various processes or components as needed. In addition, the features described relative to some examples may also be combined in other examples.
[0017] At least one embodiment of the present invention discloses a method for scheduling a business server and a dedicated server, such as Figure 1 As shown, the following steps are included: Step 100: Initialize the deep reinforcement learning agent. The specific implementation is as follows: (1) State space construction: Constructing business feature vector , collected in real time through the business monitoring system Dimension business characteristic data, for each dimension Normalization is performed, that is, ,in is the historical mean of the feature, is the historical standard deviation of the feature.
[0018] Building a server status matrix , from the server cluster Node collection key indicators (CPU usage, memory usage, disk IO, etc.), forming dimensional state matrix, and each element of the state matrix is standardized to interval.
[0019] Joint state space, forming a state vector by feature concatenation , , Represents the feature concatenation operation, and the state vector dimension is , ensuring that the dimensions between features are aligned.
[0020] (2) Action space definition: Deployment Actions : Discrete action coding is used, and each action corresponds to the deployment relationship from a specific business to the target server, such as Indicates that the business flow Assign to server group .
[0021] Scheduling Actions : Contains operations such as resource reservation and priority adjustment, such as Indicates that 20% of CPU resources are reserved for high-priority services.
[0022] Action Space One-hot encoding is used, dimension Determine based on the actual server cluster size (50-200 discrete actions) and establish an action validity verification mechanism.
[0023] (3) Deep neural network construction: The network structure adopts a fully connected DNN, which contains 3 hidden layers (the number of neurons is 256, 128, and 64 respectively), and the number of nodes in the input layer matches the state space dimension.
[0024] The number of nodes in the output layer is , the activation function uses ReLU (output layer linear activation), and the parameter initialization uses the Xavier method.
[0025] Synchronously build the target network and adopt a soft update strategy ,in , is the current network parameter, is the target network parameter, and the soft update strategy refers to the target network parameter Will The ratio is slowly updated to the current network parameters .
[0026] (4) Exploration strategy initialization: set up -greedy strategy parameters: initial exploration rate , exponential decay factor , minimum exploration rate .
[0027] The experience replay pool adopts a priority experience replay mechanism, the capacity is adjusted to 100,000 experiences, and the sampling strategy is priority sampling based on TD error.
[0028] Step 200: The agent interacts with the environment. The agent observes the current business characteristics. and server status , combined into the current state . Choose an action based on the current strategy And execute, get environmental feedback rewards and the new state The environmental feedback reward formula is: ; in: is the reward value for the environment feedback, , , , is the weight coefficient, and each coefficient satisfies , , If it is a negative number, the optimal weight combination is determined by Bayesian optimization; It represents the processing efficiency, which is calculated as follows: ; in, is the number of successfully deployed services, is the total number of business requests; Indicates resource utilization, which is calculated as follows: ; in, is the number of server cluster nodes, Indicates The amount of resources used by server nodes, Indicates Total resources of server nodes; represents the delay penalty, which is calculated as: ; in, is the number of businesses, Indicates The processing time of each business, Indicates The service level agreement of each business stipulates the time; Represents the overload penalty, which is calculated as: ; is the number of server cluster nodes, is the resource usage threshold that is set. Indicates The amount of resources used by server nodes, Indicates The total resources of server nodes.
[0029] Step 300, experience storage and sampling. Deposit into experience pool (add termination status flag ), using a stratified sampling strategy: 70% of the samples come from recent experience and 30% from historical experience.
[0030] Step 400, network training and updating: (1) Calculate the target Q value. The calculation method is: ; in, Indicates the calculated The target Q value, represents the first sampled from the experience replay pool The environmental feedback reward value in the transferred data, represents the discount factor, Represents all actions in the action space Take the maximum value operation, Indicates that the target network is in status Next, for action The Q value of is the target network parameter; (2) The mean square error loss function is used, and its calculation method is: ; in, Represents the loss function, which is used to measure the current network parameters The difference between the estimated Q value and the target Q value is Represents the number of transfer data sampled from the experience replay pool, Indicates that the main Q network is in state Next, for action Q value estimation of ; (3) Use Adam optimizer to update parameters and learning rate ; (4) Synchronously update the target network parameters every 10,000 steps .
[0031] Step 500, strategy optimization and model deployment: (1) Use periodic strategy evaluation: test the performance of the current strategy every 5000 steps; (2) The value decays linearly to ; (3) The convergence condition is set as the average environmental feedback reward change rate is less than 2% and the Q value fluctuation is less than 10%; (4) Save the model parameters that perform best on the validation set; The risk control assessment method based on multi-dimensional perception and enterprise data quantification realizes intelligent scheduling of server resources through a deep reinforcement learning framework, effectively controls server load while ensuring business processing efficiency, and builds a scalable intelligent operation and maintenance infrastructure.
[0032] In one embodiment of the present invention, in order to solve the technical problem that when facing a sudden peak of business traffic, it is necessary not only to realize reasonable deployment and scheduling of servers, but also to be able to quickly predict the duration of the peak and the specific demand for resources, so as to dynamically allocate resources and expand and shrink capacity in advance to avoid the degradation of business processing performance, the following solution is provided: Step 601: Build a traffic prediction model. ,in Indicates the business flow value of different time periods in the past. Indicates the total amount of historical data, and uses time series analysis algorithms (such as ARIMA model) or recurrent neural networks in deep learning (such as LSTM) to build traffic prediction models The model input is historical business flow data , the output is the predicted future business flow value and peak duration and resource requirements .
[0033] Step 602: Expand the state space and action space. Based on step 100, expand the state space to : ; in Represents Cartesian product operation, and the action space adds resource expansion and contraction related actions to form a new action space , such as applying for a new server, releasing a server, etc. Indicates the number of original actions, Indicates the number of new actions. Initialize the deep neural network (such as the Q network in the improved DQN), whose input is the state space , the output is the Q value of each action , initialize network parameters .
[0034] Step 603, the agent observes the current state Then, select an action based on the current strategy And execute. Among them, Indicates the current business characteristics. Indicates the current server status. Indicates the current predicted flow value, Indicates the duration of the current peak. Indicates the current resource demand. Traffic reward after executing the action And new business features , Server Status , predict traffic value , Peak duration and resource requirements , thus obtaining a new state Traffic Rewards The setting of needs to consider the rationality of resource expansion and contraction and business processing performance. The calculation method is: ; in, , , is the weight coefficient and satisfies , Indicates the business processing performance reward, which is calculated as follows: ; in For the The processing time of a business request, is the processing time threshold; Indicates the resource expansion and contraction efficiency reward, and the calculation method is: ; in, is the predicted resource demand, is the amount of resources actually used; Represents the resource cost reward, which is calculated as: ; in, is the cost of resources used, The upper limit of the budget cost.
[0035] Step 604, transfer each step Store to experience replay pool In the process, a batch of transfer data is randomly sampled from the experience replay pool. ,in Indicates the state obtained by sampling, represents the sampled action, Indicates the traffic reward obtained by sampling, represents the next state obtained by sampling, From 1 to , Indicates the number of sampled data.
[0036] Step 605: Update the Q network parameters according to the objective function of the deep reinforcement learning algorithm (such as the improved DQN). The calculation method is: ; in Indicates the traffic reward obtained by sampling, represents the next state obtained by sampling, is a discount factor used to measure the importance of future traffic rewards. are the target network parameters, Indicates that among all possible actions Select the action with the largest Q value. By minimizing the loss function, the calculation method is: ; in, Indicates the state obtained by sampling, represents the sampled action, is the number of sampled data, Is the main Q network in state Next, for action The Q value of Represents the main Q network parameters, and uses the gradient descent algorithm to update the Q network parameters .
[0037] Step 606: Based on the strategy learned by the agent, when a peak in business traffic is predicted, the agent will schedule resource requests in advance. and peak duration , combined with the information about resource expansion and contraction provided by the cloud resource provider ,in Indicates resource configuration information, Indicates the total number of available resource configurations, such as the scalable server types, quantity, cost, and scaling time, and performs resource scaling actions, such as applying for new server resources or releasing resources that are no longer needed.
[0038] By building a traffic prediction model and expanding the state space and action space, the intelligent agent can learn resource expansion and contraction strategies. After learning and training, the intelligent agent can reasonably allocate resources dynamically and scale up and down in advance when facing sudden business traffic peaks, effectively avoiding the decline of business processing performance, improving the stability and reliability of business processing, and optimizing resource usage costs.
[0039] In one embodiment of the present invention, the following solution is provided for the technical problem of optimizing server deployment and scheduling strategies, realizing cross-data center collaborative operation and maintenance management, and reducing overall operating costs in a multi-data center environment according to factors such as the geographical distribution of different data centers, network latency, and energy costs: Step 701, reconstruct the state space and action space. Based on step 100, introduce the geographic information of each data center ,in Indicates the geographic location coordinates of the data center; network latency information between data centers ,in , From 1 to , Indicates data center To the data center network latency; energy cost information for each data center ,in Indicates data center Energy cost per unit time. Reconstructing the state space into : ; The action space is redefined as the set of server deployment and scheduling actions across data centers. For example, moving a business from a data center Migrate to the data center Initialize a deep neural network (such as the Actor-Critic network in the policy gradient-based A2C algorithm), and the Actor network input is the state space , the output is the probability of each action , the Critic network input is the state space , the output is the state value , initialize the Actor network parameters and Critic network parameters .
[0040] Step 702: The agent observes the current business characteristics , Server Status , Geographic Information , network delay information and energy cost information , combined into the current state: ; According to the action probability output by the Actor network Sample Select an Action And execute, cost reward after executing the action And new business features , Server Status , Geographic Information , network delay information and energy cost information , thus obtaining a new state : ; Cost Incentive The setting is based on business objectives and operating costs, and the calculation method is: ; in, , , is the weight coefficient and satisfies , Represents the network delay reward, which is calculated as: ; in For Data Center arrive network latency, For Data Center arrive business traffic; represents the energy cost bonus, which is calculated as: ; For Data Center The unit energy cost, For Data Center energy consumption; Represents the load balancing reward, which is calculated as follows: ; For Data Center The load factor, is the average load factor of all data centers.
[0041] Step 703, transfer each step Store in experience pool middle.
[0042] Step 704: Update the Actor-Critic network parameters based on a policy gradient algorithm (such as A2C). Calculate the advantage function , The calculation method is: ; in is a discount factor that measures the importance of future cost rewards, Indicates the Critic network's response to the next state The value assessment of Indicates the Critic network's response to the current state The value assessment of the Actor network parameters is as follows: ; in is the learning rate of the Actor network, Express Find the gradient to update the Actor network parameters. Indicates that the Actor network is in state Next, take action Probability The natural logarithm of ; the parameter update formula of the Critic network is: ; in is the learning rate of the Critic network, Express Find the gradient to update the Critic network parameters.
[0043] Step 705, continuously repeat steps 702-704, so that the intelligent agent continuously optimizes the strategy through learning to adapt to different business characteristics, server status and various data center related factors in the multi-data center environment, and find the optimal cross-data center server deployment and scheduling strategy.
[0044] By reconstructing the state space and action space, introducing multi-data center related information, and using a policy gradient-based algorithm for learning and optimization, the agent can learn cross-data center server deployment and scheduling strategies that take into account factors such as geographic distribution, network latency, and energy costs. After implementing this solution, collaborative operation and maintenance management of multiple data centers can be achieved, effectively reducing overall operating costs, improving business processing efficiency and resource utilization in a multi-data center environment, and ensuring the service quality of the business.
[0045] In one embodiment of the present invention, the following solution is provided for the technical problem of dynamically adjusting server deployment and scheduling strategies to ensure safe operation of services based on the security level of the services and the security threats faced by the servers while taking into account the security of the services: Step 801: Expand the state space and action space. Based on step 100, introduce the security level information of the service. ,in Indicates business Security level (such as high, medium, low); security threat information faced by the server ,in Represents a server node Security threat indicators faced (such as DDoS attack risk score, number of vulnerabilities, etc.). Expand the state space to , the action space is adjusted to a set of server deployment and scheduling actions that take security factors into consideration , such as deploying high-security services to servers with low security threats. Initialize a deep neural network (such as the Dueling DQN network based on the dual Q network), whose input is the state space , the output is the Q value of each action , initialize network parameters .
[0046] Step 802: The agent observes the current business characteristics , Server Status , Security level information and security threat information , combined into the current state : ; According to the current strategy (such as Greedy strategy) chooses an action And execute, safety reward after executing the action And new business features , Server Status , Security level information and security threat information , thus obtaining a new state : ; Safety Rewards The setting must take into account both business security and processing efficiency. The calculation method is: ; in, , , is the weight coefficient and satisfies , Represents the security level matching reward, which is calculated as follows: ; in For Business The safety level, For Server The safety protection level, Indicates business Whether to deploy on the server The 0-1 variable on 1, No is 0; Indicates the business processing efficiency reward, which is calculated as follows: ; in For Business processing delay; Represents the security threat reward, which is calculated as: ; in, For Server security threat indicators; Step 803, transfer each step Store to experience replay pool In the process, a batch of transfer data is randomly sampled from the experience replay pool. , .
[0047] Step 804: Update network parameters according to the Dueling DQN algorithm based on dual Q networks. Two Q networks (main Q network and target Q network) are used respectively, where Indicates the status of the main Q network pair and actions The Q value of Represents the target Q network pair state and actions The target Q value is calculated as: ; in is the discount factor, are the target network parameters (regularly synchronized with the main Q network parameters). By minimizing the loss function: ; Update the main Q network parameters using the gradient descent algorithm .
[0048] Step 805, continuously repeating steps 802-804, so that the intelligent agent continuously optimizes the strategy through learning to adapt to different business security levels and server security threat conditions, and find the optimal server deployment and scheduling strategy that takes into account both security and efficiency.
[0049] By expanding the state space and adjusting the action space, integrating the business security level and server security threat information, and using the Dueling DQN algorithm based on the dual Q network for learning, the agent can learn server deployment and scheduling strategies that take security into consideration. After the implementation of this solution, server deployment and scheduling can be optimized while ensuring the safe operation of the business, avoiding business interruptions or data leakage caused by security threats, while maintaining high business processing efficiency and resource utilization.
[0050] Take the order processing system of an e-commerce platform as an example. The platform has 10 business servers. The scheduling optimization process during the "Double 11" event is as follows: Figure 2-Figure 5 As shown; Through scheduling optimization of the deep reinforcement learning model, the system showed good adaptability during peak hours: The average server load was reduced from 85% to 65% while maintaining high resource utilization; The order processing success rate increased to 99.2%, and the user experience was significantly improved; System response time was reduced by 67.9%, from 2.8 seconds to 0.9 seconds; The resource utilization balance was improved by 26.4%, effectively avoiding single-point server overload; In at least one embodiment of the present invention, a computer storage medium is provided for storing computer-readable instructions, which can execute the aforementioned method for scheduling a business server and a dedicated server when the computer-readable instructions are read.
[0051] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation mode. The above-mentioned specific implementation mode is merely illustrative and not restrictive. Under the guidance of this embodiment, ordinary technicians in this field can also make more forms of equivalent embodiments, all of which are protected by this embodiment.
Claims
1. A method for scheduling a business server and a dedicated server, characterized in that: include: Initialize the deep reinforcement learning agent, including constructing the state space, defining the action space, building a deep neural network, and initializing the exploration strategy; Interact the agent with the environment to obtain environmental feedback rewards and new states; The transfer data of the interaction between the agent and the environment is stored in the experience pool and sampled; Calculate the target Q value and loss function based on the sampled data and update the network parameters; Optimize strategies and deploy models based on preset conditions.
2. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: The action space includes: Deployment action, used to determine the distribution relationship between services and servers; Scheduling actions, used to make resource reservations and priority adjustments; Deployment Actions : Using discrete action coding, each action corresponds to the deployment relationship from a specific business to the target server; Scheduling Actions : Includes resource reservation and priority adjustment operations; Action Space One-hot encoding is used, dimension The value is determined based on the actual server cluster size.
3. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: The state space consists of: Business feature vectors, which are collected and normalized in real time by a business monitoring system; A server status matrix, wherein the server status matrix includes key indicator data of server cluster nodes; The steps to generate the state space include: Constructing business feature vector , collected in real time through the business monitoring system Dimensional business feature data, for each dimension Normalization is performed, that is, ,in is the historical mean of the feature, is the historical standard deviation of the feature; Building a server status matrix , from the server cluster Node collection key indicators, forming dimensional state matrix, and each element of the state matrix is standardized to interval; Joint state space, forming a state vector by feature concatenation , , Represents the feature concatenation operation, and the state vector dimension is , ensuring that the dimensions between features are aligned.
4. A method for scheduling business servers and dedicated servers according to claim 3, characterized in that: The agent observes the current business feature vector and server state matrix, combines them into the current state; selects and executes actions according to the current strategy, and obtains environmental feedback rewards and new states.
5. A method for scheduling business servers and dedicated servers according to claim 4, characterized in that: The environmental feedback reward formula is: ; in: is the reward value for the environment feedback, , , , is the weight coefficient, and each coefficient satisfies , , is a negative number, It represents the processing efficiency, which is calculated as follows: ; in, is the number of successfully deployed services, is the total number of business requests; Indicates resource utilization, which is calculated as follows: ; in, is the number of server cluster nodes, Indicates The amount of resources used by server nodes, Indicates Total resources of server nodes; represents the delay penalty, which is calculated as: ; in, is the number of businesses, Indicates The processing time of each business, Indicates The service level agreement of each business stipulates the time; Represents the overload penalty, which is calculated as: ; is the number of server cluster nodes, is the resource usage threshold that is set. Indicates The amount of resources used by server nodes, Indicates The total resources of server nodes.
6. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: The target Q value is calculated as follows: ; in, Indicates the calculated The target Q value, represents the first sampled from the experience replay pool The environmental feedback reward value in the transferred data, represents the discount factor, Represents all actions in the action space Take the maximum value operation, Indicates that the target network is in state Next, for action The Q value of is the target network parameter; The loss function is calculated as: ; in, Represents the loss function, which is used to measure the current network parameters The difference between the estimated Q value and the target Q value is Represents the number of transfer data sampled from the experience replay pool, Indicates that the main Q network is in state Next, for action The Q value of .
7. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: Also includes: Build a traffic prediction model to predict business traffic peaks and resource requirements; Expand the state space and action space, and add information related to resource expansion and contraction; Dynamic resource allocation is performed based on the prediction results.
8. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: Also includes: Introducing geographical information, network latency information, and energy cost information of multiple data centers; Reconstruct state space and action space to achieve cross-data center collaborative scheduling; Optimize deployment strategies across data centers and reduce overall operating costs.
9. A method for scheduling business servers and dedicated servers according to claim 1, characterized in that: Also includes: Introducing business security level information and server security threat information; Expand the state space and action space, and add safety-related scheduling strategies; Comprehensive optimization is performed based on security level matching, processing efficiency and threat level.
10. A computer storage medium, characterized in that: It is used to store computer-readable instructions, and when the computer-readable instructions are read, a business server and dedicated server scheduling method as described in any one of claims 1-9 can be executed.
Citation Information
Patent Citations
Mobile edge computing system task scheduling method based on migration and reinforcement learning
CN111858009A
Dynamic load balancing deep strong speaking learning resource scheduling method in edge environment
CN116048801A
Edge computing unloading and resource allocation method based on multi-agent reinforcement learning
CN116321293A
Dynamic edge computing server placement method based on deep reinforcement learning
CN117311952A
Time-varying task scheduling method and system based on constraint near-end strategy optimization
CN117851056A
Cited By
Scheduling strategy selection large model training method based on reinforcement learning
CN120525020A