Distributed system flow intelligent hierarchical scheduling method and system based on reinforcement learning
By adopting a dual Q network structure based on reinforcement learning in a distributed system for intelligent hierarchical scheduling of traffic, the problem that traditional methods are difficult to adapt to dynamic traffic changes is solved, and efficient resource utilization and system stability are achieved.
Patent Information
- Application Number
- CN202510027249.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional traffic scheduling methods are difficult to adapt to dynamically changing traffic loads and system resource conditions, resulting in low resource utilization, degraded service quality, and even system crashes.
The intelligent hierarchical scheduling method of distributed system traffic based on reinforcement learning is adopted to obtain traffic characteristics by building a two-layer neural network structure, and the dual Q network structure is used to generate intelligent scheduling strategies, and interface priority and resource configuration are dynamically adjusted.
It realizes adaptive hierarchical scheduling of distributed system traffic, improves resource utilization and system stability, and can effectively deal with traffic fluctuations and emergencies.
Smart Images

Figure CN119996326A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to traffic classification technology, and in particular to a distributed system traffic intelligent classification scheduling method and system based on reinforcement learning. Background Art
[0002] Distributed systems play a vital role in modern Internet applications. They improve the throughput, reliability and scalability of the system by distributing tasks to multiple nodes. However, as the scale of the system and the number of user requests continue to grow, how to effectively manage and schedule traffic has become a key challenge. Traditional traffic scheduling methods are usually based on predefined rules or static configurations, which are difficult to adapt to dynamically changing traffic loads and system resource conditions, which may lead to low resource utilization, degraded service quality, and even system crashes.
[0003] Difficult to adapt to dynamically changing traffic: Traditional traffic scheduling methods usually rely on pre-set rules or static configurations and lack the ability to perceive and adapt to real-time traffic changes. When traffic patterns change, these methods find it difficult to dynamically adjust scheduling strategies, resulting in unbalanced resource allocation and degraded service performance.
[0004] Lack of refined traffic management: Existing methods usually treat all traffic as equally important, without differentiating the quality of service requirements of different types of requests. This may affect the performance of critical business traffic and fail to meet the service level agreements of different users.
[0005] Insufficient online learning capabilities: Traditional traffic scheduling methods usually require manual intervention or offline training, and it is difficult to automatically adjust and optimize according to the state of the system during operation. This limits the system's ability to adapt to dynamic environments and increases operation and maintenance costs. Summary of the invention
[0006] The embodiments of the present invention provide a distributed system traffic intelligent hierarchical scheduling method and system based on reinforcement learning, which can solve the problems in the prior art.
[0007] According to a first aspect of the embodiments of the present invention,
[0008] Provides a distributed system traffic intelligent hierarchical scheduling method based on reinforcement learning, including:
[0009] Construct a two-layer neural network structure to obtain the traffic characteristics of the distributed system. The first layer of the two-layer neural network structure uses a long short-term memory network to collect interface call data in the distributed system in real time, record the concurrent request volume, system throughput, and resource utilization rate. The second layer of the two-layer neural network structure uses a temporal convolutional network to extract the time dimension characteristics of the interface call data and generate a multidimensional state vector for reinforcement learning. The multidimensional state vector is filtered through an attention mechanism network to obtain the core traffic characteristics.
[0010] The core traffic features are input into a reinforcement learning model, wherein the reinforcement learning model adopts a dual Q network structure, wherein the first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefit, and stores historical decision data through an experience replay mechanism, and calculates the objective function value of the dual Q network structure based on a temporal difference algorithm, and optimizes the objective function value using a deep deterministic policy gradient method, and outputs an intelligent hierarchical scheduling strategy, wherein the intelligent hierarchical scheduling strategy includes an interface priority parameter, a processing queue configuration parameter, and a resource allocation ratio parameter;
[0011] A multi-level traffic control mechanism is established based on the intelligent hierarchical scheduling strategy, and the rate of interface requests of different priorities is limited by a distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter, and regulates the processing resources according to the resource allocation ratio parameter. When the system detects performance fluctuations, the real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to realize adaptive hierarchical scheduling of distributed system traffic.
[0012] The core traffic features are input into a reinforcement learning model. The reinforcement learning model adopts a dual Q network structure. The first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefits. The historical decision data stored through the experience playback mechanism includes:
[0013] The core traffic features are input into a reinforcement learning model using a dual Q network structure, wherein the first Q network of the dual Q network structure is used to evaluate the interface priority allocation strategy under the current system state, wherein the first Q network uses a four-layer fully connected structure, wherein the input layer receives the core traffic features, the two hidden layers contain sixty-four and thirty-two neurons respectively and use a ReLU activation function, and the output layer calculates the interface priority allocation probability through a Softmax function, and generates the interface priority allocation strategy at the current moment according to the interface priority allocation probability;
[0014] The interface priority allocation strategy and the core traffic characteristics are input into the second Q network of the dual Q network structure. The second Q network is used to evaluate future system benefits. The second Q network adopts a long short-term memory network structure and includes one hundred and twenty-eight memory units. The future state of the system is predicted according to the interface priority allocation strategy and the core traffic characteristics. The state action value function is calculated based on the future state of the system, and the advantage function value is calculated in combination with the current state value function. The resource allocation strategy is determined according to the advantage function value.
[0015] A layered experience playback mechanism is constructed, which includes a short-term memory pool and a long-term memory pool. The short-term memory pool uses a sliding window method to update and store the latest 10,000 decision records in real time. Each decision record includes the core traffic characteristics, the interface priority allocation strategy, the resource configuration strategy and the corresponding system state migration information. The long-term memory pool stores 50,000 historical decision data that have been evaluated for importance.
[0016] The objective function value of the dual Q network structure is calculated based on the temporal difference algorithm, and the objective function value is optimized by the deep deterministic policy gradient method. The output intelligent hierarchical scheduling strategy includes:
[0017] Collecting state information in a distributed system, constructing a state-action transfer sequence based on the state information, wherein the state-action transfer sequence includes a current state, an execution action, an immediate reward, and a next state, and inputting the state-action transfer sequence into a dual Q network structure, wherein the dual Q network structure includes a first Q network and a second Q network, wherein the first Q network includes a first online Q network and a first target Q network, and the second Q network includes a second online Q network and a second target Q network, and wherein the first online Q network and the second online Q network use the same four-layer fully connected structure;
[0018] Inputting the next state in the state-action transfer sequence into the first target Q network and the second target Q network for evaluation respectively, comparing the output value of the first target Q network with the output value of the second target Q network, selecting the minimum value of the two output values as the target network output value, and combining the target network output value with the immediate reward in the state-action transfer sequence and a preset discount factor to generate a temporal difference target value;
[0019] Inputting the current state and the executed action in the state-action transfer sequence into the first online Q network and the second online Q network respectively, calculating a first mean square error loss according to the output value of the first online Q network and the time series difference target value, calculating a second mean square error loss according to the output value of the second online Q network and the time series difference target value, and combining the first mean square error loss and the second mean square error loss to construct an objective function of double Q learning;
[0020] Calculating the gradient of the action value function with respect to the action based on the first online Q network, multiplying the gradient of the action value function with respect to the action by the gradient of the policy network with respect to the state to obtain a policy gradient, updating the policy network parameters through an Adam optimizer according to the policy gradient, and outputting a deterministic action based on the updated policy network parameters;
[0021] Adding exploration noise that obeys normal distribution to the deterministic action generates a final action, the standard deviation of the exploration noise decays exponentially with the increase of training rounds, and generating an intelligent hierarchical scheduling strategy based on the final action.
[0022] A multi-level flow control mechanism is established based on the intelligent hierarchical scheduling strategy, and a rate limit is performed on interface requests of different priorities through a distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, and allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter. The processing resources are regulated according to the resource allocation ratio parameter, including:
[0023] Establishing a multi-level flow control mechanism based on the intelligent hierarchical scheduling strategy, wherein the multi-level flow control mechanism includes a distributed token bucket algorithm module, a processing queue allocation module and a resource regulation module;
[0024] Input the interface priority parameter into the distributed token bucket algorithm module, construct a multi-level token bucket structure based on the interface priority parameter, each priority level in the multi-level token bucket structure corresponds to an independent token bucket, and determine the capacity of each independent token bucket according to the product of the historical peak request volume of the interface and the priority adjustment factor, and the priority adjustment factor increases as the priority level increases;
[0025] A centralized token pool is constructed by using a distributed cache server. The centralized token pool is used to perform storage and allocation of tokens. The token generation rate is dynamically adjusted according to the interface priority parameter. The token generation rate is calculated by multiplying the base rate by the priority weight. The priority weight increases as the priority level increases. The token generation configuration is written into the centralized token pool.
[0026] The distributed token bucket algorithm module is used to rate limit interface requests of different priorities, atomic operation instructions are used to execute token generation and set the token validity period, and Lua scripts are used to implement the combined operation of token checking and token acquisition, and the processing rate of interface requests is controlled according to the token acquisition result;
[0027] Allocate the rate-limited interface requests to the corresponding priority queues according to the processing queue configuration parameters, wherein the priority queues include a real-time processing queue, a batch processing queue, and a low-priority queue. The real-time processing queue adopts a synchronous processing mode, the batch processing queue adopts an asynchronous batch processing mode, and the low-priority queue adopts a request queuing mode;
[0028] The processing resources are regulated according to the resource allocation ratio parameters, and the number of processing threads and processing time slices of each priority level are calculated. The number of processing threads is determined by the total number of threads and the resource ratio, and the processing time slice is determined by the benchmark time slice and the priority level. The number of processing threads and the processing time slice are allocated to each priority queue.
[0029] The distributed token bucket algorithm module is used to rate limit interface requests of different priorities, atomic operation instructions are used to execute token generation and set the token validity period, and Lua scripts are used to implement the combined operation of token checking and token acquisition. The processing rate of interface requests is controlled according to the token acquisition result, including:
[0030] Receiving an interface request, parsing priority identification information in the interface request, and generating a request identification code according to the priority identification information, wherein the request identification code includes interface type information and priority level information;
[0031] The request identification code is passed into a distributed token bucket algorithm module, the distributed token bucket algorithm module determines a corresponding token generation rate and token bucket capacity parameters according to the priority level information, and creates a corresponding token bucket data structure based on the request identification code;
[0032] The current token count field, the timestamp field and the validity period field are set in the token bucket data structure, the token generation is performed using an atomic operation instruction, and the time difference between the current time and the timestamp field is multiplied by the token generation rate parameter to calculate the number of newly added tokens;
[0033] Writing the number of newly added tokens into the current token count field through an atomic operation instruction, updating the current time into the timestamp field, setting the token validity period based on the priority level information, and writing the token expiration time into the validity period field;
[0034] Write a Lua script for token checking and token acquisition, the Lua script takes the request identification code and the number of tokens required for the request as input parameters, first verifies the validity of the validity period field inside the script, and then checks whether the current token count field meets the number of tokens required for the request;
[0035] The Lua script is called to perform a combined operation. When the validity period field has not expired and the current token count field is greater than the number of tokens required for the request, the number of tokens required for the request is deducted from the current token count field through an atomic operation, and a token acquisition success flag is returned. The processing rate of the interface request is controlled according to the token acquisition success flag.
[0036] When the system detects performance fluctuations, the real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to achieve adaptive hierarchical scheduling of distributed system traffic, including:
[0037] Collecting performance data of the distributed system, performing statistical analysis on the performance data based on a sliding time window to obtain performance statistical features; constructing a performance fluctuation detection model based on the performance statistical features, wherein the performance fluctuation detection model calculates a performance fluctuation score through a response time standard deviation, a success rate change rate, a resource usage fluctuation degree, and a queue backlog growth rate; when the performance fluctuation score exceeds a preset threshold, the performance statistical features are input into the dual Q network structure as a system state;
[0038] Combining the system state with the configuration parameters of the current intelligent hierarchical scheduling strategy to form a state-action pair, calculating an immediate reward value based on the performance fluctuation score, and constructing the system state, the configuration parameters and the immediate reward value into a state transition sequence;
[0039] assigning a priority weight to the state transition sequence according to the performance fluctuation score, storing the state transition sequence and its priority weight into a priority experience replay pool, and sampling training data from the priority experience replay pool according to the priority weight;
[0040] Inputting the training data into the dual Q network for online learning, calculating the Q value of the target state through the target Q network, updating the online Q network parameters using the Q value of the target state and the instant reward value, and synchronously updating the target Q network parameters using an exponential sliding average method;
[0041] Evaluate the current intelligent hierarchical scheduling strategy based on the updated online Q network parameters, generate optimized scheduling strategy parameters, wherein the optimized scheduling strategy parameters include priority allocation weights, resource quota ratios, and queue processing parameters, and introduce exploration noise into the optimized scheduling strategy parameters;
[0042] Applying the optimized scheduling policy parameters to the distributed system, collecting performance data after application, calculating a new performance fluctuation score, and triggering a policy rollback mechanism when the new performance fluctuation score is higher than the original performance fluctuation score;
[0043] The changing trend of the performance data is continuously monitored, the priority weight is dynamically adjusted according to the new performance fluctuation score, and the intelligent hierarchical scheduling strategy is optimized through continuous iteration to achieve adaptive scheduling control of distributed system traffic.
[0044] According to a second aspect of the embodiments of the present invention,
[0045] Provides a distributed system traffic intelligent hierarchical scheduling system based on reinforcement learning, including:
[0046] The first unit is used to construct a two-layer neural network structure to obtain the traffic characteristics of the distributed system. The first layer of the two-layer neural network structure uses a long short-term memory network to collect interface call data in the distributed system in real time, record the concurrent request volume, system throughput, and resource utilization rate. The second layer of the two-layer neural network structure uses a temporal convolutional network to extract the time dimension characteristics of the interface call data and generate a multidimensional state vector for reinforcement learning. The multidimensional state vector is filtered through an attention mechanism network to obtain the core traffic characteristics;
[0047] The second unit is used to input the core traffic characteristics into a reinforcement learning model, wherein the reinforcement learning model adopts a dual Q network structure, wherein the first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefit, stores historical decision data through an experience replay mechanism, calculates the objective function value of the dual Q network structure based on a temporal difference algorithm, optimizes the objective function value using a deep deterministic policy gradient method, and outputs an intelligent hierarchical scheduling strategy, wherein the intelligent hierarchical scheduling strategy includes an interface priority parameter, a processing queue configuration parameter, and a resource allocation ratio parameter;
[0048] The third unit is used to establish a multi-level flow control mechanism based on the intelligent hierarchical scheduling strategy, and rate limit the interface requests of different priorities through the distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter, and regulates the processing resources according to the resource allocation ratio parameter. When the system detects performance fluctuations, the real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to realize adaptive hierarchical scheduling of distributed system traffic.
[0049] According to a third aspect of the embodiments of the present invention,
[0050] An electronic device is provided, comprising:
[0051] processor;
[0052] a memory for storing processor-executable instructions;
[0053] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0054] A fourth aspect of the embodiments of the present invention is:
[0055] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0056] The beneficial effects of this application are as follows:
[0057] 1. Improve resource utilization: Through the dynamic resource allocation strategy of the reinforcement learning model, the resource allocation ratio can be intelligently adjusted according to traffic characteristics and system status to avoid resource waste and improve resource utilization.
[0058] 2. Enhance system stability: Based on the multi-level flow control mechanism and distributed token bucket algorithm, it can effectively limit the rate of interface requests of different priorities, prevent system overload, and improve system stability and robustness.
[0059] 3. Realize adaptive scheduling: This method can perform online learning and policy optimization based on real-time monitoring data, realize adaptive hierarchical scheduling of distributed system traffic, and effectively deal with traffic fluctuations and emergencies. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 A schematic diagram of a flow chart of a distributed system traffic intelligent hierarchical scheduling method based on reinforcement learning according to an embodiment of the present invention;
[0061] Figure 2 It is a structural diagram of a distributed system traffic intelligent hierarchical scheduling system based on reinforcement learning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0064] Figure 1 FIG. 1 is a flow chart of a distributed system traffic intelligent hierarchical scheduling method based on reinforcement learning according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0065] S11. Construct a two-layer neural network structure to obtain the traffic characteristics of the distributed system. The first layer of the two-layer neural network structure uses a long short-term memory network to collect interface call data in the distributed system in real time, record the concurrent request volume, system throughput, and resource utilization rate. The second layer of the two-layer neural network structure uses a temporal convolutional network to extract the time dimension characteristics of the interface call data and generate a multidimensional state vector for reinforcement learning. The multidimensional state vector is filtered through an attention mechanism network to obtain the core traffic characteristics.
[0066] S12. Input the core traffic features into a reinforcement learning model, wherein the reinforcement learning model adopts a dual Q network structure, wherein the first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefit, stores historical decision data through an experience replay mechanism, calculates the objective function value of the dual Q network structure based on a temporal difference algorithm, optimizes the objective function value using a deep deterministic policy gradient method, and outputs an intelligent hierarchical scheduling strategy, wherein the intelligent hierarchical scheduling strategy includes an interface priority parameter, a processing queue configuration parameter, and a resource allocation ratio parameter;
[0067] S13. A multi-level traffic control mechanism is established based on the intelligent hierarchical scheduling strategy, and the rate of interface requests of different priorities is limited by a distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter, and adjusts the processing resources according to the resource allocation ratio parameter. When the system detects performance fluctuations, the real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to achieve adaptive hierarchical scheduling of distributed system traffic.
[0068] In an optional implementation, the core traffic features are input into a reinforcement learning model, the reinforcement learning model adopts a dual Q network structure, the first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefit, and the historical decision data is stored through the experience playback mechanism, including:
[0069] The core traffic features are input into a reinforcement learning model using a dual Q network structure, wherein the first Q network of the dual Q network structure is used to evaluate the interface priority allocation strategy under the current system state, wherein the first Q network uses a four-layer fully connected structure, wherein the input layer receives the core traffic features, the two hidden layers contain sixty-four and thirty-two neurons respectively and use a ReLU activation function, and the output layer calculates the interface priority allocation probability through a Softmax function, and generates the interface priority allocation strategy at the current moment according to the interface priority allocation probability;
[0070] The interface priority allocation strategy and the core traffic characteristics are input into the second Q network of the dual Q network structure. The second Q network is used to evaluate future system benefits. The second Q network adopts a long short-term memory network structure and includes one hundred and twenty-eight memory units. The future state of the system is predicted according to the interface priority allocation strategy and the core traffic characteristics. The state action value function is calculated based on the future state of the system, and the advantage function value is calculated in combination with the current state value function. The resource allocation strategy is determined according to the advantage function value.
[0071] A layered experience playback mechanism is constructed, which includes a short-term memory pool and a long-term memory pool. The short-term memory pool uses a sliding window method to update and store the latest 10,000 decision records in real time. Each decision record includes the core traffic characteristics, the interface priority allocation strategy, the resource configuration strategy and the corresponding system state migration information. The long-term memory pool stores 50,000 historical decision data that have been evaluated for importance.
[0072] The dynamic resource allocation method based on reinforcement learning aims to dynamically adjust interface priority and resource configuration according to core traffic characteristics, thereby improving the overall performance of the system. The core of this method is to adopt a reinforcement learning model with a dual Q network structure and combine it with a layered experience replay mechanism for training and optimization.
[0073] First, extract the core traffic features. The core traffic features may include but are not limited to: traffic size, traffic source, traffic destination, traffic type, request delay, error rate, etc. For example, the traffic size in the past hour can be recorded as the traffic size feature; the traffic source can be divided into domestic traffic and overseas traffic as the traffic source feature. These features will be used as the input of the reinforcement learning model.
[0074] Next, the extracted core traffic features are input into the first Q network. The network adopts a four-layer fully connected structure. The input layer receives the core traffic features. For example, assuming there are 5 core traffic features, the number of neurons in the input layer is 5. The first hidden layer contains 64 neurons and uses the ReLU activation function. The second hidden layer contains 32 neurons and also uses the ReLU activation function. The number of neurons in the output layer is equal to the number of interfaces, and the Softmax function is used to calculate the interface priority assignment probability. For example, assuming that the system has 3 interfaces, the output layer has 3 neurons, and the Softmax function will output 3 probability values, representing the priority of each interface. Based on these probability values, the interface priority assignment strategy at the current moment is generated. For example, the interface with the highest probability will get the highest priority.
[0075] Then, the interface priority allocation strategy and core traffic characteristics are input into the second Q network. The network adopts the long short-term memory network (LSTM) structure and contains 128 memory units. The LSTM network can capture long-term dependencies in time series data and is suitable for processing traffic data. The second Q network predicts the future state of the system based on the interface priority allocation strategy and core traffic characteristics. For example, the load of each interface in the next hour is predicted. Based on the predicted future state of the system, the state action value function is calculated, and the advantage function value is calculated in combination with the current state value function. The advantage function value reflects the advantage of the current strategy over other strategies. The resource allocation strategy is determined based on the advantage function value. For example, the strategy with the highest advantage function value will be selected to allocate computing resources, storage resources, network resources, etc.
[0076] In order to train the reinforcement learning model, a hierarchical experience replay mechanism is constructed. The mechanism includes a short-term memory pool and a long-term memory pool. The short-term memory pool uses a sliding window method to update and store the latest 10,000 decision records in real time. Each decision record contains core traffic characteristics, interface priority allocation strategy, resource configuration strategy and corresponding system state migration information. For example, the load change of the system after a certain strategy is adopted is recorded. The long-term memory pool stores 50,000 historical decision data that have been evaluated for importance. Importance evaluation is used to measure the value of historical data, and more valuable data will be retained in the long-term memory pool. Through the experience replay mechanism, the reinforcement learning model can learn from historical data and continuously optimize the strategy.
[0077] Specific data case: Assume that the system has three interfaces A, B, and C. The extracted core traffic features include traffic size and traffic type. The current traffic size is 1000, and the traffic type is video streaming. The interface priority allocation probability output by the first Q network is A: 0.5, B: 0.3, C: 0.2. Therefore, interface A gets the highest priority. The second Q network predicts that the load of interfaces A, B, and C will be 80%, 50%, and 30% respectively in the next hour. Based on these predicted values, calculate the state action value function and advantage function value. Assume that the final resource allocation strategy is: allocate 60% of computing resources to interface A, 30% of computing resources to interface B, and 10% of computing resources to interface C. Store these data and system state migration information in the experience playback mechanism.
[0078] The solution of this application can:
[0079] Improve resource utilization: By dynamically adjusting resource configuration, resources can be concentrated on more important interfaces to avoid resource waste, thereby improving resource utilization. Improve system performance: By optimizing interface priority and resource configuration, request latency can be reduced and throughput can be increased, thereby improving overall system performance. Enhance system stability: By predicting future system status and adjusting resource configuration in advance, system overload can be avoided and system stability can be enhanced.
[0080] In an optional implementation, the objective function value of the dual Q network structure is calculated based on a temporal difference algorithm, the objective function value is optimized using a deep deterministic policy gradient method, and the output intelligent hierarchical scheduling strategy includes:
[0081] Collecting state information in a distributed system, constructing a state-action transfer sequence based on the state information, wherein the state-action transfer sequence includes a current state, an execution action, an immediate reward, and a next state, and inputting the state-action transfer sequence into a dual Q network structure, wherein the dual Q network structure includes a first Q network and a second Q network, wherein the first Q network includes a first online Q network and a first target Q network, and the second Q network includes a second online Q network and a second target Q network, and wherein the first online Q network and the second online Q network use the same four-layer fully connected structure;
[0082] Inputting the next state in the state-action transfer sequence into the first target Q network and the second target Q network for evaluation respectively, comparing the output value of the first target Q network with the output value of the second target Q network, selecting the minimum value of the two output values as the target network output value, and combining the target network output value with the immediate reward in the state-action transfer sequence and a preset discount factor to generate a temporal difference target value;
[0083] Inputting the current state and the executed action in the state-action transfer sequence into the first online Q network and the second online Q network respectively, calculating a first mean square error loss according to the output value of the first online Q network and the time series difference target value, calculating a second mean square error loss according to the output value of the second online Q network and the time series difference target value, and combining the first mean square error loss and the second mean square error loss to construct an objective function of double Q learning;
[0084] Calculating the gradient of the action value function with respect to the action based on the first online Q network, multiplying the gradient of the action value function with respect to the action by the gradient of the policy network with respect to the state to obtain a policy gradient, updating the policy network parameters through an Adam optimizer according to the policy gradient, and outputting a deterministic action based on the updated policy network parameters;
[0085] Adding exploration noise that obeys normal distribution to the deterministic action generates a final action, the standard deviation of the exploration noise decays exponentially with the increase of training rounds, and generating an intelligent hierarchical scheduling strategy based on the final action.
[0086] The distributed system intelligent hierarchical scheduling method includes the following steps:
[0087] First, collect the distributed system status information. In this step, you need to comprehensively collect information such as the load status, resource usage, and task queue length of each node in the system. For example, you can collect indicators such as node CPU usage, memory occupancy, network IO, disk IO, and the number of tasks currently being executed and the number of tasks waiting to be executed. This information will be used to build a complete description of the system status.
[0088] Next, based on the collected state information, a state-action transition sequence is constructed. Each state-action transition sequence contains four elements: the current state, the executed action, the immediate reward obtained, and the resulting next state. For example, the current state is "node A CPU usage 80%, memory occupancy 60%, task queue length 10", the executed action is "schedule a new task to node B", the immediate reward obtained is "task completion time reduced by 10%", and the resulting next state is "node A CPU usage 70%, memory occupancy 50%, task queue length 9, node B CPU usage 60%, memory occupancy 70%, task queue length 1".
[0089] Then, the state-action transfer sequence is input into a dual Q network structure. The structure contains two Q networks: the first Q network and the second Q network. Each Q network contains an online network and a target network. All online networks use the same four-layer fully connected neural network structure. Taking the first online Q network as an example, the number of input layer neurons is equal to the dimension of the state information. For example, in the above example, the state information includes the CPU usage, memory occupancy and task queue length of nodes A and B, with a total of 6 dimensions, so the number of input layer neurons is 6. The number of hidden layer neurons can be set according to actual conditions, for example, 128, 64, and 32 respectively. The number of output layer neurons is equal to the number of optional actions. For example, you can choose to schedule the task to node A, node B or node C, then the number of output layer neurons is 3. The second Q network structure is exactly the same as the first Q network structure.
[0090] The next state in the state-action transition sequence is input into the first target Q network and the second target Q network for evaluation. The output values of the two networks are compared, and the smaller value is selected as the target network output value. The target network output value is combined with the immediate reward in the state-action transition sequence and a preset discount factor (e.g., 0.95) to generate a temporal difference target value.
[0091] The current state and the executed action in the state-action transfer sequence are input to the first online Q network and the second online Q network respectively. The mean square error between the output value of the first online Q network and the target value of the time series difference is calculated as the first mean square error loss. Similarly, the mean square error between the output value of the second online Q network and the target value of the time series difference is calculated as the second mean square error loss. These two mean square error losses are combined to construct the objective function of dual Q learning.
[0092] Based on the first online Q network, the gradient of the action-value function with respect to the action is calculated, and the gradient is multiplied by the gradient of the policy network with respect to the state to obtain the policy gradient. The policy network is also a neural network, whose input is the system state and output is the action. The Adam optimizer is used to update the policy network parameters according to the policy gradient. The deterministic action is output based on the updated policy network parameters.
[0093] Add exploration noise that follows a normal distribution to the deterministic action to generate the final action. The standard deviation of the exploration noise decays exponentially with the increase of training rounds. For example, the initial standard deviation is set to 1, and decays by 0.9 times every 1000 training rounds. Generate an intelligent hierarchical scheduling strategy based on the final action, such as scheduling tasks to specified nodes.
[0094] The solution of this application can:
[0095] Improve resource utilization: Through intelligent hierarchical scheduling, system resources can be used more effectively, resource waste can be avoided, and overall resource utilization can be improved. Reduce task completion time: This method can dynamically adjust the scheduling strategy according to the system status and assign tasks to the most suitable nodes for execution, thereby reducing task completion time and improving system efficiency. Enhance system stability: Through intelligent scheduling, the system load can be balanced to avoid overloading of individual nodes, thereby improving system stability and reliability.
[0096] In an optional implementation, a multi-level flow control mechanism is established based on the intelligent hierarchical scheduling strategy, and a rate limit is performed on interface requests of different priorities through a distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, and allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter. The processing resources are regulated according to the resource allocation ratio parameter, including:
[0097] Establishing a multi-level flow control mechanism based on the intelligent hierarchical scheduling strategy, wherein the multi-level flow control mechanism includes a distributed token bucket algorithm module, a processing queue allocation module and a resource regulation module;
[0098] Input the interface priority parameter into the distributed token bucket algorithm module, construct a multi-level token bucket structure based on the interface priority parameter, each priority level in the multi-level token bucket structure corresponds to an independent token bucket, and determine the capacity of each independent token bucket according to the product of the historical peak request volume of the interface and the priority adjustment factor, and the priority adjustment factor increases as the priority level increases;
[0099] A centralized token pool is constructed by using a distributed cache server. The centralized token pool is used to perform storage and allocation of tokens. The token generation rate is dynamically adjusted according to the interface priority parameter. The token generation rate is calculated by multiplying the base rate by the priority weight. The priority weight increases as the priority level increases. The token generation configuration is written into the centralized token pool.
[0100] The distributed token bucket algorithm module is used to rate limit interface requests of different priorities, atomic operation instructions are used to execute token generation and set the token validity period, and Lua scripts are used to implement the combined operation of token checking and token acquisition, and the processing rate of interface requests is controlled according to the token acquisition result;
[0101] Allocate the rate-limited interface requests to the corresponding priority queues according to the processing queue configuration parameters, wherein the priority queues include a real-time processing queue, a batch processing queue, and a low-priority queue. The real-time processing queue adopts a synchronous processing mode, the batch processing queue adopts an asynchronous batch processing mode, and the low-priority queue adopts a request queuing mode;
[0102] The processing resources are regulated according to the resource allocation ratio parameters, and the number of processing threads and processing time slices of each priority level are calculated. The number of processing threads is determined by the total number of threads and the resource ratio, and the processing time slice is determined by the benchmark time slice and the priority level. The number of processing threads and the processing time slice are allocated to each priority queue.
[0103] A multi-level flow control method based on intelligent hierarchical scheduling strategy is used to control the processing rate and resource allocation of interface requests with different priorities.
[0104] First, obtain input information such as interface priority parameters, historical peak request volume, processing queue configuration parameters, and resource allocation ratio parameters. Assume that the interface priority is divided into three levels: high, medium, and low, corresponding to values 3, 2, and 1, respectively.
[0105] Then, a multi-level token bucket structure is constructed. Each priority corresponds to an independent token bucket. For example, the historical peak request volume of the high priority interface is 1000 times / second, the historical peak request volume of the medium priority interface is 500 times / second, and the historical peak request volume of the low priority interface is 200 times / second. The priority adjustment factors are set to 3 for high priority, 2 for medium priority, and 1 for low priority. Then the capacity of the high priority token bucket is 1000*3=3000, the capacity of the medium priority token bucket is 500*2=1000, and the capacity of the low priority token bucket is 200*1=200.
[0106] Next, use a distributed cache server (such as Redis) to build a centralized token pool for storing and distributing tokens. Assume that the base rate is 100 tokens / second, and the weights of high, medium, and low priorities are 3, 2, and 1, respectively. The high-priority token generation rate is 100*3=300 / second, the medium-priority token generation rate is 100*2=200 / second, and the low-priority token generation rate is 100*1=100 / second. Write these configurations to Redis. Use atomic operation instructions (such as INCRBY) to generate tokens, and set the token validity period, for example, 1 second. Use Lua scripts to implement the combined operation of token checking and acquisition to ensure atomicity.
[0107] When an interface request arrives, an attempt is made to obtain a token from the corresponding token bucket according to the interface priority parameter. For example, when a high-priority interface request arrives, an attempt is made to obtain a token from the high-priority token bucket. If the acquisition is successful, the request is allowed to continue processing; if the acquisition fails, the request is rejected or other processing is performed, such as delayed retry.
[0108] According to the processing queue configuration parameters, the requests that obtain the token are assigned to the corresponding priority queue. Assume that the configuration parameters specify that high-priority requests enter the real-time processing queue, medium-priority requests enter the batch processing queue, and low-priority requests enter the low-priority queue. The real-time processing queue uses synchronous processing to process requests immediately. The batch processing queue uses asynchronous batch processing to process requests uniformly after accumulating a certain number of requests. The low-priority queue uses request queuing to wait for resources to be idle before processing.
[0109] Finally, the processing resources are regulated according to the resource allocation ratio parameters. Assuming that the total number of threads is 100, the resource proportions of high, medium and low priorities are 50%, 30% and 20% respectively. Then the number of high priority processing threads is 100*50%=50, the number of medium priority processing threads is 100*30%=30, and the number of low priority processing threads is 100*20%=20. Assuming that the base time slice is 10ms, the high priority processing time slice is 10ms*3=30ms, the medium priority processing time slice is 10ms*2=20ms, and the low priority processing time slice is 10ms*1=10ms. The calculated number of threads and time slices are allocated to each priority queue to control the processing capacity of each queue.
[0110] The solution of this application can:
[0111] Improve resource utilization: Dynamically adjust the token generation rate and resource allocation ratio according to the request priority, give priority to the processing of high-priority requests, avoid resource waste, and improve overall resource utilization. Ensure service stability: Through a multi-level flow control mechanism, effectively limit the request rate of interfaces with different priorities, prevent system overload, and ensure service stability and reliability. Flexible priority control: Support custom interface priority, priority adjustment factor, priority weight, resource allocation ratio and other parameters, and provide flexible priority control strategies to meet the needs of different business scenarios.
[0112] In an optional implementation, the distributed token bucket algorithm module is used to rate limit interface requests of different priorities, atomic operation instructions are used to execute token generation and set token validity period, and Lua scripts are used to implement combined operations of token checking and token acquisition. Controlling the processing rate of interface requests according to the token acquisition result includes:
[0113] Receiving an interface request, parsing priority identification information in the interface request, and generating a request identification code according to the priority identification information, wherein the request identification code includes interface type information and priority level information;
[0114] The request identification code is passed into a distributed token bucket algorithm module, the distributed token bucket algorithm module determines a corresponding token generation rate and token bucket capacity parameters according to the priority level information, and creates a corresponding token bucket data structure based on the request identification code;
[0115] The current token count field, the timestamp field and the validity period field are set in the token bucket data structure, the token generation is performed using an atomic operation instruction, and the time difference between the current time and the timestamp field is multiplied by the token generation rate parameter to calculate the number of newly added tokens;
[0116] Writing the number of newly added tokens into the current token count field through an atomic operation instruction, updating the current time into the timestamp field, setting the token validity period based on the priority level information, and writing the token expiration time into the validity period field;
[0117] Write a Lua script for token checking and token acquisition, the Lua script takes the request identification code and the number of tokens required for the request as input parameters, first verifies the validity of the validity period field inside the script, and then checks whether the current token count field meets the number of tokens required for the request;
[0118] The Lua script is called to perform a combined operation. When the validity period field has not expired and the current token count field is greater than the number of tokens required for the request, the number of tokens required for the request is deducted from the current token count field through an atomic operation, and a token acquisition success flag is returned. The processing rate of the interface request is controlled according to the token acquisition success flag.
[0119] A distributed token bucket algorithm module is a method to implement rate limiting of interface requests of different priorities. Its core lies in using atomic operations and Lua scripts to achieve efficient token management and request control.
[0120] First, receive the incoming interface request and parse the priority identification information contained in the request. For example, a request may contain a priority identification of "high" or "low". Generate a unique request identification code based on the parsed priority identification information and interface type information. For example, if the interface type is "video" and the priority is "high", the generated request identification code may be "video_high".
[0121] Next, the generated request identification code is passed to the distributed token bucket algorithm module. The module determines the corresponding token generation rate and token bucket capacity parameters based on the priority level information contained in the request identification code. For example, for a "high" priority request, the token generation rate is 100 tokens per second and the token bucket capacity is 200 tokens; for a "low" priority request, the token generation rate is 50 tokens per second and the token bucket capacity is 100 tokens. The module creates the corresponding token bucket data structure in the distributed cache based on the request identification code. The data structure contains three fields: current token count, timestamp, and validity period.
[0122] Then, perform token generation in the token bucket data structure. Use atomic operation instructions to obtain the current time and calculate the time difference between the current time and the timestamp field. For example, if the current time is 1678886400 seconds and the value of the timestamp field is 1678886300 seconds, the time difference is 100 seconds. Multiply the time difference by the token generation rate to calculate the number of newly added tokens. For example, if the token generation rate of "high" priority is 100 tokens per second, the number of newly added tokens is 10,000. Use atomic operation instructions to add the number of newly added tokens to the current token count field and update the current time to the timestamp field. At the same time, set the token validity period according to the priority level information, and write the token expiration time into the validity period field. For example, if the token validity period of "high" priority is 60 seconds, the token expiration time is 1678886460 seconds.
[0123] After that, write a Lua script for token checking and token acquisition. The script takes the request identification code and the number of tokens required for the request as input parameters. For example, the request identification code is "video_high" and the number of tokens required for the request is 10. Inside the script, first verify the validity of the validity period field, for example, check whether the current time is less than the token expiration time. Then, check whether the current token count field meets the number of tokens required for the request, for example, check whether the current token count is greater than or equal to 10.
[0124] Finally, call the Lua script to perform the combined operation. If the validity period field has not expired and the current token count field is greater than the number of tokens required for the request, use the atomic operation instruction to deduct the number of tokens required for the request from the current token count field, for example, deduct 10 from the current token count, and return the token acquisition success flag. Control the processing rate of the interface request based on the token acquisition success flag. If the token acquisition is successful, the interface request is allowed to continue processing; if the token acquisition fails, the interface request is rejected.
[0125] The solution of this application can:
[0126] Fine-grained control: By setting different token generation rates and token bucket capacities for requests of different priorities, fine-grained control of requests of different priorities can be achieved to ensure the processing speed of high-priority requests. High performance: Using atomic operation instructions to perform token generation, token checking, and token acquisition operations avoids lock contention and improves the concurrent processing capabilities of the system. Using Lua scripts to combine token checking and token acquisition operations reduces the number of network requests and further improves performance. Distributed support: Storing the token bucket data structure in a distributed cache can achieve rate limiting in a distributed environment and facilitate horizontal expansion.
[0127] In an optional implementation, when the system detects performance fluctuations, real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to achieve adaptive hierarchical scheduling of distributed system traffic, including:
[0128] Collecting performance data of the distributed system, performing statistical analysis on the performance data based on a sliding time window to obtain performance statistical features; constructing a performance fluctuation detection model based on the performance statistical features, wherein the performance fluctuation detection model calculates a performance fluctuation score through a response time standard deviation, a success rate change rate, a resource usage fluctuation degree, and a queue backlog growth rate; when the performance fluctuation score exceeds a preset threshold, the performance statistical features are input into the dual Q network structure as a system state;
[0129] Combining the system state with the configuration parameters of the current intelligent hierarchical scheduling strategy to form a state-action pair, calculating an immediate reward value based on the performance fluctuation score, and constructing the system state, the configuration parameters and the immediate reward value into a state transition sequence;
[0130] assigning a priority weight to the state transition sequence according to the performance fluctuation score, storing the state transition sequence and its priority weight into a priority experience replay pool, and sampling training data from the priority experience replay pool according to the priority weight;
[0131] Inputting the training data into the dual Q network for online learning, calculating the Q value of the target state through the target Q network, updating the online Q network parameters using the Q value of the target state and the instant reward value, and synchronously updating the target Q network parameters using an exponential sliding average method;
[0132] Evaluate the current intelligent hierarchical scheduling strategy based on the updated online Q network parameters, generate optimized scheduling strategy parameters, wherein the optimized scheduling strategy parameters include priority allocation weights, resource quota ratios, and queue processing parameters, and introduce exploration noise into the optimized scheduling strategy parameters;
[0133] Applying the optimized scheduling policy parameters to the distributed system, collecting performance data after application, calculating a new performance fluctuation score, and triggering a policy rollback mechanism when the new performance fluctuation score is higher than the original performance fluctuation score;
[0134] The changing trend of the performance data is continuously monitored, the priority weight is dynamically adjusted according to the new performance fluctuation score, and the intelligent hierarchical scheduling strategy is optimized through continuous iteration to achieve adaptive scheduling control of distributed system traffic.
[0135] An adaptive hierarchical scheduling method is used to optimize the traffic scheduling of a distributed system, and its specific implementation method is as follows:
[0136] First, collect the performance data of the distributed system. These performance data include the response time, success rate, CPU usage, memory usage, queue backlog length, etc. of each node. For example, collect the above performance indicators of all nodes every 1 second.
[0137] Next, statistical analysis is performed on the performance data based on the sliding time window to obtain performance statistical characteristics. For example, a 60-second sliding time window is used to calculate the average, standard deviation, maximum, and minimum values of each performance indicator in each time window. Assume that in a certain time window, the average response time of node A is 50 milliseconds, the standard deviation is 5 milliseconds; the success rate is 99%, and the change rate is -0.1%; the average CPU usage is 70%, and the volatility is 10%; the average queue backlog length is 100, and the growth rate is 5%.
[0138] Then, a performance fluctuation detection model is constructed based on the performance statistical characteristics. The model calculates the performance fluctuation score through the standard deviation of response time, the success rate change rate, the resource usage fluctuation degree and the queue backlog growth rate. For example, the performance fluctuation score can be calculated by multiplying the above four indicators by the preset weight coefficients and then adding them. Assuming that the weight coefficients are 0.2, 0.3, 0.2 and 0.3 respectively, the performance fluctuation score of node A is: 0.2*5+0.3*(-0.1)+0.2*10+0.3*5=3.67.
[0139] When the performance fluctuation score exceeds a preset threshold, for example, the threshold is set to 3, the performance statistical characteristics are input as the system state into the dual Q network structure. At this time, the performance statistical characteristics of node A (50 milliseconds, 5 milliseconds, 99%, -0.1%, 70%, 10%, 100, 5%) will be input as the system state.
[0140] The system state is combined with the configuration parameters of the current intelligent hierarchical scheduling strategy to form a state-action pair. Assume that the current policy parameters are: priority allocation weight is 0.5, resource quota ratio is 0.8, and queue processing parameter is 2. Then the state-action pair is: (system state, 0.5, 0.8, 2).
[0141] The immediate reward value is calculated based on the performance volatility score. For example, the reward value can be set to the negative of the performance volatility score, i.e. -3.67. The system state, configuration parameters, and immediate reward value are constructed as a state transition sequence: (system state, 0.5, 0.8, 2, -3.67).
[0142] Assign priority weights to state transition sequences based on the performance volatility score. For example, the weight can be set to the absolute value of the performance volatility score, i.e. 3.67. Store the state transition sequence and its priority weights in the priority experience replay pool.
[0143] The training data is sampled from the priority experience replay pool according to the priority weight. For example, the state transition sequence with a high priority weight is more likely to be sampled.
[0144] The training data is input into the dual Q network for online learning. The Q value of the target state is calculated through the target Q network, and the Q value of the target state and the instant reward value are used to update the online Q network parameters, and the exponential sliding average method is used to synchronously update the target Q network parameters.
[0145] Based on the updated online Q network parameters, the current intelligent hierarchical scheduling strategy is evaluated to generate optimized scheduling strategy parameters, for example, the priority allocation weight is adjusted to 0.6, the resource quota ratio is adjusted to 0.9, and the queue processing parameter is adjusted to 3. Exploration noise is introduced into the optimized scheduling strategy parameters, for example, a small random perturbation is added to each parameter.
[0146] Apply the optimized scheduling policy parameters to the distributed system, collect the performance data after application, and calculate the new performance fluctuation score. Assume that the new performance fluctuation score is 2.5.
[0147] When the new performance fluctuation score is higher than the original performance fluctuation score, for example, 2.5>3.67 does not hold, the policy rollback mechanism is not triggered. Otherwise, it rolls back to the previous policy parameters.
[0148] Continuously monitor the changing trends of performance data, dynamically adjust the priority weights according to the new performance fluctuation scores, and achieve adaptive scheduling control of distributed system traffic by continuously iterating and optimizing the intelligent hierarchical scheduling strategy.
[0149] The solution of this application can:
[0150] Improve system stability: Through real-time monitoring and dynamic adjustment of scheduling strategies, you can effectively respond to performance fluctuations, avoid system crashes or service interruptions, and thus improve system stability and reliability. Optimize resource utilization: Through intelligent hierarchical scheduling, resources can be dynamically allocated according to traffic priority and resource requirements to avoid resource waste and improve resource utilization. Improve user experience: Through adaptive scheduling control, scheduling strategies can be dynamically adjusted according to traffic changes to ensure the service quality of high-priority traffic, thereby improving user experience.
[0151] Figure 2 FIG. 1 is a schematic diagram of the structure of a distributed system traffic intelligent hierarchical scheduling system based on reinforcement learning according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0152] The first unit is used to construct a two-layer neural network structure to obtain the traffic characteristics of the distributed system. The first layer of the two-layer neural network structure uses a long short-term memory network to collect interface call data in the distributed system in real time, record the concurrent request volume, system throughput, and resource utilization rate. The second layer of the two-layer neural network structure uses a temporal convolutional network to extract the time dimension characteristics of the interface call data and generate a multidimensional state vector for reinforcement learning. The multidimensional state vector is filtered through an attention mechanism network to obtain the core traffic characteristics;
[0153] The second unit is used to input the core traffic characteristics into a reinforcement learning model, wherein the reinforcement learning model adopts a dual Q network structure, wherein the first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefit, stores historical decision data through an experience replay mechanism, calculates the objective function value of the dual Q network structure based on a temporal difference algorithm, optimizes the objective function value using a deep deterministic policy gradient method, and outputs an intelligent hierarchical scheduling strategy, wherein the intelligent hierarchical scheduling strategy includes an interface priority parameter, a processing queue configuration parameter, and a resource allocation ratio parameter;
[0154] The third unit is used to establish a multi-level flow control mechanism based on the intelligent hierarchical scheduling strategy, and rate limit the interface requests of different priorities through the distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter, and regulates the processing resources according to the resource allocation ratio parameter. When the system detects performance fluctuations, the real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to realize adaptive hierarchical scheduling of distributed system traffic.
[0155] According to a third aspect of the embodiments of the present invention,
[0156] An electronic device is provided, comprising:
[0157] processor;
[0158] a memory for storing processor-executable instructions;
[0159] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0160] A fourth aspect of the embodiments of the present invention is:
[0161] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0162] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A distributed system traffic intelligent hierarchical scheduling method based on reinforcement learning, characterized in that: include: Construct a two-layer neural network structure to obtain the traffic characteristics of the distributed system. The first layer of the two-layer neural network structure uses a long short-term memory network to collect interface call data in the distributed system in real time, record the concurrent request volume, system throughput, and resource utilization rate. The second layer of the two-layer neural network structure uses a temporal convolutional network to extract the time dimension characteristics of the interface call data and generate a multidimensional state vector for reinforcement learning. The multidimensional state vector is filtered through an attention mechanism network to obtain the core traffic characteristics. The core traffic features are input into a reinforcement learning model, wherein the reinforcement learning model adopts a dual Q network structure, wherein the first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefit, and stores historical decision data through an experience replay mechanism, and calculates the objective function value of the dual Q network structure based on a temporal difference algorithm, and optimizes the objective function value using a deep deterministic policy gradient method, and outputs an intelligent hierarchical scheduling strategy, wherein the intelligent hierarchical scheduling strategy includes an interface priority parameter, a processing queue configuration parameter, and a resource allocation ratio parameter; A multi-level traffic control mechanism is established based on the intelligent hierarchical scheduling strategy, and the rate of interface requests of different priorities is limited by a distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter, and regulates the processing resources according to the resource allocation ratio parameter. When the system detects performance fluctuations, the real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to realize adaptive hierarchical scheduling of distributed system traffic.
2. The method according to claim 1, characterized in that The core traffic features are input into a reinforcement learning model. The reinforcement learning model adopts a dual Q network structure. The first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefits. The historical decision data stored through the experience playback mechanism includes: The core traffic features are input into a reinforcement learning model using a dual Q network structure, wherein the first Q network of the dual Q network structure is used to evaluate the interface priority allocation strategy under the current system state, wherein the first Q network uses a four-layer fully connected structure, wherein the input layer receives the core traffic features, the two hidden layers contain sixty-four and thirty-two neurons respectively and use a ReLU activation function, and the output layer calculates the interface priority allocation probability through a Softmax function, and generates the interface priority allocation strategy at the current moment according to the interface priority allocation probability; The interface priority allocation strategy and the core traffic characteristics are input into the second Q network of the dual Q network structure. The second Q network is used to evaluate future system benefits. The second Q network adopts a long short-term memory network structure and includes one hundred and twenty-eight memory units. The future state of the system is predicted according to the interface priority allocation strategy and the core traffic characteristics. The state action value function is calculated based on the future state of the system, and the advantage function value is calculated in combination with the current state value function. The resource allocation strategy is determined according to the advantage function value. A layered experience playback mechanism is constructed, which includes a short-term memory pool and a long-term memory pool. The short-term memory pool uses a sliding window method to update and store the latest 10,000 decision records in real time. Each decision record includes the core traffic characteristics, the interface priority allocation strategy, the resource configuration strategy and the corresponding system state migration information. The long-term memory pool stores 50,000 historical decision data that have been evaluated for importance.
3. The method according to claim 1, characterized in that The objective function value of the dual Q network structure is calculated based on the temporal difference algorithm, and the objective function value is optimized by the deep deterministic policy gradient method. The output intelligent hierarchical scheduling strategy includes: Collecting state information in a distributed system, constructing a state-action transfer sequence based on the state information, wherein the state-action transfer sequence includes a current state, an execution action, an immediate reward, and a next state, and inputting the state-action transfer sequence into a dual Q network structure, wherein the dual Q network structure includes a first Q network and a second Q network, wherein the first Q network includes a first online Q network and a first target Q network, and the second Q network includes a second online Q network and a second target Q network, and wherein the first online Q network and the second online Q network use the same four-layer fully connected structure; Inputting the next state in the state-action transfer sequence into the first target Q network and the second target Q network for evaluation respectively, comparing the output value of the first target Q network with the output value of the second target Q network, selecting the minimum value of the two output values as the target network output value, and combining the target network output value with the immediate reward in the state-action transfer sequence and a preset discount factor to generate a temporal difference target value; Inputting the current state and the executed action in the state-action transfer sequence into the first online Q network and the second online Q network respectively, calculating a first mean square error loss according to the output value of the first online Q network and the time series difference target value, calculating a second mean square error loss according to the output value of the second online Q network and the time series difference target value, and combining the first mean square error loss and the second mean square error loss to construct an objective function of double Q learning; Calculating the gradient of the action value function with respect to the action based on the first online Q network, multiplying the gradient of the action value function with respect to the action by the gradient of the policy network with respect to the state to obtain a policy gradient, updating the policy network parameters through an Adam optimizer according to the policy gradient, and outputting a deterministic action based on the updated policy network parameters; Adding exploration noise that obeys normal distribution to the deterministic action generates a final action, the standard deviation of the exploration noise decays exponentially with the increase of training rounds, and generating an intelligent hierarchical scheduling strategy based on the final action.
4. The method according to claim 1, characterized in that: A multi-level flow control mechanism is established based on the intelligent hierarchical scheduling strategy, and a rate limit is performed on interface requests of different priorities through a distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, and allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter. The processing resources are regulated according to the resource allocation ratio parameter, including: Establishing a multi-level flow control mechanism based on the intelligent hierarchical scheduling strategy, wherein the multi-level flow control mechanism includes a distributed token bucket algorithm module, a processing queue allocation module and a resource regulation module; Input the interface priority parameter into the distributed token bucket algorithm module, construct a multi-level token bucket structure based on the interface priority parameter, each priority level in the multi-level token bucket structure corresponds to an independent token bucket, and determine the capacity of each independent token bucket according to the product of the historical peak request volume of the interface and the priority adjustment factor, and the priority adjustment factor increases as the priority level increases; A centralized token pool is constructed by using a distributed cache server. The centralized token pool is used to perform storage and allocation of tokens. The token generation rate is dynamically adjusted according to the interface priority parameter. The token generation rate is calculated by multiplying the base rate by the priority weight. The priority weight increases as the priority level increases. The token generation configuration is written into the centralized token pool. The distributed token bucket algorithm module is used to rate limit interface requests of different priorities, atomic operation instructions are used to execute token generation and set the token validity period, and Lua scripts are used to implement the combined operation of token checking and token acquisition, and the processing rate of interface requests is controlled according to the token acquisition result; Allocate the rate-limited interface requests to the corresponding priority queues according to the processing queue configuration parameters, wherein the priority queues include a real-time processing queue, a batch processing queue, and a low-priority queue. The real-time processing queue adopts a synchronous processing mode, the batch processing queue adopts an asynchronous batch processing mode, and the low-priority queue adopts a request queuing mode; The processing resources are regulated according to the resource allocation ratio parameters, and the number of processing threads and processing time slices of each priority level are calculated. The number of processing threads is determined by the total number of threads and the resource ratio, and the processing time slice is determined by the benchmark time slice and the priority level. The number of processing threads and the processing time slice are allocated to each priority queue.
5. The method according to claim 4, characterized in that The distributed token bucket algorithm module is used to rate limit interface requests of different priorities, atomic operation instructions are used to execute token generation and set the token validity period, and Lua scripts are used to implement the combined operation of token checking and token acquisition. The processing rate of interface requests is controlled according to the token acquisition result, including: Receiving an interface request, parsing priority identification information in the interface request, and generating a request identification code according to the priority identification information, wherein the request identification code includes interface type information and priority level information; The request identification code is passed into a distributed token bucket algorithm module, the distributed token bucket algorithm module determines a corresponding token generation rate and token bucket capacity parameters according to the priority level information, and creates a corresponding token bucket data structure based on the request identification code; The current token count field, the timestamp field and the validity period field are set in the token bucket data structure, the token generation is performed using an atomic operation instruction, and the time difference between the current time and the timestamp field is multiplied by the token generation rate parameter to calculate the number of newly added tokens; Writing the number of newly added tokens into the current token count field through an atomic operation instruction, updating the current time into the timestamp field, setting the token validity period based on the priority level information, and writing the token expiration time into the validity period field; Write a Lua script for token checking and token acquisition, the Lua script takes the request identification code and the number of tokens required for the request as input parameters, first verifies the validity of the validity period field inside the script, and then checks whether the current token count field meets the number of tokens required for the request; The Lua script is called to perform a combined operation. When the validity period field has not expired and the current token count field is greater than the number of tokens required for the request, the number of tokens required for the request is deducted from the current token count field through an atomic operation, and a token acquisition success flag is returned. The processing rate of the interface request is controlled according to the token acquisition success flag.
6. The method according to claim 1, characterized in that When the system detects performance fluctuations, the real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to achieve adaptive hierarchical scheduling of distributed system traffic, including: Collecting performance data of the distributed system, performing statistical analysis on the performance data based on a sliding time window to obtain performance statistical features; constructing a performance fluctuation detection model based on the performance statistical features, wherein the performance fluctuation detection model calculates a performance fluctuation score through a response time standard deviation, a success rate change rate, a resource usage fluctuation degree, and a queue backlog growth rate; when the performance fluctuation score exceeds a preset threshold, the performance statistical features are input into the dual Q network structure as a system state; Combining the system state with the configuration parameters of the current intelligent hierarchical scheduling strategy to form a state-action pair, calculating an immediate reward value based on the performance fluctuation score, and constructing the system state, the configuration parameters and the immediate reward value into a state transition sequence; assigning a priority weight to the state transition sequence according to the performance fluctuation score, storing the state transition sequence and its priority weight into a priority experience replay pool, and sampling training data from the priority experience replay pool according to the priority weight; Inputting the training data into the dual Q network for online learning, calculating the Q value of the target state through the target Q network, updating the online Q network parameters using the Q value of the target state and the instant reward value, and synchronously updating the target Q network parameters using an exponential sliding average method; Evaluate the current intelligent hierarchical scheduling strategy based on the updated online Q network parameters, generate optimized scheduling strategy parameters, wherein the optimized scheduling strategy parameters include priority allocation weights, resource quota ratios, and queue processing parameters, and introduce exploration noise into the optimized scheduling strategy parameters; Applying the optimized scheduling policy parameters to the distributed system, collecting performance data after application, calculating a new performance fluctuation score, and triggering a policy rollback mechanism when the new performance fluctuation score is higher than the original performance fluctuation score; The changing trend of the performance data is continuously monitored, the priority weight is dynamically adjusted according to the new performance fluctuation score, and the intelligent hierarchical scheduling strategy is optimized through continuous iteration to achieve adaptive scheduling control of distributed system traffic.
7. A distributed system traffic intelligent hierarchical scheduling system based on reinforcement learning, used to implement the method described in any one of claims 1 to 6, characterized in that: include: The first unit is used to construct a two-layer neural network structure to obtain the traffic characteristics of the distributed system. The first layer of the two-layer neural network structure uses a long short-term memory network to collect interface call data in the distributed system in real time, record the concurrent request volume, system throughput, and resource utilization rate. The second layer of the two-layer neural network structure uses a temporal convolutional network to extract the time dimension characteristics of the interface call data and generate a multidimensional state vector for reinforcement learning. The multidimensional state vector is filtered through an attention mechanism network to obtain the core traffic characteristics; The second unit is used to input the core traffic characteristics into a reinforcement learning model, wherein the reinforcement learning model adopts a dual Q network structure, wherein the first Q network evaluates the interface priority allocation strategy based on the current system state, and the second Q network evaluates the resource allocation strategy based on the future system benefit, stores historical decision data through an experience replay mechanism, calculates the objective function value of the dual Q network structure based on a temporal difference algorithm, optimizes the objective function value using a deep deterministic policy gradient method, and outputs an intelligent hierarchical scheduling strategy, wherein the intelligent hierarchical scheduling strategy includes an interface priority parameter, a processing queue configuration parameter, and a resource allocation ratio parameter; The third unit is used to establish a multi-level flow control mechanism based on the intelligent hierarchical scheduling strategy, and rate limit the interface requests of different priorities through the distributed token bucket algorithm. The distributed token bucket algorithm dynamically adjusts the token generation rate according to the interface priority parameter, allocates the interface request to the corresponding priority queue according to the processing queue configuration parameter, and regulates the processing resources according to the resource allocation ratio parameter. When the system detects performance fluctuations, the real-time monitoring data is fed back to the dual Q network structure for online learning, and the intelligent hierarchical scheduling strategy is continuously optimized to realize adaptive hierarchical scheduling of distributed system traffic.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Adaptive event priority scheduling method and system
CN120743481A
Adaptive event priority scheduling method and system
CN120743481B
Multistage security isolation system and method for data library gateway
CN120956518A
Strategy optimization method and device and storage medium
CN121094170A
End-side large model inference method and system for low-altitude robots
CN122572693A