Training method and system of scheduling model, scheduling method and system, storage medium
By using a deep reinforcement learning-based scheduling model, leveraging partially observable Markov decision processes and training sample sets, and combining critic and actor network structures, the scheduling problem under resource consumption constraints is solved, achieving efficient scheduling that maximizes throughput in complex environments.
Patent Information
- Application Number
- CN202110512654.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-11
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2041-05-11
AI Technical Summary
In scheduling operations, maximizing throughput while meeting resource consumption constraints is a pressing issue. In particular, under the influence of factors such as the distribution of arrival flows of scheduled objects, the randomness of service channel status changes, and user resource competition caused by multiple users sharing the system resource pool, existing technologies struggle to achieve efficient scheduling.
A deep reinforcement learning-based scheduling model is adopted. Through a partially observable Markov decision process, the scheduling model is trained using a training sample set. The parameters of the scheduling model are updated by combining critic network and actor network structures to take into account environmental information in continuous time periods, maximize throughput, and meet resource consumption constraints.
It maximizes throughput under resource consumption constraints, improves the training accuracy and execution efficiency of the scheduling model, and enables efficient scheduling in complex environments.
Smart Images

Figure CN113240003B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of resource allocation technology, specifically to a training method and system for a scheduling model, a scheduling method and system, and a storage medium. Background Technology
[0002] In applications involving resource allocation, how to perform scheduling to achieve optimal resource allocation has always been a concern. However, due to factors such as the distribution of arrival flows of scheduled objects, the randomness of service channel state changes serving scheduled objects, competition for resources among different users due to multiple users sharing the system resource pool, and resource consumption limitations, maximizing throughput while meeting resource consumption constraints is a problem that urgently needs to be solved in scheduling operations. Summary of the Invention
[0003] Therefore, the purpose of this application is to provide a training method and system for a scheduling model, a scheduling method and system, and a storage medium to solve the technical problem of maximizing throughput in scheduling operations while satisfying resource consumption.
[0004] To achieve the above and other related objectives, the first aspect of this application discloses a method for training a scheduling model, comprising the following steps: obtaining a training sample set within at least one preset duration, the training sample set including training samples within each preset duration, the training samples including scheduling data in each time slot within the preset duration, the scheduling data including environmental information, resource allocation amounts of scheduling objects obtained by running the scheduling model, and throughput obtained after performing scheduling operations based on the resource allocation amounts; and training the scheduling model based on the training sample set to update the parameters of the scheduling model.
[0005] The second aspect of this application discloses a training system for a scheduling model, comprising: an acquisition module for acquiring a training sample set within at least a preset time period, the training sample set including training samples within each preset time period, the training samples including scheduling data in each time slot within the preset time period, the scheduling data including environmental information, resource allocation amounts of scheduling objects obtained by running the scheduling model, and throughput obtained after performing scheduling operations based on the resource allocation amounts; and a training module for training the scheduling model based on the training sample set to update the parameters of the scheduling model.
[0006] The third aspect disclosed in this application provides a scheduling method, comprising the following steps: receiving a scheduling object; obtaining the resource allocation amount of the scheduling object using a scheduling model trained by the above-described training method; and performing a scheduling operation on the scheduling object according to the resource allocation amount.
[0007] The fourth aspect disclosed in this application provides a scheduling system, comprising: an input unit for receiving a scheduling object; a scheduling model trained according to the above-described training method for obtaining the resource allocation amount of the scheduling object; and a scheduling operation performed on the scheduling object according to the resource allocation amount.
[0008] The fifth aspect disclosed in this application provides an electronic device, comprising: at least one memory for storing at least one program; and at least one processor connected to the at least one memory for executing and implementing the training method of the scheduling model described above, or executing and implementing the scheduling method described above, when running the at least one program.
[0009] The sixth aspect disclosed in this application provides a cloud server system, comprising: at least one storage device for storing at least one program; and at least one processing device connected to the storage device for executing and implementing the training method of the scheduling model described above, or executing and implementing the scheduling method described above, when running the at least one program.
[0010] The seventh aspect disclosed in this application provides a computer-readable storage medium storing at least one program, which, when executed by a processor, executes and implements the training method of the scheduling model described above, or executes and implements the scheduling method described above.
[0011] In summary, this application provides a training method and system for a scheduling model, a scheduling method and system, and a storage medium. The scheduling model is trained using a training sample set within a preset time period to update the parameters of the scheduling model, so that when using the scheduling model to perform scheduling operations, the goal of maximizing throughput while satisfying resource consumption is achieved.
[0012] Other aspects and advantages of this application will readily be apparent to those skilled in the art from the detailed description below. Only exemplary embodiments of this application are shown and described in the following detailed description. As will be appreciated by those skilled in the art, the content of this application enables them to make modifications to the disclosed specific embodiments without departing from the spirit and scope of the invention to which this application pertains. Accordingly, the descriptions in the accompanying drawings and specification of this application are merely exemplary and not restrictive. Attached Figure Description
[0013] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention can be better understood by referring to the exemplary embodiments and drawings described in detail below. A brief description of the drawings is as follows:
[0014] Figure 1 The diagram shown is a flowchart of the training method for the scheduling model of this application in one embodiment.
[0015] Figure 2 The diagram shown is a flowchart of step S110 in one embodiment of the training method of the scheduling model of this application.
[0016] Figure 3 The diagram shows a critic network structure used in the scheduling model of this application.
[0017] Figure 4 The diagram shown illustrates the actor network structure used in the scheduling model of this application.
[0018] Figure 5 The diagram shown is a flowchart of the training method for the scheduling model of this application in another embodiment.
[0019] Figure 6 The diagram shown is a structural schematic of the training system for the scheduling model of this application in one embodiment.
[0020] Figure 7 The diagram shown is a flowchart of the scheduling method of this application in one embodiment.
[0021] Figure 8 The diagram shown is a structural schematic of the scheduling system of this application in one embodiment.
[0022] Figure 9 The diagram shown is a structural schematic of the electronic device of this application in one embodiment.
[0023] Figure 10 The diagram shown is a structural schematic of the cloud server system of this application in one embodiment. Detailed Implementation
[0024] The following specific embodiments illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification.
[0025] In the following description, reference is made to the accompanying drawings, which illustrate several embodiments of the present application. It should be understood that other embodiments may also be used, and changes in module or unit composition, electrical and operational aspects may be made without departing from the spirit and scope of the disclosure of this application. The following detailed description should not be considered limiting, and the scope of the embodiments of this application is defined only by the claims of the published patents. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application.
[0026] While the terms first, second, etc., are used in some instances herein to describe various elements, information, or parameters, these elements or parameters should not be limited by these terms. These terms are used only to distinguish one element or parameter from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the various described embodiments. Both the first element and the second element describe a single element, but they are not the same element unless the context otherwise explicitly indicates otherwise. Depending on the context, the word "if," as used herein, may be interpreted as "when" or "when...".
[0027] Furthermore, as used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms “comprising,” “including,” indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not preclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms “or” and “and / or” as used herein are to be interpreted inclusively, or mean any one or any combination thereof. Thus, “A, B, or C” or “A, B, and / or C” means “any one of: A; B; C; A and B; A and C; B and C; A, B, and C.” Exceptions to this definition occur only when combinations of elements, functions, steps, or operations are inherently mutually exclusive in some way.
[0028] In applications involving resource allocation, the rational scheduling of resources is a crucial issue. Scheduling problems are constrained by various factors, such as the distribution of arrival flows of scheduled objects, the randomness of service channel state changes serving these objects, competition for resources among different users due to shared system resource pools, resource consumption limitations, and latency constraints. These factors make scheduling problems difficult to solve simply. For example, in wireless communication, mobile users receive data packets from relays via wireless channels, and transmission antennas are responsible for transmitting these packets to users. Since transmission antennas are limited by their transmit power in practice, scheduling data packets to send as many packets as possible to users while meeting these power limitations is a significant challenge.
[0029] Currently, with the continuous development of machine learning, reinforcement learning, unlike supervised and unsupervised learning in traditional machine learning, is widely used because it can learn through the interaction between an agent and its environment. Specifically, in the interaction with the environment, the agent continuously learns the dynamics of the environment based on the different rewards obtained from different actions, in order to adapt to the environment and maximize cumulative rewards. Subsequently, deep neural networks were introduced on the basis of reinforcement learning, resulting in deep reinforcement learning. Due to the advantage of deep reinforcement learning in adaptive learning under specific environments, it can be used to solve user scheduling problems, aiming to maximize throughput while meeting resource consumption requirements in scheduling operations.
[0030] In view of this, this application provides a training method for a scheduling model, wherein the function of the scheduling model is to allocate resources to scheduling objects and then perform scheduling operations according to the allocated resource amount. In some embodiments, the scheduling model is a scheduling device or scheduler that allocates resources to scheduling objects by running a scheduling algorithm based on deep reinforcement learning. In other embodiments, the neural network in the scheduling model uses a neural network with memory rather than a simple fully connected network. For example, the neural network of the scheduling model is based on a recurrent neural network (RNN) structure, two typical variants of which include the Long Short-Term Memory (LSTM) structure and the Gated Recurrent Unit (GRU) structure.
[0031] Generally, Markov Decision Processes (MDPs) are used to represent MDPs in reinforcement learning. Specifically, a Markov Decision Process can be represented using quadruples. It means that, among them, Represents the set of all actions of the intelligent agent; Represents the set of states of the environment; Represents the payoff function, i.e. It is a constant or a random number; The state transition matrix represents the environment, i.e. binary pair This is also known as the model of this Markov decision process. γ represents the discount factor. At each time t, the agent observes the environmental state s. t And take action a t The environment then transitions to the next state s. t+1 And the agent gains an instantaneous benefit r t .
[0032] However, for scheduling problems, the Markov property of the environment is often difficult to satisfy, meaning that the agent cannot fully observe the environmental state s at time t. t Only its subset o can be observed t Therefore, for the scheduling problem, this application introduces a Partially Observable Markov Decision Process (POMDP) into reinforcement learning. Specifically, in this application, the Markov decision process can be represented by a quintuple. It means that, among them, Represents the set of observed variables. Represents the set of states of the environment. Represents the set of all actions of the intelligent agent; Represents the payoff function, i.e. It is a constant or a random number; The state transition matrix represents the environment, i.e. binary pair This is also referred to as the model of this partially observable Markov decision process. γ represents the discount factor. In the partially observable Markov decision process for the scheduling problem in this application, the agent's goal is to learn a deterministic policy π(s) that results in a long-term discount gain. Maximum, where T represents a preset duration.
[0033] In other words, the problem to be solved by the scheduling model in this application can be characterized by a partially observable Markov decision process and solved using an algorithm based on deep reinforcement learning. Specifically, the agent's action a... t This represents the resource allocation amount of the scheduling object obtained by running the scheduling model; the environmental variable o observed by the agent. t This represents the environmental information observed by the scheduling model. The environmental information refers to information directly observed by the scheduling model. Under fully observable conditions, this environmental information can be characterized by the environmental state, denoted by s. t This indicates that, under partially observable conditions, the environmental information is derived from environmental variables o. t Representation; Agent's gain r t This represents the throughput obtained by the scheduling model after performing scheduling operations.
[0034] Considering the partial observability of the environment in scheduling operations, if the scheduling model is trained solely based on environmental information at a certain moment, the resource allocation of the scheduled objects at that moment obtained through the running scheduling model based on that environmental information, and the throughput achieved after performing scheduling operations based on the resource allocation, the environmental information obtained at each moment cannot fully reflect the current environment. Therefore, the scheduling model trained based on the above information is difficult to achieve efficient scheduling. In view of this, in order to enable the deep reinforcement learning-based scheduling model to perform efficient scheduling while satisfying resource consumption and maximizing throughput, while fully considering the partial observability of the environment, this application provides a training method for a scheduling model. The scheduling model trained by this method can maximize throughput when performing scheduling operations. The main idea is to train the scheduling model by observing environmental information over a continuous time period, the resource allocation of the scheduled objects at that moment obtained through the running scheduling model based on that environmental information, and the throughput achieved after performing scheduling operations based on the resource allocation.
[0035] Please refer to Figure 1 The figure shows a flowchart of the training method of the scheduling model of this application in one embodiment. As shown in the figure, the training method of the scheduling model includes steps S100 and S110.
[0036] In step S100, a training sample set within at least one preset time period is obtained. The training sample set includes training samples within each preset time period, and each training sample includes scheduling data for each time slot within the preset time period. The scheduling data includes environmental information, resource allocation amounts for scheduling objects obtained through running the scheduling model, and throughput obtained after performing scheduling operations based on the resource allocation amounts.
[0037] The training process of the scheduling model can be divided into multiple preset durations based on the interaction between the scheduling model and the environment. The training method of the scheduling model is executed for each preset duration, and this process is repeated until the entire training process is completed. The preset durations can be set based on the unit quantity used by the scheduling model to make scheduling decisions, or they can be set by those skilled in the art in which the scheduling model is applied, based on experience. For a given scheduling problem, the length of each preset duration must remain consistent throughout the entire training process of the scheduling model. In some embodiments, the multiple preset durations can also be referred to as multiple segments, and the training process of the scheduling model is divided into multiple segments, the length of each segment corresponding to the length of a preset duration.
[0038] A time slot is a division of a preset duration, corresponding to the unit of time within which the scheduling model makes scheduling decisions. For example, if the scheduling model makes a scheduling decision at every moment, the length of the preset duration can be set to T, meaning it includes T moments, and the time slot is each moment. Alternatively, if the scheduling model makes a scheduling decision every 24 hours (one day), the length of the preset duration can be set to N days, and the time slot is each day (24 hours).
[0039] For ease of description, the following example assumes that the scheduling model makes a scheduling decision at every time t. The preset duration is set to T, meaning it includes T time slots, and each time slot represents a time slot. Correspondingly, the training samples are a collection of scheduling data for each time slot within a preset duration. The training sample set is a collection of training samples for multiple preset durations, including multiple consecutive scheduling data sets of duration T. In other words, the training sample set is a collection of consecutive scheduling data sets of duration T, where each T-time-long data set contains scheduling data for T time slots.
[0040] Scheduling data includes environmental information, resource allocation amounts for scheduled objects obtained through the running scheduling model, and throughput obtained by the scheduling model after performing scheduling operations based on the resource allocation amounts. Environmental information refers to the information observed when the scheduling model interacts with the environment. Generally, environmental information includes the state of the scheduled object and the state of the service channel, where the service channel is the channel that allocates resources to the scheduled object. In some embodiments, the state of the scheduled object includes information for distinguishing one scheduled object from another, such as information indicating the queue the scheduled object is in, information about which user the scheduled object is to be scheduled to, and constraint information of the scheduled object itself, such as the delay status information of the scheduled object. The state of the service channel includes information about the service channel that allocates resources to the scheduled object. For example, in the field of wireless communication scheduling, the state of the service channel includes information indicating the state of the wireless channel between the transmitting antenna and the destination user; in the field of video streaming scheduling, the state of the service channel includes information indicating the state of the communication link between the video buffer and the user terminal device; in the field of order delivery scheduling, the state of the service channel includes information indicating the state of the physical road between the delivery station and the delivery destination.
[0041] The resource allocation for a scheduled object is generated by the scheduling model based on observed environmental information. Throughput is the instantaneous gain obtained by the scheduling model after performing scheduling operations based on the acquired resource allocation. In this invention, instantaneous gain refers to instantaneous weighted throughput. The throughput is related to the state of the scheduled object and the number of successfully scheduled objects.
[0042] In some embodiments, step S100, obtaining a training sample set within at least a preset duration, may include: obtaining the resource allocation and instantaneous weighted throughput of the scheduled object in the corresponding time slot by running the scheduling model based on the environmental information in each time slot, and caching the training samples. In a specific example, firstly, the scheduling model and replay cache may be randomly initialized. The replay cache may be pre-configured in the scheduling model or may communicate with the scheduling model to provide the dataset stored in the replay cache to the scheduling model. Then, for a preset duration of T, the scheduling model obtains the resource allocation and instantaneous weighted throughput of the scheduled object at the first time slot based on the environmental information at the first time slot, and the same applies to the second time slot based on the environmental information at the second time slot, and so on until the resource allocation and instantaneous weighted throughput at time T are obtained. The scheduling data at each time slot are then stored in the replay cache as a sequence, serving as a training sample.
[0043] When the training process of the scheduling model is divided into multiple preset durations, the scheduling data also includes an indicator d to represent whether the current preset duration has ended. t In some embodiments, the end of the current preset duration can be indicated in the scheduling data as 1 or 0. For example, if the current duration ends after the resource allocation of the scheduled object is obtained through the scheduling model based on environmental information at the current moment, then 1, i.e., d, is displayed in the scheduling data. t =1, otherwise display 0, i.e., d t =0. In one embodiment, assuming the preset duration is T, the end of the current duration can be indicated by comparing the current time t with the preset duration T. If t > T, then d t =1, otherwise d t =0.
[0044] In actual operation, after obtaining the scheduling data at each moment within a preset time period, it is cached in the replay buffer as a sequence as a training sample. The above steps of obtaining training samples are repeated to obtain a training sample set consisting of multiple training samples. The one or more training samples serve as the sample input for training the scheduling model. Here, the cached training samples are data sequences including environmental information, resource allocation, weighted throughput, and indicators representing whether the current preset time period has ended.
[0045] In step S110, a scheduling model is trained based on the training sample set to update the parameters of the scheduling model.
[0046] In this system, training samples are used as input to the scheduling model, which is then trained using deep reinforcement learning based on partial Markov decisions to update its parameters. For example, the parameters of the scheduling model are the weights of the neural network structure. In some embodiments, the optimal parameters can be computed, i.e., an analytical method can be used to update the scheduling model's parameters; this method is more suitable for scheduling problems in simple environments. In other embodiments, gradient descent can be used to update the scheduling model's parameters, enabling the scheduling model to achieve efficient scheduling.
[0047] The training method of the scheduling model in this application trains the scheduling model with a training sample set. While fully considering the partial observability of the environment, it can mine the unobservable hidden information in the environment during the training process based on the information of the agent's continuous interaction with the environment over a period of time. This enables the scheduling model to maximize throughput while satisfying resource consumption constraints when performing scheduling operations.
[0048] When training a scheduling model based on a training sample set, in some embodiments, the scheduling model can be trained using all current training samples stored in the replay cache. In other embodiments, a subset of training samples can be sampled from the replay cache to train the scheduling model, thereby reducing computational complexity.
[0049] Since reinforcement learning involves continuous learning during the interaction between the agent and the environment, to improve the training accuracy of the scheduling model, this application allows the training of the scheduling model based on a training sample set to begin after the initial run of the scheduling model to obtain multiple training samples within a preset time period. In some embodiments, during the training of the scheduling model, the scheduling model and replay cache are first initialized, and then multiple training samples are obtained according to the steps described above and stored in the replay cache. After multiple training samples are stored in the replay cache, the training and learning process of the scheduling model begins. The scheduling model interacts with the environment to obtain new training samples on one hand, and updates its parameters using the cached training samples on the other, repeating this process cyclically. Based on this, please refer to... Figure 2 The figure shows a flowchart of step S110 in the training method of the scheduling model of this application in one embodiment. As shown in the figure, step S110 includes steps S111 to S113.
[0050] In step S111, the training sample set is sampled.
[0051] During the training process, the training sample data stored in the replay cache of the scheduling model continuously increases. In order to introduce updated training sample data during training, while reducing computational complexity and improving training accuracy, the stored training sample set is randomly sampled to obtain training samples for training the scheduling model.
[0052] In step S112, the parameters of the scheduling model are updated using gradient descent based on the sampled data.
[0053] Considering the partial observability of the environment in scheduling problems, in order to learn implicit information that is not observable in the environment based on environmental variables, gradient descent is used in some embodiments to update the parameters of the scheduling model. Gradient descent can be achieved by using operators such as softmax or max operators to calculate the gradient.
[0054] In step S113, steps S111 and S112 are repeated until the preset number of updates is met.
[0055] In a training operation within a preset duration, the number of updates to the scheduling model can be set. Then, based on the set number of updates, the process of sampling from the training sample set and updating the scheduling model's parameters using gradient descent based on the sampled data is repeated until the set number of updates is reached. The number of updates can be set by those skilled in the art in which the scheduling model is applied, based on experience. In some embodiments, the number of updates is set to two; after two updates, the next preset duration begins. It should be noted that during the acquisition of training samples within the next preset duration, the scheduling model's parameters are updated. The updated scheduling model obtains the resource allocation and throughput of the scheduled objects based on environmental information, and this process is repeated.
[0056] In the case of a deep neural network structure for the scheduling model, to improve the performance of the scheduling algorithm, this application's scheduling model includes two types of network structures: a critic network structure and an actor network structure. The scheduling model combines these two types of network structures to obtain neural networks with different structures. For ease of description, the instance of the critic network is denoted as Q, and all its weight parameters are denoted as θ. The instance of the actor network is denoted as π, and all its weight parameters are denoted as...
[0057] In some embodiments, the critic network structure includes a fully connected and long short-term memory (LSTM) dual-branch structure. The actor network structure includes a fully connected and LSM dual-branch structure.
[0058] Please see Figure 3 The figure shows a schematic diagram of the critic network structure used in the scheduling model of this application. As shown, the input of the critic network includes: environmental information at the current time t, which in some examples is the environmental variable o. t The resource allocation amount 'a' of the scheduling object in the scheduling model at the current time t. t ; and the resource allocation amount (a) of the scheduling object by the scheduling model at the previous time step (t-1).t-1 In some embodiments, the concatenated vector (o) is used. t ,a t ), (o t ,a t-1 The input is in the form of (), and then fed into two branches: the first fully connected (FC) branch and the first long short-term memory (LSTM) branch. In the first fully connected branch, the input (o) is... t ,a t The input first passes through a fully connected layer, then through a rectified linear unit activation function (ReLU), and finally outputs. In the first long short-term memory branch, the input (o...)... t ,a t-1 The output first passes through a second fully connected layer, then a second rectified linear unit activation function (ReLU), and then into the first long short-term memory layer before being output. The outputs of the first fully connected branch and the first long short-term memory branch are then concatenated in a first concatenation layer. The concatenated output first passes through a third fully connected layer, then a third rectified linear unit activation function (ReLU), and then a fourth fully connected layer before outputting the action 'a' to be performed in the current state. t The reward Q(o) obtained later t ,a t ).
[0059] Please see Figure 4 The figure shows a schematic diagram of the actor network structure used in the scheduling model of this application. As shown, the input of the actor network includes: environmental information at the current time t, which in some examples is the environmental variable o. t The resource allocation amount (a) of the scheduling object in the scheduling model at the previous time step (t-1). t-1 In some embodiments, the concatenated vector (o) is used. t ,a t-1 The input is in the form of (), and then fed into two branches: the second fully connected branch and the second long short-term memory branch. In the second fully connected branch, the input (o) is... t ,a t-1 First, it passes through the fifth fully connected layer, then the fifth rectified linear unit (ReLU) activation function, and finally the output. In the second long short-term memory branch, the input (o t ,a t-1First, the data passes through the sixth fully connected layer, then the sixth rectified linear unit activation function, and then the second long short-term memory layer before outputting. The outputs of the second fully connected branch and the second long short-term memory branch are then concatenated through the second splicing layer, passed through a seventh fully connected layer, and finally through the hyperbolic tangent (tanh) activation function to output the resource allocation amount 'a' of the scheduling object at time t. t .
[0060] When the scheduling model includes a critic network structure and an actor network structure, the step of training the scheduling model based on a training sample set includes: training the critic network structure and the actor network structure based on the training sample set. That is, training the weight parameters θ of the critic network structure and the weight parameters of the actor network structure based on the training sample set. Train and update the algorithm. For example, the target value of the critic network can be calculated using the softmax operator to improve algorithm performance. Alternatively, the target value of the critic network can be calculated using the max operator.
[0061] Since the parameter training of the actor network is influenced by the output of the critic network, in order to update the actor network parameters more stably, in some embodiments, the step of training the scheduling model based on the training sample set includes: training the critic network structure based on the training sample set; and training the actor network structure using the critic network structure trained at least once. That is, during the training of the scheduling model based on the training sample set, the parameters of the critic network in the scheduling model are updated sequentially, while the parameters of the actor network in the scheduling model are updated intermittently. The number of intervals is given empirically, and a good interval is generally 2. For example, if the training process of the scheduling model is divided into M durations based on the interaction between the scheduling model and the environment, it can be set to update the critic network parameters and actor network parameters only during an even number of durations. Then, if the training process is divided into six durations, the weight parameters of the critic network and actor network are updated simultaneously only during the second, fourth, and sixth durations when training the scheduling model based on the training sample set.
[0062] The training method of the scheduling model in this application samples the training dataset and uses gradient descent to train the scheduling model, which can asymptotically approach the optimal scheduling strategy and uncover implicit information in the environment, thereby improving the efficiency of the scheduling model in performing scheduling operations and maximizing throughput. Furthermore, the intervalic updates to the actor network improve the stability of the scheduling model training.
[0063] In practical scheduling problems, scheduling operations are constrained by resource consumption. In order to provide a hard guarantee of resource consumption constraints when the scheduling model performs scheduling operations, the training method of the scheduling model in this application further includes the step of determining whether the obtained resource allocation meets the preset constraints. If the constraints are met, the scheduling operation is performed based on the obtained resource allocation; if the constraints are not met, the scheduling operation is not performed based on the obtained resource allocation.
[0064] The constraints can be set by those skilled in the art in which the scheduling model is applied, based on actual environmental constraints. For example, the constraints may include instantaneous resource consumption constraints, average resource consumption constraints, and service channel capacity constraints. In some embodiments, the constraint is an instantaneous resource consumption constraint. In one example, the constraint can be set to an instantaneous resource allocation less than or equal to a fixed value. When the instantaneous resource consumption of the current scheduling object obtained by the scheduling model is less than or equal to the fixed value, the scheduling model performs a scheduling operation based on the resource allocation; when the instantaneous resource consumption of the current scheduling object obtained by the scheduling model is greater than the fixed value, the scheduling model does not perform a scheduling operation. In other embodiments, the constraint satisfies an average resource consumption constraint. In one example, the constraint is set to the average of the resource allocation obtained at the current moment and the cumulative used resource allocation being less than or equal to a preset value. Here, the resource allocation obtained at the current moment refers to the resource allocation obtained by the scheduling model at the current moment, and the cumulative used resource allocation refers to the sum of the resource allocations at all moments prior to the current moment. The preset value is set according to the constraints of the actual scheduling problem; for example, in the field of wireless communication scheduling, the preset value is set based on the average power of the transmitting antenna. In other words, the constraint is set to the average resource consumption being less than or equal to a fixed value. When the average cumulative resource consumption of the scheduled object obtained by the scheduling model is less than or equal to the fixed value, the scheduling model performs a scheduling operation based on the resource allocation; when the average cumulative resource consumption of the scheduled object obtained by the scheduling model is greater than the fixed value, the scheduling model does not perform a scheduling operation. The average cumulative resource consumption is obtained by adding the resource allocation obtained at the current time to the sum of the resource allocations at all previous times and then averaging the results. In a specific example, the scheduling model obtains the resource allocation 'a' of the scheduled object based on the environmental information at time t. t Then, after passing through the internal function f, the action f(a) that satisfies the average resource consumption constraint is output. t The scheduling model executes the task within the learning environment; in other words, it obtains the resource allocation amount 'a' for the scheduled object based on the environmental information at time 't'. t Then, if the average resource consumption constraint is satisfied, i.e. If the scheduling operation fails, the operation will be executed; otherwise, actions that do not meet the average resource consumption constraint will be reset to zero. P refers to the average resources consumed by the scheduling model. avg It refers to the average resource consumption, as shown in formula (1).
[0065]
[0066] Correspondingly, the data constituting the training samples is the output data after optimization by the internal function f. In other words, when the average resource consumption constraint is not met, the resource allocation of the scheduling object and the throughput obtained after performing scheduling operations based on the resource allocation are set to zero in the scheduling data that enters the learning environment and is cached as training samples. This enables the scheduling model to consciously perceive and comply with the constraints during the training process.
[0067] Furthermore, considering the latency issue in scheduling operations, latency constraints can be set for scheduling objects. These latency constraints can be set by those skilled in the art in which the scheduling model is applied, based on actual system requirements. In this case, the training method of the scheduling model of this application further includes the following steps: caching scheduling objects and processing the cached scheduling objects according to the set latency constraints. In some embodiments, the step of processing the cached scheduling objects according to the latency constraints includes serving scheduling objects that meet the latency constraints and discarding scheduling objects that do not meet the latency constraints. In some embodiments, the same latency constraint is set for scheduling objects of the same user. Therefore, the user's scheduling objects must be served within the set latency time after arriving at the cache; otherwise, the scheduling objects are discarded. This allows the scheduling model to maximize throughput while satisfying resource consumption and latency constraints when performing scheduling operations.
[0068] In a specific example, the scheduling model makes a scheduling decision at every time step. Based on this, the training process of the scheduling model is set to include M preset durations, also known as M segments, where each preset duration includes T time steps, i.e., each segment has a length of T. Furthermore, the scheduling model is updated twice for each segment; for the entire training process, the actor network parameters are updated when the number of segments is a multiple of q, in this example q = 2. (See also...) Figure 5 The figure shows a flowchart of the training method for the scheduling model of this application in another embodiment. As shown in the figure, the training method for the scheduling model includes steps S200 to S250.
[0069] Training of the scheduling model begins based on the above settings. In step S200, the scheduling model and replay buffer are randomly initialized.
[0070] In step S210, it is determined whether the current number of segments is greater than the total number of segments M. If it is greater than M, the training process ends; otherwise, step S220 is executed. In this example, the training starts from the number of segments 1.
[0071] In step S220, it is determined whether the current time t is greater than the segment length T. If it is greater than T, step S230 is executed; otherwise, steps S221 and S222 are executed. In this example, it starts from t=1.
[0072] In step S221, the scheduling model is run to obtain scheduling data and perform scheduling operations based on the observed environmental information. In some embodiments, in order to ensure that the scheduling operations meet the constraints, the scheduling model needs to be optimized by a constraint adaptation module before executing the scheduling operations based on the environmental information. That is, the module determines whether to execute the scheduling operation and obtains the constraint-processed scheduling data based on the constraints.
[0073] In step S222, the next time is obtained based on the current time, i.e., t = t + 1, and the process returns to step S220.
[0074] In step S230, the scheduling data obtained in step S221 is stored in the replay buffer as training samples in the form of a sequence.
[0075] In step S240, it is determined whether the current update count i is greater than the preset update count 2. If yes, step S250 is executed; otherwise, steps S241 to S246 are executed. In this example, it starts from i = 1.
[0076] In step S241, a portion of the training samples are sampled from the training sample set cached in step S230 as training sample input.
[0077] In step S242, the softmax operator is used to calculate the critic network target value.
[0078] In step S243, the critic network parameters are updated.
[0079] In step S244, it is determined whether the current number of segments is a multiple of 2. If so, steps S245 and S246 are executed; otherwise, step S246 is executed directly.
[0080] In step S245, the actor network parameters are updated.
[0081] In step S246, the next update count is obtained based on the current update count, i.e., i = i + 1, and the process returns to step S240.
[0082] In step S250, the next segment number is obtained based on the current segment number, that is, segment number = current segment number + 1, and the process returns to step S210.
[0083] In a specific example, the scheduling model of this application employs a neural network based on a reinforcement learning scheduling algorithm. During the training phase of the scheduling model, the scheduling algorithm guides the agent (corresponding to the scheduling model in this application) to interact with the environment. In each interaction segment (corresponding to the preset duration in this application) within each interaction time slot, the agent continuously explores the environment and maximizes its cumulative gain (corresponding to the throughput in this application) by utilizing the characteristics of the environment. After a large number of interaction segments, the training phase of the agent ends, resulting in a deep neural network that can be practically deployed.
[0084] For example, the scheduling model of this application can perform scheduling operations for N users, taking the scheduling model making a scheduling decision at every time t as an example. In some embodiments, the scheduling model also includes a caching module. The cache module This is used for caching scheduling objects. In scheduling operations, time is divided into discrete moments t∈{0,1,2,...}, and scheduling objects are served in discrete packets. After scheduling objects arrive at the scheduling model, these scheduling objects are first cached in the caching module. In the process, the data is then selected by the scheduling algorithm at an appropriate time to be served on the service channel, with the aim of sending it to the destination user of the scheduled object. (Cache module) There are N unordered queues, each corresponding to one of the N users. At time t, for user i, queue q... i The number of arriving scheduled objects is A. i (t), with weight β i Its corresponding service channel status is c i (t). Each scheduling object belonging to user i has the same latency constraint τ. i In other words, the user's scheduling object needs to be processed by the cache module within τ after it arrives. i Served within a specified timeframe, otherwise it will time out and be discarded. Therefore, the caching module... The state of each scheduled object can be uniquely determined by the user index and the remaining delay (i, τ). Specifically, for user i's queue q... i Its state at time t is in The representation of queue q i The number of users with a remaining latency of τ.
[0085] At each time t, the scheduling model must make a scheduling decision. For user i's queue q... i Its scheduling decision at time t is in The representation of queue q iThe remaining latency of a scheduled object is τ, representing the service resources available. Whether a scheduled object can be successfully sent depends on its service resources and the channel status. Higher service resources increase the probability of a scheduled object successfully reaching its destination user; failed scheduled objects will be discarded and not retransmitted. At time t, the number of scheduled objects successfully sent by user i is denoted as D. i (t). Then, from time 0 to time t, the weighted throughput of the scheduling model is: Where, β i This is the weight of user i's throughput. The average resource consumed by the scheduling model is...
[0086] Based on this, regarding the training process of the scheduling model, given a segment length T, for a time t∈{1,…,T} within a segment: the agent's action is... This represents the resource allocation for each state scheduling object in the cache; the environmental variables observed by the agent are... This represents the state of the scheduling objects in each queue in the cache, as well as the state of the service channels, as observed by the agent. The observed environment variables are partial observations of the complete environment variables; d t Indicates whether the current segment has ended: if the agent takes action a t After that, this segment ends, then d t =1, otherwise d t =0; the agent's instantaneous gain is This represents the instantaneous weighted throughput. Additionally, for any given time, the state transition probability matrix... It remains unknown. Let γ be 1, meaning the agent's optimization objective is...
[0087] Furthermore, based on the aforementioned critic network and actor network structures, in this example, the number of neurons in the fully connected layers and long short-term memory layers of both network templates is set to L. The final result is eight deep neural network instances generated as follows: two instances, actor network π1 and actor network π2, generated from the actor network template; two instances, critic network Q1 and critic network Q2, generated from the critic network template; and two actor network instances and two target network copies of the critic network instances: the target actor network. Target Actor Network Target Critics Network Target Critics Network
[0088] During the training of the scheduling model, the inputs to the scheduling algorithm include: number of training segments M, segment length T, network width L, action noise variance σ, number of training samples K, number of sampling noise samples J, sampling noise variance σ′, noise margin c, inverse temperature β, and sampling distribution. The policy delay is q, and the target network update rate is τ. The output obtained from the scheduling algorithm includes: the trained actor network π1, actor network π2, critic network Q1, and critic network Q2. Now, combined with... Figure 5 The specific process of the scheduling algorithm is shown.
[0089] First, initialize the actor network, critic network, and corresponding target actor and target critic networks. Specifically, use random weight parameters. Initialize actor networks π1 and π2; initialize critic networks Q1 and Q2 with random weight parameters θ1 and θ2; initialize critic networks Q1 and Q2 with random weight parameters θ2. Initialize the target actor network Target Actor Network With random weight parameters Initialize the target critic network Target Critics Network
[0090] Secondly, initialize the replay cache.
[0091] Then, repeat the following loop for fragment numbers from 1 to M:
[0092] a) Repeat the following cycle from time t to T:
[0093] i) The agent observes environmental variables o t ;
[0094] ii) Let h be the historical observation at time t. t ←(h t-1, a t-1, o t );
[0095] iii) Let action a t ←g(π1, π2)+∈, where It follows a normal distribution with parameters (0, σ).
[0096]
[0097] iv) The agent executes the action f(a) optimized by the constraint adaptation module. t ) into the learning environment;
[0098] b) The sequence (o1, a1, r1, d1, ..., o Ta T r T d T Store in cache
[0099] c) Repeat the following loop for i from 1 to 2:
[0100] i) Randomly retrieve from cache Sampling K sequence samples {(o1, a1, r1, ..., o T a T r T )};
[0101] ii) Repeat the following cycle from 1 to T for time t:
[0102] 1) Randomly sample J noise samples ∈′, where That is, it follows a normal distribution with parameters (0, σ′);
[0103] 2) in
[0104]
[0105] 3)
[0106] 4) Calculate according to the following formula definition
[0107]
[0108] Where β is the inverse temperature, It is a sampling distribution. (The result is...)
[0109]
[0110] iii) According to Bellman loss:
[0111]
[0112] Update the critic network weight parameters θ using gradient descent i
[0113] iv) If the number of fragments is a multiple of q, then
[0114] 1) Based on the gradient of the iterative policy:
[0115]
[0116] Updating actor network weight parameters using gradient descent
[0117] 2) Update the target network weight parameters according to the following formula:
[0118]
[0119]
[0120] The scheduling model trained using the above training method can strictly satisfy the average resource consumption constraint. Where P avg Under resource and latency constraints, efficient scheduling is performed to achieve near-maximum weighted average throughput.
[0121] This application also provides a training system for a scheduling model; please refer to [link / reference]. Figure 6 The figure shows a schematic diagram of the training system for the scheduling model of this application in one embodiment. As shown in the figure, the training system for the scheduling model includes an acquisition module 10 and a training module 11.
[0122] The acquisition module 10 is used to acquire a training sample set within at least one preset time period. The training sample set includes training samples within each preset time period. The training samples include scheduling data in each time slot within the preset time period. The scheduling data includes environmental information, the resource allocation amount of the scheduling object obtained by running the scheduling model, and the throughput obtained after performing scheduling operations based on the resource allocation amount.
[0123] The interaction between the scheduling model and the environment during training can be divided into multiple preset durations. The training method of the scheduling model described in this application is executed for each preset duration, and this process is repeated until the entire training process is completed. The preset durations can be set based on the unit quantity in which the scheduling model makes scheduling decisions, or they can be set by those skilled in the art in which the scheduling model is applied, based on experience. In this application, each preset duration has the same length throughout the entire training process of the scheduling model. In some embodiments, the multiple preset durations can also be referred to as multiple segments, and the training process of the scheduling model is divided into multiple segments, the length of each segment corresponding to the length of the preset duration.
[0124] A time slot is a division of a preset duration, corresponding to the unit of time within which the scheduling model makes scheduling decisions. For example, if the scheduling model makes a scheduling decision at every moment, the length of the preset duration can be set to T, meaning it includes T moments, and the time slot is each moment. Alternatively, if the scheduling model makes a scheduling decision every 24 hours (one day), the length of the preset duration can be set to N days, and the time slot is each day (24 hours).
[0125] Scheduling data includes environmental information, resource allocation amounts for scheduled objects obtained through the running scheduling model, and throughput obtained by the scheduling model after performing scheduling operations based on the resource allocation amounts. Environmental information refers to the information observed when the scheduling model interacts with the environment. Generally, environmental information includes the state of the scheduled object and the state of the service channel, where the service channel is the channel that allocates resources to the scheduled object. In some embodiments, the state of the scheduled object includes information for distinguishing one scheduled object from another, such as information indicating the queue the scheduled object is in, information about which user the scheduled object is to be scheduled to, and constraint information of the scheduled object itself, such as the delay status information of the scheduled object. The state of the service channel includes information about the service channel that allocates resources to the scheduled object. For example, in the field of wireless communication scheduling, the state of the service channel includes information indicating the state of the wireless channel between the transmitting antenna and the destination user; in the field of video streaming scheduling, the state of the service channel includes information indicating the state of the communication link between the video buffer and the user terminal device; in the field of order delivery scheduling, the state of the service channel includes information indicating the state of the physical road between the delivery station and the delivery destination.
[0126] The resource allocation for a scheduled object is generated by the scheduling model based on observed environmental information. Throughput is the instantaneous gain obtained by the scheduling model after performing scheduling operations based on the acquired resource allocation. In this invention, instantaneous gain refers to instantaneous weighted throughput. The throughput is related to the state of the scheduled object and the number of successfully scheduled objects.
[0127] In some embodiments, the acquisition module includes: an acquisition unit, configured to obtain the resource allocation amount and instantaneous weighted throughput of the scheduling object in the corresponding time slot by running the scheduling model based on the environmental information in each time slot; and a caching unit, configured to cache the training samples.
[0128] In a specific example, firstly, the scheduling model and replay buffer can be randomly initialized. The replay buffer can be pre-configured within the scheduling model or can communicate with the scheduling model to provide it with the dataset stored in the replay buffer. Then, for a preset duration T, the scheduling model obtains the resource allocation and instantaneous weighted throughput of the scheduled object at the first time step based on the environmental information. Similarly, it obtains the resource allocation and instantaneous weighted throughput at the second time step based on the environmental information, and so on, until the resource allocation and instantaneous weighted throughput at time T are obtained. The scheduling data for each time step are then stored sequentially in the replay buffer as a training sample.
[0129] When the training process of the scheduling model is divided into multiple preset durations, the scheduling data also includes an indicator d to represent whether the current preset duration has ended. t In some embodiments, the end of the current preset duration can be indicated in the scheduling data as 1 or 0. For example, if the current duration ends after the resource allocation of the scheduled object is obtained through the scheduling model based on environmental information at the current moment, then 1, i.e., d, is displayed in the scheduling data. t =1, otherwise display 0, i.e., d t =0. In one embodiment, assuming the preset duration is T, the end of the current duration can be indicated by comparing the current time t with the preset duration T. If t > T, then d t =1, otherwise d t =0.
[0130] In actual operation, after obtaining the scheduling data at each moment within a preset time period, it is cached in the replay buffer as a sequence as a training sample. The above steps of obtaining training samples are repeated to obtain a training sample set consisting of multiple training samples. The one or more training samples serve as the sample input for training the scheduling model. Here, the cached training samples are data sequences including environmental information, resource allocation, weighted throughput, and indicators representing whether the current preset time period has ended.
[0131] The training module 11 is used to train the scheduling model based on the training sample set to update the parameters of the scheduling model.
[0132] In this system, training samples are used as input to the scheduling model, which is then trained using deep reinforcement learning based on partial Markov decisions to update its parameters. For example, the parameters of the scheduling model are the weights of the neural network structure. In some embodiments, the optimal parameters can be computed, i.e., an analytical method can be used to update the scheduling model's parameters; this method is more suitable for scheduling problems in simple environments. In other embodiments, gradient descent can be used to update the scheduling model's parameters, enabling the scheduling model to achieve efficient scheduling.
[0133] The training system of the scheduling model in this application trains the scheduling model with a training sample set. While fully considering the partial observability of the environment, it can mine hidden information that is not observable in the environment during the training process based on multiple environmental information observed over a period of time. This enables the scheduling model to maximize throughput while satisfying resource consumption when performing scheduling operations.
[0134] When training a scheduling model based on a training sample set, in some embodiments, the scheduling model can be trained using all current training samples stored in the replay cache. In other embodiments, a subset of training samples can be sampled from the replay cache to train the scheduling model, thereby reducing computational complexity.
[0135] Since reinforcement learning involves continuous learning during the interaction between the agent and the environment, to improve the training accuracy of the scheduling model, this application allows the training of the scheduling model based on a training sample set to begin after the initial run of the scheduling model to obtain multiple training samples within a preset time period. In some embodiments, during the training of the scheduling model, the scheduling model and replay cache are first initialized, and then multiple training samples are obtained according to the steps described above and stored in the replay cache. After multiple training samples are stored in the replay cache, the training and learning process of the scheduling model begins. The scheduling model interacts with the environment to obtain new training samples on one hand, and updates its parameters using the cached training samples on the other hand, repeating this process cyclically. Based on this, the training module includes: a sampling unit for sampling the training sample set; and an update unit for updating the parameters of the scheduling model using gradient descent based on the sampled data. The training module performs the following steps: Figure 2 As shown in the description, it will not be repeated here.
[0136] In the case of a deep neural network structure in the scheduling model, to improve the performance of the scheduling algorithm, the scheduling model of this application includes a critic network structure and an actor network structure. The training module is used to train the critic network structure and the actor network structure based on the training sample set. The critic network structure and the actor network structure are as follows: Figure 3 , Figure 4 As shown in the description, it will not be repeated here.
[0137] Since the parameter training of the actor network is influenced by the output of the critic network, in order to update the parameters of the actor network more stably, in some embodiments, the training module includes: a first training unit for training the critic network structure based on the training sample set; and a second training unit for training the actor network structure using the critic network structure after at least one training iteration.
[0138] In practical scheduling problems, scheduling operations are constrained by resource consumption. To ensure that the scheduling model provides a hard guarantee of resource consumption constraints when performing scheduling operations, the training system of the scheduling model in this application further includes a constraint adaptation module, used to determine whether the obtained resource allocation meets preset constraints. If the constraints are met, the scheduling operation is performed based on the resource allocation; if the constraints are not met, the scheduling operation is not performed based on the resource allocation. The constraints can be set by those skilled in the art in which the scheduling model is applied, based on actual environmental constraints. The constraints may include instantaneous resource consumption constraints, average resource consumption constraints, service channel capacity constraints, etc. In some embodiments, the constraint is set to the average of the resource allocation obtained at the current moment and the cumulative used resource allocation being less than or equal to a preset value. The resource allocation obtained at the current moment refers to the resource allocation obtained through the scheduling model at the current moment, and the cumulative used resource allocation refers to the sum of the resource allocations at all moments before the current moment. The preset value is set according to the constraints of the actual scheduling problem; for example, in the field of wireless communication scheduling, the preset value is set based on the average power of the transmitting antenna.
[0139] Furthermore, considering the latency issue in scheduling operations, latency constraints can be set for the scheduling objects. These constraints can be set by those skilled in the art in which the scheduling model is applied, based on experience, or according to actual system requirements. In this case, the training system for the scheduling model of this application further includes: a caching module for caching the scheduling objects; and a processing module for processing the cached scheduling objects according to the latency constraints.
[0140] Here, the working method of each module in the training system of the scheduling model of this application is the same or similar to the corresponding steps in the training method of the above-mentioned scheduling model, and will not be repeated here.
[0141] This application also provides a scheduling method; please refer to [link / reference]. Figure 7 The figure shows a flowchart of the scheduling method of this application in one embodiment. As shown in the figure, the scheduling method includes steps S300 to S320.
[0142] In step S300, the scheduling object is received.
[0143] In this context, the scheduling object is the object to be sent to the user, and its nature depends on the application scenario of the scheduling operation. In scheduling operations, time is divided into discrete time slots, and the scheduling object is served in discrete packets. For example, in the field of wireless communication scheduling, the scheduling object is the wireless communication data packet; in the field of video stream scheduling, the scheduling object is the video stream data packet; and in the field of order delivery scheduling, the scheduling object is the delivery order.
[0144] In step S310, the resource allocation amount of the scheduling object is obtained by using the scheduling model trained by the above training method.
[0145] In some embodiments, when a time delay constraint is set, after receiving a scheduling object, the scheduling object is first cached in the caching module of the scheduling model, and then the scheduling model executes the scheduling operation at an appropriate time. If the scheduling object cached in the caching module is not served within the set time delay, the scheduling object is discarded.
[0146] After receiving a scheduling object, the scheduling model obtains the resource allocation amount for the scheduling object based on the observed environmental information. Specifically, the scheduling model is used as follows: using the actor network π1, actor network π2, critic network Q1, and critic network Q2 trained during the training phase, at time t, the following steps are performed:
[0147] a) The agent observes environmental variables o t ;
[0148] b) Let h be the historical observation at time t. t ←(h t - 1, a t - 1, o t );
[0149] c) Let action a t ←g(π1, π2), where
[0150]
[0151] d) The agent executes the action f(a) optimized by the constraint adaptation module. t ) into the learning environment;
[0152] e) Proceed to the next time step and repeat step a).
[0153] The scheduling model described above can be applied to various scenarios involving scheduling problems to maximize throughput while satisfying resource consumption constraints and / or latency limitations. For example, in the field of wireless communication scheduling, mobile users receive data packets from relays via wireless channels, and transmission antennas are responsible for transmitting these data packets to users. The scheduling object is the wireless communication data packet, which has latency limitations; if a data packet arrives at the relay and is not transmitted within the time limit, it will be discarded. Furthermore, the transmission antenna is limited by its average transmission power. Therefore, using the scheduling model trained according to this application for scheduling operations can maximize the number of successfully scheduled data packets while satisfying the average transmission power and latency limitations.
[0154] For example, in the field of video stream scheduling, mobile users stream online video via wireless networks, and routers are responsible for transmitting video data streams to users. The scheduling object is the video stream data packet. Since low-latency video-on-demand services are crucial for user experience, video stream data packets have latency constraints. Furthermore, transmission resources such as bandwidth are limited. Therefore, using the scheduling model trained in this application can maximize the successfully scheduled video traffic while satisfying bandwidth and latency constraints.
[0155] For example, in the field of order delivery scheduling, users order items on an online platform, and the platform is responsible for delivery. The scheduling object is the delivery order. To improve user experience, ensuring timely delivery is crucial, thus order delivery has time delay constraints. However, delivery capacity is often limited. Therefore, using the scheduling model trained in this application can maximize the number of successfully scheduled orders while satisfying both delivery capacity and time delay constraints.
[0156] In step S320, the scheduling objects are scheduled according to the obtained resource allocation.
[0157] This application also provides a scheduling system; please refer to [link / reference]. Figure 8 The figure shows a schematic diagram of the scheduling system in one embodiment of this application. As shown, the scheduling system includes an input unit 20 and a scheduling model 21. The input unit 20 is used to receive scheduling objects. The scheduling model 21 is used to obtain the resource allocation amount of the scheduling object and perform scheduling operations on the scheduling object according to the resource allocation amount.
[0158] The working methods of each module in the scheduling system of this application are the same or similar to the corresponding steps in the above scheduling method, and will not be repeated here.
[0159] This application also provides an electronic device. Please refer to [link / reference]. Figure 9 The figure shows a schematic diagram of the structure of an electronic device according to an embodiment of the present application. As shown, the electronic device includes at least one memory 30 and at least one processor 31.
[0160] In this embodiment, the electronic device is, for example, an electronic device loaded with an APP application or capable of accessing web pages / websites. The electronic device includes components such as a memory, a memory controller, one or more processing units (CPUs), peripheral interfaces, RF circuitry, audio circuitry, speakers, microphones, an input / output (I / O) subsystem, a display screen, other output or control devices, and external ports. These components communicate via one or more communication buses or signal lines. The electronic device includes, but is not limited to, personal computers such as desktop computers, laptops, tablets, smartphones, and smart TVs. The electronic device can also be an electronic device consisting of a host with multiple virtual machines and corresponding human-computer interaction devices (such as touch screens, keyboards, and mice) for each virtual machine.
[0161] The at least one memory is used to store at least one program; in embodiments, the memory may include high-speed random access memory and may also include non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. In some embodiments, the memory may also include memory located remotely from one or more processors, such as network-attached memory accessed via RF circuitry or external ports and communication networks, wherein the communication network may be the Internet, one or more intranets, local area networks, wide area networks, storage area networks, etc., or suitable combinations thereof. The memory controller can control access to the memory by other components of the device, such as the CPU and peripheral interfaces.
[0162] In one embodiment, the at least one processor is connected to the at least one memory and is used to execute and implement at least one embodiment as described above for the training method of the scheduling model when running the at least one program, such as... Figure 1 , Figure 2 as well as Figure 5 The embodiments described herein; or, performing and implementing at least one embodiment as described above for the scheduling method, such as... Figure 7 The embodiments described herein. In these embodiments, the processor is operatively coupled to memory and / or non-volatile storage devices. More specifically, the processor can execute instructions stored in the memory and / or non-volatile storage devices to perform operations in a computing device, such as generating image data and / or transmitting image data to an electronic display. Thus, the processor may include one or more general-purpose microprocessors, one or more special-purpose processors, one or more field-programmable logic arrays, or any combination thereof.
[0163] This application also provides a cloud server system. Please refer to [link / reference]. Figure 10The figure shows a schematic diagram of the structure of the cloud server system in one embodiment of the present application. As shown in the figure, the cloud server system includes at least one storage device 40 and at least one processing device 41.
[0164] In some embodiments of this application, the cloud server system can be deployed on one or more physical servers based on factors such as functionality and load. When distributed across multiple physical servers, the server can consist of servers based on a cloud architecture. For example, cloud-based servers include public cloud servers and private cloud servers, where public or private cloud servers include Software-as-a-Service (SaaS), Platform-as-a-Service (PaaS), and Infrastructure-as-a-Service (IaaS). Examples of private cloud servers include Meituan Cloud Computing Service Platform, Alibaba Cloud Computing Service Platform, Amazon Cloud Computing Service Platform, Baidu Cloud Computing Platform, and Tencent Cloud Computing Platform. The server can also consist of distributed or centralized server clusters. For example, a server cluster consists of at least one physical server. Each physical server is configured with multiple virtual servers, and each virtual server runs at least one functional module of the restaurant merchant information management server. The virtual servers communicate with each other via a network.
[0165] In one embodiment, the at least one storage device is used to store at least one program, and the at least one processing device is connected to the at least one storage device for executing and implementing at least one embodiment as described above for the training method of the scheduling model when running the at least one program, such as... Figure 1 , Figure 2 as well as Figure 5 The embodiments described herein; or, performing and implementing at least one embodiment as described above for the scheduling method, such as... Figure 7 The embodiments described herein.
[0166] This application also provides a computer-readable and writable storage medium storing at least one program, which, when executed by a processor, implements at least one embodiment as described above for the training method of the scheduling model, such as... Figure 1 , Figure 2 as well as Figure 5 The embodiments described herein; or perform and implement at least one embodiment as described above for the scheduling method, such as Figure 7 The embodiments described herein.
[0167] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0168] In the embodiments provided in this application, the computer-readable and writable storage medium may include read-only memory, random access memory, EEPROM, CD-ROM or other optical disc storage devices, disk storage devices or other magnetic storage devices, flash memory, USB flash drive, portable hard drive, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Additionally, any connection may be appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable and writable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are intended for non-transient, tangible storage media. The disks and optical discs used in the application include compact discs (CDs), laser discs, optical discs, digital multifunction discs (DVDs), floppy disks, and Blu-ray discs, where disks typically copy data magnetically, while optical discs use lasers to copy data optically.
[0169] In one or more exemplary aspects, the functionality described by the computer program of the methods of this application can be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored or transmitted as one or more instructions or code onto a computer-readable medium. The steps of the methods or algorithms disclosed in this application can be embodied in processor-executable software modules, wherein the processor-executable software modules can reside on a tangible, non-transitory computer-readable and writable storage medium. The tangible, non-transitory computer-readable and writable storage medium can be any available medium accessible to a computer.
[0170] The flowcharts and block diagrams in the foregoing figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Accordingly, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0171] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A training method of a scheduling model, characterized by, The method comprises the following steps: obtaining a training sample set in a preset time period, the training sample set comprising training samples in the preset time period, the training samples comprising scheduling data in each time slot in the preset time period, the scheduling data comprising environment information, a resource allocation amount of a scheduling object obtained by running a scheduling model, and a throughput obtained after performing a scheduling operation based on the resource allocation amount; wherein the environment information comprises a state of the scheduling object and a state of a service channel, the service channel being a channel for allocating resources to the scheduling object; the step of obtaining the training sample set in the preset time period comprises: obtaining the resource allocation amount of the scheduling object in the corresponding time slot and the instantaneous weighted throughput in the corresponding time slot by running the scheduling model according to the environment information in each time slot; and caching the training samples; wherein the scheduling model comprises two actor networks and two critic networks, the critic network structure comprising a first fully connected branch and a first long short-term memory branch, the first fully connected branch and the first long short-term memory branch being connected in sequence through a first splicing layer, a third fully connected layer, a third rectified linear unit activation function, and a fourth fully connected layer; wherein the first fully connected branch comprises a first fully connected layer and a first rectified linear unit activation function in sequence, and the first long short-term memory branch comprises a second fully connected layer, a second rectified linear unit activation function, and a first long short-term memory layer in sequence; the critic network structure comprises a first fully connected branch and a first long short-term memory branch, the first fully connected branch and the first long short-term memory branch being connected in sequence through a first splicing layer, a third fully connected layer, a third rectified linear unit activation function, and a fourth fully connected layer; wherein the first fully connected branch comprises a first fully connected layer and a first rectified linear unit activation function in sequence, and the first long short-term memory branch comprises a second fully connected layer, a second rectified linear unit activation function, and a first long short-term memory layer in sequence; and training the scheduling model based on the training sample set to update the parameters of the scheduling model; the step of training the scheduling model based on the training sample set comprises: sampling the training sample set; updating the parameters of the scheduling model in a gradient descent manner based on the sampled data; and repeating the above steps until a preset number of updates is met to complete training in a preset time period, so that the parameters of the scheduling model are updated during the acquisition process of the training samples in the next preset time period; during training of the scheduling model based on the training sample set, the parameters of the critic network in the scheduling model are updated sequentially, and the parameters of the actor network in the scheduling model are updated intermittently; wherein the training of the critic network target value is realized by means of a softmax operator; The method further comprises a step of determining whether the obtained resource allocation satisfies a preset constraint condition; the constraint condition is an average resource consumption constraint, and if the constraint condition is satisfied, a scheduling operation is performed based on the resource allocation; if the constraint condition is not satisfied, the action that does not satisfy the constraint condition is set to zero, so that in entering a learning environment and buffering scheduling data as training samples, the resource allocation of the scheduling object and the throughput obtained after performing the scheduling operation based on the resource allocation are set to zero, and the constraint condition comprises that the average of the sum of the resource allocation obtained at the current time and the accumulated used resource allocation is less than or equal to a preset value; wherein the resource allocation obtained at the current time refers to the resource allocation obtained at the current time through the scheduling model, and the accumulated used resource allocation refers to the sum of the resource allocation at all times before the current time.
2. The method of claim 1, wherein, The method further comprises: buffering the scheduling object; and processing the buffered scheduling object according to the delay limit condition.
3. A training system of a scheduling model, characterized by, The method further comprises: an obtaining module, configured to obtain a training sample set within at least one preset time length, the training sample set comprising training samples within each preset time length, the training samples comprising scheduling data at each time slot within the preset time length, the scheduling data comprising environment information, resource allocation of a scheduling object obtained through running the scheduling model, and throughput obtained after performing a scheduling operation based on the resource allocation; wherein the environment information comprises a state of the scheduling object and a state of a service channel, the service channel being a channel for allocating resources to the scheduling object; the obtaining module comprises: an obtaining unit, configured to obtain, according to the environment information at each time slot, the resource allocation of the scheduling object at the corresponding time slot and the instantaneous weighted throughput through running the scheduling model; and a buffering unit, configured to buffer the training samples; wherein the scheduling model comprises two actor networks and two critic networks, the critic network structure comprising a first fully connected branch and a first long short-term memory branch, the first fully connected branch and the first long short-term memory branch being connected in sequence through a first splicing layer, a third fully connected layer, a third rectified linear unit activation function and a fourth fully connected layer; wherein the first fully connected branch comprises a first fully connected layer and a first rectified linear unit activation function in sequence, and the first long short-term memory branch comprises a second fully connected layer, a second rectified linear unit activation function and a first long short-term memory layer in sequence; and the critic network structure comprises a first fully connected branch and a first long short-term memory branch, the first fully connected branch and the first long short-term memory branch being connected in sequence through a first splicing layer, a third fully connected layer, a third rectified linear unit activation function and a fourth fully connected layer; wherein the first fully connected branch comprises a first fully connected layer and a first rectified linear unit activation function in sequence, and the first long short-term memory branch comprises a second fully connected layer, a second rectified linear unit activation function and a first long short-term memory layer in sequence. The training module is configured to train the scheduling model based on the training sample set to update parameters of the scheduling model; the training module comprises a sampling unit configured to sample the training sample set; and an updating unit configured to update the parameters of the scheduling model in a gradient descent manner based on the sampled data; the training module further repeats the sampling and parameter updating until a preset number of updates is reached to complete training within a preset time period, so that the parameters of the scheduling model are updated during the acquisition of the training sample in the next preset time period; the training module comprises a first training unit configured to perform successive updating of parameters of a critic network in the scheduling model during training of the scheduling model based on the training sample set; and a second training unit configured to perform interval updating of parameters of an actor network in the scheduling model; wherein the training of the critic network target value is achieved by means of a softmax operator; The constraint adaptation module is further configured to determine whether the obtained resource allocation meets a preset constraint condition; the constraint condition is an average resource consumption constraint; if the constraint condition is met, the scheduling operation is performed based on the resource allocation; if the constraint condition is not met, the action that does not meet the constraint condition is set to zero, so that the resource allocation of the scheduling object and the throughput obtained after the scheduling operation based on the resource allocation are set to zero when entering the learning environment and caching the scheduling data as the training sample; the constraint condition comprises that the average of the sum of the obtained resource allocation at the current time and the cumulative used resource allocation is less than or equal to a preset value; wherein the obtained resource allocation at the current time refers to the resource allocation obtained by the scheduling model at the current time, and the cumulative used resource allocation refers to the sum of the resource allocation at all times before the current time.
4. The system for training of scheduling models according to claim 3, characterized in that, The scheduling model training system further comprises: a caching module configured to cache the scheduling object; and a processing module configured to process the cached scheduling object according to the time delay limit condition.
5. A scheduling method characterized by, The method comprises the following steps: receiving a scheduling object; training the obtained scheduling model by using the training method according to any one of claims 1-2 to obtain a resource allocation of the scheduling object; and performing a scheduling operation on the scheduling object according to the resource allocation. The scheduling object is a wireless communication data packet.
6. The scheduling method of claim 5, wherein, The scheduling object is a video stream data packet.
7. The scheduling method of claim 5, wherein, The scheduling object is a delivery order.
8. The scheduling method of claim 5, wherein, The method comprises:
9. A dispatch system characterized by, an input unit configured to receive a scheduling object; a scheduling model trained by using the training method according to any one of claims 1-2, used to obtain a resource allocation of the scheduling object, and perform a scheduling operation on the scheduling object according to the resource allocation. The method comprises:
10. An electronic device, comprising: at least one memory configured to store at least one program; at least one processor connected to the at least one memory, configured to run the at least one program to perform and implement the training method of the scheduling model according to any one of claims 1-2, or perform and implement the scheduling method according to any one of claims 5-8. The method comprises:
11. A cloud server system, characterized by at least one storage device, configured to store at least one program; at least one processing device, connected to the storage device, configured to execute the at least one program to perform and implement the training method of the scheduling model according to any one of claims 1 to 2, or perform and implement the scheduling method according to any one of claims 5 to 8.
12. A computer-readable storage medium, characterized in that, at least one program, stored in the storage device, and executed by the processor to perform and implement the training method of the scheduling model according to any one of claims 1 to 2, or perform and implement the scheduling method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Multi-agent reinforcement learning scheduling method and system, and electronic device
CN109947567A
Iterative reinforcement learning method for high-efficiency value function of shared recurrent neural network
CN111582441A