Computer-implemented method, computing device, radio resource scheduler of radio communication network and computer software
By using transfer models and machine learning models in the system to optimize action selection, the probability of negative results is reduced, and the problem that the MDP framework cannot avoid catastrophic results is solved, and more efficient resource allocation and radio resource scheduling is achieved.
Patent Information
- Application Number
- CN202380082933.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-05
- Filing Date
- 2023-06-23
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art cannot effectively reduce the probability of negative long-term results when selecting system actions, especially in the case of uncertainty in the long-term effects of the action, and the traditional Markov decision-making process (MDP) framework cannot effectively avoid potential catastrophic results.
Through a computer-implemented method, using transfer models and machine learning models, the calculation system's action selection in different states is ensured that the selected action is along a gain path with a high probability of at least equal to or higher than a predetermined value, reducing the probability of a low gain path, combining binning rules and iterative computing optimization strategies to reduce the occurrence of negative results.
It effectively reduces the probability of negative results in the system, optimizes resource allocation and radio resource scheduling, and improves the system's decision-making efficiency in uncertain environments.
Smart Images

Figure CN120303970A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of selecting actions to be performed by a system. More specifically, the present disclosure relates to the field of selecting an optimal action of a system when performing an action whose output is not deterministically known. Background Art
[0002] Many systems can be modeled to determine which actions should be performed by the system to optimize the output. Examples of such systems include, for example, robots, schedulers, or more generally any physical system capable of performing actions to achieve a goal.
[0003] In many systems, the output of the same action may result in different outputs. For example, if a robot is instructed to search for an object when its battery level is low, it is generally not possible to know in advance whether the robot will find the object or stop because its battery has been completely discharged. In another example, if a scheduler of a radio communication network grants radio resources to a user terminal to send a message, the user terminal may successfully send the message or may not be able to send the message, depending on unpredictable network conditions. Some outputs (such as the robot finding an object or successfully sending a message) are considered desirable, while other outputs (such as the robot's battery being completely discharged or failing to send a message) are considered undesirable. Thus, the general purpose of controlling a system is to improve the probability of obtaining the most desirable output.
[0004] However, the situation becomes complex when consecutive actions lead to long-term outputs that cannot be deterministically predicted. For example, a scheduler of a radio communication network can authorize radio resources to different user terminals without knowing in advance whether the message transmission of the user terminals will be successful. At the same time, some user terminals may be bound by application limitations (e.g., limitations from an application requesting the transmission of a message), and miss the limitation if the message is sent too late, which is considered a highly undesirable output. Therefore, the optimization of the decision-making process of such a system is a complex problem.
[0005] A Markov Decision Process (MDP) provides a framework for solving such problems. The MDP models a system using actions, rewards, and state transition probabilities. At each time step, the system is in a current state, in which it can perform multiple different actions. When performing an action, a state transition can occur from the current state to one or more subsequent states with different probabilities. The probability of the state transition occurring depends on the current state and the action performed.
[0006] Each state transition is also associated with a reward, depending on the desirability of the state transition, which can be positive or negative. Thus, different paths of consecutive actions and state transitions can be defined. Each path can be associated with a gain defined by the consecutive rewards of the state transitions of the path and a probability corresponding to the combination of the probabilities of the state transitions of the path. Thus, the MDP framework allows associating to each path whether the long-term outcome of the defined path is a desired gain and the probability that the path occurs when a series of actions are performed in the consecutive states of the path.
[0007] For example, R.S. Sutton and A.G. Barto describe in "Reinforcement Learning: An Introduction (Second Edition)" (The MIT Press, 2020) decision algorithms based on MDP. Decision algorithms based on MDP typically select, in the current state, the action that maximizes the expected gain. More specifically, they calculate the average gain for each state and determine, for each possible action in the current state, the gain associated with each possible subsequent state and the associated probability. Then, they select the action that provides the highest expected gain.
[0008] If such decision processes generally allow increasing the expected gain of the system, they cannot take into account the probability of actions that lead to potentially catastrophic outcomes. In fact, an action associated with a high expected gain may lead not only to paths with high gain, but also to many paths with very high gain, and few paths with low probability and very low gain. Thus, decision processes based on maximum expected gain cannot minimize the likelihood of potentially catastrophic outcomes.
[0009] Therefore, in a system where the long-term effects of each action are not known in a deterministic manner in advance, a decision process is needed that reduces the probability of selecting actions that lead to very negative long-term outcomes. SUMMARY OF THE INVENTION
[0010] The present disclosure improves the situation.
[0011] A computer-implemented method is proposed for selecting an action from a set of possible actions performed in a system, when the system is in a current state within a set of possible states at a time step, the method comprising: performing the selection of the action that, from the current state, guides the system with the highest probability along a path corresponding to a series of actions and state transitions in consecutive time steps and having a gain of at least equal to or higher than a predetermined value; causing the system to perform the action.
[0012] By "system", we designate a representation of a real-world system that can be modeled according to different states and perform different actions, each action possibly causing the system to switch from an initial state to one or more possible subsequent states.
[0013] By "path", we specify a sequence of successive transitions between successive states of the system at successive time steps. Within a path, each transition may or may not be associated with an action.
[0014] By "along a path by the system", we specify that a sequence of states and state transitions of the system at successive time steps corresponds to a sequence of states and state transitions defined by the path.
[0015] By "guiding the system along a path with a probability", we specify that when the system is in a state and performs an action, it is expected to have a given probability of following that path.
[0016] By "gain", we specify a value associated with a path that indicates the extent to which the successive states defined by the path correspond to good or bad outcomes.
[0017] By "a path generated from the current state upon performing an action", we specify that the initial state of the first state transition is the current state and its first transition upon performing the action is a path possible from the current state.
[0018] This allows for the selection in the system of an action that exhibits the lowest probability of resulting in a gain below a predetermined threshold, where each action can lead to different outcomes with the defined probabilities.
[0019] Thus, this allows for the performance of an action with the highest possible probability of avoiding outcomes considered bad in the system, where the long-term effects of each action are not known in advance in a deterministic manner.
[0020] On the other hand, a computing device is proposed that includes at least one processing unit configured to implement at least a part of the method as defined herein.
[0021] On the other hand, a radio resource scheduler for a radio communication network is proposed that includes at least one processing unit configured to implement at least a part of the method as defined herein.
[0022] On the other hand, a computer software including instructions is proposed for implementing at least a part of the method defined herein when the software is executed by a processor.
[0023] On the other hand, a computer-readable non-transitory recording medium is proposed on which software is registered to implement the method as defined herein when the software is executed by a processor.
[0024] The following features may optionally be implemented individually or in combination with each other:
[0025] Advantageously, the selection is based on: retrieving a transition model that defines, for at least one state transition from an initial state to a subsequent state between two consecutive time steps, the probability of performing a transition from the initial state to the subsequent state when performing an action belonging to the set of possible actions performed in the initial state, and a reward associated with the transition; and calculating, for multiple paths using at least the transition model: the gain of the multiple paths; the probability of the multiple paths.
[0026] By "transition model", we specify a model that defines state transitions and determines the probability and reward associated with each state transition. At least one of the probabilities of the state transitions can depend on the actions performed by the system in the state.
[0027] By "reward", we specify a value that indicates the desirability of the transition. The reward can be positive if the transition results in a desired outcome, or negative if the transition results in an undesired outcome.
[0028] By "obtaining", we specify different means of obtaining the probability and gain for each of the multiple paths. For example, the obtaining step can include:
[0029] - calculating the probability and gain;
[0030] - reading or receiving pre-calculated probability and gain during a training phase;
[0031] - and the like.
[0032] This allows taking into account the probability and outcome associated with the possible transitions used to determine the gain of the paths, and thus obtaining a calculation of the gain of the paths that accurately takes into account the evolution of the state of the system.
[0033] These features are equally advantageous and also propose a method that separately includes the above steps.
[0034] Accordingly, a computer-implemented method is also proposed, which includes: retrieving a transition model that defines, for at least one state transition from an initial state to a subsequent state between two consecutive time steps, the probability of performing a transition from the initial state to the subsequent state when performing an action belonging to the set of possible actions performed in the initial state, and a reward associated with the transition; and calculating, for multiple paths using at least the transition model: the gain of the multiple paths; the probability of the multiple paths.
[0035] Advantageously, the gains of the plurality of paths are calculated as the cumulative rewards of the consecutive state transitions and action sequences for each path in the plurality of paths in a continuous time state; the probabilities of the plurality of paths are calculated as the combined probabilities of each combination of state transitions and actions for each path in the plurality of paths in a continuous time state.
[0036] By "cumulative reward", we specify the sum of the rewards corresponding to the consecutive state transitions of a path. The reward can depend only on the state transition, or on both the state transition and the corresponding action. The sum can be a weighted sum. For example, the weight associated with each reward can decrease with the time step.
[0037] By "combined probability", we specify the probability of a path defined as the combination of the probabilities of state transitions. The combined probability can generally be the product of the probabilities of consecutive state transitions.
[0038] This allows the use of all transitions in a finite or infinite time window to determine the gains and probabilities of paths, and thus obtain a reliable calculation of the gains and probabilities of paths.
[0039] Advantageously, a binning rule is applied to the reference gains of the paths generated from each state to obtain the size of the reference set associated with each state, including the gains of the paths generated from the state and the probabilities of said paths to a predetermined size.
[0040] By "binning rule", we specify the rule for combining multiple paths associated with similar gains into a single path with a probability equal to the cumulative probability of the paths belonging to the bin.
[0041] This allows the number of paths to be considered to be limited by combining paths associated with similar gains. Thus, this allows the computational complexity to be limited while obtaining reliable results for the gains and probabilities.
[0042] Advantageously, using a model to calculate the gains and probabilities includes performing iterative calculations of the reference set of states for each state in the set of possible states, the reference set of states including the reference gains of the paths generated from the state when performing an action, and the reference probabilities of the system following the paths generated from the state when performing the action, and applying the binning rule at each iteration of the iterative calculation.
[0043] This allows the gains and probabilities of paths starting from each possible state to be calculated recursively while maintaining the size of the set representing the probabilities and gains of paths starting from each state, and thus making recursive calculation possible even when the number of possible paths is large.
[0044] Advantageously, the iterative calculation includes: for each state in the set of possible states, obtaining an initial reference set; until a stopping criterion is met: for each state in the set of possible states; selecting, according to a policy, an action to be performed in the state; for each subsequent state having a non-zero transition probability from the state using the selected action, redefining the initial reference set of the state by concatenating the following: the reference gain in the reference set of the subsequent state, multiplied by a discount factor and added to the reward for the transition from the state to the subsequent state; the reference probability of the subsequent state, multiplied by the non-zero transition probability; reducing the size of the reference set of the state to the predetermined size by applying a binning rule to the reference gain of the state
[0045] By "policy", we specify a set of rules that define which action should be performed in which state of the system.
[0046] This provides an efficient way to implement the recursive calculation while ensuring compliance with the stopping criterion at the end of the calculation.
[0047] Advantageously, the iterative calculation includes: for each state in the set of possible states, obtaining an initial reference set; until a stopping criterion is met: for each state in the set of possible states; selecting, according to a policy, an action to be performed in the state; for each subsequent state having a non-zero transition probability from the state using the selected action, redefining the initial reference set of the state by concatenating the following: the reference gain in the reference set of the subsequent state multiplied by a discount factor and added to the reward for the transition from the state to the subsequent state; the reference probability of the subsequent state multiplied by the non-zero transition probability; reducing the size of the reference set of the state to the predetermined size by applying the binning rule to the reference gain of the state.
[0048] This allows obtaining a policy optimized for the goal of obtaining the highest probability of selecting an action that results in a gain higher than or at least equal to a threshold.
[0049] Advantageously, the stopping criterion is selected from the group including at least one of the following: a maximum number of iterations; a criterion for the convergence of the gain between two consecutive iterations.
[0050] Advantageously, the transition model belongs to a Markov decision process model, where the gain of a path is calculated as the weighted sum of the rewards of each consecutive transition in the path, where the weight is a discount factor included between zero and one that is the power of the transition exponent.
[0051] By "discount factor", we specify a factor that reduces the relative importance of the reward corresponding to more time steps according to the time difference between the reward and the current time step.
[0052] The Markov decision process model provides an effective way to calculate the gain, where continuous rewards are considered and the weight factor associated with the reward decreases over time.
[0053] Advantageously, the step of using the transition model further includes using a machine learning model that is trained to predict path probabilities based on the states of the paths generated from the state to have a gain belonging to each of a plurality of intervals of possible gains.
[0054] This allows only the trained model to be stored and thus reduces the size of the model parameters while providing an accurate prediction of the probabilities.
[0055] Advantageously, the machine learning model is iteratively trained using an error signal calculated as a cost function based on the prediction of the path probability starting from the initial state and the prediction of the path probability starting from the subsequent state multiplied by the transition probability from the initial state to the subsequent state when performing the action, the action guiding the system from the initial state along a path with a gain of at least equal to or higher than a predetermined value with the highest probability.
[0056] This allows the machine learning model to be effectively trained to converge to a prediction consistent with the optimal action selection.
[0057] Advantageously, the iterative calculation includes: until a stop criterion is met: performing a loop that selects a plurality of states from the set of possible states, and for each state in the plurality of states: using the machine learning model to calculate the prediction of the path probability starting from the initial state; for each possible action in the state: retrieving from the transition model the probability of entering each subsequent state when performing the possible action and the associated reward in the transition model; using the machine learning model to calculate the prediction of the path probability starting from each subsequent state; calculating the prediction of the path probability starting from the initial state when performing the possible action based on the prediction of the path probability from each subsequent state and the probability of entering each subsequent state and the associated reward when performing the possible action; when performing the possible action, selecting an action that guides the system from the initial state along a path with a gain of at least equal to or higher than a predetermined value with the highest probability based on the prediction of the path probability from the initial state; calculating the value of the cost function based on the prediction of the path probability from the initial state and the prediction of the path probability starting from the initial state when performing the possible action, the possible action guiding the system from the initial state along a path with a gain of at least equal to or higher than a predetermined value with the highest probability; updating the machine learning model based on the plurality of values of the cost function calculated for the plurality of states
[0058] This provides an effective way to train a machine learning model to calculate the probability of a path for each action in each state and thus determine the action to be selected.
[0059] Advantageously, the method is implemented by a resource allocator of the system, where: the set of possible actions corresponds to possible resource allocations in the state; the reward is defined based on the resource allocation of the system and the result of the state transition.
[0060] By "resource allocator", we specify a module capable of allocating physical or computational resources of a physical or computational system.
[0061] This allows the resource allocator to allocate resources so as to reduce the probability of resource allocations that lead to negative outcomes for the system.
[0062] Advantageously, the method is implemented by a radio resource scheduler of a radio communication network, where: the set of possible actions corresponds to possible allocations of radio resources to users of the radio communication network; at least one probability of performing a state transition depends on the probability that a user of the radio communication network successfully transmits a data packet through the network when radio resources are allocated to the user; at least one reward depends on compliance with a time limit for transmitting at least one data packet.
[0063] By "time limit for transmitting at least one data packet", we specify any limit relative to the time when one or more data packets should be transmitted. The time limit can be, for example, a limit relative to the time limit for transmitting and / or receiving data packets, such as Round Trip Time (RTT), Packet Delay Budget (PDB), etc. The time limit can be, for example, a time limit set by an application that needs to transmit data packets. In this case, it can be referred to as an "application limit".
[0064] This allows the radio resource scheduler to allocate radio resources to users of the radio communication network so as to reduce the likelihood of negative outcomes associated with failures to meet telecommunication constraints.
[0065] Other features, details, and advantages will be shown in the following detailed description and the drawings. Description of the Drawings
[0066] Figure 1
[0067] Figure 1 Shows an example of a policy-dependent gain in an exemplary system that includes a mobile robot modeled according to the MDP framework, and the present invention can be implemented for its control.
[0068] Figure 2
[0069] Figure 2 An example of a radio resource scheduler of a radio communication network in which the present invention can implement its control.
[0070] Figure 3
[0071] Figure 3 An example of a computer-implemented method according to multiple embodiments of the present invention.
[0072] Figure 4
[0073] Figure 4 A first example of the calculation of a reference gain and a probability according to an embodiment.
[0074] Figure 5
[0075] Figure 5 A second example of the calculation of a reference gain and a probability according to an embodiment.
[0076] Figure 6
[0077] Figure 6 A third example of the calculation of a reference gain and a probability according to an embodiment.
[0078] Figure 7
[0079] Figure 7 An example of multiple paths in an embodiment of the present invention.
[0080] Figure 8
[0081] Figure 8 Represents a part of a grid that shows possible paths in multiple embodiments of an embodiment of the present invention. Detailed Description
[0082] For the sake of facilitating the presentation of the present invention, some embodiments will be presented with respect to a model that conforms to the Markov decision process (MDP) framework. However, it should be noted that this is provided only by way of non-limiting example, and the present invention is not limited to examples of this framework. For the sake of facilitating the presentation of the present invention, some notations commonly used in the MDP framework are reminded below. It should be noted that the following definitions are provided only by way of non-limiting examples of the definitions in multiple embodiments of the present invention.
[0083] In the MDP framework, the gain at time t can be recorded, for example, as G t , and is expressed as:
[0084]
[0085] (Equation 1)
[0086] where:
[0087] -r i is the reward received at each time step i;
[0088] -λ represents the discount, 0 < λ < 1, which means that the weight assigned to the reward decreases over time towards 0.
[0089] It is also possible to calculate the gain in a finite time window. In this case, λ = 1 can be used.
[0090] In the MDP framework, the policy π represents one or more decision rules to be used at each time step. In other words, the policy π defines the action to be taken by the system at each time step, which depends at least on the state s ∈ S, where S is the set of possible system states at time step t of the system.
[0091] In some cases, multiple different policies are available. Therefore, one of the problems solved in a system modeled by the MDP framework is to select the best policy from the set of possible policies.
[0092] A solution widely used in the prior art consists of selecting the policy that maximizes the expected gain E[G t . The expected gain is calculated based on the actions taken at each time step according to the policy and the probabilities of performing different state transitions in the model. Here, the probabilities of the transitions model the randomness in the environment of the system.
[0093] Since a given policy π deterministically selects the action in a state, the value of the state s ∈ S obtained using the given policy π can be defined as the expected gain of the given system being in state s at time t and using the given policy π:
[0094] v π (s) = E[G t | s]
[0095] (Equation 2)
[0096] If
[0097]
[0098] then the policy π * is considered optimal.
[0099] In other words, if a policy provides a better value for each possible state than any possible alternative policy, then that policy is considered the best.
[0100] Under the optimal policy π * the value of state s can be expressed via the Bellman equation as:
[0101]
[0102] where:
[0103] - A s is the set of allowed actions in state s;
[0104] - p(s j | s, a) represents the transition probability from state s to state s j given action a;
[0105] - r(s j , s) is the reward obtained when transitioning from state s to state s j ;
[0106] - R(s, A) = E[r|s, a] = ∑ j p(s j | s, a)r(sj, s)
[0107] Given π * , the optimal action in state s is obtained as follows:
[0108]
[0109] Within this framework, pre - existing policies can be compared.
[0110] Policies can also be found. In particular, a policy that defines the maximum expected gain of the optimization system can be defined. Standard methods for obtaining the optimal policy include calculating the value which involves calculating as:
[0111]
[0112] This method is described, for example, in "Reinforcement Learning: An Introduction (Second Edition)" by R.S. Sutton and A.G. Barto (MIT Press, 2020).
[0113] where v n converges to We refer to the obtained policy as the maximum expected gain policy.
[0114] As described above, the policy definition allows obtaining a policy that provides the maximum expected gain, but it fails to consider whether the policy may lead to very low gains in some cases. In other words, it is possible to obtain a maximum expected gain policy that has a very high probability of leading to state transitions with high gains, but may lead to state transitions with very low gains in some cases. This situation is problematic because low gains may correspond to catastrophic outcomes in the physical system.
[0115] Now refer to Figure 1 。
[0116] Figure 1 represents an example of the gain depending on the policy in an exemplary system that includes a mobile robot modeled according to the MDP framework, and the control of which can be implemented by the present invention.
[0117] Figure 1 The example of
[0118] - A mobile robot operating on a battery should collect empty cans;
[0119] - The robot can be in two battery states S = {s1 = "low", s2 = "high"} that form a set of possible states;
[0120] - In each of the two states, the possible actions form a set of possible actions A = {a1 = "search", a2 = "wait", a3 = "recharge"};
[0121] - In the case of the action "wait", the robot remains in the same state with probability p(s i |s i , a2) = 1 and obtains a reward r wait ;
[0122] - With the action "search", the robot transitions from the state "high" to the state "low" with probability p(s1|s2, a1) = 1 - β and obtains a reward r search > r wait . In the state "low", it also remains in the same state with probability β and obtains a reward r search , but otherwise obtains a negative reward r rescue , and returns to the state "high". Finally, with the action "recharge", the robot enters the state "high" with probability p(s2|s i , a3) = 1 and obtains a reward of 0.
[0123] For the simulation, the following parameters are used: r rescue = -1, r wait = 0.4, r searchα = 0.9, β = 0.8, and discount λ = 0.8.
[0124] In this example, two strategies are considered:
[0125] - The maximum expected gain strategy consists of performing the action "search" in both states, i.e., the robot always searches regardless of the state;
[0126] - An alternative strategy, where the robot implements the action "wait" in the state "low".
[0127] Based on this, when the robot is initially in states s1 = "low" and s2 = "high", the gain and probability are calculated for each possible sequence of state transitions for the two strategies.
[0128] Figure 1 Denotes the probability that a given strategy provides a gain higher than a quantity. More specifically:
[0129] - The horizontal axis represents the number x;
[0130] - The vertical axis represents the probability P(G t > x) that the gain is higher than x when using the strategy.
[0131] Figure 1 Denotes the probability that the obtained gain is higher than different quantities for different cases:
[0132] - Pmax high Denotes the probability of the maximum expected gain strategy when the robot is initially in state s2 = "high";
[0133] - Pmax low Denotes the probability of the maximum expected gain strategy when the robot is initially in state s1 = "low";
[0134] - Palt denotes the probability of the alternative strategy regardless of the initial state of the robot.
[0135] More specifically, Figure 1 Denotes the empirical complement cumulative density function (CCDF) obtained by simulating the two strategies.
[0136] Figure 1 Shows:
[0137] - The alternative strategy provides a gain that is always equal to 2;
[0138] - The maximum expected gain strategy has a very high probability of resulting in a gain higher than 2, but also has a non - zero probability of resulting in a gain lower than 2 in some cases.
[0139] Situations where the maximum expected gain strategy results in a gain less than 2 typically correspond to the situation where the robot's battery discharges prematurely and the robot needs to be rescued prematurely.
[0140] This example shows that if the maximum expected gain strategy generally provides good results, it may lead to negative results in some cases. In particular, the maximum expected gain strategy is not optimal for maximizing p(G t >α|s1) (where α < 2).
[0141] The present invention can be used, for example, to improve the strategies used by such robots.
[0142] Now refer to Figure 2 .
[0143] The present invention can be used to define strategies for many different purposes, particularly for resource allocation. For example, a resource allocator can use the present invention to optimize resource allocation. The resource allocator can allocate physical or computing resources of a physical or computing system. For example, the resource allocator can allocate computing, memory, network rerouting, etc.
[0144] The present invention can be used, for example, to improve the control of radio resource schedulers in radio communication networks.
[0145] Figure 2 An example of a radio resource scheduler for a radio communication network for control in which the present invention can be implemented is shown.
[0146] In Figure 2 's example, the radio resource scheduler is able to allocate a single radio resource to a single user out of 6 different users at each time step. The user to whom the radio resource is allocated is able to send a data packet. However, the data packet transmission has a probability of failure, which depends on multiple factors, such as the user itself, the state of the radio communication network, etc. At the same time, each user needs to comply with a time limit. For example, each user may need to successfully send a data packet periodically. For example, the time limit can be an application limit, for example, a limit set by the application that needs to send the data packet.
[0147] In Figure 2 's example, the horizontal axis represents consecutive time steps. The current time step is time step 8, and the resource allocator has allocated a radio resource to user 5, as shown by the black dot in the figure. The vertical axis represents different users. Each user is associated with a horizontal line, where the square represents the number of remaining time steps before they need to successfully send a data packet.
[0148] In this example:
[0149] - Users 1 and 5 need to send data packets by time step 8 at the latest;
[0150] - User 2 needs to send a data packet no later than time step 11;
[0151] - User 3 needs to send a data packet no later than time step 12;
[0152] - User 4 needs to send a data packet no later than time step 14.
[0153] Therefore, in this example:
[0154] - The set of possible actions includes allocating radio resources to users 1, 2, 3, 4, 5, or 6, respectively;
[0155] - The state of the system includes the number of time steps remaining for each user before a data packet needs to be successfully sent;
[0156] - When a user sends a data packet, the data packet transmission may succeed with probability p or fail with probability 1 - p;
[0157] - If the time limit is not met, i.e., in the current case, if a user cannot successfully transmit a data packet before the transmission deadline, it is a highly negative reward. For example, if the time limit is an application limit, this may cause an error in the application.
[0158] As will be explained in more detail below, the method according to the present invention allows obtaining a strategy that increases the probability of having a gain higher than a threshold. In other words, it reduces the probability of having a gain lower than the threshold, which can, for example, correspond to an undesirable result.
[0159] In Figure 2 's example, the present invention allows, for example, obtaining a strategy that reduces the probability of having a low gain and thus reduces the probability of missing the application deadline.
[0160] However, Figure 2 's example is provided only by way of a non - limiting example of a radio resource allocator that can implement the present invention, and the present invention can be applied to other resource allocators in different embodiments.
[0161] Now refer to Figure 3 .
[0162] Figure 3 represents an example of a computer - implemented method P3 according to multiple embodiments of the present invention.
[0163] Method P3 is a method for the system to select an action from the set of possible actions A(s) in the current state s of the set of possible states S of the system at time step t.
[0164] For example, method P3 can be implemented to control a system, such as Figure 1 orFigure 2 The system illustrated in
[0165] Thus, method P3 generally represents a stage called the "inference stage", which is the stage of selecting actions to control the system.
[0166] In multiple embodiments of the present invention, some steps can be implemented within the inference stage and / or in the training stage on which the inference stage depends. The training stage can, for example, constitute determining the path probability or policy applied to the control system.
[0167] The following disclosure can indicate, when applicable, the method steps that can be implemented during the inference or training stage.
[0168] Returning to Figure 3 , method P3 further includes a first step S33: selecting an action in the current state, which action guides the system from the current state with the highest probability along a path corresponding to a series of actions and state transitions in consecutive time steps and having a gain of at least equal to or higher than a predetermined value.
[0169] For example, the first step S33 can include selecting the action that exhibits the highest sum of probabilities of the paths generated from the current state when performing an action that results in a gain of at least equal to or higher than a predetermined value.
[0170] In other words, step S33 does not include selecting the action that leads to the maximum expected gain as in the prior art, but includes minimizing the probability of the path that leads to a gain lower than, or equal to or lower than the defined limit value.
[0171] Method P3 further includes a second step S34 of causing the system to execute the selected action.
[0172] Thus, method P3 causes the system to execute an action that reduces the probability of a path having a gain lower than, or equal to or lower than a predetermined limit value, which predetermined limit value can be referred to as α in the remainder of this disclosure. Since the gain can be defined according to the desirability of the results of the actions performed by the system and the state transitions, this method limits the probability of ultimately leading to an undesirable result. For example, if this method is implemented by a resource allocator, this method can reduce the probability of resource allocation that leads to the generation of highly undesirable results (such as loss of communication limits, loss of computing resources for successful execution of programs, etc.).
[0173] Determining the probability along a path with a gain that is at least equal to or higher than a predetermined value can be performed in different ways. The steps described below can be performed partially or fully during the training phase. As will be explained in more detail below, the gain and probability can be calculated for a predetermined policy or for a policy iteratively defined by selecting the best action in each state during the training and / or inference phase.
[0174] One solution associated with the action probability of guiding the system along a path corresponding to a sequence of actions and state transitions in consecutive time steps and having a gain that is at least equal to or higher than a predetermined value includes: running or simulating the system while performing a sequence of actions, and associating a gain depending on the result of the running or simulation with the path followed by the system during the running or simulation of the system.
[0175] For example, in Figure 2 the example of the scheduler shown, the scheduler can be simulated or run a large number of times, where different actions correspond, for example, to multiple different policies, or to an exhaustive selection of all possible actions in each state that grants radio resources to different users, and then each resulting path followed by the system (i.e., the consecutive states and state transitions of the system) can be associated with:
[0176] - The probability of observing the system following the path from the state at the time of performing the action. Such a probability can be calculated by dividing the number of times the system follows the path from the state at the time the action occurs by the number of times the system performs the action in that state;
[0177] - A gain determined based on the result of the path. For example, in Figure 2 the example, each time an application deadline is missed, a negative reward is added to the gain of the path.
[0178] Another option for determining the probability includes calculating the probability using a transition model before step S33). This option is represented by Figure 3 steps S31 and S32 in
[0179] Although steps S31 and S32 are represented in Figure 3 as preparatory steps of method P3, they can be performed partially or fully by a training method executed before method P3.
[0180] The first preparatory step S31, performed by method P3 or during the training phase, includes: retrieving a transition model and the reward associated with the transition, the transition model defining the probability of performing a transition from an initial state to a subsequent state when performing an action belonging to the set of possible actions to be performed in the initial state for at least one state transition between two consecutive time steps.
[0181] For example, step S31 may include retrieving a model that conforms to the MDP framework, the transition probability p(s j | s, a), and the reward r(s j , s).
[0182] The method P3 or a second preliminary step S32 performed during the training phase includes: obtaining, for a plurality of paths corresponding to a series of actions and state transitions in consecutive time steps, probabilities and gains respectively corresponding to the plurality of paths, the probabilities and gains being calculated by using the model to calculate: the gains of the plurality of paths; the probabilities of the plurality of paths.
[0183] In other words, step S32 allows obtaining, for each of the plurality of paths, the gain of the path and the probability of the path along which it is taken. The gain and probability of the path can be calculated during the execution of method P3 or pre-calculated during the training phase. Examples of such determination are provided in Figure 4 , Figure 5 , Figure 6 and Figure 7 .
[0184] Now refer to Figure 7 .
[0185] Figure 7 represents an example of a plurality of paths in an embodiment of the present invention.
[0186] Figure 7 The example of
[0187] Figure 7 is applicable to but not limited to implementing the method within the MDP framework.
[0188] More specifically,
[0189] - calculating the gains of the plurality of paths as the cumulative rewards of each combination of state transitions and actions of each of the plurality of paths in consecutive time states;
[0190] - calculating the probabilities of the plurality of paths as the combined probabilities of each combination of state transitions and actions of each of the plurality of paths in consecutive time states.
[0191] Figure 7 Specific examples of such calculations are provided.
[0192] In the example of Figure 7 , considering 3 consecutive time steps t, t + 1, and t + 2, and the system can be in 6 different states s 17 , s27 , s 37 , s 47 , s 57 and s 67 .
[0193] Figure 7 Several examples of transitions are shown in Figure 7 which only shows a small subset of the possible actions. For example:
[0194] - When the system is in state s 27 , two actions a 217 and a 227 are possible;
[0195] - When the system is in state s 17 , two actions a 117 and a 127 are possible;
[0196] - When the system is in state s 27 , action a 217 can result in two transitions:
[0197] ο A transition from state s 27 to state s 17 with probability p 217 and reward r = 1;
[0198] ο A transition from state s 27 to state s 47 with probability 1 - p 217 and reward r = 0;
[0199] - When the system is in state s 27 , action a 227 can result in two transitions:
[0200] ο A transition from state s 27 to state s 57 with probability p 227 ;
[0201] ο A transition from state s 27 to state s 67 with probability 1 - p 227 ;
[0202] - When the system is in state s 17 , action a 117 can result in two transitions:
[0203] ο A transition from state s 17 to state s 27 with probability p 117and reward r = 0;
[0204] ο The transition from state s 17 to state s 47 has a probability of 1 - p 117 and reward r = -1;
[0205] - When the system is in state s 17 action a 127 can result in two transitions:
[0206] ο The transition from state s 27 to state s 17 has a probability of p 127 ;
[0207] ο The transition from state s 27 to state s 47 has a probability of 1 - p 127 .
[0208] In the Figure 7 example, with discount λ = 1;
[0209] - The first example of the path 27 generated by state s is characterized by:
[0210] ο The first state transition from state s 217 to state s 27 when performing action a 17 has a probability of p 217 and reward r = 1;
[0211] ο The second state transition from state s 117 to state s 17 when performing action a 27 has a probability of p 117 and reward r = 0;
[0212] ο Probability
[0213] ο Gain
[0214] - The first example of the path 27 generated by state s is characterized by:
[0215] ο The first state transition from state s 217 to state s 27 when performing action a 17 has a probability of p 217 and reward r = 1;
[0216] ο When performing action a 117 from state s 17 to state s 47 the second state transition, with probability 1 - p 117 and reward r = -1;
[0217] ο Probability
[0218] ο Gain
[0219] - Obviously, more examples of paths can be defined, for example, by using the state transition from state s 217 when performing action a 27 to state s 47 using the state transition from state s 227 when performing action a 27 to state s 57 using the state transition from state s 127 when performing action a 17 to state s 67 and so on.
[0220] More generally:
[0221] - The path probability of the k-th path generated from state s can be labeled as P s (k), and the path gain of this path can be labeled as G s (k);
[0222] - The path probability of the k-th path generated from state s when performing action a can be labeled as P s,a (k), and the path gain of this path can be labeled as G s,a (k).
[0223] Figure 7 The examples depicted in are provided by means of simple illustrative and non-limiting examples, and more examples of transitions and paths can be envisioned. For example:
[0224] - An action can result in more than two possible state transitions. For example, an action can result in a single state transition or more than 2 state transitions with a probability equal to 1; The sum of the state transition probabilities from the same state for the same action should be equal to 1;
[0225] - Two different actions may result in the same state transition, possibly with different probabilities and rewards. For example, in Figure 7 the example of 227 action a 27 can also result in the state transition from state s 17 to state s 217Probabilities and rewards that can be different from 1;
[0226] - State transitions can include transitions from a state to itself, e.g., from state s 27 to s 27 ;
[0227] - A path can include more than 2 transitions;
[0228] - The gain and probability of a path can be calculated over a finite or infinite time horizon, i.e., within a finite time window or an infinite time window;
[0229] - In the above example, the gain is calculated as the sum of the rewards. It can also be calculated as a weighted sum, e.g., by applying a discount factor to the power of the transition exponent, as shown in, e.g., Equation 1 above;
[0230] In the remainder of this disclosure, the state at time t from which an action can be taken will be referred to as the "current state", and in a transition, the state from which the transition occurs is referred to as the "initial state", and the state to which the transition occurs is referred to as the "subsequent state". For example, in a transition from state s 27 to state s 17 s 27 is the initial state and s 17 is the final state.
[0231] The gain and probability of a path can be calculated in different ways. For example, the complete combination of all possible paths can be calculated over a time horizon. Alternatively, as will be explained below, the gain and probability can be calculated iteratively using binning rules specifically.
[0232] Thus, an action can be selected, for example, by calculating the following:
[0233]
[0234] where
[0235] - s is the current state;
[0236] - a is an action belonging to the set of actions A s that can be executed in the current state;
[0237] - α is the gain limit;
[0238] - P s,a (k) is the occurrence probability of the k-th possible path generated from the current state s when the action a is executed;
[0239] - 1{G s,a (k)>α} is equal to
[0240] ο1, if the gain G s,a (k) of the k-th possible path generated from the current state s when performing action a is higher than α;
[0241] ο0, otherwise.
[0242] -a * is the selected action, which exhibits a higher value Q(s, a, α) among all possible actions.
[0243] Therefore, at the output of state S33, an action is selected that guides the system along the path corresponding to a series of actions and state transitions in consecutive time steps with a gain of at least equal to or higher than a predetermined value α with the highest probability.
[0244] As described above, according to different embodiments of the present invention, some steps can be performed during the training phase or the inference phase. Therefore, according to different embodiments of the present invention, the output of the training phase and the input of the inference phase can be different.
[0245] For example, in multiple embodiments of the present invention, the training phase finally calculates the gain G s,a (k) and the occurrence probability P s,a (k) of each k-th possible path generated from each state s when performing action a.
[0246] Therefore, in such an example, the inference phase receives all the values G s,a (k) and P s,a (k) as inputs, and calculates the following equations 8 and 9 at step S33 to select an action that guides the system along the path corresponding to a series of actions and state transitions in consecutive time steps with a gain of at least equal to or higher than a predetermined value with the highest probability.
[0247] In other embodiments of the present invention, assuming that the best action is selected at each state, the training phase finally calculates the gain G s (k) and the occurrence probability P s (k) of each k-th possible path generated from each state s.
[0248] Therefore, in such an example, the inference phase receives all the values G s (k) and P s (k) and the transition model as inputs. At step S33, the probability of the path generated from the current state when performing an action can be calculated based on:
[0249] - The probability of performing a transition from the current state to each subsequent state in the transition model;
[0250] - The probability P along each path generated from each subsequent steps (k).
[0251] Meanwhile, the gain associated with each path generated from the current state when performing an action can be calculated based on the following:
[0252] - The reward associated with each transition from the current state to each subsequent state in the transition model;
[0253] - The gain G of each path generated from each subsequent state s (k), a discount factor can be applied.
[0254] Thus, in these embodiments, step S33 includes calculating the value G of the current state s,a (k) and P s,a (k), and then using equations 8 and 9 to select the action to be performed.
[0255] In other embodiments of the present invention, the training phase fully determines the action to be performed in each state. Thus, the output of the training phase is the optimal policy π*, and step S33 includes applying the policy π* to select the action to be performed in the current state. The policy π* can be expressed, for example, in the form of a table associated with each state of the action to be performed.
[0256] It should be noted that, in this case, the policy is predetermined for a given limit value α. Therefore, the use of the policy during the inference phase must be optimized for the given limit value α. In contrast, in embodiments that rely on the use of equation 8 and 9 or similar equations, the parameter α can be set during the inference phase during the inference phase.
[0257] Now refer to Figure 8 .
[0258] Figure 8 A part of the grid representing possible paths in multiple embodiments of the embodiments of the present invention.
[0259] In multiple embodiments of the present invention, the probability and gain of a path are calculated for a path defined by a series of consecutive states.
[0260] In other embodiments of the present invention, each state is associated with a reference set, which includes the reference gain of the paths generated from the state and the reference probability along the paths from the state, and a binning rule is applied to the reference gain of the paths generated from each state to assign the size of the reference set associated with each state, which includes the gain of the paths generated from the state and the probability of the paths to a predetermined size. The reference set, gain, and probability thus define the possible paths resulting from the state by their respective gains and probabilities.
[0261] In other words, paths with similar gains are placed in "bins" that define gain intervals. Thus, in such a representation, a state or a combination of a state and an action is associated with the probability of guiding the system along the paths in each bin (i.e., paths with gains within the defined intervals). For computational purposes, all paths falling within the defined bins can be associated with the same gain, such as the gain at the center of the bin. In this representation, the exact state transitions used by the paths are irrelevant, and the only relevant values are the gains and probabilities of the different bins.
[0262] The bins can be predefined such that the gain intervals forming each bin are known in advance. In this case, the values G s,a (k) and G s (k) are actually the same for all states because the bins are predefined for all states. For example, multiple bins can be predefined by the user around an expected gain from the states, or intervals separated by rules between a minimum gain and a maximum gain can be defined.
[0263] In multiple embodiments of the present invention, the calculation of the gain and probability of a path thus depends on iterative calculations and binning rules that allow maintaining the size of the reference set of gains and probabilities at a predefined size K. The principles of the iterative calculations and binning rules will be further explained below.
[0264] First, at each iteration, a number of paths are generated from each state. Recursive calculations within all possible states and paths may result in too many possible combinations and paths. To limit the number of combinations while maintaining good results, paths with similar probabilities can be merged in order to obtain K paths of state paths. The probabilities of the K paths resulting from a state can be represented as:
[0265] P s =[P s (1),..., P s (K)]
[0266] (Equation 10)
[0267] The associated gains can be represented as:
[0268] G s =[G s (1),..., G s (K)]
[0269] (Equation 11)
[0270] Thus, a state can be represented by the reference set Ω s as:
[0271] Ω s ={Ps , G s}
[0272] (Equation 12)
[0273] Therefore, within the framework of the present invention, the value of state s can be defined as:
[0274]
[0275] In other words, the value of the state can be expressed as the probability that when the system is in the state, it will ultimately follow a path with a gain higher than α.
[0276] It is worth noting that the above symbols do not depend on action a. This is because, in the iterative calculation, the action is automatically associated with each state by applying a predetermined policy or by automatically selecting the action that maximizes the state value.
[0277] In Figure 8 's example, in the initial state s 18 and two possible subsequent states s′ 18 and s′ 28 a grid is represented between. The two possible state transitions are as follows:
[0278] - The transition from state s 18 to state s′ 18 has a probability p and a reward r(s′ 18 );
[0279] - The transition from state s 18 to state s′ 28 has a probability (1 - p) and a reward r(s′ 28 ).
[0280] In the iteration, the reference set 18 of state s′ and the reference set 28 of state s′ will be used to calculate the reference set s 18 . The reference sets and were previously estimated, and the vectors and both have a size of K.
[0281] The reference set 18 of state s can initially be calculated as:
[0282]
[0283] where λ is the discount factor, 0 < λ ≤ 1.
[0284] At this stage, the set has a size equal to 2*K (since there are two successor states; for multiple successor states n, it will generally be n times K).
[0285] To limit the size of the set, the binning rule is then applied to Then the binning rule merges some paths to transform the vector into a vector of size K More generally, at each iteration and for each path, the binning rule can be applied to transform the vector G s into a vector of size K and the vectors of probabilities can be merged accordingly. For example, paths with a gain difference less than ∈ can be merged. Paths can also be merged within a predefined bin, where a bin is an interval of possible gains. The intervals characterizing the bins (e.g., their size and center) can be predetermined. For example, for state s, the bins can be defined such that
[0286] Assume that the binning function has been applied to state s′ 18 and s′ 28 , and thus the vectors of gains associated with these two states are and and the reference gain vector for state s 18 can be calculated by:
[0287]
[0288] Then:
[0289]
[0290] where Bin(G s , x) represents applying the binning rule to vector x.
[0291] To provide a specific example of the binning rule, let us assume r(s′ 18 ) = 2, r(s′ 28 ) = 0, and λ = 0.8. Thus, we get The binning rule can separate the paths with a gain close to 1 (exponent 3 in ) from the 3 paths with a gain close to 3 (exponents 1, 2, and 4 in ), thus obtaining
[0292] A specific example of the iterative calculation of the gains and probabilities of paths will now be presented below.
[0293] Now refer to Figure 4 .
[0294] Figure 4 An iterative method P4 that represents calculating the gain and probability of a path given an input policy π. Thus, method P4 can be executed during the training phase. The reference gain and probability thus obtained can then be used during the inference phase, for example, by method P3.
[0295] Method P4 includes a first step S40: For each state in the set of possible states S, obtain an initial reference set where and both have a size of K.
[0296] Then, a loop S44 is executed for each state s in the set of possible states S. For each state, the following steps are performed for each state.
[0297] In a first sub-step S41, an action a is executed in state s according to the policy π.
[0298] The second sub-step S42 includes redefining the initial reference set of the state, which is done by concatenating the following operations for each subsequent state that has a non-zero transition probability from the state using the selected action: multiplying the reference gain in the reference set of the subsequent state by the discount factor and adding the reward for the transition from the state to the subsequent state, and multiplying the reference probability of the subsequent state by the non-zero transition probability.
[0299] For each state with a non-zero transition probability p(s jk |s, a), the gain can be calculated as:[[]]
[0300]
[0301] Then, the probability is:[[]]
[0302]
[0303] Then, the iteration includes sub-step S43: reducing the size of the reference set of the state to the predetermined size K by applying a binning rule to the reference gain of the state.
[0304] When sub-steps S41 to S43 are executed for a state, sub-step S44 verifies whether they need to be executed for another state. If this is the case, sub-steps S41 to S43 are executed for another state. Otherwise, step S45 is executed to verify whether the stopping criterion is met.
[0305] If the stopping criterion is not met, a new iteration of sub-steps S41 to S43 is performed for all states.
[0306] Otherwise, the method is complete, and the reference gain and probability are calculated for each state.
[0307] Thus, this provides an efficient way to calculate the reference gain and probability for a given policy.
[0308] Of course, method P4 can be repeated for multiple different policies in order to select the policy that allows maximizing the probability of a gain higher than a predetermined value in a given state.
[0309] Now refer to Figure 5 .
[0310] Figure 5 Denote a computer-implemented method P5 for calculating the reference gain and probability.
[0311] Method P5 does not take a predetermined policy as input, but instead iteratively constructs an optimal policy by selecting an optimal action for each state at each iteration.
[0312] Accordingly, sub-step S41 is replaced by sub-steps S51 and S52, while the other steps and sub-steps are similar to those of method P4.
[0313] Step S51 is intended to determine the probability P s,a (k) of having a gain in each bin k in state s when performing action a. To this end, step S51 may include, for each possible action in the state, obtaining the probability of entering each subsequent state when performing the possible action in state s. Such a probability p(s′|s, a) can be retrieved, for example, from the transition model. These probabilities associated with the reward of each transition and the reference set of each subsequent state s’ allow determining the probability P (k) of having a gain in each bin k in state s when performing action a. s,a (k).
[0314] Step S52 includes selecting the action among the possible actions that guides the system from the state along a path corresponding to a sequence of actions and state transitions in consecutive time steps and having a gain of at least equal to or higher than a predetermined value with the highest probability.
[0315] For example, given the probability of transitioning from a state to possible subsequent states and the reference set of subsequent states, the probability that an action results in a gain of at least equal to or higher than a predetermined value α can be calculated:
[0316]
[0317] Then, the action a to be performed can be selected * as:
[0318]
[0319] For the application of the binning rule, the reference probability is thus those obtained by using the selected action a * obtained.
[0320] The value of the state can also be defined as:
[0321] v(s, α) = Q(s, a * , α)
[0322] (Equation 22)
[0323] In practice, this equation is equivalent to Equation 13.
[0324] Therefore, the optimized policy is constructed iteratively.
[0325] For each of methods P4 and P5, different stopping criteria can be used. For example, the stopping criteria can be based on:
[0326] - The maximum number of iterations;
[0327] - The criterion for the gain to converge between two consecutive iterations;
[0328] - And so on.
[0329] Now refer to Figure 6 .
[0330] Figure 6 Represents a third example of a computer-implemented method P6 for calculating reference gain and probability.
[0331] In the Figure 6 example, step S32 of using the transition model also includes using a machine learning model that is trained to predict the path probability from the states of the paths generated by the state to have a gain that decreases in each of a plurality of possible gain intervals.
[0332] Therefore, the machine learning model is trained to predict the probability along the path from a given state assuming that the optimal action is performed at each state.
[0333] For example, an error signal can be used to iteratively train the machine learning model, and the error signal is calculated as a cost function based on the prediction of the path probability from the initial state and the prediction of the path probability from the subsequent state multiplied by the transition probability from the initial state to the subsequent state when performing the action, and the action guides the system from the initial state along the path with a gain of at least equal to or higher than a predetermined value with the highest probability.
[0334] For example, the learning phase may include training a machine learning model to predict, from an initial state s, the probability P s (k) that the system follows a path with a gain within bin k from that state s. In the example of Figure 6 , the probability thus does not represent the probability of following a defined path, but rather the probability of following a path where the gain is within the defined bin. The bins may be predefined. For example, multiple bins may be predefined by the user around an expected gain from a state, or intervals separated by rules between a minimum and a maximum gain may be defined.
[0335] Then, in the inference phase, the probability of a path generated from the current state when an action is executed can be calculated based on:
[0336] - the probability of performing a transition from the current state to each subsequent state in the transition model;
[0337] - the probability P s (k) along each path at each subsequent step.
[0338] Meanwhile, the gain associated with each path generated from the current state when an action is executed can be calculated based on:
[0339] - the reward associated with each transition from the current state to each subsequent state in the transition model;
[0340] - the gain G s (k) along each path for each subsequent step, for which a discount factor may be applied.
[0341] Thus, the inference phase may include calculating the values G s,a (k) and P s,a (k) for the current state, and then using equations 8 and 9 to select the action to be performed.
[0342] In the example of Figure 6 , the training method P6 includes a double loop for training a machine learning model to minimize a cost function:
[0343] Loop S67 loops over multiple states to calculate the values of the cost function for the states.
[0344] Loop S67 may loop through all states or a sub - part thereof, for example, to obtain multiple values of the cost function for multiple states, respectively.
[0345] At the output of loop S67; step S68 includes updating the machine learning model based on the multiple values of the cost function calculated for the multiple states.
[0346] At the output of step S68, step S69 includes verifying whether a stop criterion is met. If the stop criterion S69 is met, the training phase terminates, and the machine learning model can be used for inference in method P3.
[0347] Otherwise, a new loop of steps S61 to S68 is executed until the stop criterion is met.
[0348] The stop criterion can be, for example, with respect to one or more of the number of iterations or a convergence criterion. The stop criterion can include, for example, one or more of the following:
[0349] - Comparison of the number of iterations with a threshold;
[0350] - Comparison of the average value of the cost function at the updated output of step S68;
[0351] - Comparison of the difference in the average value of the cost function between two consecutive time steps;
[0352] - And so on.
[0353] Within loop S67, steps S61 to S66 determine the value of the cost function for the selected state.
[0354] Step S61 includes calculating a prediction of the path probability from the initial state using the machine learning model. Thus, the currently trained machine learning model is applied to the selected state.
[0355] Then, for each possible action that can be performed from the selected state, steps S62 to S64 are repeated.
[0356] Step S62 includes retrieving from the transition model the probability of entering each subsequent state when performing the possible action and the associated reward in the transition model.
[0357] Thus, at the output of step S62, all state transitions to each possible subsequent state, as well as the associated probabilities of the transitions and the rewards for the transitions, are known.
[0358] Step S63 includes using the machine learning model to calculate a prediction of the path probability from each of the subsequent states.
[0359] Thus, at the output of step S63, at the current learning of the model, for each subsequent state, it is known what the probability is of the path from the subsequent state along a path with a gain in each bin.
[0360] Step S64 includes calculating a prediction of the path probability from the initial state when performing the possible action based on the prediction of the path probability from each of the subsequent states and the probability of entering each subsequent state when performing the possible action and the associated reward.
[0361] In step S64, when performing each possible action, the probability of transition, the reward of transition, and the probability and gain of each bin in each subsequent state can thus be used to calculate the probability of a path from the selected state along a path with gain in each bin.
[0362] Step S65 includes selecting, from the prediction of the path probability from the initial state when performing the possible action, an action that guides the system from the initial state along a path with a gain of at least equal to or higher than a predetermined value with the highest probability.
[0363] In other words, at the current training of the model, the best action for this state is selected at step S65. Equations 8 and 9 can be used, for example, in step S65.
[0364] Step S66 includes calculating the value of a cost function based on the prediction of the path probability from the initial state and the prediction of the path probability from the initial state when performing the possible action, the possible action guiding the system from the initial state along a path with a gain of at least equal to or higher than a predetermined value with the highest probability.
[0365] Substantially, the cost function includes the difference between two calculations of estimating path probability:
[0366] - Obtaining one path probability by directly applying a machine learning model to the selected state;
[0367] - Obtaining one path probability by applying the machine learning model to a subsequent state and combining this information with the transition probability and the reward from the transition model.
[0368] If the machine learning model is properly trained, the difference between the two calculations will be low.
[0369] In fact, this means continuously updating and training the machine learning model based on the cost function of the model to provide path probabilities associated with bins consistent with the transition model. Thus, during the inference phase, the machine learning model can be directly used to predict the probability of a path from the current state along each bin. Therefore, the selection of an action at step S33 during the inference phase will actually guide the system along a path corresponding to a series of actions and state transitions in consecutive time steps and having a gain of at least equal to or higher than a predetermined value with the highest probability.
[0370] Accordingly, method P6 provides an efficient way to train a machine learning engine to predict the probability that when performing an action in a state, guiding the system along a path whose gain falls within a defined bin. However, in an embodiment of the present invention, the probability of guiding the system along a path with a defined gain is provided only by way of example of the application of a machine learning model.
[0371] The present disclosure is not limited to the methods, apparatuses, and computational programs described herein, which are merely examples. The present invention encompasses every alternative contemplated by those skilled in the art upon reading this text.
Claims
1. A computer-implemented method for selecting an action from a set of possible actions executed in a system when the system is in a current state within a set of possible states in a time step, the computer-implemented method comprising: - Performing the selection of the action that guides the system from the current state with the highest probability along a path corresponding to a series of actions and state transitions in consecutive time steps and having a gain of at least equal to or higher than a predetermined value; - Causing the system to execute the action.
2. The computer-implemented method according to claim 1, wherein, The selection is based on: - Retrieving a transition model that defines, for at least one state transition from an initial state to a subsequent state between two consecutive time steps, the probability of performing a transition from the initial state to the subsequent state when performing an action belonging to the set of possible actions executed in the initial state, and the reward associated with the transition; - Using at least the transition model to calculate for multiple paths: ○ The gain of the multiple paths; ○ The probability of the multiple paths.
3. The computer-implemented method according to claim 2, wherein - The gain of the multiple paths is calculated as the cumulative reward of the consecutive state transitions and action sequences of each path in the multiple paths in consecutive time states; - The probability of the multiple paths is calculated as the combined probability of each combination of state transitions and actions of each path in the multiple paths in consecutive time states.
4. The computer-implemented method according to claim 3, wherein, Applying a binning rule to the reference gain of the paths generated from each state to make the size of the reference set associated with each state a predetermined size, the reference set including the gain of the paths generated from the state and the probability of the paths.
5. The computer-implemented method according to claim 4, wherein, Using the model to calculate the gain and probability includes: for each state in the set of possible states, performing an iterative calculation of the reference set of the state, the reference set including the reference gain of the paths generated from the state when performing an action, and the reference probability of causing the system to follow the paths generated from the state when performing the action, and applying the binning rule at each iteration of the iterative calculation.
6. The computer-implemented method according to one of claims 5, wherein, The iterative calculation includes: - For each state in the set of possible states, obtaining an initial reference set; - Until a stop criterion is met: ○ For each state in the set of possible states; ■ Selecting an action to be performed in the state according to a policy; ■ For each subsequent state having a non-zero transition probability from the state using the selected action, redefining the initial reference set of the state by concatenating: ● The reference gain in the reference set of the subsequent state multiplied by a discount factor and added to the reward of the transition from the state to the subsequent state; ● The reference probability of the subsequent state multiplied by the non-zero transition probability; ■ Reducing the size of the reference set of the state to the predetermined size by applying the binning rule to the reference gain of the state.
7. The computer-implemented method according to one of claims 5, wherein, The iterative calculation includes: - For each state in the set of possible states, obtaining an initial reference set; - Until a stop criterion is met: ○For each state in the set of possible states: ■For each possible action in the state, obtain the probability of entering each subsequent state when performing the possible action in the state; ■Select an action among the possible actions, the action guiding the system from the state along a path corresponding to a sequence of actions and state transitions in consecutive time steps with a gain at least equal to or higher than a predetermined value with the highest probability; ■For each subsequent state having a non - zero transition probability from the state using the selected action, re - define the initial reference set of the state by concatenating: ●The reference gain in the reference set of the subsequent state multiplied by a discount factor and added to the reward for the transition from the state to the subsequent state; ●The reference probability of the subsequent state multiplied by the non - zero transition probability, ■Reduce the size of the reference set of the state to the predetermined size by applying the binning rule to the reference gain of the state.
8. The computer-implemented method according to claim 2, wherein, The step of using the transition model further includes using a machine learning model, which is trained to predict path probabilities according to the states of the path generated from the state, with gains belonging to each of a plurality of intervals of possible gains.
9. The computer-implemented method according to claim 8, wherein, Iteratively train the machine learning model using an error signal calculated as a cost function, the cost function being based on the prediction of the path probability starting from the initial state and the prediction of the path probability starting from the subsequent state multiplied by the transition probability from the initial state to the subsequent state when performing the action, the action guiding the system from the initial state along a path with a gain at least equal to or higher than a predetermined value with the highest probability.
10. The computer-implemented method according to claim 9, wherein, The iterative calculation includes: -Until a stopping criterion is met: ○Execute a loop that selects a plurality of states from the set of possible states, and for each state in the plurality of states: ■Use the machine learning model to calculate the prediction of the path probability starting from the initial state; ■For each possible action in the state: ●Retrieve from the transition model the probability of entering each subsequent state when performing the possible action and the associated reward in the transition model; ●Use the machine learning model to calculate the prediction of the path probability starting from each subsequent state; ●Calculate the prediction of the path probability starting from the initial state when performing the possible action according to the prediction of the path probability from each subsequent state and the probability of entering each subsequent state when performing the possible action and the associated reward; ■When performing the possible action, select an action that guides the system from the initial state along a path with a gain at least equal to or higher than a predetermined value with the highest probability according to the prediction of the path probability from the initial state. ■ Calculate the value of the cost function based on the prediction of the path probability from the initial state and the prediction of the path probability starting from the initial state when performing the possible action, where the possible action guides the system from the initial state along a path with a gain of at least equal to or higher than a predetermined value with the highest probability; ○ Update the machine learning model based on the multiple values of the cost function calculated for the multiple states.
11. The computer-implemented method according to any one of the preceding claims, the computer-implemented method being implemented by a resource allocator of the system, wherein: - The set of possible actions corresponds to possible resource allocations in the state; - Define the reward according to the result of the resource allocation and state transition of the system.
12. The computer-implemented method according to the preceding claim, the computer-implemented method being implemented by a radio resource scheduler of a radio communication network, wherein: - The set of possible actions corresponds to possible allocations of radio resources to users of the radio communication network; - At least one probability of performing a state transition depends on the probability that a user successfully transmits a data packet through the radio communication network when radio resources are allocated to the user of the radio communication network; - At least one reward depends on compliance with a time limit for transmitting at least one data packet.
13. A computing device, the computing device including at least one processing unit configured to implement the computer-implemented method according to one of claims 1 to 12.
14. A radio resource scheduler of a radio communication network, the radio resource scheduler including at least one processing unit configured to implement the computer-implemented method according to one of claims 1 to 12.
15. A computer software, the computer software including instructions that, when executed by a processor, implement at least a part of the computer-implemented method according to one of claims 1 to 12.