Communication systems and communication methods
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON TELEGRAPH & TELEPHONE CORP
- Filing Date
- 2023-04-18
- Publication Date
- 2026-07-31
AI Technical Summary
【0011】 開示の技術によれば、強化学習によりスケジューリングを行う通信装置において、学習における収束の速度、及び精度を向上させることが可能となる。
Smart Images

Figure 0007898113000026 
Figure 0007898113000027 
Figure 0007898113000028
Abstract
Description
[Technical Field]
[0001] This invention relates to packet scheduling in wireless communication systems. [Background technology]
[0002] Currently, wireless communication systems have evolved into heterogeneous networks with multi-band, multi-access systems. In cellular communications, fifth-generation mobile communication (5G) has been put into practical use, utilizing a wide range of frequencies from below 1 GHz to the millimeter wave band, and providing a world where cells of various sizes, from small cells to macrocells, are superimposed.
[0003] Furthermore, Wi-Fi, another representative wireless access system, utilizes the 2.4 / 5 / 60GHz bands, and the use of the 6GHz band is also being considered. Wireless terminals such as smartphones generally have interfaces that support both cellular and Wi-Fi access, and each interface supports multiple bands. It is becoming common for terminals to select a wireless base station to connect to from multiple frequencies and access methods, and dual connectivity, where a single terminal can integrate and utilize multiple base stations, is also being implemented.
[0004] In such a heterogeneous environment, controlling and optimizing which interface a terminal uses and which base station it selects is effective for the efficient use of system resources.
[0005] Furthermore, as an advancement of 5G, the goal is to realize communication functions for ultra-reliable and ultra-low latency applications that have not been widely used in conventional wireless communication, such as uRLLC (Ultra-Reliable and Low Latency Communications).
[0006] As one means of achieving high reliability (low packet loss) and low latency, there is a method, as disclosed in Non-Patent Document 1, that uses reinforcement learning to optimize the network used to send packets with higher reliability. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] THL Dinh, M. Kaneko, K. Kawamura, T. Moriyama and Y. Takatori, "Improving Reliability by Risk-Averse Reinforcement Learning over Sub6GHz / mmWave Integrated Networks," ICC 2022 - IEEE International Conference on Communications, 2022, pp. 3178-3183, doi: 10.1109 / ICC45855.2022.9839175 [Overview of the project] [Problems that the invention aims to solve]
[0008] The technology disclosed in Non-Patent Document 1 makes it possible to optimize the communication line used by a communication device through reinforcement learning. However, when there are multiple communication devices, each communication device cannot obtain information about the surrounding communication devices, which leads to a problem where the convergence of learning takes a long time and the accuracy of learning deteriorates. Note that "communication device" refers to, for example, a wireless base station, a wireless terminal, or both a wireless base station and a wireless terminal.
[0009] The present invention has been made in view of the above points, and aims to provide a technology for improving the speed and accuracy of convergence in learning in a communication device that performs scheduling by reinforcement learning. [Means for solving the problem]
[0010] According to the disclosed technology, multiple communication devices perform scheduling using reinforcement learning. A communication system comprising: and an aggregation device that is communicatively connected to the plurality of communication devices, The aggregation device is The aforementioned multiple communication devices An information gathering unit that collects feedback information from, A reward calculation unit that uses the feedback information to calculate the overall reward for the multiple communication devices, The system includes an information distribution unit that distributes the total reward to the multiple communication devices, Each of the plurality of communication devices performs the reinforcement learning using both the total reward distributed from the aggregation device and the reward calculated by the communication device itself. Communication system It will be provided. [Effects of the Invention]
[0011] According to the disclosed technology, it is possible to improve the speed and accuracy of convergence during learning in a communication device that performs scheduling using reinforcement learning. [Brief explanation of the drawing]
[0012] [Figure 1] This is a diagram showing an example configuration of a wireless communication system. [Figure 2] This is a diagram illustrating the configuration of a wireless base station (or wireless terminal). [Figure 3] This is a diagram illustrating the configuration of a wireless base station (or wireless terminal). [Figure 4] This is a flowchart illustrating the operation overview. [Figure 5] This is a diagram showing the configuration of the aggregation device. [Figure 6] This is a sequence diagram illustrating the operation of the aggregation device. [Figure 7] This figure shows an example of how to calculate the total reward. [Figure 8] This is a diagram used to explain the system model. [Figure 9] This is a diagram to explain reinforcement learning. [Figure 10] This is a diagram of Algorithm 1. [Figure 11] This figure shows an example of the device's hardware configuration. [Modes for carrying out the invention]
[0013] Hereinafter, embodiments of the present invention (this embodiment) will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the embodiments described below.
[0014] (Example system configuration) Figure 1 shows an example of the configuration of a wireless communication system in this embodiment. As shown in Figure 1, the system includes multiple wireless base stations 100, multiple wireless terminals 200, and an aggregation device 300. In the example in Figure 1, the aggregation device 300 is connected to the internet.
[0015] In the example shown in Figure 1, the wireless base station 100 is connected to the aggregation device 300, but the wireless terminal 200 may also be connected to the aggregation device 300. Alternatively, both the wireless base station 100 and the wireless terminal 200 may be connected to the aggregation device 300.
[0016] In this embodiment, a wireless base station 100 equipped with multiple wireless interfaces determines, using a reinforcement learning method described later, which wireless interface to transmit packets to a device (wireless terminal), and the number of packets to transmit through that wireless interface, and then performs the transmission. The determination of the wireless interface and the number of packets may also be called scheduling.
[0017] The method according to this embodiment can also be applied to the wireless terminal 200. The wireless base station and the wireless terminal may be collectively referred to as communication equipment. The operation related to the aggregation device 300 will be described later.
[0018] Furthermore, in the specific examples described later, the wireless interface is explained as being of two types: Sub-6GHz and mmWave, but the wireless interface is not limited to these. Also, "wireless interface" may be interpreted as "frequency." In other words, in this embodiment, in a configuration where multiple frequencies are aggregated and used, the selection of frequencies and the determination of the number of packets can be realized by the reinforcement learning method described later.
[0019] Figure 2 shows an example configuration of the wireless base station 100. The wireless terminal 200 may also have a configuration similar to that shown in Figure 2.
[0020] As shown in Figure 2, the wireless base station 100 includes a communication I / F unit 110, a control unit 120, a wireless communication unit 130, and an antenna 101.
[0021] The wireless communication unit 130 includes a scheduler unit 140, a receiver unit 131, a wireless communication signal generation unit 132, and an RF unit 135. The scheduler unit 140 includes a reinforcement learning unit 150, a communication quality measurement unit 141, an overall wireless resource allocation calculation unit 142, and an individual wireless resource allocation calculation unit 143. The "individual wireless resource allocation calculation unit 143, receiver unit 131, wireless communication signal generation unit 132, RF unit 135, and antenna 101" are provided in quantities equal to the number of wireless interfaces. However, any one of the "individual wireless resource allocation calculation unit 143, receiver unit 131, wireless communication signal generation unit 132, RF unit 135, and antenna 101" may be shared by multiple interfaces. Also, the "individual wireless resource allocation calculation unit 143, receiver unit 131, wireless communication signal generation unit 132, RF unit 135, and antenna 101" may be called a "wireless interface".
[0022] The reinforcement learning unit 150 includes a Q-table management unit 151, a state calculation unit 152, a reward calculation unit 153, and a risk assessment unit 154. The operation of each unit is as follows.
[0023] The communication interface unit 110 communicates with the aggregation device 300. The control unit 120, for example, includes a CPU and memory and controls the entire device. The wireless communication unit 130 performs operations related to wireless communication.
[0024] The scheduler unit 140 performs packet scheduling and other operations. The receiving unit 131 receives signals from other communication devices (e.g., feedback from wireless terminals) via the antenna and RF unit. The wireless communication signal generation unit 132 generates signals to be transmitted wirelessly from the data of the packets to be transmitted. The RF unit 135 performs processing such as mounting the signals onto a carrier wave. The scheduler unit 140 can also be implemented using a computer and a program, and the program can be recorded on a recording medium or provided via a network.
[0025] The communication quality measurement unit 141 measures communication quality (e.g., packet loss rate) based, for example, on the number of transmitted packets and feedback from the communication partner (e.g., ACK / NACK). In this embodiment, it is possible to handle situations where instantaneous CSI feedback (ACK / NACK, etc.) cannot be obtained from each device, but sporadic CSI feedback can be obtained, and statistical values of communication quality (average value across all devices, etc.) can be obtained from the sporadic CSI feedback.
[0026] The overall wireless resource allocation calculation unit 142 determines the amount to allocate to each wireless interface based on the action determined by the reinforcement learning unit 150 for each frame, relative to the total number of packets to be transmitted. The individual wireless resource allocation calculation unit 143 also determines the amount of wireless resources corresponding to the number of packets transmitted on the relevant wireless interface (the wireless interface connected to the individual wireless resource allocation calculation unit 143), based on the action determined by the reinforcement learning unit 150 for each frame.
[0027] The wireless base station 100 (or wireless terminal 200) can also be represented by the configuration shown in Figure 3. As shown in Figure 3, the wireless base station 100 has a reinforcement learning unit 10, a transmitting unit 20, and a receiving unit 30. The reinforcement learning unit 10 performs the same processing as the reinforcement learning unit 150. The transmitting unit 20 performs processing related to transmission (e.g., calculating transmission resource allocation, packet transmission), and the receiving unit 30 performs processing related to reception (e.g., receiving feedback, calculating communication quality).
[0028] (Regarding the reinforcement learning unit 150) In this embodiment, the wireless base station 100 (or wireless terminal 200) employs a configuration that aggregates multiple wireless interfaces (or multiple frequencies).
[0029] By equipping the scheduler unit 140, which allocates wireless resources to transmission packets of each wireless interface, with a reinforcement learning unit 150, reinforcement learning is applied (1), allowing the system to autonomously learn and perform the optimal connection to obtain the desired communication quality. Furthermore, by using a risk-averse learning method (Non-Patent Literature 1) (2) based on parallel updating of multiple Q tables (including a single Q table), it is possible to select actions that prioritize the reliability of communication.
[0030] Regarding the application of reinforcement learning as described in (1) above, in this embodiment, the state s(t) is the Satisfaction Level of each wireless terminal based on packet loss rate information (detected from ACK feedback) for each wireless terminal at each wireless interface, and the action a(t) is the combination of wireless interfaces to be used for each device (wireless terminal if the source is a wireless base station) and packet scheduling (number of packets to be transmitted at each wireless interface). In this embodiment, the optimal action a(t) for each device is learned from the state s(t) using Risk-Averse Average Q-learning.
[0031] In this embodiment, it is assumed that, for example, uRLLC is used. In that case, it is possible that instantaneous CSI feedback cannot be used in order to maintain low latency. In this embodiment, the wireless interface selection and packet scheduling method are designed so that good risk-averse learning can be performed even when the instantaneous channel state is unknown.
[0032] Regarding the Risk-averse learning method described in (2) above, the equations (equations (11) and (12) described later) that show the concept of an evaluation function that responds to Risk (magnitude of variance) in Risk-Averse Learning are modified to reflect the decrease in reward for high-risk behavior by adding a term that responds quickly to the variance (risk) of the cumulative reward sumr. The term that responds to the variance of the cumulative reward sumr and reflects it in the evaluation is the second term (the term containing Var) in equation (12) (the Taylor expansion of equation (11) described later).
[0033] As will be explained in more detail later, in this embodiment, the instantaneous reward reflects the average packet reception success rate across all devices, as well as penalties due to risk conditions (e.g., failure to meet QoS targets such as reliability and latency).
[0034] Furthermore, in this embodiment, the aggregation device 300 calculates the total reward as the reward described above based on feedback from multiple wireless base stations 100 (or multiple wireless terminals 200), and transmits the calculated total reward to the multiple wireless base stations 100 (or multiple wireless terminals 200).
[0035] In the reinforcement learning unit 150 shown in Figure 2, the Q-table management unit 151 maintains, initializes, and updates the Q-table. The state calculation unit 152 calculates the state s(t). The reward calculation unit 153 transmits feedback information to the aggregation device 300 and receives the total reward from the aggregation device 300. The reward calculation unit 153 can also calculate the reward r for s(t) and a(t) itself. The risk assessment unit 154 calculates an evaluation function based on the Q-table and selects an action. The calculation of the evaluation function may be performed by the reward calculation unit 153.
[0036] Here, we will explain the operational overview of the wireless base station 100 related to reinforcement learning with reference to the flowchart in Figure 4.
[0037] In S101, the state calculation unit 152 obtains the Satisfaction Level of each wireless terminal based on packet loss rate information (detected from ACK feedback) for each wireless terminal at each wireless interface, and calculates the state s(t).
[0038] In S102, the risk assessment unit 154 determines action a using the ε-greedy method based on multiple Q tables (or a single Q table) managed by the Q table management unit 151.
[0039] In S103, the reinforcement learning unit 150 notifies the overall wireless resource allocation calculation unit 142, the individual wireless resource allocation calculation unit 143, etc., of the determined action a, and the wireless base station 100 executes action a.
[0040] In S104, the communication quality measurement unit 141 acquires packet loss information, and this packet loss information is passed to the reward calculation unit 153 in the reinforcement learning unit 150.
[0041] In S105, the reward calculation unit 153 transmits feedback information to the aggregation device 300 and obtains the total reward from the aggregation device 300. In S106, the Q table management unit 151 updates multiple Q tables (or a single Q table).
[0042] (Regarding the operation of the aggregation device 300) In the following, as an example, we will describe the case where a wireless base station 100 is connected to an aggregation device 300, as shown in Figure 1. However, even when a wireless terminal 200 is connected to the aggregation device 300, the operation described below (operation using the aggregation device 300) can be applied by replacing the wireless base station 100 with the wireless terminal 200.
[0043] In this embodiment, as shown in Figure 1, an aggregation device 300 capable of communicating with the wireless base station 100 is arranged. Each wireless base station 100 uses the reinforcement learning described above to schedule packet transmission to multiple wireless interfaces.
[0044] Each wireless base station 100 transmits feedback information obtained through reinforcement learning to the aggregation device 300. The aggregation device 300 calculates the overall reward based on the feedback information received from each wireless base station 100 and transmits the calculated overall reward to each wireless base station 100.
[0045] Each wireless base station 100 updates a multi-Q table (or a single Q table) based on the total reward received from the aggregation device 300, and selects an action by referring to the updated multi-Q table (or single Q table).
[0046] (Example configuration of aggregation device 300) Figure 5 shows an example of the configuration of the aggregation device 300. As shown in Figure 5, the aggregation device 300 has a communication I / F unit 310, an information collection unit 320, a reward calculation unit 330, and an information distribution unit 340.
[0047] The communication interface unit 310 performs data communication with each wireless base station 100. The communication method between the aggregation device 300 and the wireless base station 100 may be wireless or wired.
[0048] The information gathering unit 320 collects feedback information from each wireless base station 100 via the communication interface unit 310. The information distribution unit 340 distributes the total reward to each wireless base station 100 via the communication interface unit 310. The reward calculation unit 330 calculates the total reward based on the feedback information collected by the information gathering unit 320.
[0049] (Example of system operation) Next, referring to the sequence diagram in Figure 6, the operation when using the aggregation device 300 in the wireless communication system according to this embodiment will be explained. The sequence shown in Figure 6 is executed at predetermined time intervals (for example, every frame or the execution cycle of reinforcement learning). In reality, there are multiple wireless base stations 100, but Figure 6 shows only one wireless base station 100. The operation shown in Figure 6 is executed for each wireless base station 100.
[0050] Furthermore, while Figure 6 shows, as an example, the operation of a wireless base station 100 communicating with an aggregation device 300, the wireless base station 100 in Figure 6 may be replaced with a wireless terminal 200. In other words, the operation of the wireless terminal 200 communicating with the aggregation device 300 is the same as the operation shown in Figure 6.
[0051] <s201> In S201, each wireless base station 100 performs scheduling using a reinforcement learning method. Specifically, each wireless base station 100 selects the user for each wireless interface and determines the number of packets to transmit for each wireless interface.
[0052] In this reinforcement learning-based scheduling (action selection), multiple Q-tables (or a single Q-table) updated based on the overall reward are used.
[0053] <s202> In S202, each radio base station 100 calculates the risk state, the number of successfully received packets, and the number of transmitted packets based on the feedback (ACK / NACK) from the communication partner (here, the radio terminal 100), and transmits these as feedback information to the aggregation device 300.
[0054] Here, the risk state, the number of successfully received packets, and the number of transmitted packets are represented as follows. Note that the example here assumes the system model (model using Sub-6GHz and mmWave) described later.
[0055] Index representing the risk state: u k ν (t) Number of successfully received packets: Ω k ν (t) Number of transmitted packets: l k ν (t) k represents the terminal (device), ν represents the radio interface (Sub or mW). t indicates the target frame. At this time, u k ν (t) is determined by Equation (14) described later. In Equation (14), ρ is the packet loss rate, and ρ max is the required packet loss rate.
[0056] Each radio base station 100 transmits the above information for each radio terminal and each radio interface as feedback information to the aggregation device 300. The information collection unit 320 of the aggregation device 300 acquires the feedback information transmitted from each radio base station 100.
[0057] <s203> In S203, the reward calculation unit 330 of the aggregation device 300 calculates the overall reward, which is the overall reward for all wireless base stations 100, using the feedback information collected from each wireless base station 100.
[0058] The total reward is calculated, for example, by the formula shown in Figure 7. This formula also assumes the system model described later (a model using Sub-6GHz and mmWave). In the formula in Figure 7, b represents a radio base station.
[0059] As shown in Figure 7, the overall reward r is the average over time t of the sum of the "average packet reception success rate of all devices and the penalty due to the risk state at each radio IF" for each radio base station.
[0060] In the example in Figure 7, the calculation of the average packet reception success rate is divided into cases based on 'a', which represents the action. As will be described later, in this system model example, 'a' is one of the values 0, 1, or 2. If 'a' is anything other than 2, the value shown in A in Figure 7 is used, and if 'a' is 2, the larger of B and C is used.
[0061] <s204> In S204, the information distribution unit 340 of the aggregation device 300 distributes the total reward calculated in S203 to each wireless base station 100.
[0062] <s205> In S205, each wireless base station 100 updates multiple Q tables (or a single Q table) using the total reward received from the aggregation device 300 by the reinforcement learning method of this embodiment.
[0063] <s206> In S206, each wireless base station 100 selects a new state and reflects it in the processing.
[0064] The operation of the wireless base station 100 in this embodiment (particularly the operation by the reinforcement learning unit 150) will be described in more detail below using an example of a specific wireless interface. In the following, in order to make the explanation of the reinforcement learning process in this system model easier to understand, an example of operation when a reward is calculated by a single wireless base station 100 is shown.
[0065] (System Model) In this embodiment, as shown in Figure 8, we will explain using downlink (DL) transmission in a wireless network consisting of multiple APs (Access Points) accommodating multiple devices as an example. Each AP is assumed to be equipped with Sub-6GHz and mmWave (millimeter wave) interfaces. Each AP corresponds to a wireless base station 100. The devices correspond to wireless terminals 200. In the following description, we will assume that the wireless base station 100 performs the reinforcement learning operation according to this embodiment, but the wireless terminal 200 can also perform the same operation.
[0066] As shown in Figure 8, AP b sends the desired packet to the device set K. The device set K also receives DL interference from all other APs b'≠b.
[0067] At the start of each scheduling frame t, AP b is L to each device k∈K. k Assume there are (t) packets. Each packet l∈L k (t) is the size of d bits, which is sent to device k∈K.
[0068] AP b transmits these packets via N subchannels on the Sub-6GHz interface and M beams on the mmWave interface. Each Sub-6GHz subchannel or each mmWave beam can be assigned to a unique device in each scheduling time frame. Multiple devices can be supported in each frame via different subchannels in Sub-6GHz or different beams in mmWave.
[0069] In the Sub-6GHz band, the signal-to-interference + noise ratio (SINR) from AP b to device k in sub-channel n is:
[0070]
number
[0071] For the mmWave interface, analog beamforming is assumed, and the transmitted beam width and beam direction from AP b to device k on beam m are θ, respectively. bkm and β bkm This is expressed as follows, and is adjusted according to the target device k and time frame t in each beam m.
[0072] For the sake of simplification, without loss of generality, the received beam gain G in device k is given. k Rx Assume that is fixed. To maximize the rate obtained, θ bkm It is set to the narrowest beam width, β bkm This is given by the line-of-sight (LoS) direction from AP b to device k. Therefore, the SINR of beam m in device k housed in AP b is given as follows:
[0073]
number
[0074]
number
[0075]
number
[0076] Therefore, the feasible rates for device k housed in AP b are as follows:
[0077]
number
[0078] The number of packets sent to device k in frame t of interface ν is l k ν (t)∈{0,…,L k (t)} is written as L k (t) is the total number of packets queued in frame t, so l k sub (t) + l k mW (t) ≤ L k (t) The number of packets that device k successfully received on each interface is Ω. k ν (t) can be calculated by AP b based on the ACK feedback of device k as follows:
[0079]
number
[0080]
number
[0081]
number
[0082] Here, r bk ν (t) is unknown in AP, so l k,max ν This is unknown in AP. Therefore, l k ν (t)≦l k,max ν If (t) is true, that is, if the number of transmitted packets on the subchannel or beam allocated to device k is less than the number of packets that device k can receive, then it is assumed that all these packets are received successfully and their ACKs are fed back to the AP. However, l k ν (t)≧l k,max ν If (t), then l k ν (t)-l k,max ν (t) The packet will enter a NACK state.
[0083] Based on the above, the PLR (packet loss rate) of device k in frame t is defined as the average packet loss occurrences across both interfaces up to frame t, as follows.
[0084]
number
[0085]
number
[0086]
number
[0087] (Regarding the Markov decision process (MDP)) The goal here is to define the individual PLR constraints of each device (here, ρ max The objective is to maximize the long-term PSR averaged across all devices while satisfying the following conditions: t This represents the PLR satisfaction level (and ACK feedback state) for all devices. Action a t This involves interface selection and packet scheduling for all devices. In this embodiment, the action is optimized by obtaining a reward r(t) based on the state s(t) and action a(t), and maximizing the objective function.
[0088] Each AP (radio base station) is an agent that makes decisions regarding interface selection and packet scheduling. In each frame t, the AP determines the current state s t I know. State s t This consists of the current PLR satisfaction level of the device associated with the AP and their feedback state in the previous frame t-1. t Based on this, AP takes action a t It takes. That is, the AP determines the number of packets in each interface of each device in the current frame t, and obtains the immediate reward r t from the environment, and transitions to a new state s t+1 .
[0089] Since information such as the current CSI and interface statistics is unknown, the AP does not have knowledge of the transition probability P(s t+1 |s t , a t ). In this embodiment, this problem is solved using the framework of RL (reinforcement learning).
[0090] (Risk-Averse Reinforcement Learning) In order to best satisfy the strict reliability requirements, in this embodiment, an approach of RSRL (Risk-Sensitive Reinforcement Learning) called risk-averse average Q-learning (RAQL: Risk-Averse Average Q-learning) is used. Compared with traditional RL methods that aim to maximize the expected return like QL, RSRL introduces the concept of risk, and that risk is linked to the variance of the reward. RAQL achieves a further reduction in variance, thereby reducing the risk.
[0091] Instead of taking the expected reward as the objective function like traditional RL, the expected utility of the following reward is used as the objective function.
[0092]
Equation
[0093]
number
[0094] The meanings of the symbols in equations (11) and (12) above are as follows.
[0095] J π : The average utility function (immediate reward r) with policy π in a Markov decision process. t (Discounted sum) Π: policy (strategy) E π,h : Expected value under policy π and state h of the wireless channel (propagation path, etc.) r t : Immediate reward value in process t β: parameter Var[]: Variance of [] Order O():() As will be described later, in this embodiment, multiple Q tables are learned simultaneously by using equation (22) as the update rule. Then, the sample variances of these Q tables are used as an approximation of the true variance. From this variance, a risk aversion^Q table is calculated and used for action selection.
[0096] (RAQL-based interface selection and packet scheduling method) Next, the RAQL-based algorithm executed by the AP (wireless base station 100) in this embodiment will be described in detail. The state space and action space are defined as follows.
[0097] The state s(t) is the current QoS satisfaction level of the PLR for all devices k∈Κ in frame t, and the most recent ACK feedback for the packet sent to frame t-1, as shown in equations (13) and (14) below. s(t) may be defined as not including ACK feedback.
[0098] [Number] Here,
[0099] [Number] is.
[0100] Action: a(t) indicates the interface selection for which packets of each device should be transmitted. To avoid the explosion of the action space size and make the proposed method scalable, in this embodiment, as described below, the interface selection task and the packet scheduling task are aggregated into three actions a k (t) for device k. The AP does not have the knowledge of the instantaneous CSI, but it is appropriate to assume that the long-term CSI such as the average path loss or the average SINR is known through sporadic feedback.
[0101] Therefore, each AP can perform subchannel and beam allocation based on the average CSI of each device. In this case, all subchannels are equivalent for each device, and thus the AP can randomly select each subchannel to be allocated to each device. And the scheduling task of each AP corresponds to determining the number of packets to be transmitted for each subchannel in each device. Frame length T s The maximum number of packets transmitted from AP b during the period and successfully received by device k can be estimated as in the following equation (15).
[0102] [Number] ~ r bk ν is the known average rate of device k at interface ν. Each action a k (t) is as follows.
[0103] a k (t)=0: Only the Sub-6GHz interface is used, and the number of transmitted packets is
[0104]
number
[0105] a k (t)=1: Only the mmWave interface is used, and the number of transmitted packets is
[0106]
number
[0107] a k (t)=2: Both the Sub-6GHz interface and the mmWave interface are used, but mmWave is given higher priority to maximize the number of transmitted packets by utilizing the high data rate.
[0108]
number
[0109]
number
[0110]
number
[0111]
number
[0112] The meanings of each symbol in equation (21) are as follows:
[0113] r(s(t),a(t)): Immediate reward value in process t Ω k sub (τ): Number of packets successfully transmitted via the Sub6GHz interface Ω k mW (τ): Number of packets successfully transmitted via millimeter-wave interface. l k sub (τ): Number of packets transmitted via Sub6GHz interface l k mW (τ): Number of packets transmitted via millimeter-wave interface u k sub (t): The packet loss rate ρ at the Sub6GHz interface is equal to the required quality ρ max Variables that change depending on whether they reach a certain level. u k mW (t): Packet loss rate ρ at millimeter-wave interface is equal to the required quality ρ max Variables that change depending on whether they reach a certain level. As is clear from equation (14), u k ν If (t)=0, that is, if device k is in a risk state that does not satisfy the PLR in equation (14), a penalty is imposed on the reward.
[0114] Furthermore, when using the aggregation device 300, as previously explained, the aggregation device 300 calculates the total reward shown in Figure 7 based on the feedback information from each wireless base station 100 and distributes it to each wireless base station 100.
[0115] In this embodiment, the RAQL-based interface selection and packet scheduling method is executed by algorithm 1 shown in Figure 10. That is, the wireless base station 100 executes this algorithm, for example, by running a program on its CPU. The meaning of each symbol is as follows.
[0116] ε: Search rate λ: Attenuation rate I: Number of Q tables λ p : Risk control parameters Q:Q Table V: Number of Q table updates α: learning rate In Algorithm 1, AP first initializes I Q tables along with table V, which counts the number of choices for each action a under state s. The corresponding learning rate α is also initialized to 0, and the algorithm starts from a random state (lines 1-2).
[0117] In each frame t, a Q table is randomly selected and used to calculate the risk avoidance^Q table using equation (24) described later (rows 3-5). Unlike conventional QL, RAQL updates the Q function using equation (22) below.
[0118]
number
[0119]
number
[0120]
number
[0121] Next, given the current state and the search rate ε, an action a(t) is selected using the ε-greedy strategy. The AP sends a packet based on the selected action and receives an immediate reward (equation (21)) (lines 6-9). Then, the environment transitions to a new state (lines 10-16). This process is repeated until the maximum number of frames T is reached.
[0122] When using the aggregation device 300, the overall reward shown in Figure 7 is used as the reward in algorithm 1 of Figure 10. Alternatively, both the overall reward and the reward for the wireless base station 100 alone, calculated using equation (21), may be used as the reward.
[0123] (Example hardware configuration) The aggregation device 300, the wireless base station 100, and the wireless terminal 200 can all be implemented, for example, by having a computer run a program. This computer may be a physical computer or a virtual machine on the cloud. Hereinafter, the aggregation device 300, the wireless base station 100, and the wireless terminal 200 will be collectively referred to as the device.
[0124] In other words, the device can be realized by using hardware resources such as the CPU and memory built into a computer to execute a program corresponding to the processing performed by the device. The program can be recorded on a computer-readable recording medium (such as portable memory), saved, and distributed. It can also be provided via a network, such as the Internet or email.
[0125] Figure 11 shows an example of the hardware configuration of the computer described above. The computer in Figure 11 has a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., all of which are interconnected by a bus BS. Note that the communication device may not include a display device 1006.
[0126] The program that enables processing on the computer is provided, for example, on a recording medium 1001 such as a CD-ROM or memory card. When the recording medium 1001 containing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001; it may also be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files and data.
[0127] The memory device 1003 reads and stores a program from the auxiliary storage device 1002 when a program startup command is received. The CPU 1004 implements the functions related to the memory device 1003 according to the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network. The display device 1006 displays a GUI (Graphical User Interface) or the like, run by a program. The input device 1007 consists of a keyboard and mouse, buttons, or a touch panel, and is used to input various operation commands. The output device 1008 outputs the calculation results.
[0128] (Effects of the embodiment) The technology according to this embodiment makes it possible to improve the convergence speed and accuracy of learning in communication devices when there are multiple communication devices that perform scheduling by reinforcement learning.
[0129] (Note 1) This specification discloses at least the following communication devices and communication methods: (Section 1) A communication device that performs wireless communication using multiple wireless interfaces, A wireless interface that transmits packets to a certain device, and a reinforcement learning unit that determines the number of packets to transmit to the device via the wireless interface using risk-avoidance reinforcement learning, A transmitting unit that transmits a number of packets determined by the reinforcement learning unit to the device. A communication device equipped with the following features. (Section 2) The reinforcement learning unit learns actions for each state through risk-avoidance type reinforcement learning, where the satisfaction level based on the packet loss rate of each wireless terminal at each wireless interface is used as the state, and the combination of wireless interfaces used by each device and the number of packets transmitted at each wireless interface are used as the actions. The communication device described in paragraph 1. (Section 3) The reinforcement learning unit further includes a receiving unit that receives feedback from multiple devices that are destinations for packets, The reinforcement learning unit calculates the packet loss rate based on the feedback. The communication device described in paragraph 2. (Section 4) The reinforcement learning unit calculates an immediate reward based on the average packet reception success rate for all devices and the penalty for the risk state where the QoS target value is not met. Using past immediate rewards, it calculates a policy that maximizes the average utility function so as to reflect the decrease in rewards for high-risk behaviors. A communication device as described in any one of paragraphs 1 through 3. (Section 5) The communication device comprises a first wireless interface and a second wireless interface that performs communication at a higher data rate than the first wireless interface. The action selected by the reinforcement learning unit is one of three actions: using only the first wireless interface, using only the second wireless interface, or preferentially using the second wireless interface. A communication device as described in any one of paragraphs 1 through 4. (Section 6) A communication device as described in any one of paragraphs 1 through 5, and a communication system including the said device. (Section 7) A communication method performed by a communication device that performs wireless communication using multiple wireless interfaces, A wireless interface for transmitting packets to a certain device, and a reinforcement learning step for determining the number of packets to transmit to the device via the wireless interface using risk-avoidance reinforcement learning. A transmission step which involves sending the number of packets determined by the reinforcement learning step to the device. A communication method that includes the following features.
[0130] (Note 2) Furthermore, this specification discloses the following aggregation devices, communication systems, communication methods, and storage media. (Additional note 1) An information collection unit that collects feedback information from multiple communication devices that perform scheduling using reinforcement learning, A reward calculation unit that uses the feedback information to calculate the overall reward for the multiple communication devices, An information distribution unit distributes the aforementioned total reward to the aforementioned multiple communication devices. A consolidation device equipped with the following features. (Additional note 2) The feedback information includes an indicator representing the risk status, the number of successfully received packets, and the number of transmitted packets. The aggregation device described in Appendix 1. (Additional note 3) The reward calculation unit calculates the overall reward by calculating the sum of the average packet reception success rate of all devices and the penalty due to the risk status at each wireless interface for the multiple communication devices. A consolidation device as described in Appendix 1 or 2. (Additional note 4) A communication system including the aggregation device described in any one of the appendices 1 to 3, and the plurality of communication devices. (Additional note 5) A method of communication performed by a computer, An information gathering step that collects feedback information from multiple communication devices that perform scheduling using reinforcement learning, A reward calculation step that uses the feedback information to calculate the total reward for the multiple communication devices, An information distribution step which distributes the total reward to the multiple communication devices. A communication method that includes the following features. (Additional note 6) A non-temporary storage medium storing a program for causing a computer to function as a component in any one of the aggregation devices described in any one of the appendices 1 to 3.
[0131] Although this embodiment has been described above, the present invention is not limited to this specific embodiment, and various modifications and changes are possible within the scope of the gist of the invention as described in the claims. [Explanation of Symbols]
[0132] 100 wireless base stations 101 Antenna 110 Communication I / F section 120 Control Unit 130 Wireless Communication Section 131 Receiving Unit 132 Wireless communication signal generation unit 135 RF section 140 Scheduler Section 141 Communication Quality Measurement Unit 142 Overall Wireless Resource Allocation Calculation Unit 143 Individual Wireless Resource Allocation Calculation Unit 150 Reinforcement Learning Department 151 Q Table Management Department 152 State Calculation Unit 153 Compensation Calculation Department 154 Risk Assessment Department 200 wireless terminals 300 Aggregation device 310 Communication I / F section 320 Information Gathering Department 330 Remuneration Calculation Department 340 Information Distribution Department 1000 drive unit 1001 Recording media 1002 Auxiliary storage 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device
Claims
1. A communication system comprising: a plurality of communication devices that perform scheduling using reinforcement learning; and an aggregation device that is communicatively connected to the plurality of communication devices, The aggregation device is An information collection unit that collects feedback information from the aforementioned multiple communication devices, A reward calculation unit that uses the feedback information to calculate the overall reward for the multiple communication devices, The system includes an information distribution unit that distributes the total reward to the multiple communication devices, Each of the plurality of communication devices performs the reinforcement learning using both the total reward distributed from the aggregation device and the reward calculated by the communication device itself. Communication system.
2. The feedback information includes an indicator representing the risk status, the number of successfully received packets, and the number of transmitted packets. The communication system according to claim 1.
3. The reward calculation unit calculates the overall reward by calculating the sum of the average packet reception success rate of all devices and the penalty due to the risk status at each wireless interface for the multiple communication devices. A communication system as described in claim 1.
4. A communication method in a communication system comprising a plurality of communication devices that perform scheduling using reinforcement learning, and an aggregation device that is communicatively connected to the plurality of communication devices, The aggregation device includes an information gathering step of collecting feedback information from the plurality of communication devices, The aggregation device performs a reward calculation step in which it calculates the overall reward for the plurality of communication devices using the feedback information, The aggregation device includes an information distribution step of distributing the total reward to the plurality of communication devices, Each of the plurality of communication devices performs the reinforcement learning using both the total reward distributed from the aggregation device and the reward calculated by the communication device itself. Communication method.