Communication device, communication system, and communication method

The communication device uses risk-averse reinforcement learning to optimize packet scheduling across multiple wireless interfaces, addressing inefficiencies in resource utilization and environmental adaptability, ensuring high-quality and efficient wireless communication.

JP7720587B2Active Publication Date: 2025-08-08NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022020561
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-08-08
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

Existing wireless communication systems face inefficiencies in resource utilization due to fixed redundant wireless interfaces, leading to suboptimal use of resources and inability to adapt to environmental changes, particularly in heterogeneous networks with varying QoS requirements.

Method used

A communication device employing risk-averse reinforcement learning to dynamically determine the number of packets transmitted through multiple wireless interfaces, optimizing packet scheduling and resource allocation to balance communication quality and efficiency.

Benefits of technology

The solution enhances wireless resource utilization efficiency while maintaining desired communication quality, adapting to environmental changes and meeting stringent reliability and latency requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007720587000026
    Figure 0007720587000026
  • Figure 0007720587000027
    Figure 0007720587000027
  • Figure 0007720587000028
    Figure 0007720587000028
Patent Text Reader

Abstract

To provide a technique for achieving both desired communication quality and improvement in radio resource usage efficiency while following changes in the environment.SOLUTION: A communication apparatus that performs wireless communication using a plurality of wireless interfaces includes: the wireless interfaces that transmit packets to a certain device; a reinforcement learning unit that uses risk-averse reinforcement learning to determine the number of packets to be transmitted to the device by the wireless interfaces; and a transmission unit that transmits, to the device, packets whose number is determined by the reinforcement learning unit.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to packet scheduling in wireless communication systems. [Background technology]

[0002] Currently, wireless communication systems have evolved to become heterogeneous networks based on multi-band, multi-access systems. In cellular communications, the fifth generation mobile communications (5G) has been put into practical use, and a wide range of frequencies, from sub-1 GHz to millimeter wave bands, are being used, and we are entering a world where cells of various sizes, from small cells to macrocells, are provided in an overlapping manner.

[0003] Wireless LAN, another typical wireless access system, also uses radio frequencies in the 2.4, 5, and 60 GHz bands, with the use of the 6 GHz band also being considered. Wireless devices such as smartphones generally have interfaces that support both cellular and wireless LAN access, and each interface supports multiple bands. It is becoming common for devices to select a wireless base station to connect to from multiple frequencies and access methods, and it is also becoming common for a single device to use multiple base stations in a unified manner, such as dual connectivity.

[0004] In such a heterogeneous environment, it is effective to control and optimize which base station a terminal selects and which I / F on the entire system in order to make effective use of system resources.

[0005] Furthermore, as 5G advances, the goal is to realize communication functions for ultra-high reliability and ultra-low latency applications, such as uRLLC (Ultra-Reliable and Low Latency Communications), which have not been widely used in conventional wireless communications.

[0006] One of the conventional techniques for achieving high reliability (low packet loss) and low latency is a method of redundantly transmitting the same data over multiple wireless interfaces and multiple bands, and then combining the data on the receiving side (e.g., Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Cisco Parallel Redundancy Protocol Over Wireless https: / / www.cisco.com / c / ja_jp / td / docs / wireless / outdoor_industrial / iw3702 / technote / b_prp_dg.html [Non-patent document 2] Yue Gao, Kry Yik Chau Lui, Pablo Hernandez-Leal, "Robust Risk-Sensitive Reinforcement Learning Agents for Trading Markets," RL4RealLife Workshop in Int. Conf. on Machine Learning (ICML), 2021. Summary of the Invention [Problem to be solved by the invention]

[0008] The technology in Non-Patent Document 1 basically sets fixed redundant wireless I / Fs or bands according to the required QoS level, which may result in the use of more wireless resources than necessary, resulting in poor wireless resource utilization efficiency.Furthermore, it is not possible to flexibly reflect the amount of resources required in response to changes in the environment.

[0009] The present invention has been made in view of the above points, and aims to provide a technology for achieving both desired communication quality and improved utilization efficiency of wireless resources while adapting to changes in the environment. [Means for solving the problem]

[0010] According to the disclosed technology, A communication device that performs wireless communication using a plurality of wireless interfaces, a wireless interface for transmitting packets to a device; and a reinforcement learning unit for determining, using risk-averse reinforcement learning, the number of packets to be transmitted to the device via the wireless interface; a transmitting unit that transmits the number of packets determined by the reinforcement learning unit to the device; A communication device comprising: The reinforcement learning unit learns actions for the states by risk-averse reinforcement learning, where a satisfaction level based on a packet loss rate in each wireless interface is defined as a state, and a combination of wireless interfaces used by each device and the number of packets transmitted in each wireless interface are defined as actions. A communication device is provided. [Effects of the Invention]

[0011] The disclosed technology provides a technology for achieving both desired communication quality and improved utilization efficiency of wireless resources while adapting to environmental changes. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a wireless communication system. [Figure 2] FIG. 1 is a diagram illustrating the configuration of a wireless base station (or a wireless terminal). [Figure 3] FIG. 1 is a diagram illustrating the configuration of a wireless base station (or a wireless terminal). [Figure 4] 10 is a flowchart showing an outline of an operation. [Figure 5] FIG. 1 is a diagram for explaining a system model. [Figure 6] FIG. 1 is a diagram illustrating reinforcement learning. [Figure 7] FIG. 1 illustrates Algorithm 1. [Figure 8] FIG. 2 illustrates an example of a hardware configuration of the apparatus. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, an embodiment of the present invention (the present embodiment) will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment.

[0014] (System configuration example) An example of the configuration of a wireless communication system according to this embodiment is shown in Fig. 1. As shown in Fig. 1, the system includes a wireless base station 100 and a plurality of wireless terminals 200. In the example of Fig. 1, the wireless base station 100 is connected to the Internet.

[0015] In this embodiment, a wireless base station 100 having multiple wireless interfaces determines, for packets to be transmitted to a device (wireless terminal), the wireless interface through which the packets will be transmitted and the number of packets to be transmitted through that wireless interface, using a reinforcement learning technique described later, and then transmits the packets. Note that determining the number of packets may also be called packet scheduling. However, the technique according to this embodiment can also be applied to a wireless terminal 200. The wireless base station and wireless terminal may be collectively called a communication apparatus.

[0016] In addition, in the specific examples described below, two types of wireless interfaces, Sub-6 GHz and mmWave, are described, but the wireless interface is not limited to these. Furthermore, the "wireless interface" may be interpreted as a "frequency." In other words, in this embodiment, in a form in which multiple frequencies are aggregated and used, frequency selection and packet number determination can be achieved by a reinforcement learning technique described below.

[0017] 2 shows an example of the configuration of the radio base station 100. The radio terminal 200 may also have a similar configuration to that shown in FIG.

[0018] As shown in FIG. 2, the wireless base station 100 includes a communication I / F unit 110, a control unit 120, a wireless communication unit 130, and an antenna 101.

[0019] The wireless communication unit 130 includes a scheduler unit 140, a receiving unit 131, a wireless communication signal generation unit 132, and an RF unit 135. The scheduler unit 140 includes a reinforcement learning unit 150, a communication quality measurement unit 141, a total wireless resource allocation calculation unit 142, and an individual wireless resource allocation calculation unit 143. The number of "individual wireless resource allocation calculation units 143, receiving units 131, wireless communication signal generation units 132, RF units 135, and antennas 101" is equal to the number of wireless interfaces. However, any one of the "individual wireless resource allocation calculation units 143, receiving units 131, wireless communication signal generation units 132, RF units 135, and antennas 101" may be shared by multiple interfaces. Furthermore, the "individual wireless resource allocation calculation units 143, receiving units 131, wireless communication signal generation units 132, RF units 135, and antennas 101" may be referred to as a "wireless interface."

[0020] The reinforcement learning unit 150 includes a Q table management unit 151, a state calculation unit 152, a reward calculation unit 153, and a risk evaluation unit 154. The operation of each unit is as follows.

[0021] The communication I / F unit 110 performs communication with, for example, the Internet, etc. The control unit 120 includes, for example, a CPU and a memory, and controls the entire device. The wireless communication unit 130 performs operations related to wireless communication.

[0022] The scheduler unit 140 performs packet scheduling and the like. The receiver unit 131 receives signals from other communication devices (e.g., feedback from a wireless terminal) via an antenna and RF unit. The wireless communication signal generator unit 132 generates signals to be transmitted wirelessly from the data of the packets to be transmitted. The RF unit 135 performs processing such as placing the signals on carrier waves. The scheduler unit 140 can also be realized by a computer and a program, and the program can be recorded on a recording medium or provided via a network.

[0023] The communication quality measurement unit 141 measures communication quality (e.g., packet loss rate) based on, for example, the number of transmitted packets and feedback (e.g., ACK / NACK) from the communication partner. Note that in this embodiment, it is assumed that instantaneous CSI feedback (ACK / NACK, etc.) cannot be obtained from each device, but sporadic CSI feedback is available, and communication quality statistics (e.g., average value across all devices) can be obtained from the sporadic CSI feedback.

[0024] The total radio resource allocation calculation unit 142 determines the amount to be allocated to each radio interface with respect to the total number of packets to be transmitted, based on the behavior determined for each frame by the reinforcement learning unit 150. Furthermore, the individual radio resource allocation calculation unit 143 determines the amount of radio resources corresponding to the number of packets to be transmitted in the relevant radio interface (the radio interface connected to the individual radio resource allocation calculation unit 143), based on the behavior determined for each frame by the reinforcement learning unit 150.

[0025] The wireless base station 100 (or the wireless terminal 200) can also be represented by the configuration shown in Fig. 3. As shown in Fig. 3, the wireless base station 100 has a reinforcement learning unit 10, a transmitter 20, and a receiver 30. The reinforcement learning unit 10 performs the same processing as the reinforcement learning unit 150. The transmitter 20 performs processing related to transmission (e.g., calculation of transmission resource allocation, packet transmission), and the receiver 30 performs processing related to reception (e.g., feedback reception, communication quality calculation).

[0026] (About the reinforcement learning unit 150) In this embodiment, a configuration is adopted in which the radio base station 100 (or the radio terminal 200) aggregates a plurality of radio interfaces (or a plurality of frequencies).

[0027] The scheduler unit 140, which allocates wireless resources to packets transmitted from each wireless interface, is equipped with a reinforcement learning unit 150, which applies reinforcement learning (1) and autonomously learns and performs optimal connections to obtain the desired communication quality. At the same time, a risk-averse learning method (Non-Patent Document 2) (2) based on parallel updating of multiple Q tables (including a single Q table) is used, enabling behavior selection that prioritizes communication reliability.

[0028] Regarding the application of reinforcement learning in (1) above, in this embodiment, the state s(t) is the satisfaction level based on packet loss rate information (detected from ACK feedback) at each wireless interface, and the action a(t) is the combination of wireless interfaces to be used for each device (wireless terminal when the source is a wireless base station) and packet scheduling (number of packets to be transmitted at each wireless interface). In this embodiment, the optimal action a(t) for each device is learned from the state s(t) using risk-averse average Q-learning.

[0029] In the case of uRLLC assumed in this embodiment, instantaneous CSI feedback cannot be used to maintain low latency. In this embodiment, the radio interface selection and packet scheduling method are designed to enable good risk-averse learning even when the instantaneous channel state is unknown.

[0030] Regarding the risk-averse learning method in (2) above, the equations (Equations (11) and (12) described below) that show the concept of the evaluation function that responds to risk (magnitude of variance) in risk-averse learning include a term that quickly responds to the variance (risk) of past reward r, thereby reflecting the decrease in reward for high-risk behavior. The term that reacts to the variance of past reward r and reflects it in the evaluation is the second term (the term with Var) in Equation (12) described below (the equation obtained by Taylor expansion of Equation (11)).

[0031] As will be explained in detail later, in this embodiment, the instantaneous reward reflects the average packet reception success rate across all devices and the penalty due to a risk state (e.g., a state in which QoS targets such as reliability and delay are not achieved).

[0032] In the reinforcement learning unit 150 shown in FIG. 2, the Q table management unit 151 holds, initializes, updates, etc. the Q table. The state calculation unit 152 calculates the state s(t). The reward calculation unit 153 calculates the reward r for s(t) and a(t). The risk assessment unit 154 calculates an evaluation function based on the Q table and selects an action. Note that the calculation of the evaluation function may be performed by the reward calculation unit 153.

[0033] Here, an outline of the operation of the radio base station 100 related to reinforcement learning will be described with reference to the flowchart of FIG.

[0034] In S101, the state calculation unit 152 acquires a satisfaction level based on information about the packet loss rate (detected from ACK feedback) in each wireless interface, and calculates the state s(t).

[0035] In S102, the risk assessment unit 154 determines an action a based on the multiple Q tables (or a single Q table) managed in the Q table management unit 151 using the ε-greedy method.

[0036] In S103, the reinforcement learning unit 150 notifies the total radio resource allocation calculation unit 142, the individual radio resource allocation calculation unit 143, etc. of the determined action a, and the radio base station 100 then executes the action a.

[0037] In S104 , the communication quality measurement unit 141 acquires packet loss information, and passes the packet loss information to the reward calculation unit 153 in the reinforcement learning unit 150 .

[0038] In S105, the reward calculation unit 153 calculates the reward. In S106, the Q table management unit 151 updates the multiple Q tables (or the single Q table).

[0039] Hereinafter, the operation of the radio base station 100 in this embodiment (particularly the operation of the reinforcement learning unit 150) will be described in more detail using an example in which a specific radio interface is used.

[0040] (System Model) In this embodiment, as shown in Fig. 5, downlink (DL) transmission in a wireless network configured with multiple APs accommodating multiple devices will be described as an example. Each AP is assumed to have a Sub-6 GHz and mmWave (millimeter wave) interface. Each AP corresponds to a wireless base station 100. A device corresponds to a wireless terminal 200. In the following description, the wireless base station 100 will be described as performing the reinforcement learning operation according to the present invention, but the wireless terminal 200 can also perform the same operation.

[0041] As shown in Figure 5, AP b transmits a desired packet to a set of devices K. Also, the set of devices K receives DL interference from all other APs b' ≠ b.

[0042] At the beginning of each scheduling frame t, AP b assigns L k Let there be (t) packets, where each packet l∈L k (t) is d bits in size and is sent to device k∈K.

[0043] AP b transmits these packets over N subchannels on the sub-6GHz interface and M beams on the mmWave interface. Each sub-6GHz subchannel or mmWave beam can be assigned to a unique device in each scheduling time frame. Multiple devices can be supported in each frame via different subchannels in sub-6GHz and different beams in mmWave.

[0044] In the sub-6GHz band, the signal-to-interference-plus-noise ratio (SINR) from AP b to device k on subchannel n is:

[0045]

number

[0046] For the mmWave interface, analog beamforming is assumed, and the transmission beam width and beam direction on beam m from AP b to device k are θ bkm and β bkm and is adjusted according to the target device k and time frame t in each beam m.

[0047] For simplicity and without loss of generality, let the receive beam gain G at device k be k Rx is assumed to be fixed. To maximize the rate obtained, θ bkm is set to the narrowest beam width, and β bkm is given by the line-of-sight (LoS) direction from AP b to device k. Therefore, the SINR of beam m at device k served by AP b is given as follows:

[0048]

number

[0049]

number

[0050]

number

[0051] Therefore, the achievable rate for device k accommodated by AP b is:

[0052]

number

[0053] Let l be the number of packets sent to device k in frame t on interface ν. k ν (t)∈{0,…,L k (t)}. L k (t) is the total number of queued packets at frame t, so l k sub (t)+l k mW (t)≦L k (t). The number of packets successfully received by device k on each interface is Ω. k ν (t) can be calculated by AP b based on the ACK feedback of device k as follows:

[0054]

number

[0055]

number

[0056]

number

[0057] where r bk ν Since (t) is unknown in AP, l k,max ν is unknown in the AP. k ν (t)≦l k,max ν If (t), that is, the number of transmitted packets in the assigned subchannel or beam of device k is smaller than the number of packets that device k can receive, we assume that all these packets are successfully received and their ACKs are fed back to the AP. However, k ν (t) ≥ l k,max ν If (t), then l k ν (t)-l k,max ν (t) The packet is in NACK state.

[0058] Based on the above, the packet loss rate (PLR) of device k at frame t is defined as the average packet loss rate across both interfaces up to frame t, as follows:

[0059]

number

[0060]

number

[0061]

number

[0062] (About Markov Decision Processes (MDPs)) The goal here is to maximize the long-term PSR averaged over all devices (here ρ) while satisfying the individual PLR constraints of each device. max ) This problem can be modeled as an MDP characterized by a state space, an action space, transition probabilities, and a reward function, as shown in Figure 6. In Figure 5, the state s t is the PLR satisfaction level (and ACK feedback status) for all devices. t is the interface selection and packet scheduling for all devices. In this embodiment, the reward r(t) is obtained based on the state s(t) and the action a(t), and the action is optimized by maximizing the objective function.

[0063] Each AP (wireless base station) is an agent that makes interface selection and packet scheduling decisions. At each frame t, the AP determines its current state s t Know the state t consists of the current PLR satisfaction levels of devices associated with the AP and their feedback states in the previous frame t-1. t Based on this, AP will take action t That is, the AP determines the number of packets in each interface of each device in the current frame t and obtains the immediate reward r from the environment. t Get the new state s t+1 Transition to.

[0064] Since the information such as the instantaneous CSI and interface statistics is unknown, the AP determines the transition probability P(s t+1 |s t ,a tIn this embodiment, this problem is solved by using a framework of RL (reinforcement learning).

[0065] (Risk-Averse Reinforcement Learning) To best satisfy stringent reliability requirements, we use a Risk-Sensitive Reinforcement Learning (RSRL) approach called Risk-Averse Average Q-learning (RAQL). Compared with traditional RL methods that aim to maximize expected returns, such as QL, RSRL introduces the concept of risk, which is linked to the variance of rewards. RAQL achieves further variance reduction, thereby reducing risk.

[0066] Instead of taking the expected reward as the objective function as in traditional RL, we use the expected utility of the reward as the objective function:

[0067]

number

[0068]

number

[0069] The symbols in the above formulas (11) and (12) have the following meanings:

[0070] J π: The average utility function (immediate reward r) under policy π in a Markov decision process t (discount sum) Π: Policy E π,h : Policy π, expected value under the state h of the wireless channel (propagation path, etc.) r t : Immediate reward value in process t β: parameter Var[]: Variance of [] Order of O():() As will be described later, in this embodiment, multiple Q-tables are simultaneously trained by using Equation (22) as an update rule. The sample variances of these Q-tables are then used as an approximation of the true variance. From these variances, a risk-averse ^Q-table is calculated and used for behavior selection.

[0071] (RAQL-based interface selection and packet scheduling method) Next, the algorithm based on RAQL executed by the AP (wireless base station 100) in this embodiment will be described in detail. The state space and the action space are defined as follows.

[0072] State: s(t) is the current QoS satisfaction level of the PLR for all devices k∈K at frame t and the most recent ACK feedback for packets sent at frame t−1, as shown in Equations (13) and (14) below. ACK feedback may not be included in s(t).

[0073]

number

[0074]

number

[0075] Action: a(t) denotes the interface selection for each device to transmit its packets. To avoid the explosion of the action space size and make the proposed method scalable, we divide the interface selection task and packet scheduling task into three actions a(t) for device k, as explained below. k The AP does not have knowledge of the instantaneous CSI, but it is reasonable to assume that it knows the long-term CSI, such as the average path loss or the average SINR, due to sporadic feedback.

[0076] Therefore, each AP can assign subchannels and beams based on the average CSI of each device. In this case, all subchannels are equivalent for each device, so the AP can randomly select each subchannel to be assigned to each device. The scheduling task of each AP is then equivalent to determining the number of packets to be transmitted per subchannel for each device. The frame length T s The maximum number of packets transmitted from AP b and successfully received by device k during the period can be estimated as shown in equation (15) below.

[0077]

number

[0078] a k (t)=0: Only the Sub-6GHz interface is used, and the number of transmitted packets is

[0079]

number

[0080] a k(t)=1: Only the mmWave interface is used, and the number of transmitted packets is

[0081]

number

[0082] a k (t)=2: Both the Sub-6GHz interface and the mmWave interface are used, but the mmWave interface is prioritized to maximize the number of packets transmitted by taking advantage of its high data rate.

[0083]

number

[0084]

number

[0085]

number

[0086]

number

[0087] The meaning of each symbol in formula (21) is as follows:

[0088] r(s(t),a(t)): immediate reward value in process t Ω k sub (τ): Number of packets successfully sent via Sub6GHZ I / F Ω k mW (τ): Number of packets successfully transmitted via the millimeter wave interface l k sub (τ): Number of packets sent via Sub6GHZ I / F l k mW (τ): Number of packets sent over the millimeter wave interface u k sub (t): Packet loss rate ρ at Sub6GH I / F is required quality ρ max Variables that change depending on whether u k mW (t): Packet loss rate ρ at millimeter wave I / F is the required quality ρ max Variables that change depending on whether As is clear from equation (14), u k ν If (t)=0, i.e., device k is in a risk state that does not satisfy the PLR in equation (14), the reward is penalized.

[0089] The RAQL-based interface selection and packet scheduling method in this embodiment is performed by Algorithm 1 shown in Fig. 7. That is, the radio base station 100 executes this algorithm by, for example, running a program on the CPU. The meanings of the symbols are as follows:

[0090] ε: Search rate λ: Attenuation rate Number of I:Q tables λ p : Risk control parameters Q:Q Table V:Q table update count α: learning rate In Algorithm 1, the AP first initializes I Q-tables along with table V, which counts the number of selections of each action a under state s. The corresponding learning rate α is also initialized to 0, and the algorithm starts from a random state (lines 1-2).

[0091] At each frame t, a Q-table is randomly selected, and the Q-table is used to calculate the risk aversion ^Q-table by Equation (24) described below (lines 3-5). Unlike conventional QL, in RAQL, the Q-function is updated by Equation (22) below.

[0092]

number

[0093]

number

[0094]

number

[0095] Next, given the current state and search rate ε, an action a(t) is selected using the ε-greedy strategy. The AP transmits a packet based on the selected action and receives an immediate reward (Equation (21)) (lines 6-9). The environment then transitions to a new state (lines 10-16). This process is repeated until the maximum number of frames T is reached.

[0096] (Example of hardware configuration) Both the radio base station 100 and the radio terminal 200 can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud. Hereinafter, the radio base station 100 and the radio terminal 200 will be collectively referred to as communication devices.

[0097] That is, the communication device can be realized by using hardware resources such as a CPU and memory built into a computer to execute a program corresponding to the processing performed by the communication device. The program can be recorded on a computer-readable recording medium (such as a portable memory) and stored or distributed. The program can also be provided via a network such as the Internet or email.

[0098] Fig. 8 is a diagram showing an example of the hardware configuration of the computer. The computer in Fig. 8 has a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., which are all connected to each other via a bus BS. Note that the communication device may not have the display device 1006.

[0099] A program for realizing processing on the computer is provided by a recording medium 1001 such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.

[0100] The memory device 1003 reads and stores the program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes functions related to the communication device in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, or the like, and is used to input various operation instructions. The output device 1008 outputs the results of calculations.

[0101] (Effects of the embodiment) The technology according to this embodiment can provide a technology for achieving both desired communication quality and improved utilization efficiency of wireless resources while adapting to changes in the environment.

[0102] (Addendum) This specification discloses at least the following communication devices and communication methods. (Section 1) A communication device that performs wireless communication using a plurality of wireless interfaces, a wireless interface for transmitting packets to a device; and a reinforcement learning unit for determining, using risk-averse reinforcement learning, the number of packets to be transmitted to the device via the wireless interface; a transmitting unit that transmits the number of packets determined by the reinforcement learning unit to the device; A communication device comprising: (Section 2) The reinforcement learning unit learns actions for the states by risk-averse reinforcement learning, where a satisfaction level based on a packet loss rate in each wireless interface is defined as a state, and a combination of wireless interfaces used by each device and the number of packets transmitted in each wireless interface are defined as actions. 2. The communication device according to claim 1. (Section 3) the reinforcement learning unit further includes a receiving unit that receives feedback from a plurality of devices that are packet destinations; The reinforcement learning unit calculates the packet loss rate based on the feedback. 3. The communication device according to claim 2. (Section 4) The reinforcement learning unit calculates an immediate reward based on the average packet reception success rate for all devices and a penalty due to a risk state in which the QoS target value is not achieved, and calculates a policy that maximizes the average utility function using past immediate rewards so as to reflect a decrease in reward for high-risk behavior. 1. A communication device according to claim 1, wherein the communication device is a communication device having a plurality of communication paths. (Section 5) the communication device includes a first wireless interface and a second wireless interface that performs communication at a data rate higher than that of the first wireless interface; The behavior selected by the reinforcement learning unit is one of three behaviors: using only the first wireless interface, using only the second wireless interface, and using the second wireless interface preferentially. 10. A communication device according to claim 1, wherein the communication device is a communication device having a plurality of communication lines. (Section 6) A communication system including the communication device according to any one of claims 1 to 5. (Section 7) A communication method executed by a communication device that performs wireless communication using a plurality of wireless interfaces, a reinforcement learning step of determining a wireless interface for transmitting packets to a certain device and the number of packets to be transmitted to the device via the wireless interface using risk-averse reinforcement learning; a transmitting step of transmitting the number of packets determined by the reinforcement learning step to the device; A communication method comprising:

[0103] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims. [Explanation of symbols]

[0104] 100 wireless base stations 101 Antenna 110 Communication I / F section 120 control section 130 Radio Communication Department 131 Receiving unit 132 Wireless communication signal generator 135 RF section 140 Scheduler 141 Communication Quality Measurement Unit 142 Total radio resource allocation calculation unit 143 Individual radio resource allocation calculation unit 150 Reinforcement Learning Department 151 Q Table Management Department 152 State calculation unit 153 Remuneration Calculation Department 154 Risk Assessment Department 200 Wireless Terminals 1000 Drive Device 1001 Recording media 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input Device 1008 Output Device

Claims

1. A communication device that performs wireless communication using multiple wireless interfaces, a wireless interface for transmitting packets to a certain device, and a reinforcement learning unit for determining the number of packets to be transmitted to the device via the wireless interface using risk-averse reinforcement learning; a transmitting unit that transmits the number of packets determined by the reinforcement learning unit to the device; A communication device comprising: The reinforcement learning unit learns actions for the states by risk-averse reinforcement learning, where a satisfaction level based on a packet loss rate in each wireless interface is defined as a state, and a combination of wireless interfaces used by each device and the number of packets transmitted in each wireless interface are defined as actions. Communication equipment.

2. the reinforcement learning unit further includes a receiving unit that receives feedback from a plurality of devices that are packet destinations; The reinforcement learning unit calculates the packet loss rate based on the feedback. The communication device according to claim 1 .

3. A communication device that performs wireless communication using multiple wireless interfaces, a wireless interface for transmitting packets to a certain device, and a reinforcement learning unit for determining the number of packets to be transmitted to the device via the wireless interface using risk-averse reinforcement learning; a transmitting unit that transmits the number of packets determined by the reinforcement learning unit to the device; A communication device comprising: The reinforcement learning unit calculates an immediate reward based on the average packet reception success rate for all devices and a penalty due to a risk state in which the QoS target value is not achieved, and calculates a policy that maximizes the average utility function using past immediate rewards so as to reflect a decrease in reward for high-risk behavior. Communication equipment.

4. A communication device that performs wireless communication using multiple wireless interfaces, a wireless interface for transmitting packets to a certain device, and a reinforcement learning unit for determining the number of packets to be transmitted to the device via the wireless interface using risk-averse reinforcement learning; a transmitting unit that transmits the number of packets determined by the reinforcement learning unit to the device; A communication device comprising: the communication device includes a first wireless interface and a second wireless interface that performs communication at a data rate higher than that of the first wireless interface; The behavior selected by the reinforcement learning unit is one of three behaviors: using only the first wireless interface, using only the second wireless interface, and using the second wireless interface preferentially. Communication equipment.

5. A communication system including a communication apparatus according to any one of claims 1 to 4 and said device.

6. A communication method executed by a communication device that performs wireless communication using a plurality of wireless interfaces, a reinforcement learning step of determining a wireless interface for transmitting packets to a certain device and the number of packets to be transmitted to the device via the wireless interface using risk-averse reinforcement learning; a transmitting step of transmitting the number of packets determined by the reinforcement learning step to the device; A communication method comprising: In the reinforcement learning step, the communication device learns an action for the state by risk-averse reinforcement learning, where a satisfaction level based on a packet loss rate in each wireless interface is set as a state, and a combination of wireless interfaces used by each device and the number of packets transmitted in each wireless interface are set as actions. Communication method.

7. A communication method executed by a communication device that performs wireless communication using a plurality of wireless interfaces, a reinforcement learning step of determining a wireless interface for transmitting packets to a certain device and the number of packets to be transmitted to the device via the wireless interface using risk-averse reinforcement learning; a transmitting step of transmitting the number of packets determined by the reinforcement learning step to the device; A communication method comprising: In the reinforcement learning step, the communication device calculates an immediate reward based on an average packet reception success rate for all devices and a penalty due to a risk state in which the QoS target value is not achieved, and calculates a policy that maximizes an average utility function using past immediate rewards so as to reflect a decrease in reward for high-risk behavior. Communication method.

8. A communication method executed by a communication device that performs wireless communication using multiple wireless interfaces, comprising: a reinforcement learning step of determining a wireless interface for transmitting packets to a certain device and the number of packets to be transmitted to the device via the wireless interface using risk-averse reinforcement learning; a transmitting step of transmitting the number of packets determined by the reinforcement learning step to the device; A communication method comprising: the communication device includes a first wireless interface and a second wireless interface that performs communication at a data rate higher than that of the first wireless interface; The behavior selected by the reinforcement learning step is one of three behaviors: using only the first wireless interface, using only the second wireless interface, and using the second wireless interface preferentially. Communication method.

Citation Information

Patent Citations

  • Radio communication network and radio apparatus used therefor

    JP2009171353A

  • Reliable device-to-device communication

    WO2021121541A1