Method for resource control in wireless communication - Patents.com
The reinforcement learning-based resource control method addresses the limitations of conventional machine learning by using delay times as rewards to adaptively adjust wireless communication parameters, significantly reducing delay times and enhancing communication efficiency in dynamic environments.
Patent Information
- Application Number
- JP2021200086
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-09
AI Technical Summary
Conventional machine learning-based radio resource control in wireless communication systems fails to consider delay times, leading to potential selection of communication control parameters that result in large delays, and is limited to pre-learned scenarios, making it ineffective in dynamic environments.
A reinforcement learning-based resource control method that calculates states and rewards based on delay times, allowing for adaptive adjustment of parameters such as MCS and CW values to minimize delay times, and integrates reinforcement learning for both MCS and CW control.
The proposed solution effectively reduces large delay times during control communication by learning optimal parameter settings based on real-time delay feedback, improving communication efficiency and adaptability in dynamic environments.
Smart Images

Figure 0007675633000003 
Figure 0007675633000004 
Figure 0007675633000005
Abstract
Description
[Technical field]
[0001] An embodiment of the present invention relates to a method for resource control in wireless communication. By law Regarding. [Background technology]
[0002] With the advent of the Internet of Things (IoT) era, there are an increasing number of cases where cloud computing systems and edge computing systems are being introduced in offices, factories, etc. Various objects in an office or factory function as stations (STAs) and perform wireless communication with base stations (access points [APs]) installed in the office or factory.
[0003] For example, if the station is a device carried by an employee, or if an obstacle accidentally appears between the station and the access point, wireless communication that was successful up until that point may suddenly fail. Also, if CSMA / CA (Carrier Sense Multiple Access with Collision Avoidance) is used as the access method, packet collisions, known as collisions, may occur in many-to-one wireless communication between stations and access points.
[0004] For this reason, for example, in each station that transmits packets to an access point, resource control that adaptively determines parameters for wireless communication is important. For this reason, the use of machine learning for resource control in wireless communication (radio resource control) is also being considered. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2021 / 117373 Summary of the Invention [Problem to be solved by the invention]
[0006] For example, in Patent Document 1, pre-learning is performed using communication environment information to determine communication control parameters. Here, the communication control parameters are basically MIMO (Multiple-Input and Multiple-Output) precoding parameters. In order to know the communication environment information, a reference signal or the like is required, and when Massive MIMO or the like is taken into consideration, the overhead becomes enormous. When it is determined that the communication environment is one in which changes are easy to predict, the overhead is reduced by performing MIMO precoding using the machine learning results.
[0007] However, while this method takes into consideration overhead reduction, it does not consider delay time, which is important when communicating control signals, etc. Here, the delay time refers to the time from when a packet transmission is started until an ACK is returned, or the time until a timeout occurs without an ACK being returned. In addition, since this method only considers utilizing pre-learned content when using machine learning, if the pre-learned situation deviates from the current situation, the expected effect cannot be obtained.
[0008] As described above, in the conventional method of using machine learning for radio resource control, since the delay time is not taken into consideration, there is a possibility that communication control parameters that will result in a large delay time are selected. Furthermore, since the conventional method is limited to a method of utilizing the results of pre-learning, there is a high possibility that desired characteristics will not be obtained for situations that deviate from the pre-learned situation.
[0009] The problem to be solved by the present invention is to provide a method for resource control in wireless communication that uses reinforcement learning using delay time as a reward for resource control, thereby suppressing large delay times that become a serious problem during control communication. Law The purpose is to provide. [Means for solving the problem]
[0010] According to an embodiment, a method of resource control in wireless communication calculates a state for reinforcement learning to be used for resource control based on a parameter for resource control applied at the time of transmitting the most recent packet and a communication environment observed before transmitting a packet to be transmitted next, including a communication environment observed with the transmission of the most recent packet. The method calculates a reward for reinforcement learning based on a delay time, which is a time from when the transmission of a packet is started until an ACK is returned or a timeout occurs without an ACK being returned, and whether or not the transmission of the packet is successful. The method performs reinforcement learning to select an action using the calculated state and reward, with the action option being what adjustment to make to the parameter applied at the time of transmitting the most recent packet for the transmission of the packet to be transmitted next. The method determines a parameter to be applied to the transmission of the packet to be transmitted next based on the result of reinforcement learning. The parameters include an MCS value and a CW value. The method calculates one state and a reward for the MCS value and the CW value, defines an action option that combines an action for the MCS value and an action for the CW value, and performs reinforcement learning for both the MCS value and the CW value in a joint manner. [Brief description of the drawings]
[0011] [Figure 1] FIG. 2 is a block diagram showing an example of a configuration related to radio resource control of the communication device according to the first embodiment. [Diagram 2] FIG. 2 is a diagram showing an example of a wireless communication environment of the communication device according to the first embodiment. [Diagram 3] FIG. 2 is a diagram for explaining an example of radio resource control of the communication device according to the first embodiment. [Figure 4] 4 is a diagram for explaining a delay time used as a reward for reinforcement learning by the communication device according to the first embodiment. [Figure 5A] FIG. 4 is a diagram for explaining a first example of a reward definition for the communication device according to the first embodiment. [Figure 5B] FIG. 2 is a diagram showing an example of a reward definition that may be included in a first example of a reward definition for a communication device according to the first embodiment. [Figure 6A] FIG. 11 is a first diagram for explaining a second example of a reward definition for the communication device of the first embodiment. [Figure 6B]FIG. 2 is a second diagram for explaining a second example of a reward definition for the communication device of the first embodiment. [Figure 7] FIG. 11 is a diagram for explaining a third example of a reward definition for the communication device according to the first embodiment. [Figure 8] FIG. 13 is a diagram showing a derived example of the third example of the reward definition of the communication device in the first embodiment. [Figure 9] 5 is a flowchart showing the flow of operations related to radio resources of the communication device of the first embodiment. [Figure 10] FIG. 11 is a block diagram showing an example of a configuration related to radio resource control of a communication device according to a second embodiment. [Figure 11] FIG. 1 is a diagram showing an example of an MCS. [Figure 12] FIG. 11 is a block diagram showing an example of a configuration related to radio resource control of a communication device according to a third embodiment. [Figure 13] FIG. 1 is a diagram showing an example of general CW control. [Figure 14] FIG. 13 is a block diagram showing an example of a configuration related to radio resource control of a communication device according to a fourth embodiment. [Figure 15] FIG. 13 is a block diagram showing an example of a configuration related to radio resource control of a communication device according to a fifth embodiment. [Figure 16] FIG. 13 is a block diagram showing an example of a configuration related to radio resource control of a communication device according to a sixth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, embodiments will be described with reference to the drawings. (First embodiment) First, the first embodiment will be described.
[0013] FIG. 1 is a block diagram showing an example of a configuration related to radio resource control of a communication device 1 according to the first embodiment.
[0014] The communication device 1 performs wireless communication conforming to the IEEE 802.11ac standard as a station with an access point. FIG. 2 shows an example of an environment in which the communication device (STA) 1 performs wireless communication with a base station (AP) 2.
[0015] The example shown in FIG. 2 is assumed to be a factory where products are manufactured by robots. The communication device 1 may be a communication device mounted on a robot, a communication device attached to a product being manufactured, a communication device mounted on an electronic device for a worker to check the operating status of the robot, a communication device mounted on a server that centrally controls the robot, or a communication device mounted on a camera for monitoring the inside of the factory. Here, it is assumed that these communication devices 1 perform wireless communication with the same base station 2. The communication devices 1 attempt to start transmitting packets to the base station 2 at timings determined by each of them, without communicating with other communication devices. In addition, the base station 2 that deals with multiple communication devices 1 does not perform control such as allocating the right to transmit packets to the multiple communication devices 1.
[0016] FIG. 3 is a diagram for explaining an example of radio resource control that can be performed in the communication device 1 of this embodiment that performs radio communication with the base station 2 in the environment shown in FIG.
[0017] The communication device 1 may perform MCS (Modulation and Coding Scheme) control as radio resource control (a1). The MCS value indicates a combination of a modulation scheme and a coding rate. For example, when the distance between the station and the access point is short and errors that cause uncorrectable data corruption are unlikely to occur, the communication time can be shortened by applying an MCS value with a high communication rate. On the other hand, when the distance between the station and the access point is long and errors are likely to occur, the occurrence of errors can be suppressed by applying an MCS value with a low communication rate.
[0018] Furthermore, the communication device 1 may perform a CW (Contention Window) control as radio resource control (a2). Each station can detect whether or not another station is transmitting a packet to the access point. While a station is transmitting a packet to the access point, another station that is about to transmit a packet to the access point waits for the end of the currently ongoing packet transmission. When the transmission of this packet ends, if multiple stations that have been waiting transmit packets to the access point at the same time, a collision of packets called a collision occurs. To prevent collision, a station randomly selects a backoff value indicating a waiting time from a predetermined range, and waits for the waiting time indicated by the selected backoff value. If no other station transmits a packet to the access point during this waiting, the station transmits a packet to the access point after waiting. In other words, the station that selects the shortest backoff value among the multiple stations acquires the right to transmit a packet to the access point. The predetermined range from which the backoff value can be selected is the CW. If another station transmits a packet to the access point during standby, the station stops counting the waiting time indicated by the selected backoff value, and resumes counting when the other station finishes transmitting the packet to the access point.
[0019] However, if two stations happen to select the same backoff value, a collision will occur between the two stations. In this case, to reduce the probability that the same backoff value will be selected again between the two stations, the two stations will generally double the size of their CWs. For example, if a collision occurs three times, the size of the CW will be 8 (2 3 ) and returns to its initial value when a collision is avoided.
[0020] By the way, a station that has sent a packet to an access point detects the occurrence of a collision when the access point does not return an ACK. Therefore, the station may mistakenly recognize the absence of an ACK as a collision when the ACK is not returned due to other factors, such as the MCS being too high, causing uncorrectable data corruption and the access point being unable to recognize the source station. In such a case, doubling the size of the CW is a complete waste of time and could only bring about adverse effects (increased delay time).
[0021] Therefore, when an error occurs in which an ACK is not returned, the communication device 1 of the present embodiment determines whether or not to double the CW size as CW control, rather than automatically doubling the CW size as in the conventional method.
[0022] The communication device 1 of the present embodiment executes radio resource control, which may include the above-mentioned MCS control and CW control, by using reinforcement learning. Reinforcement learning is a method of learning an action for obtaining the highest reward in a certain state by considering a model in which a certain reward is obtained when a certain action is taken in a certain state, and repeating trials of observing the reward obtained by a combination of a state and an action in this model. Moreover, the communication device 1 of the present embodiment uses a delay time, which is important when communicating a control signal, as a reward for this reinforcement learning. The delay time in the communication device 1 of the present embodiment will be described with reference to FIG. 4.
[0023] In FIG. 4, it is assumed that stations STA-A and STA-B transmit packets to the access point AP every 100 milliseconds.
[0024] Assume that at time t1, station STA-A and station STA-B attempt to start transmitting packets to the access point AP. Also assume that at this time, station STA-A and station STA-B happen to select the same backoff value. As a result, a collision occurs at time t2. Station STA-A and station STA-B detect the occurrence of a collision due to an ACK timeout at time t3. Station STA-A and station STA-B observe the time (time t1-t3) from when they attempt to start transmitting packets to the access point AP to the ACK timeout as the delay time (delay time A-1, B-1).
[0025] Stations STA-A and STA-B, which caused the collision, reselect backoff values, for example by doubling the CW size, and wait for the waiting time indicated by this backoff value after a predetermined time has elapsed. Here, it is assumed that station STA-A has selected a backoff value smaller than station STA-B. As a result, station STA-A starts transmitting packets to access point AP at time t4. Meanwhile, station STA-B stops counting its waiting time.
[0026] The collision is avoided, the packet from station STA-A is received by access point AP, and at time t5, an ACK for this packet is returned from access point AP to station STA-A. Station STA-A measures the time (time t1-t5) from when it starts transmitting a packet to access point AP until the ACK is returned as the delay time (delay time A-2).
[0027] When station STA-B detects this ACK return, it restarts counting the waiting time that it had stopped after a predetermined time has elapsed, and starts transmitting a packet to access point AP at time t6. Also, at time t7, an ACK for this packet is returned from access point AP to station STA-B. Like station STA-A, station STA-B measures the time (time t1-t7) from when it tries to start transmitting a packet to access point AP until the ACK is returned as the delay time (delay time B-2).
[0028] By including the time during which other stations are communicating in the delay time, such as delay time B-2 shown in Figure 4, the influence from other stations and the impact on other stations are taken into account in reinforcement learning in which the delay time is used as a reward.
[0029] In addition, by providing rewards for learning that reduces delay time, the overall delay time can be reduced even though stations (communication devices 1) each learn independently without centralized control by an access point, for example.
[0030] The communication device 1 of the present embodiment, which uses reinforcement learning using delay time as a reward for radio resource control, is not limited to a station that transmits packets to an access point. The communication device 1 may be any communication device that transmits packets.
[0031] Based on the above, returning to FIG. 1, an example of the configuration relating to radio resource control of the communication device 1 of this embodiment will be described.
[0032] 1, the communication device 1 has a communication unit 11, a communication environment observation unit 12, a delay time measurement unit 13, a state calculation unit 14, a reward calculation unit 15, a learning unit 16, and a parameter (behavior) determination unit 17. Each of these units may be realized by a processor executing a program, or may be realized as hardware such as an electronic circuit.
[0033] The communication unit 11 executes transmission of packets, reception of ACK, reception of a synchronization signal and a reference signal (for example, broadcast from an access point), etc. The communication unit 11 can also execute reception of packets and transmission of ACK.
[0034] The communication environment observation unit 12 observes the communication environment with the communication partner (here, an access point) from the ACK, synchronization signal, reference signal, etc. received by the communication unit 11. The communication environment information obtained by the communication environment observation unit 12 includes a CRC (Cyclic Redundancy Check), an SNR (Signal to Noise Ratio), an EVM (Error Vector Magnitude), communication success / failure, propagation path response characteristics, etc.
[0035] The delay time measurement unit 13 measures the time from when the communication unit 11 attempts to start transmitting a packet (not from when it starts; not time t2, time t4, time t6, etc. in FIG. 4, but time t1) until the communication unit 11 receives an ACK. The delay time measurement unit 13 also measures the time from when the communication unit 11 attempts to start transmitting a packet until the ACK timeout. In other words, the delay time measurement unit 13 measures the delay time.
[0036] The state calculation unit 14 extracts and holds information to be used as a state for reinforcement learning from the communication environment information output from the communication environment observation unit 12. The state calculation unit 14 may be a storage unit that stores only the latest communication environment information itself output from the communication environment observation unit 12. The state calculation unit 14 also holds parameters (wireless communication parameters 100) determined by a parameter (action) determination unit 17, which will be described later. These parameters are parameters for wireless resource control, including an MCS value, a CW value, and the like, which are applied by the communication unit 11 when transmitting a packet.
[0037] The reward calculation unit 15 calculates a reward for reinforcement learning based on the communication success or failure included in the communication environment information output from the communication environment observation unit 12 and the delay time measured by the delay time measurement unit 13. In the communication device 1 of this embodiment, the reward is defined so that it becomes smaller as the delay time becomes longer.
[0038] Here, an example of the definition of this reward will be described. FIG. 5A shows a first example of a reward definition.
[0039] In the communication device 1 of this embodiment, the learning unit 16 performs learning to suppress the occurrence of a large delay time, and defines the reward so that the longer the delay time, the smaller the reward, as shown in Fig. 5A. Specifically, the first reward value, which is the reward when the delay time is the first delay time, should be equal to or greater than the second reward value, which is the reward when the delay time is the second delay time that is longer than the first delay time. By defining the reward in this way, the learning unit 16 can perform learning so as to make a selection that will result in a higher reward, i.e., a selection that will result in a shorter delay time.
[0040] In addition, since the first reward value, which is the reward when the delay time is the first delay time, is equal to or greater than the second reward value, which is the reward when the delay time is the second delay time that is longer than the first delay time, the definition shown in Fig. 5B may include a constant value regardless of the change in the delay time. In Fig. 5A and Fig. 5B, the reward is expressed in the form of a linear function (straight line) with respect to the delay time, but it may be a function of a degree of a quadratic function or higher, such as a curved line.
[0041] As in this first example, by defining a shorter delay time as the reward, the effect is that long delay times are less likely to occur as learning progresses.
[0042] 6A and 6B show a second example of a reward definition. In the second example, a delay threshold is introduced. In other words, a threshold is set and rewards are given differently depending on whether the delay is greater or smaller than the threshold.
[0043] In the example shown in Figure 6A, when the delay time is smaller than the threshold, the reward is inversely proportional to the delay time, and when the delay time is larger than the threshold, the reward is a straight line with a slight slope. In the example shown in Figure 6B, different fixed values are given before and after the threshold. In both cases, the reward value for a long delay time is less than the reward value for a short delay time.
[0044] As in this second example, by changing the formula for calculating the reward before and after the threshold, effective learning is possible.
[0045] FIG. 7 shows a third example of a reward definition. In the third example, two thresholds are set, and different reward values are given in three intervals, and in particular, the reward is set to 0 between two thresholds (first threshold, second threshold). By defining it in this way, it is possible to give a reward that is unlikely to have a large effect on learning when the delay time falls between them.
[0046] As in this third example, by setting the reward value between the thresholds to 0 and adjusting the reward weight, appropriate learning is possible.
[0047] FIG. 8 shows a derivation of the third example of the reward definition. In this derivation, it is defined that a positive reward is given when the value is equal to or less than the first threshold, and a negative reward is given when the value is equal to or more than the second threshold. This is a derivation of the third example, and clearly shows that a reward other than 0 may be defined between the first and second thresholds, not limited to 0.
[0048] It is preferable that the various threshold values shown in Figures 6A, 6B, 7, and 8 have larger values, for example, as the number of stations (communication devices 1) performing packet processing for the same access point increases. This is because, as the number of stations increases, packet collisions become more likely to occur, and further, because the time during which other stations are communicating is also counted as delay time, if the threshold value is too small, a delay time greater than the second threshold value will always occur, increasing the possibility that the desired learning operation will not be performed.
[0049] It is also preferable to define that different reward calculation formulas are used when communication is successful and when communication is unsuccessful. When an error occurs, it is not possible to determine whether the error is due to a packet collision or an error due to an MCS that is too high, so in that case, it may be better to give a fixed reward separate from the delay time to obtain better control results. Therefore, it is preferable to add such a definition.
[0050] Below (Example 1) and (Example 2) are examples of reward calculation formulas. Here, DTh is the threshold and Td is the delay time. (Example 1)
number
number
[0051] Returning to FIG. 1, the description of an example of the configuration relating to radio resource control of the communication device 1 of the present embodiment will be continued.
[0052] The learning unit 16 performs learning to select an action that is estimated to have the highest reward fed back from the reward calculation unit 15 in the state input from the state calculation unit 14, by repeating the process of selecting an action for the state input from the state calculation unit 14 and feeding back a reward from the reward calculation unit 15. Here, the selection of an action means determining how to change the parameters that were applied at the time of the most recent packet transmission. The options for the action include leaving the action as it is.
[0053] As described above, the reward is defined to be smaller as the delay time becomes longer. Therefore, the learning unit 16 selects an action for optimizing the parameters to the parameters that cause the least possible delay time in the environment in which the communication device 1 is placed.
[0054] The parameter (behavior) determination unit 17 determines wireless communication parameters 100 to be applied by the communication unit 11 when transmitting packets, based on the learning result of the learning unit 16. The communication unit 11 applies the wireless communication parameters 100 determined by the parameter (behavior) determination unit 17, and transmits packets, for example, to an access point. The parameter (behavior) determination unit 17 also supplies the determined wireless communication parameters 100 to the state calculation unit 14.
[0055] FIG. 9 is a flowchart showing the flow of operations related to radio resources of the communication device 1 of the first embodiment.
[0056] The communication device 1 observes a communication environment and acquires communication environment information (S101). The communication device 1 calculates a state for reinforcement learning from an applied radio resource control parameter and the acquired communication environment information (S102).
[0057] Furthermore, the communication device 1 measures a delay time (S103) in parallel with, for example, S101 to S102, and calculates a reward for reinforcement learning from the measured delay time and success or failure of communication (S104).
[0058] Then, the communication device 1 executes reinforcement learning to learn an action to obtain the highest reward in a certain state using the delay time as a reward (S105). Note that the reward is defined to be smaller as the delay time becomes longer. Also, the action is how to change the applied radio resource control parameter.
[0059] The communication device 1 determines the radio resource control parameters to be applied based on the results of the reinforcement learning (S106). Then, the communication device 1 applies the determined radio resource control parameters to perform wireless communication (S107). The communication device 1 adaptively controls the radio resource control parameters by repeating the above steps S101 to S107.
[0060] Note that there is a one-packet delay between S101-S102, S105-S107 and S103-S104. Specifically, the reward calculated in S103-S104 is given for the action selected at the time of transmitting the previous packet. That is, there is a one-packet delay between the state input from the state calculation unit 14 to the learning unit 16 and the reward fed back from the reward calculation unit 15 to the learning unit 16 in FIG. 1. The learning unit 16 accumulates the state input from the state calculation unit 14, and performs reinforcement learning using the state input at the time of transmitting the previous packet, while determining an action for the radio resource control parameter to be applied to the next packet transmission from the newly input state.
[0061] As described above, in the communication device 1 of the first embodiment, by using reinforcement learning using delay time as a reward for radio resource control, the device can automatically learn the environment in which it is installed and select communication parameters that do not cause a large delay time in that environment.
[0062] In other words, the communication device 1 of the first embodiment realizes suppression of a large delay time that becomes a serious problem during control communication and the like.
[0063] Second embodiment Next, a second embodiment will be described.
[0064] FIG. 10 is a block diagram showing an example of a configuration related to radio resource control of the communication device 1 of the second embodiment.
[0065] In the radio resource control described in the first embodiment, a wide variety of control parameters may exist, but in the second embodiment, the control parameters are limited to MCS and adjusted. Therefore, as shown in Fig. 10, the configuration of the communication device 1 in the second embodiment is different from the configuration of the communication device 1 in the first embodiment (see Fig. 1) in that the learning unit 16 is changed to an MCS control learning unit 16A, the parameter (action) (action) determining unit 17 is changed to an MCS determining unit 17A, and the wireless communication parameter 100 is changed to an MCS value 101. The other components are the same as those in the first embodiment. However, the information used in the state calculating unit 14 among the communication environment information output from the communication environment observing unit 12 may be different between the configuration of the first embodiment and the configuration of the second embodiment.
[0066] Fig. 11 shows an example of MCS values. Specifically, Fig. 11 shows MCS values established in IEEE 802.11ac. As shown in Fig. 11, the MCS value indicates a combination of a modulation method and a coding rate. If communication is performed using a high MCS value, it is possible to transmit a large amount of data in a short time, but errors are more likely to occur when the signal strength is low. On the other hand, if communication is performed using a low MCS value, errors are less likely to occur, but the time required to transmit data is longer.
[0067] Therefore, in the second embodiment, in order to select an appropriate MCS value that does not cause a large delay time in the environment in which the device is installed, the device learns to select a behavioral choice between setting a higher value than the previous MCS value, setting a lower value, or not changing the MCS value. At this time, there is no particular limit to the amount by which the MCS value is changed, and the value may be changed by two or more values.
[0068] In this way, the communication device 1 of the second embodiment enables MCS control that does not cause long delay times by using reinforcement learning for radio resource control to adjust the MCS value depending on the state using the delay time as a reward.
[0069] Third embodiment Next, a third embodiment will be described.
[0070] FIG. 12 is a block diagram showing an example of a configuration related to radio resource control of the communication device 1 of the second embodiment.
[0071] In the second embodiment, the control parameter is adjusted by being limited to MCS. In contrast, in the third embodiment, the control parameter is adjusted by being limited to CW. Therefore, as shown in FIG. 12, the configuration of the communication device 1 of the third embodiment is different from the configuration of the communication device 1 of the first embodiment (see FIG. 1) in that the learning unit 16 is changed to a CW control learning unit 16B, the parameter (action) determination unit 17 is changed to a CW determination unit 17B, and the wireless communication parameter 100 is changed to a CW value 102. The other components are the same as those of the first embodiment. However, the information used by the state calculation unit 14 among the communication environment information output from the communication environment observation unit 12 may be different between the configuration of the first embodiment and the configuration of the second embodiment.
[0072] An example of a typical CW control is shown in Figure 13. As shown in Figure 13, the CW value is a value that doubles each time a communication error occurs and a retransmission is performed, and a value randomly selected from this range is selected as the backoff value. Allowing a waiting time equal to the backoff time has the effect of lowering the probability of simultaneous communication with other terminals, and the larger the CW value, the lower the probability of simultaneous communication, making it less likely that errors will occur due to packet collisions. On the other hand, there is a high possibility that the waiting time will be long, and as a result, delay times tend to be longer.
[0073] Therefore, in the third embodiment, in order to select an appropriate CW value so as not to cause a large delay time in the installation environment, the system learns to select a behavioral choice between setting a higher CW value than before or leaving it as it is. At this time, there is no particular limit to the amount by which the CW value is changed, and it may be changed by four times or more at a time.
[0074] In this way, the communication device 1 of the third embodiment enables CW control that does not cause long delay times by using reinforcement learning for radio resource control to adjust the CW value according to the state using delay times as rewards.
[0075] (Fourth embodiment) Next, a fourth embodiment will be described.
[0076] FIG. 14 is a block diagram showing an example of a configuration related to radio resource control of the communication device 1 of the fourth embodiment.
[0077] In the second embodiment, control by reinforcement learning was performed for MCS. In the third embodiment, control by reinforcement learning was performed for CW. In the fourth embodiment, learning for MCS control and CW control is performed independently and in parallel, and the MCS value and CW value obtained as a result are utilized for radio resource control. Therefore, as shown in FIG. 14, the configuration of the communication device 1 of the fourth embodiment is the same as the configuration of the communication device 1 of the first embodiment (see FIG. 1), in that the communication unit 11, the communication environment observation unit 12, and the delay time measurement unit 13 are common, but there are components related to learning for MCS control and CW control ([1] state calculation unit 14-1, reward calculation unit 15-1, MCS control learning unit 16A, MCS determination unit 17A, [2] state calculation unit 14-2, reward calculation unit 15-2, CW control learning unit 16B, CW determination unit 17B), and different learning is performed in parallel. With this configuration, in the fourth embodiment, since both MCS and CW can be controlled simultaneously using reinforcement learning, the effect of suppressing the occurrence of a large delay time is further enhanced. Also, the configuration of the communication device 1 in the second embodiment and the configuration of the communication device 1 in the third embodiment can be used as is.
[0078] In this way, the communication device 1 of the fourth embodiment enables a combination of MCS control and CW control that does not cause a long delay time.
[0079] Fifth embodiment Next, a fifth embodiment will be described.
[0080] FIG. 15 is a block diagram showing an example of a configuration related to radio resource control of the communication device 1 of the fourth embodiment.
[0081] In the fourth embodiment, an example in which the configuration of the second embodiment and the configuration of the third embodiment are simultaneously used has been described, but the fifth embodiment has a configuration in which MCS and CW are learned collectively. As shown in Fig. 15, compared to the configuration of the communication device 1 of the first embodiment (see Fig. 1), the learning unit 16 is changed to an MCS control & CW control learning unit 16C, the parameter (action) determination unit 17 is changed to an MCS & CW determination unit 17C, and the wireless communication parameter 100 is changed to an MCSS value & CW value 103. The other components are the same as those of claim 1.
[0082] The changes compared to the first embodiment are the same as those in the second and third embodiments, but in the fifth embodiment, since learning is performed to perform MCS control and CW control collectively, the options for possible actions are MCS control options × CW control options. Furthermore, the communication environment information to be considered as a state must also include information necessary for both MCS control and CW control. Therefore, compared to the configurations of the second to fourth embodiments, the configuration of the fifth embodiment tends to have a larger number of states and action options.
[0083] That is, the communication device 1 of the fifth embodiment can adaptively perform more detailed radio resource control according to the environment in which it is installed.
[0084] Sixth embodiment Next, a sixth embodiment will be described.
[0085] FIG. 16 is a block diagram showing an example of a configuration related to radio resource control of the communication device 1 of the sixth embodiment.
[0086] In radio resource control using reinforcement learning, it is possible that actions that are not the desired actions are continuously selected in the learning process. In order to avoid such a situation, it is more stable to add a mechanism for switching to heuristic radio resource control (conventional method) in a situation where errors continue to occur repeatedly or a selection that causes too long a delay time is continuously made. Therefore, as shown in FIG. 16, a switching unit 18 is added to the configuration of the communication device 1 of the sixth embodiment compared to the configuration of the communication device 1 of the first embodiment (see FIG. 1).
[0087] The switching unit 18 has a heuristic resource control unit 181. When an error occurs, the heuristic resource control unit 181 performs general resource control such as lowering the MCS value or automatically doubling the CW size. After transmitting a packet using parameters determined based on the results of reinforcement learning, the switching unit 18 monitors the observation results of the communication environment observation unit 12 and the measurement results of the delay time measurement unit 13, and applies the wireless communication parameters determined by the heuristic resource control unit 181 to the wireless communication of the communication unit 11, instead of the wireless communication parameters 100 determined by the parameter (action) determination unit 17, when determining that the learning unit 16 has continuously selected an action that is not a desired action.
[0088] This allows the communication device 1 of the sixth embodiment to deal with the continuous selection of undesired actions in the learning process, which may occur when reinforcement learning is used for radio resource control.
[0089] As described above, the communication device 1 of each embodiment uses reinforcement learning using the delay time as a reward for resource control, thereby realizing suppression of large delay times that can become a serious problem during control communication, etc.
[0090] The present invention is not limited to the above-described embodiment, and the components can be modified and embodied in the implementation stage without departing from the gist of the invention. In addition, various inventions can be formed by appropriately combining the multiple components disclosed in the above-described embodiment. For example, some components may be deleted from all the components shown in the embodiment. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]
[0091] 1...communication device, 2...base station, 11...communication unit, 12...communication environment observation unit, 13...delay time measurement unit, 14...state calculation unit, 15...reward calculation unit, 16...learning unit, 16A...MCS control learning unit, 16B...CW control learning unit, 16C...MCS control & CW control learning unit, 17...parameter (action) determination unit, 17A...MCS determination unit, 17B...CW determination unit, 17C...MCS & CW determination unit, 18...switching unit, 181...heuristic resource control unit.
Claims
1. 1. A method of resource control in wireless communication, comprising: Calculating a state for reinforcement learning to be used for the resource control based on a parameter for the resource control applied at the time of transmitting the most recent packet and a communication environment observed up to the time of transmitting the next packet, including the communication environment observed accompanying the transmission of the most recent packet; Calculating a reward for the reinforcement learning based on a delay time, which is a time from when the packet transmission is started until an ACK is returned or a time until a timeout occurs without an ACK being returned, and on whether the packet transmission is successful or not; performing reinforcement learning for selecting an action using the calculated state and reward, with respect to the next packet to be transmitted, the action being a choice of what adjustment to make to the parameters that were applied when the most recent packet was transmitted; determining the parameters to be applied to the next packet to be transmitted based on the results of the reinforcement learning; The parameters include an MCS value and a CW value; Calculating one state and one reward for the MCS value and the CW value, defining an option for the action that combines an action for the MCS value and an action for the CW value, and performing the reinforcement learning for both the MCS value and the CW value in an integrated manner. method.
2. 1. A method of resource control in wireless communication, comprising: Calculating a state for reinforcement learning to be used for the resource control based on a parameter for the resource control applied at the time of transmitting the most recent packet and a communication environment observed up to the time of transmitting the next packet, including the communication environment observed accompanying the transmission of the most recent packet; Calculating a reward for the reinforcement learning based on a delay time, which is a time from when the packet transmission is started until an ACK is returned or a time until a timeout occurs without an ACK being returned, and on whether the packet transmission is successful or not; performing reinforcement learning for selecting an action using the calculated state and reward, with respect to the next packet to be transmitted, the action being a choice of what adjustment to make to the parameters that were applied when the most recent packet was transmitted; determining the parameters to be applied to the next packet to be transmitted based on the results of the reinforcement learning; switching to another resource control method that operates in accordance with a predetermined rule when it is determined that the resource control using the reinforcement learning is not operating normally after transmitting a packet using the parameters determined based on the result of the reinforcement learning; method.
3. The parameters include a Modulation and Coding Scheme (MCS) value; The action options include increasing, decreasing, or leaving the MCS value unchanged. The method of claim 2.
4. The parameters include a CW (Contention Window) value, The action options include increasing the CW value or leaving it unchanged. The method of claim 2.
5. The parameters include an MCS value and a CW value; Calculating the state and the reward for each of the MCS value and the CW value, and defining the action options, and performing the reinforcement learning for the MCS value and the reinforcement learning for the CW value independently and in parallel; The method of claim 2.
6. The method of claim 2, wherein the communication environment includes information on at least one of a Cyclic Redundancy Check (CRC), a Signal to Noise Ratio (SNR), an Error Vector Magnitude (EVM), communication success or failure, or a propagation path response characteristic.
7. The method according to claim 2, wherein a first reward value, which is a reward when the delay time is a first time, is equal to or greater than a second reward value, which is a reward when the delay time is a second time that is longer than the first time.
8. The method according to claim 7 , wherein the reward is calculated using a different formula when the delay time is equal to or less than a threshold and when the delay time is equal to or more than the threshold.
9. The method of claim 7 , wherein the reward is zero if the delay time is between a first threshold and a second threshold.
10. 8. The method of claim 7, wherein a positive reward is given if the delay time is less than or equal to a first threshold, and a negative reward is given if the delay time is greater than or equal to a second threshold that is greater than the first threshold.
11. The method according to claim 9 or 10, wherein the first threshold value and the second threshold value have larger values as the number of communication devices to which packets are sent to the same destination increases.
12. The method of claim 2 , wherein calculating the reward includes using different formulas for successful and unsuccessful packet transmissions.
13. The method according to claim 2 , further comprising determining that the resource control using the reinforcement learning is not operating normally when the number of consecutive errors in which an ACK is not returned exceeds or is equal to or exceeds a threshold value.
14. The method according to claim 2 , further comprising determining that the resource control using the reinforcement learning is not operating normally if the delay time is equal to or exceeds a threshold value.
15. The method according to claim 2, wherein the other resource control method performs an operation of lowering an MCS value included in the parameter or increasing a CW value when the number of consecutive errors in which an ACK is not returned exceeds or is equal to or exceeds a threshold value.
16. The method according to claim 2, wherein the other resource control method performs an operation of increasing the MCS value when the MCS value included in the parameters is determined to be small for the observed communication environment when the delay time is equal to or exceeds a threshold value continuously.
Citation Information
Patent Citations
Realization method for Q learning based vehicle-mounted network media access control (MAC) protocol
CN105306176A
Data transmission method and system, hardware system and computer storage medium
CN112422240A
Access method of user equipment, user equipment and adjustment method of contention window
JP2015091132A
Wireless communication device, wireless communication system and interference determination method
JP2017130726A
Method for adjusting contention window size in wireless access system supporting unlicensed band, and device for supporting the same
JP2019118153A