Enhanced distributed channel access based on reinforcement learning

By adopting an enhanced distributed channel access method based on reinforcement learning in wireless communication devices, the problems of low delay and spectrum efficiency in channel access in the prior art are solved, and higher data rate and spectrum efficiency are achieved.

CN120052050APending Publication Date: 2025-05-27QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072885.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-08-31
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the process of channel access, existing wireless communication systems have problems with delay, low spectrum efficiency and data rate. Especially in the selection of contention window (CW) parameters, it is difficult for the device to maintain near the optimal CW value, resulting in low spectrum efficiency and data rate.

Method used

Using a reinforcement learning (RL)-based enhanced distributed channel access (EDCA) method, by receiving information associated with the RL model, the device may perform a distributed channel access process and send a protocol data unit (PDU) during a time slot based on the output of the RL model.

Benefits of technology

By supporting an RL model associated with performing a distributed channel access process, wireless communication devices can more effectively select channel access parameters, reduce transmission failures, improve spectral efficiency, and reduce delay and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120052050A_ABST
    Figure CN120052050A_ABST
Patent Text Reader

Abstract

The present disclosure provides methods, components, devices, and systems for obtaining one or more parameters associated with a channel access procedure using a reinforcement learning (RL) model. Some aspects are more particularly directed to mechanisms in which a wireless communication device can receive information associated with the RL model and transmit a protocol data unit (PDU) during a time slot based on an output of the model. The wireless communication device can perform a distributed channel access procedure according to the information using the RL model, and can further transmit the PDU during the time slot based on the output of the RL model according to the distributed channel access procedure. The information associated with the RL model can indicate or configure the RL model, or can indicate whether the wireless communication is allowed to retrain the RL model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This patent application claims priority to U.S. Patent Application No. 17 / 969,591, titled "REINFORCEMENT LEARNING-BASED ENHANCED DISTRIBUTED CHANNEL ACCESS," filed on October 19, 2022, by Naik et al., which is assigned to the assignee of the present application and is hereby incorporated by reference in its entirety. Technical Field

[0003] The following relates to wireless communication, including enhanced distributed channel access (EDCA) based on reinforcement learning (RL). Background Art

[0004] Wireless communication systems are widely deployed to provide various types of communication content, such as voice, video, packet data, messaging, broadcasting, and so on. These systems may be multi-access systems capable of supporting communication with multiple users by sharing available system resources, such as time, frequency, and power. A wireless network (e.g., a WLAN, such as a Wi-Fi (such as Institute of Electrical and Electronics Engineers (IEEE) 802.11) network) may include an AP that can communicate with one or more stations (STAs) or mobile devices. The AP may be coupled to a network such as the Internet and may enable the mobile devices to communicate via the network (or communicate with other devices coupled to the access point). Wireless devices can communicate bidirectionally with network devices. For example, in a WLAN, an STA can communicate with an associated AP via the DL and UL. The DL (or forward link) may refer to the communication link from the AP to the station, while the UL (or reverse link) may refer to the communication link from the station to the AP. Summary of the Invention

[0005] The systems, methods, and devices of the present disclosure each have several innovative aspects, none of which alone is solely responsible for the desired attributes disclosed herein.

[0006] One innovative aspect of the subject matter described in the present disclosure can be implemented in a method for wireless communication at a wireless communication device. The method may include: receiving information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access process at the wireless communication device in a wireless local area network according to the information; and transmitting a protocol data unit during a time slot based on the output of the reinforcement learning model according to the distributed channel access process.

[0007] Another innovative aspect of the subject matter described in this disclosure can be implemented in an apparatus for wireless communication at a wireless communication device. The apparatus can include a processor, a memory coupled to the processor, and instructions stored in the memory. The instructions can be executable by the processor to cause the apparatus to: receive information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at the wireless communication device in a wireless local area network according to the information; and transmit a protocol data unit according to the distributed channel access procedure and during a time slot based on an output of the reinforcement learning model.

[0008] Another innovative aspect of the subject matter described in this disclosure can be implemented in another apparatus for wireless communication at a wireless communication device. The apparatus can include: means for receiving information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at the wireless communication device in a wireless local area network according to the information; and means for transmitting a protocol data unit according to the distributed channel access procedure and during a time slot based on an output of the reinforcement learning model.

[0009] Another innovative aspect of the subject matter described in this disclosure can be implemented in a non-transitory computer-readable medium storing code for wireless communication at a wireless communication device. The code can include instructions executable by a processor to: receive information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at the wireless communication device in a wireless local area network according to the information; and transmit a protocol data unit according to the distributed channel access procedure and during a time slot based on an output of the reinforcement learning model.

[0010] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, receiving the information associated with the reinforcement learning model can include operations, features, means, or instructions for: receiving an indication that the wireless communication device is permitted to develop the reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

[0011] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, receiving the information associated with the reinforcement learning model can include operations, features, means, or instructions for: receiving information associated with the reinforcement learning model and an indication of whether the wireless communication device is permitted to retrain the reinforcement learning model for the distributed channel access procedure, the method further including selectively retraining the reinforcement learning model based on whether the wireless communication device is permitted to retrain the reinforcement learning model.

[0012] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, receiving the information associated with the reinforcement learning model may include operations, features, components, or instructions for performing the following: receiving an indication that the wireless communication device is permitted to retrain the reinforcement learning model for the distributed channel access procedure, where the reinforcement learning model may be pre-loaded at the wireless communication device, and the method further includes retraining the reinforcement learning model based on the wireless communication device being permitted to retrain the reinforcement learning model.

[0013] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, receiving the information associated with the reinforcement learning model may include operations, features, components, or instructions for performing the following: receiving an indication of a policy associated with training or retraining the reinforcement learning model, and the method further includes training or retraining the reinforcement learning model according to the policy.

[0014] One innovative aspect of the subject matter described in this disclosure may be implemented in a method for wireless communication at a wireless communication device. The method may include: sending information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network according to the information; and receiving a protocol data unit from the second wireless communication device according to the distributed channel access procedure and during a time slot based on an output of the reinforcement learning model.

[0015] Another innovative aspect of the subject matter described in this disclosure may be implemented in an apparatus for wireless communication at a wireless communication device. The apparatus may include a processor, a memory coupled to the processor, and instructions stored in the memory. The instructions may be executable by the processor to cause the apparatus to: send information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network according to the information; and receive a protocol data unit from the second wireless communication device according to the distributed channel access procedure and during a time slot based on an output of the reinforcement learning model.

[0016] Another innovative aspect of the subject matter described in this disclosure may be implemented in another apparatus for wireless communication at a wireless communication device. The apparatus may include: means for sending information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network according to the information; and means for receiving a protocol data unit from the second wireless communication device according to the distributed channel access procedure and during a time slot based on an output of the reinforcement learning model.

[0017] Another innovative aspect of the subject matter described in this disclosure can be implemented in a non-transitory computer-readable medium storing code for wireless communication at a wireless communication device. The code can include instructions executable by a processor to perform the following operations: sending information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network according to the information; and receiving a protocol data unit from the second wireless communication device during a time slot based on an output of the reinforcement learning model according to the distributed channel access procedure.

[0018] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, sending the information associated with the reinforcement learning model can include operations, features, components, or instructions for performing the following operations: sending an indication that the second wireless communication device is allowed to develop the reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

[0019] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, sending the information associated with the reinforcement learning model can include operations, features, components, or instructions for performing the following operations: sending the information associated with the reinforcement learning model and an indication of whether the second wireless communication device is allowed to retrain the reinforcement learning model for the distributed channel access procedure.

[0020] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, sending the information associated with the reinforcement learning model can include operations, features, components, or instructions for performing the following operations: sending an indication that the second wireless communication device is allowed to retrain the reinforcement learning model for the distributed channel access procedure, where the reinforcement learning model can be pre-loaded at the second wireless communication device.

[0021] In some examples of the methods, apparatuses, and non-transitory computer-readable media described herein, sending the information associated with the reinforcement learning model can include operations, features, components, or instructions for performing the following operations: sending an indication of a policy associated with training or retraining the reinforcement learning model.

[0022] Details of one or more specific implementations of the subject matter described in this disclosure are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following drawings are not drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1A schematic diagram of an example wireless communication network that supports reinforcement learning (RL)-based enhanced distributed channel access (EDCA) in accordance with aspects of the present disclosure is shown.

[0024] Figure 2 An example signaling diagram that supports RL-based EDCA in accordance with one or more aspects of the present disclosure is shown.

[0025] Figure 3 An example RL model that supports RL-based EDCA in accordance with one or more aspects of the present disclosure is shown.

[0026] Figure 4 and Figure 5 An example RL process that supports RL-based EDCA in accordance with one or more aspects of the present disclosure is shown.

[0027] Figure 6 An example process flow that supports RL-based EDCA in accordance with one or more aspects of the present disclosure is shown.

[0028] Figure 7 A flowchart that illustrates an example process that can be performed by a wireless AP that supports RL-based EDCA in accordance with one or more aspects of the present disclosure is shown.

[0029] Figure 8 A flowchart that illustrates an example process that can be performed by a wireless STA that supports RL-based EDCA in accordance with one or more aspects of the present disclosure is shown.

[0030] Figure 9 and Figure 10 A block diagram of an example wireless communication device that supports RL-based EDCA in accordance with one or more aspects of the present disclosure is shown.

[0031] Like reference numerals and names in different figures represent like elements. Detailed Description

[0032] The following description is directed to certain specific examples in order to describe innovative aspects of the present disclosure. However, one of ordinary skill in the art will readily recognize that the teachings herein can be applied in many different ways. Some or all of the described examples can be in accordance with the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards, IEEE 802.15 standards, such as Bluetooth as defined by the Bluetooth Special Interest Group (SIG) implemented in any device, system, or network that transmits and receives radio frequency (RF) signals according to one or more of the standards, such as the Long Term Evolution (LTE), 3G, 4G, or 5G (New Radio (NR)) standards released by the 3rd Generation Partnership Project (3GPP). The described examples can be implemented in any device, system, or network capable of transmitting and receiving RF signals according to one or more of the following technologies or techniques: Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Space Division Multiple Access (SDMA), Rate Splitting Multiple Access (RSMA), Multi-User Shared Access (MUSA), Single-User (SU) Multiple-Input Multiple-Output (MIMO), and Multi-User (MU)-MIMO. The described examples can also be implemented using other wireless communication protocols or RF signals suitable for use in one or more of a Wireless Personal Area Network (WPAN), Wireless Local Area Network (WLAN), Wireless Wide Area Network (WWAN), Wireless Metropolitan Area Network (WMAN), or Internet of Things (IoT) network.

[0033] In some Wi-Fi systems, one or more devices may contend for channel access according to a channel access technology such as Enhanced Distributed Channel Access (EDCA). EDCA may be associated with two main parameters, which include a contention window (CW) and a backoff (BO) counter. According to EDCA, a device may initialize the CW to a minimum value (such as CW min ), select a random BO (RBO) in the range [0, CWmin - 1], decrement the RBO by 1 for each time slot sensed as idle by the device, and perform a transmission when RBO = 0. If the transmission fails, the device may set the CW to the smaller of twice the previously used CW and a maximum value (such as CW max ). If a subsequent transmission attempt is successful, the device may re-initialize the CW to the minimum value (such as CW min ). In other words, each time a transmission fails, the device may double the CW, and once a transmission is successful, the device may reset to the minimum CW value. Thus, the device has a relatively low likelihood of keeping the CW near the "optimal" CW (the "optimal" CW can be understood as a CW having a high likelihood of facilitating successful transmission while maintaining a suitable latency), unless the optimal CW duration is the minimum CW value. Instead, the device may have a high likelihood of scaling the CW proportionally towards the optimal CW value by first experiencing one or more transmission failures. This channel access design may introduce latency and may be associated with relatively low spectral efficiency and data rate.

[0034] Various aspects of the present disclosure generally relate to communicating protocol data units (PDUs) using a reinforcement learning (RL) model associated with a channel access procedure. Some aspects more specifically relate to one or more configuration- or signaling-based mechanisms according to which a wireless communication device may receive information associated with the RL model and transmit a PDU during a time slot based on an output of the model. In some particular implementations, a wireless communication device (which may be an example of a station (STA)) may use the RL model to perform a distributed channel access procedure based on the information and may further transmit a PDU during a time slot based on an output of the RL model during the distributed channel access procedure. By supporting such an RL model associated with performing a distributed channel access procedure, a wireless communication device may have a higher likelihood of maintaining a CW that promotes successful transmission while maintaining an appropriate latency relative to doubling the CW each time a transmission fails and resetting to a minimum CW value once a transmission is successful, as described above. This may enable lower latency, higher data rates, greater spectral efficiency, and lower power consumption.

[0035] Certain aspects of the subject matter may be implemented to achieve one or more of the following potential advantages. In some particular implementations, by supporting such an RL model associated with performing a distributed channel access procedure, a wireless communication device may use RL to obtain, identify, select, or otherwise determine relatively better channel access parameters than those that the wireless communication device may otherwise use in a channel access procedure based on a non-RL model. For example, according to using an RL model associated with a distributed channel access procedure, a wireless communication device may sense or identify a time slot during which a transmission is relatively more likely to result in a successful transmission from the wireless communication device. Thus, the wireless communication device may experience fewer transmission failures, which may enable lower latency, higher data rates, greater spectral efficiency, and lower power consumption.

[0036] Various additional aspects of the present disclosure relate to information associated with the RL model that a wireless communication device may receive. In some particular implementations, the information may include an indication that the wireless communication device is permitted to develop an RL model at the wireless communication device and use the RL model for a distributed channel access procedure once it is developed. Additionally or alternatively, the wireless communication device may be configured with an RL model via the information such that the information may indicate a complete RL model or various aspects of the RL model. Additionally or alternatively, the wireless communication device may receive an indication as to whether the wireless communication device is permitted to retrain the RL model or an indication as to what parameters the wireless communication device may obtain using the RL model.

[0037] Aspects of the present subject matter may also be implemented to achieve one or more of the following potential advantages. For example, according to support signaling mechanisms by which a wireless communication device may receive an indication of how the wireless communication device is permitted or expected to use an RL model, a network controller (such as an access point (AP)) may enable the RL model to be used for channel access in a controlled manner. Thus, network deployment may balance the lower latency, higher data rate, greater spectral efficiency, and lower power consumption associated with using an RL model with channel access fairness. For example, the network controller may restrict a wireless communication device from retraining the RL model to avoid a situation where the wireless communication device dominates other RL - incapable devices in channel access. Alternatively, the network controller may enable a wireless communication device to retrain the RL model in a deployment where all local devices have RL capabilities such that these devices may compete for the channel on an equal footing. Thus, a network employing such signaling mechanisms may balance RL - based channel access effectiveness with network - wide channel access fairness to ensure that various device types can receive service.

[0038] Figure 1 FIG. shows a schematic diagram of an example wireless communication network 100 that supports RL - based EDCA in accordance with aspects of the present disclosure. According to some aspects, the wireless communication network 100 may be an example of a wireless local area network (WLAN) (such as a Wi - Fi network) (and will be referred to hereinafter as WLAN 100). For example, WLAN 100 may be a network that implements at least one of the IEEE 802.11 wireless communication protocol standards family (such as the standards defined by the IEEE 802.11 - 2020 specification or its revisions, including but not limited to 802.11ay, 802.11ax, 802.11az, 802.11ba, 802.11bd, 802.11be, 802.11bf, and the 802.11 revisions associated with Wi - Fi8). WLAN 100 may include a number of wireless communication devices, such as wireless AP 102 and multiple wireless STAs 104. Although Figure 1 only one AP 102 is shown, WLAN 100 may also include multiple APs 102. Figure 1 The illustrated AP 102 may represent various different types of APs, including but not limited to enterprise - level APs, single - band APs, dual - band APs, stand - alone APs, software - enabled APs (soft APs), and multi - link APs. The coverage area and capacity of a cellular network (such as LTE or 5G NR) may be further improved by small cells supported by APs acting as small base stations. Additionally, a dedicated cellular network may be established by a wireless regional network using small cells.

[0039] Each of the STAs 104 may also be referred to as a mobile station (MS), mobile device, mobile phone, wireless phone, access terminal (AT), user equipment (UE), subscriber station (SS), or subscriber unit, etc. The STAs 104 may represent various devices, such as mobile phones, personal digital assistants (PDAs), other handheld devices, netbooks, notebook computers, tablet computers, laptop computers, chromebooks, extended reality (XR) headsets, wearable devices, display devices (such as TVs (including smart TVs), computer monitors, navigation systems, etc.), music or other audio or stereo devices, remote control devices ("remote controls"), printers, kitchen appliances (including smart refrigerators) or other household appliances, remote keys (such as for passive keyless entry and start (PKES) systems), Internet of Things (IoT) devices, and transportation vehicles, etc. Various STAs 104 in the network can communicate with each other via the AP 102.

[0040] A single AP 102 and the associated set of STAs 104 may be referred to as a basic service set (BSS), which is managed by the corresponding AP 102. Figure 1 An example coverage area 108 of the AP 102 is additionally shown, and this example coverage area may represent the basic service area (BSA) of the WLAN 100. The BSS can be identified or indicated to users by a service set identifier (SSID), and can also be identified or indicated to other devices by a basic service set identifier (BSSID), which can be the media access control (MAC) address of the AP 102. The AP 102 may periodically broadcast a beacon frame ("beacon") including the BSSID so that any STA 104 within the wireless range of the AP 102 can "associate" or re-associate with the AP 102 to establish a corresponding communication link 106 (also referred to hereinafter as a "Wi-Fi link") with the AP 102 or maintain the communication link 106 with the AP. For example, the beacon may include an identification or indication of the primary channel used by the corresponding AP 102 and a timing synchronization function for establishing or maintaining timing synchronization with the AP 102. The AP 102 can provide access to an external network to various STAs 104 in the WLAN via the corresponding communication link 106.

[0041] To establish a communication link 106 with an AP 102, each of the STAs 104 is configured to perform a passive or active scanning operation ("scanning") on a frequency channel in one or more frequency bands, such as the 2.4 GHz, 5 GHz, 6 GHz, or 60 GHz frequency bands. To perform passive scanning, the STA 104 listens for beacons transmitted by the corresponding AP 102 at periodic time intervals, referred to as target beacon transmission times (TBTTs) (measured in time units (TUs), where one TU may be equal to 1024 microseconds (μs)). To perform active scanning, the STA 104 generates probe requests and sequentially transmits these probe requests on each channel to be scanned, and listens for probe responses from the AP 102. Each STA 104 may identify, determine, ascertain, or select the AP 102 with which to associate based on the scanning information obtained through passive or active scanning, and perform authentication and association operations to establish a communication link 106 with the selected AP 102. The AP 102 assigns an association identifier (AID) to the STA 104 at the end of the association operation, and the AP 102 uses this association identifier (AID) to track the STA 104.

[0042] Since wireless networks are becoming increasingly common, the STA 104 may have the opportunity to choose among a number of BSSs within the range of the STA or among multiple APs 102 that together form an extended service set (ESS) (including multiple connected BSSs). Extended network stations associated with the WLAN 100 may be connected to a wired or wireless distribution system that allows multiple APs 102 to be connected in such an ESS. Thus, the STA 104 may be covered by more than one AP 102 and may be associated with different APs 102 at different times for different transmissions. Additionally, after associating with an AP 102, the STA 104 may also periodically scan its surroundings to look for a more suitable AP 102 with which to associate. For example, a STA 104 that is moving relative to its associated AP 102 may perform a "roaming" scan to look for another AP 102 with more desirable network characteristics, such as a greater received signal strength indicator (RSSI) or a reduced traffic load.

[0043] In some embodiments, STA 104 may form a network without an AP 102 or other equipment other than STA 104 itself. An example of such a network is an ad hoc network (or wireless ad hoc network). An ad hoc network may alternatively be referred to as a mesh network or a peer-to-peer (P2P) network. In some embodiments, an ad hoc network may be implemented within a larger wireless network such as WLAN 100. In such an example, although STA 104 may be able to communicate with each other through AP 102 using communication link 106, STA 104 may also communicate directly with each other via direct wireless communication link 110. Additionally, two STA 104 may communicate via direct communication link 110 regardless of whether the two STA 104 are associated with and served by the same AP 102. In such an ad hoc system, one or more STA 104 may assume the role played by AP 102 in a BSS. Such a STA 104 may be referred to as a group owner (GO) and may coordinate transmissions within the ad hoc network. Examples of direct wireless communication link 110 include Wi-Fi direct connections, connections established through the use of Wi-Fi tunneling direct link setup (TDLS) links, and other P2P group connections.

[0044] AP 102 and STA 104 may operate and communicate (via respective communication link 106) according to one or more of the IEEE 802.11 wireless communication protocol standards. These standards define the WLAN radio and baseband protocols for the PHY and MAC layers. AP 102 and STA 104 send and receive wireless communications (also referred to hereinafter as "Wi-Fi communications" or "wireless packets") to and from each other in the form of PHY protocol data units (PPDUs). AP 102 and STA 104 in WLAN 100 may send PPDUs on an unlicensed spectrum, which may be a part of a spectrum including bands traditionally used by Wi-Fi technology such as the 2.4 GHz band, 5 GHz band, 60 GHz band, 3.6 GHz band, and 900 MHz band. Some examples of AP 102 and STA 104 described herein may also communicate in other bands that support both licensed and unlicensed communications such as the 5.9 GHz band and 6 GHz band. AP 102 and STA 104 may also communicate on other bands such as shared licensed bands where multiple operators may have licenses to operate in one or more of the same or overlapping bands.

[0045] Each frequency band in the frequency band may include multiple sub - bands or frequency channels. For example, PPDUs compliant with the IEEE 802.11n, 802.11ac, 802.11ax, and 802.11be standard revisions may be transmitted on the 2.4 GHz, 5 GHz, or 6 GHz frequency bands, where each frequency band is divided into multiple 20 - MHz channels. Thus, these PPDUs are transmitted on physical channels with a minimum bandwidth of 20 MHz, but larger channels can be formed through channel bonding. For example, PPDUs may be transmitted on physical channels with a bandwidth of 40 MHz, 80 MHz, 160 MHz, or 320 MHz by bonding multiple 20 - MHz channels together.

[0046] Each PPDU is a composite structure that includes a PHY preamble and a payload in the form of a PHY service data unit (PSDU). The information provided in the preamble can be used by the receiving device to decode the subsequent data in the PSDU. In instances where the PPDU is transmitted on a bonded channel, the preamble field may be replicated and transmitted in each of the multiple component channels. The PHY preamble may include both a legacy part (or "legacy preamble") and a non - legacy part (or "non - legacy preamble"). The legacy preamble can be used for other purposes such as packet detection, automatic gain control, and channel estimation. The legacy preamble is also typically used to maintain compatibility with legacy devices. The format, decoding, and the information provided in the non - legacy part of the preamble are associated with the specific IEEE 802.11 protocol to be used for transmitting the payload.

[0047] Access to the shared wireless medium is generally controlled by the distributed coordination function (DCF). With DCF, there is generally no centralized master device that allocates the time and frequency resources of the shared wireless medium. Instead, before a wireless communication device (such as AP102 or STA 104) is permitted to transmit data, the wireless communication device may wait for a specific time and contend for access to the wireless medium at that specific time. DCF is implemented by using time intervals, including the slot time (or "slot interval") and the inter - frame space (IFS). The IFS provides priority access for control frames for proper network operation. Transmissions can start at slot boundaries. There are different variants of the IFS, including the short IFS (SIFS), the distributed IFS (DIFS), the extended IFS (EIFS), and the arbitration IFS (AIFS). The values for the slot time and the IFS can be provided by appropriate standard specifications, such as one or more of the IEEE 802.11 wireless communication protocol standard family.

[0048] In some specific implementations, a wireless communication device may implement DCF by using Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) technology. According to such technology, before transmitting data, the wireless communication device may perform a Clear Channel Assessment (CCA) and may determine (such as identify, detect, ascertain, calculate, or compute) that the relevant wireless channel is idle. CCA includes both physical (PHY-level) carrier sensing and virtual (MAC-level) carrier sensing. Physical carrier sensing is accomplished via a measurement of the received signal strength of a valid frame, which is then compared with a threshold to determine (such as identify, detect, ascertain, calculate, or compute) whether the channel is busy. For example, if the received signal strength of the detected preamble is higher than the threshold, the medium is considered busy. Physical carrier sensing also includes energy detection. Energy detection involves measuring the total energy received by the wireless communication device, regardless of whether the received signal represents a valid frame. If the detected total energy is higher than the threshold, the medium is considered busy.

[0049] Virtual carrier sensing is implemented via the use of a Network Allocation Vector (NAV), which effectively serves as the duration remaining before the wireless communication device can contend for access, even in the absence of detected symbols or even if the detected energy is below the relevant threshold. Each time a valid frame not addressed to the wireless communication device is received, the NAV is reset. When the NAV reaches 0, the wireless communication device performs physical carrier sensing. If the channel remains idle within the appropriate IFS, the wireless communication device initiates a backoff timer, which represents the time duration during which the device senses the medium as idle before being permitted to transmit. If the channel remains idle until the backoff timer expires, the wireless communication device becomes the holder (or "owner") of the TxOP and may begin transmitting. The TxOP is the time duration during which the wireless communication device can transmit frames on the channel after it has won contention for the wireless medium. The TxOP duration may be indicated in the U-SIG field of the PPDU. On the other hand, if one or more of the carrier sensing mechanisms in the carrier sensing mechanism indicate that the channel is busy, the MAC controller within the wireless communication device will not permit transmission.

[0050] Each time the wireless communication device generates a new PPDU for transmission in a new TxOP, the wireless communication device randomly selects a new backoff timer duration. The available distribution of numbers that can be randomly selected for the backoff timer is referred to as the contention window (CW). There are different CW and TxOP durations for each of the following four access categories (AC): Voice (AC_VO), Video (AC_VI), Background (AC_BK), and Best Effort (AC_BE). This enables prioritization of specific types of traffic in the network.

[0051] Some APs 102 and STAs 104 can implement spatial reuse techniques. For example, APs 102 and STAs 104 configured to communicate using IEEE 802.11ax or 802.11be can be configured with a BSS color. APs 102 associated with different BSSs can be associated with different BSS colors. The BSS color is a digital identifier of the corresponding BSS of the AP (such as a 6-bit field carried by the SIG field). Each STA 104 can learn its own BSS color after associating with the corresponding AP. The BSS color information is conveyed at both the PHY sublayer and the MAC sublayer. If an AP 102 or STA 104 detects, obtains, selects, or identifies a wireless packet from another wireless communication device during contention access, the AP 102 or STA 104 can apply different contention parameters based on whether the wireless packet is sent by another wireless communication device within its BSS or sent to the other wireless communication device or sent from a wireless communication device from an overlapping BSS (OBSS), as determined, identified, ascertained, or calculated by the BSS color indication in the preamble of the wireless packet. For example, if the BSS color associated with the wireless packet is the same as the BSS color of the AP 102 or STA, the AP 102 or STA 104 can use a first received signal strength indication (RSSI) detection threshold when performing a clear channel assessment (CCA) on the wireless channel. However, if the BSS color associated with the wireless packet is different from the BSS color of the AP 102 or STA, the AP 102 or STA 104 can use a second RSSI detection threshold instead of the first RSSI detection threshold when performing a CCA on the wireless channel, and the second RSSI detection threshold is greater than the first RSSI detection threshold. In this way, the criteria for winning contention are relaxed when interfering transmissions are associated with an OBSS.

[0052] In some deployments, the WLAN 100 or devices of the WLAN 100 may support reinforcement learning, such as via artificial intelligence (AI) or machine learning (ML). Such AI / ML-capable devices of the WLAN 100 may employ AI / ML-based operations across one or more various use embodiments, including for optimizing, improving, or otherwise facilitating 802.11 features. As described herein, RL may perform tasks based on a given set of state, action, and reward. For example, in the case of a system given an agent that interacts with an environment as well as state, action, and reward, a wireless communication device (such as AP 102 or STA 104) may learn a policy (such as learning a mapping from state to action). The agent may learn through experience (such as by taking or executing actions and observing rewards or updated states). Examples of RL models may include deep Q network (DQN), policy gradient, actor-critic techniques, contextual multi-armed bandit (MAB), or non-contextual MAB, and so on.

[0053] A state may be a representation of the current environment (such as the current environment associated with a task). In other words, a set of one or more states may inform an RL agent of information associated with the current situation. An action may be something that the RL agent can execute to change one or more states. A reward may be associated with the utility of the RL agent for performing a "correct" action. For example, a state may inform the RL agent of the current situation, and a reward may signal to the RL agent a state that may be desired or associated with those states. Thus, given a set of one or more states, one or more actions, and one or more rewards, the RL agent may learn a policy, which may refer to a function according to which one or more actions can be executed in each of various states to increase (e.g., maximize) the reward. In an example RL process, the RL agent may execute a first action based on a first state and may receive a first reward associated with the execution of the first action from the first state. The RL agent may also identify a second state based on executing the first action from the first state, may execute a second action based on the second state, and may receive a second reward associated with the execution of the second action from the second state. The RL agent may track the first reward and the second reward and may learn which actions from which states result in increased rewards.

[0054] As more devices are capable of implementing RL functions (such as AI algorithms or ML algorithms), some networks can utilize RL to select, identify, calculate, estimate, or otherwise determine one or more channel access parameters. For example, wireless communication can use the output of an RL model to obtain one or more parameters associated with channel access and can use the one or more parameters as part of the channel access process. In some embodiments, wireless communication can use RL-based (such as RL-determined) parameters as part of a network-supported channel access process (such as an EDCA process), or can use RL-based parameters to bypass a network-supported channel access process.

[0055] Figure 2 Illustrates an example of a signaling diagram 200 that supports RL-based EDCA according to one or more aspects of the present disclosure. The signaling diagram 200 can be implemented or be implemented to achieve or facilitate aspects of the WLAN 100. For example, the signaling diagram 200 illustrates communication between a device 205 and a device 210, which can be examples of the devices described by Figure 1 Illustrated and referenced Figure 1 For example. The device 205 can be an example of the AP 102 or the STA 104. Similarly, the device 210 can be an example of the AP 102 or the STA 104. In some embodiments, the devices 205 and 210 can support one or more configuration- or signaling-based mechanisms according to which the device 205 can conditionally utilize some or all of the AI / ML capabilities of the device 210 based on signaling from the device 205. The device 210 can send signaling to the device 205 via a communication link 215, and the device 205 can send signaling to the device 210 via a communication link 220.

[0056] In some scenarios, one or more devices, which can be examples of 802.11 devices, can attempt to obtain channel access. In some embodiments, one or more devices can use the EDCA protocol to access the wireless medium. In such embodiments, each device can maintain CW and BO counters for each access category (AC) in a set of access categories (ACs). For example, a device can support one or more ACs. In some embodiments, a device can support four ACs including voice (VO), video (VI), best effort (BE), and background (BK). Different access categories can be associated with different priorities. For example, voice can be associated with the highest priority, video with the second highest priority, best effort with the third highest priority, and background with the lowest priority.

[0057] As part of the EDCA protocol for accessing the wireless medium, a device can initialize the CW to CW min and in [0, CW min-1] uniformly select BO. During each idle time slot, the device may decrement BO by 1, and if BO = 0, the device may transmit. If the transmission fails, the device may set CW = min(2*CW, CW max ). If the transmission is successful, the device may re-initialize CW to CW min . In other words, each time the transmission fails, the device may double CW, and once the transmission is successful, the device may reset it to the minimum CW value. Therefore, the possibility that the device keeps the CW duration near the "optimal" CW duration may be relatively low. For a given network configuration, the selection of the "optimal" CW value may achieve relatively few collisions (such as minimum collisions). The network configuration may involve or include the number of STAs 104 in the BSS, the number of overlapping BSSs, or the traffic load, and other aspects.

[0058] A device (such as an 802.11 device) may hover around the optimal value through the EDCA protocol, rather than being able to keep the CW duration near the "optimal" CW duration, but most of the time the device may pick a sub-optimal value for the CW duration. For example, if for a given network, the optimal CW = 2 4 *CW min , then the EDCA protocol may take four collisions for the device to scale up to such a CW value. In addition, after each successful transmission, the CW value is reset to CW min . Therefore, after each successful transmission, the device may take another four collisions to reach that value, which will result in a relatively long delay. Further, for example, if CW = 25 is the optimal CW value for a given network, the device may actually never be able to use CW = 25 according to the EDCA protocol because the CW value is calculated or selected as a power of 2.

[0059] Therefore, in some specific embodiments, the device 205 may use an RL-based model to determine or make channel access decisions. In other words, the device 205 may support AI / ML capabilities to control channel access in some networks (such as 802.11 networks). In some specific embodiments, to support and facilitate RL-based channel access, the device 210 may send RL model information 225 to the device 205. The RL model information 225 may include information associated with the RL model 230, and the device 205 may use this information to determine or make one or more channel access decisions. The content of the RL model information 225 may vary with the specific embodiment or deployment scenario. In some aspects, the device 210 may generate and send the RL model information 225 as part of controlling the use of the RL model 230 (such as AI / ML) in the BSS, or to support one or more available options for model generation and use, or both.

[0060] In some specific implementations, for example, the device 205 that uses the RL model 230 for channel access may use RL to develop the RL model 230 from scratch (such as through experience), and may derive interference from the RL model 230. In such specific implementations, the RL model information 225 may include an indication that the device 205 is allowed to develop its own RL model 230. Additionally or alternatively, the device 205 may use the RL model 230 downloaded from another STA (such as the device 210) as is, and may use the downloaded RL model 230 to make inferences. For example, the device 205 may download the RL model 230 from the device 210 (such as the AP 102), and may use the RL model 230 to make inferences. In such specific implementations where the device 205 downloads the RL model 230 from the device 210, the RL model information 225 may include the complete RL model 230 (such as the algorithm associated with the RL model 230), and may further indicate whether the device 205 is allowed to retrain the RL model 230. In some specific implementations, the RL model information 225 may indicate that the device 205 is not allowed to retrain the RL model 230. In some other examples, the RL model information 225 may indicate that the device 205 is allowed to retrain the RL model 230. In such examples, the device 205 may perform retraining (such as refinement) of the RL model 230 downloaded from another STA (such as the device 210), and may use the RL model 230 to make inferences. In other words, the device 205 may download the RL model 230 from the device 210 (such as the AP 102), and may use experience (such as transitions) to refine the RL model 230 and use the model to make inferences.

[0061] Additionally or alternatively, device 205 may use the locally pre-trained RL model 230 as is (such as without retraining). For example, device 205 may be pre-loaded with the pre-trained RL model 230, and device 205 may use the RL model 230 for inference without adjusting the RL model 230. In such embodiments, the RL model information 225 may indicate that device 205 is permitted to use the pre-trained RL model 230. In some embodiments, such as in embodiments where device 205 is equipped with multiple pre-trained RL models, the RL model information 225 may indicate the RL model 230 from among the multiple pre-trained RL models equipped at device 205. In some embodiments, device 205 may use RL to perform retraining (such as refinement) of the pre-trained model 230. For example, device 205 may be pre-loaded with the pre-trained RL model 230, and may use experience to refine the RL model 230 and use the RL model 230 for inference. In such embodiments, the RL model information 225 may indicate that device 205 is permitted to use the pre-trained RL model 230, and may further indicate that device 205 is permitted to retrain or refine the pre-trained RL model 230.

[0062] In some embodiments, a device 210 (such as a control entity) in a BSS may permit the use of an ML model (such as RL model 230) downloaded by device 210 to STA 104. For example, different RL models at different STAs 104 may result in different channel access probabilities, which may lead to unfairness. Thus, in such embodiments, device 210 may include in the RL model information 225 an indication that device 205 is permitted to use the RL model 230 indicated by device 210 and is not permitted to use other RL models 230 in the BSS. In some embodiments, device 210 (such as a control entity in a BSS) may not permit adjustment of the downloaded RL model 230. For example, adjusting the RL model may result in different channel access probabilities at different STAs 104, which may lead to unfairness. Thus, in such embodiments, devices 205 and 210 may support downloading of the RL model 230 without permitting retraining in the BSS.

[0063] In some other specific implementations, the device 210 (such as a control entity) may allow the downloaded RL model 230 to be adjusted in a specific and controlled manner. For example, the device 210 may announce the values of parameters associated with the RL technique (these parameters may be referred to as hyperparameters) or the policies that can be used to adjust the RL model 230. Such hyperparameters may include one or more of β, γ, ε, where β may be a learning rate, γ may be a discount factor, and ε may be a parameter associated with the ε-greedy policy. In some aspects, the learning rate may determine the degree to which newly acquired information overrides old information (where a factor of 0 may cause the agent not to learn anything, and a factor of 1 may cause the agent to consider only the most recent information), the discount factor may determine the importance of future rewards (where a discount factor of 0 may cause the agent to consider only the current reward, and a factor close to 1 may cause the agent to strive for long-term high rewards), and ε may be associated with the probability of selecting an action corresponding to the highest Q value (e.g., probability (1 - ε)). For example, the device 210 may indicate via the RL model information 225 that the ε-greedy policy (and only the ε-greedy policy) can be used to retrain or adjust the RL model 230. In some specific implementations, the device 210 may announce the criteria under which a controlled entity (such as the device 205) may initiate the retraining of the RL model 230. In other words, if one or more conditions or criteria are met, the device 205 may retrain the RL model 230, where the device 210 may indicate such conditions or criteria via the RL model information 225.

[0064] In some aspects, a set of available options for RL-based training or retraining of the RL model 230 may be indicated via the RL model information 225. The available policies during training or retraining may be one of random, randomly within a certain range (e.g., the sending probability may be uniformly random in (0.2, 0.7)), greedy, ε-greedy, Boltzmann (such as Softmax), or the EDCA protocol. In a specific implementation where the training or retraining policy is the EDCA protocol, when the RL model 230 is not used to derive inferences, the device 205 may use the EDCA protocol, but shift e t =(S t , S t+1 , A t , R t ) may be stored and used to train the RL model 230 (such as DQN). In some specific implementations, the device 210 may also indicate via the RL model information 225 a set of available or allowed policies during inference. In some aspects, the device 210 may indicate to the device 205 that the policy during inference may be greedy.

[0065] Thus, in some specific implementations, device 205 may use RL model 230 for channel access inference (such as for one or more decisions associated with channel access). Based on the inference at device 205, it is expected that an agent (such as an agent associated with AI / ML of device 205) may possess RL model 230, such as a Q-table. In an example specific implementation where RL model 230 is a Q-table for channel access, the agent may calculate state S t =[S t,0 , S t,1 , S t,2 and the agent may select an action according to α = argmax a Q(S t , a), where the action selection strategy may be greedy (such as device 205 may select the action with the highest Q value). The agent may calculate CW = 2 α * CW min , and uniformly select RBO in [0, CW - 1]. In other words, device 205 may use the RL model to obtain the channel access parameter α as an output, and may use the channel access parameter α to calculate the CW value. In some aspects, the parameter α may be referred to as the CW level. If the measured or sensed medium is busy, the device may freeze the countdown of RBO, and may count down (such as decrement) RBO for each idle time slot and transmit when RBO = 0. For example, when RBO = 0, device 205 may transmit PDU 240.

[0066] Additionally or alternatively, device 205 may use RL model 230 to obtain one or more other channel access parameters. In other words, device 205 may derive one or more actions from the output of RL model 230, and these actions may be based on one or more channel access parameters. In some specific implementations, for example, device 205 may obtain the transmission probability in an idle time slot (such as 0 < p T < 1) based on the output of RL model 230. In such specific implementations, device 205 may obtain p T from RL model 230 according to the set of states, and if the time slot is idle, device 205 may transmit PDU 240 with a probability of p T . Similarly, device 205 may have a probability of waiting for the next idle time slot (1 - p T) Additionally or alternatively, device 205 may obtain an RBO value based on the output of RL model 230 for selection after PPDU transmission. For example, the output of RL model 230 may directly give the RBO value, and device 205 may use this RBO value accordingly during the EDCA process. Additionally or alternatively, device 205 may obtain a CW value based on the output of RL model 230 for selection after PPDU transmission. For example, the output of RL model 230 may directly give the CW value, and device 205 may use this CW value accordingly during the EDCA process. Additionally or alternatively, device 205 may obtain a set of parameters α, T based on the output of RL model 230 such that CW = T α *CW min . For example, device 205 may use CW = T during the EDCA process α *CW min to calculate the CW value. In some specific implementations, device 210 may indicate to device 205 which parameters device 205 may obtain via RL model 230 using RL model information 225.

[0067] The agent may determine a reward (such as rt) based on the success or failure of PPDU transmission, and the agent may store this reward in a reward buffer. The agent may calculate the next state S t+1 = [S t+1,0 , S t+1,1 , S t+1,2 and may store the experience e t = (S t , S t+1 , A t , R t ). In some specific implementations, the agent may refrain from updating RL model 230 (such as a Q-table) and may repeat such steps for the next PPDU transmission. According to using RL model 230 to obtain channel access decisions, device 205 may support conditional retraining. For example, it may be desirable for the agent to have a trained Q-table, and during inference, the agent tracks the average reward in the reward buffer. If the average reward buffer is less than a threshold reward (such as less than δ), then the agent may initiate retraining of RL model 230. In some aspects, the retraining may be performed similarly to the initial training process, except that during retraining, the initial Q-table is not all zeros. Instead, the initial Q-table is the current Q-table.

[0068] As used herein, depending on the context, "meeting a threshold" may mean a value greater than the threshold, greater than or equal to the threshold, less than the threshold, less than or equal to the threshold, equal to the threshold, not equal to the threshold, etc.

[0069] Device 205 may derive one or more rewards from one or more of various communication parameters that may be measured, calculated, or determined at device 205 or indicated to device 205 from device 210. Such communication parameters may include throughput metrics, latency metrics, number of collisions, packet / frame delivery ratio, or packet / frame loss ratio. In some aspects, the throughput metric may refer to the throughput observed in a last time period (such as in the last 10 seconds). In some aspects, the latency metric may refer to the average or a specific percentile latency observed for a last specific number of aggregated MAC PDUs (A-MPDUs). For example, the latency metric may be the average or 95th percentile latency observed for the last 100 A-MPDUs. In some aspects, the number of collisions may refer to the proportion of last specific number of transmissions that resulted in a collision and may be indicated by device 210 (such as AP 102). For example, the number of collisions may refer to the proportion of the last 100 transmissions that resulted in a collision. In some aspects, the packet / frame delivery ratio may refer to the ratio of successfully delivered packets / frames to the total number of packets / frames transmitted within a given time interval. In some aspects, the packet / frame loss ratio may refer to the reciprocal of the packet / frame delivery ratio.

[0070] In some embodiments, the state that device 205 may use as an input into RL model 230 may include one or more parameters associated with the environment of device 205. Device 205 may measure, determine, or calculate one or more of the state parameters, or may receive an indication of one or more of the state parameters (such as from device 210), or any combination thereof. In some aspects, device 205 may receive an indication of one or more state parameters from device 210 via parameter indication 235. These states may include one or more of the following: the proportion of successful last ′X′ transmissions, the proportion of last ′X′ transmissions that resulted in a collision, the number of interruptions during an RBO countdown, the number of channel busy periods in the last ′Y′ seconds, the number of idle periods in the last ′Y′ seconds, the number of unique receiver address (RA) field values observed within a specific time window, the number of unique transmitter address (TA) field values observed within a specific time window, the number of non-ML STAs 104 in a BSS, or the number of active STAs 104 (such as the number of contenders that may be indicated by device 210). In some aspects, the generation or determination of the number of non-ML STAs 104 in a BSS may be associated with: device 210 transmitting an indication of a list of MAC addresses with ML capabilities, device 205 observing variants of one or more A-control fields, or ML-based STA 104 using different (such as reserved for other fields) values for one or more specific fields (such as the BSS color field).

[0071] Figure 3Illustrates an example of an RL model 300 that supports RL-based EDCA according to one or more aspects of the present disclosure. The RL model 300 may be implemented or implemented to achieve or facilitate aspects of the WLAN 100 or the signaling diagram 200. For example, as illustrated and referenced by Figure 2 Illustrated and referenced Figure 2 As described, the device 205 may employ the RL model 300 to obtain an output associated with one or more channel access parameters or decisions. Further for example, the RL model 300 may be an example of the RL model 230 as illustrated and referenced by Figure 2 Illustrated and referenced Figure 2 As described.

[0072] The RL model 300 includes an agent 305 that can interact with an environment 310. For example, the RL model 300 may be a decision-making model associated with channel access and may interact with the environment to output decisions associated with channel access. In some embodiments, the agent 305 may output information associated with an action 315 (such as At) based on a reward 320 (such as R t ) and a state 325 (such as S t ). The action 315 may interact with the environment 310, affect the environment, or be affected by the environment to produce an updated state 325 (such as S t+1 ) and a fresh reward 320. The device 205 may employ the RL model 300 to obtain an output associated with the channel access process and may, in some embodiments, be permitted to retrain or refine the RL model 300.

[0073] To describe the RL algorithm, various components of the RL model 300 may be defined. Such components of the RL model 300 may include RL techniques, states 325, actions 315, policies, rewards 320, and inference processes. In some aspects, given an RL technique, training the RL model 300 may follow any supported process. In some embodiments, the RL technique may be Q-learning, deep Q-learning, or double Q-learning techniques, according to which the device 205 (or the agent 305) learns Q-values for state-action pairs. In some other embodiments, the RL technique may be policy gradient (such as REINFORCE), according to which the device 205 (or the agent 305) learns a policy. In some other embodiments, the RL technique may be an actor-critic technique (such as AC, A2C, or A3C), according to which the device 205 (or the agent 305) learns Q-values and a policy. In some other embodiments, the RL technique may be contextual MAB or non-contextual MAB. As described herein, context may be the same as state, and non-contextual MAB may not make the associated policy based on observations of the state (such as the RL model 300 not taking the state 325 as an input).

[0074] According to the specific implementation described herein, the AP 102 may provide a downloadable trained RL model 300 for use by one or more STAs 104. Additionally, in some specific implementations, the STA 104 may transmit one or more of the state variables to a peer STA 104 for calculating an updated state 325. In such specific implementations, the AP 102 may send an indication of the number of non-ML STAs or the number of active contenders or both in the BSS to the ML STA 104 (such as device 205). In some specific implementations, the AP 102 may include the indication in the A-control subfield. Additionally or alternatively, the STA 104 may piggyback (such as include or multiplex) an acknowledgment (ACK) with information related to the failure reason of one or more MPDUs. Such information related to the failure reason may indicate, for example, that the failure reason is a poor choice of modulation and coding scheme (MCS) or a conflict. Further, in some specific implementations, the first STA 104 may transmit one or more of the rewards 320 to the second STA 104 to assist the first STA 104 in training or retraining the RL model 300. For example, the AP 102 may piggyback (such as include or multiplex) information related to the signal-to-interference-plus-noise ratio (SINR) of the received PPDU.

[0075] Additionally, in some specific implementations, the AP 102 may change one or more EDCA values or parameters for the non-ML STA 104 based on how many ML devices and non-ML devices are in a given geographical area. For example, if a given area is associated with one or more ML devices that are permitted to use the RL model for channel access, the AP 102 may adjust or modify one or more EDCA values for the non-ML STA 104 to facilitate a fair channel access probability for the non-ML STA 104. In other words, the AP 102 may adjust or modify one or more EDCA values or parameters such that the non-ML STA 104 is more capable of competing with STAs 104 that are permitted to use ML-based channel access.

[0076] Figure 4 An example of an RL process 400 that supports RL-based EDCA is illustrated in accordance with one or more aspects of the present disclosure. The RL process 400 may be implemented or be implemented to achieve or facilitate aspects of the WLAN 100, the signaling diagram 200, or the RL model 300. For example, a device 205 (such as the AP 102 or the STA 104) may execute the RL process 400 as part of a Q-learning process. In some specific implementations, the device 205 may execute the RL process 400 to train or retrain the RL model 230 or the RL model 300, as shown by Figure 2 and Figure 3 illustrated and referenced by Figure 2 andFigure 3 as described

[0077] At 405, the device 205 may initialize the Q-table. For example, the device 205 may initialize the Q-table to zero such that each (state, action) pair is associated with a Q-value of 0. An example of an initialized Q-table is illustrated below by Table 1.

[0078] <![CDATA[A 0 > <![CDATA[A 1 > ... <![CDATA[A K > <![CDATA[S 0 > 0 0 ... 0 <![CDATA[S 1 > 0 0 ... 0 ... ... ... ... ... <![CDATA[S M > 0 0 ... 0

[0079] Table 1: Example of an Initialized Q-Table

[0080] At 410, the device 205 may observe the current state, which may be represented as S t . For example, the device 205 may observe the current state S of the environment of the device 205 t .

[0081] At 415, the device 205 may select an action At and perform the corresponding function. In some aspects, performing the corresponding function of the action At may be referred to as "playing". For example, the action At may be selected by a random, greedy, ε-greedy, or Boltzmann (Softmax) policy. The Boltzmann (Softmax) policy may be described by , where Q may be the Q-value and τ is the temperature factor that determines the probability of performing an action other than the action with the highest Q-value.

[0082] At 420, the device 205 may receive a reward Rt and move to the next state S t+1 .

[0083] At 425, the device 205 may use the Bellman Equation to calculate an updated Q-value. The Bellman Equation may be in the following form: Q(S t , A t ) = (1 - β)Q(S t , A t ) + β * (R t + λ * max a Q(S t+1 , a)), Q(S t , A t ) = Q(S t , A t ) + β * (R t + λ * max a Q(S t+1 , a) - Q(S t , A t )) or Q(S t , A t ) = (1 - α)Q(S t, A t ) + α * (R t + λ * max a Q(S t+1 , a)), or any combination thereof, where Q can be a Q - value, β can be a learning rate, S can be a state, A can be an action, R can be a reward, and α can be a CW level.

[0084] At 430, the device 205 can update the Q - table. Table 2 shows an example updated Q - table. The device can repeat the RL process 400 starting at 410.

[0085] <![CDATA[A 0 > <![CDATA[A 1 > ... <![CDATA[A j > ... <![CDATA[A K > <![CDATA[S 0 > <![CDATA[Q(S 0 ,A 0 )]]> <![CDATA[Q(S 0 ,A 1 )]]> ... <![CDATA[Q(S 0 ,A j )]]> ... <![CDATA[Q(S 0 ,A K )]]> <![CDATA[S 1 > <![CDATA[Q(S 1 ,A 0 )]]> <![CDATA[Q(S 1 ,A 1 )]]> ... <![CDATA[Q(S 1 ,A j )]]> ... <![CDATA[Q(S 1 ,A K )]]> ... ... ... ... ... ... ... <![CDATA[S j > <![CDATA[Q(S j ,A 0 )]]> <![CDATA[Q(S j ,A 1 )]]> ... <![CDATA[Q(S j , A j )]]> ... <![CDATA[Q(S j , A K )]]> ... ... ... ... ... ... ... <![CDATA[S M > <![CDATA[Q(S M ,A 0 )]]> <![CDATA[Q(S M ,A 1 )]]> ... <![CDATA[Q(S M ,A j )]]> ... <![CDATA[Q(S M ,A K )]]>

[0086] Table 2: Example updated Q - table

[0087] As described herein, the components of RL can include an agent, an environment, a state, an action, a reward, and a policy. In an example of the RL technique of Q - learning, the agent can be the STA 104 that learns the RL model and the environment (the environment can be an 802.11 network). The state can be a vector of three observations. In one example, the first observation can include the number of interruptions in the latest or most recent RBO countdown, the second observation can include the number of unique TA fields observed in the last or most recent 5 seconds, and the third observation can include the number of the last 10 PPDUs for which the STA 104 received an ACK. The action can be to select α such that CW = 2 α *CW min . The action space can be {0, 1, 2,..., K}, where K can be defined such that CW max = 2 K *CW min . In some embodiments, the reward can be the reciprocal of the average latency of one or more MPDUs in the transmitted PPDU, where if the PPDU transmission fails, the reward can be equal to 0. In some embodiments, the policy can be ε - greedy during training and greedy during inference. For an ε - greedy policy, the agent can select the action with the highest Q - value with a probability of (1 - ε). With a probability of ε, the agent can uniformly select one of the other actions. For example, if there are K + 1 actions and one action has the highest Q - value, the agent can select one of the other K actions, each with a probability of ε / K.

[0088] Figure 5Illustrates an example of RL process 500 that supports RL-based EDCA according to one or more aspects of the present disclosure. RL process 500 may implement or be implemented to achieve or facilitate aspects of WLAN 100, signaling diagram 200, reinforcement learning model 300, or RL process 400. For example, device 205 (such as AP 102 or STA 104) may execute RL process 500 as part of a deep Q-learning (DQN) process. In such examples, device 205 may use state 505 as an input into neural network 510 to learn, derive, generate, or otherwise approximate a Q-table, which may output a set of one or more Q-value actions 515, including Q-value action 1, Q-value action 2, and so on up to Q-value action N.

[0089] State 505 may be associated with or include one or more observations. For example, state 505 may be associated with or include a first observation (such as state s 0 ), a second observation (such as state s 1 ), and a third observation (such as state s 3 ). Additionally, in such examples, one or more Q-value actions 515 may include Q(α = 0) values, Q(α = 1) values, and so on up to Q(α = K) values.

[0090] For example, the agent may have an untrained RL model at a first time representable by t = 0. State 505 may be represented by S 0 = [S 01 , S 02 ,..., S 0M . Initially, the agent may pick a random action and values for both the learning rate β and the discount factor γ. At each time representable by t ≥ 0, the agent may execute an action task. During the action task, the agent may select an action A t according to a policy. The agent may receive a reward R t and move to the next step S t+1 . The agent may also record the transition experience e t = (S t , S t+1 , A t , R t ) and store it in a buffer.

[0091] At each time t ≥ 0, the agent may also execute a learning task. During the learning task, the agent may compute a loss L, which may be represented by L = (Q Target - Q(S t , A t )) 2 , where Q Target = R t+γmax a Q(S t+1 , a). The agent can compute the gradient of the loss, which can be represented by , where Φ can be the parameters of a deep neural network (DNN). The agent can backpropagate the gradient to minimize the loss. For example, the previous DNN parameters Φ′ can be updated to be represented by . If Q Target is considered the true label and Q(S t , A t ) is the predicted label, the learning can be the same as that of a DNN for supervised learning. This learning process can simulate the temporal-difference Bellman equation used at 425. If, for example, the agent has more experience and thus uses smaller parameter updates, the learning rate β can decay over time. After sufficient iterations, the DQN model learned by the agent can provide a sufficiently accurate representation of the converged Q-table. The resulting DQN can be referred to as the trained model and can be used for channel access.

[0092] Figure 6 illustrates an example of a process flow 600 that supports RL-based EDCA according to one or more aspects of the present disclosure. The process flow 600 can implement or be implemented to achieve aspects of the WLAN 100, the signaling diagram 200, the RL model 300, the RL process 400, or the RL process 500. For example, the process flow 600 illustrates communication between the device 205 and the device 210, which is illustrated by Figure 2 and described with reference to Figure 2 for these devices.

[0093] In the following description of the process flow 600, operations (such as reporting or providing) can be performed in an order different from the shown order, or operations performed by the example devices can be performed in a different order or at different times. For example, a particular operation can also be omitted from the process flow 600, or other operations can be added to the process flow 600. Additionally, although some operations or signaling are shown as occurring at different times for discussion purposes, these operations can actually occur simultaneously.

[0094] At 605, device 205 may receive information associated with an RL model from device 210. In some embodiments, the RL model may be associated with performing a distributed channel access procedure at device 205 in a WLAN based on this information. The information associated with the RL model may include various contents according to decisions at a control entity (such as device 210). In some embodiments, the information may include an indication that device 205 is allowed (e.g., permitted or authorized) to develop the RL model and use the RL model for the distributed channel access procedure. In some embodiments, the information may include the configuration of the RL model and an indication of whether device 205 is allowed to retrain the RL model. In some embodiments, the information may include an indication that device 205 is allowed to retrain the RL model for the distributed channel access procedure, where the RL model may be pre-loaded or pre-configured at device 205.

[0095] In some embodiments, the information may include an indication of which channel access parameters device 205 may use the RL model to determine. In some embodiments, the information may include an indication of one or more parameters (such as one or more hyperparameters) associated with an RL technique that device 205 is to follow when training or retraining the RL model. In some embodiments, the information may include an indication of one or more rewards associated with the RL model. In such embodiments, for example, device 210 may indicate what rewards device 205 may obtain from the RL model, may indicate an upper limit associated with one or more rewards (e.g., to control potential retraining of the RL model by device 205), or any combination thereof.

[0096] In some embodiments, the information may include an indication of a strategy that the device may use for training or retraining the RL model. Example strategies may include random, random within a range (e.g., uniformly random within a specified range), greedy, ε-greedy, Boltzmann (such as Softmax), or one or more of the EDCA protocol. Additionally, the strategy may refer to how device 205 (e.g., the agent of the RL model used by device 205) maps states to actions or learns to map states to actions to increase (e.g., maximize) one or more rewards, or is associated with these operations.

[0097] At 610, device 205 may receive an indication of one or more parameters associated with the environment of device 205 from device 210. In some embodiments, the output of the RL model may be associated with using one or more parameters as inputs into the RL model. The one or more parameters include one or more of the following: the proportion of successfully transmitted PDU's in the most recent set of PDU's, the proportion of PDU's in the most recent set that result in a collision, the number of interruptions during the RBO countdown during the distributed channel access procedure, the number of channel busy periods during a most recent time period, the number of channel idle periods during a most recent time period, the number of unique RA field values observed during a time window, the number of unique TA field values observed during a time window, the number of devices within the BSS that are not able to use reinforcement learning for channel access, or the number of active devices within the BSS.

[0098] At 615, device 205 may develop an RL model. For example, if device 210 indicates that device 205 is permitted to develop an RL model, then device 205 may develop an RL model. In some aspects, developing an RL model may equivalently be understood as initially creating the RL model (e.g., initially training or learning the RL model). For example, if an RL model has not been created or configured for device 205, then device 210 may instruct device 205 (e.g., based on information communicated at 605 and / or 610) to develop or initially create the model. Device 205 may initially develop an RL model for use based on the information communicated at 605 and / or 610, rather than receiving an RL model or an indication of an RL model.

[0099] At 620, device 205 may (re)train the RL model. For example, if device 210 indicates that device 205 is permitted to train or retrain (such as refine) the RL model, then device 210 may (re)train the RL model. Generally, device 205 may selectively train or retrain the RL model based on whether device 210 indicates that device 205 is permitted to train or retrain the RL model.

[0100] At 625, device 205 may derive one or more rewards associated with the RL model based on one or more communication parameters associated with communications to or from device 205. In some embodiments, device 205 may derive one or more rewards based on one or more of the following: SINR, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted PDUs to the total number of PDUs transmitted, or ratio of the number of unsuccessfully transmitted PDUs to the total number of PDUs transmitted. Device 205 may store the one or more rewards in a reward buffer and, in some aspects, may track the average reward over time. In some embodiments, if device 205 is permitted, device 205 may trigger retraining of the RL model in the event that the average reward fails to meet a threshold. In some embodiments, the rewards derived by device 205 may be associated with the initial training or development of the RL model.

[0101] At 630, device 205 may perform a distributed channel access procedure (depending on the type of output of the RL model) and may transmit a PDU during a time slot based on the output of the RL model. For example, device 205 may obtain a transmission probability, an RBO value, a CW duration, or values of one or more parameters (e.g., values of one or both of α or T) as an output of the RL model, device 205 may calculate the CW duration therefrom, and device 205 may identify the time slot during which to transmit the PDU based on any of such outputs of the RL model. In some embodiments, device 205 may identify the time slot during which to transmit the PDU based on achieving a suitable reward associated with the RL model, where such suitable reward may be achieved during the initial development or retraining of the RL model at one or both of device 205 and device 210. In some embodiments, device 205 may use the output of the RL model to make channel access decisions for one or more transmissions. For example, device 205 may always be permitted to use the RL model for channel access decision-making, or may be permitted to use the RL model for channel access decision-making at a specified time or under specified conditions (e.g., as indicated by device 210).

[0102] In addition, as described herein, time slots based on the output of the RL model can be selected, identified, or otherwise used for the transmission of PDUs in various ways depending on the type of output of the RL model. For example, if the RL model outputs a transmission probability, the time slot can be based on the output of the RL model according to the actual transmission of the PDU during the time slot being associated with the transmission probability. Further for example, if the RL model outputs an RBO counter value, the time slot can be based on the output of the RL model according to the device 205 transmitting the PDU once RBO = 0 (e.g., according to the EDCA protocol), where the RBO counter value of 0 can be associated with a specific time slot based on the initial value of the RBO counter and the number of busy time slots since the start of the CW. Further for example, if the RL model outputs the CW duration or the value of one or more parameters by which the device 205 can calculate the CW duration, the time slot can be based on the output of the RL model according to the device 205 transmitting the PDU during the time slot within the CW duration. In addition, the device 205 can transmit the PDU according to the distributed channel access procedure by transmitting the PDU when the RBO counter value is equal to 0 or by otherwise transmitting the PDU during a time slot sensed as idle (e.g., not busy or not used by another device).

[0103] Figure 7 FIG. 700 is a flow chart illustrating an example process 700 that can be performed at a wireless STA supporting RL-based EDCA in accordance with one or more aspects of the present disclosure. The operations of process 700 can be implemented by a wireless STA or its components as described herein. For example, process 700 can be performed by a wireless communication device that acts as or operates within a wireless STA (such as the wireless communication device 900 described with reference to Figure 9 ). In some specific implementations, process 700 can be performed by a wireless STA (such as one of the STAs 104 described with reference to Figure 1 ).

[0104] At 702, the method can include receiving information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at a wireless communication device in a wireless local area network according to the information. The operation of 705 can be performed according to the examples disclosed herein.

[0105] At 704, the method can include transmitting a protocol data unit according to the distributed channel access procedure and during a time slot based on the output of the reinforcement learning model. The operation of 704 can be performed according to the examples disclosed herein.

[0106] Figure 8FIG. 800 is a flow chart showing an example process 800 that can be performed at a wireless AP supporting RL-based EDCA in accordance with one or more aspects of the present disclosure. The operations of process 800 can be implemented by a wireless AP or its components as described herein. For example, process 800 can be performed by a wireless communication device (such as the wireless communication device 1000 described with reference to Figure 10 ), which acts as a wireless AP or operates within a wireless AP. In some specific implementations, process 800 can be performed by a wireless AP (such as one of the APs 102 described with reference to Figure 1 ).

[0107] At 802, the method can include transmitting information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network based on the information. The operation of 802 can be performed according to the examples disclosed herein.

[0108] At 804, the method can include receiving a protocol data unit from the second wireless communication device during a time slot based on the output of the reinforcement learning model and according to the distributed channel access procedure. The operation of 804 can be performed according to the examples disclosed herein.

[0109] Figure 9 FIG. 900 is a block diagram showing an example wireless communication device 900 that supports RL-based EDCA in accordance with some aspects of the present disclosure. In some specific implementations, the wireless communication device 900 is configured to or operable to perform process 700 described with reference to Figure 7 . In various examples, the wireless communication device 900 can be a chip, SoC, chipset, package, or device that can include one or more modems (such as a Wi-Fi (IEEE 802.11) modem or a cellular modem (such as a 3GPP 4G LTE or 5G compatible modem), one or more processors, processing blocks, or processing elements (collectively referred to as "processors"); one or more radio components (collectively referred to as "radio components"); and one or more memories or storage blocks (collectively referred to as "memory").

[0110] In some specific implementations, the wireless communication device 900 can be a device for a STA (such as the STA described with reference to Figure 1The device in the described STA104). In some other examples, the wireless communication device 900 can be a STA that includes such a chip, SoC, chipset, package, or device and multiple antennas. The wireless communication device 900 is capable of sending and receiving wireless communications, for example, in the form of wireless packets. For example, the wireless communication device can be configured or operable to send and receive packets in the form of physical layer PPDUs and MPDUs that follow one or more of the IEEE 802.11 wireless communication protocol standards. In some specific implementations, the wireless communication device 900 further includes an application processor or can be coupled to an application processor, and the application processor can be further coupled to another memory. In some specific implementations, the wireless communication device 900 further includes a user interface (UI) (such as a touch screen or keypad) and a display, and the display can be integrated with the UI to form a touch screen display. In some specific implementations, the wireless communication device 900 can further include one or more sensors, such as one or more inertial sensors, accelerometers, temperature sensors, pressure sensors, or altitude sensors.

[0111] The wireless communication device 900 includes an RL model component 902, a PDU component 904, an RL model development component 906, an RL model training component 908, an RL model policy component 910, and an RL model reward component 912. Portions of one or more of the components 902, 904, 906, 908, 910, and 912 can be implemented at least partially in hardware or firmware. For example, the PDU component 904 can be implemented at least partially by a modem. In some specific implementations, at least some of the components 902, 904, 906, 908, 910, and 912 are implemented at least partially by a processor and implemented as software stored in a memory. For example, portions of one or more of the components 902, 904, 906, 908, 910, and 912 can be implemented as non-transitory instructions (or "code") that can be executed by a processor to perform the functions or operations of the corresponding modules.

[0112] In some specific implementations, the processor may be a component of a processing system. A processing system generally refers to a system or a series of machines or components that receive inputs and process these inputs to generate a set of outputs (which can be passed to other systems or components such as device 900). For example, the processing system of device 900 may refer to a system that includes various other components or sub-components of device 900 (such as a processor, or a transceiver, or a communication manager, or a combination of other components or components of device 900). The processing system of device 900 may interface with other components of device 900 and may process information (such as inputs or signals) received from other components or output information to other components. For example, a chip or a modem of device 900 may include a processing system, a first interface for outputting information, and a second interface for obtaining information. In some specific implementations, the first interface may refer to the interface between the processing system of the chip or the modem and the transmitter, such that device 900 can transmit the information output from the chip or the modem. In some specific implementations, the second interface may refer to the interface between the processing system of the chip or the modem and the receiver, such that device 900 can obtain information or signal inputs, and the information can be passed to the processing system. Those of ordinary skill in the art will readily recognize that the first interface may also obtain information or signal inputs, and the second interface may also output information or signal outputs.

[0113] The RL model component 902 may be capable of, configured to, or operable to receive information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at a wireless communication device in a wireless local area network. The PDU component 904 may be capable of, configured to, or operable to send a protocol data unit during a time slot based on the output of the reinforcement learning model and according to the distributed channel access procedure.

[0114] In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model development component 906 may be capable of, configured to, or operable to receive an indication that the wireless communication device is allowed to develop a reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

[0115] In some specific implementations, the RL model development component 906 may be capable of, configured to, or operable to develop a reinforcement learning model at the wireless communication device according to the received indication that the wireless communication device is allowed to develop a reinforcement learning model.

[0116] In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model component 902 may be capable of, configured to, or operable to receive a configuration associated with the reinforcement learning model and an indication of whether the wireless communication device is allowed to retrain the reinforcement learning model for a distributed channel access process. The method further includes. In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model training component 908 may be capable of, configured to, or operable to selectively retrain the reinforcement learning model based on whether the wireless communication device is allowed to retrain the reinforcement learning model.

[0117] In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model training component 908 may be capable of, configured to, or operable to receive an indication that the wireless communication device is allowed to retrain the reinforcement learning model for a distributed channel access process, where the reinforcement learning model is pre-loaded at the wireless communication device. The method further includes. In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model training component 908 may be capable of, configured to, or operable to retrain the reinforcement learning model based on the wireless communication device being allowed to retrain the reinforcement learning model.

[0118] In some specific implementations, the RL model component 902 may be capable of, configured to, or operable to receive an indication that the wireless communication device is allowed to use the reinforcement learning model to obtain one or more of the following: a transmission probability, a backoff counter, a duration of a contention window, or one or more parameters associated with the duration of the contention window, where the output of the reinforcement learning model includes a transmission probability, a backoff counter, a duration of a contention window, or one or more parameters associated with the duration of the contention window.

[0119] In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model policy component 910 may be capable of, configured to, or operable to receive an indication of a policy associated with training or retraining the reinforcement learning model. The method further includes. In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model training component 908 may be capable of, configured to, or operable to train or retrain the reinforcement learning model according to the policy.

[0120] In some specific implementations, the policy includes an enhanced distributed channel access protocol or a distributed coordination function protocol. In some specific implementations, one or more parameters associated with using the enhanced distributed channel access protocol or the distributed coordination function protocol are stored in a buffer of the wireless communication device. In some specific implementations, training or retraining the reinforcement learning model is based on one or more parameters.

[0121] In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model component 902 may be capable of, configured to, or operable to receive an indication of one or more parameters associated with a reinforcement learning technique that a wireless communication device is to follow when training or retraining a reinforcement learning model, where the reinforcement learning technique is associated with Q-learning techniques, policy gradients, actor-critic techniques, or contextual multi-armed bandit (MAB) or non-contextual MAB techniques. The method further includes. In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model training component 908 may be capable of, configured to, or operable to train or retrain a reinforcement learning model based on one or more parameters associated with the reinforcement learning technique.

[0122] In some specific implementations, to support transmitting a protocol data unit (PDU) according to a distributed channel access procedure and during a time slot based on the output of a reinforcement learning model, the PDU component 904 may be capable of, configured to, or operable to attempt to transmit the PDU during one or more idle time slots according to a transmission probability, where the transmission probability is the output of the reinforcement learning model.

[0123] In some specific implementations, to support transmitting a protocol data unit (PDU) according to a distributed channel access procedure and during a time slot based on the output of a reinforcement learning model, the PDU component 904 may be capable of, configured to, or operable to transmit the PDU during the time slot according to the expiration of a backoff counter, where the backoff counter is the output of the reinforcement learning model.

[0124] In some specific implementations, to support transmitting a protocol data unit (PDU) according to a distributed channel access procedure and during a time slot based on the output of a reinforcement learning model, the PDU component 904 may be capable of, configured to, or operable to transmit the PDU during a contention window, where the duration of the contention window is associated with the output of the reinforcement learning model.

[0125] In some specific implementations, the output of the reinforcement learning model includes the absolute value of the duration of the contention window or one or more parameters associated with a multiplication factor. In some specific implementations, the duration of the contention window is equal to the product of the minimum contention window duration and the multiplication factor.

[0126] In some specific implementations, to support receiving information associated with a reinforcement learning model, the RL model reward component 912 may be capable of, configured to, or operable to receive an indication of one or more rewards associated with the reinforcement learning model, where the one or more rewards include one or more of the following: signal-to-interference-plus-noise ratio, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted protocol data units to the total number of protocol data units transmitted or ratio of the number of unsuccessfully transmitted protocol data units to the total number of protocol data units transmitted, where the output of the reinforcement learning model is at least partially based on the one or more rewards.

[0127] In some specific implementations, one or more rewards are stored in a reward buffer associated with the reinforcement learning model.

[0128] In some specific implementations, the RL model reward component 912 may be capable of, configured to, or operable to derive one or more rewards associated with the reinforcement learning model based on one or more of the following: signal-to-interference-plus-noise ratio, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted protocol data units to the total number of protocol data units transmitted or ratio of the number of unsuccessfully transmitted protocol data units to the total number of protocol data units transmitted, where the output of the reinforcement learning model is at least partially based on the one or more rewards.

[0129] In some specific implementations, one or more updated parameters associated with the environment of a wireless communication device are calculated based on transmitting protocol data units according to a distributed channel access procedure. In some specific implementations, the one or more updated parameters are stored at the wireless communication device. In some specific implementations, potential retraining of the reinforcement learning model is associated with the one or more updated parameters.

[0130] In some specific implementations, the RL model component 902 may be capable of, configured to, or operable to identify an updated state associated with the wireless communication device based on transmitting protocol data units. In some specific implementations, the RL model component 902 may be capable of, configured to, or operable to input the updated state into the reinforcement learning model to obtain an updated output of the reinforcement learning model. In some specific implementations, the PDU component 904 may be capable of, configured to, or operable to transmit a second protocol data unit according to the distributed channel access procedure and during a second time slot based on the updated output of the reinforcement learning model.

[0131] In some specific implementations, the RL model component 902 may be capable of, configured to, or operable to receive an indication of one or more parameters associated with the environment of the wireless communication device, where the output of the reinforcement learning model is associated with using the one or more parameters as inputs into the reinforcement learning model.

[0132] In some specific implementations, one or more parameters include one or more of the following: the proportion of successfully transmitted protocol data units in the most recent set, the proportion of protocol data units in the most recent set that result in collisions, the number of interruptions during the random backoff countdown period during the distributed channel access process, the number of channel busy periods during the most recent time period, the number of channel idle periods during the most recent time period, the number of unique receiver address field values observed during a time window, the number of unique transmitter address field values observed during a time window, the number of devices within the basic service set that are unable to use reinforcement learning for channel access, or the number of active devices within the basic service set.

[0133] In some specific implementations, the reinforcement learning model is a decision-making model developed based on interactions with the environment. In some specific implementations, the output of the reinforcement learning model includes decisions associated with channel access.

[0134] Figure 10 A block diagram of an example wireless communication device 1000 that supports RL-based EDCA in accordance with some aspects of the present disclosure is shown. In some specific implementations, the wireless communication device 1000 is configured to or operable to perform the process 800 described with reference to Figure 8 In various examples, the wireless communication device 1000 may be a chip, SoC, chipset, package, or device that may include the following: one or more modems (such as a Wi-Fi (IEEE 802.11) modem or a cellular modem (such as a 3GPP 4G LTE or 5G compatible modem); one or more processors, processing blocks, or processing elements (collectively referred to as "processors"); one or more radio components (collectively referred to as "radio components"); and one or more memories or storage blocks (collectively referred to as "memory").

[0135] In some specific implementations, the wireless communication device 1000 may be for an AP (such as with reference to Figure 1The device in the described AP102). In some other examples, the wireless communication device 1000 may be an AP that includes such a chip, SoC, chipset, package, or device, as well as multiple antennas. The wireless communication device 1000 is capable of sending and receiving wireless communications, for example, in the form of wireless packets. For example, the wireless communication device 1000 may be configured or operable to send and receive packets in the form of physical layer PPDUs and MPDUs that conform to one or more of the IEEE 802.11 wireless communication protocol standards. In some specific implementations, the wireless communication device 1000 further includes an application processor or is coupled to an application processor, which may be further coupled to another memory. In some specific implementations, the wireless communication device 1000 further includes at least one external network interface that enables communication with a core network or a backhaul network to obtain access to an external network including the Internet.

[0136] The wireless communication device 1000 includes an RL model component 1002, a PDU component 1004, an RL model development component 1006, an RL model training component 1008, an RL model policy component 1010, and an RL model reward component 1012, or any combination thereof. Portions of one or more of the components 1002, 1004, 1006, 1008, 1010, and 1012 may be at least partially implemented in hardware or firmware. For example, the PDU component 1004 may be at least partially implemented by a modem. In some specific implementations, at least some of the components 1002, 1004, 1006, 1008, 1010, and 1012 are at least partially implemented by a processor and implemented as software stored in a memory. For example, portions of one or more of the components 1002, 1004, 1006, 1008, 1010, and 1012 may be implemented as non-transitory instructions (or "code") executable by a processor to perform the functions or operations of the corresponding modules.

[0137] In some specific implementations, the processor may be a component of a processing system. A processing system generally refers to a system or a series of machines or components that receive inputs and process these inputs to produce a set of outputs (which can be passed to other systems or components such as device 1000). For example, the processing system of device 1000 may refer to a system that includes various other components or sub-components of device 1000 (such as a processor, or a transceiver, or a communication manager, or a combination of other components or components of device 1000). The processing system of device 1000 may interface with other components of device 1000 and may process information (such as inputs or signals) received from other components or output information to other components. For example, a chip or a modem of device 1000 may include a processing system, a first interface for outputting information, and a second interface for obtaining information. In some specific implementations, the first interface may refer to the interface between the processing system of the chip or the modem and the transmitter, such that device 1000 can transmit the information output from the chip or the modem. In some specific implementations, the second interface may refer to the interface between the processing system of the chip or the modem and the receiver, such that device 1000 can obtain information or signal inputs, and the information can be passed to the processing system. Those of ordinary skill in the art will readily recognize that the first interface can also obtain information or signal inputs, and the second interface can also output information or signal outputs.

[0138] The RL model component 1002 may be capable of, configured to, or operable to send information associated with a reinforcement learning model, where the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network according to the information. The PDU component 1004 may be capable of, configured to, or operable to receive a protocol data unit from the second wireless communication device during a time slot based on the output of the reinforcement learning model according to the distributed channel access procedure.

[0139] In some specific implementations, to support sending information associated with a reinforcement learning model, the RL model development component 1006 may be capable of, configured to, or operable to send an indication that the second wireless communication device is allowed to develop a reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

[0140] In some specific implementations, to support sending information associated with a reinforcement learning model, the RL model component 1002 may be capable of, configured to, or operable to send information associated with the reinforcement learning model and an indication of whether the second wireless communication device is allowed to retrain the reinforcement learning model for the distributed channel access procedure.

[0141] In some specific implementations, to support sending information associated with a reinforcement learning model, the RL model training component 1008 may be capable of, configured to, or operable to send an indication that a second wireless communication device is permitted to retrain a reinforcement learning model for a distributed channel access process, where the reinforcement learning model is pre-loaded at the second wireless communication device.

[0142] In some specific implementations, the RL model component 1002 may be capable of, configured to, or operable to send an indication that a second wireless communication device is permitted to use a reinforcement learning model to obtain one or more of the following: a transmission probability, a backoff counter, a duration of a contention window, or one or more parameters associated with the duration of the contention window, where the output of the reinforcement learning model includes a transmission probability, a backoff counter, a duration of a contention window, or one or more parameters associated with the duration of the contention window.

[0143] In some specific implementations, to support sending information associated with a reinforcement learning model, the RL model policy component 1010 may be capable of, configured to, or operable to send an indication of a policy associated with training or retraining a reinforcement learning model.

[0144] In some specific implementations, the policy includes an enhanced distributed channel access protocol or a distributed coordination function protocol.

[0145] In some specific implementations, to support sending information associated with a reinforcement learning model, the RL model component 1002 may be capable of, configured to, or operable to send an indication of one or more parameters associated with a reinforcement learning technique that a second wireless communication device is to follow when training or retraining a reinforcement learning model, where the reinforcement learning technique is associated with a Q-learning technique, policy gradient, actor-critic technique, or context multi-armed bandit (MAB) or non-contextual MAB technique.

[0146] In some specific implementations, to support sending information associated with a reinforcement learning model, the RL model reward component 1012 may be capable of, configured to, or operable to send an indication of one or more rewards associated with the reinforcement learning model, where the one or more rewards include one or more of the following: a signal-to-interference-plus-noise ratio, a throughput metric, a latency metric, a number of collisions, a ratio of the number of successfully transmitted protocol data units to the total number of transmitted protocol data units, or a ratio of the number of unsuccessfully transmitted protocol data units to the total number of transmitted protocol data units, where the output of the reinforcement learning model is at least partially based on the one or more rewards.

[0147] In some specific implementations, the PDU component 1004 may be capable of, configured to, or operable to receive a second protocol data unit according to a distributed channel access procedure and during a second time slot based on an updated output of a reinforcement learning model.

[0148] In some specific implementations, the RL model component 1002 may be capable of, configured to, or operable to send an indication of one or more parameters associated with the environment of a second wireless communication device, wherein the output of the reinforcement learning model is associated with using the one or more parameters as inputs into the reinforcement learning model.

[0149] In some specific implementations, the one or more parameters include one or more of the following: the proportion of protocol data units in a most recent set that were successfully transmitted, the proportion of protocol data units in a most recent set that resulted in a collision, the number of interruptions during a random backoff countdown during a distributed channel access procedure, the number of channel busy periods during a most recent time period, the number of channel idle periods during a most recent time period, the number of unique receiver address field values observed during a time window, the number of unique transmitter address field values observed during a time window, the number of devices within a basic service set that are not able to use reinforcement learning for channel access, or the number of active devices within a basic service set.

[0150] In some specific implementations, the reinforcement learning model is a decision-making model developed based on interactions with the environment. In some specific implementations, the output of the reinforcement learning model includes a decision associated with channel access.

[0151] Specific implementation examples are described in the following numbered clauses:

[0152] Clause 1: A method for wireless communication at a wireless communication device, comprising: receiving information associated with a reinforcement learning model, wherein the reinforcement learning model is associated with performing a distributed channel access procedure at the wireless communication device in a wireless local area network according to the information; and transmitting a protocol data unit according to the distributed channel access procedure and during a time slot at least partially based on an output of the reinforcement learning model.

[0153] Clause 2: The method according to clause 1, wherein receiving the information associated with the reinforcement learning model comprises: receiving an indication that the wireless communication device is permitted to develop the reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

[0154] Clause 3: The method according to clause 2, further comprising: developing the reinforcement learning model at the wireless communication device according to the received indication that the wireless communication device is permitted to develop the reinforcement learning model.

[0155] Clause 4: The method according to any one of Clauses 1 to 3, wherein receiving the information associated with the reinforcement learning model includes: receiving the information associated with the reinforcement learning model and an indication of whether the wireless communication device is permitted to retrain the reinforcement learning model for the distributed channel access process, and the method further includes: selectively retraining the reinforcement learning model at least in part based on whether the wireless communication device is permitted to retrain the reinforcement learning model.

[0156] Clause 5: The method according to any one of Clauses 1 to 4, wherein receiving the information associated with the reinforcement learning model includes: receiving an indication that the wireless communication device is permitted to retrain the reinforcement learning model for the distributed channel access process, wherein the reinforcement learning model is pre-loaded at the wireless communication device, and the method further includes: retraining the reinforcement learning model at least in part based on the wireless communication device being permitted to retrain the reinforcement learning model.

[0157] Clause 6: The method according to any one of Clauses 1 to 5, further including: receiving an indication that the wireless communication device is permitted to use the reinforcement learning model to obtain one or more of the following: a transmission probability, a backoff counter, a duration of a contention window, or one or more parameters associated with the duration of the contention window, wherein an output of the reinforcement learning model includes the transmission probability, the backoff counter, the duration of the contention window, or the one or more parameters associated with the duration of the contention window.

[0158] Clause 7: The method according to any one of Clauses 1 to 6, wherein receiving the information associated with the reinforcement learning model includes: receiving an indication of a policy associated with training or retraining the reinforcement learning model, and the method further includes: training or retraining the reinforcement learning model according to the policy.

[0159] Clause 8: The method according to Clause 7, wherein the policy includes an enhanced distributed channel access protocol or a distributed coordination function protocol, one or more parameters associated with using the enhanced distributed channel access protocol or the distributed coordination function protocol are stored in a buffer of the wireless communication device, and training or retraining the reinforcement learning model is at least in part based on the one or more parameters.

[0160] Clause 9: The method according to any one of Clauses 1 to 8, wherein receiving the information associated with the reinforcement learning model includes: receiving an indication of one or more parameters associated with a reinforcement learning technique to be followed by the wireless communication device when training or retraining the reinforcement learning model, wherein the reinforcement learning technique is associated with a Q-learning technique, policy gradient, actor-critic technique, or context multi-armed bandit (MAB) or non-contextual MAB technique, and the method further includes: training or retraining the reinforcement learning model according to the one or more parameters associated with the reinforcement learning technique.

[0161] Clause 10: The method according to any one of Clauses 1 to 9, wherein transmitting the protocol data unit during the time slot based at least in part on the output of the reinforcement learning model according to the distributed channel access procedure includes: attempting to transmit the protocol data unit during one or more idle time slots according to a transmission probability, wherein the transmission probability is the output of the reinforcement learning model.

[0162] Clause 11: The method according to any one of Clauses 1 to 10, wherein transmitting the protocol data unit during the time slot based at least in part on the output of the reinforcement learning model according to the distributed channel access procedure includes: transmitting the protocol data unit during the time slot according to the expiration of a backoff counter, wherein the backoff counter is the output of the reinforcement learning model.

[0163] Clause 12: The method according to any one of Clauses 1 to 11, wherein transmitting the protocol data unit during the time slot based at least in part on the output of the reinforcement learning model according to the distributed channel access procedure includes: transmitting the protocol data unit during a contention window, wherein the duration of the contention window is associated with the output of the reinforcement learning model.

[0164] Clause 13: The method according to Clause 12, wherein the output of the reinforcement learning model includes the absolute value of the duration of the contention window or one or more parameters associated with a multiplication factor, and the duration of the contention window is equal to the product of a minimum contention window duration and the multiplication factor.

[0165] Clause 14: The method according to any one of Clauses 1 to 13, wherein receiving the information associated with the reinforcement learning model includes: receiving an indication of one or more rewards associated with the reinforcement learning model, wherein the one or more rewards include one or more of the following: signal-to-interference-plus-noise ratio, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted protocol data units to the total number of transmitted protocol data units, or ratio of the number of unsuccessfully transmitted protocol data units to the total number of transmitted protocol data units, and wherein the output of the reinforcement learning model is at least partially based on the one or more rewards.

[0166] Clause 15: The method according to Clause 14, wherein the one or more rewards are stored in a reward buffer associated with the reinforcement learning model.

[0167] Clause 16: The method according to any one of Clauses 1 to 15, further comprising: deriving one or more rewards associated with the reinforcement learning model based on one or more of the following: signal-to-interference-plus-noise ratio, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted protocol data units to the total number of transmitted protocol data units, or ratio of the number of unsuccessfully transmitted protocol data units to the total number of transmitted protocol data units, and wherein the output of the reinforcement learning model is at least partially based on the one or more rewards.

[0168] Clause 17: The method according to any one of Clauses 1 to 16, wherein one or more updated parameters associated with the environment of the wireless communication device are calculated at least partially based on transmitting the protocol data unit according to the distributed channel access procedure, the one or more updated parameters are stored at the wireless communication device, and potential retraining of the reinforcement learning model is associated with the one or more updated parameters.

[0169] Clause 18: The method according to any one of Clauses 1 to 17, further comprising: obtaining an updated state associated with the wireless communication device based on transmitting the protocol data unit; inputting the updated state into the reinforcement learning model to obtain an updated output of the reinforcement learning model; and transmitting a second protocol data unit according to the distributed channel access procedure and during a second time slot at least partially based on the updated output of the reinforcement learning model.

[0170] Clause 19: The method according to any one of Clauses 1 to 18 further comprises: receiving an indication of one or more parameters associated with the environment of the wireless communication device, wherein the output of the reinforcement learning model is associated with using the one or more parameters as inputs into the reinforcement learning model.

[0171] Clause 20: The method according to Clause 19, wherein the one or more parameters comprise one or more of the following: the proportion of protocol data units in a most recent set that were successfully transmitted, the proportion of the most recent set of protocol data units that resulted in a collision, the number of interruptions during a random backoff countdown during the distributed channel access procedure, the number of channel busy periods during a most recent time period, the number of channel idle periods during the most recent time period, the number of unique receiver address field values observed during a time window, the number of unique transmitter address field values observed during the time window, the number of devices within a basic service set that are not able to use reinforcement learning for channel access, or the number of active devices within the basic service set.

[0172] Clause 21: The method according to any one of Clauses 1 to 20, wherein the reinforcement learning model is a decision-making model developed based on interactions with the environment, and the output of the reinforcement learning model comprises a decision associated with channel access.

[0173] Clause 22: A method for wireless communication at a wireless communication device, comprising: transmitting information associated with a reinforcement learning model, wherein the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network based on the information; and receiving a protocol data unit from the second wireless communication device during a time slot according to the distributed channel access procedure and at least partially based on the output of the reinforcement learning model.

[0174] Clause 23: The method according to Clause 22, wherein transmitting the information associated with the reinforcement learning model comprises: transmitting an indication that the second wireless communication device is permitted to develop the reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

[0175] Clause 24: The method according to any one of Clauses 22 to 23, wherein transmitting the information associated with the reinforcement learning model comprises: transmitting information associated with the reinforcement learning model and an indication as to whether the second wireless communication device is permitted to retrain the reinforcement learning model for the distributed channel access procedure.

[0176] Clause 25: The method according to any one of Clauses 22 to 24, wherein transmitting the information associated with the reinforcement learning model includes: transmitting an indication that the second wireless communication device is allowed to retrain the reinforcement learning model for the distributed channel access process, wherein the reinforcement learning model is pre-loaded at the second wireless communication device.

[0177] Clause 26: The method according to any one of Clauses 22 to 25, further comprising: transmitting an indication that the second wireless communication device is allowed to use the reinforcement learning model to obtain one or more of the following: transmission probability, backoff counter, duration of the contention window, or one or more parameters associated with the duration of the contention window, wherein the output of the reinforcement learning model includes the transmission probability, the backoff counter, the duration of the contention window, or the one or more parameters associated with the duration of the contention window.

[0178] Clause 27: The method according to any one of Clauses 22 to 26, wherein transmitting the information associated with the reinforcement learning model includes: transmitting an indication of a policy associated with training or retraining the reinforcement learning model.

[0179] Clause 28: The method according to Clause 27, wherein the policy includes an enhanced distributed channel access protocol or a distributed coordination function protocol.

[0180] Clause 29: The method according to any one of Clauses 22 to 28, wherein transmitting the information associated with the reinforcement learning model includes: transmitting an indication of one or more parameters associated with a reinforcement learning technique that the second wireless communication device is to follow when training or retraining the reinforcement learning model, wherein the reinforcement learning technique is associated with a Q-learning technique, policy gradient, actor-critic technique, or contextual multi-armed bandit (MAB) or non-contextual MAB technique.

[0181] Clause 30: The method according to any one of Clauses 22 to 29, wherein transmitting the information associated with the reinforcement learning model includes: transmitting an indication of one or more rewards associated with the reinforcement learning model, wherein the one or more rewards include one or more of the following: signal-to-interference-plus-noise ratio, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted protocol data units to the total number of protocol data units transmitted, or ratio of the number of unsuccessfully transmitted protocol data units to the total number of protocol data units transmitted, wherein the output of the reinforcement learning model is at least partially based on the one or more rewards.

[0182] Clause 31: The method according to any one of Clauses 22 to 30 further comprises: receiving a second protocol data unit during a second time slot according to the distributed channel access procedure and at least partially based on an updated output of the reinforcement learning model.

[0183] Clause 32: The method according to any one of Clauses 22 to 31 further comprises: transmitting an indication of one or more parameters associated with the environment of the second wireless communication device, wherein the output of the reinforcement learning model is associated with using the one or more parameters as inputs into the reinforcement learning model.

[0184] Clause 33: The method according to Clause 32, wherein the one or more parameters comprise one or more of the following: a proportion of protocol data units successfully transmitted in a most recent set of protocol data units, a proportion of the most recent set of protocol data units that resulted in a collision, a number of interruptions during a random backoff countdown during the distributed channel access procedure, a number of channel busy periods during a most recent time period, a number of channel idle periods during the most recent time period, a number of unique receiver address field values observed during a time window, a number of unique transmitter address field values observed during the time window, a number of devices within a basic service set that are not able to perform channel access using reinforcement learning, or a number of active devices within the basic service set.

[0185] Clause 34: The method according to any one of Clauses 22 to 33, wherein the reinforcement learning model is a decision-making model developed based on an interaction with the environment, and the output of the reinforcement learning model comprises a decision associated with channel access.

[0186] Clause 35: An apparatus for wireless communication at a wireless communication device, comprising: a processor; a memory coupled to the processor; and instructions stored in the memory and executable by the processor to cause the apparatus to perform the method according to any one of Clauses 1 to 21.

[0187] Clause 36: An apparatus for wireless communication at a wireless communication device, comprising at least one component for performing the method according to any one of Clauses 1 to 21.

[0188] Clause 37: A non-transitory computer-readable medium storing code for wireless communication at a wireless communication device, the code comprising instructions executable by a processor to perform the method according to any one of Clauses 1 to 21.

[0189] Clause 38: An apparatus for wireless communication at a wireless communication device, comprising: a processor; a memory coupled to the processor; and instructions stored in the memory and executable by the processor to cause the apparatus to perform the method according to any one of Clauses 22 to 34.

[0190] Clause 39: An apparatus for wireless communication at a wireless communication device, comprising at least one component for performing the method according to any one of Clauses 22 to 34.

[0191] Clause 40: A non-transitory computer-readable medium storing code for wireless communication at a wireless communication device, the code comprising instructions executable by a processor to perform the method according to any one of Clauses 22 to 34.

[0192] As used herein, the term "determine" or "determination" encompasses a variety of actions, and thus, "determine" may include computing, calculating, processing, deriving, researching, looking up (such as looking up in a table, database, or other data structure), reasoning, ascertaining, and similar actions. Additionally, "determine" may include receiving (such as receiving information), accessing (such as accessing data stored in a memory), sending (such as sending information), etc. Additionally, "determine" may include parsing, selecting, obtaining, picking, establishing, and other such similar actions.

[0193] As used herein, the phrase referring to "at least one" of a list of items refers to any combination of those items (which includes a single member). For example, "at least one of a, b, or c" is intended to cover: a, b, c, a - b, a - c, b - c, and a - b - c. As used herein, unless otherwise explicitly indicated, "or" is intended to be interpreted in an inclusive sense. For example, "a or b" may include only a, only b, or a combination of a and b.

[0194] As used herein, unless otherwise explicitly indicated, "based on" is intended to be interpreted in an inclusive sense. For example, unless otherwise explicitly indicated, "based on" may be used interchangeably with "at least partially based on", "associated with", or "in accordance with". Specifically, unless the phrase means "only based on 'one'" or the like in the context, whether it is "based on 'one'" or "at least partially based on 'one'" can be based on "one" alone or on a combination of "one" and one or more other factors, conditions, or information.

[0195] The various illustrative components, logical components, logical blocks, modules, circuits, operations, and algorithmic processes described in connection with the examples disclosed herein can be implemented as electronic hardware, firmware, software, or any combination of hardware, firmware, or software, including the structures disclosed in this specification and structural equivalents thereof. This interchangeability of hardware, firmware, and software has been described generally in terms of their functionality and illustrated in the various illustrative components, blocks, modules, circuits, and processes described above. Whether such functionality is implemented in hardware, firmware, or software depends upon the particular application and design constraints imposed on the overall system.

[0196] For those of ordinary skill in the art, various modifications to the examples described in this disclosure will be readily apparent, and the general principles defined herein can be applied to other examples without departing from the spirit or scope of the disclosure. Thus, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the broadest scope consistent with the disclosure, the principles disclosed herein, and the novel features.

[0197] Additionally, the various features described in the context of separate examples in this specification can also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple examples. Thus, although the features may be described above as acting in a particular combination and even initially claimed as such, one or more features from the claimed combination can in some implementations be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.

[0198] Similarly, although the operations are depicted in the figures in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. Additionally, the figures may schematically depict one or more example processes in the form of a flowchart or a flow diagram. However, other operations not depicted can be incorporated into the example processes schematically illustrated. For example, one or more additional operations can be performed before, after, concurrently with, or between any of the operations illustrated. In some environments, multitasking and parallel processing may be advantageous. Further, the separation of the various system components described in the examples above should not be understood as requiring such separation in all examples, but rather it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

Claims

1. An apparatus for wireless communication at a wireless communication device, comprising: a processor; a memory coupled to the processor; and instructions stored in the memory and executable by the processor to cause the apparatus to: receive information associated with a reinforcement learning model, wherein the reinforcement learning model is associated with performing a distributed channel access procedure at the wireless communication device in a wireless local area network according to the information; and transmit a protocol data unit according to the distributed channel access procedure and during a time slot based at least in part on an output of the reinforcement learning model.

2. The apparatus according to claim 1, wherein the instructions for receiving the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: receive an indication that the wireless communication device is permitted to develop the reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

3. The apparatus according to claim 2, wherein the instructions are further executable by the processor to cause the apparatus to: develop the reinforcement learning model at the wireless communication device according to the received indication that the wireless communication device is permitted to develop the reinforcement learning model.

4. The apparatus according to claim 1, wherein the instructions for receiving the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: receive a configuration associated with the reinforcement learning model and an indication of whether the wireless communication device is permitted to retrain the reinforcement learning model for the distributed channel access procedure, wherein the instructions are further executable by the processor to cause the apparatus to: selectively retrain the reinforcement learning model at least in part based on whether the wireless communication device is permitted to retrain the reinforcement learning model.

5. The apparatus according to claim 1, wherein the instructions for receiving the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: receive an indication that the wireless communication device is permitted to retrain the reinforcement learning model for the distributed channel access procedure, wherein the reinforcement learning model is pre-loaded at the wireless communication device, wherein the instructions are further executable by the processor to cause the apparatus to: retrain the reinforcement learning model at least in part based on the indication that the wireless communication device is permitted to retrain the reinforcement learning model.

6. The apparatus according to claim 1, wherein the instructions are further executable by the processor to cause the apparatus to: receive an indication that the wireless communication device is permitted to use the reinforcement learning model to obtain one or more of: a transmission probability, a backoff counter, a duration of a contention window, or one or more parameters associated with the duration of the contention window, wherein the output of the reinforcement learning model includes the transmission probability, the backoff counter, the duration of the contention window, or the one or more parameters associated with the duration of the contention window.

7. The apparatus according to claim 1, wherein the instructions for receiving the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: Receive an indication of a policy associated with training or retraining the reinforcement learning model, wherein the instructions are further executable by the processor to cause the apparatus to: Train or retrain the reinforcement learning model according to the policy.

8. The apparatus according to claim 1, wherein the instructions for receiving the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: Receive an indication of one or more parameters associated with a reinforcement learning technique to be followed by the wireless communication device when training or retraining the reinforcement learning model, wherein the reinforcement learning technique is associated with a Q-learning technique, policy gradient, actor-critic technique, or contextual multi-armed bandit (MAB) or non-contextual MAB technique, wherein the instructions are further executable by the processor to cause the apparatus to: Train or retrain the reinforcement learning model according to the one or more parameters associated with the reinforcement learning technique.

9. The apparatus according to claim 1, wherein the instructions for transmitting the protocol data unit according to the distributed channel access procedure and during the time slot based at least in part on the output of the reinforcement learning model are executable by the processor to cause the apparatus to: Attempt to transmit the protocol data unit during one or more idle time slots according to a transmission probability, wherein the transmission probability is the output of the reinforcement learning model.

10. The apparatus according to claim 1, wherein the instructions for transmitting the protocol data unit according to the distributed channel access procedure and during the time slot based at least in part on the output of the reinforcement learning model are executable by the processor to cause the apparatus to: Transmit the protocol data unit during the time slot according to the expiration of a backoff counter, wherein the backoff counter is the output of the reinforcement learning model.

11. The apparatus according to claim 1, wherein the instructions for transmitting the protocol data unit according to the distributed channel access procedure and during the time slot based at least in part on the output of the reinforcement learning model are executable by the processor to cause the apparatus to: Transmit the protocol data unit during a contention window, wherein the duration of the contention window is associated with the output of the reinforcement learning model.

12. The apparatus according to claim 1, wherein the instructions are further executable by the processor to cause the apparatus to: Derive one or more rewards associated with the reinforcement learning model based on one or more of the following: signal-to-interference-plus-noise ratio, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted protocol data units to the total number of transmitted protocol data units, or ratio of the number of unsuccessfully transmitted protocol data units to the total number of transmitted protocol data units, wherein the output of the reinforcement learning model is at least partially based on the one or more rewards.

13. The apparatus according to claim 1, wherein the instructions are further executable by the processor to cause the apparatus to: Obtain an updated state associated with the wireless communication device based on transmitting the protocol data unit; Input the updated state into the reinforcement learning model to obtain an updated output of the reinforcement learning model; and Transmit a second protocol data unit according to the distributed channel access procedure and during a second time slot that is at least partially based on the updated output of the reinforcement learning model.

14. An apparatus for wireless communication at a wireless communication device, comprising: a processor; a memory coupled to the processor; and instructions stored in the memory and executable by the processor to cause the apparatus to: Transmit information associated with a reinforcement learning model, wherein the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network according to the information; and Receive a protocol data unit from the second wireless communication device according to the distributed channel access procedure and during a time slot that is at least partially based on the output of the reinforcement learning model.

15. The apparatus according to claim 14, wherein the instructions for transmitting the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: Transmit an indication that the second wireless communication device is allowed to develop the reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

16. The apparatus according to claim 14, wherein the instructions for transmitting the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: Transmit information associated with the reinforcement learning model and an indication of whether the second wireless communication device is allowed to retrain the reinforcement learning model for the distributed channel access procedure.

17. The apparatus according to claim 14, wherein the instructions for transmitting the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: Transmit an indication that the second wireless communication device is allowed to retrain the reinforcement learning model for the distributed channel access procedure, wherein the reinforcement learning model is pre-loaded at the second wireless communication device.

18. The apparatus according to claim 14, wherein the instructions for transmitting the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: Send an indication of one or more parameters associated with a reinforcement learning technique to be followed by the second wireless communication device when training or retraining the reinforcement learning model, wherein the reinforcement learning technique is associated with a Q-learning technique, policy gradient, actor-critic technique, or contextual multi-armed bandit (MAB) or non-contextual MAB technique.

19. The apparatus of claim 14, wherein the instructions for sending the information associated with the reinforcement learning model are executable by the processor to cause the apparatus to: Send an indication of one or more rewards associated with the reinforcement learning model, wherein the one or more rewards include one or more of the following: signal-to-interference plus noise ratio, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted protocol data units to the total number of protocol data units transmitted, or ratio of the number of unsuccessfully transmitted protocol data units to the total number of protocol data units transmitted, and wherein the output of the reinforcement learning model is at least partially based on the one or more rewards.

20. The apparatus of claim 14, wherein the instructions are further executable by the processor to cause the apparatus to: Send an indication of one or more parameters associated with the environment of the second wireless communication device, wherein the output of the reinforcement learning model is associated with using the one or more parameters as inputs into the reinforcement learning model.

21. A method for wireless communication at a wireless communication device, comprising: Receiving information associated with a reinforcement learning model, wherein the reinforcement learning model is associated with performing a distributed channel access procedure at the wireless communication device in a wireless local area network according to the information; and Transmitting a protocol data unit according to the distributed channel access procedure and during a time slot that is at least partially based on the output of the reinforcement learning model.

22. The method of claim 21, wherein transmitting the protocol data unit according to the distributed channel access procedure and during the time slot that is at least partially based on the output of the reinforcement learning model comprises: Attempting to transmit the protocol data unit during one or more idle time slots according to a transmission probability, wherein the transmission probability is the output of the reinforcement learning model.

23. The method of claim 21, wherein transmitting the protocol data unit according to the distributed channel access procedure and during the time slot that is at least partially based on the output of the reinforcement learning model comprises: Transmitting the protocol data unit during the time slot according to the expiration of a backoff counter, wherein the backoff counter is the output of the reinforcement learning model.

24. The method of claim 21, wherein transmitting the protocol data unit according to the distributed channel access procedure and during the time slot that is at least partially based on the output of the reinforcement learning model comprises: Transmitting the protocol data unit during a contention window, wherein the duration of the contention window is associated with the output of the reinforcement learning model.

25. The method according to claim 21, wherein receiving the information associated with the reinforcement learning model comprises: receiving an indication of one or more rewards associated with the reinforcement learning model, wherein the one or more rewards include one or more of the following: signal-to-interference-plus-noise ratio, throughput metric, latency metric, number of collisions, ratio of the number of successfully transmitted protocol data units to the total number of transmitted protocol data units, or ratio of the number of unsuccessfully transmitted protocol data units to the total number of transmitted protocol data units, and wherein the output of the reinforcement learning model is at least partially based on the one or more rewards.

26. The method according to claim 21, further comprises: receiving an indication of one or more parameters associated with the environment of the wireless communication device, wherein the output of the reinforcement learning model is associated with using the one or more parameters as inputs into the reinforcement learning model.

27. A method for wireless communication at a wireless communication device, comprises: transmitting information associated with a reinforcement learning model, wherein the reinforcement learning model is associated with performing a distributed channel access procedure at a second wireless communication device in a wireless local area network according to the information; and receiving a protocol data unit from the second wireless communication device during a time slot according to the distributed channel access procedure and at least partially based on the output of the reinforcement learning model.

28. The method according to claim 27, wherein transmitting the information associated with the reinforcement learning model comprises: transmitting an indication that the second wireless communication device is allowed to develop the reinforcement learning model and use the reinforcement learning model for the distributed channel access procedure.

29. The method according to claim 27, wherein transmitting the information associated with the reinforcement learning model comprises: transmitting information associated with the reinforcement learning model and an indication of whether the second wireless communication device is allowed to retrain the reinforcement learning model for the distributed channel access procedure.

30. The method according to claim 27, wherein transmitting the information associated with the reinforcement learning model comprises: transmitting an indication that the second wireless communication device is allowed to retrain the reinforcement learning model for the distributed channel access procedure, wherein the reinforcement learning model is pre-loaded at the second wireless communication device.