Channel access method and related device - Patents.com
The channel access method improves Wi-Fi system throughput and reduces latency by using neural networks trained with operation information from multiple stations to predictively manage channel access, addressing the inefficiencies of current CSMA/CA mechanisms.
Patent Information
- Application Number
- JP2023577777
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-17
- Filing Date
- 2022-06-14
- Publication Date
- 2025-05-22
- Estimated Expiration
- 2042-06-14
AI Technical Summary
Current Wi-Fi systems using CSMA/CA mechanisms for channel access suffer from low system throughput and high latency due to unpredictable channel access behavior among stations, leading to increased collisions and inefficient data transmission.
A channel access method where an access point receives operation information from multiple stations, uses this information to train a neural network for each station, and transmits the training results back to the stations, enabling them to predictively determine whether to access the channel, thereby improving throughput and reducing latency.
The proposed method enhances system throughput and reduces latency by improving the predictive ability of stations to access the channel, thereby minimizing collisions and optimizing data transmission efficiency.
Smart Images

Figure 0007681731000071 
Figure 0007681731000072 
Figure 0007681731000073
Abstract
Description
[Technical field]
[0001] This application claims priority to Chinese Patent Application No. 202110673131.6, entitled “Channel Access Method and Related Apparatus,” filed with the China State Intellectual Property Office on June 17, 2021, which is incorporated herein by reference in its entirety.
[0002] The present application relates to the field of communication technologies, and in particular to a channel access method and related apparatus. [Background technology]
[0003] In wireless networks, such as short-range / wireless local area networks (Wireless Fidelity, Wi-Fi), channels for data transmission are shared. When multiple stations (STAs) in a certain area transmit packets to the same access point (AP), collisions occur and data transmission fails.
[0004] Currently, Wi-Fi systems use a carrier sense multiple access / collision avoidance (CSMA / CA) mechanism to avoid collisions on a shared channel. Specifically, when a packet arrives, a STA with sensing capability senses the channel status within a random duration. If the channel is idle within the random duration, the STA accesses the channel.
[0005] The method of avoiding collisions on a shared channel by using CSMA / CA mechanism can be considered as a collision resolution algorithm, i.e., expecting to achieve collision resolution effect through complete randomization. In other words, each STA in this method has no ability to predict whether other STAs will access the channel. Therefore, the system throughput is low and the latency is high. Summary of the Invention [Means for solving the problem]
[0006] SUMMARY OF THE PRESENT APPLICATION The embodiments of the present application provide a channel access method and related apparatus for improving system throughput and reducing latency.
[0007] According to a first aspect, an embodiment of the present application provides a channel access method, in which an access point (AP) receives operation information reported separately by N stations (STAs), and the N operation information is used to determine a training result of a first neural network of each STA. The AP determines a training result of the first neural network of each STA based on the N operation information, and transmits the training result of the first neural network of each STA to a corresponding STA.
[0008] It can be seen that the training result of the first neural network of each STA is determined based on the operation information reported by the N STAs, not only on the operation information of that STA, which can improve the prediction ability of the first neural network and help improve the ability of the STA to predict whether to access the channel, thereby improving the system throughput and reducing the delay.
[0009] In an optional embodiment, the operation information indicates an operation for a period of time, and the operation is transmission or skipping transmission. The period is the time between the time when the STA last successfully reports the operation information and the current time. In other words, the operation is the operation of transmitting a packet or skipping the transmission of a packet by the STA since the STA last successfully reports the operation information.
[0010] In an optional embodiment, the AP may further receive carrier sensing result information or packet transmission result information reported by the N STAs separately. The carrier sensing result information includes a carrier sensing result, and the packet transmission result information includes a packet transmission result. Thus, the AP determining a training result of the first neural network of each STA based on the N pieces of operation information is the AP determining a training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of carrier sensing result information, or the AP determining a training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of packet transmission result information.
[0011] It can be seen that each STA may further report carrier sensing result information or packet transmission result information to the AP. Thus, the AP can directly train the first neural network of each STA based on the N pieces of operation information and the N pieces of carrier sensing result information, or directly train the first neural network of each STA based on the N pieces of operation information and the N pieces of packet transmission result information, thereby helping to reduce the processing complexity of the AP.
[0012] In one optional implementation, the training results are neural network parameters or gradients, and the neural network parameters / gradients are used by the corresponding STA to update the first neural network.
[0013] In an optional embodiment, when the AP receives operation information reported separately by N STAs, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield, and data 1 subfield to data T subfield, where T is a positive integer.
[0014] The time indication subfield indicates the time when the STA successfully received the first response information last time. The first response information is response information transmitted when the AP successfully received the operation information transmitted by the STA. In other words, the first response information is response information received when the STA successfully reported the operation information last time, and the response information may be acknowledgement ACK information. The data 1 subfield indicates the operation performed in the first slot after the STA successfully received the first response information last time. In other words, the data 1 subfield indicates the operation performed in the first slot after the STA successfully reported the operation information last time. The data T subfield indicates the operation performed in the Tth slot after the STA successfully received the first response information last time, and the Tth slot is also the last slot before the STA is currently reporting the operation information.
[0015] It can be seen that for N STAs, the operation information reported by each STA is carried in the first frame, and the operation information reported by each STA to the AP includes the time when the STA last successfully reported its operation information, and the operation from the first slot to the Tth slot after the operation information was last successfully reported.
[0016] In another optional embodiment, when the AP receives operation information reported separately by N STAs, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield, an operation 1 subfield, a time 1 subfield, ..., an operation P subfield, and a time P subfield, where P is a positive integer.
[0017] The time indication subfield indicates the time when the STA last successfully received the first response information. The first response information is the response information sent when the AP successfully receives the operation information sent by the STA. In other words, the time indication subfield indicates the time when the STA last successfully reported the operation information.
[0018] The operation 1 subfield indicates the first operation after the STA has last successfully received the first response information. The operation P subfield indicates the Pth operation between the time when the STA has last successfully received the first response information and the current time. In other words, the operation 1 subfield indicates the first operation after the STA has last successfully reported the operation information, and the operation P subfield indicates the last operation between the time when the STA has last successfully reported the operation information and the current time.
[0019] The Time 1 subfield indicates the duration of operation 1 or the end time of operation 1. The Time P subfield indicates the duration of operation P or the end time of operation P. If the Time 1 subfield indicates the duration of operation 1 and the Time P subfield indicates the duration of operation P, different operations have different meanings represented by their durations. If the operation is a send operation, the duration represents the packet length of the packet to be sent. If the operation is a skip send operation, the duration represents the duration to skip the transmission of the packet.
[0020] It can be seen that for N STAs, the operation information reported by each STA is carried in the first frame, and the operation information reported by each STA to the AP includes the time when the STA last successfully reported its operation information, each operation since the STA last successfully reported its operation information, and the duration or end time of each operation.
[0021] In yet another optional embodiment, when the AP receives operation information reported separately by N STAs, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time 1 indication subfield, an operation 1 subfield, ..., a time P indication subfield, and an operation P subfield, where P is a positive integer.
[0022] The action 1 subfield indicates the first action after the STA has last successfully received the first response information. The action P subfield indicates the Pth action between the time when the STA last successfully received the first response information and the current time. The first response information is response information transmitted when the AP successfully receives the action information transmitted by the STA. In other words, the action 1 subfield indicates the first action after the STA has last successfully reported the action information, and the action P subfield indicates the last action between the time when the STA last successfully reported the action information and the current time. The time 1 indication subfield indicates the start time of action 1. The time P indication subfield indicates the start time of action P.
[0023] It can be seen that for N STAs, the operation information reported by each STA is carried in the first frame, and the operation information reported by each STA to the AP includes each operation since the STA last successfully reported its operation information, and the start time of each operation.
[0024] In yet another optional embodiment, when the AP receives operation information reported separately by N STAs, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time 1 indication subfield, a duration 1 subfield, ..., a time K indication subfield, and a duration K subfield, where K is a positive integer.
[0025] The Time 1 Indication subfield indicates the start time / end time of Operation 1. Operation 1 is a transmission operation performed when the STA transmits a packet for the first time after the STA successfully receives the first response information last time and does not receive the second response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The second response information is response information transmitted when the AP successfully receives the packet transmitted by the STA. The Duration 1 subfield indicates the duration of Operation 1.
[0026] The Time K Indication subfield indicates the start time / end time of operation K. Operation K is a transmission operation performed when the STA transmits a packet for the Kth time after the STA has previously successfully received the first response information and has not received the second response information. The Duration K subfield indicates the duration of operation K.
[0027] It can be seen that for N STAs, the operation information reported by each STA is carried in the first frame, and the operation information reported by each STA to the AP includes the start time / end time of the transmission operation for each time the STA fails to transmit a packet after last successfully reporting the operation information, and the duration of the transmitted packet for each time the packet fails to be transmitted.
[0028] In yet another optional embodiment, when the AP receives operation information reported separately by N STAs, the operation information is carried in an operation details field of a first frame reported by the STA, where the operation details field includes a first time 1 indication subfield, a second time 1 indication subfield, ..., a first time K indication subfield, and a second time K indication subfield, where K is a positive integer.
[0029] The first time 1 subfield indicates the start time of operation 1. The first time K subfield indicates the start time of operation K. Operation 1 is a transmission operation performed when the STA transmits a packet for the first time after successfully receiving the first response information last time and does not receive the second response information. Operation K is a transmission operation performed when the STA transmits a packet for the Kth time after successfully receiving the first response information last time and does not receive the second response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The second response information is response information transmitted when the AP successfully receives the packet transmitted by the STA. In other words, operation 1 is an operation in which the corresponding STA fails to transmit a packet for the first time after successfully reporting the operation information last time, and operation K is an operation in which the STA fails to transmit a packet for the Kth time after successfully reporting the operation information last time.
[0030] The second time 1 indication subfield indicates the end time of action 1. The second time K indication subfield indicates the end time of action K.
[0031] It can be seen that for N STAs, the operation information reported by each STA is carried in the first frame, and the operation information reported by each STA to the AP includes the start time and end time of each transmission operation that the STA fails to transmit a packet after it last successfully reported its operation information.
[0032] In a further optional embodiment, when the AP receives the operation information and carrier sensing result information reported separately by N STAs, the operation information and carrier sensing result information are carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield, and a data 1 subfield to a data T subfield, where T is a positive integer.
[0033] The time indication subfield indicates the time when the STA last received the first response information normally. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the STA.
[0034] The Data 1 subfield indicates the carrier sensing result and the operation to be performed in the first slot after the STA last successfully received the first response information. The Data T subfield indicates the carrier sensing result and the operation to be performed in the Tth slot after the STA last successfully received the first response information.
[0035] It can be seen that for N STAs, the operational information and carrier sensing result information reported by the STAs is carried in the first frame, and the information reported by each STA to the AP includes the time when the STA last successfully reported its operational information, as well as the carrier sensing result and the operation performed in each slot since the STA last successfully reported its operational information.
[0036] In a further optional embodiment, when the AP receives the operation information and packet transmission result information reported by N STAs separately, the operation information and packet transmission result information are carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield, and a data 1 subfield to a data T subfield, where T is a positive integer.
[0037] The time indication subfield indicates the time when the STA last received the first response information normally. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the STA.
[0038] The Data 1 subfield indicates the packet transmission result and the operation to be performed in the first slot after the STA last successfully received the first response information. The Data T subfield indicates the packet transmission result and the operation to be performed in the Tth slot after the STA last successfully received the first response information.
[0039] It can be seen that for N STAs, the operation information and packet transmission result information reported by the STAs is carried in the first frame, and the information reported by each STA to the AP includes the time when the STA last successfully reported its operation information, as well as the packet transmission results and operations performed in each slot since the STA last successfully reported its operation information.
[0040] In one optional embodiment, the AP determining a training result of the first neural network of each STA based on the N pieces of operational information includes: the AP inputting status information of each STA into a first neural network of the corresponding STA to obtain an output of the first neural network; the AP inputting the output of each first neural network into a second neural network to obtain an output of the second neural network, where the output of the second neural network represents an expected reward within a preset time; and the AP training a third neural network based on the output of the second neural network and a reward function, and determining a training result of each first neural network by minimizing a loss function of the third neural network, where the third neural network includes each of the first neural network and the second neural network.
[0041] Status information of the STA is obtained based on the behavior information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of behavior information, and a reward function is determined based on the N pieces of behavior information.
[0042] Further, status information of the STA is obtained based on the operation information and the carrier sensing result information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of operation information and the N pieces of carrier sensing result information, and a reward function is determined based on the N pieces of operation information and the N pieces of carrier sensing result information.
[0043] Alternatively, status information of the STA is obtained based on operation information and packet transmission result information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of operation information and the N pieces of packet transmission result information, and a reward function is determined based on the N pieces of operation information and the N pieces of packet transmission result information.
[0044] It can be seen that the AP first inputs the status information obtained based on the information reported by each STA into the first neural network of the STA to obtain the output of each first neural network, then inputs the outputs of the N first neural networks into the second neural network to obtain the output of the second neural network, and then trains the third neural network based on the loss function to finally obtain the training result of the first neural network. The training result of the first neural network of each STA is determined based on the information reported by the N STAs, not only on the information of that STA. This helps improve the ability of each STA to predict the channel access behavior of other STAs.
[0045] In one optional implementation, when the AP determines that the first STA successfully transmits the packet based on the N pieces of operation information, the AP sets the value of the reward function to 1. The first STA is the STA among the N STAs that has the longest time interval between the time when the second response information is last successfully received and the current time.
[0046] It can be seen that, based on the information reported by the N STAs, the AP sets the value of the reward function to 1 when it determines that the STA has the longest interval since the last time a packet was successfully transmitted.
[0047] In another optional embodiment, when the second STA determines based on the N pieces of operation information that the second STA successfully transmits the packet, the AP sets the value of the reward function to a first duration -1. The second STA is a STA other than the first STA among the N STAs, and the first STA is a STA among the N STAs that has the longest time interval between the time when the second response information was last successfully received and the current time. The first duration is the duration between the time when the second STA last successfully received the second response information and the current time.
[0048] It can be seen that when the AP determines, based on the information reported by the N STAs, that a STA other than the STA with the longest interval since successfully transmitting a packet successfully transmits a packet, the AP sets the value of the reward function to the time interval (since the STA last successfully transmitted a packet) -1.
[0049] In yet another optional implementation, when M STAs among the N STAs determine to transmit packets in the same slot based on N pieces of operation information, the AP sets the value of the reward function to −1, where M is a positive integer less than or equal to N. It can be seen that when some STAs among the N STAs determine to transmit packets in the same slot based on information reported by the N STAs, the AP sets the reward function to −1.
[0050] In yet another optional implementation, when determining based on the N operational information that none of the N STAs transmit packets in the same slot, the AP sets the value of the reward function to 0. It can be seen that when determining based on the information reported by the N STAs that none of the N STAs transmit packets in the same slot, the AP sets the value of the reward function to 0.
[0051] In an optional embodiment, N STAs share neural network parameters. In this case, AP sends the training result of the first neural network of each STA to corresponding STA, AP broadcasts the training result of the first neural network to N STAs. It can be seen that when N STAs share neural network parameters, AP obtains the same training result by training each first neural network according to the information reported by N STAs, and AP can inform each STA of the training result by broadcasting, thereby reducing the signaling overhead of the system.
[0052] In another optional embodiment, S STAs among the N STAs share neural network parameters, where S is a positive integer equal to or less than N. The AP transmits the training results of the first neural network of each STA to the corresponding STA, which means that the AP multicasts the training results of the first neural network corresponding to the S STAs to the S STAs, and unicasts the training results of the (NS) first neural networks to the corresponding STAs. It can be seen that when some STAs among the N STAs share neural network parameters, the AP can notify some STAs of the training results corresponding to the shared neural network parameters by multicast, and unicast the training results corresponding to the non-shared neural network parameters to other STAs in a unicast manner. In this way, the training results of the STAs sharing one neural network parameter are notified by multicast, which can also reduce system overhead.
[0053] In yet another optional embodiment, when the N STAs do not share neural network parameters, the training result of each first neural network is unicast to the corresponding STA.
[0054] According to a second aspect, the present application further provides a channel access method. The channel access method according to this aspect corresponds to the channel access method according to the first aspect, and the channel access method according to this aspect is described from the station STA side. In the method, the station STA reports operation information to an access point AP, and the operation information is used to determine a training result of a first neural network, and the first neural network is a neural network of the STA. The STA receives a training result of the first neural network from the AP, and the training result of the first neural network is obtained based on the operation information, and the training result of the first neural network is used to update the first neural network to determine whether the STA accesses the channel. The STA updates the first neural network based on the training result of the first neural network, and when it senses that the channel is idle, it determines whether to access the channel based on the updated first neural network and the current status information.
[0055] In this embodiment of the present application, the STA reports operation information to the AP, and receives a training result obtained by training the first neural network based on the operation information by the AP, whereby the STA updates the first neural network based on the training result, and when it senses that the channel is idle, it can be seen that it determines whether to access the channel based on the updated first neural network and the sensed operation information. The training result for updating each first neural network is determined by the AP based on the operation information reported by the N STAs, so that the first neural network has better predictability. If the STA determines whether to access the channel based on the updated first neural network, the accuracy of determining whether to access the channel or skip accessing will be better. This improves the communication system throughput and reduces communication latency.
[0056] In an optional embodiment, the STA further reports carrier sensing result information or packet transmission result information to the AP, and the carrier sensing result information or packet transmission result information is used to determine the training result of the first neural network. In addition to reporting the operation information to the AP, the STA may further report carrier sensing result information or packet transmission result information to the AP, so that the AP can directly train the first neural network according to the information reported by the N STAs, thereby reducing the processing complexity of the AP.
[0057] In one optional implementation, the training results are neural network parameters or gradients, and the carrier sensing result information or the packet transmission result information is used to determine the training results of the first neural network.
[0058] In one optional embodiment, when a STA reports operation information, the operation information is carried in an operation details field of the first frame. The operation details field includes a time indication subfield and data 1 subfield to data T subfield, where T is a positive integer.
[0059] The time indication subfield indicates the time when the STA successfully received the first response information last time. The first response information is response information transmitted when the AP successfully received the operation information transmitted by the STA. In other words, the first response information is response information received when the STA successfully reported the operation information last time, and the response information may be acknowledgement ACK information. The data 1 subfield indicates the operation performed in the first slot after the STA successfully received the first response information last time. In other words, the data 1 subfield indicates the operation performed in the first slot after the STA successfully reported the operation information last time. The data T subfield indicates the operation performed in the Tth slot after the STA successfully received the first response information last time, and the Tth slot is also the last slot before the STA is currently reporting the operation information.
[0060] It can be seen that the operation information reported by the STA is carried in the first frame, and the operation information reported by each STA to the AP includes the operation at the time the STA last successfully reported its operation information, and from the first slot to the Tth slot after the operation information was last successfully reported.
[0061] In another optional embodiment, when a STA reports operation information, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield, an operation 1 subfield, a time 1 subfield, ..., an operation P subfield, and a time P subfield, where P is a positive integer.
[0062] The time indication subfield indicates the time when the STA last successfully received the first response information. The first response information is the response information sent when the AP successfully receives the operation information sent by the STA. In other words, the time indication subfield indicates the time when the STA last successfully reported the operation information.
[0063] The operation 1 subfield indicates the first operation after the STA has last successfully received the first response information. The operation P subfield indicates the Pth operation between the time when the STA has last successfully received the first response information and the current time. In other words, the operation 1 subfield indicates the first operation after the STA has last successfully reported the operation information, and the operation P subfield indicates the last operation between the time when the STA has last successfully reported the operation information and the current time.
[0064] The Time 1 subfield indicates the duration or end time of Operation 1. The Time P subfield indicates the duration or end time of Operation P. When the Time 1 subfield indicates the duration of Operation 1 and the Time P subfield indicates the duration of Operation P, different operations have different meanings represented by their durations. When the operation is a transmission operation, the duration represents the packet length of the packet to be transmitted. When the operation is a transmission skip operation, the duration represents the duration of skipping the transmission of the packet.
[0065] The operation information reported by the STA is carried in the first frame. It can be known that the operation information reported by the STA to the AP includes the time when the STA last normally reported the operation information, each operation after the STA last normally reported the operation information, and the duration or end time of each operation.
[0066] In yet another optional embodiment, when the STA reports operation information, the operation information is carried in the operation details field of the first frame reported by the STA. The operation details field includes a Time 1 indication subfield, an Operation 1 subfield, …, a Time P indication subfield, and an Operation P subfield, where P is a positive integer.
[0067] The Operation 1 subfield indicates the first operation after the STA last normally received the first response information. The Operation P subfield indicates the Pth operation between the time when the STA last normally received the first response information and the current time. The first response information is the response information transmitted when the AP normally receives the operation information transmitted by the STA. In other words, the Operation 1 subfield indicates the first operation after the STA last normally reported the operation information, and the Operation P subfield indicates the last operation between the time when the STA last normally reported the operation information and the current time. The Time 1 indication subfield indicates the start time of Operation 1. The Time P indication subfield indicates the start time of Operation P.
[0068] It can be seen that the operation information reported by the STAs is carried in the first frame, and the operation information reported by each STA to the AP includes each operation since the STA last successfully reported its operation information, and the start time of each operation.
[0069] In yet another optional embodiment, when a STA reports operation information, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time 1 indication subfield, a duration 1 subfield, ..., a time K indication subfield, and a duration K subfield, where K is a positive integer.
[0070] The Time 1 Indication subfield indicates the start time / end time of Operation 1. Operation 1 is a transmission operation performed when the STA transmits a packet for the first time after the STA successfully receives the first response information last time and does not receive the second response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The second response information is response information transmitted when the AP successfully receives the packet transmitted by the STA. The Duration 1 subfield indicates the duration of Operation 1.
[0071] The Time K Indication subfield indicates the start time / end time of operation K. Operation K is a transmission operation performed when the STA transmits a packet for the Kth time after the STA has previously successfully received the first response information and has not received the second response information. The Duration K subfield indicates the duration of operation K.
[0072] It can be seen that the operation information reported by the STA is carried in the first frame, and the operation information reported by the STA to the AP includes the start time / end time of the transmission operation for each failed packet transmission after the STA last successfully reported the operation information, and the duration of each failed packet transmission.
[0073] In yet another optional embodiment, when a STA reports operation information, the operation information is carried in an operation details field of a first frame reported by the STA, where the operation details field includes a first time 1 indication subfield, a second time 1 indication subfield, ..., a first time K indication subfield, and a second time K indication subfield, where K is a positive integer.
[0074] The first time 1 subfield indicates the start time of operation 1. The first time K subfield indicates the start time of operation K. Operation 1 is a transmission operation performed when the STA transmits a packet for the first time after successfully receiving the first response information last time and does not receive the second response information. Operation K is a transmission operation performed when the STA transmits a packet for the Kth time after successfully receiving the first response information last time and does not receive the second response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The second response information is response information transmitted when the AP successfully receives the packet transmitted by the STA. In other words, operation 1 is an operation in which the corresponding STA fails to transmit a packet for the first time after successfully reporting the operation information last time, and operation K is an operation in which the STA fails to transmit a packet for the Kth time after successfully reporting the operation information last time.
[0075] The second time 1 indication subfield indicates the end time of action 1. The second time K indication subfield indicates the end time of action K.
[0076] It can be seen that the operational information reported by the STA is carried in the first frame, and the operational information reported by the STA to the AP includes the start time and end time of each transmission operation that fails to transmit a packet after the STA last successfully reported the operational information.
[0077] In a further optional embodiment, when a STA reports operation information and carrier sensing result information, the operation information and carrier sensing result information are carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield, and data 1 subfield through data T subfield, where T is a positive integer.
[0078] The time indication subfield indicates the time when the STA last received the first response information normally. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the STA.
[0079] The Data 1 subfield indicates the carrier sensing result and the operation to be performed in the first slot after the STA last successfully received the first response information. The Data T subfield indicates the carrier sensing result and the operation to be performed in the Tth slot after the STA last successfully received the first response information.
[0080] It can be seen that the operational information and carrier sensing result information reported by the STA is carried in the first frame, and the information reported by the STA to the AP includes the time when the STA last successfully reported its operational information, as well as the carrier sensing result and the operation performed in each slot since the STA last successfully reported its operational information.
[0081] In a further optional embodiment, when the STA reports the operation information and the packet transmission result information, the operation information and the packet transmission result information are carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield, and a data 1 subfield to a data T subfield, where T is a positive integer.
[0082] The time indication subfield indicates the time when the STA last received the first response information normally. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the STA.
[0083] The Data 1 subfield indicates the packet transmission result and the operation to be performed in the first slot after the STA last successfully received the first response information. The Data T subfield indicates the packet transmission result and the operation to be performed in the Tth slot after the STA last successfully received the first response information.
[0084] It can be seen that the operation information and packet transmission result information reported by the STA is carried in the first frame, and the information reported by the STA to the AP includes the time when the STA last successfully reported the operation information, as well as packet transmission result information and the operations performed in each slot since the STA last successfully reported the operation information.
[0085] In one optional embodiment, the STA updates the first neural network based on the training result of the first neural network, and when it senses that the channel is idle, determining whether to access the channel based on the updated first neural network and current status information of the STA includes: the STA inputs the current status information of the STA to the updated first neural network and outputs a first value and a second value, where the first value represents an expected reward obtained by accessing the channel and the second value represents an expected reward obtained by skipping access to the channel; and if the first value is greater than the second value, the STA determines to access the channel, or if the first value is less than the second value, the STA determines to skip accessing the channel.
[0086] When sensing that the channel is idle, the STA inputs the sensed operation information into the updated first neural network to obtain an expected reward for accessing the channel and an expected reward for skipping access to the channel, and finds that if the expected reward for accessing the channel is greater than the expected reward for skipping access to the channel, it decides to access the channel.
[0087] According to a third aspect, the present application further provides a communication device. The communication device has some or all of the functions of implementing an AP according to the first aspect, or has some or all of the functions of implementing a STA according to the second aspect. For example, the functions of the communication device may have the functions of an AP according to some or all of the embodiments of the first aspect of the present application, or may have the functions of independently implementing any embodiment of the present application. The functions may be implemented by hardware, or may be implemented by hardware executing corresponding software. The hardware or software includes one or more units or modules corresponding to these functions.
[0088] In one possible design, the structure of the communication device may include a processing unit and a communication unit. The processing unit is configured to support the communication device in performing corresponding functions in the aforementioned methods. The communication unit is configured to support communication between the communication device and other communication devices. The communication device may further include a storage unit. The storage unit is configured to be coupled to the processing unit and the communication unit, and the storage unit stores program instructions and data required for the communication device.
[0089] In one embodiment, the communication device comprises: A communication unit configured to receive operation information separately reported by N stations STAs, the N pieces of operation information being used to determine a training result of a first neural network of each STA, where N is a positive integer; A processing unit configured to determine a training result of a first neural network of each STA based on the N pieces of operation information, The communication unit is further configured to transmit the training result of the first neural network of each STA to the corresponding STA; Includes.
[0090] In addition, for other optional implementations of the communication device in this aspect, please refer to the relevant contents of the first aspect, and details will not be described again in this specification.
[0091] In another embodiment, the communication device comprises: a communication unit configured to report operational information to an access point AP, the operational information being used to determine a training result of the first neural network of the processing unit; The communication unit is further configured to receive a training result of the first neural network from the AP, and the training result of the first neural network is used to update the first neural network so that the processing unit determines whether to access the channel; and a processing unit configured to update the first neural network based on a training result of the first neural network, and when detecting that the channel is idle, determine whether to access the channel based on the updated first neural network and current status information of the processing unit; Includes.
[0092] In addition, for other optional implementations of the communication device in this aspect, please refer to the relevant contents of the second aspect, and details will not be described again in this specification.
[0093] For example, the communication unit may be a transceiver or a communication interface, the storage unit may be a memory, and the processing unit may be a processor.
[0094] In one embodiment, the communication device comprises: A transceiver configured to receive operational information reported separately by N station STAs, the N pieces of operational information being used to determine a training result of a first neural network of each STA, where N is a positive integer; a processor configured to determine a training result of a first neural network of each STA based on the N pieces of operation information; The transceiver is further configured to transmit a training result of the first neural network of each STA to a corresponding STA; Includes.
[0095] In addition, for other optional implementations of the communication device in this aspect, please refer to the relevant contents of the first aspect, and details will not be described again in this specification.
[0096] In another embodiment, the communication device comprises: a transceiver configured to report operational information to an access point AP, the operational information being used to determine a training result of the first neural network of the processor; a transceiver further configured to receive a training result of the first neural network from the AP, the training result of the first neural network being used to update the first neural network so that the processor determines whether to access the channel; a processor configured to update a first neural network based on a training result of the first neural network, and when detecting that the channel is idle, determine whether to access the channel based on the updated first neural network of the processor and current status information; Includes.
[0097] In addition, for other optional implementations of the communication device in this aspect, please refer to the relevant contents of the second aspect, and details will not be described again in this specification.
[0098] In another embodiment, the communication device is a chip or a chip system. The processing unit may be represented as a processing circuit or a logic circuit. The communication unit may be an input / output interface, an interface circuit, an output circuit, an input circuit, a pin, associated circuitry, etc. on the chip or chip system.
[0099] In one implementation process, the processor may be configured to perform, for example, but not limited to, baseband-related processing, and the transceiver may be configured to perform, for example, but not limited to, radio frequency reception and transmission. The aforementioned components may be separately located on independent chips, or at least some or all of the components may be located on the same chip. For example, the processor may be divided into an analog baseband processor and a digital baseband processor. The analog baseband processor and the transceiver may be integrated on the same chip, and the digital baseband processor may be located on an independent chip. With the continuous development of integrated circuit technology, more and more components may be integrated on the same chip. For example, a digital baseband processor and multiple application processors (including, for example, but not limited to, a graphics processing unit, a multimedia processor, etc.) may be integrated on the same chip. Such a chip may be called a System on a Chip (SoC). Whether the components are separately located on different chips or integrated on one or more chips usually depends on the requirements of the product design. The embodiment of the aforementioned components is not limited in this embodiment of the present application.
[0100] According to a fourth aspect, the present application further provides a processor configured to perform the aforementioned methods. In the processes of performing these methods, the processes of transmitting the aforementioned information and receiving the aforementioned information in the aforementioned methods may be understood as the processes of outputting the aforementioned information by the processor and receiving the aforementioned input information by the processor. When outputting the information, the processor outputs the information to the transceiver, which causes the transceiver to transmit. After the information is output by the processor, other processing may need to be performed on the information before the information arrives at the transceiver. Similarly, when the processor receives the aforementioned input information, the transceiver receives the aforementioned information and inputs the aforementioned information to the processor. Furthermore, after the transceiver receives the aforementioned information, other processing may need to be performed on the information before the information is input to the processor.
[0101] Based on the above principles, for example, the reporting of operational information referred to in the above methods may be understood as the processor outputting the operational information.
[0102] Unless otherwise specified, i.e., where operations such as transmit, send, and receive related to the processor do not contradict the actual function or internal logic of the operations in the associated description, all operations may be generally understood as operations such as output, receive, and input of the processor, rather than operations such as transmit, send, and receive performed directly by radio frequency circuits and antennas.
[0103] In one implementation process, the processor may be a processor specifically configured to perform these methods, or may be a processor that executes computer instructions in memory to perform these methods, such as a general-purpose processor. The memory may be a non-transitory memory, such as a Read Only Memory (ROM). The memory and the processor may be integrated on the same chip, or may be located separately on different chips. The type of memory and the manner of arranging the memory and the processor are not limited in this embodiment of the present application.
[0104] According to a fifth aspect, the present application further provides a communication system. The system includes at least one AP and at least two STAs in the above-mentioned aspects. In another possible design, the system may further include other devices that interact with the AP and the STAs in the solution provided in the present application.
[0105] According to a sixth aspect, the present application provides a computer-readable storage medium configured to store instructions which, when executed by a communications device, perform a method according to any one of the first and second aspects.
[0106] According to a seventh aspect, the present application further provides a computer program product comprising instructions, which when run on a communications device, enable the communications device to perform a method according to any one of the first or second aspects.
[0107] According to an eighth aspect, the present application provides a chip system. The chip system includes a processor and an interface. The interface is configured to obtain a program or instruction. The processor is configured to call the program or instruction to implement or support the AP in implementing the function in the first aspect, or to call the program or instruction to implement or support the STA in implementing the function in the second aspect, for example, in determining or processing at least one of the data and information in the above-mentioned method. In one possible design, the chip system further includes a memory. The memory is configured to store the program instructions and data required for the terminal. The chip system may include a chip, or may include a chip and another discrete component.
[0108] According to a ninth aspect, the present application provides a communications apparatus including a processor configured to execute a computer program or executable instructions stored in a memory, whereby when the computer program or executable instructions are executed the apparatus is enabled to perform a method according to the first aspect and any one of the possible implementations of the first aspect.
[0109] In one possible implementation, the processor and memory are integrated together.
[0110] In another possible implementation, the memory is located external to the communication device.
[0111] According to a tenth aspect, the present application provides a communications apparatus including a processor configured to execute a computer program or executable instructions stored in a memory, whereby when the computer program or executable instructions are executed the apparatus is enabled to perform a method according to the second aspect and any one of the possible implementations of the second aspect.
[0112] In one possible implementation, the processor and memory are integrated together.
[0113] In another possible implementation, the memory is located external to the communication device. [Brief description of the drawings]
[0114]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6(a)
Figure 6(b)
Figure 6(c)
Figure 6(d)
Figure 6(e)
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
[0115] The following clearly and completely describes the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application.
[0116] In order to better understand the channel access method disclosed in the embodiments of the present application, a communication system to which the embodiments of the present application are applicable will first be described.
[0117] 1.Communication Systems FIG. 1 is a schematic diagram of the structure of a communication system according to an embodiment of the present application. The communication system may include, but is not limited to, one access point (AP) and two stations (STA). The number and form of devices shown in FIG. 1 are used as an example and do not constitute limitations on the embodiments of the present application. In practical applications, two or more APs and three or more STAs may be included. The communication system shown in FIG. 1 is described by using an example in which AP101, STA1021, and STA1022 are used, and AP101 can provide wireless services to STA1021 and STA1022. In FIG. 1, an example is used in which AP101 is a base station and STA1021 and STA1022 are mobile phones.
[0118] In this embodiment of the present application, the communication system may be a wireless local area network (WLAN), a cellular network, or other wireless communication system supporting parallel transmission over multiple links. The embodiment of the present application is mainly described by using an IEEE 802.11 deployed network. Various aspects in the present application may be extended to other networks using various standards or protocols, such as BLUETOOTH (registered trademark), high performance radio LAN (HIPERLAN) (a wireless standard similar to the IEEE 802.11 standard used primarily in Europe), a wide area network (WAN), a personal area network (PAN), or other networks known or developed in the future. Thus, various aspects provided in the present application are applicable to any suitable wireless network, regardless of coverage and wireless access protocol.
[0119] In an embodiment of the present application, the STA may have wireless receiving and transmitting capabilities, support 802.11 series protocols, and communicate with an AP or other STAs. For example, the STA may be any user communication device that allows a user to communicate with an AP and further communicate with a WLAN, including but not limited to user equipment that can be connected to a network, such as a tablet computer, a desktop computer, a laptop computer, a notebook computer, an Ultra-mobile Personal Computer (UMPC), a handheld computer, a netbook, a Personal Digital Assistant (PDA), or a mobile phone, or an Internet of Things node in the Internet of Things, or an in-vehicle communication device in the Internet of Vehicles, etc. Optionally, the STA may alternatively be a chip and processing system in the aforementioned terminal.
[0120] The AP in the embodiment of the present application is a device that provides services to STAs and may support 802.11 series protocols. For example, the AP may be a communication entity such as a communication server, a router, a switch, or a bridge. Alternatively, the AP may include various forms of macro base stations, micro base stations, relay stations, etc. Of course, the AP may alternatively be chips and processing systems in these various forms of devices to implement the methods and functions in the embodiments of the present application.
[0121] To facilitate understanding of the embodiments disclosed in this application, the following two points are explained.
[0122] (1) In the embodiments disclosed in this application, a scenario of a wireless local area network (Wireless Fidelity, Wi-Fi) network in a wireless communication network is used as an example for explanation. It should be noted that the solutions in the embodiments disclosed in this application can be applied to other wireless communication networks, and the corresponding names can be replaced with the names of the corresponding functions in other wireless communication networks.
[0123] (2) Aspects, embodiments, or features of the present application are presented in the embodiments disclosed herein by describing systems that include multiple devices, components, modules, etc. It is to be appreciated and understood that each system may include other devices, components, modules, etc. and / or may not include all of the devices, components, modules, etc. discussed with reference to the accompanying drawings. Additionally, combinations of these solutions may also be used.
[0124] 2. The technical problem to be solved by the present application. Currently, a carrier sense multiple access / collision avoidance (CSMA / CA) mechanism is used in communication systems to avoid collisions on a shared channel. That is, as shown in FIG. 2, when a packet arrives, STA1 (i.e., a CSMA / CA node) with sensing capability performs channel access by using a random backoff mechanism, i.e., senses the channel status within a random duration (Ts). If the channel is idle within the random duration, the STA accesses the channel, i.e., transmits packet y. However, only if STA2 with the same sensing capability senses the channel and the time T for STA2 to sense the channel is not equal to Ts, no collision occurs between STA1 and STA2, i.e., STA1 can transmit a packet normally. In other words, if the sensing time T of STA2 is equal to the sensing time of STA1, both STA1 and STA2 consider the channel to be idle within the sensing time, and both decide to access the channel. That is, STA1 and STA2 transmit packets simultaneously, STA1 transmits packet x, and STA2 transmits packet y, which causes a collision between STA1 and STA2 on the shared channel. As a result, neither STA1 nor STA2 can transmit packets successfully.
[0125] The CSMA / CA mechanism can be considered as a collision resolution algorithm, i.e., expecting to achieve the collision resolution effect through complete randomization. In other words, each STA in this scheme has no ability to predict whether other STAs will access the channel. Therefore, the system throughput is low and the latency is high. In addition, as the number of STAs in the network increases, the collisions in the network increase, and therefore the average backoff time of the STAs increases. This generates long transmission latency and large latency jitter. In addition, studies have shown that the theoretical upper limit of the CSMA / CA capacity is only about 85%, i.e., in the best case, there is still 15% collision between STAs. In addition, the configuration parameters of the STAs also have a large impact on the actual performance. Studies have shown that the system capacity is generally only 70%-80%. In other words, when the collisions between STAs are resolved by using the CSMA / CA mechanism in the communication system, the throughput is low.
[0126] Artificial intelligence (AI) techniques are widely used in wireless communication fields to improve communication performance and user experience. Reinforcement learning (RL) is an AI technique suitable for the channel access problem, where intelligent agents (network nodes) learn in a search process where they take actions (transmit or skip transmissions) to find an optimal policy that maximizes the expected reward (throughput) in the environment (wireless network). The online learning and model-less optimization properties of RL give it better generalization ability than traditional model-based optimization methods.
[0127] In the embodiment of the present application, the RL technique is combined with channel access. The AP uses a reinforcement learning method to train a neural network corresponding to each STA based on the operation information reported by N STAs, and obtains the training result of the neural network corresponding to each STA, so that each STA can decide whether to access the channel based on the training result, thereby improving the ability of the STA to predict whether to access the channel.
[0128] 3. Channel access method 100 (each STA reports operational information to the AP). An embodiment of the present application provides a channel access method 100. Figure 3 is a schematic interaction diagram of the channel access method 100. The channel access method 100 is described in terms of the interaction between an AP and a STA. The channel access method 100 includes, but is not limited to, the following steps:
[0129] S101: N stations STAs separately report operation information to an access point AP, and the N pieces of operation information are used to determine a training result of a first neural network of each STA, where N is a positive integer.
[0130] The AP corresponds to M STAs, where M is a positive integer greater than N. The N STAs are STAs that normally report operation information to the AP among the M STAs. For example, the AP#1 in the communication system corresponds to 10 STAs, and 8 STAs among the 10 STAs normally report operation information to the AP, in other words, the AP#1 receives operation information reported by 8 STAs among the 10 STAs. In this case, N is equal to 8.
[0131] For N STAs, each STA reports one piece of operation information to the AP. Thus, N STAs report N pieces of operation information. The operation information indicates an operation during a certain period, and the operation is transmission or skipping transmission. The period includes multiple slots. The multiple slots are multiple slots between the time when the STA last successfully reported the operation information and the current time. For example, STA1 last successfully reported the operation information at time t0, and the current time is time t1. In this case, the multiple slots are multiple slots between t0 and t1. In other words, the operation information reported by each STA includes operations in multiple slots. The operation information reported by each STA is as follows:
number
number
[0132] In addition, the operation information is carried in the first frame reported by the STA. It can be understood that each STA uses the first frame of the STA to carry the operation information, and then reports the first frame to the AP. The first frame includes a Category field and an Action Details field. The Category field indicates the category of the first frame, and the Action Details field indicates the operation information reported by the STA.
[0133] In an optional embodiment, the first frame is a management frame newly added by the STA. For example, the STA adds a management frame, namely, frame 1, and frame 1 is used to carry action information. The frame structure of frame 1 is shown in Figure 4. Frame 1 includes a Category field and an Action Details field. The Category field indicates the category of frame 1, the Action Details field indicates the action information, and the action information is carried in a training data element subfield.
[0134] In another optional embodiment, the first frame is a frame in an existing management frame in the protocol. For example, the first frame is a Quality of Service Action (QoS Action) frame, and the frame structure of the first frame is shown in FIG. 5. In this case, the category of the first frame indicated by the Category field is a QoS Action frame, and the QoS Action subfield in the Action Details field follows the Category field. The STA uses an unused value in the QoS Action field to indicate the action information to be reported, i.e., the content of the training data element subfield in the Action Details field. For example, the QoS Action field includes 2 bits, and the values 00, 01, and 11 represented by the 2 bits of the QoS Action field are used, but the value 10 is not used. In this case, the STA uses the value 10 to indicate the action information to be reported, i.e., the content of the training data element subfield in the Action Details field.
[0135] For the element format of the training data element indicating the operation information, please refer to Figure 6 (a). As shown in Figure 6 (a), the training data element includes an element identification (Element ID) subfield, a length subfield, an element ID extension subfield, and a training data subfield. When all values in the current Element ID subfield are used, the element ID subfield and the Element ID extension subfield together indicate the ID of the training data. The Length subfield indicates the length of the training data. The training data indicates the operation information reported by the STA.
[0136] When the element format of the Training data in the first frame corresponding to each STA is different, the content of the operation information reported by the STA is also different. With respect to the element format of the Training data, the following describes some optional implementations of the operation details field, that is, describes optional implementations of the operation information.
[0137] 1. The operation details field includes a time indication subfield, and data 1 to data T subfields, where T is a positive integer.
[0138] For the element format of the training data, see Fig. 6(a). The training data includes time, and data 1 to data T. The operation details field includes a time indication subfield, and data 1 to data T subfields.
[0139] The time indication subfield indicates the time when the STA last successfully received the first response information, and the time indication subfield may be implemented by using a timestamp, a sequence number, etc. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. For example, the first response information may be acknowledgement (ACK) information. That is, when the STA receives the first response information, it indicates that the STA has successfully reported the operation information. Therefore, the time indication subfield indicates the time when the STA last successfully reports the operation information.
[0140] The Data 1 subfield indicates an operation in the first slot after the STA last successfully receives the first response information. In other words, the Data 1 subfield indicates an operation in the first slot after the STA last successfully reports the operation information. The Data T subfield indicates an operation performed in the Tth slot after the STA last successfully receives the first response information. In other words, the Data T subfield indicates an operation performed by the STA in the Tth slot after the STA last successfully reports the operation information.
[0141] In other words, when each STA reports its operational information to the AP, the STA reports the time when the STA last successfully reported its operational information and the operation in each slot since the STA last successfully reported its operational information, thereby allowing the AP to obtain the operation sensed by the STA in each slot since the STA last successfully reported its operational information.
[0142] 2. The Action Details field includes a Time Indication subfield, Action 1 subfield to Action P subfield, ..., and Time 1 subfield to Time P subfield, where P is a positive integer.
[0143] For the element format of the training data, see Fig. 6(b). Unlike Fig. 6(a), the training data includes start time, action 1, time 1, ..., action P, and time P. In this case, the action details field includes a time indication subfield, an action 1 subfield, a time P subfield, ..., action P subfield, and a time P subfield.
[0144] The time indication subfield indicates the time when the STA last successfully received the first response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. In this case, the time indication subfield indicates the time when the STA last successfully reported the operation information.
[0145] The Action 1 subfield indicates the first action after the STA last successfully receives the first response information. In other words, the Action 1 subfield indicates the first action after the STA last successfully reports the action information. The Time 1 subfield indicates the duration of Action 1 or the end time of Action 1. The Action P subfield indicates the Pth action between the current time and the point at which the STA last successfully receives the first response information. In other words, the Action P subfield indicates the Pth action between the current time and the point at which the STA last successfully reports the action information. The Time P subfield indicates the duration of Action P or the end time of Action P.
[0146] It will be understood that action 1 is the first action after the STA has last successfully reported action information. If the time 1 subfield indicates the duration of action 1 and the time P subfield indicates the duration of action P, and if action 1 does not change, duration 1 is continuously accumulated, or if action 1 changes, action 2 is added, and duration 2 of action 2 is recorded, it is recorded up to the last action before the current time (i.e., action P). The STA reports the recorded action information to the AP, i.e., reports to the AP the time when the action information was last successfully reported, action 1 and its duration, action 2 and its duration, ..., and action P and its duration.
[0147] For example, if STA1 does not transmit a packet in the first slot after the operation information was reported successfully last time, operation 1 is recorded as skipping transmission. If STA1 does not transmit a packet in the first slot to the third slot, duration 1 is accumulated as 3 slots. In the fourth slot, STA1 changes the operation of skipping transmission of a packet to the operation of transmitting a packet, and STA1 adds operation 2, and operation 2 is transmitting. If the operation of transmitting a packet continues until the present time (the ninth slot), STA1 records duration 2 of operation 2 as 6 slots. Thus, the operation information reported by STA1 to the AP includes the time when STA1 reports operation information successfully last time, operation 1 is skipping transmission, the duration of skipping transmission is 3 slots, operation 2 is transmitting, and the duration of transmission is 6 slots.
[0148] In other words, each STA reports the time when the STA last successfully reports its operation information, the operations performed by the STA from the time when the STA last successfully reports its operation information to the current time, and the duration or end time of each operation. This embodiment helps the AP learn the operation behavior of each STA in each slot since the STA last successfully reports its operation information.
[0149] 3. The action information field includes a time 1 indication subfield, an action 1 subfield, ..., a time P indication subfield, and an action P subfield, where P is a positive integer.
[0150] For the element format of the training data, see Fig. 6(c). Unlike Fig. 6(a) and Fig. 6(b), the training data includes time 1, action 1, time 2, action 2, ..., time P, and action P. In this case, the action details field includes a time 1 instruction subfield, an action 1 subfield, ..., a time P instruction subfield, and an action P subfield.
[0151] The time 1 indication subfield indicates the start time of the operation 1. The operation 1 subfield indicates the first operation performed after the STA successfully received the first response information last time. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. In this case, the operation 1 subfield indicates the first operation performed after the STA successfully reported the operation information last time. The time P indication subfield indicates the start time of the operation P. The operation P subfield indicates the Pth operation between the current time and the time when the STA successfully received the first response information last time. In other words, the operation P subfield indicates the Pth operation between the current time and the time when the STA successfully transmits the operation information last time.
[0152] It will be understood that action 1 is the first action after the STA last successfully reports the action information, and time 1 marks the start time of action 1. If action 1 changes and the STA records action 2 and the start time of action 2 (time 2), then the last action among the actions from the current time to the time when the action information was last successfully reported and the start time of the action (action P and time P) are recorded, and the STA reports the recorded action information to the AP.
[0153] In other words, each STA reports to the AP the start time of each operation and each operation that has occurred since the STA last successfully reported operation information, thereby allowing the AP to obtain behavioral information regarding the STA's transmission of packets or skipping of transmissions in multiple slots based on the operations and start times of the operations reported by the STA.
[0154] 4. The Action Information field includes a Time 1 Indication subfield, a Duration 1 subfield, ..., a Time K Indication subfield, and a Duration K subfield, where K is a positive integer.
[0155] The element format of the training data may be shown in Figure 6(d). Unlike Figures 6(a) to 6(c), the training data includes time 1, duration 1, time 2, duration 2, ..., time K, and duration K. In this case, the action details field includes a time 1 instruction subfield, a duration 1 subfield, ..., a time K instruction subfield, and a duration K subfield.
[0156] The Time 1 Indication subfield indicates the start time / end time of Operation 1. Operation 1 is a transmission operation performed when the STA transmits a packet for the first time after the STA successfully receives the first response information last time and does not receive the second response information. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the STA, and the second response information is the response information transmitted when the AP successfully receives the packet transmitted by the STA. In this case, Operation 1 is an operation performed when the STA transmits a packet for the first time after the STA successfully reports the operation information last time, but fails to transmit the packet. The Duration 1 subfield indicates the duration of Operation 1. In other words, the Duration 1 subfield indicates the packet length of the packet transmitted by Operation 1.
[0157] The time K indication subfield indicates the start time / end time of operation K. Operation K is a transmission operation performed when the STA transmits a packet for the Kth time after the STA previously successfully receives the first response information and does not receive the second response information. In this case, operation K is an operation performed when the STA transmits a packet for the Kth time after the STA previously successfully reports the operation information, but fails to transmit the packet. The duration K subfield indicates the duration of operation K. In other words, the duration K subfield indicates the packet length of the packet transmitted by operation K.
[0158] This is because the AP cannot know which STA is trying to access the channel only when multiple STAs transmit packets at the same time and channel collision occurs. Therefore, each STA only needs to report operation information to the AP when it fails to transmit a packet, i.e., each STA reports the transmission operation it performed when it fails to transmit a packet, the start time / end time of the operation, and the packet length of the packet transmitted each time, so that the AP knows which STA is trying to access the channel when channel collision occurs.
[0159] 5. The Operation Information field includes a first time 1 indication subfield, a second time 1 indication subfield, ..., a first time K indication subfield, and a second time K indication subfield, where K is a positive integer.
[0160] For the element format of the training data, see Fig. 6(e). Unlike Fig. 6(a) to Fig. 6(d), the training data includes a first time 1, a second time 1, ..., a first time K, and a second time K. In this case, the operation details field includes a first time 1 instruction subfield, a second time 1 instruction subfield, ..., a first time K instruction subfield, and a second time K instruction subfield.
[0161] The first time 1 indication subfield indicates the start time of operation 1. Operation 1 is a transmission operation performed when the STA transmits a packet for the first time after the STA successfully receives the first response information last time and does not receive the second response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA, and the second response information is response information transmitted when the AP successfully receives the packet transmitted by the STA. In this case, operation 1 is an operation performed when the STA transmits a packet for the first time after the STA successfully reports the operation information last time, but fails to transmit the packet. The second time 1 indication subfield indicates the end time of operation 1.
[0162] The first time K subfield indicates the start time of operation K. Operation K is a transmission operation performed when the STA transmits a packet for the Kth time after previously successfully receiving the first response information and does not receive the second response information. In this case, operation K is an operation performed when the STA transmits a packet for the Kth time after previously successfully reporting the operation information, but fails to transmit the packet. The second time K indication subfield indicates the end time of operation K.
[0163] It can be seen that operation 1 to operation K are operations performed when a STA fails to transmit a packet after it has previously successfully reported its operation information. In this case, each STA reports to the AP the start time and end time of each failed packet transmission after it previously successfully reported its operation information, so that the AP can determine the slots in which packet transmission has failed each time and the packet length of the transmitted packet based on the start time and end time of each failed packet transmission, and further obtain the behavior information of each STA in each slot.
[0164] It can be seen that different format elements of the above five Training data fields represent different contents in the operation information reported by each STA, which makes the operation information reported by the STA to the AP more flexible.
[0165] It should be understood that the time at which each STA reports operational information to the AP is predefined by the AP. For example, the AP predefines that each STA reports operational information to the AP based on a predefined period, and then each STA reports operational information to the AP at intervals of the predefined period. In addition, the reporting time predefined by the AP for each STA may be different. For example, the AP predefines that STA1 reports operational information to the AP at intervals of predefined time 1, and STA2 reports operational information to the AP at intervals of predefined time 2.
[0166] Optionally, the time when each STA reports the operation information to the AP is notified to each STA by the AP by using signaling. For example, the AP notifies each STA of the time when it reports the operation information by using downlink control information (DCI). As another example, the AP notifies STA1 of the time #1 when STA1 reports the operation information by using DCI #1, and notifies STA2 of the time #2 when STA2 reports the operation information by using DCI #2.
[0167] S102: The AP receives operation information reported separately by the N STAs.
[0168] S103: The AP determines a training result of the first neural network of each STA based on the N pieces of motion information.
[0169] It can be understood that the AP trains the first neural network of each STA based on the N pieces of motion information to obtain the training result of the first neural network of each STA. For example, five STAs report a total of five pieces of motion information, and the five STAs correspond to the first neural network #1 to the first neural network #5, respectively. The AP trains the first neural network #1 of STA1 based on the five pieces of motion information to obtain the training result of the first neural network #1, trains the first neural network #2 of STA2 based on the five pieces of motion information to obtain the training result of the first neural network #2, and finally obtains the training result of the first neural network #5 of STA5.
[0170] It will be understood that the training result of the first neural network is the neural network parameters or gradients of the first neural network. The neural network parameters are the weights and offsets of the neurons in the first neural network. For example, the structure of the first neural network is shown in FIG. 7. The first neural network includes an input layer, an output layer, and multiple intermediate layers, and each layer includes multiple nodes. The nodes are called neurons. The neurons of two adjacent layers are connected to each other.
[0171] For two adjacent layer neurons, the output h of the lower layer neuron is the value obtained by performing an activation function on the weighted sum of all the upper layer neurons x connected to the lower layer neuron. The output can be represented by using a matrix as follows: h = f (wx + b) (1)
[0172] where w is a weight matrix, b is a bias vector, and f is an activation function. Then the output y of the n-th layer neural network can be expressed recursively as: y=f n (w n f n-1 (…)+b n ) (2)
[0173] In other words, the first neural network can be understood as a mapping relationship from input x to output y. The training process of the neural network is a process of obtaining the mapping relationship from existing data, i.e., obtaining w and b. The training result of the first neural network can be the neural network parameters w and b.
[0174] In addition, the AP may train the neural network by using the gradient descent method. Therefore, the training result of the neural network may be a gradient. The gradient is the bias of the loss function of the neural network with respect to the neural network parameters, that is, the bias of the loss function of the neural network with respect to w and b.
[0175] The neural network parameters / gradients are used by the corresponding STA to update the corresponding first neural network, that is, the neural network parameters / gradients of the STA are used to update the first neural network of the STA. For example, if the neural network parameter #1 is the neural network parameter corresponding to STA1, the neural network parameter #1 is used by STA1 to update the first neural network of STA1.
[0176] In an optional embodiment, for the AP to determine the training result of the first neural network of each STA based on N pieces of operation information, the AP inputs the status information of each STA into the first neural network of the corresponding STA to obtain the output of the first neural network; the AP inputs the output of each first neural network into a second neural network to obtain the output of the second neural network, where the output of the second neural network represents the expected reward within a preset time; and the AP trains a third neural network based on the output of the second neural network and the reward function, and determines the training result of each first neural network by minimizing the loss function of the third neural network, where the third neural network includes each first neural network and the second neural network.
[0177] Status information of the STA is obtained based on the behavior information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of behavior information, and a reward function is determined based on the N pieces of behavior information.
[0178] It can be understood that after obtaining the operation information reported by each STA, the AP determines carrier sensing result information or packet transmission result information according to each operation information, and then determines status information according to N pieces of operation information and N pieces of carrier sensing result information, or determines status information according to N pieces of operation information and N pieces of packet transmission result information. The carrier sensing result information or packet transmission result information is:
number
number
number
number
[0179]
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0180]
number
number
number
number
number
[0181] As shown in Figure 8,
number
number
number
number
number
number
number
number
number
number
[0182] It can be seen that the AP first inputs the status information obtained based on the information reported by each STA into the first neural network of the STA to obtain the output of each first neural network, then inputs the outputs of the N first neural networks into the second neural network to obtain the output of the second neural network, and then trains the third neural network based on the loss function to finally obtain the training result of the first neural network. The training result of the first neural network of each STA is determined based on the information reported by the N STAs, not only on the information of that STA. This helps improve the ability of each STA to predict the channel access behavior of other STAs.
[0183] The training process performed by the AP is described below by using an example in which the AP trains each first neural network by using a target Q neural network.
[0184] FIG. 9 is a schematic diagram of training a target Q network. FIG. 9 includes a target Q network and a prediction Q network. The structures of the target Q network and the prediction Q network are shown in FIG. 10. The neural networks shown in FIG. 10 include agent network 1 to agent network N and a mixing network. Agent network 1 to agent network N are the first neural networks of STA1 to STAN, and each agent network corresponds to one STA. The mixing network is the second neural network mentioned above.
[0185] The input of each agent network is the status information of the corresponding STA in the past period, i.e.
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0186] The AP calculates the loss function of the third neural network based on the output of the mixing network and the reward function, and trains the third neural network by minimizing the loss function, i.e., trains each agent network and the mixing network to determine the neural network parameters of each agent network. The loss function of the third neural network is as follows:
number
[0187] where r(t) is the reward function, γ is the discount factor (typically γ = 0.9), and e t represents experience, E represents the experience pool, and |E| is the number of experiences in the experience pool e t and e(t)=(s t , τ t ,a t ,r t ,s t+1 , τ t+1 ) and
number
number
number
[0188] For the process of training the third neural network by the AP, please refer to the schematic diagram shown in Figure 9. That is, the AP updates the neural network parameters of the Q network by using small batch gradient descent. The AP updates the θ - It will be understood that the neural network parameters θ of the predictive neural network are trained by fixing the parameters θ of the target neural network, and then using the loss function and the output of the mixing network. Each time the training is completed C times, the neural network parameters θ are updated to the fixed parameters θ of the target neural network. - Then, the neural network parameters of the predictive Q-network are iteratively trained. The training data for each agent network is determined by minimizing the loss function of a third neural network. Typically, C=100.
[0189] There are several optional implementations for computing the reward function of the third neural network:
[0190] 1. Set the reward function of the third neural network to 1.
[0191] It can be understood that when the first STA determines based on the operation information that the first STA successfully transmits the packet, the AP sets the value of the reward function of the third neural network to 1. The first STA is the STA among the N STAs that has the longest time interval between the time when the second response information is last successfully received and the current time, i.e., the first STA is the STA that has the longest duration since the time when the packet is last successfully received.
[0192] In other words, when the AP determines based on the N pieces of motion information that the STA with the longest duration since the last time the packet was successfully transmitted successfully transmits a packet in multiple slots, the reward function is set to 1. That is, r t =1,
number
number
[0193] 2. Set the reward function to -1 for the first duration.
[0194] When the second STA determines that the packet is successfully transmitted based on the N pieces of motion information, the AP sets the value of the reward function to the first duration -1, i.e.
number
number
[0195] 3. Set the reward function to -1.
[0196] When M STAs among N STAs decide to transmit packets in the same slot based on N pieces of behavior information, the AP sets the reward function to -1, i.e., r t It can be understood that M is equal to or smaller than N, and M is a positive integer equal to or smaller than N. In other words, when the AP determines that some STAs among the N STAs transmit packets in the same slot based on N pieces of operation information, it indicates that channel collision occurs when some STAs transmit packets in that slot, and some STAs cannot transmit packets normally, that is, the reward function is subtracted, specifically, the reward function is subtracted by 1.
[0197] 4. Set the reward function to 0.
[0198] When it is determined based on the N pieces of behavior information that none of the N STAs will transmit a packet in the same slot, the AP sets the value of the reward function to 0, i.e., r t It will be understood that = 0. In other words, when the AP determines based on N pieces of behavior information that none of all the STAs transmit a packet in one slot, no future expected reward occurs, and thus the reward function is set to 1.
[0199] In addition to the four cases mentioned above, the AP may also set the reward function to 0.
[0200] In this embodiment of the present application, if each STA reports operation information at different times, or some STAs among N STAs report operation information at different times, when AP trains neural network at present, some STAs may not report operation information, and only some STAs report latest operation information.In this case, when training neural network of each STA, AP trains the first neural network of each STA by using the operation information reported at present and the operation information reported last time by STAs that do not report operation information at present, so as to carry out intensive training of the first neural network of each STA.In addition, in this way, STAs whose operation information does not change at present do not need to report operation information, thereby reducing the signaling overhead of communication system.
[0201] Compared with the current solution in which the STA trains the neural network of the STA based on the transmission behavior and packet transmission duration observed by the STA, in this embodiment of the present application, the AP trains the first neural network of each STA based on N behavior information of the N STAs, that is, the AP refers to the behavior information of the N STAs when training the first neural network of each STA, so that the AP can better train each first neural network and obtain better training results, which makes the prediction ability of the first neural network better.
[0202] S104: The AP sends the training result of the first neural network of each STA to the corresponding STA.
[0203] S105: For each STA, the STA receives a training result of the first neural network from the AP.
[0204] S106: For each STA, the STA updates a first neural network based on the training result of the first neural network, and when it senses that the channel is idle, determines whether to access the channel based on the updated first neural network and current status information of the STA.
[0205] The STA's current status information includes the STA's operation during the past period, carrier sense results, and packet transmission results.
[0206] In an optional embodiment, as described above, the training result of the first neural network is the neural network parameters of the first neural network. In this case, the STA updates the first neural network based on the training result of the first neural network indicates that the STA updates the previous neural network parameters of the first neural network to the received neural network parameters to obtain an updated first neural network.
[0207] In another optional embodiment, as described above, the training result of the first neural network is the gradient of the first neural network. In this case, the STA updates the first neural network based on the training result of the first neural network means that the STA performs calculations on the gradient to obtain the neural network parameters of the first neural network, and then replaces the original neural network parameters of the first neural network with the neural network parameters to obtain updated neural network parameters. The process of the STA performing calculations on the gradient is represented as θ'=θ+γg, where θ' is the neural network parameters of the first neural network after updating, θ is the neural network parameters of the first neural network before updating, γ is the learning efficiency of the first neural network, and g is the gradient.
[0208] In an optional embodiment, the STA updates the first neural network based on the training result of the first neural network, and when the STA senses that the channel is idle, the STA determines whether to access the channel based on the updated first neural network and the sensed operation information includes: the STA inputs the operation information into the updated first neural network and outputs a first value and a second value, where the first value represents an expected reward obtained by accessing the channel, and the second value represents an expected reward obtained by skipping access to the channel. The STA determines to access the channel if the first value is greater than the second value, or the STA determines to skip access to the channel if the first value is less than the second value. Specifically, when the STA senses that the channel is idle, the STA determines whether to access the channel based on the first value and the second value output by the updated first neural network.
[0209] An example in which the first neural network of the STA is a part of a Q neural network is used to explain the embodiment in which, when the STA senses that the channel is idle, the STA determines whether to access the channel based on the training result of the first neural network and the operation information detected at the present time. In this case, the structure of the first neural network of the STA is shown in Figure 10. The STA uses the operation information obtained by the STA sensing the current channel as the input of the agent network to determine whether to access the channel.
number
number
number
number
number
number
[0210] In this embodiment of the present application, when the STA senses that the channel is idle, the STA may decide whether to access the channel based on the training result of the first neural network trained by the AP and the operation information sensed by the STA at the present time. The training result of the first neural network is also obtained by training the first neural network based on the operation information of each STA by the AP. The first neural network has high predictability. Therefore, in this way, when the STA decides to access the channel, the probability that the packet can be transmitted successfully is high, that is, the probability of channel collision is low. This can improve the system throughput and reduce the latency of the communication system.
[0211] Please refer to FIG. 11 for a block diagram of the implementation of this embodiment of the present application. The implementation block diagram of FIG. 11 includes a centralized training unit corresponding to the AP and a distributed execution unit corresponding to the STA. Both the centralized training unit corresponding to the AP and the distributed execution unit corresponding to the STA include a first neural network of each STA, and the neural network parameter of the first neural network is θ i It is.
[0212] The centralized training corresponding to the AP indicates that the AP trains each first neural network based on N pieces of status information obtained based on N pieces of operation information reported by N pieces of STAs to obtain a training result of each first neural network. In other words, the training result of each first neural network is obtained based on N pieces of operation information. This can improve the predictability of the first neural network. Each piece of operation information is obtained by observing the historical environment by each STA.
[0213] The distributed execution corresponding to each STA shows that after each STA obtains the training result of the first neural network distributed by the AP, the STA updates the first neural network of the STA by using the training result, and then when the STA senses that the channel is idle, the STA determines whether to access the channel based on the sensed operation information and the updated first neural network. In the manner in which the STA determines whether to access the channel based on the updated first neural network, the STA can more accurately determine whether to access the channel. This can improve system throughput and reduce system communication latency.
[0214] It will be appreciated that this embodiment of the present application is applicable to all multi-agent reinforcement learning algorithms implemented with centralized training distribution, such as the Aho-Corasick automaton algorithm, the Proximal Policy Optimization (PPO) algorithm, and the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm.
[0215] In this embodiment of the present application, N STAs report operation information to the AP. The AP determines the training result of the first neural network of each STA based on the N operation information reported by the N STAs, and transmits the training result of the first neural network of each STA to the corresponding STA, so that each STA can update the first neural network based on the training result of the first neural network, and when it senses that the channel is idle, it can determine whether to access the channel based on the updated first neural network and the sensed operation information. The AP trains the first neural network of each STA based on the N operation information, so that the first neural network has better predictability, thereby helping to improve each STA's ability to predict the channel access behavior of other STAs. That is, when each STA transmits a packet, the probability of channel collision of the STA is lower. This improves system throughput and reduces communication latency.
[0216] In addition, compared with the implementation in the current study in which the STA trains its neural network based on the historical behavior information of all STAs in the network, in this embodiment of the present application, each STA independently decides whether to access the channel based on the training result of the first neural network distributed by the AP and the historical behavior information sensed by the STA, without relying on the behavior information of other STAs other than the STA. Therefore, the actual operability of each STA is better.
[0217] In the current study, each STA may further train its neural network and report the neural network parameters obtained by training to the AP. Then, the AP processes the neural network parameters of all STAs to obtain new neural network parameters, and broadcasts the new neural network parameters to each STA. Then, the STA decides whether to access the channel based on the new neural network parameters. Compared with the one in the study, in this embodiment of the present application, the neural network of each STA is trained centrally by the AP, and each STA in the network does not need to train its neural network, that is, each STA in the network does not need to have the ability to train its neural network independently. This can reduce the interaction between each STA and the AP, and reduce the signaling overhead and computational power of the system.
[0218] FIG. 12 is a diagram of a comparison between the system throughput in this embodiment of the present application and the system throughput when the channel collision is resolved by using CSMA / CA technology. The system throughput in this embodiment of the present application is higher than the throughput when the channel collision is resolved by using CSMA / CA technology. FIG. 13 is a diagram of a comparison between the average latency of the system in this embodiment of the present application and the average latency of the system when the channel collision is resolved by using CSMA / CA technology. The average latency in this embodiment of the present application is lower than the average latency when the channel collision is resolved by using CSMA / CA technology. FIG. 14 is a diagram of a comparison between the latency jitter of the system in this embodiment of the present application and the latency jitter of the system when the channel collision is resolved by using CSMA / CA technology. The latency jitter in this embodiment of the present application is lower than the latency jitter when the channel collision is resolved by using CSMA / CA technology.
[0219] 4. Each STA reports operation information and carrier sensing result information, or each STA reports operation information and packet transmission result information.
[0220] It should be understood that in addition to reporting operation information, each STA may also report carrier sensing result information or packet transmission result information.
[0221] 1. Each STA reports its operating information and carrier sensing result information.
[0222] In other words, in addition to the operation information, each STA also reports carrier sensing result information. The carrier sensing result information includes the carrier sensing result of each of a number of slots in the current time after the STA last successfully reported the operation information. The AP receives the operation information and carrier sensing result information reported separately by the N STAs.
[0223] In this case, the N pieces of operation information and the N pieces of carrier sensing result information are carried in the operation details field of the first frame reported by the STA. The frame structure of the first frame is shown in FIG. 5. The details will not be described again. The operation details field includes a time indication subfield, and a data 1 subfield to a data T subfield, where T is a positive integer. The time indication subfield indicates the time when the STA last successfully received the first response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. In this case, the time indication subfield indicates the time when the STA last successfully reported the operation information.
[0224] The Data 1 subfield indicates the carrier sensing result and the operation performed in the first slot after the STA last successfully received the first response information. The Data T subfield indicates the carrier sensing result and the operation performed in the Tth slot after the STA last successfully received the first response information. It will be understood that the Data 1 subfield indicates the carrier sensing result and the operation performed in the first slot after the STA last successfully reported operational information. The Data T subfield indicates the carrier sensing result and the operation performed in the Tth slot after the STA last successfully reported operational information.
[0225] The above S103 in which the AP determines the training result of the first neural network of each STA based on the N pieces of operation information may be that the AP determines the training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of carrier sensing result information. It can be understood that the AP does not need to determine the carrier sensing result information based on the operation information, and can directly determine the training result of the first neural network of the STA based on the received operation information and the received carrier sensing result information. This reduces the processing complexity of the AP.
[0226] Optionally, the above S103 in which the AP determines the training result of the first neural network of each STA based on the N pieces of operation information may be that the AP determines the training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of carrier sensing result information determined based on the N pieces of operation information. In other words, in this embodiment, even if the STA reports the carrier sensing result information, the AP may further determine the training result of the first neural network based on the carrier sensing result information determined based on the operation information.
[0227] 2. Each STA reports its operation information and packet transmission result information.
[0228] In other words, in addition to the operation information, each STA also reports packet transmission result information, which includes packet transmission results obtained when the STA transmits a packet in multiple slots within the current time after the STA last successfully reports the operation information. The AP receives the operation information and carrier sensing result information reported separately by the N STAs.
[0229] In this case, the N pieces of operation information and the N pieces of packet transmission result information are carried in the operation details field of the first frame reported by the STA. The frame structure of the first frame is shown in FIG. 5. The details will not be described again. The operation details field includes a time indication subfield, and a data 1 subfield to a data T subfield, where T is a positive integer. The time indication subfield indicates the time when the STA last successfully received the first response information. The first response information is response information sent when the AP successfully receives the operation information sent by the STA. In this case, the time indication subfield indicates the time when the STA last successfully reported the operation information.
[0230] The Data 1 subfield indicates the packet transmission result and the operation performed in the first slot after the STA last successfully received the first response information. The Data T subfield indicates the packet transmission result and the operation performed in the Tth slot after the STA last successfully received the first response information. It can be understood that the Data 1 subfield indicates the packet transmission result and the operation performed in the first slot after the STA last successfully reported the operation information. The Data T subfield indicates the packet transmission result and the operation performed in the Tth slot after the STA last successfully reported the operation information.
[0231] The above S103 of the AP determining the training result of the first neural network of each STA based on the N pieces of operation information may be the AP determining the training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of packet transmission result information. It can be understood that the AP does not need to determine the packet transmission result information based on the operation information, and can directly determine the training result of the first neural network of the STA based on the received operation information and the received packet transmission result information. This reduces the processing complexity of the AP.
[0232] Optionally, the above S103 of the AP determining the training result of the first neural network of each STA based on the N pieces of operation information may be that the AP determines the training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of packet transmission result information determined based on the N pieces of operation information. In other words, in this embodiment, even if the STA reports the packet transmission result information, the AP can further determine the training result of the first neural network based on the packet transmission result information determined based on the operation information.
[0233] It should be understood that when each STA reports operation information and carrier sensing result information or reports operation information and packet transmission result information, the AP processes the N operation information and N carrier sensing result information or N operation information and N packet transmission result information reported by the N STAs in the same manner as the processing manner in the channel access method 100. Details will not be described again. For example, when each STA reports operation information and carrier sensing result information, in S103, the status information of the STA is obtained according to the operation information and carrier sensing result information of the STA, the neural network parameters of the second neural network are obtained according to the N operation information and N carrier sensing result information, and the reward function is determined according to the N operation information and N carrier sensing result information.
[0234] 5. An embodiment in which the AP distributes the training results of the first neural network to each STA.
[0235] When the neural network parameters of the first neural networks corresponding to the N STAs are the same or different, the embodiment of the AP distributing the training result of the first neural network to each STA may be different. The following describes some optional embodiments of the AP distributing the training result of the first neural network to the N STAs.
[0236] 1.N STAs share neural network parameters.
[0237] It will be understood that when N STAs share neural network parameters, the AP sending the training results of the first neural network of each STA to the corresponding STA means that the AP broadcasts the training results of the first neural network to the N STAs.
[0238] In other words, if the neural network parameters of the first neural network of each STA are the same, the training result of each first neural network determined by the AP based on the operation information reported by the N STAs is also the same. Specifically, the AP determines the training result of one first neural network based on the operation information reported by the N STAs. The AP can distribute the determined training result of the first neural network to the N STAs by multicasting. This can reduce system overhead.
[0239] 2. S STAs out of N STAs share neural network parameters.
[0240] It should be understood that S STAs among the N STAs share neural network parameters, and S is a positive integer equal to or less than N. In this case, the AP sending the training results of the first neural network of each STA to the corresponding STA means that the AP multicasts the training results of the first neural network corresponding to the S STAs to the S STAs, and unicasts the training results of the (NS) first neural networks to the corresponding STAs.
[0241] In other words, when some STAs among N STAs share neural network parameters and other STAs do not share neural network parameters, AP distributes the training result of the first neural network of the STAs that share neural network parameters to some STAs by multicast, and unicasts the training result of the first neural network of the STAs that do not share neural network parameters to corresponding STAs. This method can also reduce system overhead.
[0242] 3. The N STAs do not share neural network parameters.
[0243] It should be understood that if the neural network parameters of the N first neural networks corresponding to the N STAs are different, the training results of the first neural networks determined by the AP based on the information reported by the N STAs will also be different. Thus, the training results of the first neural networks are unicast to the corresponding STAs.
[0244] In one optional embodiment, each STA may report information indicating whether the STA and the other STAs share neural network parameters to the AP, so that the AP can determine whether some or all of the N STAs share neural network parameters based on the indication information reported by the STA, and further determine an embodiment for distributing the training results of the first neural network to each STA.
[0245] In one optional embodiment, before each STA reports operational information or before the AP sends the training results of each first neural network to the corresponding STA, the AP distributes the structure of each STA's first neural network to each STA, thereby each STA obtains the structure of the STA's first neural network.
[0246] In another optional embodiment, the first neural network of each STA is predefined by the AP. Specifically, each STA knows the structure of the first neural network of the STA and the neural network parameters of the first neural network in advance, and the AP does not need to inform each STA by using signaling. This can reduce the signaling overhead of the AP.
[0247] In yet another optional embodiment, before each STA reports operation information or before the AP transmits the training result of each first neural network to the corresponding STA, the AP distributes the structure of the multiple first neural networks to each STA. When deciding to use the structure of the first neural network, the STA reports the determined structure of the first neural network to the AP, so that the AP obtains the structure of the first neural network specifically used by each STA. In this way, each STA can flexibly select the structure of the first neural network used by the STA from the structures of the multiple first neural networks distributed by the AP.
[0248] In this embodiment of the present application, each STA may request the AP to update the training result of the STA's first neural network, and the AP may send the training result of the STA's first neural network to the STA when receiving the request information from the STA.
[0249] For the training results of the first neural networks of the N STAs, the training results of each first neural network are carried in a second frame. For the frame structure of the second frame, please refer to FIG. 15. The second frame includes an element ID subfield, a length subfield, an element ID extension subfield, and training results (neural network parameters or gradients). The second frame may be an existing management frame or a newly added management frame. For a specific embodiment, please refer to the embodiment of the first frame. The details will not be described again.
[0250] 6. Communications Equipment To implement the functions in the methods provided in the embodiments of the present application, the AP or STA may include a hardware structure and / or a software module, which implements the above-mentioned functions by using a hardware structure, a software module, or a combination of a hardware structure and a software module.Whether the functions among the above-mentioned functions are implemented by using a hardware structure, a software module, or a combination of a hardware structure and a software module depends on the specific application and design constraints of the technical solution.
[0251] As shown in Fig. 16, an embodiment of the present application provides a communication device 1600. The communication device 1600 may be a component (e.g., an integrated circuit or chip) of an AP or a component (e.g., an integrated circuit or chip) of a STA. Alternatively, the communication device 1600 may be another communication unit configured to implement the method in the method embodiment of the present application. The communication device 1600 may include a communication unit 1601 and a processing unit 1602. Optionally, the device may further include a storage unit 1603.
[0252] In one possible design, one or more units in FIG. 16 may be implemented by one or more processors, one or more processors and memories, one or more processors and transceivers, or one or more processors, memories, and transceivers. This is not limited in this embodiment of the present application. The processor, memory, and transceiver may be located separately or integrated.
[0253] The communication device 1600 has a function of implementing the AP described in the embodiment of the present application. Optionally, the communication device 1600 has a function of implementing the STA described in the embodiment of the present application. For example, the communication device 1600 includes a module or unit or means corresponding to performing the steps of the AP in the embodiment of the present application by the AP. The function or unit or means may be implemented by software, or may be implemented by hardware, or may be implemented by hardware running corresponding software, or may be implemented as a combination of software and hardware. For details, please refer to the corresponding description in the corresponding method embodiment above.
[0254] In one possible design, the communications device 1600 may include: A communication unit 1601 configured to receive operation information reported separately by N stations STAs, the N pieces of operation information being used to determine a training result of a first neural network of each STA, where N is a positive integer; A processing unit 1602 configured to determine a training result of a first neural network of each STA based on the N pieces of operation information, The communication unit 1601 is further configured to send the training result of the first neural network of each STA to the corresponding STA; may include:
[0255] In one optional implementation, the operation information indicates an operation over a period of time, the operation being a transmission or a skipped transmission.
[0256] In one optional embodiment, the communication unit 1601 is further configured to receive carrier sensing result information or packet transmission result information reported separately by the N STAs, and when determining the training result of the first neural network of each STA based on the N pieces of operation information, the processing unit 1602 is specifically configured to determine the training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of carrier sensing result information, or determine the training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of packet transmission result information.
[0257] In one optional implementation, the training results are neural network parameters or gradients, and the neural network parameters / gradients are used by the corresponding STA to update the corresponding first neural network.
[0258] In one optional embodiment, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield and data 1 subfield to data T subfield, where T is a positive integer.
[0259] The time indication subfield indicates the time when the STA last successfully received the first response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The data 1 subfield indicates the operation to be performed in the first slot after the STA last successfully received the first response information. The data T subfield indicates the operation to be performed in the Tth slot after the STA last successfully received the first response information.
[0260] In another optional embodiment, the action information is carried in an action details field of the first frame reported by the STA. The action details field includes a time indication subfield, an action 1 subfield, a time 1 subfield, ..., an action P subfield, and a time P subfield, where P is a positive integer.
[0261] The time indication subfield indicates the time when the STA last successfully received the first response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The operation 1 subfield indicates the first operation after the STA last successfully received the first response information. The time 1 subfield indicates the duration of operation 1 or the end time of operation 1. The operation P subfield indicates the Pth operation between the time when the STA last successfully received the first response information and the current time. The time P subfield indicates the duration of operation P or the end time of operation P.
[0262] In yet another optional embodiment, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time 1 indication subfield, an operation 1 subfield, ..., a time P indication subfield, and an operation P subfield, where P is a positive integer.
[0263] The time 1 indication subfield indicates the start time of operation 1. The operation 1 subfield indicates the first operation after the STA successfully received the first response information last time. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The time P indication subfield indicates the start time of operation P. The operation P subfield indicates the Pth operation between the time when the STA successfully received the first response information last time and the current time.
[0264] In yet another optional embodiment, the operation information is carried in an operation details field of the first frame reported by the STA. The operation details field includes a time 1 indication subfield, a duration 1 subfield, ..., a time K indication subfield, and a duration K subfield, where K is a positive integer.
[0265] The Time 1 Indication subfield indicates the start time / end time of Operation 1. Operation 1 is a transmission operation when the STA transmits a packet for the first time after successfully receiving the first response information last time and does not receive the second response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The second response information is response information transmitted when the AP successfully receives the packet transmitted by the STA. The Duration 1 subfield indicates the duration of Operation 1.
[0266] The Time K Indication subfield indicates the start time / end time of operation K. Operation K is a transmission operation when the STA transmits a packet for the Kth time after the STA has successfully received the first response information last time and does not receive the second response information. The Duration K subfield indicates the duration of operation K.
[0267] In yet another optional embodiment, the operation information is carried in an operation details field of the first frame reported by the STA, the operation details field including a first time 1 indication subfield, a second time 1 indication subfield, ..., a first time K indication subfield, and a second time K indication subfield, where K is a positive integer.
[0268] The first time 1 indication subfield indicates the start time of operation 1. Operation 1 is a transmission operation when the STA transmits a packet for the first time after successfully receiving the first response information last time and does not receive the second response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the STA. The second response information is response information transmitted when the AP successfully receives the packet transmitted by the STA. The second time 1 indication subfield indicates the end time of operation 1.
[0269] The first time K indication subfield indicates the start time of operation K. Operation K is a transmission operation when the STA transmits a packet for the Kth time after the STA has successfully received the first response information last time and has not received the second response information. The second time K indication subfield indicates the end time of operation K.
[0270] In a further optional embodiment, the operation information and carrier sensing result information are carried in an Operation Details field of the first frame reported by the STA. The Operation Details field includes a Time Indication subfield, and Data 1 subfield through Data T subfield, where T is a positive integer.
[0271] The time indication subfield indicates the time when the STA last received the first response information normally. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the STA.
[0272] The Data 1 subfield indicates the carrier sensing result and the operation to be performed in the first slot after the STA last successfully received the first response information. The Data T subfield indicates the carrier sensing result and the operation to be performed in the Tth slot after the STA last successfully received the first response information.
[0273] In a further optional embodiment, the operation information and packet transmission result information are carried in an operation details field of the first frame reported by the STA. The operation details field includes a time indication subfield and data 1 subfield to data T subfield, where T is a positive integer.
[0274] The time indication subfield indicates the time when the STA last received the first response information normally. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the STA.
[0275] The Data 1 subfield indicates the packet transmission result and the operation to be performed in the first slot after the STA last successfully received the first response information. The Data T subfield indicates the packet transmission result and the operation to be performed in the Tth slot after the STA last successfully received the first response information.
[0276] In one optional embodiment, when determining the training result of the first neural network of each STA based on the N pieces of operation information, the processing unit 1602 is specifically configured to: input the status information of each STA into the first neural network of the corresponding STA to obtain an output of the first neural network, input the output of each first neural network into a second neural network to obtain an output of the second neural network, the output of the second neural network representing an expected reward within a preset time, train a third neural network based on the output of the second neural network and the reward function, and determine the training result of each first neural network by minimizing a loss function of the third neural network, where the third neural network includes each of the first neural network and the second neural network.
[0277] Status information of the STA is obtained based on the behavior information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of behavior information, and a reward function is determined based on the N pieces of behavior information.
[0278] Further, status information of the STA is obtained based on the operation information and the carrier sensing result information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of operation information and the N pieces of carrier sensing result information, and a reward function is determined based on the N pieces of operation information and the N pieces of carrier sensing result information.
[0279] Alternatively, status information of the STA is obtained based on operation information and packet transmission result information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of operation information and the N pieces of packet transmission result information, and a reward function is determined based on the N pieces of operation information and the N pieces of packet transmission result information.
[0280] In one optional embodiment, the processing unit 1602 is further configured to set a value of the reward function to 1 when it determines, based on the N pieces of operation information, that the first STA successfully transmits the packet, and the first STA is the STA among the N STAs that has the longest time interval between the time when the second response information was last successfully received and the current time.
[0281] In another optional embodiment, the processing unit 1602 is further configured to set the value of the reward function to a first duration -1 when determining, based on the N pieces of operation information, that the second STA successfully transmits the packet, wherein the second STA is a STA other than the first STA among the N STAs, the first STA is a STA among the N STAs that has the longest time interval between the time when the second response information was last successfully received and the current time, and the first duration is the duration between the time when the second STA last successfully received the second response information and the current time.
[0282] In yet another optional implementation, the processing unit 1602 is further configured to set a value of the reward function to −1 when it determines, based on the N pieces of operation information, that M STAs among the N STAs transmit packets in the same slot, where M is a positive integer less than or equal to N.
[0283] In yet another optional implementation, the processing unit 1602 is further configured to set a value of the reward function to 0 when determining, based on the N pieces of operation information, that none of the N STAs transmits a packet in the same slot.
[0284] In one optional embodiment, the N STAs share neural network parameters, and when each STA transmits the training result of the first neural network to the corresponding STA, the communication unit 1601 is specifically configured to broadcast the training result of the first neural network to the N STAs.
[0285] In another optional embodiment, S STAs among the N STAs share neural network parameters, where S is a positive integer less than or equal to N, and when transmitting the training results of the first neural network of each STA to the corresponding STA, the communication unit 1601 is specifically configured to multicast the training results of the first neural networks corresponding to the S STAs to the S STAs and unicast the training results of the (NS) first neural networks to the corresponding STAs.
[0286] In yet another optional embodiment, when the N STAs do not share neural network parameters, the training result of each first neural network is unicast to the corresponding STA.
[0287] This embodiment of the present application and the above-mentioned method embodiment are based on the same concept and achieve the same technical effect. For the specific principle, please refer to the description of the above-mentioned embodiment. The details will not be described again.
[0288] In another possible design, the communications device 1600 may include: A communication unit 1601 configured to report operation information to an access point AP, the operation information being used to determine a training result of a first neural network of the processing unit; The communication unit 1601 is further configured to receive a training result of the first neural network from the AP, and the training result of the first neural network is used to update the first neural network to determine whether the processing unit accesses the channel; a processing unit 1602 configured to update a first neural network based on a training result of the first neural network, and determine whether to access the channel based on the updated first neural network and current status information of the processing unit when detecting that the channel is idle; may include:
[0289] In one optional implementation, the operation information indicates an operation over a period of time, the operation being a transmission or a skipped transmission.
[0290] In an optional embodiment, the communication unit 1601 is further configured to report the carrier sensing result information or the packet transmission result information to the AP, and the carrier sensing result information or the packet transmission result information is used to determine the training result of the first neural network of the processing unit.
[0291] In one optional implementation, the training results are neural network parameters or gradients, and the neural network parameters / gradients are used by processing unit 1602 to update the first neural network.
[0292] In one optional implementation, the motion information is carried in an motion details field of the first frame reported by processing unit 1602. The motion details field includes a time indication subfield, and data 1 subfield through data T subfield, where T is a positive integer.
[0293] The time indication subfield indicates the time when the processing unit 1602 last successfully received the first response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the processing unit 1602. The data 1 subfield indicates the operation to be performed in the first slot after the processing unit 1602 last successfully received the first response information. The data T subfield indicates the operation to be performed in the Tth slot after the processing unit 1602 last successfully received the first response information.
[0294] In another optional implementation, the action information is carried in an action details field of the first frame reported by the processing unit 1602. The action details field includes a time indication subfield, an action 1 subfield, a time 1 subfield, ..., an action P subfield, and a time P subfield, where P is a positive integer.
[0295] The time indication subfield indicates the time when the processing unit 1602 last successfully received the first response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the processing unit 1602. The operation 1 subfield indicates the first operation after the processing unit 1602 last successfully received the first response information. The time 1 subfield indicates the duration of operation 1 or the end time of operation 1. The operation P subfield indicates the Pth operation between the time when the processing unit 1602 last successfully received the first response information and the current time. The time P subfield indicates the duration of operation P or the end time of operation P.
[0296] In yet another optional implementation, the motion information is carried in a motion details field of the first frame reported by the processing unit 1602.
[0297] The action details field includes a time 1 indication subfield, an action 1 subfield, . . . , a time P indication subfield, and an action P subfield, where P is a positive integer.
[0298] The time 1 indication subfield indicates the start time of operation 1. The operation 1 subfield indicates the first operation after the processing unit 1602 successfully receives the first response information last time. The first response information is the response information sent when the AP successfully receives the operation information sent by the STA.
[0299] The time P indication subfield indicates the start time of the action P. The action P subfield indicates the Pth action between the time when the processing unit 1602 last received the first response information successfully and the current time.
[0300] In yet another optional implementation, the motion information is carried in a motion details field of the first frame reported by the processing unit 1602.
[0301] The Action Details field includes a Time 1 Indication subfield, a Duration 1 subfield, . . . , a Time K Indication subfield, and a Duration K subfield, where K is a positive integer.
[0302] The Time 1 Indication subfield indicates the start time / end time of Operation 1. Operation 1 is a transmission operation performed when the STA transmits a packet for the first time after the STA successfully receives the first response information last time and does not receive the second response information. The first response information is response information transmitted when the AP successfully receives the operation information transmitted by the processing unit 1602. The second response information is response information transmitted when the AP successfully receives the packet transmitted by the processing unit 1602. The Duration 1 subfield indicates the duration of Operation 1.
[0303] The Time K Indication subfield indicates the start time / end time of operation K. Operation K is a transmission operation when the processing unit 1602 transmits a packet for the Kth time after the previous successful reception of the first response information and does not receive the second response information. The Duration K subfield indicates the duration of operation K.
[0304] In yet another optional implementation, the motion information is carried in a motion details field of the first frame reported by the processing unit 1602.
[0305] The operation details field includes a first time 1 indication subfield, a second time 1 indication subfield, ..., a first time K indication subfield, and a second time K indication subfield, where K is a positive integer.
[0306] The first time 1 indication subfield indicates the start time of operation 1. Operation 1 is a transmission operation when the processing unit 1602 transmits a packet for the first time after the processing unit 1602 has previously successfully received the first response information and has not received the second response information. The first response information is response information transmitted when the AP has successfully received the operation information transmitted by the processing unit 1602. The second response information is response information transmitted when the AP has successfully received the packet transmitted by the processing unit 1602. The second time 1 indication subfield indicates the end time of operation 1.
[0307] The first time K indication subfield indicates the start time of operation K. Operation K is a transmission operation when the processing unit 1602 transmits a packet for the Kth time after the processing unit 1602 previously successfully receives the first response information and does not receive the second response information. The second time K indication subfield indicates the end time of operation K.
[0308] In a further optional implementation, the operation information and carrier sensing result information are carried in an operation details field of the first frame reported by the processing unit 1602. The operation details field includes a time indication subfield, and Data 1 subfield through Data T subfield, where T is a positive integer.
[0309] The time indication subfield indicates the time when the processing unit 1602 last received the first response information successfully. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the processing unit 1602.
[0310] The Data 1 subfield indicates the carrier sensing result and the action to be taken in the first slot after the processing unit 1602 last successfully received the first response information.
[0311] The data T subfield indicates the carrier sensing result and the action to be taken in the Tth slot after the processing unit 1602 last successfully received the first response information.
[0312] In a further optional embodiment, the operation information and the packet transmission result information are carried in an operation details field of the first frame reported by the processing unit 1602. The operation details field includes a time indication subfield, and data 1 subfield to data T subfield, where T is a positive integer.
[0313] The time indication subfield indicates the time when the processing unit 1602 last received the first response information successfully. The first response information is the response information transmitted when the AP successfully receives the operation information transmitted by the processing unit 1602.
[0314] The Data 1 subfield indicates the packet transmission result and the action to be taken in the first slot after the processing unit 1602 previously successfully received the first response information.
[0315] The data T subfield indicates the packet transmission result and the action to be taken in the Tth slot after the processing unit 1602 last successfully received the first response information.
[0316] In one optional embodiment, when updating the first neural network based on the training result of the first neural network, and determining whether to access the channel based on the updated first neural network and current status information of the processing unit when sensing that the channel is idle, the processing unit 1602 is specifically configured to input the current status information of the processing unit to the updated first neural network, and output a first value and a second value, where the first value represents an expected reward obtained by accessing the channel and the second value represents an expected reward obtained by skipping access to the channel, and if the first value is greater than the second value, decide to access the channel, or if the first value is less than the second value, decide to skip accessing the channel.
[0317] This embodiment of the present application and the above-mentioned method embodiment are based on the same concept and achieve the same technical effect. For the specific principle, please refer to the description of the above-mentioned embodiment. The details will not be described again.
[0318] An embodiment of the present application further provides a communication device 1700. Figure 17 is a schematic diagram of the structure of the communication device 1700. The communication device 1700 may be an AP or a STA, or may be a chip, chip system, processor, etc. that supports an AP in implementing the aforementioned method, or may be a chip, chip system, processor, etc. that supports a STA in implementing the aforementioned method. The device may be configured to implement the method described in the aforementioned method embodiment. For details, please refer to the description of the aforementioned method embodiment.
[0319] The communication device 1700 may include one or more processors 1701. The processor 1701 may be a general-purpose processor, a special-purpose processor, or the like. For example, the processor may be a baseband processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, or a central processing unit (CPU). The baseband processor may be configured to process communication protocols and communication data. The central processing unit may be configured to control the communication device (e.g., a base station, a baseband chip, a terminal, a terminal chip, a DU, or a CU), execute software programs, and process data of the software programs.
[0320] Optionally, the communication device 1700 may include one or more memories 1702. The memory 1702 may store instructions 1704, which may be executed on the processor 1701, causing the communication device 1700 to perform the methods described in the above method embodiments. Optionally, the memory 1702 may further store data. The processor 1701 and the memory 1702 may be located separately or may be integrated with each other.
[0321] The memory 1702 may include, but is not limited to, non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), Random Access Memory (RAM), Erasable Programmable Read-Only Memory (EPROM), Read-Only Memory (ROM), or Portable Read-Only Memory (Compact Disc Read-Only Memory (CD-ROM)).
[0322] Optionally, the communication device 1700 may further include a transceiver 1705 and an antenna 1706. The transceiver 1705 may be referred to as a communication unit, a transceiver machine, a transceiver circuit, etc., and is configured to perform a transceiver function. The transceiver 1705 may include a receiver and a transmitter. The receiver may be referred to as a receiving machine, a receiver circuit, etc., and is configured to perform a receiving function. The transmitter may be referred to as a transmitting machine, a transmitter circuit, etc., and is configured to perform a transmitting function.
[0323] If the communications device 1700 is an AP, the transceiver 1705 is configured to perform S102 and S104 of the channel access method 100, and the processor 1701 is configured to perform S103 of the channel access method 100.
[0324] If the communication device 1700 is a STA, the processor 1701 is configured to perform S106 of the channel access method 100, and the transceiver 1705 is configured to perform S101 and S105 of the channel access method 100.
[0325] In another possible design, the processor 1701 may include a transceiver configured to perform receiving and transmitting functions. For example, the transceiver may be a transceiver circuit, an interface, or an interface circuit. The transceiver circuit, the interface, or the interface circuit configured to perform the receiving and transmitting functions may be separate or integrated with each other. The transceiver circuit, the interface, or the interface circuit may be configured to read and write code / data, or the transceiver circuit, the interface, or the interface circuit may be configured to perform signal transmission or transfer.
[0326] In yet another possible design, optionally, the processor 1701 may store instructions 1703, which run on the processor 1701 to cause the communication device 1700 to perform the methods described in the preceding method embodiments. The instructions 1703 may be fixed within the processor 1701. In this case, the processor 1701 may be implemented by hardware.
[0327] In yet another possible design, the communication device 1700 may include a circuit. The circuit may perform the transmitting, receiving, or communication functions in the method embodiments described above. The processor and transceiver described in this embodiment of the present application may be implemented in an integrated circuit (IC), an analog IC, a radio frequency integrated circuit (RFIC), a hybrid signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver may alternatively be fabricated by using various IC technologies, such as complementary metal oxide semiconductor (CMOS), nMetal-oxide-semiconductor (NMOS), pMetal-oxide-semiconductor (PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), and gallium arsenide (GaAs).
[0328] This embodiment of the present application and the method embodiment shown in the channel access method 100 are based on the same concept and achieve the same technical effect. For the specific principle, please refer to the description of the embodiment shown in the channel access method 100. The details will not be described again.
[0329] The present application further provides a computer-readable storage medium configured to store computer software instructions, which when executed by a communication device, perform the functions of any one of the method embodiments described above.
[0330] The present application further provides a computer program product configured to store computer software instructions which, when executed by a communication device, perform the functions of any one of the method embodiments described above.
[0331] The present application further provides a computer program, which, when run on a computer, performs the functions of any one of the above method embodiments.
[0332] All or part of the above-mentioned embodiments may be implemented by using software, hardware, firmware, or any combination thereof. When software is used to implement the embodiments, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the interactions or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, or digital subscriber line (DSL)) or in a wireless manner (e.g., infrared, radio waves, or microwaves). The computer-readable storage medium may be any available medium accessible by a computer, or may be a data storage device that integrates one or more available media, such as a server or a data center. The available media may be magnetic media (e.g., floppy disks, hard disks, or magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid state disks (SSDs)).
[0333] The above description is merely a specific embodiment of the present application, and is not intended to limit the scope of protection of the present application. Any variations or replacements that are easily conceived by those skilled in the art within the technical scope disclosed in the present application shall be within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims. [Explanation of symbols]
[0334] 100 Channel Access Methods 101 Access Point AP 1021 Access Station STA 1022 Access station STA 1600 Communication Equipment 1601 Communication Unit 1602 Processing Unit 1603 Storage Unit 1700 Communication Equipment 1701 Processor 1702 Memory 1703 Command 1704 Command 1705 Transceiver 1706 Antenna
Claims
1. 1. A channel access method, the method comprising: receiving, by an access point AP, operation information reported separately by N stations STAs, the N operation information being used to determine a training result of a first neural network of each STA, where N is a positive integer; determining, by the AP, the training result of the first neural network of each STA based on the N pieces of operation information; transmitting, by the AP, the training result of the first neural network of each STA to a corresponding STA; Including, A channel access method, wherein the operation information indicates operation for a certain period of time, the certain period being the time between the time when the STA last successfully reported the operation information and the current time, and the operation being an operation of transmitting a packet by the STA or skipping transmission of a packet since the STA last successfully reported the operation information.
2. The method comprises: receiving, by the AP, carrier sensing result information or packet transmission result information reported separately by the N STAs; Further comprising: The step of determining, by the AP, the training result of the first neural network of each STA based on the N pieces of operation information includes: determining, by the AP, the training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of carrier sensing result information; or determining, by the AP, the training result of the first neural network of each STA based on the N pieces of operation information and the N pieces of packet transmission result information; 2. The method of claim 1, comprising:
3. The method of claim 1 , wherein the training results are neural network parameters or gradients, and the neural network parameters / gradients are used by the STA to update the first neural network.
4. The operation information is carried in an operation details field of a first frame reported by the STA; The operation details field includes a time indication subfield and data 1 to data T subfields, where T is a positive integer; The time indication subfield indicates a time when the STA previously received first response information normally, and the first response information is response information transmitted when the AP successfully receives operation information transmitted by the STA; The Data 1 subfield indicates an operation to be performed in a first slot after the STA previously successfully received the first response information, The method of claim 1 , wherein the data T subfield indicates an action to be taken in the Tth slot after the STA last successfully received the first response information.
5. The operation information is carried in an operation details field of a first frame reported by the STA; the action detail field includes a time indication subfield, an action 1 subfield, a time 1 subfield, ..., an action P subfield, and a time P subfield, where P is a positive integer; The time indication subfield indicates a time when the STA previously received first response information normally, and the first response information is response information transmitted when the AP successfully receives operation information transmitted by the STA; The action 1 subfield indicates a first action after the STA has previously received the first response information successfully, and the time 1 subfield indicates a duration of the action 1 or an end time of the action 1; The method of claim 1, wherein the operation P subfield indicates an operation P between the time when the STA last successfully received the first response information and the current time, and the time P subfield indicates the duration of the operation P or the end time of the operation P.
6. The operation information is carried in an operation details field of a first frame reported by the STA; the action details field includes a time 1 indication subfield, an action 1 subfield, ..., a time P indication subfield, and an action P subfield, where P is a positive integer; The time 1 indication subfield indicates a start time of operation 1, the operation 1 subfield indicates a first operation after the STA has previously successfully received first response information, and the first response information is response information transmitted when the AP has successfully received the operation information transmitted by the STA; The method of claim 1, wherein the time P indication subfield indicates a start time of an action P, and the action P subfield indicates an action P between the time when the STA last successfully received the first response information and the current time.
7. The operation information is carried in an operation details field of a first frame reported by the STA; the operation details field includes a time 1 indication subfield, a duration 1 subfield, ..., a time K indication subfield, and a duration K subfield, where K is a positive integer; The time 1 indication subfield indicates the start time / end time of operation 1, the operation 1 being a transmission operation when the STA transmits a packet for the first time after the STA has previously received the first response information successfully and does not receive the second response information, the first response information being response information transmitted when the AP has successfully received the operation information transmitted by the STA, the second response information being response information transmitted when the AP has successfully received the packet transmitted by the STA, the duration 1 subfield indicating the duration of the operation 1, The method of claim 1, wherein the time K indication subfield indicates a start time / end time of operation K, which is a transmission operation when the STA transmits a packet for the Kth time after previously successfully receiving the first response information and does not receive the second response information, and the duration K subfield indicates a duration of the operation K.
8. The operation information is carried in an operation details field of a first frame reported by the STA; the operation details field includes a first time 1 indication subfield, a second time 1 indication subfield, ..., a first time K indication subfield, and a second time K indication subfield, where K is a positive integer; The first time 1 indication subfield indicates a start time of operation 1, the operation 1 being a transmission operation when the STA transmits a packet for the first time after the STA has previously received the first response information successfully and does not receive the second response information, the first response information being response information transmitted when the AP has successfully received the operation information transmitted by the STA, the second response information being response information transmitted when the AP has successfully received the packet transmitted by the STA, the second time 1 indication subfield indicates an end time of the operation 1, The method of claim 1, wherein the first time K indication subfield indicates a start time of operation K, which is a transmission operation when the STA transmits a packet for the Kth time after previously successfully receiving the first response information and does not receive the second response information, and the second time K indication subfield indicates an end time of operation K.
9. the operational information and the carrier sensing result information are carried in an operational details field of a first frame reported by the STA; The operation details field includes a time indication subfield and data 1 to data T subfields, where T is a positive integer; The time indication subfield indicates a time when the STA previously received first response information normally, and the first response information is response information transmitted when the AP successfully receives operation information transmitted by the STA; The Data 1 subfield indicates a carrier sensing result and an operation to be performed in a first slot after the STA previously successfully received the first response information, The method of claim 2 , wherein the Data T subfield indicates a carrier sensing result and an action to be taken in the Tth slot after the STA last successfully received the first response information.
10. The operation information and the packet transmission result information are carried in an operation details field of a first frame reported by the STA; The operation details field includes a time indication subfield and data 1 to data T subfields, where T is a positive integer; The time indication subfield indicates a time when the STA previously received first response information normally, and the first response information is response information transmitted when the AP successfully receives operation information transmitted by the STA; The Data 1 subfield indicates a packet transmission result and an operation to be performed in a first slot after the STA previously successfully received the first response information, The method of claim 2 , wherein the Data T subfield indicates a packet transmission result and an action to be taken in the Tth slot after the STA has previously successfully received the first response information.
11. The method according to claim 1, receiving, by the AP, carrier sensing result information or packet transmission result information reported separately by the N STAs; Further comprising: The step of determining, by the AP, the training result of the first neural network of each STA based on the N pieces of operation information includes: inputting, by the AP, the status information of each STA into the first neural network of the corresponding STA to obtain an output of the first neural network; inputting the output of each first neural network into a second neural network by the AP to obtain an output of the second neural network, the output of the second neural network representing an expected reward within a preset time; training, by the AP, a third neural network based on the output of the second neural network and a reward function, and determining the training result of each first neural network by minimizing a loss function of the third neural network, wherein the third neural network includes each of the first neural network and the second neural network; Including, The status information of the STA is obtained based on the operation information of the STA, and neural network parameters of the second neural network are obtained based on the N pieces of operation information, and the reward function is determined based on the N pieces of operation information; or The status information of the STA is obtained based on the operation information and the carrier sensing result information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of operation information and the N pieces of carrier sensing result information, and the reward function is determined based on the N pieces of operation information and the N pieces of carrier sensing result information; or The method of claim 1, wherein the status information of the STA is obtained based on the operation information and the packet transmission result information of the STA, neural network parameters of the second neural network are obtained based on the N pieces of operation information and the N pieces of packet transmission result information, and the reward function is determined based on the N pieces of operation information and the N pieces of packet transmission result information.
12. The method comprises: setting a value of the reward function by the AP when a first STA determines to successfully transmit a packet based on the N pieces of operation information, the first STA being an STA among the N STAs having a longest time interval between a time when second response information was last successfully received and a current time, the second response information being response information transmitted when the AP successfully receives the packet transmitted by the STA; 12. The method of claim 11, further comprising:
13. The method comprises: When a second STA determines based on the N pieces of motion information that the second STA successfully transmits a packet, setting a value of the reward function to a first duration minus 1 by the AP; The second STA is a STA other than the first STA among the N STAs, the first STA is a STA among the N STAs that has the longest time interval between the time when the second response information was previously received successfully and the current time, and the second response information is response information transmitted when the AP successfully receives the packet transmitted by the STA; the first duration is a duration between a time when the second STA last successfully received the second response information and the current time; 12. The method of claim 11, further comprising:
14. The method comprises: setting a value of the reward function by the AP to −1 when M STAs among the N STAs determine to transmit packets in the same slot based on the N pieces of operation information, where M is a positive integer equal to or less than N; 12. The method of claim 11, further comprising:
15. The method comprises: setting, by the AP, a value of the reward function to 0 when determining, based on the N pieces of behavior information, that none of the N pieces of STAs transmit a packet in the same slot; 12. The method of claim 11, further comprising:
16. The N STAs share neural network parameters, and the AP transmits the training result of the first neural network of each STA to the corresponding STA, broadcasting, by the AP, the training result of the first neural network to the N STAs; 2. The method of claim 1, comprising:
17. S STAs among the N STAs share neural network parameters, where S is a positive integer less than or equal to N, and the step of transmitting the training result of the first neural network of each STA to the corresponding STA by the AP includes: multicasting, by the AP, the training results of the first neural networks corresponding to the S STAs to the S STAs, and unicasting the training results of the (N-S) first neural networks to the corresponding STAs.
2. The method of claim 1, comprising:
18. When the N STAs do not share neural network parameters, the training result of each first neural network is unicast to the corresponding STA. The method of claim 1.
19. 1. A channel access method, comprising: reporting operational information by a station STA to an access point AP, the operational information being used to determine a training result of a first neural network of the station; receiving, by the STA, the training result of the first neural network from the AP, the training result of the first neural network being used to update the first neural network to determine whether the STA should access a channel; updating, by the STA, the first neural network based on the training result of the first neural network, and determining whether to access the channel based on the updated first neural network and current status information of the STA when sensing that the channel is idle; Including, A channel access method, wherein the operation information indicates operation for a certain period of time, the certain period being the time between the time when the STA last successfully reported the operation information and the current time, and the operation being an operation of transmitting a packet by the STA or skipping transmission of a packet since the STA last successfully reported the operation information.
20. The method comprises: reporting carrier sensing result information or packet transmission result information by the STA to the AP, the carrier sensing result information or the packet transmission result information being used to determine the training result of the first neural network of the STA; 20. The method of claim 19, further comprising:
21. the training results are neural network parameters or gradients; 20. The method of claim 19, wherein the neural network parameters / gradients are used by the STA to update the first neural network.
22. The operation information is carried in an operation details field of a first frame reported by the STA; The operation details field includes a time indication subfield and data 1 to data T subfields, where T is a positive integer; The time indication subfield indicates a time when the STA previously received first response information normally, and the first response information is response information transmitted when the AP successfully receives operation information transmitted by the STA; The Data 1 subfield indicates an operation to be performed in a first slot after the STA previously successfully received the first response information, 20. The method of claim 19, wherein the Data T subfield indicates an action to be taken in the Tth slot after the STA last successfully received the first response information.
23. The operation information is carried in an operation details field of a first frame reported by the STA; the action detail field includes a time indication subfield, an action 1 subfield, a time 1 subfield, ..., an action P subfield, and a time P subfield, where P is a positive integer; The time indication subfield indicates a time when the STA previously received first response information normally, and the first response information is response information transmitted when the AP successfully receives operation information transmitted by the STA; The action 1 subfield indicates a first action after the STA has previously received the first response information successfully, and the time 1 subfield indicates a duration of the action 1 or an end time of the action 1; The method of claim 19, wherein the operation P subfield indicates an operation P between the time the STA last successfully received the first response information and the current time, and the time P subfield indicates a duration of the operation P or an end time of the operation P.
24. The operation information is carried in an operation details field of a first frame reported by the STA; the action details field includes a time 1 indication subfield, an action 1 subfield, ..., a time P indication subfield, and an action P subfield, where P is a positive integer; The time 1 indication subfield indicates a start time of operation 1, the operation 1 subfield indicates a first operation after the STA has previously successfully received first response information, and the first response information is response information transmitted when the AP has successfully received the operation information transmitted by the STA; The method of claim 19, wherein the time P indication subfield indicates a start time of an action P, and the action P subfield indicates an action P between the time when the STA last successfully received the first response information and the current time.
25. The operation information is carried in an operation details field of a first frame reported by the STA; the operation details field includes a time 1 indication subfield, a duration 1 subfield, ..., a time K indication subfield, and a duration K subfield, where K is a positive integer; The time 1 indication subfield indicates the start time / end time of operation 1, the operation 1 being a transmission operation when the STA transmits a packet for the first time after the STA has previously received the first response information successfully and does not receive the second response information, the first response information being response information transmitted when the AP has successfully received the operation information transmitted by the STA, the second response information being response information transmitted when the AP has successfully received the packet transmitted by the STA, the duration 1 subfield indicating the duration of the operation 1, The method of claim 19, wherein the time K indication subfield indicates a start time / end time of operation K, the operation K being a transmission operation when the STA transmits a packet for the Kth time after previously successfully receiving the first response information and does not receive the second response information, and the duration K subfield indicates a duration of the operation K.
26. The operation information is carried in an operation details field of a first frame reported by the STA; the operation details field includes a first time 1 indication subfield, a second time 1 indication subfield, ..., a first time K indication subfield, and a second time K indication subfield, where K is a positive integer; The first time 1 indication subfield indicates a start time of operation 1, the operation 1 being a transmission operation when the STA transmits a packet for the first time after the STA has previously received the first response information successfully and does not receive the second response information, the first response information being response information transmitted when the AP has successfully received the operation information transmitted by the STA, the second response information being response information transmitted when the AP has successfully received the packet transmitted by the STA, the second time 1 indication subfield indicates an end time of the operation 1, The method of claim 19, wherein the first time K indication subfield indicates a start time of operation K, which is a transmission operation when the STA transmits a packet for the Kth time after previously successfully receiving the first response information and does not receive the second response information, and the second time K indication subfield indicates an end time of operation K.
27. the operational information and the carrier sensing result information are carried in an operational details field of a first frame reported by the STA; The operation details field includes a time indication subfield and data 1 to data T subfields, where T is a positive integer; The time indication subfield indicates a time when the STA previously received first response information normally, and the first response information is response information transmitted when the AP successfully receives operation information transmitted by the STA; The Data 1 subfield indicates a carrier sensing result and an operation to be performed in a first slot after the STA previously successfully received the first response information, 21. The method of claim 20, wherein the Data T subfield indicates a carrier sensing result and an action to be taken in the Tth slot after the STA last successfully received the first response information.
28. The operation information and the packet transmission result information are carried in an operation details field of a first frame reported by the STA; The operation details field includes a time indication subfield and data 1 to data T subfields, where T is a positive integer; The time indication subfield indicates a time when the STA previously received first response information normally, and the first response information is response information transmitted when the AP successfully receives operation information transmitted by the STA; The Data 1 subfield indicates a packet transmission result and an operation to be performed in a first slot after the STA previously successfully received the first response information, The method of claim 20, wherein the Data T subfield indicates a packet transmission result and an action to be taken in the Tth slot after the STA last successfully received the first response information.
29. The step of updating the first neural network by the STA based on the training result of the first neural network, and determining whether to access the channel based on the updated first neural network and current status information of the STA when the STA detects that the channel is idle, includes: inputting, by the STA, the current status information of the STA into the updated first neural network to output a first value and a second value, the first value representing an expected reward obtained by accessing the channel, and the second value representing an expected reward obtained by skipping access to the channel; determining, by the STA, to access the channel if the first value is greater than the second value; or determining, by the STA, to skip access to the channel if the first value is less than the second value; 20. The method of claim 19, comprising:
30. A communication device configured to perform a method according to any one of claims 1 to 18 or any one of claims 19 to 29.
31. A communications device comprising a processor and a transceiver, the transceiver configured to communicate with other communications devices, the processor configured to execute a program such that the communications device performs a method according to any one of claims 1 to 18 or such that the communications device performs a method according to any one of claims 19 to 29.
32. 30. A computer readable storage medium having instructions stored thereon which, when executed on a computer, perform the method of any one of claims 1 to 18 or perform the method of any one of claims 19 to 29.
33. 30. A computer program comprising instructions which, when said computer program is run on a computer, causes the method of any one of claims 1 to 18 to be performed or the method of any one of claims 19 to 29 to be performed.
Citation Information
Patent Citations
WSN anti-interference method and device based on reinforcement learning, equipment and medium
CN109861720A
Unmanned aerial vehicle CSMA access method based on adaptive adjustment strategy
CN111050413A
Method for allocating resource using machine learning in a wireless network and recording medium for performing the method
KR1020200081630A
WLAN service method and WLAN system
US20140119291A1
Wireless contention reduction
US20190021114A1