A Dual-Band WLAN Optimization Scheme Based on Reinforcement Learning

By employing a dual-band WLAN optimization scheme based on reinforcement learning, the power configuration of 2.4G and 5G radio frequencies is collected and optimized, solving the problem of uneven user experience in large-scale networking and achieving better network performance and user experience.

CN114501494BActive Publication Date: 2026-04-03FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing WLAN optimization solutions lack a joint power optimization scheme for 2.4G and 5G radio frequencies in large-scale networking, resulting in an uneven user experience.

Method used

A dual-band WLAN optimization scheme based on reinforcement learning is adopted. The AP module collects radio frequency information, and the AI ​​optimization module performs data processing and reinforcement learning decision-making to optimize the power configuration of 2.4G and 5G radio frequencies and achieve collaborative optimization.

Benefits of technology

It effectively improves the balance of user experience between 2.4G and 5G radio frequencies, enhances the overall network performance and user experience, and avoids the limitations of traditional optimization schemes that only focus on coverage and signal strength.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114501494B_ABST
    Figure CN114501494B_ABST
Patent Text Reader

Abstract

This invention provides a dual-band WLAN optimization scheme based on reinforcement learning, relating to the field of wireless local area network (WLAN) technology. The scheme includes the following steps: S1. Data reporting: Relevant radio frequency (RF) information is collected by the AP module and reported to the data acquisition module of the AI ​​optimization module. This information is divided into 2.4G RF information and 5G RF information from 2.4G and 5G terminals, respectively. S2. Information processing: After receiving the complete RF information, the data acquisition module of the AI ​​optimization module processes it into a unified format and sends it to the benefit evaluation module of the AI ​​optimization module. This unified format can be the collected AP reporting information from the entire network. This invention provides a dual-band WLAN optimization scheme based on reinforcement learning. This scheme can optimize the network through reinforcement learning and user experience feedback, and also enables coordinated optimization of 2.4G and 5G RF, thereby achieving the effect that the more times the network is optimized, the better the overall user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless local area network (WLAN) technology, specifically to a dual-band WLAN optimization scheme based on reinforcement learning. Background Technology

[0002] WLAN technology allows home and business users to conveniently and flexibly access the network and provides high-speed upload / download speeds. This feature has brought WLAN into thousands of households. There are two definitions of WLAN: a broad one and a narrow one. In a broad sense, WLAN is a network that uses wireless channels of various radio waves (such as lasers and infrared rays) to replace part or all of the transmission medium of wired local area networks. In a narrow sense, WLAN is a wireless local area network based on the IEEE 802.11 series of standards, using high-frequency radio frequency (such as radio electromagnetic waves in the 2.4GHz or 5GHz band) as the transmission medium. There are two basic architectures of WLAN: one is the FTAAP architecture, also called autonomous network architecture, and the other is the AC+FITAP architecture, also called centralized network architecture. The home wireless routers we are most familiar with use the FTAAP architecture, which is also called fat AP by many people. FITAP stands for FIT Access Point, which is also called thin access point by many people.

[0003] Normally, an AP will simultaneously activate both the 2.4GHz and 5GHz radio frequencies. The 2.4GHz signal has a narrower bandwidth and fewer selectable channels, resulting in a more congested wireless environment with greater interference. The 5GHz signal has a wider bandwidth, a larger selectable channel range, a cleaner wireless environment with less interference, more stable network speed, and higher wireless rates. However, the higher frequency of 5GHz leads to greater signal attenuation during propagation, resulting in significantly less wall penetration and a shorter propagation distance compared to 2.4GHz. Load balancing technology can only achieve load balancing between different APs on the same radio frequency, such as load balancing of the 5GHz radio frequencies of AP1 and AP2, but it cannot achieve load balancing between the 5GHz radio frequency of AP1 and the 2.4GHz radio frequency of AP1. Even if it could, it would not be suitable for large-scale AC+FITAP networking. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] To address the shortcomings of existing technologies, this invention provides a dual-band WLAN optimization scheme based on reinforcement learning, which solves the problem that industry WLAN optimization schemes lack a joint power optimization scheme for 2.4G and 5G in large-scale networking.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the present invention provides the following technical solution: a dual-band WLAN optimization scheme based on reinforcement learning, comprising the following steps:

[0008] S1. Data Reporting

[0009] The data acquisition module collects relevant radio frequency information through the AP module and reports it to the AI ​​optimization module. This information is divided into 2.4G radio frequency information and 5G radio frequency information by the 2.4G terminal and the 5G terminal.

[0010] S2. Information Organization

[0011] After receiving complete radio frequency information, the data acquisition module of the AI ​​tuning module organizes it into a unified format and sends it to the benefit evaluation module of the AI ​​tuning module. This unified format can be the collection of AP reporting information from the entire network and then sending it at a fixed time frequency.

[0012] S3. Organize and transmit

[0013] The AI ​​optimization module's benefit evaluation module organizes the received data into state and reward and then sends it to the AI ​​optimization module's reinforcement learning decision-making module.

[0014] S4. Data Processing

[0015] After receiving the reward and state, the reinforcement learning decision module of the AI ​​tuning module will perform two processes: a. optimize its own decision model based on the reward, b. output the action based on the state, including the power of each 2.4G radio frequency and the power of each 5G radio frequency;

[0016] S5. Final Processing

[0017] After receiving the action from the reinforcement learning decision module of the AI ​​tuning module, the AC module sends a power configuration command to the 2.4G and 5G radio frequencies of each AP. After receiving the power configuration, the AP module changes the power.

[0018] Preferably, the 2.4G radio frequency information in S1 specifically refers to the number of 2.4G radio frequencies in the entire network, and the power, channel, bandwidth, throughput, channel utilization, co-channel interference rate, number of connected terminals, average downlink RSSI of terminals, average uplink RSSI of terminals, average throughput of terminals, average connection rate of terminals, average latency of terminals, and average packet loss rate of terminals.

[0019] Preferably, the 5G radio frequency information in S1 specifically includes the number of 5G radio frequencies in the entire network, and the power, channel, bandwidth, throughput, channel utilization, co-channel interference rate, number of connected terminals, average downlink RSSI of the terminal, average uplink RSSI of the terminal, average throughput of the terminal, average connection rate of the terminal, average latency of the terminal, and average packet loss rate of the terminal.

[0020] Preferably, in S3, state specifically refers to the number of 2.4G radios in the entire network, the power, channel, bandwidth, throughput, number of connected terminals, and average terminal throughput of each 2.4G radio, the number of 5G radios in the entire network, and the power, channel, bandwidth, throughput, number of connected terminals, and average terminal throughput of each 5G radio.

[0021] Preferably, the reward in S3 specifically includes the channel utilization, co-channel interference rate, average downlink RSSI, average uplink RSSI, average connection rate, average latency, and average packet loss rate of each 2.4G radio frequency, and the channel utilization, co-channel interference rate, number of connected terminals, average downlink RSSI, average uplink RSSI, average connection rate, average latency, and average packet loss rate of each 5G radio frequency.

[0022] Preferably, the data acquisition module in S1 is used to collect various radio frequency information, including the AP module's power, bandwidth, channel, throughput, channel utilization, co-channel interference rate, number of 2.4G radio frequency connected terminals, number of 5G radio frequency connected terminals, terminal speed, packet loss rate, and latency.

[0023] Preferably, the reinforcement learning decision module in S3 is the core module of the AI ​​tuning module. It contains a decision model built based on historical tuning data. It reports the current state based on the benefit evaluation module and gives an action that tends to tune the network better, as well as a recommended power configuration for each radio frequency. At the same time, it also optimizes its decision model based on whether the reward reported by the benefit evaluation module is positive or negative, so that the network gets better and better with each jump.

[0024] Preferably, the steps can be performed cyclically, and the cycle interval can be adjusted as needed.

[0025] (III) Beneficial Effects

[0026] This invention provides a dual-band WLAN optimization scheme based on reinforcement learning. It has the following beneficial effects:

[0027] 1. This invention provides a dual-band WLAN optimization scheme based on reinforcement learning. This invention simultaneously collects the configuration and user experience parameters of 2.4G radio frequency and 5G radio frequency, and makes decisions and optimizes the model based on these two aspects of information in the reinforcement learning decision module. This can effectively avoid the imbalance between good user experience of 2.4G or 5G radio frequency and poor user experience of 5G or 2.4G radio frequency, and ensure that the experience of both is improved.

[0028] 2. This invention provides a dual-band WLAN optimization scheme based on reinforcement learning. This invention collects rich user experience parameters and then uses reinforcement learning for optimization. The final result is to bring a better user experience, rather than the traditional optimization scheme that only focuses on the network layer indicators of coverage and signal strength. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the process of the present invention;

[0030] Figure 2 This is a schematic diagram of the AI ​​optimization module process of the present invention.

[0031] The modules include: 1. AI optimization module; 2. AP module; 3. 2.4G terminal; 4. 5G terminal; 5. AC module; 6. Data acquisition module; 7. Benefit evaluation module; and 8. Reinforcement learning decision-making module. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Example:

[0034] like Figure 1-2 As shown, this embodiment of the invention provides a dual-band WLAN optimization scheme based on reinforcement learning, including the following steps:

[0035] S1. Data Reporting

[0036] The relevant radio frequency information is collected by AP module 2 and reported to data acquisition module 6 of AI tuning module 1. This information is divided into 2.4G radio frequency information and 5G radio frequency information by 2.4G terminal 3 and 5G terminal 4.

[0037] S2. Information Organization

[0038] After receiving complete radio frequency information, the data acquisition module 6 of AI tuning module 1 organizes it into a unified format and sends it to the benefit evaluation module 7 of AI tuning module 1. This unified format can be the collection of AP reporting information from the entire network and then sending it at a fixed time frequency.

[0039] S3. Organize and transmit

[0040] After the AI ​​tuning module 1's benefit evaluation module 7 organizes the received data into state and reward, it sends it to the reinforcement learning decision module 8 of the AI ​​tuning module 1.

[0041] S4. Data Processing

[0042] After receiving the reward and state, the reinforcement learning decision module 8 of AI tuning module 1 will perform two processes: a. optimize its own decision model based on the reward, b. output the action based on the state, including the power of each 2.4G radio frequency and the power of each 5G radio frequency;

[0043] S5. Final Processing

[0044] After receiving the action from the reinforcement learning decision module 8 of the AI ​​tuning module 1, the AC module 5 sends power configuration commands to the 2.4G and 5G radio frequencies of each AP. After receiving the power configuration, the AP module 2 changes the power. By combining the above steps, the 2.4G and 5G radio frequencies can be coordinated and optimized. At the same time, the 2.4G and 5G radio frequency information is collected, and the reinforcement learning decision module 8 uses the two types of information to generate the power configuration of 2.4G and 5G.

[0045] The 2.4G radio frequency information in S1 specifically includes the number of 2.4G radio frequencies in the entire network, and the power, channel, bandwidth, throughput, channel utilization, co-channel interference rate, number of connected terminals, average downlink RSSI, average uplink RSSI, average throughput, average connection rate, average latency, and average packet loss rate of each 2.4G radio frequency.

[0046] The 5G radio frequency information in S1 specifically includes the number of 5G radio frequencies in the entire network, and the power, channel, bandwidth, throughput, channel utilization, co-channel interference rate, number of connected terminals, average downlink RSSI, average uplink RSSI, average throughput, average connection rate, average latency, and average packet loss rate of each 5G radio frequency.

[0047] In S3, "state" specifically refers to the number of 2.4G radios in the entire network, and the power, channel, bandwidth, throughput, number of connected terminals, and average terminal throughput of each 2.4G radio. It also refers to the number of 5G radios in the entire network, and the power, channel, bandwidth, throughput, number of connected terminals, and average terminal throughput of each 5G radio.

[0048] In S3, the reward specifically includes the channel utilization, co-channel interference rate, average downlink RSSI, average uplink RSSI, average connection rate, average latency, and average packet loss rate for each 2.4G radio frequency, and the channel utilization, co-channel interference rate, number of connected terminals, average downlink RSSI, average uplink RSSI, average connection rate, average latency, and average packet loss rate for each 5G radio frequency. The reward calculation formula is based on the various reward metrics in this study. The reward is calculated by summing the gains of the parameters relative to the previous reward parameter. For example, if the current time is t, then reward = f(U(2.4G RF channel utilization t – 2.4G RF channel utilization t-1) + U(2.4G RF co-channel interference rate t – 2.4G RF co-channel interference rate t-1) + … + U(2.4G terminal average packet loss rate t – 2.4G terminal average packet loss rate t-1) + U(5G RF channel utilization t – 5G RF channel utilization t-1) + … + U(5G terminal average packet loss rate t – 5G terminal average packet loss rate t- ... 4. Average packet loss rate t-1), where f represents the gain of each parameter multiplied by a fixed coefficient, such as f = 0.1xU1 + 0.1xU2 + 0.05xU3 + ... The specific coefficient to be multiplied for each parameter gain can be determined by the network administrator. For example, if a WLAN network prioritizes 5G user speed, then U(average connection rate of 5G RF terminals t – average connection rate of 5G RF terminals t-1) can be multiplied by a larger coefficient. U represents the normalization of the gain of each parameter. Normalization means normalizing the parameters... The gain becomes a value between 0 and 1, for example, Yt = average connection rate of 5G RF terminal t – average connection rate of 5G RF terminal t-1, U(Yt) = (Yt – min(5G RF terminal connection rate)) / max(5G RF terminal connection rate) - min(5G RF terminal connection rate)), min(5G RF terminal connection rate): the minimum 5G terminal connection rate that may occur in the network, max(5G RF terminal connection rate): the maximum 5G terminal connection rate that may occur in the network.

[0049] In S1, the data acquisition module 6 is used to collect various radio frequency information, including the power, bandwidth, channel, throughput, channel utilization, co-channel interference rate of AP module 2, the number of 2.4G radio frequency connected terminals, the number of 5G radio frequency connected terminals, and the terminal's speed, packet loss rate, and latency.

[0050] The reinforcement learning decision module 8 in S3 is the core module of the AI ​​tuning module 1. It contains a decision model built based on historical tuning data. It reports the current state from the reward evaluation module 7 and provides an action that tends to improve the network, as well as a recommended power configuration for each radio frequency. At the same time, it optimizes its decision model based on whether the reward reported by the reward evaluation module 7 is positive or negative, so that the network can improve with each iteration. This process can be repeated cyclically, and the interval between iterations can be adjusted as needed. The reward calculation formula of the reward evaluation module 7 can also be calculated based on the current reward parameter value instead of the difference in reward parameters.

[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A dual-band WLAN optimization scheme based on reinforcement learning, characterized in that, Includes the following steps: S1. Data Reporting The relevant radio frequency information is collected by the AP module (2) and reported to the data acquisition module (6) of the AI ​​tuning module (1). This information is divided into 2.4G radio frequency information and 5G radio frequency information by the 2.4G terminal (3) and the 5G terminal (4). The 2.4G radio frequency information specifically includes the number of 2.4G radio frequencies in the entire network, and the power, channel, bandwidth, throughput, channel utilization, co-channel interference rate, number of connected terminals, average downlink RSSI, average uplink RSSI, average throughput, average connection rate, average latency, and average packet loss rate for each 2.4G radio frequency. The 5G radio frequency information specifically includes the number of 5G radio frequencies in the entire network, and the power, channel, bandwidth, throughput, channel utilization, co-channel interference rate, number of connected terminals, average downlink RSSI, average uplink RSSI, average throughput, average connection rate, average latency, and average packet loss rate for each 5G radio frequency. S2. Information Organization After receiving the complete radio frequency information, the data acquisition module (6) of the AI ​​tuning module (1) organizes it into a unified format and sends it to the benefit evaluation module (7) of the AI ​​tuning module (1). This unified format can be the collection of AP reporting information from the entire network and then sent at a fixed time frequency. S3. Organize and transmit The AI ​​tuning module (1) and the benefit evaluation module (7) organize the received data into state and reward and send it to the reinforcement learning decision module (8) of the AI ​​tuning module (1). The state specifically refers to the number of 2.4G radios in the entire network, and the power, channel, bandwidth, throughput, number of connected terminals, and average terminal throughput of each 2.4G radio. The state also refers to the number of 5G radios in the entire network, and the power, channel, bandwidth, throughput, number of connected terminals, and average terminal throughput of each 5G radio. The reward specifically refers to the channel utilization, co-channel interference rate, average downlink RSSI, average uplink RSSI, average connection rate, average latency, and average packet loss rate of each 2.4G radio. The reinforcement learning decision module (8) in S3 is the core module of the AI ​​tuning module (1). It has a decision model built based on historical tuning data. It will report the current state based on the benefit evaluation module (7) and give an action that tends to tune the network better, the recommended power configuration of each radio frequency. At the same time, it will also optimize its decision model based on whether the reward reported by the benefit evaluation module (7) is positive or negative, so that the network can improve as it jumps. The reward calculation formula is based on the sum of the gains of each reward parameter in the current period relative to the previous reward parameter: The current time is t, reward = f(U(2.4G radio frequency channel utilization) t -2.4G radio frequency channel utilization t−1 )+U(2.4G radio frequency co-channel interference rate t -2.4G radio frequency co-channel interference rate t−1 )+U(2.4G terminal average downlink RSSI t -2.4G terminal average downlink RSSI t−1 )+U(2.4G terminal average uplink RSSI t -2.4G terminal average uplink RSSI t−1 )+U(2.4G terminal average connection speed t -2.4G average connection speed t−1 )+U(2.4G terminal average latency t -2.4G terminal average latency t−1 )+U(2.4G terminal average packet loss rate) t -2.4G terminal average packet loss rate t−1 )+U(5G radio frequency channel utilization) t -5G radio frequency channel utilization t−1 )+U(5G radio frequency co-channel interference rate) t -5G radio frequency co-frequency interference rate t−1 )+U(5G terminal average downlink RSSI t - Average downlink RSS of 5G terminals It−1 )+U(5G terminal average uplink RSSI) t - Average uplink RSSI of 5G terminals t−1 )+U(Average connection speed of 5G terminals) t - Average connection speed of 5G terminals t−1 )+U(5G terminal average latency) t - Average latency of 5G terminals t−1 )+U(5G terminal average packet loss rate) t - Average packet loss rate of 5G terminals t−1 ), where f represents the gain of each parameter multiplied by a fixed coefficient; The size of the coefficient by which the gain of each parameter needs to be multiplied is determined by the network administrator's configuration: When the WLAN network prioritizes 5G user speeds, U(5G RF terminal average connection rate t - 5G RF terminal average connection rate t-1) is multiplied by a larger coefficient. U represents the normalization of the gain for each parameter, which means converting the parameter gain to a value between 0 and 1. When Yt = 5G RF terminal average connection rate t – 5G RF terminal average connection rate t1, U(Yt) = (Yt - min(5G RF terminal connection rate)) / max(5G RF terminal connection rate) min(5G RF terminal connection rate)); where min(5G RF terminal connection rate) is the minimum 5G terminal connection rate in the network, and max(5G RF terminal connection rate) is the maximum 5G terminal connection rate in the network. S4. Data Processing After receiving the reward and state, the reinforcement learning decision module (8) of the AI ​​tuning module (1) will perform two processes: a. optimize its own decision model according to the reward, b. output the action according to the state, including the power of each 2.4G radio frequency and the power of each 5G radio frequency; S5. Final Processing After receiving the action from the reinforcement learning decision module (8) of the AI ​​tuning module (1), the AC module (5) sends a power configuration command to the 2.4G and 5G radio frequencies of each AP. After receiving the power configuration, the AP module (2) changes the power.

2. The dual-band WLAN optimization scheme based on reinforcement learning according to claim 1, characterized in that: The data acquisition module (6) in S1 is used to collect various radio frequency information, including the power, bandwidth, channel, throughput, channel utilization, co-channel interference rate of the AP module (2), the number of 2.4G radio frequency connected terminals, the number of 5G radio frequency connected terminals, and the terminal's rate, packet loss rate, and latency.

3. The dual-band WLAN optimization scheme based on reinforcement learning according to claim 1, characterized in that: The steps described can be repeated cyclically, and the interval between cycles can be adjusted as needed.