Multi-agent reinforcement learning-based WiFi (Wireless Fidelity) 7 network access point topology adjustment method and system, terminal equipment and medium
The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning dynamically adjusts the access point topology, solving the problems of load imbalance, resource waste, and high transmission latency in WiFi networks, and achieving more efficient resource utilization and energy management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 深圳开鸿数字产业发展有限公司
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-01
AI Technical Summary
Existing WiFi networks are inadequate in terms of load balancing, energy efficiency, resource sharing, and adaptability, and cannot adapt to dynamic fluctuations in user density, resulting in problems such as load imbalance, resource waste, and high transmission latency.
A WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning is adopted. By deploying a pre-trained execution network locally at the access point, topology adjustment actions are generated and verified, including merging, splitting, load sharing and multi-link operations, to dynamically respond to network state changes and optimize resource utilization and coverage balance.
It enables access points to have autonomous adaptive capabilities, reducing transmission latency, improving resource utilization and coverage balance, reducing energy consumption, and adapting to the dynamic environmental requirements of WiFi 7 networks.
Smart Images

Figure CN121968151A_ABST
Abstract
Description
A method, system, terminal device, and medium for WiFi 7 network access point topology adjustment based on multi-agent reinforcement learning. Technical Field
[0001] This invention relates to the field of wireless network configuration technology, and in particular to a method, system, terminal device, and medium for adjusting the topology of WiFi 7 network access points based on multi-agent reinforcement learning. Background Technology
[0002] With the explosion of data-intensive applications such as virtual reality and cloud gaming, WiFi networks need to meet the requirements of ultra-high peak throughput and ultra-low latency. WiFi 7 was born to meet this need. It integrates core technologies such as 4096-QAM modulation, multi-link operation (MLO), and preamble punching, providing the hardware foundation for the next generation of wireless communication.
[0003] However, existing WiFi network control strategies still have key limitations. First, their load balancing is inefficient, with some access points overloaded while neighboring devices remain idle during peak hours. Second, their energy utilization is low, with idle access points continuously operating at high power consumption. Furthermore, their static cell configuration cannot adapt to dynamic fluctuations in user density. Simultaneously, their resource sharing capabilities between neighboring access points are limited, resulting in insufficient spectrum utilization. Additionally, they lack autonomous adaptive capabilities, relying on manual or centralized scheduling, leading to lag in response. Even with WiFi 7's advanced hardware features, traditional static control logic still struggles to unleash its potential and cannot meet dynamic adaptation requirements.
[0004] Therefore, there is an urgent need for a dynamic WiFi 7 network access point cell configuration method to fill the gap in existing technology. Summary of the Invention
[0005] The technical problem this invention aims to solve is that, in the field of wireless network configuration technology, existing static access point configuration and inefficient scheduling lead to problems such as load imbalance, resource waste, high transmission latency, and low energy efficiency, which cannot adapt to dynamic network environments and the potential of WiFi 7 features. Therefore, an effective solution is urgently needed to address these technical problems.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: Firstly, the present invention provides a WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning, applied to WiFi 7 network access points. The method includes: acquiring network status; generating a topology adjustment action to be executed based on the network status using a pre-trained execution network; wherein the pre-trained execution network is locally deployed at the access point, and the pre-trained execution network is trained through multi-agent reinforcement learning; verifying the topology adjustment action to be executed, and if the verification passes, executing the topology adjustment action; and acquiring the network status after executing the topology adjustment action.
[0007] In one implementation, obtaining the network status includes: the access point collecting its own operational data and acquiring operational data of neighboring access points through communication links; integrating its own operational data with the operational data of neighboring access points to obtain an initial network status; filtering and normalizing the initial network status to form a final network status; wherein the network status includes the number of users connected to the access point, current traffic load, remaining energy, average load of neighboring access points, normalized effective bandwidth, multi-link operation enabled status, current cell size, frequency band used, and time normalization factor.
[0008] In one implementation, the topology adjustment action is an access point merging action, which includes: after the local access point and the adjacent access point cooperate, the local access point goes into hibernation and transfers the user.
[0009] In one implementation, the verification of the topology adjustment action to be executed includes: based on the network state, obtaining the current traffic load of the local access point, the average load of adjacent access points, and the remaining energy of the local access point; comparing the current traffic load of the local access point with a preset low load threshold, comparing the average load of adjacent access points with a preset coordination threshold, and comparing the remaining energy of the local access point with a preset energy-saving threshold; if all three comparison results satisfy the less than or equal to relationship, then the verification result of the merging action is obtained.
[0010] In one implementation, the execution of the topology adjustment action to be performed includes: the local access point sending a coordination request to the adjacent access point through multi-link operation and receiving a resource idle response returned by the adjacent access point; generating a user transfer list based on user QoS priority, and smoothly switching the users in the list to the adjacent access point through the links of the multi-link operation; after all users have been transferred, the local access point shuts down the signal transceiver module and enters a sleep mode.
[0011] In one implementation, the topology adjustment action is an access point splitting action, which includes: activating the newly added access point and partitioning the load and coverage of the original access point.
[0012] In one implementation, the verification of the topology adjustment action to be executed includes: based on the network status, obtaining the current traffic load, user density, and coverage overlap rate of the local access point; comparing the current traffic load of the local access point with a preset high load threshold, comparing the user density with a preset density threshold, and comparing the coverage overlap rate with a preset overlap threshold; if the load comparison satisfies a greater than or equal to relationship, or the density comparison satisfies a greater than or equal to relationship, and the overlap rate comparison satisfies a less than or equal to relationship, then the verification result of the splitting action is obtained.
[0013] In one implementation, the execution of the topology adjustment action to be performed includes: the controller receiving a split request from the original access point and activating a preset new access point; the original access point and the new access point dividing the coverage area through frequency band negotiation; the load partitioning of the original access point to the new access point based on user distribution data and QoS requirements; and enabling preamble punching technology to scan and avoid spectrum interference, thereby completing the split configuration of load and coverage area.
[0014] In one implementation, the WiFi 7 network includes several access points and a controller. After obtaining the network status after performing the topology adjustment action, the method further includes: each access point reporting the performed topology adjustment action and the network status after performing the topology adjustment action to the controller; the controller obtaining the global network status; and the controller optimizing the execution network of each access point based on the topology adjustment action performed by each access point, the network status after each access point performs the topology adjustment action, and the global network status.
[0015] In one implementation, the controller optimizes the execution network of each access point based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state. This includes optimizing the execution network of each access point through value evaluation using multi-agent reinforcement learning, based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state. The multi-agent reinforcement learning includes a collaboration between centralized evaluation and decentralized execution. The centralized evaluation generates a global value evaluation, and the local execution network of each access point implements the decentralized execution.
[0016] In one implementation, the centralized evaluation includes: deploying a centralized evaluation network based on a dynamic cell multi-agent reinforcement learning framework; and calculating the global value assessment of each action by inputting the global network state and the actions of each access point through the evaluation network.
[0017] In one implementation, the optimization of the execution network at each access point based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state, through value evaluation using multi-agent reinforcement learning, includes: global value evaluation based on the output of a centralized evaluation network; updating the parameters of the execution network through temporal differential error calculation; the controller sending the updated execution network parameters to the corresponding access points; and the access points receiving and updating the parameters of their local execution networks.
[0018] Secondly, embodiments of the present invention also provide a WiFi 7 network access point topology adjustment system based on multi-agent reinforcement learning. The system includes: a network state acquisition module for acquiring network state; an action generation module for generating a topology adjustment action to be executed based on the network state using a pre-trained execution network; wherein the pre-trained execution network is locally deployed at the access point and is trained through multi-agent reinforcement learning; an action verification execution module for verifying the topology adjustment action to be executed, and executing the topology adjustment action if the verification passes; and a post-action network state acquisition module for acquiring the network state after executing the topology adjustment action.
[0019] Thirdly, embodiments of the present invention also provide a terminal device, the terminal device including a memory, a processor, and a WiFi 7 network access point topology adjustment program based on multi-agent reinforcement learning stored in the memory and executable on the processor. When the processor executes the WiFi 7 network access point topology adjustment program based on multi-agent reinforcement learning, it implements the steps of the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning as described in any of the above schemes.
[0020] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a WiFi 7 network access point topology adjustment program based on multi-agent reinforcement learning. When the WiFi 7 network access point topology adjustment program based on multi-agent reinforcement learning is executed by a processor, it implements the steps of the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning as described in any of the above schemes.
[0021] Beneficial Effects: This invention discloses a method, system, terminal device, and medium for WiFi 7 network access point topology adjustment based on multi-agent reinforcement learning, relating to the field of wireless network configuration technology and applied to WiFi 7 network access points. The method first acquires the network state and, based on this state, generates a topology adjustment action to be executed through a pre-trained execution network. The pre-trained execution network is locally deployed at the access point and is trained using multi-agent reinforcement learning. Subsequently, the topology adjustment action to be executed is verified; if the verification passes, the action is executed. Finally, the network state after executing the topology adjustment action is obtained. This invention utilizes a pre-trained execution network deployed locally at the access point and autonomously generates topology adjustment actions based on multi-agent reinforcement learning. Decentralized execution reduces transmission latency and improves real-time adjustment performance. Simultaneously, an action verification step ensures the rationality of execution and adapts to the multi-link and anti-interference characteristics of WiFi 7 networks. Furthermore, it can dynamically respond to changes in network load and user distribution, optimize access point resource utilization and coverage balance, and reduce energy consumption. In addition, the controller also optimizes the network globally to continuously improve topology adjustment adaptability and ensure user QoS requirements and network operation stability. Attached Figure Description
[0022] Figure 1 is a flowchart of a specific implementation of the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0023] Figure 2 is a schematic diagram of the network architecture of the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0024] Figure 3 is a schematic diagram of cell merging in the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0025] Figure 4 is a schematic diagram of cell splitting in the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0026] Figure 5 is a schematic diagram of resource sharing in the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0027] Figure 6 shows the relationship between throughput and training rounds under different traffic loads for the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in the embodiments of the present invention.
[0028] Figure 7 shows the relationship between packet delivery rate and training rounds under different traffic loads for the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in the embodiments of the present invention.
[0029] Figure 8 shows the relationship between average latency and training rounds under different traffic loads for the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in the embodiments of the present invention.
[0030] Figure 9 shows the training metrics of reward, loss, learning rate, and stability for the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in this embodiment of the invention.
[0031] Figure 10 is a feature attention map of the DC-MARL scheme of the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0032] Figure 11 is a feature attention map of the WiFi 7 scheme based on the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in the embodiment of the present invention.
[0033] Figure 12 is a heatmap of the average traffic load of the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0034] Figure 13 is a schematic diagram of the WiFi 7 network access point topology adjustment and application process based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0035] Figure 14 is a schematic diagram of a WiFi 7 network access point cell configuration device based on dynamic cell multi-agent reinforcement learning provided in an embodiment of the present invention.
[0036] Figure 15 is a block diagram illustrating the internal structure of the terminal device provided in an embodiment of the present invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0038] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content, operations, or steps, nor does it require execution in the described order. For example, some operations or steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0039] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0040] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. For example, "first control information" and "second control information" are only used to distinguish different control information and do not limit their order.
[0041] Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or the order of execution, and that the words "first" and "second" do not necessarily imply that they are different.
[0042] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0043] As a core infrastructure for global connectivity, Wireless Local Area Networks (WLANs) not only provide seamless internet access for mobile devices in high-density urban scenarios but also supplement cellular network coverage blind spots, ensuring uninterrupted access to core services, online resources, and digital platforms. With the development of the digital economy, the explosive growth of data-intensive applications has placed new demands on WiFi network performance. Applications such as high-definition video, virtual reality, augmented reality, cloud gaming, remote work, and wireless screen projection not only require ultra-high-speed transmission capabilities but also have extremely low latency requirements. Furthermore, professional scenarios such as industrial automation and telemedicine also have high requirements for data transmission reliability to replace wired connections. To address these challenges, the next-generation WiFi standard, WiFi 7, has been proposed. Its goal is to support higher peak throughput and optimize latency and jitter under worst-case conditions, meeting the performance requirements of future applications through new Physical Layer (PHY) and Medium Access Control (MAC) technologies.
[0044] As a new wireless communication standard, WiFi 7 integrates several key innovative technologies, significantly improving network performance. First, 4096-QAM technology, compared to WiFi 6's 1024-QAM, encodes 12 bits of data per symbol, improving spectral efficiency by approximately 20%. Second, Multi-Link Operation (MLO) allows devices to simultaneously establish transmission links in multiple frequency bands (2.4GHz, 5GHz, and 6GHz), achieving higher throughput, interference-resistant handover, and concurrent transmission and reception through aggregation, handover, and duplex modes. Third, the Enhanced Multi-Resource Units (MRU) mechanism dynamically combines discontinuous resource units, avoiding bandwidth waste caused by spectrum fragmentation and improving throughput in congested environments. Fourth, Preamble Puncturing technology selectively disables occupied subcarriers, enabling high-throughput transmission using remaining spectrum in complex interference environments and reducing bandwidth waste. Fifth, Advanced Hybrid ARQ (HARQ) improves Packet Delivery Rate (PDR) by combining retransmission with soft combining. In addition, WiFi 7 incorporates features such as Time-Sensitive Networking (TSN) and Restricted Target Wake Time (R-TWT) to further optimize deterministic latency and energy efficiency.
[0045] Despite the advanced technological potential of WiFi 7, existing networks still have many key limitations that make it difficult to fully realize its performance advantages. Traditional WiFi systems lack a sophisticated dynamic load balancing mechanism. During peak hours, some access points (APs) become overloaded and congested, while others remain idle, resulting in poor overall network performance. Energy efficiency is low, with underutilized APs continuing to consume significant amounts of power, leading to unnecessary energy waste. Static cell size configurations cannot adapt to dynamically changing user densities, causing severe congestion in high-density areas and insufficient resource utilization in sparse areas. Resource sharing between adjacent APs is limited, resulting in low bandwidth utilization and significantly increased latency under peak load. Most existing systems rely on centralized control or manual intervention for configuration adjustments, lacking autonomous adaptability. This fails to meet the scalability requirements of large-scale networks and cannot quickly respond to dynamic network conditions such as user movement, traffic fluctuations, and interference changes. Furthermore, existing solutions lack sufficient integration of WiFi 7's core features, failing to leverage intelligent decision-making mechanisms to coordinate technologies such as 4096-QAM, MLO, and MRU. This makes it difficult to translate the theoretical performance of WiFi 7 into stable gains in practical applications, failing to fully meet the comprehensive requirements of high throughput, low latency, high reliability, and low energy consumption.
[0046] It is understandable that existing WiFi 7 networks suffer from problems such as low load balancing efficiency, unreasonable energy consumption control, lack of dynamic adaptability in cell configuration, insufficient resource sharing capabilities, lack of self-adaptation and scalability, insufficient QoS (Quality of Service) guarantee capabilities, and insufficient collaborative utilization of WiFi 7 features.
[0047] Therefore, this embodiment models WiFi 7 dynamic cell control as a multi-agent Markov Decision Process (MDP) integrating real-time user load, interference, and energy indicators. To achieve optimal decision-making in this MDP, this embodiment proposes a Dynamic Cell Multi-Agent Reinforcement Learning (DC-MARL) framework. This framework is based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm and extends it to form a core policy optimization algorithm (MATD3) adapted to multi-agent collaborative scenarios. As a policy optimization algorithm under the DC-MARL architecture, MATD3 has a dual Q-evaluation network and an adaptive target update mechanism, which can effectively eliminate the Q-value overestimation problem and ensure stable policy convergence in a dynamic network environment with multiple APs.
[0048] The network architecture of this embodiment is shown in Figure 2, which includes a controller and several access points. The access points are responsible for the network connection of several user devices.
[0049] Within this DC-MARL framework, four strategies—cell merging, cell splitting, cell adjustment, and access point coordination—can be used, as shown in Figures 3, 4, 5, and 6 respectively, enabling intelligent decision-making based on network traffic. Furthermore, WiFi 7-specific physical layer and media access control layer mechanisms, such as 4096 quadrature amplitude modulation, multi-link operation, and preamble punching, are integrated into the learning environment to achieve realistic performance evaluation. Finally, extensive simulation experiments demonstrate that the proposed framework significantly outperforms WiFi 7 benchmark solutions in terms of cumulative reward, throughput, latency, and energy efficiency.
[0050] This embodiment provides a WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning, which is applied to the access point of a WiFi 7 network, as shown in Figure 1. The method specifically includes the following steps: Step S100, obtaining network status.
[0051] In this embodiment, the access point (AP) is the core device in the WiFi 7 network that provides wireless access services to user terminals. It has the functions of signal transmission and reception, data forwarding and status awareness. It can be deployed in high-density scenarios such as shopping malls and office buildings to directly establish a wireless connection with user terminals and transmit data.
[0052] Network status is a comprehensive data set reflecting the operational status of an access point and the characteristics of its surrounding environment. It covers key indicators such as the number of connected users, traffic load, and remaining energy, comprehensively depicting the operational status of the access point and the dynamic changes in the network environment. Actions are configuration optimization operations that the access point can perform, which may include merging, splitting, load sharing, resizing, and enabling Multi-Link Operation (MLO). These actions directly affect the cell configuration and resource allocation of the access point and are means to optimize network performance.
[0053] In one implementation, obtaining the network status specifically includes the following steps: Step S110, the access point collects its own operating data and obtains the operating data of neighboring access points through the communication link; Step S120, the access point integrates its own operating data with the operating data of neighboring access points to obtain an initial network status; Step S130, the initial network status is filtered and normalized to form a final network status; wherein, the network status includes the number of users connected to the access point, current traffic load, remaining energy, average load of neighboring access points, normalized effective bandwidth, multi-link operation enabled status, current cell size, frequency band used, and time normalization factor.
[0054] In this embodiment, obtaining the network status of each access point includes three core steps: collection, integration, filtering, and normalization, ultimately forming a standardized network status vector that can support action generation.
[0055] The first step is data acquisition. Specifically, each access point collects its own operational data through built-in sensors and communication modules, including the number of connected users, current traffic load, and remaining energy. Simultaneously, it acquires operational data from neighboring access points via communication links, such as the MLO link in WiFi 7, including the average load of neighboring access points and the status of multi-link operation. This ensures that access points can not only perceive their own status but also obtain information about the surrounding network environment, providing data support for generating collaborative actions.
[0056] Secondly, there is data integration. Specifically, the access point integrates its own operational data with the operational data of neighboring access points to form the initial network state. The dimensions of the initial network state cover the core indicators of the access point's operation and are represented using a network state vector: Equation (1)
[0057] in, Indicates the first Each access point at time The local state vector, Given a 9-dimensional real number space, the definitions and physical meanings of the parameters in each dimension are as follows: This indicates the number of users currently associated with this access point, reflecting the service pressure on the access point. This represents the average user load of adjacent access points, obtained through real-time data interaction with surrounding APs, and is used to estimate the congestion of the surrounding network. This represents the normalized energy consumption of the access point, with a value range of [0,1]. It is calculated from the ratio of the current remaining energy of the access point to the maximum rated energy, reflecting the energy consumption status of the equipment. This represents the normalized coverage radius of the access point, with a value range of [0,1]. It is adjusted in real time by dynamic cell operations and corresponds to the standardized result of the actual physical coverage area. This represents a binary frequency band indicator, with a value of 0 or 1, used to specify the frequency band currently in which the access point is operating; Indicates the multi-link operation (MLO) status, 0 for disabled and 1 for enabled, which is associated with the multi-band concurrent transmission capability of WiFi 7; Represents traffic intensity, normalized to [0,1], calculated from the ratio of the data packet arrival rate to the service rate of the current access point, reflecting the level of transmission busyness; This represents the average packet delay, normalized to [0,1]. It is obtained by statistically analyzing the average time from the generation to the reception of data packets per unit time, reflecting the transmission timeliness. This represents the normalized time period characteristics, with a value range of [0,1]. It is calculated by the ratio of the current time to the total time period of the day and is used to capture the daily variation patterns of user demand and network behavior, such as the difference between peak and off-peak periods.
[0058] When the access point collects its own operational data, it obtains real-time data through hardware units such as the built-in user connection detection module, traffic statistics module, and energy consumption monitoring module. , , It acquires its own status parameters. When obtaining operational data from adjacent access points, it uses a dedicated communication link between the access points to achieve this. , Real-time sharing of status information of neighboring nodes.
[0059] Finally, there is the filtering and normalization process. Specifically, the initial network status of the access points is filtered to remove outliers, such as data exceeding reasonable ranges due to sensor malfunctions. Then, each indicator is normalized, mapping all data to the [0,1] interval to eliminate numerical magnitude differences caused by different units and avoid affecting the decision-making accuracy of subsequent controllers. For example, the number of users is normalized to a reasonable range of "0-50 people", and the latency is normalized to a range of "0-10ms".
[0060] Following the dimensional order of equation (1), the preprocessed parameters are combined to form the local state vector. This allows the vectors to comprehensively and accurately reflect the real-time status of the access point itself and the surrounding local network, providing high-quality basic data for the controller to construct the global network status.
[0061] The comprehensiveness of data acquisition ensures that the network state vector can fully characterize the access point itself and its surrounding environment, providing a foundation for accurate action generation. The standardized state vector design of Equation (1) unifies the data structure and adapts to the input requirements of the execution network. Filtering and normalization processes improve data quality, avoid interference from outliers and magnitude differences in training, and accelerate model convergence. At the same time, the design of multi-dimensional state vectors not only covers the core characteristics of WiFi7 networks but also takes into account network performance and device status, enabling decisions on diverse actions such as merging and splitting.
[0062] Step S200: Based on the network state, generate a topology adjustment action to be executed through a pre-trained execution network; wherein the pre-trained execution network is locally deployed at the access point, and the pre-trained execution network is trained through multi-agent reinforcement learning.
[0063] In this embodiment, a decentralized action generation mechanism involving multiple agents is proposed. Each access point independently generates actions based on locally collected network conditions, without waiting for global scheduling instructions from the controller. This mechanism is well-suited to the widely distributed and highly dynamic characteristics of WiFi 7 network access points. Specifically, decentralized execution reduces communication overhead between the controller and access points, minimizes action generation latency, and ensures that access points can respond in real-time to changes in local network conditions, such as sudden user growth. Independent exploration by each access point enriches the diversity of experience data, avoids the limitations of local exploration by a single access point, and makes the overall network configuration strategy more adaptable. Furthermore, the diversified design of action types covers core scenarios of cell topology adjustment and resource scheduling, and can specifically address key issues in traditional networks such as load imbalance and resource waste.
[0064] The execution network is a decentralized neural network deployed at each access point. Its core function is to generate adapted configuration actions based on the network state collected locally at the access point. Its input is a network state vector, and its output is a mixed action vector containing discrete action log probabilities and continuous adjustment amounts.
[0065] In one implementation, the topology adjustment action is an access point merging action, which specifically includes the following steps: Step S210a, after the local access point and the adjacent access point cooperate, the local access point goes into hibernation and transfers the user.
[0066] In this embodiment, Figure 3 illustrates the network structure of the merging action. Specifically, the merging action integrates resources through inter-access point collaboration, achieving energy saving and efficient spectrum utilization in low-load scenarios. When a local access point detects insufficient resource utilization, it coordinates with neighboring access points to completely transfer the users it serves before entering a dormant mode. This reduces the energy consumption of idle devices and avoids fragmented occupation of spectrum resources by scattered access points.
[0067] Specifically, the triggering of a merge action needs to be based on the load status of the access point, and the rationality of the operation is ensured through discrete action decision-making and feasibility verification. The execution network of the access point generates the logarithmic probability of the merge action based on features such as the number of users in the local state vector and the average load of adjacent access points. After filtering by softmax operation and argmax function, it determines whether to initiate a merge. During the process, the execution network identifies low-load scenarios through the strategy learned and trained, and actively initiates merges to optimize resource allocation.
[0068] The merging action optimizes energy consumption. Dormant access points can completely shut down their signal transceiver modules, stopping unnecessary energy consumption. Simulation results show that merging significantly reduces overall network energy consumption under low-load scenarios. Secondly, it improves spectral efficiency. Reducing scattered access points results in more regular coverage areas between adjacent access points, reduced spectral interference, and improved effective bandwidth utilization.
[0069] In one implementation, the topology adjustment action is an access point splitting action, which specifically includes the following steps: Step S210b, activate the newly added access point and partition the load and coverage of the original access point.
[0070] In this embodiment, Figure 4 illustrates the network structure of the splitting operation. The splitting operation is a core topology adjustment method for handling high-load scenarios. By activating new access points and rationally partitioning load and coverage areas, dynamic expansion of network capacity is achieved, alleviating local congestion. Overloaded access points are split into two or more independent access points, each taking on a portion of the load and coverage area, allowing the network topology to adapt to dynamic changes in user density.
[0071] The splitting action is triggered by the execution network based on local state vector decisions. The execution network at the access point extracts features such as the number of users, traffic intensity, and user density from the state vector, outputs the logarithmic probability of the splitting action through neural network calculations, and combines softmax and argmax operations to determine whether to initiate a split. The decision-making process identifies high-load scenarios and proactively initiates splitting to expand capacity.
[0072] Specifically, the splitting process involves two steps. First, the new access point is activated, and load and coverage are partitioned. The new access point is typically a pre-installed backup device in the network, normally in a low-power standby state. Upon receiving the splitting command, it can be activated and enter working mode within tens of milliseconds. Load and coverage partitioning is based on user distribution data and geographical information. A spatial balancing algorithm is used to divide the coverage area of the original access point into two non-overlapping areas. At the same time, according to user QoS requirements and traffic characteristics, the load is evenly distributed between the original access point and the new access point. For example, in a high-density office area, when the number of users at a certain access point exceeds the high load threshold, the splitting process activates the backup access point, dividing the office area into two coverage zones along the corridors. The original access point handles the load for users on the north side, and the new access point handles the load for users on the south side, achieving load balancing.
[0073] The splitting action significantly improves network capacity. After splitting, the total service capacity of the two access points increases, effectively alleviating congestion in high-load scenarios and greatly reducing transmission latency. Overloaded access points typically have high queuing latency, but after splitting, the queuing latency of each access point can be reasonably reduced, meeting the low latency requirements of WiFi 7.
[0074] Step S300: Verify the topology adjustment action to be executed. If the verification passes, execute the topology adjustment action to be executed.
[0075] In this embodiment, the actions of each access point generated by the network are executed. The core is to generate a hybrid action vector and transform it into a specific action, thereby achieving the synergy between discrete topology adjustment and continuous parameter optimization.
[0076] First, the hybrid action vector is generated. Specifically, the main execution network of each access point calculates and outputs the hybrid action vector through forward propagation based on the network state vector in equation (1). The dimension of the hybrid action vector can be adjusted based on the optional actions. Here, the optional actions are listed and represented using a 6-dimensional vector: Equation (2)
[0077] Among them, the first 5 dimensions The logits represent the discrete actions, corresponding to five preset discrete topology operations: Merge, Split, Share Load, Resize, and Enable MLO; the sixth dimension... This is the adjustment amount for the continuous cell radius, and its value range is... , This is the preset maximum radius adjustment range.
[0078] Subsequently, actions are calculated and executed based on the hybrid action vector.
[0079] The hybrid action vector design combines discrete topology adjustment and continuous parameter optimization. It can solve the load imbalance problem at the network topology level through actions such as merging and splitting, and optimize the coverage range through continuous radius adjustment, adapting to the complex optimization needs of WiFi 7 networks.
[0080] Actions are calculated and executed based on hybrid action vectors. Specifically, the softmax operation and argmax function are used to accurately select discrete actions, ensuring the rationality and efficiency of action execution.
[0081] First, the softmax operation is performed. Specifically, after the mixed action vector is generated, it is transformed into a specific executable action through subsequent processing. First, the logarithmic probabilities of the first 5 discrete actions are subjected to a softmax operation to transform them into selection probabilities between 0 and 1, expressed as: Equation (3)
[0082] in, For the first The probability of choosing a discrete action is given, and the sum of all probabilities is 1. A softmax operation is performed on the 5-dimensional discrete action log odds in the mixed action vector to transform the unnormalized log odds into a probability distribution between 0 and 1, with the sum of the probabilities of all discrete actions being 1. The log odds themselves are unbounded real numbers; they are mapped to positive numbers using an exponential function, and then divided by the sum of all exponents to achieve probability normalization. For example, if the log odds of splitting the action in the mixed action vector are 2.1, the log odds of merging the action are 1.3, and the log odds of other actions are all less than 1, then after the softmax operation, the probability of splitting the action will be significantly higher than that of other actions, reflecting its optimality in the current state.
[0083] Then, the discrete action with the highest probability is selected by the argmax function, which is expressed as: Equation (4)
[0084] If the selected discrete action is to adjust the size, that is... Then combine with continuous adjustment amount Perform cell radius adjustment; for other discrete actions, directly execute the corresponding topology operation, such as merging or splitting. The argmax function iterates through the five action probabilities obtained from the softmax operation and selects the discrete action with the highest probability as the final action to be executed. This ensures that the access point can choose the action that is most likely to optimize network performance in the current state, avoiding the execution of invalid actions. For example, when the probability of the load sharing action is 0.6, and the probabilities of other actions are all less than 0.2, the argmax function will directly select the load sharing action, guiding the access point to transfer some of the load to adjacent access points, alleviating its own congestion.
[0085] The softmax operation provides a probabilistic basis for action selection, making it flexible and avoiding overexploration of single actions. It also provides probabilistic interpretability, preventing insufficient exploration caused by deterministic choices. The argmax function ensures the selection of the highest-priority action, guaranteeing targeted configuration and improving efficiency. This process is closely integrated with the execution network's output, achieving a smooth transformation from action vectors to specific actions, ensuring the precise implementation of the network's optimization goals. For example, when the execution network learns during training that splitting actions is better under high load, it will output a higher log-probability of splitting actions. After softmax and argmax processing, the access point will prioritize splitting actions, optimizing network performance. This design allows access points to flexibly select action types based on network conditions, such as prioritizing splitting actions under high load and merging actions under low load and high energy consumption, achieving dynamic adaptive configuration.
[0086] In one implementation, the topology adjustment action is an access point merging action. The verification of the topology adjustment action to be executed specifically includes the following steps: Step S310a: Based on the network status, obtain the current traffic load of the local access point, the average load of adjacent access points, and the remaining energy of the local access point; Step S320a: Compare the current traffic load of the local access point with a preset low load threshold, compare the average load of adjacent access points with a preset coordination threshold, and compare the remaining energy of the local access point with a preset energy-saving threshold; Step S330a: If all three comparison results satisfy the less than or equal to relationship, the verification result of the merging action is obtained.
[0087] In this embodiment, the rationality of the merging action is verified. By quantitatively determining the load conditions, it is ensured that the merging action will not lead to service interruption or resource allocation imbalance. The verification process is based on the merging rules, combined with the load status of the local and adjacent access points, to form clear feasibility constraints, expressed as: Equation (15)
[0088] in, This is a flag indicating the execution of the merge action; a value of 1 indicates that the merge will be executed, and a value of 0 indicates that it will not be executed. This indicates that the discrete actions generated by the network are performed as merging. This represents the current number of users at the local access point. The average number of users per adjacent access point. This is the preset low load threshold.
[0089] Equation (15) defines the core triggering condition for the merging action. The merging action can only be executed when the load of both the local access point and the adjacent access point is below the threshold, thus fundamentally avoiding the service quality degradation caused by merging in scenarios with unbalanced load.
[0090] Equation (15) can be used alone as the verification of the topology adjustment action to be performed, or additional compliance verification can be added on the basis of the verification of Equation (15).
[0091] Specifically, this can be achieved from the network state vector. Based on this, three parameters are extracted: the normalized traffic intensity in the state vector corresponding to the current traffic load of the local access point, the average load of adjacent access points, and the complementary value of the normalized energy consumption corresponding to the remaining energy of the local access point, i.e., the remaining energy percentage.
[0092] Subsequently, the current traffic load of the local access point is compared with the preset low load threshold to ensure that the local access point is in a low load state, so as to avoid a sudden increase in the load of the remaining access points after merging. The average load of adjacent access points is compared with the preset coordination threshold to ensure that adjacent access points have the redundancy to handle additional users. Finally, the remaining energy of the local access point is compared with the preset energy saving threshold, and the merging of access points with lower remaining energy is prioritized to maximize the energy saving effect.
[0093] The verification passes when all comparisons satisfy the less than or equal to relationship. The verification logic is combined with the merging rules of equation (15) to achieve quantitative judgment of low load state, while the verification of remaining energy further optimizes the execution priority of merging action.
[0094] This verification avoids the service quality degradation caused by blind merging. For example, if the load of an adjacent access point is close to the threshold, the verification will fail and merging can be prevented to prevent new load congestion. It also achieves the accuracy of energy consumption optimization. By filtering by the remaining energy threshold, access points with tight energy consumption are given priority to hibernate, thus extending the overall network endurance.
[0095] In one implementation, the topology adjustment action is an access point merging action, and the execution of the topology adjustment action to be executed specifically includes the following steps: Step S340a: The local access point sends a coordination request to the adjacent access point through multi-link operation and receives a resource idle response returned by the adjacent access point; Step S350a: A user transfer list is generated based on user QoS priority, and the users in the list are smoothly switched to the adjacent access point through the links of the multi-link operation; Step S360a: After all users have been transferred, the local access point shuts down the signal transceiver module and enters sleep mode.
[0096] In this embodiment, the merging process revolves around two objectives: seamless user transfer and secure hibernation. Multi-Link Operation (MLO) is used to assist in ensuring the continuity of user services while guaranteeing that access point hibernation does not affect overall network coverage. The execution process consists of three steps, each of which can be integrated with the features of WiFi 7 to ensure optimal performance.
[0097] First, there's the interaction between the collaboration request and response. The local access point sends a collaboration request to neighboring access points via the WiFi 7 MLO link. This link is independent of the data transmission link, featuring low latency and high reliability. The collaboration request includes key information such as the number of users on the local access point, QoS requirement distribution, and coverage boundaries, facilitating neighboring access points' assessment of resource redundancy. Upon receiving the request, a neighboring access point determines its capacity based on its current load and remaining resource capacity. If resources are available, it returns a resource availability response, containing details such as the number of users that can be accommodated and the recommended switching link. Leveraging the multi-link concurrency characteristics of MLO, the collaboration interaction avoids impacting data transmission, ensuring that the user's current network experience remains uninterrupted.
[0098] Following this is a smooth user migration. The local access point generates a user migration list based on user QoS priorities, prioritizing high-real-time services such as cloud gaming and non-real-time services like file downloads. The migration process is achieved through MLO's seamless handover capability. The user device maintains connections with both the local access point and the target adjacent access point, completing the handover during data transmission intervals. The handover latency is controlled within a preset threshold, making it imperceptible to the user. For each user, the local access point synchronizes the current transmission status with the target access point, ensuring uninterrupted and non-duplicated data transmission after the migration. This fully leverages the multi-link collaborative advantages of WiFi 7's MLO, resolving packet loss and lag issues that easily occur in traditional single-link handover.
[0099] Finally, the local access point goes into sleep mode. After all users have migrated, the local access point reports the migration completion status to the controller via the MLO link and receives the sleep command from the controller. It then shuts down the signal transceiver module, retaining only the low-power communication module for receiving wake-up commands, and enters deep sleep mode. In sleep mode, network power consumption is significantly reduced. Simultaneously, the local access point synchronizes its sleep state to all surrounding access points, facilitating adjustments to coverage and load balancing strategies by these surrounding access points and preventing coverage blind spots.
[0100] During the merge process, the user experience is seamless. MLO's seamless handover and transmission status synchronization ensure uninterrupted user services, with a high handover success rate and significant energy reduction. Deep sleep mode greatly reduces energy consumption of idle access points. Furthermore, network stability is improved; the sleep state synchronization mechanism avoids coverage blind spots and load imbalances, guaranteeing the overall service quality of the network.
[0101] In one implementation, the topology adjustment action is an access point splitting action. The verification of the topology adjustment action to be executed specifically includes the following steps: Step S310b: Based on the network status, obtain the current traffic load, user density, and coverage overlap rate of the local access point; Step S320b: Compare the current traffic load of the local access point with a preset high load threshold, compare the user density with a preset density threshold, and compare the coverage overlap rate with a preset overlap threshold; Step S330b: If the load comparison satisfies a greater than or equal to relationship, or the density comparison satisfies a greater than or equal to relationship, and the overlap rate comparison satisfies a less than or equal to relationship, then the verification result of the splitting action is obtained.
[0102] In this embodiment, to ensure the effectiveness of capacity expansion and network stability, the splitting action needs to be verified. Load, density, and coverage are used for collaborative judgment to avoid problems such as uneven load, overlapping coverage, or blind spots after splitting. The verification process is based on the splitting rules, combined with network state parameters to form quantitative constraints. The rules are expressed as: Equation (16)
[0103] in, This is a flag indicating whether the splitting action will be executed. A value of 1 indicates that the splitting will be executed, and a value of 0 indicates that it will not be executed. This indicates that the discrete actions generated by the network are splitting. This represents the current number of users at the local access point. This is a preset high load threshold.
[0104] Equation (16) clearly defines the core triggering condition for the split action. The split action can only be executed when the load of the local access point exceeds the threshold, ensuring that the split action specifically addresses the overload problem and avoids meaningless network topology changes.
[0105] Equation (16) can be used alone as the verification of the topology adjustment action to be performed, or additional compliance verification can be added on the basis of the verification of Equation (16).
[0106] Specifically, the compliance verification uses the normalized traffic intensity in the state vector corresponding to the current traffic load of the local access point. This parameter directly reflects the load pressure of the access point and the user density. The user density can be calculated by the ratio of the number of users within the coverage area of the access point to the coverage area, reflecting the density of users in the area. The coverage overlap rate is calculated by the ratio of the coverage overlap area of the original access point and the surrounding access points to the coverage area of the original access point, reflecting the utilization efficiency of coverage resources. This verification complements the splitting rule of Equation (16). The verification first compares the load. The current traffic load of the local access point is greater than or equal to the preset high load threshold, ensuring that the splitting action is for overload scenarios. Secondly, it compares the density. The user density is greater than or equal to the preset density threshold. Even if the load does not reach the threshold, the excessive user density may lead to increased interference, so the splitting is still reasonable. Finally, it compares the overlap rate. The coverage overlap rate is less than or equal to the preset overlap threshold, ensuring that the coverage area of the newly added access point after splitting will not overlap excessively with the surrounding access points, avoiding waste of spectrum resources and increased interference.
[0107] The above rules ensure that the splitting action is targeted at addressing overload or high-density issues, while avoiding the negative impact of excessive overlap and interference after splitting. For example, if the load of an access point exceeds the threshold, but the coverage overlap rate of surrounding access points is already at a high level, the verification will fail to avoid further exacerbating interference after splitting. If the load of an access point does not reach the threshold, but the user density is extremely high and the coverage overlap rate is low, the verification will pass, and the splitting will alleviate interference and improve the user experience.
[0108] In one implementation, the topology adjustment action is an access point splitting action, and the execution of the topology adjustment action to be executed specifically includes the following steps: Step S340b: The controller receives the original access point splitting request and activates the preset new access point; Step S350b: The original access point and the new access point divide the coverage area through frequency band negotiation; Step S360b: Based on user distribution data and QoS requirements, the load partitioning of the original access point is allocated to the new access point; Step S370b: Preamble punching technology is enabled to scan and avoid spectrum interference, and the load and coverage area splitting configuration is completed.
[0109] In this embodiment, the purpose of the splitting operation is to rapidly expand capacity, balance the load, and reduce interference. Leveraging the technical characteristics of WiFi 7, the topology adjustment is completed step-by-step to ensure a significant improvement in network performance after execution. The execution process consists of four steps, each deeply integrated with the PHY / MAC layer characteristics of WiFi 7 to maximize the splitting effect.
[0110] The first step is activating the new access point. After receiving the split request from the original access point, the controller immediately sends an activation command to the preset backup access point. The backup access point uses WiFi 7's fast startup mechanism to complete hardware initialization, frequency band synchronization, parameter configuration, and other processes, entering a ready state. During the activation process, the controller sends the original access point's network configuration information, such as SSID, security key, and QoS policy, to the backup access point to ensure seamless switching for user devices without the need to reconfigure network connections.
[0111] The next step is coverage area allocation. The original access point and the new access point negotiate their respective operating frequency bands, typically using a dual-band (5GHz and 6GHz) collaborative mode. The original access point occupies the 5GHz band, and the new access point occupies the 6GHz band, leveraging the advantages of the 6GHz band's low interference and high bandwidth to improve service quality. Coverage area allocation is based on geographic information and user distribution data, employing a spatial clustering algorithm to divide the original access point's coverage area into two continuous and non-overlapping sub-regions. A certain amount of edge overlap is reserved during the allocation process to avoid coverage blind spots. Simultaneously, user equipment in the edge overlap area can maintain a weak connection with both access points simultaneously through MLO technology, ensuring seamless handover during mobility.
[0112] Next comes load balancing. Based on user QoS requirements and traffic characteristics, a priority-weighted allocation algorithm is used to distribute the load. High-priority users are preferentially allocated to the less interference-prone 6GHz band, i.e., the new access point, while medium- and low-priority users are allocated to the 5GHz band, i.e., the original access point. Simultaneously, based on each user's historical traffic data, it is ensured that the difference in traffic load between the two access points does not exceed a preset threshold. During the allocation process, the original access point synchronizes the user's transmission status information with the new access point to ensure the continuity of data transmission after load transfer.
[0113] Finally, interference avoidance and configuration activation are implemented. WiFi 7 preamble punching technology is enabled. The original access point and the new access point scan for interference within their respective frequency bands, marking occupied subcarriers and disabling them through punching. Transmission is then performed using only available subcarriers, effectively avoiding spectrum interference. Simultaneously, both access points dynamically combine discontinuous resource units through a Multi-Resource Unit (MRU) mechanism, improving bandwidth utilization. After configuration, the original and new access points synchronously broadcast their coverage area and operating frequency band information to surrounding access points. These access points then adjust their transmit power, channel selection, and other transmission parameters based on this information, further reducing overall network interference.
[0114] The splitting mechanism enables rapid capacity expansion, meeting the real-time needs of highly dynamic scenarios, and provides excellent load balancing. The load difference between the two access points after splitting is minimal, preventing the emergence of new overload points. Furthermore, it achieves low interference; preamble punching and frequency band separation technologies reduce interference levels after splitting, and the user experience is seamless. Rapid activation and seamless switching technologies ensure uninterrupted user services.
[0115] Step S400: Obtain the network status after performing the topology adjustment action.
[0116] In this embodiment, the network state after the action is executed is the comprehensive state data updated after the access point completes the configuration operation. It can directly reflect the impact of the action execution on network operation and is the basis for evaluating the effect of the action. The reward is a quantitative evaluation index designed based on global network performance indicators. It is used to guide the iterative optimization of model parameters. Its calculation integrates core performance dimensions such as network throughput, packet delivery rate, transmission delay, and access point energy consumption, and can achieve multi-objective collaborative optimization.
[0117] This embodiment also proposes a closed-loop design for real-time status feedback and multi-dimensional reward calculation. Real-time acquisition of the network state after an action is executed ensures the accuracy and timeliness of reward calculation, enabling the model to quickly perceive the impact of actions on the network. The multi-dimensional reward design overcomes the limitations of traditional single-index optimization, balancing objectives such as throughput improvement, latency reduction, and energy consumption optimization through weight allocation, avoiding network performance imbalances caused by optimizing a single index. For example, the normalization of transmission latency in the reward function ensures consistent evaluation of latency indicators under different traffic loads, guiding the model to generate configuration actions that balance real-time performance and energy efficiency, providing reliable value guidance for subsequent Q-value calculation and parameter updates.
[0118] Optionally, the reward can be calculated by obtaining the network status after the access point performs an action. Specifically, by fusing multi-dimensional performance indicators, a clear optimization guide is provided to the model, achieving a synergistic improvement in throughput, latency, and energy consumption.
[0119] First, data collection is performed. Specifically, after the access point executes its action, the controller collects two types of data through the communication link: one is the network status of each access point after execution, including the updated number of users, load, coverage radius, etc., corresponding to the state vector update in equation (1); the other is the global performance indicators of the WiFi 7 network, specifically including transmission delay, network throughput, packet delivery ratio (PDR), and access point energy consumption. These indicators comprehensively cover network service quality and equipment operating efficiency, and can comprehensively reflect the effect of action execution.
[0120] Next is the normalization of transmission delay. Specifically, since the reasonable range of transmission delay varies greatly under different traffic loads, directly using it for reward calculation will lead to evaluation bias. Therefore, it is necessary to normalize the transmission delay to obtain the normalized delay. The normalization uses a bounded function, expressed as: Equation (6)
[0121] in, Access point Normalization delay, For the original transmission delay, This is the preset maximum acceptable latency. When the original latency is less than... When the normalized delay is constant, the normalized delay has a linear relationship with the original delay; when the original delay exceeds a certain threshold... At this time, the normalized delay is fixed at 1 to avoid the excessive impact of extreme delay on the reward.
[0122] against In the case where the normalized delay is determined, the sensitivity of the reward to the original delay can be quantified by equation (7): Equation (7)
[0123] In this formula, Indicates reward For the original delay The partial derivative reflects the change in reward for every unit change in the original delay; The weight of the delay metric in the reward function. The maximum acceptable delay is given by equation (7), which reveals the sensitivity of rewards to changes in delay. The larger the value, the smaller the absolute value of the sensitivity, and the smoother the response of the reward to delay fluctuations. The smaller the value, the larger the absolute value of the sensitivity, and the more significant the penalty for latency changes. This sensitivity adjustment mechanism makes the penalty for latency during training more controllable, avoiding training oscillations caused by differences in the reasonable range of latency in different scenarios.
[0124] Finally, the reward is calculated using a fusion method. Specifically, the reward function integrates the global performance metric and the normalized latency, and is expressed as a weighted sum, as shown in equation (5).
[0125] in, To normalize throughput, To normalize the group delivery rate, Normalized energy consumption; , , , These are non-negative weights used to balance the importance of different metrics. The design logic of the reward function is: increasing throughput and group delivery rate will increase the reward, while increasing latency and energy consumption will decrease the reward, guiding the model to generate actions that take into account multiple objectives.
[0126] Comprehensive data collection ensures that reward calculations reflect the true effects of actions, avoiding evaluation biases caused by single metrics. Normalization makes different metrics comparable, ensuring the stability of reward calculations. The weighted fusion reward function achieves multi-objective collaborative optimization, preventing the model from overemphasizing one metric while neglecting others. For example, when an action increases throughput but significantly increases energy consumption, the reward function reduces the value of that action through the penalty effect of energy consumption weight, guiding the model to select a more balanced action configuration.
[0127] In one implementation, the WiFi 7 network includes several access points and a controller. After obtaining the network status after performing the topology adjustment action, the process further includes the following steps: Step S510: Each access point reports the performed topology adjustment action and the network status after performing the topology adjustment action to the controller; Step S520: The controller obtains the global network status; Step S530: The controller optimizes the execution network of each access point based on the topology adjustment action performed by each access point, the network status of each access point after performing the topology adjustment action, and the global network status.
[0128] In this embodiment, the controller is the global coordination unit of the entire network. In the HarmonyOS scenario, it can act as a super device, responsible for integrating the status information of all access points, formulating global optimization strategies, issuing action commands, and iteratively updating model parameters. It is the core hub for realizing multi-access point collaborative optimization.
[0129] Each access point reports the executed topology adjustment actions and the network state after the actions to the controller. The controller generates a global network state and calculates the corresponding Q value. Based on the Q value, the execution network at each access point is optimized.
[0130] The Q-value is a quantitative evaluation of the value of an action. It represents the expected cumulative reward that can be obtained in the future after performing an action in the current network state. Its purpose is to provide a quantitative standard for the quality of actions and help the model distinguish the long-term value of actions with different configurations. The calculation of the Q-value is based on the evaluation network's modeling of the mapping relationship between network state, action, and reward. It integrates the weighted sum of current reward and future reward, reflecting the long-term optimization effect of the action.
[0131] This embodiment proposes a multi-agent collaborative Q-value calculation logic. The network evaluation calculates the Q-value based on the global network state and the actions of all access points, rather than evaluating the value of a single access point's action in isolation. This logic considers the synergistic effect of multiple access point actions, avoiding the global performance degradation caused by local optima of a single access point. For example, when an access point performs a merging action, the Q-value calculation simultaneously considers its impact on the load distribution of adjacent access points and network interference, ensuring that the action selection aligns with the global optimization objective. The introduction of the Q-value allows the model to break free from the limitations of traditional rule-based configuration, learning optimal action selection strategies adapted to dynamic network environments through a data-driven approach, providing a core basis for subsequent parameter updates.
[0132] The Q-value is calculated by evaluating the network, integrating the global network status with the actions and rewards of all access points, and calculating the current Q-value and target Q-value through the main evaluation network and the target evaluation network respectively, providing a valuable basis for parameter updates.
[0133] The first step is the construction of the global network state. Specifically, the controller collects the network state after all access points have performed actions and integrates it to form the global network state. ,in This represents the number of access points. The global network status reflects the operational condition of the entire WiFi 7 network, providing a holistic perspective for centralized evaluation.
[0134] Then the main evaluation network calculates the current Q-value. Specifically, it considers the global network state and the actions of all access points. and the calculated reward Input the main evaluation network. The dual-Q structure of the main evaluation network will calculate the two Q values separately. and 2. Finally, the minimum value between the two is selected as the current Q-value output by the main evaluation network. It effectively suppresses overestimation of the Q value, ensuring the accuracy of the current Q value.
[0135] Finally, the target evaluation network calculates the target Q-value. Specifically, this involves evaluating the global network state and the target actions generated by the target execution network. And reward input target evaluation network. Target action The target execution network generates actions based on the next state, and also includes discrete actions and continuous adjustments. The target evaluation network also uses a dual-Q structure to calculate two target Q values. and The minimum value is selected as the final target Q value. .
[0136] The construction of the global network state ensures that Q-value calculations take into account the synergistic effects of access point actions, avoiding misjudgments caused by local perspectives. The dual-Q structure of the main evaluation network and the target evaluation network respectively guarantees the accuracy of the current Q-value and the target Q-value, providing reliable data for subsequent time-series difference error calculations. For example, when multiple access points collaboratively execute load-sharing actions, the global network state can reflect the balance of load distribution. The main evaluation network will output a higher current Q-value, while the target evaluation network predicts subsequent rewards based on the next state, guiding the model to continuously select collaborative actions.
[0137] In one implementation, the controller optimizes the execution network of each access point based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state. Specifically, this includes the following steps: Step S531: Based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state, the controller optimizes the execution network of each access point through value evaluation using multi-agent reinforcement learning. The multi-agent reinforcement learning includes a collaboration between centralized evaluation and decentralized execution. The centralized evaluation generates a global value evaluation, and the local execution network of each access point implements the decentralized execution.
[0138] In this embodiment, the WiFi7 network in the DC-MARL architecture includes a controller and several access points, forming a centralized evaluation and decentralized execution architecture, wherein the execution network is deployed at each access point.
[0139] From an implementation perspective, the controller, acting as the global coordination hub, possesses powerful computing and data integration capabilities. It can collect network status and action data from all access points to construct a global network view. The evaluation network, deployed on the controller, can assess the collaborative value of all access point actions from a global perspective, avoiding the local optimum trap caused by decentralized evaluation. For example, when multiple access points simultaneously execute split actions, the centralized evaluation network can identify potential load overlap and interference issues, guiding action coordination through Q-value adjustments.
[0140] The execution network is deployed at each access point, enabling each access point to generate actions in real time based on locally collected network status data, without waiting for remote commands from the controller. This design significantly reduces communication latency and adapts to the high-dynamic, low-latency configuration requirements of WiFi 7 networks. For example, when the number of users at an access point suddenly surges, the local execution network can quickly generate split actions to alleviate load pressure in a timely manner, without relying on global scheduling by the controller.
[0141] In this architecture, centralized evaluation ensures consistency in global optimization and avoids conflicts in access point actions, while decentralized execution guarantees the real-time and flexible nature of action generation, improving the network's response speed to dynamic environments. The combination of these two approaches enables the model to possess both global optimization capabilities and the real-time response capabilities of individual access points, fully leveraging the hardware performance potential of WiFi 7 networks and solving the core problems of high latency in traditional centralized scheduling and poor coordination in decentralized configuration.
[0142] In one implementation, the centralized evaluation includes: step S5311, deploying a centralized evaluation network based on a dynamic cell multi-agent reinforcement learning framework; step S5312, inputting the global network state and the actions of each access point through the evaluation network, and calculating the global value assessment of each action.
[0143] In this embodiment, the evaluation network is a centralized neural network deployed on the controller, used to evaluate the long-term value of the actions generated by the execution network and provide a quantitative basis for parameter updates. Its inputs cover the global network state and the actions of all access points, and the output is the Q value corresponding to the action, i.e., the action value estimate.
[0144] The evaluation network consists of a main evaluation network and a target evaluation network. The target evaluation network is a copy network with the same structure as the main evaluation network. The main evaluation network adopts a centralized double-Q evaluation network architecture, which contains two evaluation sub-networks with the same structure and independent parameters. Before training begins, the initial parameters of the main evaluation network are completely assigned to the target evaluation network.
[0145] This embodiment proposes a dual-Q structure main evaluation network. The two evaluation sub-networks have identical structures, both being neural networks containing three fully connected layers. Specifically, the input layer dimension is the sum of the global network state vector dimension and the action vector dimensions of all access points, the hidden layers use the ReLU activation function, and the output layer is a one-dimensional Q-value. The initial parameters of both sub-networks follow a normal distribution, ensuring the randomness and consistency of the initial parameters while avoiding training problems caused by excessively large or small parameters.
[0146] The target evaluation network has the same structure as the main evaluation network, including the number of network layers, the number of neurons per layer, and the type of activation function; the only difference is the parameter update mechanism. The parameter assignment operation before training begins involves completely copying the random initial weights and biases of the main evaluation network to the target evaluation network. This ensures that the target network maintains parameter synchronization with the main network from the initial training phase, laying the foundation for stable calculation of the target Q-value later.
[0147] The technical advantage of the dual-Q evaluation network lies in its effective suppression of Q-value overestimation. Traditional single-Q networks are prone to Q-value overestimation due to noise or sample bias, leading the model to select suboptimal actions. In this embodiment, two independent evaluation sub-networks calculate Q-values separately, and the minimum of the two is taken as the output Q-value of the main evaluation network. A double-validation mechanism filters out overestimated Q-values, ensuring the accuracy of action value assessment. The target evaluation network provides a stable target Q-value reference. Specifically, the parameters of the main evaluation network are updated rapidly during training, while the parameters of the target evaluation network are updated later, using a soft update mechanism to adjust slowly, avoiding target Q-value oscillations caused by fluctuations in the main network parameters, thus improving training stability. The combination of the two makes the Q-value calculation of the evaluation network both accurate and stable, providing a reliable basis for subsequent parameter updates, accelerating model convergence, and improving final configuration performance.
[0148] In one implementation, the optimization of the execution network of each access point based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state, through value evaluation using multi-agent reinforcement learning, specifically includes the following steps: Step S5313, global value evaluation based on the output of a centralized evaluation network; Step S5314, updating the parameters of the execution network through temporal differential error calculation; Step S5315, the controller sends the updated execution network parameters to the corresponding access point; Step S5316, the access point receives and updates the parameters of its local execution network.
[0149] In this embodiment, the execution network includes a main execution network and a target execution network. The target execution network is a replica network with the same structure as the main execution network. The main execution network is a decentralized execution network. The main execution network of each access point outputs the corresponding action through the network state. Before training begins, the initial parameters of the main execution network are completely assigned to the target execution network.
[0150] The decentralization of the main execution network is reflected in the independent deployment of a set at each access point, enabling action generation without relying on state data from other access points or controllers. Its network structure can be a three-layer fully connected layer. Specifically, the input layer dimension is the dimension of the access point's local network state vector, the hidden layer uses the ReLU activation function, and the output layer is a hybrid action vector. This structure maps local states to the action space, ensuring the targeted and efficient generation of actions.
[0151] The target execution network has the same structure as the main execution network, including the number of neurons in each layer, activation functions, and output layer design; the only difference is the parameter update mechanism. The parameter assignment operation before training begins involves completely copying the random initial parameters of the main execution network at each access point to the corresponding target execution network, ensuring that the initial state of the target execution network is consistent with that of the main execution network.
[0152] The technical advantages of the decentralized master execution network are reflected in its real-time performance and flexibility. Specifically, each access point quickly generates actions based on its local state, reducing dependence on communication links and avoiding scheduling delays caused by centralized execution. For example, when an access point detects that the load of a neighboring access point is low and its own remaining energy is insufficient, the local master execution network can quickly generate a merge action to reduce energy consumption in a timely manner, without waiting for the controller's global decision. The technical advantage of the target execution network is that it provides stable target actions: its parameter updates lag behind the master network, and it is slowly adjusted through a soft update mechanism. The generated target actions are not affected by fluctuations in the master network parameters, providing reliable input for the calculation of the target Q-value, avoiding distortion of the target Q-value caused by rapid changes in the master execution network parameters, and ensuring the stability of the training process.
[0153] Based on the Q value, the parameters of the evaluation network are iteratively updated by minimizing the loss function, and the parameters of the execution network are iteratively updated based on the optimized evaluation network.
[0154] The parameters of the evaluation network and the execution network are updated iteratively based on the Q-value. Specifically, the prerequisite for updating the parameters of the main evaluation network is to first complete the preparatory work, namely the generation of the target action, and then the parameters of the main execution network and the soft update of the target network are implemented in sequence to achieve stable optimization of the model.
[0155] First, there is the generation of the target action. During training, when the target execution network generates the target action based on the next state, Gaussian noise is added to improve training stability, specifically expressed as: Equation (8)
[0156] in, Zero-mean Gaussian noise, cropped to the range [ δ, δ], The standard deviation of noise. This is the noise clipping threshold; clipping is used to prevent excessive noise from distorting the target's motion. Execute the network for the target. Its parameters, This represents the next network state after the action is performed. The generation of the target action is the basis for subsequent calculation of the target Q-value, providing stable action input for updating the parameters of the main evaluation network.
[0157] After generating the target action, the parameters of the main evaluation network are updated first, followed by the parameter update of the main execution network. Specifically, a delayed update period is set, and the parameter update of the main execution network is only triggered when the parameter update count of the main evaluation network reaches the period. This design provides a convergence time window for the Q-value prediction of the evaluation network, ensuring that the execution network is optimized based on a relatively stable value assessment.
[0158] The noise mechanism in Equation (8) injects appropriate perturbation into the target action, avoiding overfitting caused by the target action being too simple and improving the generalization ability of the model. The delayed update mechanism ensures the stability of the main execution network update, avoiding frequent changes in the execution network strategy due to fluctuations in the evaluation network parameters, and laying the foundation for smooth convergence of the subsequent overall parameter update.
[0159] The parameters of the main execution network are iteratively updated based on the optimized main evaluation network. Specifically, the parameters of the main evaluation network are first optimized by calculating the target Q value and constructing the loss function. Then, the parameters of the main execution network are adjusted by guiding the gradient of the deterministic policy to ensure that the actions generated by the execution network can maximize the Q value.
[0160] First, the parameters of the main evaluation network are optimized. Specifically, based on the target action and the current reward generated above, the target Q value is calculated through a dual-objective evaluation network, and the minimum of the two values is taken to suppress overestimation bias, expressed as: Equation (9)
[0161] in, For the target Q value, For the current reward, As a discount factor, it balances the importance of current and future rewards; For the first A target evaluation subnetwork, Its parameters, The target action is defined. Calculating the target Q-value provides a reliable reference standard for constructing the loss function.
[0162] Subsequently, a loss function is constructed with the objective of minimizing the mean square error between the predicted Q value and the target Q value of the main evaluation network. The expression is: Equation (10).
[0163] in, For the first The loss value of the individual evaluation subnetwork. The parameters of the main evaluation subnetwork are used. The main evaluation network is used to predict the Q-value of the current state-action pair. The buffer serves as an experience replay buffer, and the expected computation is based on batch samples sampled from this buffer. Batch samples are sampled from the experience replay buffer, and the gradient of the loss function with respect to the parameters of the main evaluation network is calculated using the gradient descent algorithm. The parameters are then adjusted along the gradient descent direction until the loss function value converges to a preset range, thus completing the optimization of the main evaluation network parameters.
[0164] After the main evaluation network parameters are optimized, the main execution network parameter update phase will begin. This step uses a dual-Q network with a minimum design to suppress Q-value overestimation, the mean squared loss function to accurately characterize prediction bias, and the gradient descent algorithm to ensure smooth convergence of the main evaluation network parameters, ultimately achieving accurate Q-value prediction and providing a value assessment basis for the optimization of the main execution network.
[0165] The parameters of the main execution network are iteratively updated based on the optimized main evaluation network. Specifically, the parameter adjustment is guided by a deterministic policy gradient to ensure that the actions generated by the execution network can maximize the Q value.
[0166] The first step is setting a delayed update period. Specifically, to avoid interference from fluctuations in the parameters of the main evaluation network on the updates of the execution network, a delayed update period is set. The parameter update of the main execution network is only triggered when the parameter update count of the main evaluation network reaches the set period. This design provides a convergence time window for the Q-value prediction of the evaluation network, ensuring that the execution network is optimized based on a relatively stable value assessment.
[0167] The second step is to calculate the gradient of the deterministic policy. Specifically, based on the Q-value of the optimized main evaluation network output, the parameter gradient of the main execution network is calculated, as shown in equation (11).
[0168] in, To execute the network's policy gradient, The parameters for the main execution network; The gradient of the Q-value with respect to the action reflects the degree to which changes in the action affect the Q-value. The gradient of the action with respect to the network parameters reflects the impact of parameter changes on action generation; the expectation operation is based on state samples sampled from the empirical replay buffer. .
[0169] Finally, the parameters of the main execution network are updated. Specifically, the parameters are adjusted along the ascending direction of the deterministic policy gradient to ensure that the actions generated by the main execution network continuously improve the Q-value. The step size of the parameter update is controlled by the learning rate to avoid action oscillations caused by excessive parameter adjustments, ensuring smooth policy optimization of the execution network. After the main execution network parameter update is completed, the soft update phase of the target network begins.
[0170] The deterministic policy gradient provides a clear optimization direction for updating the execution network parameters, ensuring that the action generation policy remains consistent with the value assessment of the main evaluation network. The delayed update mechanism improves training stability, avoiding frequent changes in the execution network policy due to fluctuations in the evaluation network parameters, and laying the foundation for subsequent soft updates to the target network.
[0171] Based on the updated parameters of the main evaluation network and the main execution network, the target network is softly updated. Specifically, the parameters of the target network are slowly adjusted through a weighted sum mechanism to ensure the stability of the training process.
[0172] The first step is setting the soft update coefficient. Specifically, the soft update coefficient... The value of this coefficient can be set to the range of 0.001 to 0.01. This coefficient directly determines the degree of influence of the main network parameters on the target network parameters. The smaller the coefficient, the smoother the update of the target network parameters and the higher the training stability; the larger the coefficient, the faster the target network parameters follow the main network, but it may introduce the risk of oscillation. This embodiment selects a smaller soft update coefficient to balance training stability and convergence efficiency.
[0173] The second step is the calculation of the weighted sum. Specifically, based on the soft update coefficient, the weighted sum of the parameters of the main evaluation network and the target evaluation network, and the main execution network and the target execution network are calculated respectively, as shown in equation (12).
[0174] Equation (13)
[0175] in, For the updated target evaluation subnetwork parameters, The updated parameters for the main evaluation subnetwork; Execute network parameters for the updated target. These are the updated parameters of the main execution network. The weighted sum formula shows that only a small portion of the new parameters of the target network come from the main network, while the vast majority retain the original parameters, ensuring a smooth parameter change.
[0176] Finally, the target network parameters are updated. Specifically, the two weighted sums are used as the new target evaluation network parameters and target execution network parameters, respectively, overwriting the original parameters and completing the soft update. This process ensures that the target network parameters can slowly keep up with the optimization results of the main network, avoiding sudden parameter changes that could cause oscillations in the target Q-value and target actions, and providing a stable reference standard for the next round of training.
[0177] The soft update mechanism avoids drastic fluctuations in the target network parameters, keeping the target Q-value and target action stable, and providing a reliable basis for the next round of parameter updates in the main evaluation network and the main execution network. The slow adjustment of parameter updates makes the training process converge smoothly, avoiding model oscillations caused by fluctuations in the main network parameters, improving the robustness of the entire training process, and ensuring that the model can continuously learn the optimal policy in a dynamic WiFi 7 network environment.
[0178] In summary, this invention discloses a WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning. Based on a dynamic cell multi-agent reinforcement learning framework and combined with the core characteristics of WiFi 7, it achieves intelligent optimization of access point cell configuration. Compared with existing technologies, it significantly improves network throughput and energy efficiency, reduces transmission latency, and adapts to user and load fluctuations through dynamic access point action strategies, solving the problems of load imbalance and resource waste, ensuring service continuity and spectrum utilization, adapting to high-density network scenarios, and providing highly reliable and high-performance wireless connectivity support for data-intensive applications.
[0179] The specific training and application steps of the aforementioned DC-MARL framework are shown in Figure 13. In this process, the network is first built, then the state of each access point is observed, followed by a training round. In the subsequent training round, it is first determined whether Distributed Multi-Agent Reinforcement Learning (DC-MARL) is enabled. If not enabled, the process sequentially initializes the Q-value metric, processes state-action pairs, evaluates rewards, and updates the Q-table to obtain the optimized value. If enabled, the DC-MARL framework is entered, and it is sequentially determined whether to perform actions such as merging, splitting, resizing, and load sharing. After generating rewards, these are processed by the decision engine, while simultaneously incorporating data from the experience replay buffer.
[0180] Based on the DC-MARL framework, this embodiment can also be combined with the characteristics of WiFi 7 networks. Each access point integrates core WiFi 7 features to optimize transmission performance; these core WiFi 7 features include 4096 quadrature amplitude modulation technology, preamble punching technology, hybrid automatic repeat request technology, and resource unit allocation based on the number of service users.
[0181] It consists of a set of N autonomous access points (APs), labeled as a set. Equipped with DC-MARL optimization and aggressive mode. Each AP can access one or more data sources from the collection. The frequency band.
[0182] For multi-link operation (MLO), MLO supports concurrent transmission across multiple frequency bands, and its basic effective bandwidth formula is: Equation (19)
[0183] in, It is the MLO coordination overhead factor. It represents the percentage of bandwidth loss caused by preamble puncturing. Access point The number of active MLO links.
[0184] To understand how multi-link operation improves throughput, consider the incremental gain when enabling a new MLO link: Equation (20)
[0185] like and The change is small, and the first-order approximation is: Equation (21)
[0186] Furthermore, WiFi 7 introduces a more advanced modulation scheme, 4096-QAM, allowing each symbol to carry higher data density. This paper employs a high-level analytical model for WiFi 7 characteristics. 4096-QAM is represented as a fixed PHY rate improvement relative to the WiFi 6 benchmark; preamble puncturing reduces the effective bandwidth according to the puncturing ratio; and HARQ is modeled as a multiplicative improvement in packet delivery rate (PDR). These approximations allow for comparison of the relative performance of different schemes under consistent assumptions without needing to capture all the protocol details of IEEE 802.11be.
[0187] The modulation gain of 4096-QAM is calculated as follows: Equation (22)
[0188] Therefore, 4096-QAM offers a 20% throughput improvement compared to WiFi 6. Furthermore, preamble punching allows the AP to avoid congested sub-channels within the wide channel, thus the effective bandwidth is: Equation (23)
[0189] in, It is full bandwidth. This is due to the spectral perforation ratio caused by interference. This ensures high throughput even in interference environments, where the effective throughput becomes: Equation (24)
[0190] To improve reliability, Hybrid Automatic Repeat Request (HARQ) is adopted, and retransmission is performed through soft combining when packet transmission fails. This technology can improve packet delivery rate: Equation (25)
[0191] in, This is the baseline packet delivery rate without HARQ. Ensure HARQ is enabled when a packet fails. It is an efficiency gain factor. HARQ can typically improve PDR by 10-15%. In addition to core PHY / MAC improvements, the framework also leverages the enhanced system management features of WiFi 7 to achieve fair resource allocation and energy saving. Furthermore, enhanced OFDMA and Resource Unit (RU) allocation in WiFi 7 improve OFDMA performance through finer RU allocation and multi-user scheduling. If the AP is... If each user is provided with an equal RU allocation, then: Equation (26)
[0192] Therefore, the throughput for each user is: Equation (27)
[0193] The above formula ensures fair resource sharing and avoids starvation. The AP can set a running timer for each user, and users can enter sleep mode at other times. Therefore, the energy efficiency improvement for each user is as follows: Equation (28)
[0194] The above formula calculates the optimal sleep duration for devices within the DC-MARL framework, dynamically adjusting based on current throughput performance, access point load conditions, and cell size characteristics. By allowing longer sleep times when network conditions permit, while maintaining responsive operation during high-demand periods, it intelligently balances energy efficiency and service quality.
[0195] During the resizing and load balancing process, the AP can switch between frequency bands according to coverage area and traffic congestion: Equation (29)
[0196] The path loss is calculated using a frequency band-specific logarithmic propagation path loss model: Equation (30)
[0197] Where 32.4 is the fixed offset of the free space path loss model. It is the operating frequency. It is the exponential path loss factor. Cell size adjustment and optimal MLO band are selected based on path loss.
[0198] Each AP generates a Poisson distributed flow with a mean of λ. Assume that the uplink and downlink flows are uniformly distributed: Equation (30)
[0199] Due to competition and protocol overhead, delivery throughput will decrease by a factor. Therefore, the delivery throughput in each direction is: Equation (31)
[0200] Packet delivery rate depends on interference, link quality, and traffic load. For the DC-MARL scheme, PDR is modeled as: Equation (32)
[0201] in, Indicates the baseline reliability based on interference. Indicates the MLO enhancement factor. Indicates the spatial flow factor. Indicates traffic load. and It is the shape parameter of the Sigmoid.
[0202] The above formulas directly correspond to the performance quantification of WiFi 7's core features. Access points integrate these features and calculate and optimize transmission parameters according to the formulas to improve transmission performance, working in synergy with cell configuration optimization strategies.
[0203] Based on this, a specific simulation experiment was designed to verify the performance of this embodiment. The relevant parameters of the simulation experiment are shown in Table 1.
[0204] Table 1
[0205] The simulation experiment was conducted in a 100m×100m indoor WiFi 7 network environment, deploying 6 randomly distributed access points (APs) supporting the 2.4GHz, 5GHz, and 6GHz frequency bands. Poisson distribution was used to simulate user traffic (3-10 users per AP initially), and traffic load was covered at five levels: 200MHz, 400MHz, 600MHz, 800MHz, and 1000MHz. The experiment used the dual-delay deep deterministic policy gradient (TD3) algorithm to train a dynamic cell multi-agent reinforcement learning (DC-MARL) model, with the actuator learning rate set accordingly. evaluator learning rate The discount factor was 0.95, and the training run consisted of 500 rounds (20 seconds per round). The comparison schemes included the standard WiFi 7 scheme and the aggressive WiFi 7 MLO scheme. The core test metrics were network throughput, packet delivery rate (PDR), transmission latency, training stability, and feature attention distribution.
[0206] The experiment was conducted in three phases. The first phase involved environment deployment and parameter configuration, building a simulation network according to the aforementioned conditions. The second phase involved model training, where the DC-MARL agent learned strategies such as cell merging, splitting, and load sharing through interactive learning. The third phase involved performance testing, collecting performance data for each scheme under different traffic loads and generating corresponding performance curves and feature attention maps.
[0207] The experimental results are shown in Figures 6 to 11.
[0208] Specifically, Figure 6 illustrates the relationship between throughput and training epochs under different traffic loads, showing the throughput trends of the three schemes under low, medium, and high loads. The DC-MARL enhanced scheme converges the fastest (stable within 100 epochs), achieving a throughput exceeding 10Gbps under low load (200MHz), maintaining 33-35Gbps under medium load (400-600MHz), and breaking through 30Gbps under high load (800-1000MHz), with a peak of 38Gbps. The standard WiFi7 scheme achieves a throughput of only 21-25Gbps under high load, while the WiFi7 aggressive MLO scheme achieves approximately 25-30Gbps. The results demonstrate that DC-MARL significantly improves network throughput under different loads through dynamic cell configuration and MLO collaboration.
[0209] Figure 7 illustrates the relationship between Packet Delivery Rate (PDR) and training rounds under different traffic loads, reflecting changes in data transmission reliability. The PDR of the DC-MARL series solutions is consistently higher than that of traditional solutions. Under low load, the PDR of the DC-MARL enhanced solution reaches 99%, maintains 96% under medium load, and remains above 95% under high load; the PDR of the standard WiFi7 solution drops to 87% under high load, while the WiFi7 aggressive MLO solution reaches approximately 90%. DC-MARL's interference avoidance and HARQ collaborative strategy effectively reduces packet loss caused by congestion.
[0210] Figure 8 illustrates the relationship between average latency and training rounds under different traffic loads, demonstrating the optimization effect on transmission timeliness. The DC-MARL solution maintains low latency throughout. It achieves approximately 2.5-3.0ms under low load, 3.5-4.0ms under medium load, and no more than 4.5ms under high load; in contrast, traditional solutions experience latency spikes to 6-15ms under high load. DC-MARL reduces queuing delays and interference through dynamic cell splitting and intelligent frequency band switching.
[0211] Figure 9 illustrates the training metrics of reward, loss, learning rate, and stability, specifically demonstrating the dynamic process of model training. The cumulative reward of the DC-MARL enhancement scheme eventually reaches 800, significantly higher than that of the traditional scheme (below 400). The loss of the double-Q evaluation network converges to [value missing] within 100 rounds. The magnitude of the algorithm avoids overestimation of the Q-value. A controllable decay strategy for the learning rate balances exploration and exploitation; the reward variance is less than 100, demonstrating better stability than traditional schemes (variance > 250), validating the model's training stability and convergence efficiency.
[0212] Figure 10 illustrates the feature attention of the DC-MARL scheme, showing the attention weight distribution of DC-MARL agents. Collaborative features such as neighbor load, effective bandwidth, and MLO enabled status account for over 60%, while energy indicators account for 8-12%. This indicates that the model can focus on the global network state and avoid local optima traps by collaboratively optimizing resource allocation through multi-AP collaboration.
[0213] Figure 11 illustrates the attention characteristics of WiFi 7 solutions. It shows that traditional WiFi 7 solutions focus on self-reference indicators such as their own load and frequency band (accounting for >70%), lacking attention to neighbor status and global collaboration, resulting in rigid resource allocation and difficulty in adapting to dynamic load changes.
[0214] Therefore, it can be concluded that the DC-MARL framework, through multi-agent collaboration and dynamic cell configuration, significantly outperforms traditional WiFi7 solutions in throughput (28%-52% improvement), PDR (5%-8% improvement), transmission latency (25%-30% reduction), and training stability. It can adapt to high-density, high-load network scenarios and fully unleash the hardware performance potential of WiFi7.
[0215] In addition, performance verification under different average traffic loads (Mean Traffic) was set up, as shown in Figure 12. This experiment verified the model's adaptability under different traffic loads of 5-40 Mbps, comparing it with the traditional static strategy and the single-agent strategy. The results show that under light load, the latency of this model is reduced by 12%-15%; under medium load, the throughput is improved by 28%-33% compared with the traditional strategy, and the PDR is maintained above 99.2%; under heavy load, the traditional strategy suffers from severe congestion (PDR 82%), while this model still achieves a throughput of 38.5 Mbps, a PDR of 95.7%, and a latency of less than 22 ms. These results indicate that the model can dynamically adapt to traffic changes, and the congestion mitigation effect is significant under medium and heavy loads.
[0216] An experiment using feature attention was conducted, quantifying the contribution of state features through gradient and permutation methods. The results showed that the access point queue length (32%), user signal-to-noise ratio (28%), and current throughput (21%) were the core features, contributing a total of 81%, while marginal features accounted for only 4%. The feature attention mechanism can focus on core states, reduce redundant interference, improve decision accuracy, and provide a basis for lightweight model optimization.
[0217] The ablation study, using the complete model as a baseline, removed core components to verify the performance impact: without centralized commentators or experience replay, throughput decreased by 21%-27% and training oscillations occurred; without feature attention or soft updates to the target network, the number of convergence epochs increased by 15%-22%, and latency increased by 18%-35%. The experiments confirmed that the synergy of each component is indispensable, validating the rationality of the DC-MARL framework and the MATD3 algorithm architecture. The ablation study results are shown in Table 2.
[0218] Table 2
[0219] A queuing-theoretic interpretation experiment was conducted, mapping access points (APs) to service desks and data packets to customers. The model dynamically configured the corresponding service desk resource reallocation. The experiment showed that the model reduced the average AP queue length by 42%-51%, decreased user waiting time by 38%-45%, and improved service desk utilization balance to 85%. From a queuing-theoretic perspective, this validated the core logic of the solution in addressing the mismatch between service rate and arrival rate, demonstrating the theoretical rationale behind the technical solution.
[0220] As shown in Figure 14, this embodiment of the invention provides a WiFi 7 network access point topology adjustment system based on multi-agent reinforcement learning. The system is applied to a controller and several access points. The controller and the access points are connected to the network. The training system includes: a network state acquisition module 10, an action generation module 20, an action verification and execution module 30, and a post-action network state acquisition module 40.
[0221] Specifically, the network state acquisition module 10 is used to acquire the network state; the action generation module 20 is used to generate a topology adjustment action to be executed based on the network state and through a pre-trained execution network; wherein, the pre-trained execution network is locally deployed at the access point and is trained through multi-agent reinforcement learning; the action verification execution module 30 is used to verify the topology adjustment action to be executed, and if the verification passes, the topology adjustment action to be executed is executed; the post-action network state acquisition module 40 is used to acquire the network state after executing the topology adjustment action.
[0222] Based on the above embodiments, the present invention also provides a terminal device, the principle block diagram of which is shown in Figure 15. The terminal device includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected via a system bus. The processor of the terminal device provides computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the terminal device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning. The display screen of the terminal device can be a liquid crystal display screen or an e-ink display screen. The temperature sensor of the terminal device is pre-installed inside the terminal device for detecting the operating temperature of internal components.
[0223] Those skilled in the art will understand that the principle block diagram shown in Figure 15 is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. A specific terminal device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0224] In one embodiment, a terminal device is provided, including a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs including instructions for performing operations as described in the embodiments of the methods above.
[0225] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0226] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0227] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning, characterized in that, An access point for a WiFi 7 network, the method comprising: acquiring network status; generating a topology adjustment action to be executed based on the network status using a pre-trained execution network; wherein the pre-trained execution network is locally deployed at the access point and is trained through multi-agent reinforcement learning; verifying the topology adjustment action to be executed, and executing the topology adjustment action if the verification passes; and acquiring the network status after executing the topology adjustment action.
2. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, The process of obtaining network status includes: the access point collecting its own operational data and acquiring operational data from neighboring access points through communication links; integrating its own operational data with the operational data from neighboring access points to obtain an initial network status; filtering and normalizing the initial network status to form a final network status; wherein, the network status includes the number of users connected to the access point, current traffic load, remaining energy, average load of neighboring access points, normalized effective bandwidth, multi-link operation enabled status, current cell size, frequency band used, and time normalization factor.
3. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, The topology adjustment action is an access point merging action, which includes: after the local access point and the adjacent access point cooperate, the local access point goes into hibernation and transfers the user.
4. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 3, characterized in that, The verification of the topology adjustment action to be executed includes: based on the network status, obtaining the current traffic load of the local access point, the average load of adjacent access points, and the remaining energy of the local access point; comparing the current traffic load of the local access point with a preset low load threshold, comparing the average load of adjacent access points with a preset coordination threshold, and comparing the remaining energy of the local access point with a preset energy-saving threshold; if all three comparison results satisfy the less than or equal to relationship, the verification result of the merging action is obtained.
5. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 3, characterized in that, The execution of the topology adjustment action to be performed includes: the local access point sending a coordination request to the adjacent access point through multi-link operation and receiving a resource idle response returned by the adjacent access point; generating a user transfer list based on user QoS priority, and smoothly switching the users in the list to the adjacent access point through the links of the multi-link operation; after all users have been transferred, the local access point shuts down the signal transceiver module and enters sleep mode.
6. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, The topology adjustment action is an access point splitting action, which includes: activating the new access point and partitioning the load and coverage of the original access point.
7. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 6, characterized in that, The verification of the topology adjustment action to be executed includes: based on the network status, obtaining the current traffic load, user density, and coverage overlap rate of the local access point; comparing the current traffic load of the local access point with a preset high load threshold, comparing the user density with a preset density threshold, and comparing the coverage overlap rate with a preset overlap threshold; if the load comparison satisfies a greater than or equal to relationship, or the density comparison satisfies a greater than or equal to relationship, and the overlap rate comparison satisfies a less than or equal to relationship, then the verification result of the splitting action is obtained.
8. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 6, characterized in that, The execution of the topology adjustment action to be performed includes: the controller receiving the splitting request of the original access point and activating the preset new access point; the original access point and the new access point dividing the coverage area through frequency band negotiation; the load partitioning of the original access point to the new access point based on user distribution data and QoS requirements; and enabling preamble punching technology to scan and avoid spectrum interference, thereby completing the splitting configuration of load and coverage area.
9. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, The WiFi 7 network includes several access points and a controller. After obtaining the network status after performing the topology adjustment action, the process further includes: each access point reporting the performed topology adjustment action and the network status after performing the topology adjustment action to the controller; the controller obtaining the global network status; and the controller optimizing the execution network of each access point based on the topology adjustment action performed by each access point, the network status after each access point performs the topology adjustment action, and the global network status.
10. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 9, characterized in that, The controller optimizes the execution network of each access point based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state. This optimization includes: optimizing the execution network of each access point through value evaluation using multi-agent reinforcement learning, based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state. The multi-agent reinforcement learning includes the collaboration of centralized evaluation and decentralized execution. The centralized evaluation generates a global value evaluation, and the local execution network of each access point implements the decentralized execution.
11. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 10, characterized in that, The centralized evaluation includes: deploying a centralized evaluation network based on a dynamic cell multi-agent reinforcement learning framework; and calculating the global value assessment of each action by inputting the global network state and the actions of each access point through the evaluation network.
12. The WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning according to claim 11, characterized in that, The optimization of the execution network at each access point is achieved through value evaluation using multi-agent reinforcement learning, based on the topology adjustment actions performed by each access point, the network state after the topology adjustment actions performed by each access point, and the global network state. This includes: global value evaluation based on the output of a centralized evaluation network; updating the parameters of the execution network through temporal differential error calculation; the controller sending the updated execution network parameters to the corresponding access points; and the access points receiving and updating the parameters of their local execution networks.
13. A WiFi 7 network access point topology adjustment system based on multi-agent reinforcement learning, characterized in that, An access point for a WiFi 7 network includes: a network status acquisition module for acquiring network status; an action generation module for generating a topology adjustment action to be executed based on the network status using a pre-trained execution network; wherein the pre-trained execution network is locally deployed at the access point and is trained through multi-agent reinforcement learning; an action verification execution module for verifying the topology adjustment action to be executed, and executing the topology adjustment action if the verification passes; and a post-action network status acquisition module for acquiring the network status after executing the topology adjustment action.
14. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a WiFi 7 network access point topology adjustment program based on multi-agent reinforcement learning stored in the memory and executable on the processor. When the processor executes the WiFi 7 network access point topology adjustment program based on multi-agent reinforcement learning, it implements the steps of the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning as described in any one of claims 1-12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a WiFi 7 network access point topology adjustment program based on multi-agent reinforcement learning. When the WiFi 7 network access point topology adjustment program based on multi-agent reinforcement learning is executed by a processor, it implements the steps of the WiFi 7 network access point topology adjustment method based on multi-agent reinforcement learning as described in any one of claims 1-12.