Multi-agent reinforcement learning-based WiFi (Wireless Fidelity) 7 network access point multi-link activation method and system, terminal equipment and medium

The multi-link activation method for WiFi 7 network access points using multi-agent reinforcement learning solves the problems of load balancing, high energy consumption, and resource waste in WiFi networks, and realizes a solution for efficient network resource optimization and transmission performance, as well as adaptability and network technology applications.

CN121968152APending Publication Date: 2026-05-01深圳开鸿数字产业发展有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
深圳开鸿数字产业发展有限公司
Filing Date
2026-01-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing WiFi networks are inadequate in terms of load balancing, energy efficiency, resource sharing, and adaptability, and cannot adapt to dynamic fluctuations in user density, resulting in high transmission latency, high energy consumption, and resource waste.

Method used

A multi-link activation method for WiFi 7 network access points based on multi-agent reinforcement learning is adopted. By deploying a pre-trained execution network locally at the access point, multi-link activation actions are generated and verified to achieve load balancing and resource optimization. Combined with WiFi 7 multi-link operation, 4096-QAM modulation and preamble punching technology, the network configuration is dynamically adjusted.

Benefits of technology

It improves network transmission and energy efficiency, reduces latency, optimizes resource utilization, adapts to dynamic network environments, and ensures user QoS requirements and network stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121968152A_ABST
    Figure CN121968152A_ABST
Patent Text Reader

Abstract

The invention discloses a WiFi (Wireless Fidelity) 7 network access point multi-link activation method and system based on multi-agent reinforcement learning, terminal equipment and a medium, and relates to the technical field of wireless network configuration. The method comprises the steps that an access point obtains a network state, a to-be-executed multi-link activation action is generated through a locally-deployed pre-training execution network based on the state, and the execution network is obtained through multi-agent reinforcement learning training; and executing the action after the action passes the verification, and then obtaining the network state after the execution. According to the method, the access point locally and autonomously generates the multi-link activation action to adapt to the WiFi (Wireless Fidelity) network characteristics, so that the real-time performance and the adaptability of multi-link activation are improved, the utilization rate of network resources is optimized, and meanwhile, the user service quality and the network operation stability are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless network configuration technology, and in particular to a method, system, terminal device, and medium for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning. Background Technology

[0002] With the explosion of data-intensive applications such as virtual reality and cloud gaming, WiFi networks need to meet the requirements of ultra-high peak throughput and ultra-low latency. WiFi 7 was born to meet this need. It integrates core technologies such as 4096-QAM modulation, multi-link operation (MLO), and preamble punching, providing the hardware foundation for the next generation of wireless communication.

[0003] However, existing WiFi network control strategies still have key limitations. First, their load balancing is inefficient, with some access points overloaded while neighboring devices remain idle during peak hours. Second, their energy utilization is low, with idle access points continuously operating at high power consumption. Furthermore, their static cell configuration cannot adapt to dynamic fluctuations in user density. Simultaneously, their resource sharing capabilities between neighboring access points are limited, resulting in insufficient spectrum utilization. Additionally, they lack autonomous adaptive capabilities, relying on manual or centralized scheduling, leading to lag in response. Even with WiFi 7's advanced hardware features, traditional static control logic still struggles to unleash its potential and cannot meet dynamic adaptation requirements.

[0004] Therefore, there is an urgent need for a dynamic WiFi 7 network access point cell configuration method to fill the gap in existing technology. Summary of the Invention

[0005] The technical problem this invention aims to solve is that, in the field of wireless network configuration technology, existing static access point configuration and inefficient scheduling lead to problems such as load imbalance, resource waste, high transmission latency, and low energy efficiency, which cannot adapt to dynamic network environments and the potential of WiFi 7 features. Therefore, an effective solution is urgently needed to address these technical problems.

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a multi-link activation method for WiFi 7 network access points based on multi-agent reinforcement learning, applicable to WiFi 7 network access points, the method comprising: Get network status; Based on the network state, a multi-link activation action to be executed is generated through a pre-trained execution network; wherein, the pre-trained execution network is locally deployed at the access point, and the pre-trained execution network is trained through multi-agent reinforcement learning; The multi-link activation action to be executed is verified. If the verification passes, the multi-link activation action to be executed is executed. Obtain the network status after executing the multi-link activation action.

[0007] In one implementation, obtaining the network status includes: The access point collects its own operational data and obtains the operational data of adjacent access points through the communication link. By integrating its own operational data with the operational data of adjacent access points, the initial network state is obtained; The initial network states are filtered and normalized, and then combined to form the final network states; The network status includes the number of users connected to the access point, current traffic load, remaining energy, average load of adjacent access points, normalized effective bandwidth, multi-link operation enabled status, current cell size, frequency band used, and time normalization factor.

[0008] In one implementation, the multi-link activation action includes: Enable WiFi 7 multi-link operation, activate transmission links of at least two frequency bands, and optimize transmission throughput, anti-interference ability and transmission stability through multi-band operation.

[0009] In one implementation, the verification of the multi-link activation action to be executed includes: Based on the network status, obtain the single-band throughput, traffic load, channel interference intensity, and multi-link coordination overhead estimate of the local access point. The single-band throughput is compared with a preset demand threshold, the traffic load is compared with a preset activation threshold, the channel interference intensity is compared with a preset interference threshold, and the estimated cooperative overhead is compared with a preset overhead threshold. If the single-band throughput satisfies the relationship of being less than or equal to, the traffic load satisfies the relationship of being greater than or equal to, the channel interference intensity satisfies the relationship of being greater than or equal to, and the estimated cooperative overhead satisfies the relationship of being less than or equal to, then the verification result of the multi-link activation action is obtained.

[0010] In one implementation, executing the multi-link activation action to be executed includes: The local access point selects the appropriate frequency band combination based on the network status and initiates a multi-link collaborative configuration request; Configure the synchronization parameters for each link; A multi-link data transmission and reception mechanism is enabled, and data transmission is completed by integrating WiFi 7 transmission features, while simultaneously updating the multi-link activation status to adjacent access points.

[0011] In one implementation, executing the multi-link activation action to be executed further includes: Data transmission is completed by integrating WiFi 7 transmission features; these integrated WiFi 7 transmission features include: Integrated 4096 quadrature amplitude modulation technology improves transmission rate; Use preamble punching technology to avoid frequency band interference; Data transmission reliability is ensured through hybrid automatic repeat request technology.

[0012] In one implementation, executing the multi-link activation action to be executed further includes: Based on the number of users connected to the access point, resource units of multiple links are evenly allocated to ensure that the QoS requirements of each user meet the preset standards.

[0013] In one implementation, the WiFi 7 network includes several access points and a controller. After obtaining the network status after performing the multi-link activation action, the system further includes: Each access point will report the multi-link activation action it performs and the network status after performing the multi-link activation action to the controller; The controller acquires the global network status; The controller optimizes the execution network of each access point based on the multi-link activation actions performed by each access point, the network status after each access point performs multi-link activation actions, and the global network status.

[0014] In one implementation, after the execution network generates a multi-link activation action, it receives the global collaborative value assessment result of the controller's centralized evaluation network and performs secondary verification to avoid multi-access point multi-link activation conflicts.

[0015] In one implementation, the controller optimizes the execution network of each access point based on the multi-link activation actions performed by each access point, the network state after each access point performs the multi-link activation actions, and the global network state, including: Based on the multi-link activation actions performed by each access point, the network state after each access point performs multi-link activation actions, and the global network state, the network execution of each access point is optimized through the value evaluation of multi-agent reinforcement learning. The multi-agent reinforcement learning includes the collaboration of centralized evaluation and decentralized execution. The centralized evaluation generates a global value assessment, and the decentralized execution is implemented by the local execution network of each access point.

[0016] In one implementation, the centralized evaluation includes: A centralized evaluation network is deployed based on a dynamic community multi-agent reinforcement learning framework. By using the evaluation network, the global network state and the actions of each access point are input, and the global value assessment of each action is calculated.

[0017] In one implementation, the optimization of the network execution at each access point based on the multi-link activation actions performed by each access point, the network state after each access point performs the multi-link activation actions, and the global network state, through value evaluation using multi-agent reinforcement learning, includes: Global value assessment based on the output of a centralized evaluation network; The parameters of the execution network are updated by calculating the temporal difference error. The controller sends the updated execution network parameters to the corresponding access point; The access point receives and updates the parameters of the local network.

[0018] Secondly, embodiments of the present invention also provide a WiFi 7 network access point multi-link activation system based on multi-agent reinforcement learning, the system comprising: The network status acquisition module is used to acquire network status. An action generation module is used to generate multi-link activation actions to be executed based on the network state and through a pre-trained execution network; wherein the pre-trained execution network is locally deployed at the access point and is trained through multi-agent reinforcement learning; The action verification and execution module is used to verify the multi-link activation action to be executed. If the verification passes, the multi-link activation action to be executed is executed. The post-action network status acquisition module is used to acquire the network status after executing the multi-link activation action.

[0019] Thirdly, embodiments of the present invention also provide a terminal device, the terminal device including a memory, a processor, and a WiFi 7 network access point multi-link activation program based on multi-agent reinforcement learning stored in the memory and executable on the processor. When the processor executes the WiFi 7 network access point multi-link activation program based on multi-agent reinforcement learning, it implements the steps of the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning as described in any of the above schemes.

[0020] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a WiFi 7 network access point multi-link activation program based on multi-agent reinforcement learning. When the WiFi 7 network access point multi-link activation program based on multi-agent reinforcement learning is executed by a processor, it implements the steps of the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning as described in any of the above schemes.

[0021] Beneficial Effects: This invention discloses a method, system, terminal device, and medium for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning, relating to the field of wireless network configuration technology and applied to WiFi 7 network access points. The method first acquires the network state and, based on this state, generates multi-link activation actions to be executed through a pre-trained execution network. The pre-trained execution network is locally deployed at the access point and is trained using multi-agent reinforcement learning. Subsequently, the multi-link activation actions to be executed are verified; if the verification passes, the actions are executed. Finally, the network state after executing the multi-link activation actions is obtained. This invention utilizes a pre-trained execution network deployed locally at the access point and autonomously generates multi-link activation actions based on multi-agent reinforcement learning. Decentralized execution reduces transmission latency and improves real-time adjustment. Simultaneously, the action verification step ensures the rationality of execution and adapts to the multi-link and anti-interference characteristics of WiFi 7 networks. Furthermore, it can optimize access point resource utilization and coverage balance, reducing energy consumption. In addition, the controller also optimizes the network globally to continuously improve the adaptability of multi-link activation, ensuring user QoS requirements and network operation stability. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a specific implementation method for a multi-link activation method for WiFi 7 network access points based on multi-agent reinforcement learning, as provided in an embodiment of the present invention.

[0023] Figure 2 This is a schematic diagram of the network architecture of the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning provided in an embodiment of the present invention.

[0024] Figure 3 This is a schematic diagram of cell merging for a WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning, provided in an embodiment of the present invention.

[0025] Figure 4 This is a schematic diagram of cell splitting for a WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning, provided in an embodiment of the present invention.

[0026] Figure 5 This is a schematic diagram illustrating resource sharing in a multi-link activation method for WiFi 7 network access points based on multi-agent reinforcement learning, provided in an embodiment of the present invention.

[0027] Figure 6 The graph shows the relationship between throughput and training rounds under different traffic loads for the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning provided in this embodiment of the invention.

[0028] Figure 7 The graph shows the relationship between packet delivery rate and training rounds under different traffic loads for the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning provided in this embodiment of the invention.

[0029] Figure 8 The graph shows the relationship between average latency and training rounds under different traffic loads for the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning provided in this embodiment of the invention.

[0030] Figure 9 The training metrics diagram for the reward, loss, learning rate, and stability of the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning provided in the embodiments of the present invention.

[0031] Figure 10 The feature attention map is provided in the DC-MARL scheme of the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning, which is provided in the embodiments of the present invention.

[0032] Figure 11 This is a feature attention map of the WiFi 7 scheme based on the multi-agent reinforcement learning-based WiFi 7 network access point multi-link activation method provided in the embodiments of the present invention.

[0033] Figure 12 The average traffic load heatmap of the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning provided in the embodiments of the present invention.

[0034] Figure 13 This is a schematic diagram of the multi-link activation and application process of WiFi 7 network access points based on multi-agent reinforcement learning, provided in an embodiment of the present invention.

[0035] Figure 14 This is a schematic diagram of a WiFi 7 network access point cell configuration device based on dynamic cell multi-agent reinforcement learning, provided in an embodiment of the present invention.

[0036] Figure 15 This is a block diagram illustrating the internal structure of the terminal device provided in an embodiment of the present invention. Detailed Implementation

[0037] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0038] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content, operations, or steps, nor does it require execution in the described order. For example, some operations or steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0039] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0040] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. For example, "first control information" and "second control information" are only used to distinguish different control information and do not limit their order.

[0041] Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or the order of execution, and that the words "first" and "second" do not necessarily imply that they are different.

[0042] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0043] As a core infrastructure for global connectivity, Wireless Local Area Networks (WLANs) not only provide seamless internet access for mobile devices in high-density urban scenarios but also supplement cellular network coverage blind spots, ensuring uninterrupted access to core services, online resources, and digital platforms. With the development of the digital economy, the explosive growth of data-intensive applications has placed new demands on WiFi network performance. Applications such as high-definition video, virtual reality, augmented reality, cloud gaming, remote work, and wireless screen projection not only require ultra-high-speed transmission capabilities but also have extremely low latency requirements. Furthermore, professional scenarios such as industrial automation and telemedicine also have high requirements for data transmission reliability to replace wired connections. To address these challenges, the next-generation WiFi standard, WiFi 7, has been proposed. Its goal is to support higher peak throughput and optimize latency and jitter under worst-case conditions, meeting the performance requirements of future applications through new Physical Layer (PHY) and Medium Access Control (MAC) technologies.

[0044] As a new wireless communication standard, WiFi 7 integrates several key innovative technologies, significantly improving network performance. First, 4096-QAM technology, compared to WiFi 6's 1024-QAM, encodes 12 bits of data per symbol, improving spectral efficiency by approximately 20%. Second, Multi-Link Operation (MLO) allows devices to simultaneously establish transmission links in multiple frequency bands (2.4GHz, 5GHz, and 6GHz), achieving higher throughput, interference-resistant handover, and concurrent transmission and reception through aggregation, handover, and duplex modes. Third, the Enhanced Multi-Resource Units (MRU) mechanism dynamically combines discontinuous resource units, avoiding bandwidth waste caused by spectrum fragmentation and improving throughput in congested environments. Fourth, Preamble Puncturing technology selectively disables occupied subcarriers, enabling high-throughput transmission using remaining spectrum in complex interference environments and reducing bandwidth waste. Fifth, Advanced Hybrid ARQ (HARQ) improves Packet Delivery Rate (PDR) by combining retransmission with soft combining. In addition, WiFi 7 incorporates features such as Time-Sensitive Networking (TSN) and Restricted Target Wake Time (R-TWT) to further optimize deterministic latency and energy efficiency.

[0045] Despite the advanced technological potential of WiFi 7, existing networks still have many key limitations that make it difficult to fully realize its performance advantages. Traditional WiFi systems lack a sophisticated dynamic load balancing mechanism. During peak hours, some access points (APs) become overloaded and congested, while others remain idle, resulting in poor overall network performance. Energy efficiency is low, with underutilized APs continuing to consume significant amounts of power, leading to unnecessary energy waste. Static cell size configurations cannot adapt to dynamically changing user densities, causing severe congestion in high-density areas and insufficient resource utilization in sparse areas. Resource sharing between adjacent APs is limited, resulting in low bandwidth utilization and significantly increased latency under peak load. Most existing systems rely on centralized control or manual intervention for configuration adjustments, lacking autonomous adaptability. This fails to meet the scalability requirements of large-scale networks and cannot quickly respond to dynamic network conditions such as user movement, traffic fluctuations, and interference changes. Furthermore, existing solutions lack sufficient integration of WiFi 7's core features, failing to leverage intelligent decision-making mechanisms to coordinate technologies such as 4096-QAM, MLO, and MRU. This makes it difficult to translate the theoretical performance of WiFi 7 into stable gains in practical applications, failing to fully meet the comprehensive requirements of high throughput, low latency, high reliability, and low energy consumption.

[0046] It is understandable that existing WiFi 7 networks suffer from problems such as low load balancing efficiency, unreasonable energy consumption control, lack of dynamic adaptability in cell configuration, insufficient resource sharing capabilities, lack of self-adaptation and scalability, insufficient QoS (Quality of Service) guarantee capabilities, and insufficient collaborative utilization of WiFi 7 features.

[0047] Therefore, this embodiment models WiFi 7 dynamic cell control as a multi-agent Markov Decision Process (MDP) integrating real-time user load, interference, and energy indicators. To achieve optimal decision-making in this MDP, this embodiment proposes a Dynamic Cell Multi-Agent Reinforcement Learning (DC-MARL) framework. This framework is based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm and extends it to form a core policy optimization algorithm (MATD3) adapted to multi-agent collaborative scenarios. As a policy optimization algorithm under the DC-MARL architecture, MATD3 has a dual Q-evaluation network and an adaptive target update mechanism, which can effectively eliminate the Q-value overestimation problem and ensure stable policy convergence in a dynamic network environment with multiple APs.

[0048] The network architecture of this embodiment is as follows: Figure 2 As shown, it includes a controller and several access points, with each access point responsible for the network connection of several user devices.

[0049] Within this DC-MARL framework, four strategies can be used: cell merging, cell splitting, cell adjustment, and access point coordination, as follows: Figure 3 , Figure 4 , Figure 5 , Figure 6 As shown, intelligent decisions can be made based on network traffic. Furthermore, WiFi 7's unique physical layer and media access control layer mechanisms, such as 4096 quadrature amplitude modulation, multi-link operation, and preamble punching, are integrated into the learning environment to achieve realistic performance evaluation. Finally, extensive simulation experiments demonstrate that the proposed framework significantly outperforms WiFi 7 benchmark solutions in terms of cumulative reward, throughput, latency, and energy efficiency.

[0050] This embodiment provides a multi-link activation method for WiFi 7 network access points based on multi-agent reinforcement learning, applied to WiFi 7 network access points, such as... Figure 1 As shown, the method specifically includes the following steps: Step S100: Obtain network status.

[0051] In this embodiment, the access point (AP) is the core device in the WiFi 7 network that provides wireless access services to user terminals. It has the functions of signal transmission and reception, data forwarding and status awareness. It can be deployed in high-density scenarios such as shopping malls and office buildings to directly establish a wireless connection with user terminals and transmit data.

[0052] Network status is a comprehensive data set reflecting the operational status of an access point and the characteristics of its surrounding environment. It covers key indicators such as the number of connected users, traffic load, and remaining energy, comprehensively depicting the operational status of the access point and the dynamic changes in the network environment. Actions are configuration optimization operations that the access point can perform, which may include merging, splitting, load sharing, resizing, and enabling Multi-Link Operation (MLO). These actions directly affect the cell configuration and resource allocation of the access point and are means to optimize network performance.

[0053] In one implementation, obtaining the network status specifically includes the following steps: Step S110: The access point collects its own operating data and obtains the operating data of adjacent access points through the communication link; Step S120: Integrate its own operating data with the operating data of adjacent access points to obtain the initial network state; Step S130: Filter and normalize the initial network state, and combine them to form the final network state; The network status includes the number of users connected to the access point, current traffic load, remaining energy, average load of adjacent access points, normalized effective bandwidth, multi-link operation enabled status, current cell size, frequency band used, and time normalization factor.

[0054] In this embodiment, obtaining the network status of each access point includes three core steps: collection, integration, filtering, and normalization, ultimately forming a standardized network status vector that can support action generation.

[0055] The first step is data acquisition. Specifically, each access point collects its own operational data through built-in sensors and communication modules, including the number of connected users, current traffic load, and remaining energy. Simultaneously, it acquires operational data from neighboring access points via communication links, such as the MLO link in WiFi 7, including the average load of neighboring access points and the status of multi-link operation. This ensures that access points can not only perceive their own status but also obtain information about the surrounding network environment, providing data support for generating collaborative actions.

[0056] Secondly, there is data integration. Specifically, the access point integrates its own operational data with the operational data of neighboring access points to form the initial network state. The dimensions of the initial network state cover the core indicators of access point operation and are represented using a network state vector: Equation (1)

[0057] in, Indicates the first Each access point at time The local state vector, Given a 9-dimensional real number space, the definitions and physical meanings of the parameters in each dimension are as follows: This indicates the number of users currently associated with this access point, reflecting the service pressure on the access point. This represents the average user load of adjacent access points, obtained through real-time data interaction with surrounding APs, and is used to estimate the congestion of the surrounding network. This represents the normalized energy consumption of the access point, with a value range of [0,1]. It is calculated from the ratio of the current remaining energy of the access point to the maximum rated energy, reflecting the energy consumption status of the equipment. This represents the normalized coverage radius of the access point, with a value range of [0,1]. It is adjusted in real time by dynamic cell operations and corresponds to the standardized result of the actual physical coverage area. This represents a binary frequency band indicator, with a value of 0 or 1, used to specify the frequency band currently operating at the access point; Indicates the multi-link operation (MLO) status, 0 for disabled and 1 for enabled, which is associated with the multi-band concurrent transmission capability of WiFi 7; Represents traffic intensity, normalized to [0,1], calculated from the ratio of the data packet arrival rate to the service rate of the current access point, reflecting the level of transmission busyness; This represents the average packet delay, normalized to [0,1]. It is obtained by statistically analyzing the average time from the generation to the reception of data packets per unit time, reflecting the transmission timeliness. This represents the normalized time period characteristics, with a value range of [0,1]. It is calculated by the ratio of the current time to the total time period of the day and is used to capture the daily variation patterns of user demand and network behavior, such as the difference between peak and off-peak periods.

[0058] When the access point collects its own operational data, it obtains real-time data through hardware units such as the built-in user connection detection module, traffic statistics module, and energy consumption monitoring module. , , It acquires its own status parameters. When obtaining operational data from adjacent access points, it uses a dedicated communication link between the access points to achieve this. , Real-time sharing of status information of neighboring nodes.

[0059] Finally, there is the filtering and normalization process. Specifically, the initial network status of the access points is filtered to remove outliers, such as data exceeding reasonable ranges due to sensor malfunctions. Then, each indicator is normalized, mapping all data to the [0,1] interval to eliminate numerical magnitude differences caused by different units and avoid affecting the decision-making accuracy of subsequent controllers. For example, the number of users is normalized to a reasonable range of "0-50 people", and the latency is normalized to a range of "0-10ms".

[0060] Following the dimensional order of equation (1), the preprocessed parameters are combined to form the local state vector. This allows the vectors to comprehensively and accurately reflect the real-time status of the access point itself and the surrounding local network, providing high-quality basic data for the controller to construct the global network status.

[0061] The comprehensiveness of data acquisition ensures that the network state vector can fully characterize the access point itself and its surrounding environment, providing a foundation for accurate action generation. The standardized state vector design of Equation (1) unifies the data structure and adapts to the input requirements of the execution network. Filtering and normalization processes improve data quality, avoid interference from outliers and magnitude differences in training, and accelerate model convergence. At the same time, the design of multi-dimensional state vectors not only covers the core characteristics of WiFi7 networks but also takes into account network performance and device status, enabling decisions on diverse actions such as merging and splitting.

[0062] Step S200: Based on the network state, generate multi-link activation actions to be executed through a pre-trained execution network; wherein, the pre-trained execution network is locally deployed at the access point, and the pre-trained execution network is trained through multi-agent reinforcement learning.

[0063] In this embodiment, a decentralized action generation mechanism involving multiple agents is proposed. Each access point independently generates actions based on locally collected network conditions, without waiting for global scheduling instructions from the controller. This mechanism is well-suited to the widely distributed and highly dynamic characteristics of WiFi 7 network access points. Specifically, decentralized execution reduces communication overhead between the controller and access points, minimizes action generation latency, and ensures that access points can respond in real-time to changes in local network conditions, such as sudden user growth. Independent exploration by each access point enriches the diversity of experience data, avoids the limitations of local exploration by a single access point, and makes the overall network configuration strategy more adaptable. Furthermore, the diversified design of action types covers core scenarios of multi-link activation and resource scheduling in cells, and can specifically address key issues in traditional networks such as load imbalance and resource waste.

[0064] The execution network is a decentralized neural network deployed at each access point. Its core function is to generate adapted configuration actions based on the network state collected locally at the access point. Its input is a network state vector, and its output is a mixed action vector containing discrete action log probabilities and continuous adjustment amounts.

[0065] In one implementation, the multi-link activation action specifically includes the following steps: Step S210: Enable WiFi 7 multi-link operation, activate transmission links of at least two frequency bands, and optimize transmission throughput, anti-interference ability and transmission stability through multi-band.

[0066] In this embodiment, the multi-link activation action is a load balancing method adapted to the core feature of WiFi 7: multi-link operation. By allowing access points to autonomously enable or expand multi-band concurrent transmission links, it aggregates 2.4GHz, 5GHz, and 6GHz frequency band resources, resolving load bottlenecks caused by insufficient bandwidth or poor channel quality of a single link, and achieving a reasonable distribution of traffic load among multiple links. This action determines the necessity of link expansion based on real-time network status, and by coordinating multiple links to improve transmission capacity and reliability, it can unleash the hardware potential of the WiFi 7 network.

[0067] The triggering of multi-link activation actions is generated by the execution network deployed locally at the access point. The execution network extracts key features from the local state vector, including the current traffic load, the single-link channel quality indirectly reflected by the normalized effective bandwidth, the link occupancy status of adjacent access points, and the current enabled status of multi-link operations. The execution network uses a trained strategy to fuse and analyze these features, generating the logarithmic probability of multi-link activation actions. After probability transformation calculations, it selects the action with the highest probability to initiate multi-link activation.

[0068] During the execution of the operation, the access point, in coordination with the WiFi 7 multi-link operation mechanism, completes link aggregation. First, based on the current frequency band occupancy and channel quality, the optimal link combination is selected. For example, in high-load scenarios, the 5GHz and 6GHz dual links are activated first, leveraging the wide bandwidth and low interference of the 6GHz band to improve transmission rates. When channel interference is high, the 2.4GHz and 5GHz links can be activated, enhancing anti-interference capabilities through link diversity. After link activation, the access point integrates resources through the multi-link operation coordination mechanism. Effective throughput requires comprehensive consideration of multi-link coordination overhead, bandwidth loss due to preamble puncturing, and the basic physical layer rate of each link. The basic rate of each link is determined by the modulation and coding scheme and the number of spatial streams. The throughput gain brought by the added link is directly related to the link's basic rate, multi-link coordination efficiency, and puncturing loss. The executing network will learn and prioritize activating the link combination with the highest gain.

[0069] Multi-link activation significantly improves throughput. Through multi-link bandwidth aggregation, the effective transmission rate of the access point can approach the sum of the rates of each link, which is particularly outstanding in high-load scenarios. It also effectively enhances transmission reliability. The multi-link redundancy design allows for rapid switching to other links when one link is interfered with. Combined with hybrid automatic repeat request technology, it reduces packet loss. Furthermore, it optimizes load balancing adaptability. Multi-links provide more dimensions for traffic allocation, dynamically scheduling traffic based on the load of different links to avoid overloading a single link. For example, in high-bandwidth streaming scenarios, the access point can meet transmission demands by activating dual links, while link redundancy significantly reduces the risk of transmission interruption.

[0070] Step S300: Verify the multi-link activation action to be executed. If the verification passes, execute the multi-link activation action to be executed.

[0071] In this embodiment, the actions of each access point generated by the network are executed. The core is to generate a hybrid action vector and transform it into a specific action, thereby achieving the synergy between discrete multi-link activation and continuous parameter optimization.

[0072] The first step is the generation of the hybrid action vector. Specifically, the main execution network of each access point calculates and outputs the hybrid action vector through forward propagation based on the network state vector in equation (1). The dimension of the hybrid action vector can be adjusted based on the optional actions. Here, the optional actions are listed and represented using a 6-dimensional vector: Equation (2)

[0073] Among them, the first 5 dimensions The logits represent the discrete actions, corresponding to five preset discrete topology operations: Merge, Split, Share Load, Resize, and Enable MLO; the sixth dimension... This is the adjustment amount for the continuous cell radius, and its value range is... , This is the preset maximum radius adjustment range.

[0074] Subsequently, actions are calculated and executed based on the hybrid action vector.

[0075] The hybrid action vector design combines discrete multi-link activation with continuous parameter optimization. It can solve the load imbalance problem at the network topology level through actions such as merging and splitting, and optimize the coverage range through continuous radius adjustment, thus adapting to the complex optimization needs of WiFi 7 networks.

[0076] Actions are calculated and executed based on hybrid action vectors. Specifically, the softmax operation and argmax function are used to accurately select discrete actions, ensuring the rationality and efficiency of action execution.

[0077] First, there's the softmax operation. Specifically, after the mixed action vector is generated, it's transformed into specific executable actions through subsequent processing. First, a softmax operation is performed on the log-probability of the first 5 discrete actions, converting it into a selection probability between 0 and 1, expressed as: Equation (3)

[0078] in, For the first The probability of choosing a discrete action is given, and the sum of all probabilities is 1. A softmax operation is performed on the 5-dimensional discrete action log odds in the mixed action vector to transform the unnormalized log odds into a probability distribution between 0 and 1, with the sum of the probabilities of all discrete actions being 1. The log odds themselves are unbounded real numbers; they are mapped to positive numbers using an exponential function, and then divided by the sum of all exponents to achieve probability normalization. For example, if the log odds of splitting the action in the mixed action vector are 2.1, the log odds of merging the action are 1.3, and the log odds of other actions are all less than 1, then after the softmax operation, the probability of splitting the action will be significantly higher than that of other actions, reflecting its optimality in the current state.

[0079] Then, the discrete action with the highest probability is selected using the argmax function, represented as: Equation (4)

[0080] If the selected discrete action is to adjust the size, that is... Then combine with continuous adjustment amount Perform cell radius adjustment; for other discrete actions, directly execute the corresponding topology operation, such as merging or splitting. The argmax function iterates through the five action probabilities obtained from the softmax operation and selects the discrete action with the highest probability as the final action to be executed. This ensures that the access point can choose the action that is most likely to optimize network performance in the current state, avoiding the execution of invalid actions. For example, when the probability of the load sharing action is 0.6, and the probabilities of other actions are all less than 0.2, the argmax function will directly select the load sharing action, guiding the access point to transfer some of the load to adjacent access points, alleviating its own congestion.

[0081] The softmax operation provides a probabilistic basis for action selection, making it flexible and avoiding overexploration of single actions. It also provides probabilistic interpretability, preventing insufficient exploration caused by deterministic choices. The argmax function ensures the selection of the highest-priority action, guaranteeing targeted configuration and improving efficiency. This process is closely integrated with the execution network's output, achieving a smooth transformation from action vectors to specific actions, ensuring the precise implementation of the network's optimization goals. For example, when the execution network learns during training that splitting actions is better under high load, it will output a higher log-probability of splitting actions. After softmax and argmax processing, the access point will prioritize splitting actions, optimizing network performance. This design allows access points to flexibly select action types based on network conditions, such as prioritizing splitting actions under high load and merging actions under low load and high energy consumption, achieving dynamic adaptive configuration.

[0082] In one implementation, the verification of the multi-link activation action to be executed specifically includes the following steps: Step S310: Based on the network status, obtain the single-band throughput, traffic load, channel interference intensity, and multi-link cooperative overhead estimate of the local access point; Step S320: Compare the single-band throughput with a preset demand threshold, compare the traffic load with a preset activation threshold, compare the channel interference intensity with a preset interference threshold, and compare the estimated cooperative overhead with a preset overhead threshold. Step S330: If the single-band throughput satisfies the relationship of being less than or equal to, the traffic load satisfies the relationship of being greater than or equal to, the channel interference intensity satisfies the relationship of being greater than or equal to, and the estimated cooperative overhead satisfies the relationship of being less than or equal to, then the result of the verification of the multi-link activation action is obtained.

[0083] In this embodiment, the verification of multi-link activation actions ensures the necessity and feasibility of link expansion, avoiding resource waste or a surge in coordination overhead caused by blind activation. The verification process is based on parameters in the network state vector, and through quantitative threshold comparison, it filters out scenarios that truly require multi-link support, ensuring the rationality and efficiency of action execution.

[0084] Specifically, the verification extracts three parameters from the network state vector: the current traffic load of the local access point, the corresponding normalized traffic intensity, which reflects the current single link's carrying pressure; the channel quality of the target link, which is converted into normalized effective bandwidth through the link's signal-to-interference-plus-noise ratio to ensure that the activated link has usable transmission quality; and the link's idle resources, i.e., the proportion of unoccupied resource units in the target frequency band, which reflects the physical resource redundancy of the link extension.

[0085] The current traffic load is compared with a preset overload threshold. Only when the current traffic load reaches or exceeds the overload threshold does it indicate that the single link can no longer handle the current traffic, and link expansion is necessary. The normalized effective bandwidth of the target link is compared with a preset quality threshold. The normalized effective bandwidth of the target link must reach or exceed the quality threshold to ensure that the activated link can provide stable transmission services and avoid increased transmission error rates due to poor link quality. The proportion of idle resources on the link is compared with a preset redundancy threshold. The proportion of idle resources must reach or exceed the redundancy threshold to ensure that there are sufficient physical resources to support data transmission after the link is activated, and to avoid serious interference between the new link and existing links.

[0086] The verification result is determined by the condition that all three comparisons above are met before the multi-link activation action passes the verification. This forms a triple constraint of demand, quality, and resources, ensuring the targeted execution of actions while avoiding the risk of invalid activation. For example, if the current traffic load of the access point does not reach the overload threshold, but the target link quality is excellent, the verification will still fail, avoiding unnecessary link activation that increases the coordination overhead of multi-link operations; if the traffic load meets the standard but the target link has insufficient idle resources, the verification will also fail, preventing the addition of new links from exacerbating spectrum interference.

[0087] The verification process improves resource utilization efficiency by filtering invalid activation scenarios through threshold constraints. This allows multi-link operation resources to be concentrated on high-demand, high-quality, and high-redundancy scenarios, significantly increasing the effective success rate of multi-link activation while reducing system overhead. It minimizes unnecessary link negotiation and synchronization operations, allowing access point computing resources to be used more for data transmission and avoiding fragmented spectrum resource consumption. Furthermore, the verification process is fast and efficient, completed based on local state vectors without cross-access point coordination, with response latency controlled within milliseconds, adapting to the dynamic changes in WiFi 7 network traffic.

[0088] In one implementation, executing the multi-link activation action to be executed specifically includes the following steps: Step S340: The local access point selects an appropriate frequency band combination based on the network status and initiates a multi-link collaborative configuration request; Step S350: Configure the synchronization parameters for each link; Step S360: Enable the multi-link data transmission and reception mechanism, and complete data transmission by integrating WiFi 7 transmission features, and synchronously update the multi-link activation status to adjacent access points.

[0089] In this embodiment, the execution process of the multi-link activation action aims at seamless activation, efficient aggregation, and controllable interference. It deeply integrates the multi-link operation, preamble punching, and multi-resource unit features of WiFi 7, and completes link expansion and load balancing configuration in four steps to ensure that the network performance potential is quickly released after the action is executed.

[0090] The first step is link negotiation and activation. The local access point, through a multi-link operation link discovery mechanism, identifies available target frequency bands in the vicinity and sends a link activation request to the network controller. This request includes information such as current traffic characteristics, target link channel quality, and expected bandwidth requirements. Based on global spectrum occupancy, the controller returns activation permission and link configuration parameters to the access point, including modulation and coding scheme level, number of spatial streams, and resource unit allocation scheme. Upon receiving the permission, the access point initiates physical layer synchronization of the target link. Through WiFi 7's fast link establishment mechanism, link activation is completed in a very short time, ensuring seamless service for the user.

[0091] Following this is bandwidth aggregation and traffic scheduling. After a link is activated, the access point, based on a multi-link operation aggregation mode, allocates current traffic to various links according to service type and link characteristics. High-priority services are preferentially allocated to the 6GHz link with excellent channel quality, utilizing its low latency and high bandwidth characteristics; medium- and low-priority services are allocated to 5GHz or 2.4GHz links, achieving reasonable traffic distribution. Traffic scheduling aims to maximize total effective throughput, monitoring the load of each link in real time and dynamically adjusting the traffic allocation ratio to ensure that the load difference between links is controlled within a reasonable range, avoiding throughput reduction caused by single-link overload.

[0092] Following this is interference avoidance and spectrum optimization. The access point utilizes WiFi 7 preamble punching technology to scan for interference in the frequency band where the active link is located, marks occupied subcarriers, and disables these subcarriers through the punching mechanism, using only idle subcarriers for transmission, thus improving effective bandwidth utilization. Simultaneously, the access point dynamically combines discontinuous resource units through a multi-resource unit mechanism to avoid bandwidth waste caused by spectrum fragmentation, further improving spectrum utilization efficiency.

[0093] Following this is state synchronization and optimization feedback. After link activation and configuration, the access point synchronizes the updated multi-link status to the controller and adjacent access points, including information such as the number of activated links, the load of each link, and effective bandwidth. The controller adjusts the transmission parameters of surrounding access points based on the global status to avoid link interference between access points; adjacent access points optimize their own load distribution strategies based on the synchronization information, forming a regional collaborative load balancing effect. The access point continuously monitors the transmission performance of multiple links and generates optimization feedback periodically. If the channel quality of a link deteriorates or the load is too high, dynamic traffic rescheduling is triggered to ensure that multiple links are always in optimal operating condition.

[0094] The execution process enables seamless activation, with rapid link establishment and traffic scheduling resulting in extremely short service interruptions that are imperceptible to users. It also significantly improves throughput, offering substantial improvements over single-link activation in dual- or triple-link scenarios. Furthermore, it ensures effective interference control. Preamble punching and multi-resource unit mechanisms keep interference levels within a reasonable range after link activation, guaranteeing overall network stability. For example, in high-density scenarios, access points activating 5GHz and 6GHz dual links through this process significantly improve downlink throughput for single users. Simultaneously, due to effective interference control, the performance of surrounding access points experiences only a slight decrease, achieving overall regional performance optimization.

[0095] In one implementation, executing the multi-link activation action to be executed further includes the following steps: Step S370: Integrate WiFi 7 transmission features to complete data transmission; wherein, integrating WiFi 7 transmission features includes: Step S371: Integrate 4096 quadrature amplitude modulation technology to improve transmission rate; Step S372: Activate preamble punching technology to avoid frequency band interference; Step S373: Ensure data transmission reliability through hybrid automatic repeat request technology.

[0096] In this embodiment, integrating WiFi 7 transmission characteristics to complete data transmission is the core step connecting link activation and traffic scheduling during the multi-link activation process. By fully utilizing the transmission technology advantages of the WiFi 7 standard, 4096 quadrature amplitude modulation technology, preamble punching technology, and hybrid automatic repeat request technology are integrated with the activated multi-link resources. This improves the multi-link data transmission rate while effectively avoiding frequency band interference, ensuring data transmission reliability, and ensuring that the transmission performance after multi-link activation is fully released to meet the service quality requirements of various services.

[0097] Specifically, the first step is to integrate 4096 Quadrature Amplitude Modulation (4096-QAM) technology to improve transmission rates, which is one of the core rate-enhancing technologies of the WiFi 7 standard compared to its predecessors. By carrying more bits of information on a single carrier symbol, it improves the bit transmission efficiency per unit of spectrum resources. In multi-link active scenarios, adaptive integration is performed for different activated frequency band links. The access point first performs real-time detection of the channel quality of each active link to determine whether the current signal-to-interference-plus-noise ratio (SNR) of the link meets the application conditions of 4096-QAM technology. For links with excellent channel quality, especially 6GHz band links, 4096-QAM technology is prioritized; for 5GHz links with average channel quality but meeting the minimum application threshold, it is selectively enabled based on the priority of transmission services; for 2.4GHz links with poor channel quality, it is temporarily not enabled to avoid increasing the bit error rate. In practice, the physical layer module of the access point will adjust the modulation parameters according to the link adaptation results, and increase the number of bits carried by each symbol to a higher level. This will directly increase the transmission rate of a single link without increasing the spectrum bandwidth, thereby strengthening the overall bandwidth advantage after multi-link aggregation.

[0098] Subsequently, preamble puncturing technology is employed to avoid frequency band interference. The core logic of preamble puncturing technology is to identify and block interfering subcarriers within a frequency band, using only idle subcarriers for data transmission, thereby reducing the impact of external interference on transmission quality. In multi-link active scenarios, the simultaneous operation of multiple frequency band links may lead to cross-interference between frequency bands, and other wireless devices in the external environment may also interfere with each link. The access point first performs a comprehensive scan of the frequency bands where each active link is located using a spectrum scanning module, identifying the locations of subcarriers occupied by other wireless signals or experiencing strong interference, and forming a list of interfering subcarriers. Then, based on this list, the preamble puncturing mechanism is activated, marking these interfering subcarriers in the preamble portion of the data transmission, instructing the receiver to ignore signals on these subcarriers during demodulation. Simultaneously, the access point dynamically adjusts the subcarrier mapping scheme for data transmission, redistributing data to unmarked idle subcarriers to ensure data transmission integrity. In multi-link scenarios, the application considers the interference situation of each link, avoiding excessive occupation of idle resources on other links to avoid interference on one link, achieving optimal overall interference avoidance through global coordination.

[0099] Subsequently, Hybrid Automatic Repeat Request (HARQ) technology ensures data transmission reliability. Combining the advantages of forward error correction coding and HARQ, the data is first encoded using forward error correction coding during transmission. The receiving end then attempts to recover the data using error correction coding. If error correction fails, a retransmission request is sent to the sending end, which then retransmits the data, thereby reducing the probability of data loss. In multi-link active scenarios, a multi-link collaborative HARQ mechanism is adopted to adapt to the concurrent transmission characteristics of multiple links. In specific implementation, the access point configures an independent HARQ process for each active link and sets up a global HARQ coordination module. When data transmission on a link encounters an error and cannot be recovered through forward error correction coding, the receiving end sends a retransmission request to the access point. This request carries the link identifier and data packet information. Upon receiving the request, the access point's global HARQ coordination module determines the current transmission status of the link. If the link's current load is low and channel quality has been restored, it instructs the link to retransmit. If the link is still under high load or high interference, other active links with lighter loads and better channel quality can be scheduled to undertake the retransmission task, achieving dynamic allocation of retransmission resources. This collaborative retransmission mechanism not only ensures the reliability of data transmission but also fully utilizes the resource redundancy of multiple links, preventing further load increases on a single link due to retransmission.

[0100] Combining the aforementioned WiFi 7 features, transmission speeds are significantly improved. The integration of 4096-QAM technology directly enhances bit transmission efficiency per unit spectrum. Combined with the bandwidth aggregation advantages of multiple links, the overall transmission speed is substantially higher than without this technology, better meeting the high-bandwidth demands of services such as high-definition video and cloud gaming. Furthermore, interference suppression is significantly enhanced. Preamble punching technology precisely shields interfering subcarriers, reducing the impact of interference signals on data transmission, lowering the bit error rate, and effectively controlling frequency band interference during concurrent multi-link transmission, ensuring transmission stability. In addition, transmission reliability is greatly improved. The collaborative application of HARQ technology, through the combination of forward error correction and dynamic retransmission, effectively reduces the probability of data loss. In particular, the multi-link collaborative retransmission mechanism avoids the bottleneck of single-link retransmission, ensuring the integrity and timeliness of data transmission. The synergistic effect of these three factors means that the transmission performance after multi-link activation goes beyond the basic level of bandwidth aggregation; it achieves high-speed, low-interference, and high-reliability transmission goals through the integration of technical characteristics, further enhancing the practical application value of multi-link activation.

[0101] In one implementation, executing the multi-link activation action to be executed further includes the following steps: Step S380: Based on the number of users connected to the access point, allocate resource units of multiple links evenly so that the QoS requirements of each user meet the preset standards.

[0102] In this embodiment, during the multi-link activation process, resource units of multiple links are evenly allocated based on the number of connected users at the access point, ensuring user fairness and service quality. Using the current actual number of connected users at the access point as a benchmark, combined with the total amount of activated multi-link resources, resource units are evenly allocated to ensure that each user receives resources sufficient to meet their preset service quality requirements, avoiding issues such as service lag and excessive latency for some users due to uneven resource allocation.

[0103] The multi-link activation action also sequentially completes user count statistics, resource unit accounting, and balanced allocation.

[0104] Specifically, the first step is real-time statistics and status identification of the number of connected users at the access point. The access point first collects information on successfully associated user terminals in real time through the user association management module to accurately count the total number of connected users. During the statistics process, users in a dormant state, disconnected, or those merely maintaining signal association without data transmission needs are excluded. Only active users with actual business transmission needs are included in the statistics to ensure the accuracy of the allocation benchmark. Simultaneously, the access point identifies the service type of each active user and determines their corresponding quality of service (QoS) requirements. For example, it identifies the low latency requirements of cloud gaming users, the high bandwidth requirements of high-definition video users, and the medium latency and bandwidth requirements of web browsing users, and categorizes users according to their QoS requirements. The access point updates user statistics and service type tags at fixed intervals to adapt to dynamic scenarios such as user access, logout, or service switching.

[0105] Secondly, the total amount of multi-link resource units is calculated, and allocable resources are screened. A resource unit is the basic resource unit for data transmission in a WiFi 7 network, and the number and carrying capacity of resource units vary across different frequency bands. The access point first conducts a resource unit survey of all activated multi-links to calculate the total number of resource units available for data transmission on each link. During this calculation, resource units already used for link synchronization and control signaling transmission are deducted, and only effective resource units available for user service data transmission are counted. Simultaneously, the access point assesses the carrying capacity of effective resource units based on the real-time channel quality status of each link. For example, a 6GHz link with excellent channel quality has a higher resource unit carrying capacity than a 5GHz link with average channel quality. The resource units of each link are weighted according to carrying capacity to form a unified standard for the total allocable resources. This avoids the problem of uneven actual transmission capacity caused by allocating solely based on the number of resource units, ensuring that the fairness of the allocation is based on the equal value of actual transmission.

[0106] Finally, resource unit allocation and service quality verification are performed based on the number of users. The access point, based on the total number of active users, divides the converted total allocable resources into an average share, determining the basic resource unit allocation for each user. Based on this basic allocation, fine-tuning is performed, taking into account the previously identified user service types and service quality requirements. Specifically, for users with high service quality requirements, the resource unit allocation is appropriately increased based on the average share to ensure their core needs are met; for users with low service quality requirements, the allocation share can be appropriately reduced, provided it does not fall below the minimum average share, and the surplus resources are allocated to high-demand users. During fine-tuning, the principle of total balance is followed to ensure that the total resource unit allocation for all users does not exceed the total allocable resources. After allocation, the access point initiates a service quality verification mechanism, monitoring each user's service transmission parameters in real time, including transmission latency, bandwidth, and packet loss rate, to determine whether they meet the preset service quality standards. If a user's transmission parameters are found to be substandard, the access point will readjust the resource unit allocation scheme, prioritizing the user's basic service quality requirements, until the service quality of all users meets the preset standards.

[0107] The above operations significantly improve user fairness. By using an average allocation mechanism based on the number of active users, the problem of a few users consuming excessive resources and causing a decline in the user experience for others is avoided. This ensures that every user can fairly enjoy the resource benefits brought by multi-link activation. In particular, the fine-tuning and optimization combined with business types ensures fairness while taking into account differentiated needs, achieving a balance between fairness and efficiency. Furthermore, it enhances the stability of overall service quality. The precise allocation of resource units ensures that the service quality needs of each user are met, reducing business interruptions and lag caused by insufficient resources. At the same time, the dynamic user statistics and allocation adjustment mechanism allows the resource allocation scheme to quickly adapt to dynamic changes in user access and exit, maintaining the stability of overall service quality even in scenarios with large fluctuations in the number of users.

[0108] Step S400: Obtain the network status after performing the multi-link activation action.

[0109] In this embodiment, the network state after the action is executed is the comprehensive state data updated after the access point completes the configuration operation. It can directly reflect the impact of the action execution on network operation and is the basis for evaluating the effect of the action. The reward is a quantitative evaluation index designed based on global network performance indicators. It is used to guide the iterative optimization of model parameters. Its calculation integrates core performance dimensions such as network throughput, packet delivery rate, transmission delay, and access point energy consumption, and can achieve multi-objective collaborative optimization.

[0110] This embodiment also proposes a closed-loop design for real-time status feedback and multi-dimensional reward calculation. Real-time acquisition of the network state after an action is executed ensures the accuracy and timeliness of reward calculation, enabling the model to quickly perceive the impact of actions on the network. The multi-dimensional reward design overcomes the limitations of traditional single-index optimization, balancing objectives such as throughput improvement, latency reduction, and energy consumption optimization through weight allocation, avoiding network performance imbalances caused by optimizing a single index. For example, the normalization of transmission latency in the reward function ensures consistent evaluation of latency indicators under different traffic loads, guiding the model to generate configuration actions that balance real-time performance and energy efficiency, providing reliable value guidance for subsequent Q-value calculation and parameter updates.

[0111] Optionally, the reward can be calculated by obtaining the network status after the access point performs an action. Specifically, by fusing multi-dimensional performance indicators, a clear optimization guide is provided to the model, achieving a synergistic improvement in throughput, latency, and energy consumption.

[0112] First, data collection is performed. Specifically, after the access point executes its action, the controller collects two types of data through the communication link: one is the network status of each access point after execution, including the updated number of users, load, coverage radius, etc., corresponding to the state vector update in equation (1); the other is the global performance indicators of the WiFi 7 network, specifically including transmission delay, network throughput, packet delivery ratio (PDR), and access point energy consumption. These indicators comprehensively cover network service quality and equipment operating efficiency, and can comprehensively reflect the effect of action execution.

[0113] Next is the normalization of transmission delay. Specifically, since the reasonable range of transmission delay varies greatly under different traffic loads, directly using it for reward calculation will lead to evaluation bias. Therefore, it is necessary to normalize the transmission delay to obtain a normalized delay. Normalization uses a bounded function, expressed as: Equation (6)

[0114] in, Access point Normalization delay, For the original transmission delay, This is the preset maximum acceptable latency. When the original latency is less than... When the normalized delay is constant, the normalized delay has a linear relationship with the original delay; when the original delay exceeds a certain threshold... At this time, the normalized delay is fixed at 1 to avoid the excessive impact of extreme delay on the reward.

[0115] against In the case where the normalized delay is determined, the sensitivity of the reward to the original delay can be quantified by equation (7): Equation (7)

[0116] In this formula, Indicates reward For the original delay The partial derivative reflects the change in reward for every unit change in the original delay; The weight of the delay metric in the reward function. The maximum acceptable delay is given. Equation (7) reveals the sensitivity of rewards to changes in delay. The larger the value, the smaller the absolute value of the sensitivity, and the smoother the response of the reward to delay fluctuations. The smaller the value, the larger the absolute value of the sensitivity, and the more significant the penalty for latency changes. This sensitivity adjustment mechanism makes the penalty for latency during training more controllable, avoiding training oscillations caused by differences in the reasonable range of latency in different scenarios.

[0117] Finally, the reward is calculated using a weighted summation method. Specifically, the reward function combines global performance metrics and normalized latency, and is expressed as: Equation (5)

[0118] in, To normalize throughput, To normalize the group delivery rate, Normalized energy consumption; , , , These are non-negative weights used to balance the importance of different metrics. The design logic of the reward function is: increasing throughput and group delivery rate will increase the reward, while increasing latency and energy consumption will decrease the reward, guiding the model to generate actions that take into account multiple objectives.

[0119] Comprehensive data collection ensures that reward calculations reflect the true effects of actions, avoiding evaluation biases caused by single metrics. Normalization makes different metrics comparable, ensuring the stability of reward calculations. The weighted fusion reward function achieves multi-objective collaborative optimization, preventing the model from overemphasizing one metric while neglecting others. For example, when an action increases throughput but significantly increases energy consumption, the reward function reduces the value of that action through the penalty effect of energy consumption weight, guiding the model to select a more balanced action configuration.

[0120] In one implementation, the WiFi 7 network includes several access points and a controller. After obtaining the network status after performing the multi-link activation action, the process further includes the following steps: Step S510: Each access point reports the multi-link activation action it will perform and the network status after performing the multi-link activation action to the controller; Step S520: The controller acquires the global network status; Step S530: The controller optimizes the execution network of each access point based on the multi-link activation action performed by each access point, the network status after each access point performs the multi-link activation action, and the global network status.

[0121] In this embodiment, the controller is the global coordination unit of the entire network. In the HarmonyOS scenario, it can act as a super device, responsible for integrating the status information of all access points, formulating global optimization strategies, issuing action commands, and iteratively updating model parameters. It is the core hub for realizing multi-access point collaborative optimization.

[0122] Each access point reports the executed multi-link activation action and the network status after the action to the controller. The controller generates a global network status and calculates the corresponding Q value. The Q value is then used to optimize the network execution at each access point.

[0123] The Q-value is a quantitative evaluation of the value of an action. It represents the expected cumulative reward that can be obtained in the future after performing an action in the current network state. Its purpose is to provide a quantitative standard for the quality of actions and help the model distinguish the long-term value of actions with different configurations. The calculation of the Q-value is based on the evaluation network's modeling of the mapping relationship between network state, action, and reward. It integrates the weighted sum of current reward and future reward, reflecting the long-term optimization effect of the action.

[0124] This embodiment proposes a multi-agent collaborative Q-value calculation logic. The network evaluation calculates the Q-value based on the global network state and the actions of all access points, rather than evaluating the value of a single access point's action in isolation. This logic considers the synergistic effect of multiple access point actions, avoiding the global performance degradation caused by local optima of a single access point. For example, when an access point performs a merging action, the Q-value calculation simultaneously considers its impact on the load distribution of adjacent access points and network interference, ensuring that the action selection aligns with the global optimization objective. The introduction of the Q-value allows the model to break free from the limitations of traditional rule-based configuration, learning optimal action selection strategies adapted to dynamic network environments through a data-driven approach, providing a core basis for subsequent parameter updates.

[0125] The Q-value is calculated by evaluating the network, integrating the global network status with the actions and rewards of all access points, and calculating the current Q-value and target Q-value through the main evaluation network and the target evaluation network respectively, providing a valuable basis for parameter updates.

[0126] The first step is the construction of the global network state. Specifically, the controller collects the network state after all access points have performed actions and integrates it to form the global network state. ,in This represents the number of access points. The global network status reflects the operational condition of the entire WiFi 7 network, providing a holistic perspective for centralized evaluation.

[0127] Then the main evaluation network calculates the current Q-value. Specifically, it considers the global network state and the actions of all access points. and the calculated reward Input the main evaluation network. The dual-Q structure of the main evaluation network will calculate the two Q values ​​separately. and 2. Finally, the minimum value between the two is selected as the current Q-value output by the main evaluation network. It effectively suppresses overestimation of the Q value, ensuring the accuracy of the current Q value.

[0128] Finally, the target evaluation network calculates the target Q-value. Specifically, this involves evaluating the global network state and the target actions generated by the target execution network. And reward input target evaluation network. Target action The target execution network generates actions based on the next state, and also includes discrete actions and continuous adjustments. The target evaluation network also uses a dual-Q structure to calculate two target Q values. and The minimum value is selected as the final target Q value. .

[0129] The construction of the global network state ensures that Q-value calculations take into account the synergistic effects of access point actions, avoiding misjudgments caused by local perspectives. The dual-Q structure of the main evaluation network and the target evaluation network respectively guarantees the accuracy of the current Q-value and the target Q-value, providing reliable data for subsequent time-series difference error calculations. For example, when multiple access points collaboratively execute load-sharing actions, the global network state can reflect the balance of load distribution. The main evaluation network will output a higher current Q-value, while the target evaluation network predicts subsequent rewards based on the next state, guiding the model to continuously select collaborative actions.

[0130] In one implementation, after the execution network generates a multi-link activation action, it receives the global collaborative value assessment result of the controller's centralized evaluation network and performs secondary verification to avoid multi-access point multi-link activation conflicts.

[0131] In this embodiment, the secondary verification after the network generates multi-link activation actions is based on the controller's centralized evaluation network performing conflict prediction and value verification on the locally generated multi-link activation actions. The aim is to avoid problems such as frequency band occupation conflicts and resource unit competition caused by multiple access points activating multiple links simultaneously, and to ensure optimal global coordination of actions.

[0132] After generating a multi-link activation action locally at the access point, the key information of the action is first reported to the controller. Then, the controller, considering the global network status (i.e., the current status of all access points and the proposed action), calculates the global collaborative value of the action, assesses conflict risks, and provides feedback on the verification results. Finally, the access point only performs local parameter verification after passing the secondary verification.

[0133] Specifically, firstly, after a local access point generates a multi-link activation action, it reports the action details to the controller via the WiFi 7 multi-link operation low-power control link. This includes the proposed frequency band combination to be activated, the required number of resource units, the estimated bandwidth requirement, and the planned activation time interval. The reported information must be associated with core data such as its own frequency band usage status and multi-link activation status to ensure the controller obtains the complete action context. Subsequently, the controller's centralized evaluation network receives the reported actions from all access points and the global network status, comprehensively calculating the global collaborative value of each proposed action. The evaluation process assesses both the local benefits of the action itself and the global conflict risk, balancing the local benefits and the global conflict impact through preset conflict penalty weights. Then, the controller presets a global collaborative value baseline, which is set based on factors such as network scale, total spectrum resources, and typical service requirements. If the global collaborative value of the access point's proposed action reaches or exceeds the baseline, it indicates that the global collaborative benefit of the action outweighs the conflict risk, and the secondary verification passes; if it does not reach the baseline, the verification fails.

[0134] Secondary verification effectively avoids conflicts caused by multi-access point and multi-link activation. By predicting frequency band and resource unit occupancy conflicts in advance from a global perspective, it ensures that multi-link activation actions are more aligned with global resource allocation needs, reducing performance fluctuations caused by conflicts. Furthermore, it improves spectrum resource utilization. The controller, through global scheduling, ensures that multi-link activation actions from different access points stagger their use of key frequency bands and resource units, avoiding resource fragmentation and waste. In addition, it ensures network stability, preventing interference amplification caused by conflicts and maintaining stable network latency after multi-link activation without significant fluctuations, matching the low latency and high reliability characteristics of WiFi 7.

[0135] In one implementation, the controller optimizes the execution network of each access point based on the multi-link activation actions performed by each access point, the network state after each access point performs the multi-link activation actions, and the global network state. Specifically, this includes the following steps: Step S531: Based on the multi-link activation actions performed by each access point, the network state after each access point performs multi-link activation actions, and the global network state, optimize the network execution of each access point through value evaluation of multi-agent reinforcement learning. The multi-agent reinforcement learning includes the collaboration of centralized evaluation and decentralized execution. The centralized evaluation generates a global value assessment, and the decentralized execution is implemented by the local execution network of each access point.

[0136] In this embodiment, the WiFi7 network in the DC-MARL architecture includes a controller and several access points, forming a centralized evaluation and decentralized execution architecture, wherein the execution network is deployed at each access point.

[0137] From an implementation perspective, the controller, acting as the global coordination hub, possesses powerful computing and data integration capabilities. It can collect network status and action data from all access points to construct a global network view. The evaluation network, deployed on the controller, can assess the collaborative value of all access point actions from a global perspective, avoiding the local optimum trap caused by decentralized evaluation. For example, when multiple access points simultaneously execute split actions, the centralized evaluation network can identify potential load overlap and interference issues, guiding action coordination through Q-value adjustments.

[0138] The execution network is deployed at each access point, enabling each access point to generate actions in real time based on locally collected network status data, without waiting for remote commands from the controller. This design significantly reduces communication latency and adapts to the high-dynamic, low-latency configuration requirements of WiFi 7 networks. For example, when the number of users at an access point suddenly surges, the local execution network can quickly generate split actions to alleviate load pressure in a timely manner, without relying on global scheduling by the controller.

[0139] In this architecture, centralized evaluation ensures consistency in global optimization and avoids conflicts in access point actions, while decentralized execution guarantees the real-time and flexible nature of action generation, improving the network's response speed to dynamic environments. The combination of these two approaches enables the model to possess both global optimization capabilities and the real-time response capabilities of individual access points, fully leveraging the hardware performance potential of WiFi 7 networks and solving the core problems of high latency in traditional centralized scheduling and poor coordination in decentralized configuration.

[0140] In one implementation, the centralized evaluation includes: Step S5311: Deploy a centralized evaluation network based on the dynamic community multi-agent reinforcement learning framework; Step S5312: Through the evaluation network, input the global network status and the actions of each access point, and calculate the global value assessment of each action.

[0141] In this embodiment, the evaluation network is a centralized neural network deployed on the controller, used to evaluate the long-term value of the actions generated by the execution network and provide a quantitative basis for parameter updates. Its inputs cover the global network state and the actions of all access points, and the output is the Q value corresponding to the action, i.e., the action value estimate.

[0142] The evaluation network consists of a main evaluation network and a target evaluation network. The target evaluation network is a copy network with the same structure as the main evaluation network. The main evaluation network adopts a centralized double-Q evaluation network architecture, which contains two evaluation sub-networks with the same structure and independent parameters. Before training begins, the initial parameters of the main evaluation network are completely assigned to the target evaluation network.

[0143] This embodiment proposes a dual-Q structure main evaluation network. The two evaluation sub-networks have identical structures, both being neural networks containing three fully connected layers. Specifically, the input layer dimension is the sum of the global network state vector dimension and the action vector dimensions of all access points, the hidden layers use the ReLU activation function, and the output layer is a one-dimensional Q-value. The initial parameters of both sub-networks follow a normal distribution, ensuring the randomness and consistency of the initial parameters while avoiding training problems caused by excessively large or small parameters.

[0144] The target evaluation network has the same structure as the main evaluation network, including the number of network layers, the number of neurons per layer, and the type of activation function; the only difference is the parameter update mechanism. The parameter assignment operation before training begins involves completely copying the random initial weights and biases of the main evaluation network to the target evaluation network. This ensures that the target network maintains parameter synchronization with the main network from the initial training phase, laying the foundation for stable calculation of the target Q-value later.

[0145] The technical advantage of the dual-Q evaluation network lies in its effective suppression of Q-value overestimation. Traditional single-Q networks are prone to Q-value overestimation due to noise or sample bias, leading the model to select suboptimal actions. In this embodiment, two independent evaluation sub-networks calculate Q-values ​​separately, and the minimum of the two is taken as the output Q-value of the main evaluation network. A double-validation mechanism filters out overestimated Q-values, ensuring the accuracy of action value assessment. The target evaluation network provides a stable target Q-value reference. Specifically, the parameters of the main evaluation network are updated rapidly during training, while the parameters of the target evaluation network are updated later, using a soft update mechanism to adjust slowly, avoiding target Q-value oscillations caused by fluctuations in the main network parameters, thus improving training stability. The combination of the two makes the Q-value calculation of the evaluation network both accurate and stable, providing a reliable basis for subsequent parameter updates, accelerating model convergence, and improving final configuration performance.

[0146] In one implementation, the optimization of the network execution at each access point based on the multi-link activation actions performed by each access point, the network state after each access point performs the multi-link activation actions, and the global network state, through value evaluation using multi-agent reinforcement learning, specifically includes the following steps: Step S5313: Global value assessment based on the output of the centralized evaluation network; Step S5314: Update the parameters of the execution network by calculating the time difference error; Step S5315: The controller sends the updated execution network parameters to the corresponding access point; Step S5316: The access point receives and updates the parameters of the local execution network.

[0147] In this embodiment, the execution network includes a main execution network and a target execution network. The target execution network is a replica network with the same structure as the main execution network. The main execution network is a decentralized execution network. The main execution network of each access point outputs the corresponding action through the network state. Before training begins, the initial parameters of the main execution network are completely assigned to the target execution network.

[0148] The decentralization of the main execution network is reflected in the independent deployment of a set at each access point, enabling action generation without relying on state data from other access points or controllers. Its network structure can be a three-layer fully connected layer. Specifically, the input layer dimension is the dimension of the access point's local network state vector, the hidden layer uses the ReLU activation function, and the output layer is a hybrid action vector. This structure maps local states to the action space, ensuring the targeted and efficient generation of actions.

[0149] The target execution network has the same structure as the main execution network, including the number of neurons in each layer, activation functions, and output layer design; the only difference is the parameter update mechanism. The parameter assignment operation before training begins involves completely copying the random initial parameters of the main execution network at each access point to the corresponding target execution network, ensuring that the initial state of the target execution network is consistent with that of the main execution network.

[0150] The technical advantages of the decentralized master execution network are reflected in its real-time performance and flexibility. Specifically, each access point quickly generates actions based on its local state, reducing dependence on communication links and avoiding scheduling delays caused by centralized execution. For example, when an access point detects that the load of a neighboring access point is low and its own remaining energy is insufficient, the local master execution network can quickly generate a merge action to reduce energy consumption in a timely manner, without waiting for the controller's global decision. The technical advantage of the target execution network is that it provides stable target actions: its parameter updates lag behind the master network, and it is slowly adjusted through a soft update mechanism. The generated target actions are not affected by fluctuations in the master network parameters, providing reliable input for the calculation of the target Q-value, avoiding distortion of the target Q-value caused by rapid changes in the master execution network parameters, and ensuring the stability of the training process.

[0151] Based on the Q value, the parameters of the evaluation network are iteratively updated by minimizing the loss function, and the parameters of the execution network are iteratively updated based on the optimized evaluation network.

[0152] The parameters of the evaluation network and the execution network are updated iteratively based on the Q-value. Specifically, the formula logic and sequence in the technical disclosure are followed. First, the preparatory work for updating the parameters of the main evaluation network is completed, namely, the generation of the target action. Then, the parameters of the main execution network and the soft update of the target network are implemented in sequence to achieve stable optimization of the model.

[0153] The first step is the generation of the target action. During training, when the target execution network generates the target action based on the next state, Gaussian noise is added to improve training stability. Specifically, this is expressed as follows: Equation (8)

[0154] in, Zero-mean Gaussian noise, cropped to the range [ δ, δ], The standard deviation of noise. This is the noise clipping threshold; clipping is used to prevent excessive noise from distorting the target's motion. Execute the network for the target. Its parameters, This represents the next network state after the action is performed. The generation of the target action is the basis for subsequent calculation of the target Q-value, providing stable action input for updating the parameters of the main evaluation network.

[0155] After generating the target action, the parameters of the main evaluation network are updated first, followed by the parameter update of the main execution network. Specifically, a delayed update period is set, and the parameter update of the main execution network is only triggered when the parameter update count of the main evaluation network reaches the period. This design provides a convergence time window for the Q-value prediction of the evaluation network, ensuring that the execution network is optimized based on a relatively stable value assessment.

[0156] The noise mechanism in Equation (8) injects appropriate perturbation into the target action, avoiding overfitting caused by the target action being too simple and improving the generalization ability of the model. The delayed update mechanism ensures the stability of the main execution network update, avoiding frequent changes in the execution network strategy due to fluctuations in the evaluation network parameters, and laying the foundation for smooth convergence of the subsequent overall parameter update.

[0157] The parameters of the main execution network are iteratively updated based on the optimized main evaluation network. Specifically, the parameters of the main evaluation network are first optimized by calculating the target Q value and constructing the loss function. Then, the parameters of the main execution network are adjusted by guiding the gradient of the deterministic policy to ensure that the actions generated by the execution network can maximize the Q value.

[0158] First, the parameters of the main evaluation network are optimized. Specifically, based on the generated target action and current reward, the target Q-value is calculated using a dual-objective evaluation network, and the minimum of the two values ​​is taken to suppress overestimation bias, expressed as: Equation (9)

[0159] in, For the target Q value, For the current reward, As a discount factor, it balances the importance of current and future rewards; For the first A target evaluation subnetwork, Its parameters, The target action is defined. Calculating the target Q-value provides a reliable reference standard for constructing the loss function.

[0160] Subsequently, a loss function is constructed with the objective of minimizing the mean squared error between the predicted Q-value and the target Q-value of the main evaluation network. The expression is as follows: Equation (10)

[0161] in, For the first The loss value of the individual evaluation subnetwork. The parameters of the main evaluation subnetwork are used. The main evaluation network is used to predict the Q-value of the current state-action pair. The buffer serves as an experience replay buffer, and the expected computation is based on batch samples sampled from this buffer. Batch samples are sampled from the experience replay buffer, and the gradient of the loss function with respect to the parameters of the main evaluation network is calculated using the gradient descent algorithm. The parameters are then adjusted along the gradient descent direction until the loss function value converges to a preset range, thus completing the optimization of the main evaluation network parameters.

[0162] After the main evaluation network parameters are optimized, the main execution network parameter update phase will begin. This step uses a dual-Q network with a minimum design to suppress Q-value overestimation, the mean squared loss function to accurately characterize prediction bias, and the gradient descent algorithm to ensure smooth convergence of the main evaluation network parameters, ultimately achieving accurate Q-value prediction and providing a value assessment basis for the optimization of the main execution network.

[0163] The parameters of the main execution network are iteratively updated based on the optimized main evaluation network. Specifically, the parameter adjustment is guided by a deterministic policy gradient to ensure that the actions generated by the execution network can maximize the Q value.

[0164] The first step is setting a delayed update period. Specifically, to avoid interference from fluctuations in the parameters of the main evaluation network on the updates of the execution network, a delayed update period is set. The parameter update of the main execution network is only triggered when the parameter update count of the main evaluation network reaches the set period. This design provides a convergence time window for the Q-value prediction of the evaluation network, ensuring that the execution network is optimized based on a relatively stable value assessment.

[0165] The second step is to calculate the gradient of the deterministic policy. Specifically, based on the Q-value of the optimized main evaluation network output, the parameter gradient of the main execution network is calculated, as follows: Equation (11)

[0166] in, To execute the network's policy gradient, The parameters for the main execution network; The gradient of the Q-value with respect to the action reflects the degree to which changes in the action affect the Q-value. The gradient of the action with respect to the network parameters reflects the impact of parameter changes on action generation; the expectation operation is based on state samples sampled from the empirical replay buffer. .

[0167] Finally, the parameters of the main execution network are updated. Specifically, the parameters are adjusted along the ascending direction of the deterministic policy gradient to ensure that the actions generated by the main execution network continuously improve the Q-value. The step size of the parameter update is controlled by the learning rate to avoid action oscillations caused by excessive parameter adjustments, ensuring smooth policy optimization of the execution network. After the main execution network parameter update is completed, the soft update phase of the target network begins.

[0168] The deterministic policy gradient provides a clear optimization direction for updating the execution network parameters, ensuring that the action generation policy remains consistent with the value assessment of the main evaluation network. The delayed update mechanism improves training stability, avoiding frequent changes in the execution network policy due to fluctuations in the evaluation network parameters, and laying the foundation for subsequent soft updates to the target network.

[0169] Based on the updated parameters of the main evaluation network and the main execution network, the target network is softly updated. Specifically, the parameters of the target network are slowly adjusted through a weighted sum mechanism to ensure the stability of the training process.

[0170] The first step is setting the soft update coefficient. Specifically, the soft update coefficient... The value of this coefficient can be set to the range of 0.001 to 0.01. This coefficient directly determines the degree of influence of the main network parameters on the target network parameters. The smaller the coefficient, the smoother the update of the target network parameters and the higher the training stability; the larger the coefficient, the faster the target network parameters follow the main network, but it may introduce the risk of oscillation. This embodiment selects a smaller soft update coefficient to balance training stability and convergence efficiency.

[0171] The second step is the calculation of the weighted sum. Specifically, based on the soft update coefficient, the weighted sums of the parameters of the main evaluation network and the target evaluation network, and the main execution network and the target execution network are calculated separately, as follows: Equation (12)

[0172] Equation (13)

[0173] in, For the updated target evaluation subnetwork parameters, The updated main evaluation subnetwork parameters; Execute network parameters for the updated target. These are the updated parameters of the main execution network. The weighted sum formula shows that only a small portion of the new parameters of the target network come from the main network, while the vast majority retain the original parameters, ensuring a smooth parameter change.

[0174] Finally, the target network parameters are updated. Specifically, the two weighted sums are used as the new target evaluation network parameters and target execution network parameters, respectively, overwriting the original parameters and completing the soft update. This process ensures that the target network parameters can slowly keep up with the optimization results of the main network, avoiding sudden parameter changes that could cause oscillations in the target Q-value and target actions, and providing a stable reference standard for the next round of training.

[0175] The soft update mechanism avoids drastic fluctuations in the target network parameters, keeping the target Q-value and target action stable, and providing a reliable basis for the next round of parameter updates in the main evaluation network and the main execution network. The slow adjustment of parameter updates makes the training process converge smoothly, avoiding model oscillations caused by fluctuations in the main network parameters, improving the robustness of the entire training process, and ensuring that the model can continuously learn the optimal policy in a dynamic WiFi 7 network environment.

[0176] In summary, this invention discloses a multi-link activation method for WiFi 7 network access points based on multi-agent reinforcement learning. Based on a dynamic cell multi-agent reinforcement learning framework and combined with the core characteristics of WiFi 7, it achieves intelligent optimization of access point cell configuration. Compared with existing technologies, it significantly improves network throughput and energy efficiency, reduces transmission latency, and adapts to user and load fluctuations through dynamic action strategies at the access points, solving the problems of load imbalance and resource waste, ensuring service continuity and spectrum utilization, adapting to high-density network scenarios, and providing highly reliable and high-performance wireless connectivity support for data-intensive applications.

[0177] The specific training and application steps of the DC-MARL framework described above are as follows: Figure 13As shown in the diagram. In this process, the network is first established, then the state of each access point is observed, followed by the training round. In the subsequent training round, it is first determined whether Distributed Multi-Agent Reinforcement Learning (DC-MARL) is enabled. If not enabled, the process sequentially initializes the Q-value metric, processes state-action pairs, evaluates rewards, and updates the Q-table to obtain the optimized value. If enabled, the DC-MARL framework is entered, and it is sequentially determined whether to perform actions such as merging, splitting, resizing, and load sharing. After generating rewards, these are processed by the decision engine, while simultaneously incorporating data from the experience replay buffer.

[0178] Based on the DC-MARL framework, this embodiment can also be combined with the characteristics of WiFi 7 networks. Each access point integrates core WiFi 7 features to optimize transmission performance; these core WiFi 7 features include 4096 quadrature amplitude modulation technology, preamble punching technology, hybrid automatic repeat request technology, and resource unit allocation based on the number of service users.

[0179] It consists of a set of N autonomous access points (APs), labeled as a set. Equipped with DC-MARL optimization and aggressive mode. Each AP can access one or more data sources from the collection. The frequency band.

[0180] For Multi-Link Operation (MLO), MLO supports concurrent transmission across multiple frequency bands, and its basic effective bandwidth formula is: Equation (19)

[0181] in, It is the MLO coordination overhead factor. It represents the percentage of bandwidth loss caused by preamble puncturing. Access point The number of active MLO links.

[0182] To understand how multi-link operations improve throughput, consider the incremental gain when enabling a new MLO link: Equation (20)

[0183] like and The changes are relatively small, and the first-order approximation is: Equation (21)

[0184] Furthermore, WiFi 7 introduces a more advanced modulation scheme, 4096-QAM, allowing each symbol to carry higher data density. This paper employs a high-level analytical model for WiFi 7 characteristics. 4096-QAM is represented as a fixed PHY rate improvement relative to the WiFi 6 benchmark; preamble puncturing reduces the effective bandwidth according to the puncturing ratio; and HARQ is modeled as a multiplicative improvement in packet delivery rate (PDR). These approximations allow for comparison of the relative performance of different schemes under consistent assumptions without needing to capture all the protocol details of IEEE 802.11be.

[0185] The modulation gain of 4096-QAM is calculated as follows: Equation (22)

[0186] Therefore, 4096-QAM offers a 20% throughput improvement compared to WiFi 6. Furthermore, preamble punching allows the AP to avoid congested sub-channels within the wide channel, thus the effective bandwidth is: Equation (23)

[0187] in, It is full bandwidth. This is due to the spectral puncture rate caused by interference. This ensures high throughput even in interference environments, where the effective throughput becomes: Equation (24)

[0188] To improve reliability, Hybrid Automatic Repeat Request (HARQ) is used to retransmit packets when transmission fails, which improves packet delivery rate. Equation (25)

[0189] in, This is the baseline packet delivery rate without HARQ. Ensure HARQ is enabled when a packet fails. It is an efficiency gain factor. HARQ can typically improve PDR by 10-15%. In addition to core PHY / MAC improvements, the framework also leverages the enhanced system management features of WiFi 7 to achieve fair resource allocation and energy saving. Furthermore, enhanced OFDMA and Resource Unit (RU) allocation in WiFi 7 improve OFDMA performance through finer RU allocation and multi-user scheduling. If the AP is... If each user is provided with an equal RU allocation, then: Equation (26)

[0190] Therefore, the throughput per user is: Equation (27)

[0191] The above formula ensures fair resource sharing and avoids starvation. The AP can be set to run a timer for each user, and users can enter sleep mode at other times, thus improving energy efficiency for each user as follows: Equation (28)

[0192] The above formula calculates the optimal sleep duration for devices within the DC-MARL framework, dynamically adjusting based on current throughput performance, access point load conditions, and cell size characteristics. By allowing longer sleep times when network conditions permit, while maintaining responsive operation during high-demand periods, it intelligently balances energy efficiency and service quality.

[0193] During resizing and load balancing, the AP can switch between frequency bands based on coverage area and traffic congestion: Equation (29)

[0194] The path loss is calculated using a frequency band-specific logarithmic propagation path loss model: Equation (30)

[0195] Where 32.4 is the fixed offset of the free space path loss model. It is the operating frequency. It is the exponential path loss factor. Cell size adjustment and optimal MLO band are selected based on path loss.

[0196] Each AP generates a Poisson distributed flow with a mean of λ. Assume that uplink and downlink flow are uniformly distributed: Equation (30)

[0197] Due to competition and protocol overhead, delivery throughput will decrease by a factor. Therefore, the delivery throughput in each direction is: Equation (31)

[0198] Packet delivery rate depends on interference, link quality, and traffic load. For the DC-MARL scheme, PDR is modeled as follows: Equation (32)

[0199] in, Indicates the baseline reliability based on interference. Indicates the MLO enhancement factor. Represents the spatial flow factor. Indicates traffic load. and It is the shape parameter of the Sigmoid.

[0200] The above formulas directly correspond to the performance quantification of WiFi 7's core features. Access points integrate these features and calculate and optimize transmission parameters according to the formulas to improve transmission performance, working in synergy with cell configuration optimization strategies.

[0201] Based on this, a specific simulation experiment was designed to verify the performance of this embodiment. The relevant parameters of the simulation experiment are shown in Table 1.

[0202] Table 1

[0203] The simulation experiment was conducted in a 100m×100m indoor WiFi 7 network environment, deploying 6 randomly distributed access points (APs) supporting the 2.4GHz, 5GHz, and 6GHz frequency bands. Poisson distribution was used to simulate user traffic (3-10 users per AP initially), and traffic load was covered at five levels: 200MHz, 400MHz, 600MHz, 800MHz, and 1000MHz. The experiment used the dual-delay deep deterministic policy gradient (TD3) algorithm to train a dynamic cell multi-agent reinforcement learning (DC-MARL) model, with the actuator learning rate set accordingly. evaluator learning rate The discount factor was 0.95, and the training run consisted of 500 rounds (20 seconds per round). The comparison schemes included the standard WiFi 7 scheme and the aggressive WiFi 7 MLO scheme. The core test metrics were network throughput, packet delivery rate (PDR), transmission latency, training stability, and feature attention distribution.

[0204] The experiment was conducted in three phases. The first phase involved environment deployment and parameter configuration, building a simulation network according to the aforementioned conditions. The second phase involved model training, where the DC-MARL agent learned strategies such as cell merging, splitting, and load sharing through interactive learning. The third phase involved performance testing, collecting performance data for each scheme under different traffic loads and generating corresponding performance curves and feature attention maps.

[0205] Experimental results passed Figures 6 to 11 exhibit.

[0206] Specifically, Figure 6The relationship between throughput and training epochs under different traffic loads was demonstrated, showing the throughput trends of the three schemes under low, medium, and high loads. The DC-MARL enhanced scheme converged the fastest (stable within 100 epochs), achieving throughput exceeding 10Gbps under low load (200MHz), maintaining 33-35Gbps under medium load (400-600MHz), and breaking through 30Gbps under high load (800-1000MHz), with a peak of 38Gbps. The standard WiFi7 scheme achieved a throughput of only 21-25Gbps under high load, while the WiFi7 aggressive MLO scheme achieved approximately 25-30Gbps. The results indicate that DC-MARL significantly improves network throughput under different loads through dynamic cell configuration and MLO collaboration.

[0207] Figure 7 The relationship between Packet Delivery Rate (PDR) and training rounds under different traffic loads is shown, reflecting changes in data transmission reliability. The PDR of the DC-MARL series solutions is consistently higher than that of traditional solutions. Under low load, the enhanced DC-MARL solution achieves a PDR of 99%, maintains 96% under medium load, and remains above 95% under high load; the standard WiFi 7 solution's PDR drops to 87% under high load, while the WiFi 7 aggressive MLO solution reaches approximately 90%. DC-MARL's interference avoidance and HARQ collaborative strategy effectively reduces packet loss caused by congestion.

[0208] Figure 8 The study demonstrates the relationship between average latency and training rounds under different traffic loads, showcasing the optimization effect on transmission timeliness. The DC-MARL solution maintains low latency throughout. It achieves approximately 2.5-3.0ms under low load, 3.5-4.0ms under medium load, and no more than 4.5ms under high load; in contrast, traditional solutions experience latency spikes to 6-15ms under high load. DC-MARL reduces queuing delays and interference through dynamic cell splitting and intelligent frequency band switching.

[0209] Figure 9 Training metrics such as reward, loss, learning rate, and stability are presented, demonstrating the dynamic process of model training. The cumulative reward of the DC-MARL enhancement scheme eventually reaches 800, significantly higher than that of traditional schemes (below 400). The loss of the double-Q evaluation network converges to [value missing] within 100 rounds. The magnitude of the algorithm avoids overestimation of the Q-value. A controllable decay strategy for the learning rate balances exploration and exploitation; the reward variance is less than 100, demonstrating better stability than traditional schemes (variance > 250), validating the model's training stability and convergence efficiency.

[0210] Figure 10The presentation showcases the attention features of the DC-MARL scheme, revealing the attention weight distribution of DC-MARL agents. Collaborative features such as neighbor load, effective bandwidth, and MLO enable status account for over 60%, while energy metrics account for 8-12%. This demonstrates that the model can focus on the global network state and optimize resource allocation through multi-AP collaborative optimization, avoiding the trap of local optima.

[0211] Figure 11 The study demonstrates the attention characteristics of WiFi 7 solutions, showing that traditional WiFi 7 solutions focus on self-reference indicators such as their own load and frequency band (accounting for >70%), lacking attention to neighbor status and global coordination, resulting in rigid resource allocation and difficulty in adapting to dynamic load changes.

[0212] Therefore, it can be concluded that the DC-MARL framework, through multi-agent collaboration and dynamic cell configuration, significantly outperforms traditional WiFi7 solutions in throughput (28%-52% improvement), PDR (5%-8% improvement), transmission latency (25%-30% reduction), and training stability. It can adapt to high-density, high-load network scenarios and fully unleash the hardware performance potential of WiFi7.

[0213] In addition, performance verification (mean traffic) is set up under different average traffic loads, such as... Figure 12 As shown in the figure. This experiment verifies the model's adaptability under different traffic loads ranging from 5-40 Mbps, comparing it with traditional static strategies and single-agent strategies. The results show that under light load, the model reduces latency by 12%-15%; under medium load, the throughput is improved by 28%-33% compared to the traditional strategy, with PDR maintained above 99.2%; under heavy load, the traditional strategy suffers from severe congestion (PDR 82%), while the model still achieves a throughput of 38.5 Mbps, a PDR of 95.7%, and a latency of less than 22 ms. These results indicate that the model can dynamically adapt to traffic changes, and its congestion mitigation effect is significant under medium and heavy loads.

[0214] An experiment using feature attention was conducted, quantifying the contribution of state features through gradient and permutation methods. The results showed that the access point queue length (32%), user signal-to-noise ratio (28%), and current throughput (21%) were the core features, contributing a total of 81%, while marginal features accounted for only 4%. The feature attention mechanism can focus on core states, reduce redundant interference, improve decision accuracy, and provide a basis for lightweight model optimization.

[0215] The ablation study, using the complete model as a baseline, removed core components to verify the performance impact: without centralized commentators or experience replay, throughput decreased by 21%-27% and training oscillations occurred; without feature attention or soft updates to the target network, the number of convergence epochs increased by 15%-22%, and latency increased by 18%-35%. The experiments confirmed that the synergy of each component is indispensable, validating the rationality of the DC-MARL framework and the MATD3 algorithm architecture. The ablation study results are shown in Table 2.

[0216] Table 2

[0217] A queuing-theoretic interpretation experiment was conducted, mapping access points (APs) to service desks and data packets to customers. The model dynamically configured the corresponding service desk resource reallocation. The experiment showed that the model reduced the average AP queue length by 42%-51%, decreased user waiting time by 38%-45%, and improved service desk utilization balance to 85%. From a queuing-theoretic perspective, this validated the core logic of the solution in addressing the mismatch between service rate and arrival rate, demonstrating the theoretical rationale behind the technical solution.

[0218] like Figure 14 As shown in the figure, this embodiment of the invention provides a WiFi 7 network access point multi-link activation system based on multi-agent reinforcement learning. The system is applied to a controller and several access points. The controller and the access points are connected to the network. The training system includes: a network state acquisition module 10, an action generation module 20, an action verification and execution module 30, and a post-action network state acquisition module 40.

[0219] Specifically, the network state acquisition module 10 is used to acquire the network state; the action generation module 20 is used to generate a multi-link activation action to be executed based on the network state and through a pre-trained execution network; wherein, the pre-trained execution network is locally deployed at the access point and is trained through multi-agent reinforcement learning; the action verification execution module 30 is used to verify the multi-link activation action to be executed, and if the verification passes, the multi-link activation action to be executed is executed; the post-action network state acquisition module 40 is used to acquire the network state after executing the multi-link activation action.

[0220] Based on the above embodiments, the present invention also provides a terminal device, the principle block diagram of which can be as follows: Figure 15As shown, the terminal device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When executed by the processor, the computer program implements a multi-link activation method for WiFi 7 network access points based on multi-agent reinforcement learning. The display screen can be an LCD screen or an e-ink screen. The temperature sensor is pre-installed inside the terminal device to detect the operating temperature of the internal components.

[0221] Those skilled in the art will understand that Figure 15 The schematic diagram shown is only a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0222] In one embodiment, a terminal device is provided, including a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs including instructions for performing operations as described in the embodiments of the methods above.

[0223] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0224] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0225] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning, characterized in that, The method, applied to an access point in a WiFi 7 network, includes: Get network status; Based on the network state, a multi-link activation action to be executed is generated through a pre-trained execution network; wherein, the pre-trained execution network is locally deployed at the access point, and the pre-trained execution network is trained through multi-agent reinforcement learning; The multi-link activation action to be executed is verified. If the verification passes, the multi-link activation action to be executed is executed. Obtain the network status after executing the multi-link activation action.

2. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 1, characterized in that, The acquisition of network status includes: The access point collects its own operational data and obtains the operational data of adjacent access points through the communication link. By integrating its own operational data with the operational data of adjacent access points, the initial network state is obtained; The initial network states are filtered and normalized, and then combined to form the final network states; The network status includes the number of users connected to the access point, current traffic load, remaining energy, average load of adjacent access points, normalized effective bandwidth, multi-link operation enabled status, current cell size, frequency band used, and time normalization factor.

3. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 1, characterized in that, The multi-link activation action includes: Enable WiFi 7 multi-link operation, activate transmission links of at least two frequency bands, and optimize transmission throughput, anti-interference ability and transmission stability through multi-band operation.

4. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 3, characterized in that, The verification of the multi-link activation action to be executed includes: Based on the network status, obtain the single-band throughput, traffic load, channel interference intensity, and multi-link coordination overhead estimate of the local access point. The single-band throughput is compared with a preset demand threshold, the traffic load is compared with a preset activation threshold, the channel interference intensity is compared with a preset interference threshold, and the estimated cooperative overhead is compared with a preset overhead threshold. If the single-band throughput satisfies the relationship of being less than or equal to, the traffic load satisfies the relationship of being greater than or equal to, the channel interference intensity satisfies the relationship of being greater than or equal to, and the estimated cooperative overhead satisfies the relationship of being less than or equal to, then the verification result of the multi-link activation action is obtained.

5. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 3, characterized in that, The execution of the multi-link activation action to be executed includes: The local access point selects the appropriate frequency band combination based on the network status and initiates a multi-link collaborative configuration request; Configure the synchronization parameters for each link; A multi-link data transmission and reception mechanism is enabled, and data transmission is completed by integrating WiFi 7 transmission features, while simultaneously updating the multi-link activation status to adjacent access points.

6. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 5, characterized in that, The execution of the multi-link activation action to be executed further includes: Data transmission is completed by integrating WiFi 7 transmission features; these integrated WiFi 7 transmission features include: Integrated 4096 quadrature amplitude modulation technology improves transmission rate; Use preamble punching technology to avoid frequency band interference; Data transmission reliability is ensured through hybrid automatic repeat request technology.

7. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 6, characterized in that, The execution of the multi-link activation action to be executed further includes: Based on the number of users connected to the access point, resource units of multiple links are evenly allocated to ensure that the QoS requirements of each user meet the preset standards.

8. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 1, characterized in that, The WiFi 7 network includes several access points and a controller. After obtaining the network status after performing the multi-link activation action, it further includes: Each access point will report the multi-link activation action it performs and the network status after performing the multi-link activation action to the controller; The controller acquires the global network status; The controller optimizes the execution network of each access point based on the multi-link activation actions performed by each access point, the network status after each access point performs multi-link activation actions, and the global network status.

9. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 8, characterized in that, After the execution network generates a multi-link activation action, it receives the global collaborative value assessment result of the controller's centralized evaluation network and performs secondary verification to avoid multi-access point multi-link activation conflicts.

10. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 9, characterized in that, The controller optimizes the execution network of each access point based on the multi-link activation actions performed by each access point, the network state after each access point performs the multi-link activation actions, and the global network state, including: Based on the multi-link activation actions performed by each access point, the network state after each access point performs multi-link activation actions, and the global network state, the network execution of each access point is optimized through the value evaluation of multi-agent reinforcement learning. The multi-agent reinforcement learning includes the collaboration of centralized evaluation and decentralized execution. The centralized evaluation generates a global value assessment, and the decentralized execution is implemented by the local execution network of each access point.

11. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 10, characterized in that, The centralized evaluation includes: A centralized evaluation network is deployed based on a dynamic community multi-agent reinforcement learning framework. By using the evaluation network, the global network state and the actions of each access point are input, and the global value assessment of each action is calculated.

12. The method for multi-link activation of WiFi 7 network access points based on multi-agent reinforcement learning according to claim 11, characterized in that, The network execution at each access point is optimized based on the multi-link activation actions performed by each access point, the network state after each access point performs the multi-link activation actions, and the global network state, through value evaluation using multi-agent reinforcement learning, including: Global value assessment based on the output of a centralized evaluation network; The parameters of the execution network are updated by calculating the temporal difference error. The controller sends the updated execution network parameters to the corresponding access point; The access point receives and updates the parameters of the local network.

13. A multi-link activation system for WiFi 7 network access points based on multi-agent reinforcement learning, characterized in that, An access point for a WiFi 7 network, the system comprising: The network status acquisition module is used to acquire network status. An action generation module is used to generate multi-link activation actions to be executed based on the network state and through a pre-trained execution network; wherein the pre-trained execution network is locally deployed at the access point and is trained through multi-agent reinforcement learning; The action verification and execution module is used to verify the multi-link activation action to be executed. If the verification passes, the multi-link activation action to be executed is executed. The post-action network status acquisition module is used to acquire the network status after executing the multi-link activation action.

14. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a WiFi 7 network access point multi-link activation program based on multi-agent reinforcement learning, which is stored in the memory and can run on the processor. When the processor executes the WiFi 7 network access point multi-link activation program based on multi-agent reinforcement learning, it implements the steps of the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a WiFi 7 network access point multi-link activation program based on multi-agent reinforcement learning. When the WiFi 7 network access point multi-link activation program based on multi-agent reinforcement learning is executed by a processor, it implements the steps of the WiFi 7 network access point multi-link activation method based on multi-agent reinforcement learning as described in any one of claims 1-12.