A method, system, device and medium for indoor substation AP deployment based on MAPPO

By deploying APs within substations using the MAPPO algorithm, a path loss and channel overlap model is constructed to achieve joint optimization of AP location, power, channel, and bandwidth. This solves the problems of dynamic environmental adaptability and multi-dimensional parameter optimization in AP deployment within substations, thereby improving communication performance and resource utilization.

CN121240089BActive Publication Date: 2026-06-19FUXIN POWER SUPPLY COMPANY STATE GRID LIAONING ELECTRIC POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUXIN POWER SUPPLY COMPANY STATE GRID LIAONING ELECTRIC POWER
Filing Date
2025-10-20
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing AP deployment strategies cannot be flexibly adjusted within substations, making it difficult to adapt to dynamic and complex environments. They also cannot jointly optimize multi-dimensional variables such as AP location, power, channel, and bandwidth, resulting in signal strength attenuation, data transmission fluctuations, and impacting the real-time performance and integrity of the communication network.

Method used

An indoor substation AP deployment method based on MAPPO is adopted. By establishing an indoor substation system scenario model, constructing a downlink channel model with path loss and channel overlap, defining a Markov decision process, and using the Multi-Agent Proximity Policy Optimization (MAPPO) algorithm to jointly optimize AP location, power, channel, and bandwidth.

Benefits of technology

It improves the communication performance and resource utilization efficiency of wireless access points, adapts to complex electromagnetic environments, achieves coordinated optimization of system rate and coverage quality, reduces channel conflicts and power consumption, and improves communication reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121240089B_ABST
    Figure CN121240089B_ABST
Patent Text Reader

Abstract

This invention relates to the field of power communication and intelligent deployment technology, specifically to an indoor substation AP deployment method, system, device, and medium based on MAPPO. The method first establishes an indoor substation scenario model and defines the data transmission process; then, it constructs a downlink channel model integrating line-of-sight / non-line-of-sight path loss and channel overlap interference, and establishes a resource allocation model to maximize the total system rate based on the Shannon-Hartley theorem, clarifying the rules for AP power, channel, and bandwidth allocation; next, the problem is transformed into a decentralized partially observable Markov decision process, utilizing a centralized training-distributed execution architecture of the MAPPO algorithm, and through collaborative optimization of the policy network and value network, outputting an optimized deployment scheme for AP location, power, channel, and bandwidth allocation. This invention is applicable to solving the collaborative optimization problem of AP deployment and resource allocation in the complex electromagnetic environment of indoor substations, improving the transmission performance and resource utilization efficiency of power communication networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power communication and smart deployment technology. Background Technology

[0002] With the continuous advancement of smart grid construction, wireless access points (APs), as core nodes for terminal access and data transmission, directly impact the performance of communication networks due to their rational deployment. However, in practical applications, indoor substation scenarios present the following problems:

[0003] The dense distribution of metal obstacles and the complex and variable electromagnetic environment within substations cause APs to face problems such as non-line-of-sight transmission loss and channel overlap interference. AP deployment is prone to problems such as expanded coverage blind spots and increased channel conflicts, which in turn lead to signal strength attenuation and data transmission fluctuations, affecting the business continuity of real-time monitoring and remote operation and maintenance, and making it difficult to meet the requirements of smart grids for real-time and complete data transmission.

[0004] Current AP deployment strategies primarily rely on heuristic algorithms such as genetic algorithms and greedy algorithms. While these algorithms can complete basic deployment planning in simple scenarios, they have significant shortcomings. First, they operate based on preset rules, making it difficult to flexibly adjust deployment strategies and adapt to dynamic and complex environments. Second, they only optimize single-dimensional parameters and cannot jointly optimize multi-dimensional variables such as AP location, power, channel, and bandwidth, making it difficult to achieve globally optimal resource allocation. Furthermore, their search space is limited, making them prone to getting trapped in local optima. In scenarios with dense multi-AP deployments, they cannot balance coverage and communication quality, exhibiting poor algorithm flexibility and adaptability.

[0005] In recent years, Multi-Agent Reinforcement Learning (MARL) has been increasingly applied to AP optimization and deployment tasks in communication scenarios due to its advantages in multi-agent collaborative decision-making and dynamic environment adaptive learning. Existing solutions employ the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, but it uses deterministic policies to output continuous actions, and its experience replay mechanism requires a large number of training samples, resulting in high resource consumption, slow convergence speed, low utilization efficiency, and long training time, making it difficult to adapt to the dynamic requirements of AP deployment in substations. In contrast, the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm adopts a centralized training-distributed execution architecture. Through the collaborative optimization of the policy network and value network, it has advantages in sample utilization efficiency and policy update flexibility, and can better handle multi-dimensional joint optimization problems involving AP location, power, channel, and bandwidth. However, its specific adaptation to the complex electromagnetic environment of substations still has technical gaps.

[0006] Therefore, there is an urgent need for an intelligent deployment strategy for substation wireless access points that can jointly optimize multi-dimensional variables in order to improve network communication performance and resource utilization efficiency. Summary of the Invention

[0007] To overcome the problems of poor adaptability to dynamic environments and insufficient joint optimization of multi-dimensional parameters in existing AP deployment strategies, this invention provides an indoor substation AP deployment method based on MAPPO.

[0008] The technical solution adopted by this invention to achieve the above objectives is: an indoor substation AP deployment method based on MAPPO, comprising the following steps:

[0009] Based on the actual environmental characteristics of indoor substations, an indoor substation system scenario model is established;

[0010] Based on the indoor substation system scenario model, a path loss model under non-uniform medium is constructed by analyzing path loss and channel overlap, and a downlink channel model is constructed by using the path loss model and channel overlap.

[0011] Based on the downlink channel model, a resource allocation model that maximizes the total system rate is established.

[0012] Based on the resource allocation model, by defining a Markov decision process, the problem of wireless access point (AP) deployment and resource allocation is transformed into a decentralized partially observable Markov decision process (Dec-POMDP) ​​decision model.

[0013] Construct a multi-objective reward function based on the Dec-POMDP decision model;

[0014] Based on the Dec-POMDP decision model and multi-objective reward function, the multi-agent near-end policy optimization (MAPPO) algorithm is used to output AP location, power, channel and bandwidth allocation.

[0015] Preferably, based on the actual environmental characteristics of the indoor substation, an indoor substation system scenario model is established, including: defining the indoor substation scenario as a two-dimensional rectangular area. ,Will Evenly divided into A set of squares, with the geometric center of each square as the target point (terminal position). The target point is the specific location that needs to be covered by the AP in the indoor substation scenario. , The coordinates are ;deploy A service set consists of multiple access points of the same type. , Coordinates are .

[0016] Preferably, based on the indoor substation system scenario model, a path loss model under non-uniform media is constructed by analyzing path loss and channel overlap, and a downlink channel model is constructed by using the path loss model and channel overlap, including: analyzing path loss and constructing a path loss model under non-uniform media:

[0017] ;

[0018] in, For reference distance Loss at the point, , At the speed of light, For carrier frequency, for With the target point Path types between, when With the target point The path between them is a Loss path. ;when With the target point The path between them is an NLoS path. , for and The distance between them , Let be the path loss exponents for LoS and NLoS, respectively. , The mean is 0 and the standard deviation is for LoS and NLoS respectively. Gaussian random variables;

[0019] Define binary variables express Channel allocation:

[0020] ;

[0021] Channel spacing When the channel is in an overlapping state, it is defined as a non-overlapping channel; otherwise, it is a partially overlapping channel.

[0022] Preferably, based on the downlink channel model, a resource allocation model that maximizes the total system rate is established, including: defining... For target point The data rate is:

[0023] ;

[0024] in, express For target point The generated data rate (Mbps) express For target point Allocated channel bandwidth (MHz) express For target point The generated signal-to-interference-plus-noise ratio (dB) is calculated using the following formula:

[0025] ;

[0026] in, For target point from The received signal strength, for The transmission power, This refers to the noise power in the indoor environment of a substation. Other adjacent (remove (Except for) the target point The generated interference power, For two channels and The degree of overlap is 0-1.

[0027] Preferably, based on the resource allocation model, the definition of a Markov decision process includes: defining a seven-tuple of a Markov decision process (Dec-POMDP):

[0028] ;

[0029] Among them, intelligent agents For the AP set, global state Includes parameters such as AP location, power, channel, and loss; action space. , For channel set, For power discretization set, Adjusting the discretized set for position coordinates, local observation Include Information about itself and its associated target points; The state transition function is derived from the physical model. drive; For a multi-objective reward function, it is a weighted sum of the system's total rate, uniformity, and penalty term; This is the discount factor.

[0030] Preferably, a multi-objective reward function is constructed based on the Dec-POMDP decision model:

[0031] ;

[0032] in, As a reward for the total system speed, For the percentage of effective payload duration, It is a uniform reward. To restrain punishment, For indicator functions; These are the weighting coefficients.

[0033] Preferably, based on the Dec-POMDP decision model and multi-objective reward function, the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is used for training with pruning policy loss:

[0034] ;

[0035] in, For probability ratios, This is the cutting factor. The entropy coefficient, For time step The advantage function, For policy entropy;

[0036] Calculate the loss of value:

[0037] ;

[0038] in, For loss function, For state Value network output, For time step Instant rewards As a discount factor, This is a termination marker;

[0039] Calculate the advantage function:

[0040] ;

[0041] in, This is the GAE coefficient.

[0042] An indoor substation AP deployment system based on MAPPO includes:

[0043] Scene modeling module: Used to build indoor substation system scene models and define data transmission processes;

[0044] Channel model construction module: Analyzes path loss and channel overlap to construct a downlink channel model;

[0045] Resource allocation calculation module: Establishes a resource allocation model that maximizes the total system rate, determines AP power, channel allocation and bandwidth allocation rules, and outputs a resource allocation scheme;

[0046] Decision model transformation module: This module is used to transform the joint optimization problem of AP deployment and resource allocation into a decentralized partially observable Markov decision model, clarifying the state transition logic.

[0047] Reward function building module: used to construct multi-objective reward functions;

[0048] MAPPO algorithm optimization module: used for iterative optimization strategy, outputting the optimized AP location, power, channel and bandwidth allocation scheme;

[0049] Data interaction and storage module: Used to support the data flow between modules, store the parameters of each module, and provide users with a data query interface.

[0050] An indoor substation AP deployment device based on MAPPO includes a memory and a processor. The memory stores a computer program that is used to execute the above-described method when loaded by the processor.

[0051] A readable storage medium storing a computer program suitable for executing the above-described methods when loaded by a processor.

[0052] The beneficial effects of this invention are as follows:

[0053] This invention constructs a unified channel model that integrates path loss and equipment interference. Through Shapely geometric calculations and calibration with measured data, it improves the model's adaptability to the complex environment of indoor substations. This invention achieves joint optimization of wireless access point (AP) location, power, channel, and bandwidth through a hierarchical decision network and an improved reward function, thereby improving system speed and coverage quality. This invention integrates multi-dimensional joint decision-making of AP location, power, channel, and bandwidth, breaking through the limitations of traditional single-dimensional, step-by-step optimization algorithms. It solves the joint optimization problem of high-dimensional discrete action space through the hierarchical policy network of the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. To address the complex electromagnetic environment of substations, this invention constructs a unified model integrating path loss and channel interference. A centralized training-distributed execution architecture is adopted to adapt to independent AP deployment. A multi-objective reward function is used to achieve coordinated optimization of system rate, coverage balancing, and power consumption. This invention constructs a unified model integrating line-of-sight (LoS) / non-line-of-sight (NLoS) path loss and channel overlap interference, accurately characterizing the complex electromagnetic environment of substations and improving the model's adaptability to non-uniform media and dynamic interference. Based on the IEEE 802.11ax protocol and the 2.4GHz frequency band, this method is compatible with existing power communication equipment and can be promoted in the fields of power communication and intelligent deployment. It can be directly applied to AP deployment projects in indoor substations. Through dynamic resource allocation and topology balancing strategies, it improves communication reliability and resource utilization in complex environments, possessing engineering implementation and promotion value. It achieves intelligent and efficient AP deployment, providing a reference for building wireless communication systems in smart substations. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of the overall process of an embodiment of the present invention;

[0055] Figure 2 This is a schematic diagram of a system model according to an embodiment of the present invention;

[0056] Figure 3 This is a MAPPO deployment topology diagram of the main power distribution room (100m×100m, 12AP) according to an embodiment of the present invention;

[0057] Figure 4 This is a comparison chart of algorithm channel overlap rates in the main power distribution room (100m×100m, 12AP) according to an embodiment of the present invention.

[0058] Figure 5 This is a graph showing the algorithmic transmission power variation of the main power distribution room (100m×100m, 12AP) according to an embodiment of the present invention.

[0059] Figure 6 This is a graph showing the total rate change of the algorithm system in the main power distribution room (100m×100m, 12AP) according to an embodiment of the present invention.

[0060] Figure 7 This is a diagram showing the MAPPO bandwidth utilization of the main power distribution room (100m×100m, 12AP) according to an embodiment of the present invention.

[0061] Figure 8 This is a comparison chart of the algorithm round rewards for the main power distribution room (100m×100m, 12AP) according to an embodiment of the present invention;

[0062] Figure 9 This is a comparison chart of algorithm round rewards for a transformer room (100m×100m, 12AP) according to an embodiment of the present invention;

[0063] Figure 10 This is a comparison chart of the algorithm round rewards for the main power distribution room (90m×90m, 9AP) according to an embodiment of the present invention. Detailed Implementation

[0064] Embodiments of the present invention provide an indoor substation AP deployment method based on MAPPO to achieve optimal coverage of multiple target points after optimized deployment of wireless access points (APs), such as... Figure 1 As shown, it includes the following steps:

[0065] S1. Establish an indoor substation system scenario model and define the data transmission process.

[0066] Scene modeling:

[0067] Define the indoor substation scenario as a two-dimensional rectangular area. ,Will Evenly divided into A set of squares, with the geometric center of each square as the target point (terminal position). The target point is the specific location that needs to be covered by the AP in the indoor substation scenario. , The coordinates are ;deploy A service set consists of multiple access points of the same type. , Coordinates are Meanwhile, the hybrid transmission environment of line-of-sight (LoS) / non-line-of-sight (NLoS) caused by metal obstacles is also considered.

[0068] In the indoor environment of a substation, due to the presence of numerous electrical and metal obstacles, communication between an access point (AP) and a target point is not always line-of-sight; non-line-of-sight transmission also exists. Each target point is associated with only one AP with the strongest signal, communicating with it using the same channel as that AP. The APs and their associated target points form a Basic Service Set (BSS), and all nodes belonging to the same BSS operate on the same channel. In a network of multiple APs, adjacent APs may operate on partially overlapping channels or even the same channel, leading to mutual interference between adjacent BSSs.

[0069] Define the data transmission process:

[0070] Based on the IEEE 802.11 ax protocol, the downlink transmission process includes: the AP broadcasts a trigger frame to start data transmission; the destination returns a clear transmit frame (CTS) to confirm channel availability; the AP transmits downlink multi-user data packets (DL MU PPDU); and the destination returns an acknowledgment frame (ACK) to complete the transmission.

[0071] This section defines the data transmission process in a BSS (Browser Subsystem) under the IEEE 802.11 ax protocol. Since the downlink carries the majority of traffic, we only focus on the downlink, specifically the data transmission from the Access Point (AP) to the target point. First, the AP broadcasts a request to the target point to send a trigger frame, which is the start signal for data transmission, informing the UE (User Equipment) that it is ready to receive data. Then, upon receiving the request, the target point sends a Clear-To-Send (CTS) packet back to the AP to confirm that it is ready to receive data and that the corresponding frequency band is available for transmission, avoiding frequency band conflicts. Next, upon receiving the CTS, the AP transmits a Downlink Multi-User Packet Protocol Data Unit (DLMU PPDU) to the target point; this is the crucial step in the actual data transmission. Finally, after successfully receiving the data, the target point sends an Acknowledgement message back to the AP to ensure correct data reception. If no Acknowledgement message is received, data retransmission may be necessary.

[0072] S2. Based on the indoor substation system scenario model, by analyzing path loss and channel overlap, a downlink channel model is constructed that integrates line-of-sight (LoS), non-line-of-sight (NLoS) path loss and channel overlap interference.

[0073] Select the operating frequency band and complete. Channel allocation:

[0074] The IEEE 802.11ax standard specifies that access points (APs) operate in the 2.4GHz and 5GHz frequency bands. Compared to the 5GHz band, the 2.4GHz band has a longer signal wavelength, resulting in stronger penetration, longer propagation distance, and wider coverage, making it more suitable for the electromagnetic environment of substations with dense metal equipment and complex obstacles. Although the 2.4GHz band experiences more co-channel interference, its compatibility with older wireless devices and signal stability in complex obstructed environments better meet the deployment needs of substations with their multi-metal structures and dense equipment.

[0075] The IEEE 802.11ax standard specifies that access points (APs) operate in the 2.4GHz and 5GHz frequency bands. Compared to the 5GHz band, the 2.4GHz band has a longer signal wavelength, resulting in stronger penetration, longer propagation distance, and wider coverage, making it more suitable for the electromagnetic environment of substations with dense metal equipment and complex obstacles. Although the 2.4GHz band experiences more co-channel interference, its compatibility with older wireless devices and signal stability in complex obstructed environments better meet the deployment needs of substations with their multi-metal structures and dense equipment.

[0076] The 2.4 GHz band contains 11 channels (channels 1-11), with channel spacing... The channels are defined as non-overlapping channels (such as channel 1 and channel 6). The channels partially overlap (e.g., channel 1 and channel 2). This channel distribution characteristic directly affects the interference control strategy when deploying APs—in a 2.4GHz network deployment example, adjacent APs may share more than 5 channels (e.g., ...). Use channel 1 Using the channel 6 allocation method can effectively reduce co-channel and adjacent-channel interference, providing a physical layer design basis for subsequent channel optimization allocation based on the MAPPO algorithm.

[0077] The 2.4 GHz band (wavelength 0.125 m) was selected, and the channel set is as follows: Define binary variables express Channel allocation:

[0078] ;

[0079] Channel spacing When the channel is in an overlapping state, it is defined as a non-overlapping channel; otherwise, it is a partially overlapping channel.

[0080] Construct a path loss model:

[0081] Based on the characteristics of non-uniform media, consider a... With the target point The signal transmission path is determined, and a path loss model under non-uniform media is constructed:

[0082] ;

[0083] in, For reference distance Loss at the point, , At the speed of light, For carrier frequency, for With the target point Path types between, when With the target point The path between them is a Loss path. ;when With the target point The path between them is an NLoS path. , for and The distance between them , Let be the path loss exponents for LoS and NLoS, respectively. , The mean is 0 and the standard deviation is for LoS and NLoS respectively. Gaussian random variables.

[0084] S3. Based on the downlink channel model, and according to the Shannon-Hartley theorem and resource unit (RU) allocation rules, establish a resource allocation model that maximizes the total system rate to obtain AP power, channel, and bandwidth allocation:

[0085] Based on the Shannon-Hartley theorem, define For target point The data rate is:

[0086] ;

[0087] in, express For target point The generated data rate (Mbps) express For target point Allocated channel bandwidth (MHz) express For target point The generated signal-to-interference-plus-noise ratio (dB) is calculated using the following formula:

[0088] ;

[0089] in, For target point from The received signal strength, for The transmission power, This refers to the noise power in the indoor environment of a substation. Other adjacent (remove (Except for) the target point The generated interference power, For two channels and The degree of overlap at any given time, with a value between 0 and 1, represents the channel. and The greater the distance between them, The smaller.

[0090] Assuming the AP operates in the 2.4GHz band, its effective channel bandwidth is 20MHz, and the number of subcarriers is... A 20MHz channel bandwidth supports different numbers of... It supports a maximum of nine 26-tone RUs, four 52-tone RUs, two 106-tone RUs, and one 242-tone RU, with the number of RUs varying depending on the bandwidth. This means that in OFDMA transmission, a 20MHz channel bandwidth can support a maximum of nine target points. The channel bandwidth occupied by each k-tone RU (k∈K) also varies. This indicates the bandwidth corresponding to the k-tone RU:

[0091] ;

[0092] Target point Select the AP with the strongest signal strength for association and define a binary variable. The relationship between each AP and the target point is represented as follows:

[0093] ;

[0094] Define binary variables express Target points served The allocation relationship with k-tone RU is as follows:

[0095] ;

[0096] Therefore, there is , .

[0097] The bandwidth allocation rule is based on OFDMA technology and is determined according to the number of target points served by the AP. Resource Unit (RU) types are dynamically allocated, prioritizing the maximization of the AP's total 20MHz bandwidth while balancing the data rate at the target point. The specific allocation rules are as follows:

[0098] when At that time, allocation 20MHz bandwidth;

[0099] when At that time, allocation 18MHz bandwidth;

[0100] when At that time, allocation 18MHz bandwidth;

[0101] when At that time, allocation 16MHz bandwidth;

[0102] when At that time, allocation 20MHz bandwidth;

[0103] when At that time, allocation 20MHz bandwidth;

[0104] when At that time, allocation 16MHz bandwidth;

[0105] when At that time, allocation 18MHz bandwidth;

[0106] when At that time, allocation , bandwidth 18MHz.

[0107] Prioritize the use of RUs with larger tones to improve bandwidth utilization, with the total bandwidth as close to 20 MHz as possible. Sort the target points served by the AP in descending order of distance, assigning RUs with larger tones to more distant target points and RUs with smaller tones to closer target points, in order to balance rate demand.

[0108] S4. Based on the resource allocation model, by defining a seven-tuple of Markov Decision Process (Dec-POMDP), the AP deployment and resource allocation problem is transformed into a decentralized, partially observable Dec-POMDP decision model:

[0109] This paper establishes an optimization problem related to AP deployment and resource allocation, namely, maximizing the total downlink system rate while satisfying constraints on AP transmit power and deployment location, AP-target point association, and AP channel and bandwidth allocation. The introduced decision variables include: vector... This indicates the AP's transmit power. It is the first Transmit power of each AP; matrix Indicates the deployment location of the AP. yes The coordinates ( );matrix This indicates the relationship between the AP and the target point. yes With the target point The relationships between them; vectors Indicates the channel allocation of the AP. yes With channel The distribution relationship between them; tensors Indicates AP and target point RU allocation, yes For target point The assigned k-tone RU.

[0110] The problem of AP deployment and resource allocation in the indoor substation scenario is described as follows:

[0111] ;

[0112] ;

[0113] Among them, C1 is the AP transmit power constraint, C2 is the AP deployment location constraint, C3 is the binary variable constraint, including association, channel allocation, and RU allocation, C4 is the association constraint, each target point can only be associated with one AP, C5 is the channel allocation constraint, an AP must be allocated one and only one channel, C6 is the RU allocation constraint, each target point must be allocated one k-tone RU by the associated AP, and C7 is the AP service quantity constraint, the number of target points associated with an AP does not exceed the maximum service quantity. C8 is the RU bandwidth constraint, meaning the total RU bandwidth allocated by the AP to the target points it serves cannot exceed the channel bandwidth B. C9 is the SINR constraint, meaning the target point obtains the required SINR from its associated AP. Not less than the threshold .

[0114] Define the seven-tuple of a Markov decision process (Dec-POMDP):

[0115] ;

[0116] Among them, intelligent agents For the AP set, global state Includes parameters such as AP location, power, channel, and loss; action space. , For channel set, For power discretization set, Adjusting the discretized set for position coordinates, local observation Include Information about itself and its associated target points; The state transition function is derived from the physical model. drive; For a multi-objective reward function, it is a weighted sum of the system's total rate, uniformity, and penalty term; This is a discount factor used to balance long-term and short-term rewards. .

[0117] S5. Based on the Dec-POMDP decision model, a multi-objective reward function is constructed by combining system rate, AP uniformity, and constraint penalties.

[0118] The multi-objective reward function is:

[0119] ;

[0120] in, As a reward for the total system speed, , This represents the percentage of effective payload duration. It is a uniformity reward used to promote AP topology balance. ; To constrain penalties, penalties are imposed separately for invalid movement and invalid power. , For indicator functions; The weighting coefficient has a value of [value]. The weighting coefficients can be adjusted according to actual needs.

[0121] S6. Based on the Dec-POMDP decision model and multi-objective reward function, and utilizing the centralized training-distributed execution architecture of the MAPPO algorithm, the policy network and value network are collaboratively optimized to output the final optimized AP location, power, channel, and bandwidth to meet the target point. Data transmission requirements;

[0122] Building the algorithm architecture:

[0123] The MAPPO algorithm adopts a centralized training-distributed execution architecture, which is suitable for multi-AP collaborative decision-making scenarios in substations. It decomposes the AP deployment problem into a multi-agent parallel optimization task. The specific network structure design logic is as follows:

[0124] Policy network (Actor) design: Each AP acts as an independent agent, focusing on local observations. It encompasses its own location, target point information, channel state, etc. Features are extracted through a two-layer 32-dimensional MLP, and the input dimension is non-linearly mapped and dimensionality reduced. The output layer... The function generates an action probability distribution for a discrete, multi-dimensional action space. To adapt to the dynamic environment of substations, a channel change sensing module is embedded on the network input side to integrate time-series information such as channel occupancy and interference intensity in real time, thereby enhancing the strategy's response capability to time-varying interference.

[0125] Value network (Critic) design: with global state As input, integrate the AP location set. Power configuration Channel allocation and the target point signal-to-interference-plus-noise ratio Interactive information is captured through a 2-layer 32-dimensional MLP to identify complex correlations such as interference coupling and channel conflicts between APs. The output layer directly maps the total system value. This is used to quantify the long-term cumulative reward expectation under the global state. To improve the accuracy of value estimation, a dynamic reward scaling mechanism is introduced at the network output, which adaptively adjusts the value output range according to the reward distribution characteristics of the current scenario, thereby optimizing the stability of policy gradient updates.

[0126] Execute the training process:

[0127] Data collection: Initializing the policy network Value Network experience buffer Value normalizer Reset AP status and obtain initial observations. Each AP sampling action Execute actions Get rewards New observations ,storage arrive .

[0128] Training is performed using a pruning strategy loss:

[0129] ;

[0130] in, For probability ratios, This is the cutting factor. The entropy coefficient, For time step The advantage function, The policy entropy. This is achieved through gradient descent.

[0131] renew ,from Sample batch data Calculate the value loss ,renew .

[0132] Using generalized advantage estimation (GAE), the reward difference and value difference during the state transition process are fitted, and the advantage function is calculated:

[0133] ;

[0134] in, This is the GAE coefficient.

[0135] from Sample batch data Calculate the probability ratio of the old and new strategies. Based on the pruning strategy loss function, the exploration and utilization of policy updates are balanced to calculate the value loss:

[0136] ;

[0137] in, For loss function, For state Value network output, For time step Instant rewards As a discount factor, This is a termination marker.

[0138] Parameter updates use the Adam optimizer with a learning rate of [missing information]. The weight decays to The gradient clipping threshold is 0.8.

[0139] Huber loss is defined as:

[0140] ;

[0141] in, The value network parameters are updated using gradients to set an error threshold. .

[0142] After training is complete, perform the following output and verification steps:

[0143] Output deployment plan:

[0144] After training, the policy network Output the optimized deployment parameters for the AP, as follows:

[0145] Channel allocation results: Outputs the channel allocation for each AP. ,in For the corresponding 2.4GHz frequency band channels, the operating channel of each AP is clearly defined in order to make reasonable use of spectrum resources and reduce inter-channel interference.

[0146] Transmit power configuration: Provide the transmit power for each AP. ,in It is derived from a discrete power set, and by optimizing the power, it reduces unnecessary power loss and signal interference while meeting coverage requirements.

[0147] Deployment location information: Outputs the deployment location of each AP. ,in, and The location is determined based on the coordinates of the two-dimensional plane within the substation's indoor environment, combined with a training-derived position adjustment strategy. By accurately characterizing the spatial deployment coordinates of the access points (APs), and coordinating with channel and power configurations, coverage and signal quality are ensured, thereby optimizing overall communication performance.

[0148] This embodiment first constructs a system scenario, using a 100m×100m two-dimensional rectangular area as an indoor substation scenario, uniformly divided into 10×10 grids, with a target point TP at the center of the grid, deploying 12 APs and considering obstacle occlusion; then, a channel model is constructed, integrating line-of-sight and non-line-of-sight path loss models in the 2.4GHz band; next, a resource allocation model is designed based on the Shannon-Hartley theorem; then, the decision-making process is modeled as a Dec-POMDP seven-tuple, with the agent being the AP; finally, a reward function is designed, and the MAPPO algorithm is used to output the joint optimization results of AP location, power, and channel.

[0149] An indoor substation AP deployment system based on MAPPO includes:

[0150] Scene modeling module: Used to build indoor substation system scene models and define data transmission processes;

[0151] Channel model construction module: Analyzes path loss and channel overlap to construct a downlink channel model;

[0152] Resource allocation calculation module: Establishes a resource allocation model that maximizes the total system rate, determines AP power, channel allocation and bandwidth allocation rules, and outputs a resource allocation scheme;

[0153] Decision model transformation module: This module is used to transform the joint optimization problem of AP deployment and resource allocation into a decentralized partially observable Markov decision model, clarifying the state transition logic.

[0154] Reward function building module: used to construct multi-objective reward functions;

[0155] MAPPO algorithm optimization module: used for iterative optimization strategy, outputting the optimized AP location, power, channel and bandwidth allocation scheme;

[0156] Data interaction and storage module: Used to support the data flow between modules, store the parameters of each module, and provide users with a data query interface.

[0157] An indoor substation AP deployment device based on MAPPO includes a memory and a processor. The memory stores a computer program, which is executed by the processor when loaded.

[0158] A readable storage medium storing a computer program suitable for executing the above-described MAPPO-based indoor substation AP deployment method when loaded by a processor.

[0159] like Figure 2 The diagram shown is a model of the indoor substation wireless access point deployment method according to this embodiment, illustrating a WLAN in an indoor substation scenario based on IEEE 802.11ax. The system is located in a two-dimensional rectangular area. , .Will Evenly divided into There are 3 squares, each with the same length and width. The geometric center of each square is considered the target point, denoted as... Target point The coordinates are Use AP to cover TP, Optimize the deployment of APs of the same type, where the AP set is represented as... The coordinates are .

[0160] In the indoor environment of a substation, due to the presence of numerous electrical and metal obstacles, communication between an access point (AP) and a target point is not always line-of-sight; non-line-of-sight transmission also exists. Each target point is associated with only one AP with the strongest signal, communicating with it using the same channel as that AP. The APs and their associated target points form a Basic Service Set (BSS), and all nodes belonging to the same BSS operate on the same channel. In a network of multiple APs, adjacent APs may operate on partially overlapping channels or even the same channel, leading to mutual interference between adjacent BSSs.

[0161] like Figure 3-7 The image shows the simulation results for the main power distribution room scenario performance verification.

[0162] like Figure 3 The diagram shown is a MAPPO deployment topology for the main power distribution room (100m × 100m, 12 APs) in this embodiment. In the 100m × 100m main power distribution room scenario, the MAPPO algorithm achieves coordinated optimization of AP deployment and resource allocation through hierarchical decision-making. The 12 APs are evenly distributed in non-obstacle areas (avoiding obstacles such as...). Within a 50m x 50m sub-area (metal cabinet), the number of APs is ≤3, and the target point association success rate reaches 100%. Compared to the Greedy algorithm's random deployment leading to AP overlap (e.g., 3 APs clustered in the upper left corner), the MAPPO algorithm, through the uniformity term in the reward function, increases the average distance between APs to 28.5m, maintaining a minimum distance of 12.3m, effectively reducing co-channel interference. This topology avoids obstacle occlusion through Shapely geometric calculations, while simultaneously using the uniformity term in the reward function to guide AP distributed deployment, achieving synergistic optimization of coverage balance and interference suppression.

[0163] like Figure 4 The figure shows a comparison of the channel overlap rates of the algorithms in the main power distribution room (100m×100m, 12 APs) of this embodiment. The figure compares the channel overlap rate convergence curves of the MAPPO, Greedy, IAC, and IQL algorithms. The MAPPO algorithm, with its channel spacing ∆≥5 allocation strategy (e.g., AP1 selects channel 1, AP2 selects channel 6), achieves a stable average overlap rate convergence to 0.25, a 66.67% reduction compared to Greedy's 0.75, and reductions of 34.21% and 16.67% compared to IAC's 0.38 and IQL's 0.30, respectively. This performance stems from the policy network's control over channel overlap. Dynamic optimization – when the channel spacing between adjacent APs is detected At the same time, a reward and punishment mechanism was used to force a channel switch, such as changing channel 3 to channel 8, which effectively reduced co-channel and adjacent channel interference and verified the effectiveness of the 2.4GHz band non-overlapping channel allocation strategy.

[0164] like Figure 5 The figure shown is a graph illustrating the algorithmic transmit power variation in the main power distribution room (100m×100m, 12APs) of this embodiment. The MAPPO algorithm achieves fine-grained power management through RU dynamic allocation, with the average transmit power stabilized at 5dBm, a 68.75% reduction compared to the Greedy algorithm's 16dBm, and all power values ​​satisfying... Constraints, all communication links satisfy The requirements are met. Compared with the IAC algorithm, MAPPO reduces power fluctuation by 42%, thanks to the centralized value network's collaborative optimization of global disturbances, avoiding the local optima problem of the independent Actor-Critic architecture.

[0165] like Figure 6 The figure shows the overall rate change of the algorithm system in the main power distribution room (100m×100m, 12 APs) of this embodiment. The figure shows that MAPPO stabilizes at 25Mbps after training, representing improvements of 66.67%, 38.89%, and 13.64% compared to Greedy's 15Mbps, IAC's 18Mbps, and IQL's 22Mbps, respectively. Its core advantage stems from the CTDE architecture's centralized value function for capturing inter-AP channel overlap interference, combined with a dynamic RU allocation strategy—when the number of target service points... At that time, allocation This combination achieves an average channel bandwidth utilization of 90.4% for 20MHz. Furthermore, the system rate weight in the reward function... The guidance strategy prioritizes optimizing data transmission efficiency.

[0166] like Figure 7 The figure shown is a MAPPO bandwidth utilization diagram for the main power distribution room (100×100m, 12 APs) in this embodiment. This diagram illustrates the efficient utilization strategy of the MAPPO 20MHz channel. When an AP serves one target point, the bandwidth allocation is... (Bandwidth 20MHz); when serving 9 target points, allocate 9... (Bandwidth 18MHz), with an average utilization rate of 90.4%. This strategy strictly adheres to OFDMA technical specifications, such as supporting up to 9 [unclear - possibly referring to specific technologies or features] in the 20MHz band. Or 5 And the bandwidth utilization and target point rate balance are achieved through a greedy RU allocation rule.

[0167] Figure 8 — Figure 10 This is a verification of the scene adaptability and region scaling of the simulation results.

[0168] like Figure 8The figure shows a comparison of round rewards for the main power distribution room (100m×100m, 12 APs) in this embodiment. This figure reflects the long-term reward convergence performance of the algorithm. The average round reward of MAPPO is stable at around 50, which is 42%–139% higher than Greedy's 21, IAC's 30, and IQL's 35. This reward consists of system rate (60%), AP uniformity (30%), and constraint penalty (10%). The reward advantage stems from the fact that the centralized value function incorporates indicators such as total system rate and AP uniformity into the optimization objective: when AP deployment is too dense, the uniformity term in the reward function reduces the round reward, driving APs to be deployed more dispersedly to reduce interference, thereby achieving a globally optimal strategy.

[0169] like Figure 9 The image shows a comparison of round rewards for the algorithm in the transformer room (100m×100m, 12 APs) scenario of this embodiment. After switching to the 100m×100m transformer room scenario, MAPPO's average round reward stabilized at 47.5, an improvement of 25%–164% compared to Greedy's 18, IAC's 35, and IQL's 38. Dynamic path loss modeling is key: when there are obstacles between the AP and the target point, it automatically switches to the NLoS model (path loss index is adjusted from 1.45 to 3.15), and updates the value function in real time through advantage estimation. In contrast, algorithms like Greedy suffer from SINR calculation bias due to fixed parameters; when the proportion of non-line-of-sight paths exceeds 60%, the total system rate decreases by 35%. In this scenario, MAPPO uses PopArt normalization technology to address the standard deviation of multipath fading caused by metal obstacles. From 2.25 to 3.19, the round reward variance was significantly better than IQL.

[0170] like Figure 9 The diagram shows a comparison of round rewards for the algorithm in the main power distribution room (90m×90m, 9 APs) of this embodiment. After shrinking the main power distribution room to 90m×90m, the average round reward for MAPPO in the deployment of 9 APs and 81 target points reached 60, an improvement of 22%–243% compared to Greedy's 17.5, IAC's 48, and IQL's 49. The AP deployment density increased from 0.12 / 100m² to 0.11 / 81m². MAPPO avoids over-concentration through sub-region constraints (≤2 APs per 30m×30m sub-region). The power control strategy is refined: allocation is based on a distance of ≤15m from the AP. AP = 15-30m allocation AP > 30m allocation This keeps the average power at 5dBm.

[0171] This invention has been described through embodiments. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of this invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, this invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of this invention.

Claims

1. A method for deploying access points (APs) in an indoor substation based on MAPPO, characterized in that, Includes the following steps: Based on the actual environmental characteristics of indoor substations, an indoor substation system scenario model is established; Based on the indoor substation system scenario model, a path loss model under non-uniform medium is constructed by analyzing path loss and channel overlap, and a downlink channel model is constructed by using the path loss model and channel overlap. Based on the downlink channel model, a resource allocation model that maximizes the total system rate is established. Based on the resource allocation model, by defining a Markov decision process, the problem of wireless access point (AP) deployment and resource allocation is transformed into a decentralized partially observable Markov decision process Dec-POMDP decision model. Construct a multi-objective reward function based on the Dec-POMDP decision model; Based on the Dec-POMDP decision model and multi-objective reward function, the MAPPO algorithm is optimized using a multi-agent near-end policy to output AP location, power, channel and bandwidth allocation.

2. The MAPPO-based indoor substation AP deployment method of claim 1, wherein, Based on the actual environmental characteristics of indoor substations, an indoor substation system scenario model is established, including: Define the indoor substation scenario as a two-dimensional rectangular area. ,Will Evenly divided into A set of squares, with the geometric center of each square as the target point set. The target point is the specific location that needs to be covered by the AP in the indoor substation scenario. , The coordinates are ;deploy A service set consists of multiple access points of the same type. , Coordinates are .

3. The MAPPO-based indoor substation AP deployment method of claim 1, wherein, Based on an indoor substation system scenario model, a path loss model under non-uniform media is constructed by analyzing path loss and channel overlap. A downlink channel model is then constructed using the path loss model and channel overlap, including: Analyze path loss and construct a path loss model for non-uniform media: ; in, For reference distance Loss at the point, , At the speed of light, For carrier frequency, for With the target point Path types between, when With the target point The path between them is the line-of-sight (LoS) path. ;when With the target point The path between them is a non-line-of-sight (NLoS) path. , for and The distance between them , Let be the path loss exponents for LoS and NLoS, respectively. , The mean is 0 and the standard deviation is for LoS and NLoS respectively. Gaussian random variables; Defining binary variables Indicating Channel allocation: ; Channel spacing is defined as a non-overlapping channel and vice versa.

4. The MAPPO-based indoor substation AP deployment method of claim 3, wherein, Based on the downlink channel model, a resource allocation model that maximizes the total system rate is established, including: Definitions Data rate for the target point is: ; in, express For target point The generated data rate is Mbps. express For target point The allocated channel bandwidth is MHz. express For target point The generated signal-to-interference-plus-noise ratio (SIR) in dB is calculated using the following formula: ; in, For target point from The received signal strength, for The transmission power, This refers to the noise power in the indoor environment of a substation. Other adjacent For target point The generated interference power, For two channels and The degree of overlap is 0-1.

5. The indoor substation AP deployment method based on MAPPO according to claim 1, characterized in that, Based on the resource allocation model, a Markov decision process is defined, including: Define the seven-tuple of the Markov decision process Dec-POMDP: ; Among them, intelligent agents For the AP set, global state Includes parameters such as AP location, power, channel, and loss; action space. , For channel set, For power discretization set, Adjusting the discretized set for position coordinates, local observation Include Information about itself and its associated target points; The state transition function is derived from the physical model. drive; For a multi-objective reward function, it is a weighted sum of the system's total rate, uniformity, and penalty term; This is the discount factor.

6. The MAPPO-based indoor substation AP deployment method of claim 1, wherein, Based on the Dec-POMDP decision model, a multi-objective reward function is constructed: ; wherein, is a total system rate reward, is a uniformity reward, is a constraint penalty, is a weighting factor.

7. The MAPPO-based indoor substation AP deployment method of claim 1, wherein, Based on the Dec-POMDP decision model and multi-objective reward function, the MAPPO algorithm is trained using a pruning policy loss with multi-agent proximal policy optimization. ; in, For probability ratios, This is the cutting factor. The entropy coefficient, For time step The advantage function, For policy entropy; Calculate the loss of value: ; in, For loss function, For state Value network output, For time step Instant rewards As a discount factor, This is the end marker; Calculate the advantage function: ; wherein GAE coefficients are generalized advantage estimation.

8. A MAPPO-based indoor substation AP deployment system, characterized by, include: Scene modeling module: Used to build indoor substation system scene models and define data transmission processes; Channel model construction module: Analyzes path loss and channel overlap to construct a downlink channel model; Resource allocation calculation module: Establishes a resource allocation model that maximizes the total system rate, determines AP power, channel allocation and bandwidth allocation rules, and outputs a resource allocation scheme; Decision model transformation module: This module is used to transform the joint optimization problem of AP deployment and resource allocation into a decentralized partially observable Markov decision model, clarifying the state transition logic. Reward function building module: used to construct multi-objective reward functions; MAPPO algorithm optimization module: used for iterative optimization strategy, outputting the optimized AP location, power, channel and bandwidth allocation scheme; Data interaction and storage module: Used to support the data flow between modules, store the parameters of each module, and provide users with a data query interface.

9. A MAPPO-based indoor substation AP deployment apparatus, characterized by, It includes a memory and a processor, the memory being used to store a computer program, which, when loaded by the processor, is used to execute the method of any one of claims 1-7.

10. A readable storage medium, characterized by, The storage medium stores a computer program that is adapted to execute the method described in any one of claims 1-7 when loaded by a processor.