A base station site selection and antenna parameter joint optimization method

By employing a multi-agent deep reinforcement learning method, the joint optimization of base station site selection and antenna parameters is achieved, solving the problems of global optimality and computational complexity in complex scenarios in existing technologies, and realizing efficient coverage optimization and cost reduction.

CN122340495APending Publication Date: 2026-07-03CHINA INFOMRAITON CONSULTING & DESIGNING INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA INFOMRAITON CONSULTING & DESIGNING INST CO LTD
Filing Date
2026-04-01
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies struggle to guarantee global optimality and scalability in complex scenarios when optimizing coverage in residential communities. They suffer from high computational complexity, difficulty in balancing multi-dimensional indicators, and a lack of self-learning capabilities in response to dynamic changes in the environment and business.

Method used

A multi-agent deep reinforcement learning approach is adopted to jointly optimize base station site selection and antenna parameters. Through 3D raster modeling, link budget and antenna parameter constraints, coverage determination and index statistics, combined with multi-agent deep reinforcement learning algorithm, the joint optimization of base station site selection and antenna parameters is achieved.

Benefits of technology

It significantly reduces the overlap coverage ratio, increases the coverage rate, improves the training convergence speed and policy stability, achieves collaborative optimization of indoor and outdoor user perception, and reduces computing and deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122340495A_ABST
    Figure CN122340495A_ABST
Patent Text Reader

Abstract

This invention discloses a method for joint optimization of base station site selection and antenna parameters, comprising: Step 1, acquiring data of the target area and performing three-dimensional raster modeling to obtain a set of candidate sites; Step 2, calculating the received power of the reference signal received by the raster from the antenna site to ensure that the power meets the standard and that the raster effectively covers the antenna main lobe directivity requirements; Step 3, defining core performance indicators to comprehensively evaluate the antenna deployment scheme; Step 4, treating each candidate site as an intelligent agent, and based on a global-local hybrid reward mechanism, obtaining a globally approximately optimal deployment scheme through autonomous interaction and collaborative learning between the intelligent agent and the environment; Step 5, based on the trained model, inputting cell parameters and outputting the optimal base station site selection and antenna parameter setting scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication network planning and optimization technology, and in particular to a method for joint optimization of base station site selection and antenna parameters. Background Technology

[0002] Optimizing coverage in residential areas requires addressing issues such as continuous outdoor-indoor coverage, building penetration, antenna directivity constraints, and multi-site coupling. Existing technologies mainly include:

[0003] (1) Experience-driven approach: This approach relies on engineers' experience combined with building and user distribution to determine site selection and parameters. This method can quickly provide feasible solutions, but it is difficult to guarantee global optimization and scalability in complex scenarios, and it is highly sensitive to personnel experience.

[0004] (2) Heuristic and traditional optimization methods: Common methods include greedy algorithms based on maximizing coverage, simulated annealing, and genetic algorithms. These methods can improve the site selection efficiency to some extent, but they are often computationally complex and difficult to balance among multi-dimensional indicators (such as coverage, overlap rate, and user signal quality).

[0005] (3) Simulation and ray tracing method: Relying on the fine propagation model to screen candidate points, the single-point evaluation accuracy is high, but the global search overhead is huge, making it difficult to apply quickly in large-scale cells; at the same time, it lacks the ability to learn from dynamic changes in the environment and business. Summary of the Invention

[0006] Purpose of the invention: The technical problem to be solved by the present invention is to address the shortcomings of the prior art by proposing a joint optimization method for base station site selection and antenna parameters.

[0007] To address the aforementioned technical problems, this invention discloses a method for joint optimization of base station location and antenna parameters, comprising the following steps:

[0008] Step 1: Data acquisition and 3D raster modeling;

[0009] Step 1-1: Collect relevant basic data of the target cell, including:

[0010] Geographic and architectural information: building outline, floor height, roof height, etc.;

[0011] RF and antenna parameters: including the operating frequency band of the base station equipment to be deployed, single carrier transmit power, antenna gain, and horizontal and vertical 3 dB beamwidth of the antenna, etc.

[0012] Engineering and propagation parameters include feeder loss, passive device insertion loss, shadow fading margin, system interference margin, and penetration loss in indoor scenarios (such as penetration loss between glass and walls).

[0013] Steps 1-2: 3D raster modeling: (e.g.) Figure 2 As shown, a three-dimensional coordinate system is established with the north direction of the residential area as the positive Y-axis. The entire three-dimensional space of the residential area is divided into N identical and closely adjacent cuboid grids. Each grid is determined by the three-dimensional coordinates of its center position. definition, Number the grid cells; differentiate between "indoor" and "outdoor" grid cells based on building geometry, and record... ,in =1 indicates an indoor grid; otherwise, it is an outdoor grid.

[0014] Steps 1-3: Pre-set candidate antenna locations: Based on the three-dimensional gridded cell model constructed in Step 1-2, generate a set of candidate antenna deployment locations. The total number of them is M, where each candidate point is denoted as M. , Numbering candidate points; this method comprises the following two aspects:

[0015] 1) Feasible points on the roof: Based on the geographical and architectural information of the community, extract the roof plan polygons of each building. Using the geometric center of the roof as the primary candidate point, supplementary candidate points can be added at certain intervals along its edges. The installation height of each roof point is determined. Add the preset support height to the roof height;

[0016] 2) Roadside light pole locations: Considering the advantages of rapid deployment using existing infrastructure in actual projects, existing roadside light poles in the community are used as candidate locations to avoid the cost and complexity of constructing new poles. The installation height of each light pole location is determined. Take the actual measured or standard height of the light pole directly.

[0017] Step 2: Link Budget and Antenna Parameter Constraints

[0018] Step 2-1, Antenna Candidate Locations With grid The Euclidean distance is:

[0019]

[0020] To match the subsequent path loss constant, the distance is converted to km when substituted into the path loss formula;

[0021] Step 2-2: Use a logarithmic distance model to evaluate the average path loss over medium to long distances.

[0022]

[0023] in, 32.4 represents the actual operating carrier frequency of the system (MHz); 32.4 represents the carrier frequency at which... Euclidean distance (km) Calibration constant in the unit system; The higher the frequency, the greater the free space attenuation; n is the path loss exponent. For urban residential areas, n≈3.19, which reflects the additional attenuation caused by building layout, vegetation shading, etc. in residential areas. It is between pure line-of-sight and non-line-of-sight and is a reasonable and empirical value.

[0024] Steps 2-3: Subtract various engineering losses and propagation losses from the transmit power and antenna gain, and calculate the received power of the reference signal from antenna point i to grid j. :

[0025]

[0026] In the formula, P tx Where PL1 is the transmit power, G is the antenna gain, PL2 is the feed line loss, and PL3 is the passive component insertion loss. The path loss is calculated in step 2-2; PL3 and PL4 are the glass and wall penetration losses, respectively, only when the grid... =1 indicates that the calculation is performed indoors; PL5 is the shadow fading margin, and PL6 is the interference margin.

[0027] Steps 2-4: To avoid misjudging sidelobe signals as effective coverage, it is necessary to further determine whether the grid is located within the antenna's main lobe radiation range. In this embodiment, in addition to ensuring power compliance, an antenna main lobe directivity determination is added to ensure that only grids located within the antenna's main lobe beamwidth are counted as effective coverage. Figure 3 and Figure 4 As shown, first calculate from the candidate point location Pointing grid relative azimuth and relative downslope :

[0028]

[0029] in ,Will Mapped to This ensures the continuity of angle values;

[0030]

[0031] in, The angle of inclination relative to the horizontal plane, with downward being positive and upward being negative, and the range of values ​​is [not specified]. .

[0032] Let the current configuration parameter of antenna i be the azimuth angle. With downhill angle The horizontal and vertical 3 dB beamwidths of the antenna are respectively , The criterion for determining whether grid j is located within the main lobe range of antenna i is:

[0033]

[0034]

[0035] Here, represents the horizontal angular deviation of grid j relative to the center of the main lobe of antenna i; and represents the vertical angular deviation of grid j relative to the center of the main lobe of antenna i. Only when both conditions are met simultaneously is grid j considered to be within the effective coverage area of ​​the main lobe of antenna i. This mechanism internalizes the antenna's directional constraints into the optimization process and is a core element in ensuring that the solution's output results are highly consistent with real network perception.

[0036] Step 3: Coverage Determination and Indicator Statistics

[0037] This step is a crucial bridge connecting the underlying signal calculation and the upper-level optimization algorithm. Based on the power, azimuth, and downtilt angle determination criteria calculated in step 2, the coverage status of each three-dimensional grid is accurately determined. On this basis, a multi-dimensional and quantifiable network performance evaluation index system is constructed as the target of subsequent optimization algorithms, directly guiding the agent's decision-making.

[0038] Step 3-1: To ensure that network coverage quality is consistent with the actual user experience, effective coverage is defined using dual constraints of power and directionality, and a coverage determination function is given:

[0039]

[0040] in As a power threshold, to ensure that the received signal strength within the grid meets the terminal requirements, −95 dBm is a typical engineering value for the 3.5 GHz band. This value can be flexibly adjusted according to different service requirements and frequency band characteristics; directivity constraint. This ensures that the signal comes from the effective radiation range of the antenna main lobe, and avoids including unstable signals from the antenna side lobes or back, even if the power is momentarily up to standard, in the effective coverage, so that the coverage assessment results are highly consistent with the "stable and dominant" signal quality actually perceived by users. This indicates that grid j is well covered by antenna i; otherwise, it indicates that the grid is not covered or the coverage requirement is not met.

[0041] Step 3-2: Based on the decision function in Step 3-1, statistical analysis of the global coverage is performed, and a series of core performance indicators are defined. These indicators together constitute a comprehensive evaluation of the antenna deployment scheme.

[0042] Raster overlay set: Let J be the set of coverage sources for grid j, containing the antenna numbers of all antennas that can provide high-quality coverage for grid j. This indicates how many antennas simultaneously cover grid j;

[0043] Coverage count: Count the total number of grid cells in the entire target cell that are covered by at least one antenna with high quality;

[0044] Overlap count: The number of grids simultaneously covered by multiple antennas is counted. This metric is a key parameter for measuring potential co-channel interference and resource waste within the network, and optimization algorithms need to minimize it.

[0045] Coverage:

[0046]

[0047] The overlap rate is:

[0048]

[0049] The total number of grids is N. The coverage rate comprehensively measures the breadth and quality of network coverage. It is the final result after balancing coverage and interference and is the primary goal of the optimization algorithm. The overlap rate is the proportion of grids with overlapping coverage to the total number of grids. It reflects the severity of potential interference in the network and is an important indicator for measuring coverage quality.

[0050] Step 4: Multi-agent deep reinforcement learning

[0051] Step 4-1: Multi-agent environment modeling

[0052] The specific technical solution of the present invention is: an antenna location and parameter joint optimization scheme based on multi-agent deep reinforcement learning, which regards each candidate point as an agent, and through its autonomous interaction and collaborative learning with the environment, finally obtains a globally approximately optimal deployment scheme.

[0053] Intelligent agent: Each candidate antenna point i is considered as an independent intelligent agent;

[0054] Local observation of intelligent agents : Geometric and engineering attributes of the current candidate point, such as its location coordinates and installation height; the agent's previous action decision and the coverage rate under this action;

[0055] Actions of the intelligent agent : Whether the candidate point is enabled, the antenna azimuth angle setting, and the antenna downtilt angle setting; wherein, the azimuth angle and downtilt angle are both represented by discrete settings, and their value range and discrete step length can be preset according to different cell scenarios, frequency band conditions and engineering implementation requirements. This invention does not limit them.

[0056] The global environmental state s is obtained by environmental aggregation and includes observation information of all agents as well as non-local information, grid statistics, global coverage and overlap statistics, indoor weights, etc. It is only used for centralized value evaluation during the training period.

[0057] Reward Function: The reward function is a weighted sum of global and local rewards. The global reward considers coverage, overlap, and construction cost simultaneously.

[0058]

[0059] Among them, parameters Used to maximize coverage; parameters Treat the overlap rate as a deduction item to reduce overlapping coverage; The number of candidate antenna sites currently in use is adjusted by... The strategy optimizes the coverage by taking into account construction costs, aiming to achieve optimal coverage with fewer antennas.

[0060] Local rewards take into account the coverage contribution and interference penalty of the base station itself.

[0061]

[0062] Among them, the number of grids covered by agent i The number of overlapping grid cells caused by agent i N is the total number of grid cells, λ1 and λ2 This represents the local reward weighting coefficient.

[0063] To address the credit allocation problem in multi-agent cooperation, this invention employs a hybrid reward mechanism, where the reward function is a weighted average of global and local reward terms.

[0064]

[0065] in, and As adjustable weight coefficients, the local reward term is only related to the coverage contribution and interference of agent i itself. This design enables the agent to quickly learn the basic coverage strategy based on local feedback in the early stage of training, and to coordinate the cooperative avoidance among multiple base stations by relying on the global reward term in the later stage of training, thereby significantly shortening the training time while ensuring global optimality.

[0066] Step 4-2: Multi-agent deep reinforcement learning algorithm

[0067] This invention employs the existing Multi-Agent Proximal Policy Optimization (MAPPO) algorithm as the foundational deep reinforcement learning framework. However, it constructs a specific state space, action space, and hybrid reward function structure for base station location scenarios, thus forming a dedicated policy model suitable for joint optimization of wireless networks. Based on this, a hybrid reward incentive mechanism for multiple base stations is constructed. Unlike conventional reinforcement learning that uses global rewards, this invention introduces a local reward shaping architecture based on coverage contribution and interference penalty terms. Specifically, in the centralized training phase, the Critic network evaluates the overall collaborative effect based on the global state; in the distributed execution phase, the reward function r of each agent... i It is composed of a weighted average of global coverage metrics and local independent coverage metrics. It solves the convergence difficulty problem caused by unclear credit allocation in traditional multi-agent algorithms in sparse reward scenarios such as base station site selection, and realizes spontaneous collaboration among base station groups under non-interactive conditions.

[0068] Actor Network (Policy Network): Each agent has its own Actor Network, which receives local observations from the agent. Output the probability distribution in its action space. This refers to the probability of selecting each action. The agent samples actions based on this probability and executes them accordingly.

[0069] Critic Network (Value Network): During the training phase, a centralized Critic Network exists to receive global state information s and the actions of all agents. The output is the evaluation of the global state-action pair, namely the state value function V(s) or the action advantage function A(s,a).

[0070] The training process can be divided into the following steps:

[0071] 1. Interactive Sampling: All agents interact with the environment according to the current policy, collecting a large amount of experience trajectory data. The data is stored in a shared sampling trajectory cache for use in this round of policy updates.

[0072] 2. Dominance estimation: Using the Critic network and the generalized dominance estimation method, the dominance function value A(s,a) at each time step is calculated to quantify the superiority or inferiority of taking joint action a relative to the average level in the global state.

[0073] 3. Strategy Update: Update the network parameters θ of each Actor by maximizing the following clipping objective function.

[0074]

[0075] in, It represents the probability ratio between the old and new strategies. The min and clip operations work together to ensure the stability of the strategy updates, preventing the learning effect from being ruined by excessively large single-step updates.

[0076] 4. Value Update: The parameters of the Critic network are updated by minimizing the mean squared error. ,

[0077]

[0078] Where V target It is the value of the target value network, used for stable training.

[0079] Once training is complete, during the deployment phase, only the Actor network for each agent is needed, with each agent basing its local observations on its own data. It can make decisions independently without needing to communicate with other intelligent agents in real time or access the global state, making the location scheme generation process efficient, low-latency, and easy to implement in a distributed manner.

[0080] Step 5: Based on the trained model, input cell parameters and output the optimal base station site selection and antenna parameter setting scheme.

[0081] The policy model obtained through training and convergence in step 4-2 is applied to the target residential community scenario. The policy model is the Actor network (policy network) in step 4-2, whose input is the local observations of the agent defined in step 4-1, and whose output is the agent's action (whether to enable, azimuth angle setting, downtilt angle setting). Step 4-1 is only used to define the multi-agent environment and observations / actions / rewards, and is not equivalent to the policy model itself; step 4-2 trains the policy model on the environment defined in step 4-1, and step 5 calls the trained policy model for inference output. No manual parameter tuning is required; only the basic parameters of the community are input, and the engineered configuration results of antenna site selection and directivity parameters can be generated in one go, with the output aperture consistent with the evaluation aperture during the training period, facilitating direct construction and acceptance.

[0082] Step 5-1: Parameter Input and Consistency Verification

[0083] The operator only needs to provide the following information: 3D building and geographic data of the community, a preset set of candidate locations, service frequency bands and transmit power, antenna specifications, default values ​​for coverage threshold and engineering loss, and deployment budget ceiling. Based on this, the system completes the rasterization, link budget, and coverage determination in steps 1–3, and constructs local observations of the agent consistent with those in the training phase as input for model inference.

[0084] Step 5-2: Model Reasoning and Solution Generation

[0085] The system generates local observations for each candidate point agent on the target cell using the same observation construction method as during the training phase. These local observations are then input into the pre-trained Actor policy model, which outputs the corresponding actions for each agent, thereby automatically outputting the configuration for each antenna entity to be deployed.

[0086] Location selection: Select the actual location coordinates to be used from the candidate point set;

[0087] Directional parameters: Provides the azimuth and tilt angle settings for each enabled position;

[0088] Adaptive Quantity: The model will automatically determine the actual number of antennas to be used, avoiding the need to pre-define the required M antennas.

[0089] Consistency Note: If the local observation defined in step 4-1 includes the "previous action decision" and the "coverage rate under this action," then the deployment inference phase should adopt interactive rolling inference, consistent with the training phase: initialize the previous action (initialize the rule), call the coverage determination in step 3 to obtain the corresponding coverage statistics as part of the observation; then repeat several rounds of inference updates until the action is stable or the preset number of rounds is reached, and then output the final solution. If the local observation does not include the above historical items, then the observation can be constructed directly in a single round and the inference output can be completed in one go.

[0090] This process does not require manual searching and fine-tuning, and the inference time depends only on the size of the candidate points and the model size, making it suitable for large-scale engineering applications.

[0091] Step 5-3, Output Format and Deliverables

[0092] 1) Deployment list: Coordinates, azimuth, and downtilt angle of each antenna site, in CSV / Excel format, which can be used directly for construction layout and parameter activation;

[0093] 2) Coverage assessment appendix: A summary table of key indicators such as coverage, overlap, and reference signal received power, to facilitate comparison with existing or baseline schemes;

[0094] 3) Version and traceability information: Model version identifier, input parameter summary and timestamp, to support solution reproduction and subsequent auditing.

[0095] Beneficial effects:

[0096] (1) Joint optimization and reduced overlap: Compared with simulation / ray tracing and heuristic step-by-step parameter tuning, this scheme learns directly in the joint space of position and angle, which is more global and significantly reduces invalid overlap and improves effective coverage.

[0097] (2) More consistent indoor perception: Penetration loss is explicitly modeled and constrained together with the main lobe angle window, and indoor coverage is improved in sync with the average optimal RSRP.

[0098] (3) Faster implementation: The candidate point location, azimuth angle, and downtilt angle are discrete and the construction level is determined. After reasoning, only conflict resolution and quantification are needed to generate the drawing; there is zero communication during the execution period, and the calculation and deployment costs are low.

[0099] (4) Generalization and robustness: The transferable strategy across building types and frequency bands has lower recalculation and readjustment requirements than traditional solutions and has strong reusability.

[0100] By constructing a three-dimensional integrated indoor and outdoor main lobe-constrained coverage determination model and proposing a multi-agent deep reinforcement learning collaborative optimization architecture based on a global-local hybrid reward mechanism, joint decision optimization of base station site selection, azimuth angle, and downtilt angle is achieved. Compared with existing methods, this invention can significantly reduce the overlapping coverage ratio, improve coverage rate, and enhance training convergence speed and policy stability while ensuring engineering feasibility. Attached Figure Description

[0101] Figure 1 Overall process diagram.

[0102] Figure 2 A schematic diagram of the three-dimensional rasterization of the residential area.

[0103] Figure 3 Azimuth determination diagram.

[0104] Figure 4 Diagram illustrating the determination of the downtilt angle.

[0105] Figure 5 A schematic diagram of the structure of a multi-agent deep reinforcement learning algorithm.

[0106] Figure 6 The curves showing the changes in coverage and overlap coverage during the training process of the method of this invention. Detailed Implementation

[0107] To address the pain points in cell coverage optimization, such as insufficient integrated indoor / outdoor modeling, lack of joint optimization, difficulty in coordinating multiple objectives, and weak generalization ability, this invention proposes a joint optimization method for antenna location and beam parameters based on multi-agent deep reinforcement learning.

[0108] (1) Three-dimensional coverage modeling integrating indoor and outdoor: In the three-dimensional gridded cell model, penetration loss and shadow fading are introduced, and a unified coverage determination function is formed by receiving power threshold and antenna direction determination to ensure the coordinated optimization of indoor and outdoor user perception.

[0109] (2) Joint optimization of location, azimuth and downtilt angle: By modeling each antenna as an agent through multi-agent modeling, the position and antenna directivity parameters are weighed at the global scale at one time by multi-agent deep reinforcement learning algorithm, so as to reduce overlap and improve effective coverage.

[0110] (3) Endogenization of constraints that can be implemented in engineering: Engineering angle constraints such as the horizontal / vertical 3dB beamwidth of the antenna are embedded in the optimization and judgment process; the range of values ​​for azimuth and downtilt is discretely quantized to be consistent with engineering feasibility.

[0111] (4) Multi-objective reward design: Construct a coverage gain term with coverage rate as the core and a penalty term with overlap count as the core, and introduce indoor weight and received signal power to achieve an adjustable trade-off between coverage, overlap and experience.

[0112] Example:

[0113] To verify the effectiveness of the proposed method for joint optimization of base station location and antenna parameters, a typical high-rise dense residential community was selected as a simulation verification scenario, and the modeling, training and deployment scheme were performed according to the process described in steps 1 to 5 of the specification.

[0114] 1. Simulation Scene Construction

[0115] A typical high-rise dense residential community was selected as the simulation scenario. The community area is approximately 200m × 150m, and contains 6 residential buildings, numbered B1–B6. Each building has a floor height of approximately 3m and 12 floors, with an average distance of approximately 20m between buildings.

[0116] Following the three-dimensional rasterization modeling method described in steps 1-2, a three-dimensional coordinate system is established with the due north direction of the cell as the positive Y-axis, and the target cell space is divided into rectangular raster grids of uniform size. In this embodiment, the grid size is set to 10m × 10m × 3m, and each grid is represented by the three-dimensional coordinates of its center point, denoted as . There are approximately 3,600 grid points in total.

[0117] During the grid division process, the grid attributes are automatically determined based on the building's geometric structure information, and the grid is then divided into:

[0118] Interior grid: Located in the interior area of ​​a building;

[0119] Outdoor grid: Located in the open space area outside the building.

[0120] Each grid point is considered a potential wireless service location for assessing network coverage quality. This 3D gridded modeling method enables refined wireless coverage assessment in complex residential scenarios.

[0121] 2. Candidate antenna location settings and equipment parameters:

[0122] Following the candidate point generation method described in steps 1-3, a set S of candidate antenna points is preset on the roof and both sides of the road, with a total of M=24 candidate points, including 18 candidate points on the roof and 6 candidate points on the roadside lampposts. Each candidate point is denoted as . , where i is the candidate point number.

[0123] The candidate points for rooftops were determined based on the planar geometry of each building's roof. Three candidate points were set for each building, located in the central area of ​​the roof and on both sides where installation was possible, with the mounting height set at the roof height plus 2 meters. The candidate points for roadside light poles were selected from existing light pole locations on the main roads and internal roads of the community, with a uniform mounting height of 8 meters. The three-dimensional coordinates and installation height of each candidate point were recorded for subsequent distance calculations, main lobe direction determination, and action decisions.

[0124] The antenna uses a 3.5 GHz band macro-micro hybrid device with a transmit power of 20 dBm, a horizontal beamwidth of 65°, a vertical beamwidth of 15°, and an antenna gain of 15 dBi. The link budget parameters are shown in the table below:

[0125] Table 1. Link Budget

[0126] Parameter type numerical values illustrate <![CDATA[Feeder loss PL1]]> 1.0dB Typical values ​​of coaxial feeder <![CDATA[Passive device insertion loss PL2]]> 0.5dB Combiner and power divider losses <![CDATA[Shadow fading margin PL5]]> 5dB City Scene Experience Points <![CDATA[Interference Margin PL6]]> 3dB Resource competitive retention value <![CDATA[Glass penetration loss PL3]]> 3dB Typical losses of single-pane windows <![CDATA[Wall penetration loss PL4]]> 8dB Typical loss of concrete wall surface <![CDATA[Received power threshold P th > -95dBm 3.5 GHz Engineering Threshold

[0127] Based on the above parameters, the path loss between each candidate point and the grid and the main lobe directivity are calculated, thereby obtaining the initial signal power distribution matrix.

[0128] 3. Coverage determination and performance index calculation

[0129] For any candidate point i and grid j, first calculate the three-dimensional Euclidean distance between them according to step 2-1. :

[0130]

[0131] To match the subsequent path loss constant, the distance is converted to km when substituted into the path loss formula;

[0132] The logarithmic distance model is used to evaluate the average path loss over medium to long distances:

[0133]

[0134] in, 32.4 represents the actual operating carrier frequency of the system, in MHz; Euclidean distance Calibration constants in a unit system; The higher the frequency, the greater the free space attenuation; n is the path loss exponent. For urban residential areas, n≈3.19, which reflects the additional attenuation caused by building layout, vegetation shading, etc. in residential areas. It is between pure line-of-sight and non-line-of-sight and is a reasonable and empirical value.

[0135] Further, following steps 2-3, the received power of the reference signal from candidate point i to grid j is calculated. :

[0136]

[0137] In the formula, P tx Where PL1 is the transmit power, G is the antenna gain, PL2 is the feed line loss, and PL3 is the passive component insertion loss. The path loss is calculated in step 2-2; PL3 and PL4 are the glass and wall penetration losses, respectively, only when the grid... =1 indicates that the calculation is performed indoors; PL5 is the shadow fading margin, and PL6 is the interference margin.

[0138] After obtaining the received power, proceed with steps 2-4 to determine the antenna main lobe directivity. For each candidate point i and grid j, calculate the relative azimuth and relative downtilt angles respectively.

[0139]

[0140]

[0141] It is then compared with the azimuth angle, downtilt angle, and horizontal and vertical 3 dB beamwidth of the current antenna configuration.

[0142]

[0143]

[0144] A grid is considered effectively covered by candidate point i only if grid j meets the power threshold requirement and is simultaneously located within the horizontal and vertical beamlines of the antenna main lobe. Therefore, a coverage determination function f(i,j) is constructed. If the main lobe directionality constraint is satisfied, then f(i,j)=1; otherwise, f(i,j)=0.

[0145]

[0146] Based on this, the overall coverage is statistically analyzed:

[0147] Raster overlay set: Let J be the set of coverage sources for grid j, containing the antenna numbers of all antennas that can provide high-quality coverage for grid j. This indicates how many antennas simultaneously cover grid j;

[0148] Coverage count: Count the total number of grid cells in the entire target cell that are covered by at least one antenna with high quality;

[0149] Overlap count: The number of grids simultaneously covered by multiple antennas is counted. This metric is a key parameter for measuring potential co-channel interference and resource waste within the network, and optimization algorithms need to minimize it.

[0150] Coverage:

[0151]

[0152] The overlap rate is:

[0153]

[0154] The total number of grids is N. The coverage rate comprehensively measures the breadth and quality of network coverage. It is the final result after balancing coverage and interference and is the primary goal of the optimization algorithm. The overlap rate is the proportion of grids with overlapping coverage to the total number of grids. It reflects the severity of potential interference in the network and is an important indicator for measuring coverage quality.

[0155] 4. Multi-agent deep reinforcement learning modeling

[0156] In this embodiment, following step 4-1, each candidate antenna location is considered as an agent, thus the system contains a total of 24 agents. Each agent makes decisions regarding the deployment and parameters of its corresponding candidate location.

[0157] Local observations for each agent include: the three-dimensional coordinates, installation height, candidate point category, current deployment status, action result of the previous moment, and local coverage statistics of the candidate point.

[0158] The global environmental state s is used for centralized value evaluation during the training phase, and includes information such as local observations of all agents, global coverage, overlap rate, number of currently enabled antennas, and indoor coverage ratio.

[0159] 5. Motion space settings

[0160] To ensure the output results are feasible for engineering applications, this embodiment sets the action space discretized. The action of each agent consists of three parts: whether it is enabled, the azimuth angle setting, and the tilt angle setting.

[0161] in:

[0162] (1) Activate action 0 indicates that no antenna is deployed at this candidate point, and 1 indicates that an antenna is deployed.

[0163] (2) The azimuth angle is discrete in 30° intervals from 0° to 330°, for a total of 12 positions;

[0164] (3) The tilt angle is discrete in 3° intervals from -15° to 15°, with a total of 11 positions.

[0165] When a candidate point is not enabled, its azimuth and downtilt angle actions do not participate in the effective coverage calculation; when a candidate point is enabled, the system determines its final antenna configuration parameters based on the azimuth and downtilt angle settings output by the agent.

[0166] 6. Reward Function Design

[0167] This embodiment employs a hybrid reward mechanism combining global and local rewards as described in step 4-1. The global reward is:

[0168]

[0169] The local reward is:

[0170]

[0171] To address the credit allocation problem in multi-agent cooperation, this invention employs a hybrid reward mechanism, where the reward function is a weighted average of global and local reward terms.

[0172]

[0173] In this embodiment, to balance global collaborative optimization and learnability in the early stages of training, the following approach is adopted: , , , , , , The above parameters are not intended to limit the invention and can be adjusted equivalently for different cell sizes, frequency band conditions, and deployment goals.

[0174] 7. Training process of multi-agent deep reinforcement learning algorithm

[0175] This embodiment uses the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm as the training framework. Each agent has an independent Actor policy network, while a centralized shared Critic value network is set up. During the training phase, the Critic receives global state information to estimate value, while the Actors only receive their own local observations and output action probability distributions.

[0176] First, initialize the environment, Actor network, and Critic network parameters. Then, all agents interact with the simulation environment under the current policy to sample trajectory data, including local observations, joint actions, reward values, and the state at the next time step. Next, calculate the advantage function using the Critic network and the generalized advantage estimation method. Then, update the Actor network parameters by using the PPO clipping objective function and update the Critic network parameters by minimizing the value loss. Repeat the above steps until the global reward converges or the preset number of training rounds is reached.

[0177] In this embodiment, the number of training rounds is set to 3000, and the maximum number of interaction steps per round is set to 20; the learning rates of both the Actor network and the Critic network are set to... The discount factor is set to 0.99; the PPO clipping parameter is set to 0.2; the generalized advantage estimation parameter is set to 0.95; after several rounds of training, the current strategy is validated, and metrics such as coverage, overlap rate, and number of activated sites are recorded.

[0178] As training progresses, the system gradually learns to form a collaborative division of labor among candidate points, thereby achieving the joint optimization goal of "minimum number of sites, maximum coverage, and minimum overlap".

[0179] 8. Training Results and Output Scheme

[0180] Once the training converges, the target cell parameters from step 5 are input into the trained Actor policy model for decision-making. During the execution phase, each agent outputs whether to enable the system and the corresponding azimuth and downtilt angle levels based on its local observations, and the system automatically generates the final deployment plan.

[0181] In this embodiment, the trained and converged strategy model can automatically filter out the actual activation locations from the candidate point set and simultaneously provide the azimuth and downtilt angle configuration parameters for each activation point, so that the coverage and overlap rates achieve a better balance while meeting the engineering implementation constraints. Figure 6 As shown, in the cell scenario corresponding to this embodiment, the coverage and overlap coverage fluctuate significantly in the early stages of training because the strategy model is still in the exploratory phase. As the number of training steps increases, the coverage gradually improves and stabilizes, while the overlap coverage gradually decreases and converges. The strategy model ultimately selects 10 actual active locations from 24 candidate locations and outputs the corresponding azimuth and downtilt angle configurations. Under this deployment scheme, the total coverage reaches 96.2%, and the overlap rate is 24.7%. Compared with manual experience-based configuration or unoptimized baseline schemes, the method of this invention can reduce invalid overlap coverage, improve the effective coverage level, and reduce the workload of repeated manual parameter tuning.

[0182] This invention provides a method for joint optimization of base station site selection and antenna parameters. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A method for joint optimization of base station site selection and antenna parameters, characterized in that, The steps include the following: Step 1: Acquire data of the target area and perform 3D raster modeling to obtain a set of candidate points; Step 2: Calculate the received power of the reference signal received by the grid from the antenna point to ensure that the power meets the standard and that the grid effectively covers the antenna main lobe directivity requirements. Step 3: Define core performance indicators to comprehensively evaluate the antenna deployment scheme; Step 4: Treat each candidate point as an intelligent agent. Based on the global-local hybrid reward mechanism, through the autonomous interaction and collaborative learning between the intelligent agent and the environment, the global near-optimal deployment scheme is finally obtained. Step 5: Based on the trained model, input cell parameters and output the optimal base station site selection and antenna parameter setting scheme.

2. The method for joint optimization of base station site selection and antenna parameters according to claim 1, characterized in that, Step 1 specifically includes: Step 1-1: Collect relevant basic data of the target cell, including: geographic and building information, radio frequency and antenna parameters, and engineering and propagation parameters; Steps 1-2: Establish a three-dimensional coordinate system with the north direction of the residential area as the positive Y-axis. Divide the three-dimensional space of the entire residential area into N cuboid grids of the same size and closely adjacent to each other. Each grid is defined by the three-dimensional coordinates of its center position. Based on the building geometry, the grids are distinguished between indoor and outdoor areas. Steps 1-3: Preset candidate antenna locations: Based on the three-dimensional gridded cell model constructed in Step 1-2, generate a set S of candidate antenna locations, with a total number of M.

3. The method for joint optimization of base station site selection and antenna parameters according to claim 1, characterized in that, The candidate locations for antenna deployment include the following two levels: Rooftop Feasible Points: Based on the geographical and architectural information of the community, extract the roof plan polygons of each building. Using the geometric center of the roof as the primary candidate point, supplement candidate points at certain intervals along its edges. The installation height of each rooftop point is determined. Add the preset support height to the roof height; Roadside light pole locations: Existing roadside light poles in the community will be used as candidate locations, with the installation height of each light pole location determined. Take the actual measured or standard height of the light pole directly.

4. The method for joint optimization of base station site selection and antenna parameters according to claim 1, characterized in that, Step 2 specifically includes: subtracting various engineering losses and propagation losses from the transmit power and antenna gain, and calculating the received power of the reference signal received by grid j from antenna point i. Based on the premise that the received power meets the standard, the antenna main lobe directivity determination is added to ensure that only the grid located within the antenna main lobe beamwidth is counted as effective coverage.

5. The method for joint optimization of base station site selection and antenna parameters according to claim 4, characterized in that, The antenna main lobe directivity determination specifically involves: First, calculate the candidate points. Pointing grid relative azimuth and relative downslope Let the current configuration parameter of antenna i be the azimuth angle. With downslope The horizontal and vertical 3 dB beamwidths of the antenna are respectively , The criterion for determining whether grid j is located within the main lobe range of antenna i is: in, This represents the horizontal angular deviation of grid j relative to the center direction of the main lobe of antenna i; This represents the vertical angular deviation of grid j relative to the center direction of the main lobe of antenna i. When the conditions of the above two formulas are met simultaneously, grid j is considered to be within the effective coverage area of ​​the main lobe of antenna i.

6. The method for joint optimization of base station site selection and antenna parameters according to claim 5, characterized in that, Step 3 specifically includes: Step 3-1: Define effective coverage using dual constraints of received power and antenna main lobe directivity. Step 3-2: Based on the decision function in Step 3-1, statistically analyze the global coverage and define core performance indicators to comprehensively evaluate the antenna deployment scheme. The grid coverage set includes: Grid coverage set: containing the antenna numbers of all grids that can provide high-quality coverage of grid j; Coverage count: counting the total number of grids in the entire target cell that are covered by at least one antenna with high quality; Overlap count: counting the number of grids that are covered by more than one antenna simultaneously. This indicator is a key parameter for measuring potential co-channel interference and resource waste in the network, and the optimization algorithm needs to minimize it; Coverage rate; Overlap rate.

7. The method for joint optimization of base station site selection and antenna parameters according to claim 6, characterized in that, The coverage determination function corresponding to the determination in step 3-1 is: in, As a receive power threshold, the antenna main lobe directivity is constrained as follows: This ensures that the signal comes from the effective radiation range of the antenna's main lobe; This indicates that grid j is well covered by antenna i; otherwise, it indicates that the grid is not covered or the coverage requirement is not met.

8. The method for joint optimization of base station site selection and antenna parameters according to claim 7, characterized in that, In step 4, each candidate antenna point is treated as an independent agent, and a multi-agent deep reinforcement learning environment is constructed. Each agent corresponds to a candidate point and is used to make decisions on whether to enable the candidate point and the corresponding antenna parameters. The local observations of each agent include at least the location coordinates, installation height, and historical decision information of the candidate point. The actions of each agent include at least whether to enable the candidate point, the antenna azimuth angle setting, and the downtilt angle setting. The global state of the environment includes the observation information of all agents and the global coverage statistics of the target area.

9. The method for joint optimization of base station site selection and antenna parameters according to claim 8, characterized in that, The global-local hybrid reward mechanism described in step 4 is as follows: The reward function is composed of a weighted average of global and local reward terms. in, and Adjustable weighting coefficients, local reward items It is only related to the coverage contribution and interference of agent i itself; global rewards Simultaneously consider coverage rate, overlap rate, and construction cost; The local reward Consider the coverage contribution and interference penalty of the base station itself.

10. The method for joint optimization of base station site selection and antenna parameters according to claim 7, characterized in that, Step 5 specifically includes: applying the policy model trained in step 4 to the target residential community scenario, with the input being the agent's local observations and the output being the agent's actions, including whether candidate points are enabled, azimuth angle setting, and tilt angle setting.