An unmanned aerial vehicle cluster topology dynamic optimization method and system under complex terrain
By constructing a dynamic twin scene model and optimizing the topology, the problem of insufficient communication and navigation performance of UAV swarms in complex terrain was solved, and stable execution of swarm tasks was achieved.
Patent Information
- Application Number
- CN202511316860.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-09-16
AI Technical Summary
In complex terrain, the communication and navigation performance of UAV swarms is affected by terrain obstruction and signal attenuation. Traditional topologies are difficult to adapt to dynamic environments, leading to swarm coordination failure and failing to meet mission reliability requirements.
By collecting 3D terrain modeling data and real-time environmental parameters, a dynamic twin scene model is constructed. Combined with distributed reinforcement learning and reconfigurable antennas, relay node location and link priority are optimized, antenna beam parameters are dynamically adjusted, a collaborative adaptation scheme is generated, and the cluster topology is optimized in real time.
It achieves dynamic improvement in cluster communication and navigation performance under complex terrain, ensuring the stable execution of cluster tasks.
Smart Images

Figure CN120805745B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) technology, and in particular to a method and system for dynamic optimization of UAV swarm topology under complex terrain. Background Technology
[0002] In complex terrains such as canyons and urban clusters, UAV swarms are susceptible to terrain obstruction and signal attenuation. Traditional fixed topologies are ill-suited to dynamic environments—for example, obstructed areas can lead to communication link interruptions, and signal attenuation reduces navigation accuracy, ultimately causing swarm coordination failures. Existing topology optimization methods are mostly based on idealized terrain assumptions and fail to construct a dynamic correlation model between terrain and swarm status. They rely solely on static relay node deployment or fixed link priority allocation for optimization, which cannot respond in real-time to changes in terrain obstruction and signal fluctuations. Furthermore, the lack of coordinated adaptation design between antenna parameters and topology means that the beam adjustment of reconfigurable antennas is not dynamically adjusted in conjunction with terrain obstruction angles, further exacerbating communication and navigation stability issues and making it difficult to meet the reliability requirements of UAV swarm reconnaissance and rescue missions in complex terrains. Summary of the Invention
[0003] The purpose of this invention is to provide a method and system for dynamic optimization of UAV swarm topology in complex terrain, so as to overcome the shortcomings of the prior art, and to achieve dynamic improvement of swarm communication and navigation performance in complex terrain, and ensure stable execution of swarm tasks.
[0004] One embodiment of this application provides a method for dynamic optimization of UAV swarm topology under complex terrain, the method comprising:
[0005] Collect 3D terrain modeling data and real-time environmental parameters of canyons or urban clusters, generate a terrain digital twin containing occlusion areas and signal attenuation coefficients through terrain feature extraction algorithms, and construct a dynamically updated twin scene model by integrating the location and communication link status data of drone clusters.
[0006] The twin scene model is input into a distributed reinforcement learning model. With cluster communication connectivity and navigation signal integrity as joint optimization objectives, a reward decay function based on terrain occlusion is designed. A topology optimization strategy for UAV relay node location and link priority allocation is generated through multi-agent asynchronous decision-making.
[0007] Based on the relay node location and link direction in the topology optimization strategy, the beam pattern database of the reconfigurable antenna is called, and the antenna beamwidth and gain parameters are dynamically adjusted by matching the shielding angle in the terrain digital twin to generate a collaborative adaptation scheme of antenna parameters and topology.
[0008] Based on the aforementioned collaborative adaptation scheme, cluster communication quality and navigation continuity indicators are collected in real time. The indicator deviations are fed back to the terrain digital twin for scene correction, driving the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjusting the reconfigurable antenna parameters, and generating dynamic optimization results of cluster topology adapted to complex terrain changes.
[0009] Optionally, the step of collecting 3D terrain modeling data and real-time environmental parameters of canyons or urban clusters, generating a terrain digital twin including occlusion areas and signal attenuation coefficients through terrain feature extraction algorithms, and constructing a dynamically updated twin scene model by integrating the location and communication link status data of the UAV swarm, includes:
[0010] The raw data was collected by category. For canyon terrain, 3D modeling data of elevation difference, rock wall inclination angle and valley width were collected. For urban agglomeration, 3D modeling data of building height, building spacing and street direction were collected. At the same time, real-time environmental parameters were collected, including meteorological parameters such as wind speed, wind direction and precipitation intensity, and electromagnetic parameters such as interference frequency and signal interference intensity. The data were organized according to terrain type and parameter dimensions to form the raw multidimensional dataset.
[0011] The raw multidimensional data is preprocessed, and the terrain 3D modeling data is divided into grids. The grid cell size is standardized and filled with elevation values. The spatiotemporal smoothing algorithm is used to eliminate instantaneous disturbances in real-time environmental parameters, and Kalman filtering is used to correct outliers that exceed the physical reasonable range, resulting in a standardized data grid.
[0012] Topographic feature parameters are extracted, and a deep learning semantic segmentation algorithm is used to analyze the standardized data grid, identify and mark the shading areas in the terrain, and calculate the signal attenuation coefficient of each grid cell based on the electromagnetic wave propagation model and environmental electromagnetic parameters to generate a terrain feature map.
[0013] The navigation position data and communication link status data of the drone swarm acquired in real time are mapped to the corresponding grid cells of the terrain feature map according to the timestamp, and the data is updated at fixed time intervals to construct a twin scene model that includes static terrain features and dynamic swarm status.
[0014] Optionally, the step of inputting the twin scene model into a distributed reinforcement learning model, using cluster communication connectivity and navigation signal integrity as joint optimization objectives, designs a reward decay function based on terrain occlusion, and generates a topology optimization strategy for UAV relay node location selection and link priority allocation through multi-agent asynchronous decision-making, including:
[0015] Key state parameters, including terrain occlusion, relative position of UAV cluster, communication link connectivity, and navigation signal strength, are extracted from the twin scene model. A multi-dimensional state space matrix is constructed according to the UAV number and terrain grid index, and the matrix elements quantify the corresponding state parameter values.
[0016] A joint optimization objective is set, with cluster communication connectivity and navigation signal integrity as the core optimization indicators. Based on the weights of their impact on cluster topology, a comprehensive optimization indicator is calculated using weighted averages, thus clarifying the direction for indicator improvement.
[0017] The design incorporates a reward decay function, where the base reward value is positively correlated with the comprehensive optimization index; a terrain occlusion correction factor is introduced, where the reward value decreases linearly with occlusion when the occlusion of the area where the UAV is located exceeds a set value; and a fixed penalty value is set if a communication link is interrupted or the navigation and positioning are out of tolerance, thus forming a dynamic reward mechanism.
[0018] The system performs asynchronous decision-making by multi-agents, splitting the multidimensional state space matrix into local state segments for each drone. Each drone acts as an independent agent, calculating decision priorities based on its local segment. The global coordination layer of the distributed reinforcement learning model integrates the decisions of each agent, prioritizing low-occlusion areas as relay nodes, and allocating link priorities in reverse order of signal attenuation coefficients. This generates a topology optimization strategy that includes the three-dimensional coordinates of relay nodes and the link priority ranking.
[0019] Optionally, the step of calling the beam pattern database of the reconfigurable antenna based on the relay node location and link direction in the topology optimization strategy, dynamically adjusting the antenna beamwidth and gain parameters to match the shielding angle in the terrain digital twin, and generating a cooperative adaptation scheme for antenna parameters and topology includes:
[0020] The topology optimization strategy parameters are analyzed, and the identification information, three-dimensional coordinates and main communication link direction angle of each relay node are extracted from the topology optimization strategy. The coverage requirements of each link are clarified and the strategy parameter table is compiled.
[0021] Call the beam pattern database, retrieve the matching initial beam parameters of the reconfigurable antenna from the database based on the link direction angle in the strategy parameter table, and extract the antenna parameter adjustment boundary to obtain the initial antenna configuration set;
[0022] Dynamically match the shielding angle adjustment parameters, query the maximum terrain shielding angle on each communication link path from the terrain digital twin; classify and adjust the antenna parameters according to the size of the shielding angle. When the shielding angle is in a preset small range, compress the beamwidth and increase the gain. When the shielding angle is in a preset medium range, maintain parameter balance. When the shielding angle is in a preset large range, expand the beamwidth and reduce the gain to obtain the adjusted antenna parameters.
[0023] Collaborative verification of adaptability involves checking the matching between the adjusted antenna parameters and the relay node locations, and correcting any conflicting parameters. The relay node and link information in the topology optimization strategy are associated with the corresponding antenna parameters by UAV ID to generate a collaborative adaptation scheme for antenna parameters and topology.
[0024] Optionally, based on the aforementioned collaborative adaptation scheme, the real-time collection of cluster communication quality and navigation continuity indicators, the feedback of indicator deviations to the terrain digital twin for scene correction, the driving of the distributed reinforcement learning system to iteratively update the topology optimization strategy, the synchronous adjustment of reconfigurable antenna parameters, and the generation of dynamic optimization results for cluster topology adapted to complex terrain changes, including:
[0025] Load the collaborative adaptation scheme, extract the antenna parameters and topology strategy in the scheme as the benchmark configuration, clarify the expected communication coverage and navigation and positioning accuracy requirements of each UAV, and form a benchmark parameter table for the scheme;
[0026] Based on the baseline parameter table of the scheme, the performance indicators are collected, and the communication quality indicators and navigation continuity indicators of the UAV cluster are collected at preset time intervals. The actual collected values are compared with the expected values in the baseline parameter table of the scheme, and the deviation rate of each indicator is calculated to form an indicator deviation set.
[0027] Feedback indicator deviation correction twin, filter out anomalies where the indicator deviation concentration exceeds the fault tolerance range of the scheme, associate the anomaly deviation with the corresponding grid cell of the terrain digital twin, update the shading area boundary and signal attenuation coefficient of the cell, and generate the corrected terrain digital twin.
[0028] The iterative optimization strategy and parameters are input into the distributed reinforcement learning model, and the topology optimization strategy is adjusted in combination with parameter conflicts in the collaborative adaptation scheme. At the same time, the antenna parameters are corrected according to the new topology strategy, and the adjusted indicators are collected to verify whether the deviation rate meets the scheme requirements. After the target is met, the optimized topology strategy and antenna parameters are integrated to generate a cluster topology dynamic optimization result adapted to complex terrain changes.
[0029] Another embodiment of this application provides a dynamic optimization system for UAV swarm topology in complex terrain, the system comprising:
[0030] The data acquisition module is used to collect 3D terrain modeling data and real-time environmental parameters of canyons or urban clusters. It generates a terrain digital twin containing occlusion areas and signal attenuation coefficients through terrain feature extraction algorithms. It integrates the location and communication link status data of the drone swarm to build a dynamically updated twin scene model.
[0031] The generation module is used to input the twin scene model into the distributed reinforcement learning model, with cluster communication connectivity and navigation signal integrity as joint optimization objectives, design a reward decay function based on terrain occlusion, and generate a topology optimization strategy for UAV relay node location and link priority allocation through multi-agent asynchronous decision-making.
[0032] The matching module is used to call the beam pattern database of the reconfigurable antenna according to the relay node position and link direction in the topology optimization strategy, match the shielding angle in the terrain digital twin, dynamically adjust the antenna beamwidth and gain parameters, and generate a cooperative adaptation scheme of antenna parameters and topology.
[0033] The optimization module is used to collect cluster communication quality and navigation continuity indicators in real time based on the aforementioned collaborative adaptation scheme, feed back the indicator deviations to the terrain digital twin for scene correction, drive the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjust the reconfigurable antenna parameters, and generate dynamic optimization results of cluster topology that adapt to complex terrain changes.
[0034] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.
[0035] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.
[0036] Compared with existing technologies, this invention provides a method for dynamic optimization of UAV swarm topology in complex terrain. It collects 3D terrain modeling data and real-time environmental parameters from canyons or urban clusters, and integrates the location and communication link status data of the UAV swarm to construct a dynamically updated twin scene model. The twin scene model is then input into a distributed reinforcement learning model, which generates a topology optimization strategy for UAV relay node location and link priority allocation through asynchronous decision-making by multiple agents. Based on the relay node location and link direction in the topology optimization strategy, a collaborative adaptation scheme for antenna parameters and topology is generated. Based on this collaborative adaptation scheme, real-time collection of swarm communication quality and navigation continuity indicators generates a dynamic optimization result for the swarm topology adapted to changes in complex terrain. This enables dynamic improvement of swarm communication and navigation performance in complex terrain, ensuring stable execution of swarm tasks. Attached Figure Description
[0037] Figure 1 Hardware structure block diagram of a computer terminal for a method of dynamic optimization of UAV swarm topology under complex terrain provided in an embodiment of the present invention;
[0038] Figure 2 A flowchart illustrating a method for dynamic optimization of UAV swarm topology under complex terrain, provided in an embodiment of the present invention;
[0039] Figure 3 This is a schematic diagram of a dynamic optimization system for UAV swarm topology in complex terrain, provided as an embodiment of the present invention. Detailed Implementation
[0040] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0041] This invention first provides a method for dynamic optimization of UAV swarm topology under complex terrain. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers.
[0042] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for a method of dynamic topology optimization of UAV swarms in complex terrain, provided in an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0043] The non-volatile storage medium can store the operating system and computer program. The computer program includes program instructions that, when executed, cause the processor to perform a dynamic optimization method for UAV swarm topology under any complex terrain.
[0044] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0045] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to perform dynamic optimization methods for UAV swarm topology under any complex terrain.
[0046] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0047] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0048] See Figure 2 The present invention provides a method for dynamic optimization of UAV swarm topology under complex terrain, which may include the following steps:
[0049] S201 collects 3D terrain modeling data and real-time environmental parameters of canyons or urban clusters, generates a terrain digital twin containing occlusion areas and signal attenuation coefficients through terrain feature extraction algorithms, and integrates the location and communication link status data of drone clusters to construct a dynamically updated twin scene model.
[0050] Specifically, raw data can be collected in categories. For canyon terrain, 3D modeling data such as elevation difference, rock wall inclination, and valley width can be collected. For urban clusters, 3D modeling data such as building height, building spacing, and street orientation can be collected. At the same time, real-time environmental parameters can be collected, including meteorological parameters such as wind speed, wind direction, and precipitation intensity, and electromagnetic parameters such as interference frequency and signal interference intensity. The data can be organized according to terrain type and parameter dimensions to form a raw multidimensional dataset.
[0051] Raw data acquisition is the foundation for building a terrain digital twin. It must adhere to the principle of "differentiated acquisition based on terrain type + comprehensive coverage of environmental parameters" to ensure that the data accurately reflects the physical characteristics of complex terrain and the real-time environmental impact, thus addressing the problem of "data simplification leading to twin distortion." This step needs to clearly define "what to collect, how to collect, and how to organize" to provide structured input for subsequent preprocessing.
[0052] I. Classification and Collection of 3D Terrain Modeling Data:
[0053] For two typical complex terrain types, canyons and urban clusters, core parameters that characterize terrain shading and signal propagation are collected. For each parameter, the acquisition equipment, accuracy range, and sampling density are clearly defined to ensure that the data is quantifiable and reusable.
[0054] 3D modeling data of the canyon terrain:
[0055] Elevation difference: Reflects the vertical topographic undulation of the canyon (directly affecting signal obstruction). A drone-based LiDAR (model Velodyne VLP-16, ranging accuracy ±2cm, sampling density 10 points / ㎡) is used to collect data by flying back and forth along the canyon axis. For example, if the canyon floor is 1000m above sea level and the highest elevation of the two side walls is 1500m, the elevation difference is recorded as 500m (maximum) and 300m (average).
[0056] Rock wall dip angle: reflects the steepness of the rock wall (the larger the dip angle, the stronger the signal reflection and the more severe the obstruction). The rock wall plane is fitted by LiDAR point cloud data to calculate the angle between the plane and the horizontal plane (accuracy ±0.5°). For example, the rock wall dip angle on the west side is 75° and the rock wall dip angle on the east side is 60°.
[0057] Valley bottom width: Reflects the horizontal spatial scale of the canyon (the smaller the width, the more obvious the signal multipath effect). The coordinates of the two sides of the valley bottom are collected using RTK-GPS (accuracy ±1cm) and the horizontal distance is calculated. For example, if the valley bottom width of a certain canyon is between 50m and 100m, it is recorded as 50m, 55m, 60m...100m in 5m intervals.
[0058] 3D modeling data of urban agglomeration terrain:
[0059] Building height: Reflects the height of vertical obstruction sources in the city (tall buildings easily block the signals of low-altitude drones). The urban area is photographed using an oblique photography camera (model DJI P1, 50 million pixels, modeling accuracy ±5cm), and the building height data is generated by 3D modeling software (such as ContextCapture). For example, the building height in a commercial area is 30m-100m, recorded as 30m (low-rise), 60m (mid-rise), and 100m (high-rise).
[0060] Building spacing: Reflects the density of buildings in the horizontal direction (the smaller the spacing, the more difficult it is for the signal to diffract). The coordinates of the exterior walls of adjacent buildings are collected by RTK-GPS to calculate the horizontal distance. For example, if the building spacing in a residential area is 20m-50m, it is recorded as 20m (dense area) and 50m (open area).
[0061] Street orientation: Reflects the main direction of signal propagation in the city (signal attenuation is smaller along the street direction). By combining GIS maps with field measurements, the azimuth of the street is recorded (accuracy ±1°). For example, a main road is oriented east-west (azimuth 90°) and a secondary road is oriented north-south (azimuth 0°).
[0062] II. Real-time environmental parameter acquisition:
[0063] Environmental parameters directly affect the quality of UAV communication and navigation signals. It is necessary to collect two key parameters: meteorological and electromagnetic, to ensure coverage against external interference factors affecting signal propagation.
[0064] Meteorological parameters:
[0065] Wind speed: Affects the stability of the drone (indirectly affecting the stability of the communication link). A small weather station (Davis Vantage Pro2, measurement range 0-60m / s, accuracy ±0.1m / s) is installed at a high point of the terrain with a sampling frequency of 1Hz. For example, the collected values are 3m / s (light breeze) and 8m / s (strong wind).
[0066] Wind direction: Affects the signal propagation path (signal attenuation is smaller when the wind is downwind). It is collected by the wind direction sensor of the weather station (measurement range 0-360°, accuracy ±5°) and recorded as azimuth angle, for example, 30° east of northeast (wind direction angle 60°).
[0067] Rainfall intensity: Affects electromagnetic wave attenuation (rainwater increases signal absorption). It was collected using a tipping bucket rain gauge (resolution 0.1mm, accuracy ±0.5mm) and recorded according to intensity as 0.5mm / h (light rain) and 10mm / h (heavy rain).
[0068] Electromagnetic parameters:
[0069] Interference frequency: Identify electromagnetic interference sources affecting drone communication (such as WiFi, Bluetooth, industrial equipment), and use a spectrum analyzer (model Keysight N9918A, frequency range 9kHz-18GHz, resolution 1Hz) to collect data at key terrain points (such as canyon bottom, city square), for example, 2.4GHz (WiFi interference) and 5.8GHz (drone image transmission interference).
[0070] Signal interference intensity: Quantify the impact of interference sources on UAV signals. Use a signal strength meter (measurement range -120dBm to -30dBm, accuracy ±1dBm) to collect the power of the interference signal. For example, the interference intensity of 2.4GHz is -80dBm (weak interference) and the interference intensity of 5.8GHz is -60dBm (medium interference).
[0071] III. Organizing the original multidimensional dataset:
[0072] Organize the data according to the hierarchical structure of "terrain type - parameter dimension" to avoid confusion, for example:
[0073] Original multidimensional dataset (terrain type: canyon):
[0074] 3D modeling data: elevation difference {maximum 500m, average 300m}, rock wall dip angle {75° on the west side, 60° on the east side}, valley floor width {50m, 55m, ..., 100m};
[0075] Real-time environmental parameters: meteorology {wind speed 3m / s, wind direction 60°, precipitation intensity 0.5mm / h}, electromagnetic {interference frequency 2.4GHz / 5.8GHz, interference intensity -80dBm / -60dBm};
[0076] Original multidimensional dataset (terrain type: urban cluster):
[0077] 3D modeling data: Building height {30m, 60m, 100m}, building spacing {20m, 50m}, street orientation {90° (east-west), 0° (north-south)};
[0078] Real-time environmental parameters: Meteorological {wind speed 5m / s, wind direction 180° (south), precipitation intensity 0mm / h}, Electromagnetic {interference frequency 2.4GHz, interference intensity -70dBm}.
[0079] Each dataset is labeled with the collection time (accurate to the minute) and the collection device number to ensure data traceability and provide a basis for identifying outliers in subsequent preprocessing.
[0080] The raw multidimensional data is preprocessed, and the terrain 3D modeling data is divided into grids. The grid cell size is standardized and filled with elevation values. The spatiotemporal smoothing algorithm is used to eliminate instantaneous disturbances in real-time environmental parameters, and Kalman filtering is used to correct outliers that exceed the physical reasonable range, resulting in a standardized data grid.
[0081] The raw data may contain issues such as "discrete terrain data and fluctuating environmental parameters," requiring preprocessing to transform it into unified and stable standardized data. This provides reliable input for terrain feature extraction and addresses the problem of "feature extraction bias caused by low data quality." This step needs to be performed in two steps: "terrain data gridding" and "environmental parameter smoothing correction," to ensure consistent data format and reasonable numerical values.
[0082] I. Mesh generation of terrain 3D modeling data:
[0083] Gridding is the core method for transforming discrete terrain data into a continuous grid. It requires unifying the grid size and coordinate system to ensure that data from different terrain types can be compared and fused.
[0084] Mesh cell size setting: Based on signal propagation characteristics (UAV communication wavelength is usually 10cm-1m, and the mesh size needs to be smaller than the wavelength to capture details), the mesh cell size is set to 5m×5m (balancing accuracy and computational efficiency). The canyon terrain extends the mesh along the axis direction, and the urban agglomeration terrain is covered by the mesh in rectangular areas.
[0085] Coordinate System 1: A three-dimensional coordinate system combining WGS84 geodetic coordinates (latitude and longitude) and altitude is adopted. For example, the coordinates of the lower left corner of a certain canyon grid are (longitude 110.0°, latitude 25.0°, altitude 1000m), and the coordinates of the upper right corner are (longitude 110.1°, latitude 25.1°, altitude 1500m), dividing it into 20×20=400 grid units;
[0086] Elevation value filling: For each grid cell, the maximum, minimum, and average elevation values are extracted from the original LiDAR or oblique photogrammetry data and filled into the grid attributes. For example, the elevation values of a canyon grid cell (110.02°, 25.03°) are {max:1200m, min:1050m, avg:1125m}, and the elevation values of an urban cluster grid cell (120.05°, 30.02°) are {max:80m (building top), min:0m (ground), avg:40m}.
[0087] After gridding, the integrity of the grid needs to be verified to ensure that there is no missing data (missing grids are filled by interpolation of adjacent grids). For example, if a grid on the edge of a canyon has no LiDAR data, it is filled by the average elevation of the adjacent grids on the east and south sides, which is 1100m.
[0088] II. Spatiotemporal smoothing and outlier correction of real-time environmental parameters:
[0089] Environmental parameters are susceptible to transient disturbances (such as gusts of wind or sudden electromagnetic interference), and data stability must be ensured through algorithmic processing.
[0090] Spatiotemporal smoothing algorithm: The moving average algorithm is used to eliminate instantaneous disturbances in the time dimension. The window size is set to 5 sampling points (10 seconds per sampling point, 50 seconds window). The formula is: smoothed value = (x1 + x2 + x3 + x4 + x5) / 5. For example, if the original wind speed data is 3 m / s, 8 m / s (gust), 4 m / s, 3 m / s, 5 m / s, the smoothed value = (3 + 8 + 4 + 3 + 5) / 5 = 4.6 m / s, eliminating the influence of gusts. In the spatial dimension, Kriging interpolation is used to spatially complete sparsely collected electromagnetic parameters (such as those collected only at 3 points) to ensure that each grid cell has environmental parameter values.
[0091] Kalman filtering corrects outliers: For data exceeding the physically reasonable range (such as a sudden increase in precipitation intensity to 100 mm / h, far exceeding local climate levels), Kalman filtering is used for correction. Filtering parameter settings: The state variable is the environmental parameter value, the process noise covariance Q = 0.01 (reflecting the natural fluctuation of the parameter), and the measurement noise covariance R = 0.1 (reflecting sensor error). For example, if the original precipitation intensity data are 0.5 mm / h, 100 mm / h (outlier), and 0.6 mm / h, after filtering, the outlier is corrected to 0.55 mm / h, which is within the physically reasonable range (the local maximum precipitation intensity is 50 mm / h).
[0092] The standardized data grid obtained after preprocessing contains two types of attributes, namely "terrain elevation" and "environmental parameters", with a unified format and stable values, laying the foundation for subsequent terrain feature extraction.
[0093] Topographic feature parameters are extracted, and a deep learning semantic segmentation algorithm is used to analyze the standardized data grid, identify and mark the shading areas in the terrain, and calculate the signal attenuation coefficient of each grid cell based on the electromagnetic wave propagation model and environmental electromagnetic parameters to generate a terrain feature map.
[0094] Terrain feature extraction is the core step in transforming "raw data" into "usable features in a twin." It requires identifying occlusion areas through semantic segmentation and calculating attenuation coefficients through propagation models to address the problem of "inability to quantify the impact of terrain on the signal." This step needs to clarify "how to identify occlusion and how to calculate attenuation" to generate a feature map that can directly support topology optimization.
[0095] I. Occlusion Region Recognition Based on Deep Learning Semantic Segmentation:
[0096] Using the U-Net semantic segmentation model (suitable for segmenting structured data such as terrain), terrain regions in the standardized data grid are divided into two categories: "occluded regions" and "unoccluded regions." The specific process is as follows:
[0097] Data preprocessing: The elevation values of the standardized data grid are converted into grayscale images (the higher the elevation, the larger the grayscale value). For example, 1000m corresponds to grayscale 0 and 1500m corresponds to grayscale 255, forming an input image of 256×256 pixels.
[0098] Model Training and Inference: The U-Net model was trained using 10,000 images with complex terrain annotations (manually labeled occluded areas, such as canyon walls and shadows of city buildings). Training parameters: batch size 16, number of iterations 50, learning rate 0.001, and cross-entropy loss function. The pre-processed images were input into the trained model, which outputs binarized segmentation results (white for occluded areas, black for unoccluded areas). For example, the grid on the west side of the canyon wall and the grid below the city buildings were marked as occluded areas.
[0099] Optimization of occluded region boundaries: Morphological closing operations (dilation followed by erosion) are used to eliminate small holes (such as small blanks caused by windows in tall buildings) in the segmentation results. The structuring element is set to 3×3 pixels to ensure that the boundaries of the occluded region are continuous and accurate.
[0100] After identification, the segmentation accuracy needs to be verified. By comparing with real-world photos, the accuracy of identifying the occluded area must be ≥95%. For example, if the overlap between the occluded area marked on the model of a certain city area and the shadow of the high-rise building on the actual site reaches 97%, it is considered qualified.
[0101] II. Calculation of signal attenuation coefficient based on electromagnetic wave propagation model:
[0102] Using a free-space propagation model combined with environmental interference correction, the signal attenuation coefficient (unit: dB / m) of each grid cell is calculated, quantifying the signal propagation loss within the grid. Specific steps include:
[0103] Basic attenuation coefficient calculation (free space model): The formula is L0 = 32.45 + 20log(d) + 20log(f), where d is the signal propagation distance (km) and f is the signal frequency (MHz). For example, if the UAV communication frequency is f = 2400MHz and the signal propagation distance within a certain grid cell is d = 0.5km, then L0 = 32.45 + 20log0.5 + 20log2400 ≈ 32.45 - 6.02 + 67.6 ≈ 94.03dB, and the basic attenuation coefficient = 94.03dB / 500m ≈ 0.188dB / m;
[0104] Environmental interference correction: Combining the interference intensity (I, in dBm) from the electromagnetic parameters and the precipitation intensity (P, in mm / h) from the meteorological parameters, the correction formula is L=L0+α×I+β×P, where α=0.01 (interference intensity correction coefficient) and β=0.1 (precipitation intensity correction coefficient). For example, if the interference intensity I=-80dBm and the precipitation intensity P=0.5mm / h, after correction, L=94.03+0.01×(-80)+0.1×0.5≈94.03-0.8+0.05≈93.28dB, and the attenuation coefficient after correction is 93.28dB / 500m≈0.187dB / m;
[0105] Terrain shading correction: If the grid cell is a shading area, an additional shading attenuation ΔL = 10 × log (θ / 90°) is added (θ is the inclination angle of the shading area, in °). For example, if the rock wall inclination angle θ = 75°, ΔL = 10 × log (75 / 90) ≈ -0.79dB, and the final attenuation coefficient = 0.187dB / m + (-0.79dB) / 500m ≈ 0.185dB / m (the attenuation in the shading area is slightly lower due to enhanced signal reflection).
[0106] III. Generation of Topographic Feature Maps:
[0107] By integrating the occlusion area markers and signal attenuation coefficients, a visualized terrain feature map is generated. Each grid cell in the map is labeled with "occlusion status (occluded / unoccluded)" and "attenuation coefficient (dB / m)". For example:
[0108] "Topographic feature map (canyon terrain, grid size 5m×5m): "
[0109] Grid (110.02°, 25.03°): Shading status = Shading (west rock face), attenuation coefficient = 0.185dB / m;
[0110] Grid (110.05°, 25.05°): Shading state = Unshading (open area at the bottom of the valley), attenuation coefficient = 0.188dB / m;
[0111] Grid (110.08°, 25.07°): Shading status = Shaded (east side rock face), attenuation coefficient = 0.186dB / m.
[0112] The map needs to support dynamic updates (updated when merging drone data later) to provide core feature inputs for building a twin scene model.
[0113] The navigation position data and communication link status data of the drone swarm acquired in real time are mapped to the corresponding grid cells of the terrain feature map according to the timestamp, and the data is updated at fixed time intervals to construct a twin scene model that includes static terrain features and dynamic swarm status.
[0114] The twin scene model needs to integrate "static terrain" and "dynamic cluster" data, and ensure spatiotemporal synchronization between the two through timestamp mapping to solve the problem of "separation of terrain and cluster states leading to a static model". This step needs to clarify "how to acquire cluster data, how to map it, and how to update it" to achieve the dynamism and realism of the twin.
[0115] I. Real-time acquisition of drone swarm data:
[0116] Obtain core data that reflects the cluster's status and ensure that the data matches the spatiotemporal dimensions of the terrain feature map:
[0117] Navigation location data: The system uses the RTK-GPS (accuracy ±1cm) and IMU (inertial measurement unit, update frequency 100Hz) built into the UAV to fuse the UAV's three-dimensional coordinates (longitude, latitude, altitude) and timestamp (accurate to milliseconds). For example, the location data of UAV 1 is "timestamp: 2024-11-01 10:00:00.123, coordinates: (110.02°, 25.03°, 500m)";
[0118] Communication link status data: Link parameters are collected through an ad hoc network between drones, including link signal-to-noise ratio (SNR, in dB, measurement range 0-40dB), link connectivity (1 = connected, 0 = disconnected), and data transmission rate (in Mbps). For example, the link data between drone 1 and drone 2 is "timestamp: 2024-11-01 10:00:00.123, SNR=25dB, connectivity=1, rate=10Mbps".
[0119] Data is transmitted back to the ground processing system in real time via a wireless data transmission link (transmission rate 50Mbps, latency ≤100ms) to ensure data timeliness.
[0120] II. Timestamp Mapping and Grid Association:
[0121] The cluster data is precisely mapped to the grid cells of the terrain feature map according to the timestamp, ensuring that "the terrain features of the grid where the drone is located are associated with the terrain features of that grid":
[0122] Timestamp synchronization: The UAV clock is calibrated with the ground system clock using the NTP protocol (synchronization accuracy ≤1ms) to ensure that the timestamp of the cluster data is consistent with the timestamp of the terrain feature map acquisition.
[0123] Grid positioning: Based on the UAV's three-dimensional coordinates, calculate the terrain grid cell in which it is located (by matching the coordinate range). For example, if the UAV's coordinates (110.02°, 25.03°, 500m) fall within the grid (110.02°-110.07°, 25.03°-25.08°), associate the occlusion status and attenuation coefficient of that grid.
[0124] Data association: The navigation position and link status data of the UAV are added to the corresponding grid cell as "dynamic attributes". For example, the attributes of the grid (110.02°, 25.03°) are updated to "occlusion status = occlusion, attenuation coefficient = 0.185dB / m, UAV 1 position = (110.02°, 25.03°, 500m), link 1 (1-2) = SNR25dB, connectivity 1".
[0125] III. Dynamic Updates and Twin Scenario Model Construction:
[0126] Update the cluster data at fixed time intervals (set according to the frequency of terrain changes and the speed of cluster movement, for example, 30 seconds) to ensure that the model can reflect the real-time status:
[0127] Update frequency settings: Static terrain features (such as canyon walls and building heights) are updated every hour (slow changes), while dynamic cluster data is updated every 30 seconds (fast drone movement).
[0128] Model Construction: A visual twin scene is built using the Unity 3D engine, overlaying terrain feature maps (static layer) and cluster data (dynamic layer). The static layer renders terrain elevation and occlusion areas (distinguished by different colors), while the dynamic layer renders drone positions (icons) and communication links (lines, with colors reflecting SNR: green = SNR≥20dB, yellow = 10dB≤SNR<20dB, red = SNR<10dB).
[0129] Model Validation: Through field flight verification, the deviation between the UAV position and the actual position in the twin scenario is ≤1m, and the deviation between the link status and the actual SNR is ≤2dB, thus the model is deemed qualified.
[0130] The final twin scene model can not only intuitively display the static features of complex terrain, but also reflect the dynamic status of the drone swarm in real time, providing high-fidelity scene input for subsequent distributed reinforcement learning.
[0131] S202, the twin scene model is input into the distributed reinforcement learning model, and the cluster communication connectivity rate and navigation signal integrity are used as joint optimization objectives. A reward decay function based on terrain occlusion is designed, and a topology optimization strategy for UAV relay node location and link priority allocation is generated through multi-agent asynchronous decision-making.
[0132] Specifically, key state parameters, including terrain occlusion, relative position of UAV cluster, communication link connectivity, and navigation signal strength, can be extracted from the twin scene model. A multi-dimensional state space matrix can be constructed according to the UAV number and terrain grid index, and the matrix elements can be quantified to correspond to the state parameter values.
[0133] Key state parameters are the foundation of input for distributed reinforcement learning models. They need to be accurately extracted and quantified from the twin scene model to ensure that the parameters objectively reflect the dynamic relationship between "terrain and cluster". The multidimensional state space matrix is the structured carrier of the parameters, solving the problem that "dispersed parameters lead to inefficient model learning". This step needs to clarify "which parameters to extract, how to quantify them, and how to construct the matrix" to provide structured input for subsequent optimization.
[0134] I. Extraction and quantification of key state parameters:
[0135] Each parameter needs to be defined by combining the static terrain features and dynamic cluster data of the twin scene model, and avoid subjective judgment through quantifiable indicators. The specific extraction and quantification methods are as follows:
[0136] Terrain occlusion: Reflects the degree of signal occlusion in the grid cell where the UAV is located. It is quantified based on the "occlusion area percentage" in the terrain feature map generated in step two, with a value range of 0-100% (0% for no occlusion and 100% for complete occlusion). The extraction method is as follows: Locate the terrain grid cell where the UAV is located in the twin scene model, and calculate the area percentage of the occluded area (such as canyon walls or shadows of tall buildings in the city) within that cell. For example, if UAV 1 is located in grid 101, the total area of this grid is 25㎡ (5m×5m), and the occluded area is 17.5㎡, then the terrain occlusion = (17.5 / 25)×100% = 70%;
[0137] Relative position of a drone swarm: This reflects the spatial relationship between a single drone and the swarm center (affecting link connection distance). The three-dimensional Euclidean distance (in meters) between drones is calculated using the swarm's geometric center as the origin. The formula is: Where (x0, y0, z0) are the coordinates of the cluster center, and (x1, y1, z1) are the coordinates of the UAVs. For example, if the cluster center coordinates are (110.05°, 25.05°, 500m) and UAV #2 coordinates are (110.08°, 25.07°, 520m), after coordinate transformation (1° longitude ≈ 111km, 1° latitude ≈ 111km), the calculated horizontal distance is approximately 360 meters, and the height difference is 20 meters. ≈360.56 meters, quantified as 361 meters (rounded to the nearest integer);
[0138] Communication link connectivity status: Reflects the effectiveness of the link between UAVs. It is quantized by combining link connectivity with signal-to-noise ratio (SNR), with a value range of 0-1 (0 for interruption, 1 for optimal connectivity). The quantization formula is: Connectivity status value = Connectivity identifier × (SNR-10) / 30, where the connectivity identifier = 1 (connected) or 0 (interrupted), the SNR threshold is 10dB (below which the link is unstable), and the upper limit is 40dB (optimal). For example, if the link SNR between UAV 3 and UAV 4 is 25dB and the connectivity identifier is 1, then the connectivity status value = 1 × (25-10) / 30 = 0.5; if the SNR is 5dB and the connectivity identifier is 0, then the status value = 0.
[0139] Navigation signal strength: Reflects the reliability of UAV navigation and positioning. It is based on a comprehensive quantification of GPS signal strength (unit: dBm) and positioning error (unit: meters), with a value range of 0-1 (0 being the worst and 1 being the best). The quantification formula is: Navigation signal value = 0.5 × [(-signal strength - 80) / 40] + 0.5 × [1 - (positioning error / 10)], where the signal strength ranges from -120dBm (worst) to -80dBm (best), and the positioning error ranges from 0 to 10 meters (10 meters is the out-of-tolerance threshold). For example, if the signal strength of UAV No. 5 is -90dBm and the positioning error is 2 meters, then the navigation signal value = 0.5×[(90-80) / 40] + 0.5×[1-(2 / 10)] = 0.5×0.25 + 0.5×0.8 = 0.125 + 0.4 = 0.525, which is rounded to 0.53.
[0140] II. Construction of the multidimensional state space matrix:
[0141] The matrix uses "UAV ID - Terrain Grid Index" as a two-dimensional index, and the elements are the quantized values of the four parameters mentioned above, ensuring that the status of each UAV is strongly correlated with the terrain grid where it is located. The matrix dimensions are set according to the cluster size and terrain range (example: 5 UAVs, 5 core grids):
[0142] Index definition: The row index is the drone number (1-5), and the column index is the terrain grid index (101-105, corresponding to the core grid of drone activity in the twin scene).
[0143] Element padding: Each matrix element is a quadruple of "(terrain occlusion, relative position, communication link status, navigation signal strength)," for example:
[0144] Matrix element (1, 101) = (70, 350, 0.5, 0.53) → UAV No. 1 is in grid No. 101, with a coverage of 70%, a distance of 350 meters from the cluster center, a link status of 0.5, and a navigation signal of 0.53;
[0145] Matrix element (2, 102) = (30, 280, 0.8, 0.75) → UAV No. 2 is in grid No. 102, with an obscurity of 30%, a relative distance of 280 meters, a link status of 0.8, and a navigation signal of 0.75;
[0146] Matrix element (3, 103) = (90, 420, 0.2, 0.35) → UAV No. 3 is in grid No. 103, with 90% obscurity, a relative distance of 420 meters, a link status of 0.2, and a navigation signal of 0.35;
[0147] Matrix verification: Check if there are outliers in the elements (such as occlusion > 100% or negative relative distance). If so, correct them based on interpolation of adjacent elements. For example, the occlusion of element (4, 104) is 110%, which is corrected to 100% to ensure that the matrix data is reasonable.
[0148] A joint optimization objective is set, with cluster communication connectivity and navigation signal integrity as the core optimization indicators. Based on the weights of their impact on cluster topology, a comprehensive optimization indicator is calculated using weighted averages, thus clarifying the direction for indicator improvement.
[0149] The joint optimization objective serves as a "guiding principle" for reinforcement learning, requiring a balance between the importance of communication and navigation to avoid the loss of cluster functionality due to a single optimization (e.g., optimizing only communication leading to navigation inaccuracies). This step necessitates determining the weights of the indicators, the calculation methods, and the directions for improvement, ensuring that the objective is quantifiable and achievable.
[0150] I. Definition and Calculation of Core Optimization Indicators:
[0151] Cluster communication connectivity rate: This reflects the overall connectivity of links between drones within a cluster. The calculation method is "actual number of connected links / theoretical maximum number of links × 100%". The theoretical maximum number of links = n × (n-1) / 2 (where n is the number of drones), and the actual number of connected links is the number of links with a connectivity status value ≥ 0.3 (0.3 is the effective threshold for a link). For example, with 5 drones, the theoretical maximum number of links = 5 × 4 / 2 = 10, and the actual number of connected links = 7 (of which 3 links have a connectivity status value < 0.3), then the communication connectivity rate = 7 / 10 × 100% = 70%.
[0152] Navigation signal integrity: Reflects the overall reliability of navigation and positioning within the drone cluster. It is calculated as "number of drones with navigation signal values ≥ 0.5 / total number of drones × 100%" (where 0.5 is the navigation reliability threshold). For example, if 4 out of 5 drones have navigation signal values ≥ 0.5 and 1 has a value < 0.5, then the navigation signal integrity = 4 / 5 × 100% = 80%.
[0153] II. Determination of Indicator Weights and Calculation of Comprehensive Optimization Indicators:
[0154] The weights are determined based on "cluster task requirements" and "expert experience." Communication connectivity is more critical for cluster data transmission, and its weight is higher than that of navigation signal integrity. Specific settings are as follows:
[0155] Weight allocation: Cluster communication connectivity weight ω1=0.6, navigation signal integrity weight ω2=0.4 (the weights sum to 1 to ensure fairness);
[0156] The formula for the comprehensive optimization index is: Comprehensive index S = ω1 × (communication connectivity rate / 100) + ω2 × (navigation signal integrity / 100), with a value range of 0-1 (1 is the optimal value).
[0157] Example calculation: If the communication connectivity rate is 70% and the navigation signal integrity is 80%, then S = 0.6 × 0.7 + 0.4 × 0.8 = 0.42 + 0.32 = 0.74; if the communication connectivity rate is increased to 90% and the navigation integrity is increased to 85%, then S = 0.6 × 0.9 + 0.4 × 0.85 = 0.54 + 0.34 = 0.88, and the indicators are significantly improved.
[0158] III. Clarifying the Direction for Indicator Improvement:
[0159] Based on the current and target values of the comprehensive optimization indicators (usually S≥0.9 is set as the target), the direction for improvement is clarified:
[0160] In scenarios with low connectivity (e.g., S=0.74, communication connectivity rate 70%), the improvement direction is to "increase relay nodes, reduce links in obscured areas, and increase the number of connected links". For example, add relay nodes in low obscurity grids (e.g., grid 102, obscurity 30%) to improve the interrupted link between UAVs 3 and 4.
[0161] In scenarios with low navigation integrity (e.g., S=0.65, navigation integrity 60%), the improvement direction is to "move drones with poor navigation signals (e.g., No. 3, signal value 0.35) to areas with low obstruction to reduce the obstruction of navigation signals by terrain".
[0162] For scenarios where both metrics are low: prioritize improving the communication connectivity rate, which has a higher weight, and then optimize navigation integrity to avoid resource dispersion.
[0163] The design incorporates a reward decay function, where the base reward value is positively correlated with the comprehensive optimization index; a terrain occlusion correction factor is introduced, where the reward value decreases linearly with occlusion when the occlusion of the area where the UAV is located exceeds a set value; and a fixed penalty value is set if a communication link is interrupted or the navigation and positioning are out of tolerance, thus forming a dynamic reward mechanism.
[0164] The reward decay function is the "feedback mechanism" of a reinforcement learning model. It needs to address the problem of "the model's inability to perceive the impact of terrain and the consequences of failures" by "rewarding optimization and penalizing risk avoidance." The function design must take into account "optimization objectives, terrain effects, and failure penalties" to ensure that it is dynamic and reasonable.
[0165] I. Calculation of basic reward value:
[0166] The base reward value is positively correlated with the overall optimization metric S, encouraging the model to improve the metric. The formula is: Base Reward R0 = 10 × S (10 is the reward coefficient, ensuring the reward value is between 0 and 10 for easy gradient updates). Example:
[0167] When S=0.74, R0=10×0.74=7.4;
[0168] When S=0.88, R0=10×0.88=8.8;
[0169] When S=0.95 (meeting the standard), R0=10×0.95=9.5 (close to the full score bonus).
[0170] II. Design of Terrain Shading Correction Factor:
[0171] A correction factor γ is introduced to reduce the reward in high-occlusion regions (to prevent the model from selecting occluded regions as relay nodes), and the formula is as follows:
[0172] If the terrain shading α ≤ 50% (low shading), γ = 1 (no attenuation, this choice is encouraged).
[0173] If 50% < α ≤ 80% (medium occlusion), γ = 1 - (α - 50) / 100 (linear decay; the higher the occlusion, the smaller γ).
[0174] If α > 80% (high occlusion), γ = 0.2 (lowest attenuation coefficient, significantly reduces reward, avoid selection).
[0175] Example calculation:
[0176] When the drone is in the area with α=30% (low obscurity) and γ=1, the corrected reward R1=R0×γ=7.4×1=7.4;
[0177] When the drone is in the α=70% (medium obscuration) region, γ=1 - (70-50) / 100=0.8, and the corrected reward R1=7.4×0.8=5.92;
[0178] When the drone is in the area with α=85% (high obscurity) and γ=0.2, the corrected reward R1=7.4×0.2=1.48.
[0179] III. Setting Fixed Penalty Values:
[0180] For two types of faults, "communication link interruption" and "navigation and positioning errors", a fixed penalty value P is set to offset part of the reward and prevent the model from ignoring the fault risk.
[0181] Communication link interruption penalty: If the connectivity status value of a link is 0 (interrupted), P1 = -5 (the penalty is 5 for each interrupted link, and the more links there are, the heavier the penalty).
[0182] Navigation and positioning error penalty: If the navigation signal value of a certain drone is <0.3 (positioning error >7 meters, out of tolerance), P2=-3 (3 penalty per out-of-tolerance drone).
[0183] IV. Complete Calculation of the Dynamic Reward Mechanism:
[0184] The final reward value R = R1 + ΣP1 + ΣP2, example:
[0185] Scenario 1: S=0.74, α=70% (γ=0.8), 1 link is interrupted (P1=-5), 1 drone has navigation error (P2=-3), then R=5.92 -5 -3=-2.08 (negative reward, model needs to adjust strategy);
[0186] Scenario 2: S=0.88, α=30% (γ=1), 0 links interrupted, 0 aircraft out of tolerance, R=8.8 +0 +0=8.8 (positive reward, model strategy is reasonable);
[0187] Scenario 3: S=0.95, α=40% (γ=1), 0 interruptions, 0 out-of-tolerance flights, R=9.5 +0 +0=9.5 (high reward, optimal model strategy).
[0188] The reward mechanism needs to be verified through simulation to ensure that the model receives high rewards in low-occlusion, high-connectivity, and no-out-of-error scenarios, and low rewards or even penalties in other scenarios, thus guiding the model to learn the optimal policy.
[0189] The system performs asynchronous decision-making by multi-agents, splitting the multidimensional state space matrix into local state segments for each drone. Each drone acts as an independent agent, calculating decision priorities based on its local segment. The global coordination layer of the distributed reinforcement learning model integrates the decisions of each agent, prioritizing low-occlusion areas as relay nodes, and allocating link priorities in reverse order of signal attenuation coefficients. This generates a topology optimization strategy that includes the three-dimensional coordinates of relay nodes and the link priority ranking.
[0190] Multi-agent asynchronous decision-making is the core of distributed reinforcement learning. It requires each drone to make autonomous decisions, and then global coordination to avoid conflicts, thus solving the problems of "low efficiency and inability to adapt to cluster scale in centralized decision-making." This step needs to clarify "how to make local decisions, how to integrate globally, and how to generate policies" to ensure that the decisions are efficient and in line with the optimization objectives.
[0191] I. Splitting Local State Fragments:
[0192] The multidimensional state space matrix is split into independent local state segments according to the UAV number. Each segment contains only the state parameters of that UAV and the link states of its 2-3 neighboring UAVs (reducing data transmission volume and enabling asynchronous decision-making). For example:
[0193] Local segment of UAV No. 1: (self-obscuration 70%, relative position 350 meters, link status with UAVs No. 2 and No. 5 0.5 / 0.6, self-navigation signal 0.53);
[0194] Local segment of UAV No. 2: (self-obscuration 30%, relative position 280 meters, link status with UAVs No. 1 and No. 3 0.5 / 0.7, self-navigation signal 0.75);
[0195] Local segment of UAV No. 3: (self-obscuration 90%, relative position 420 meters, link status with UAVs No. 2 and No. 4 0.7 / 0.2, self-navigation signal 0.35).
[0196] After being split up, each drone makes independent decisions through its local computing unit without waiting for other drones, achieving asynchronous operation (decision interval ≤ 100ms, adapting to the dynamic movement of drones).
[0197] II. Calculation of decision priority for independent agents:
[0198] Each drone, acting as an intelligent agent, calculates its "relay node candidate priority" based on local fragments. Higher priority drones are more suitable as relay nodes (relay nodes require low obstruction, high link status, and high navigation signal strength). The calculation formula is as follows:
[0199] Priority P = (1 - α / 100) × Average Link State Value × Navigation Signal Value
[0200] Where α is the self-occlusion degree, and the average link state value is the average link state of the local segment with neighboring UAVs.
[0201] Example calculation:
[0202] UAV No. 1: α=70%, average link status=(0.5+0.6) / 2=0.55, navigation signal 0.53, P=(1-0.7)×0.55×0.53=0.3×0.2915≈0.087;
[0203] UAV No. 2: α=30%, average link state=(0.5+0.7) / 2=0.6, navigation signal 0.75, P=(1-0.3)×0.6×0.75=0.7×0.45=0.315;
[0204] UAV No. 3: α=90%, average link status=(0.7+0.2) / 2=0.45, navigation signal 0.35, P=(1-0.9)×0.45×0.35=0.1×0.1575≈0.016;
[0205] UAV No. 4 (α=40%, average link 0.65, navigation 0.68): P=0.6×0.65×0.68≈0.265;
[0206] UAV No. 5 (α=60%, average link 0.58, navigation 0.62): P=0.4×0.58×0.62≈0.143.
[0207] Priority ranking: No. 2 (0.315) > No. 4 (0.265) > No. 5 (0.143) > No. 1 (0.087) > No. 3 (0.016), No. 2 and No. 4 are the optimal relay node candidates.
[0208] III. Decision Integration at the Overall Coordination Level:
[0209] The distributed reinforcement learning model (using the FedRL framework, with the global coordination layer deployed at the ground control center) receives the priority results from each agent and integrates them according to the following rules:
[0210] Relay node selection: The top two priority drones were selected as relay nodes (a cluster of 5 drones requires 2 relay nodes to ensure coverage of all drones), namely drones 2 and 4. Their 3D coordinates were obtained using a twin scene model: drone 2 (110.08°, 25.07°, 520m) and drone 4 (110.03°, 25.02°, 510m). Simultaneously, it was verified that the selected locations were in low-obscurity areas (α=30% for drone 2, α=40% for drone 4, meeting the requirements).
[0211] Link priority allocation: Extract the signal attenuation coefficient of all links from the terrain digital twin (calculated in step 2, unit dB / m), and sort them in reverse order of "the smaller the attenuation coefficient, the higher the priority" (smaller attenuation means less signal loss and better link). For example, the link list and attenuation coefficients are: 2-1 (0.185), 2-3 (0.190), 4-5 (0.182), 4-1 (0.188), and the priority order is 4-5 (1) > 2-1 (2) > 4-1 (3) > 2-3 (4).
[0212] IV. Generation of Topology Optimization Strategies:
[0213] Integrate relay node and link priorities to generate a structured policy, including "relay node information and link priority list," as shown in the example:
[0214] "Topology optimization strategy for drone swarms in complex terrain (canyon terrain)"
[0215] Relay node configuration:
[0216] Relay Node 1: Number 2, 3D coordinates (110.08°, 25.07° N, altitude 520m), grid occlusion 30%, responsible for covering UAVs 1 and 3;
[0217] Relay Node 2: Number 4, 3D coordinates (110.03°, 25.02° N, altitude 510m), grid occlusion 40%, responsible for covering UAVs 5 and 1;
[0218] Communication link priority (in reverse order of signal attenuation coefficient):
[0219] Priority 1: Links 4 and 5 (attenuation coefficient 0.182dB / m, SNR=28dB);
[0220] Priority 2: Link 2-1 (attenuation coefficient 0.185dB / m, SNR=25dB);
[0221] Priority 3: Link 4 to Link 1 (attenuation coefficient 0.188dB / m, SNR=23dB);
[0222] Priority 4: Links 2 and 3 (attenuation coefficient 0.190dB / m, SNR=21dB);
[0223] Policy validity period: 30 seconds (must be updated at fixed intervals to adapt to drone movement).
[0224] The strategy needs to be output to subsequent steps to provide relay nodes and link directions for antenna parameter adjustment, while ensuring that the strategy is executable (relay nodes cover all drones and there are no link conflicts).
[0225] S203, based on the relay node position and link direction in the topology optimization strategy, call the beam pattern database of the reconfigurable antenna, match the shielding angle in the terrain digital twin to dynamically adjust the antenna beamwidth and gain parameters, and generate a cooperative adaptation scheme of antenna parameters and topology.
[0226] Specifically, the topology optimization strategy parameters can be analyzed, and the identification information, three-dimensional coordinates, and main communication link direction angle of each relay node can be extracted from the topology optimization strategy to clarify the coverage requirements of each link and organize them into a strategy parameter table.
[0227] Analyzing topology optimization strategy parameters serves as a bridge between "topology strategy" and "antenna parameter adjustment." It requires accurately extracting core information about relay nodes and links to ensure that subsequent antenna parameters precisely match the topology structure, avoiding adaptation deviations due to missing information. This step necessitates clearly defining "what information to extract, how to define the information, and how to organize it," providing a clear input basis for calling the beam database.
[0228] I. Extraction and Definition of Core Parameters:
[0229] From the generated topology optimization strategy, four types of key parameters are extracted. The physical meaning and quantification standard of each type of parameter must be clearly defined to avoid ambiguity.
[0230] Relay node identification information: Used to uniquely distinguish different relay nodes, including "node number" (consistent with the UAV number, such as UAV No. 2 being numbered R2 as a relay node), "topographic grid index" (associated with the terrain digital twin, such as R2 being located in grid No. 102), and "node type" (primary relay / secondary relay, the primary relay is responsible for the core link, and the secondary relay provides auxiliary coverage, such as R2 being the primary relay and R4 being the secondary relay).
[0231] 3D coordinates of relay nodes: These reflect the spatial location of the relay nodes and directly affect the coverage direction of the antenna beam. They adopt the WGS84 coordinate system and are formatted as "longitude (°), latitude (°), altitude (m)". For example, the 3D coordinates of R2 are (110.08°, 25.07°, 520m), and those of R4 are (110.03°, 25.02°, 510m).
[0232] Main communication link azimuth angle: Defines the main beam direction for communication between relay nodes and other UAVs, expressed as an azimuth angle with "true north as 0° and clockwise as positive" (accuracy ±1°). 2-3 main link azimuth angles need to be extracted for each relay node (covering the main communication targets). For example, R2's main links include a azimuth angle of 30° for "R2→UAV 1" (azimuth angle from R2 to UAV 1) and a azimuth angle of 150° for "R2→UAV 3"; R4's main links include a azimuth angle of 330° for "R4→UAV 5" and a azimuth angle of 60° for "R4→UAV 1".
[0233] Link coverage requirements: Based on the relative distance between drones (the relative positions calculated in the above steps), the communication radius (unit: m) centered on the relay node is determined to ensure coverage of the target drone without wasting energy. For example, if the relative distance between R2 and drone 1 is 180m, the coverage requirement is 200m (leaving a 20m redundancy); if the relative distance between R2 and drone 3 is 220m, the coverage requirement is 250m.
[0234] II. Textualized version of the strategy parameter table:
[0235] Because tables are prohibited, parameters must be organized using continuous text grouped by "relay node" to ensure the information is structured and easy to read:
[0236] "Topology optimization strategy parameter table (canyon terrain cluster, 5 drones, 2 relay nodes):"
[0237] Relay node R2 (primary relay, UAV 2, belonging to grid 102):
[0238] 3D coordinates: Longitude 110.08°, Latitude 25.07°, Altitude 520m;
[0239] Main communication link:
[0240] Link 1: R2 → UAV No. 1, azimuth angle 30°, coverage range required 200m;
[0241] Link 2: R2→UAV No. 3, azimuth angle 150°, coverage range required 250m;
[0242] Relay node R4 (secondary relay, UAV 4, belonging to grid 104):
[0243] 3D coordinates: Longitude 110.03°, Latitude 25.02°, Altitude 510m;
[0244] Main communication link:
[0245] Link 3: R4 → UAV No. 5, azimuth angle 330°, coverage range required 180m;
[0246] Link 4: R4 → UAV No. 1, azimuth angle 60°, coverage range required 220m.
[0247] The parameter table must include the "strategy generation time" (e.g., 2024-11-01 10:05:00) and the "validity period" (30 seconds) to ensure that subsequent steps use the latest strategy parameters.
[0248] Call the beam pattern database, retrieve the matching initial beam parameters of the reconfigurable antenna from the database based on the link direction angle in the strategy parameter table, and extract the antenna parameter adjustment boundary to obtain the initial antenna configuration set;
[0249] The beam pattern database is a core resource for storing the correspondence between "azimuth angle and beam parameters". Calling this database can quickly obtain the initial antenna parameters for the adapted link direction, avoiding blind adjustments; extracting the adjustment boundary can ensure that subsequent parameter adjustments do not exceed the hardware capabilities, solving the problem of "antenna parameters going out of range causing hardware failure".
[0250] I. Structure and Contents of the Beam Pattern Database:
[0251] The database is built based on measured data of reconfigurable antennas (such as phased array antennas, model AD9361, supporting beamwidth adjustment of 15°-120° and gain adjustment of 8dB-20dB), storing the correspondence between "link direction angle - beamwidth - gain - radiation pattern", where:
[0252] Link direction angle: Divided in 5° intervals (0°, 5°, 10°…360°), covering all possible communication directions;
[0253] Initial beam parameters: For each azimuth angle, store the "default beamwidth" (the optimal width when there is no obstruction in that direction) and the "default gain" (the optimal value to balance coverage and power consumption). For example, for azimuth angles of 0°-60° (open direction), the default beamwidth is 60° and the gain is 15dB; for azimuth angles of 60°-120° (east side of the canyon wall), the default beamwidth is 80° and the gain is 14dB.
[0254] Beam pattern data: Stores the beam pattern (energy distribution curve) corresponding to each parameter for subsequent verification of coverage, but it is not called during the initial configuration stage; only the beam width and gain are extracted.
[0255] The database uses a storage method of local caching + cloud backup. The local cache ensures that the call latency is ≤50ms, which meets the real-time requirements.
[0256] II. Initial beam parameter retrieval process:
[0257] Based on the link direction angle in the strategy parameter table, the initial parameters are set in the database using the "nearest match" method (direction angle deviation ≤ 5°), as shown in the example below:
[0258] Link 1 (R2→UAV No. 1, azimuth angle 30°): Search the database for the parameters corresponding to the azimuth angle 30°, and match the default beamwidth 60° and default gain 15dB;
[0259] Link 2 (R2→UAV No. 3, azimuth angle 150°): Search azimuth angle 150°, matched with default beamwidth 80°, default gain 14dB;
[0260] Link 3 (R4→5 UAV, azimuth angle 330°): Search azimuth angle 330° (i.e. -30°), matched with default beamwidth 70°, default gain 15dB;
[0261] Link 4 (R4→UAV No. 1, azimuth angle 60°): Search azimuth angle 60°, matched with default beamwidth 65°, default gain 15dB.
[0262] If there is no perfectly matched azimuth angle (e.g., azimuth angle 33°), then take the average of the parameters of the closest azimuth angle 30° or 35°. For example, for 33°, take the average of 30° (60° / 15dB) and 35° (62° / 15dB), with a beamwidth of 61° and a gain of 15dB.
[0263] III. Extraction of Antenna Parameter Adjustment Boundaries:
[0264] The adjustment boundary is the hardware physical limitation of the reconfigurable antenna, which needs to be extracted from the database to ensure that subsequent adjustments do not exceed the capability range. Parameters include:
[0265] Beamwidth adjustment boundaries: minimum 15° (narrow beam, high directivity), maximum 120° (wide beam, large coverage), adjustment step 5° (the smallest adjustment unit supported by hardware).
[0266] Gain adjustment boundaries: minimum 8dB (low gain, low power consumption), maximum 20dB (high gain, long coverage), adjustment step 1dB;
[0267] Correlation constraint: There is a negative correlation between beamwidth and gain (the narrower the beam, the higher the gain). For example, the maximum gain is 20dB when the beamwidth is 15°, and the minimum gain is 8dB when the beamwidth is 120°. This constraint will be marked in the database to avoid conflicts in subsequent adjustments.
[0268] IV. Formation of the initial antenna configuration set:
[0269] Integrate the retrieved initial parameters and adjustment boundaries, and organize them by "link number" to form an initial configuration set. Example:
[0270] Initial antenna configuration set (based on beam pattern database retrieval):
[0271] Link 1 (R2→1, direction angle 30°):
[0272] Initial beamwidth: 60°, adjustment range 15°-120°;
[0273] Initial gain: 15dB, adjustment range: 8dB-20dB;
[0274] Association constraint: For every 5° decrease in beamwidth, the gain can be increased by 1dB;
[0275] Link 2 (R2→3, azimuth angle 150°):
[0276] Initial beamwidth: 80°, adjustment range 15°-120°;
[0277] Initial gain: 14dB, adjustment range: 8dB-20dB;
[0278] Association constraint: For every 5° decrease in beamwidth, the gain can be increased by 1dB;
[0279] Link 3 (R4→5, azimuth angle 330°):
[0280] Initial beamwidth: 70°, adjustment range: 15°-120°;
[0281] Initial gain: 15dB, adjustment range: 8dB-20dB;
[0282] Association constraint: For every 5° decrease in beamwidth, the gain can be increased by 1dB;
[0283] Link 4 (R4→1, azimuth angle 60°):
[0284] Initial beamwidth: 65°, adjustment range: 15°-120°;
[0285] Initial gain: 15dB, adjustment range: 8dB-20dB;
[0286] Association constraint: For every 5° decrease in beamwidth, the gain can be increased by 1 dB.
[0287] The initial configuration set needs to verify whether the parameters meet the coverage requirements. For example, if the initial beamwidth of link 1 is 60° and the gain is 15dB, the signal strength within a 200m coverage area should be ≥-85dBm (communication threshold) to be considered qualified.
[0288] Dynamically match the shielding angle adjustment parameters, query the maximum terrain shielding angle on each communication link path from the terrain digital twin; classify and adjust the antenna parameters according to the size of the shielding angle. When the shielding angle is in a preset small range, compress the beamwidth and increase the gain. When the shielding angle is in a preset medium range, maintain parameter balance. When the shielding angle is in a preset large range, expand the beamwidth and reduce the gain to obtain the adjusted antenna parameters.
[0289] Terrain shielding angle is a key factor affecting beam coverage (the larger the shielding angle, the more easily the beam is blocked by the terrain). It is necessary to dynamically adjust antenna parameters to adapt to shielding conditions and solve the problem of "coverage failure due to initial parameters not considering terrain shielding." This step needs to clarify "how to query the shielding angle, how to classify and adjust it, and how to calculate the adjusted parameters" to ensure that the parameters are adapted to the terrain.
[0290] I. Method for querying the maximum terrain shielding angle:
[0291] Query the maximum shading angle on the link path from the terrain digital twin. The shading angle is defined as "the angle between the line connecting the link start point (relay node) and the highest point of the terrain on the path, and the horizontal direction of the link" (unit: °). Query steps:
[0292] Link path modeling: In the digital twin, the three-dimensional coordinates of the relay node and the target UAV are connected to generate link path segments (such as the path R2 (520m) → UAV 1 (510m)).
[0293] Identification of the highest point of terrain: Along the path segment, take a terrain sampling point every 10m, extract the elevation value of each sampling point, and find the point corresponding to the maximum value (i.e., the highest point). For example, the elevations of the sampling points on the link 1 path are 520m, 525m, 530m, 528m...510m, and the highest point has an elevation of 530m, located 80m east of R2.
[0294] Shielding angle calculation: The shielding angle θ is calculated using trigonometric functions. The formula is θ = arctan [(H_h - H_s) / d], where H_h is the elevation of the highest point, H_s is the elevation of the link's starting point, and d is the horizontal distance from the starting point to the highest point. For example, in link 1, H_h = 530m, H_s = 520m, d = 80m, θ = arctan [(530-520) / 80] = arctan (0.125) ≈ 7.13°, rounded to 7° (maximum terrain shielding angle).
[0295] Use this method to query the maximum shielding angle of all links. Example results: Link 1θ=7°, Link 2θ=45°, Link 3θ=15°, Link 4θ=30°.
[0296] II. Shading Angle Preset Range and Adjustment Rules:
[0297] Based on the impact of the shielding angle on beam coverage, the shielding angle is preset to three ranges, each corresponding to a different parameter adjustment strategy. The rule is based on "small shielding angles use narrow beams with high gain (to reduce interference), large shielding angles use wide beams with low gain (to expand coverage)":
[0298] Preset small range: 0°≤θ≤30° (small shielding, beam is not easily blocked), adjustment strategy is "compress beamwidth (reduce sidelobe interference) and increase gain (enhance signal strength)", adjustment range: beamwidth reduced by 20%, gain increased by 2dB;
[0299] Preset medium range: 30°<θ≤60° (medium shielding, part of the beam is blocked), the adjustment strategy is "maintain parameter balance (neither compress nor expand, taking into account both coverage and interference)", the adjustment range: beamwidth remains unchanged, gain remains unchanged;
[0300] Preset large range: 60°<θ≤90° (large shielding, beam is easily blocked), the adjustment strategy is "expand beam width (diffuse to cover the blocked area) and reduce gain (avoid energy waste)", the adjustment range is: beam width increased by 20% and gain decreased by 2dB.
[0301] The adjustment range must conform to the antenna adjustment boundary. If the adjustment exceeds the boundary, the boundary value is taken. For example, the initial beamwidth is 15° (minimum boundary). At small shielding angles, it cannot be compressed, only the gain is increased.
[0302] III. Example of adjusting antenna parameter calculation:
[0303] Based on the initial parameters and the occlusion angle query results, the adjusted parameters are calculated for each link:
[0304] Link 1 (θ=7°, small range):
[0305] Initial beamwidth 60° → Compressed by 20% → 60° × (1-20%) = 48° (meets the 15°-120° boundary);
[0306] Initial gain 15dB → boost 2dB → 15dB + 2dB = 17dB (meets the 8dB-20dB boundary).
[0307] Adjusted parameters: beamwidth 48°, gain 17dB;
[0308] Link 2 (θ=45°, medium range):
[0309] Initial beamwidth 80° → remain unchanged → 80°;
[0310] Initial gain 14dB → remains unchanged → 14dB;
[0311] Adjusted parameters: beamwidth 80°, gain 14dB;
[0312] Link 3 (θ=15°, small range):
[0313] Initial beamwidth 70° → Compressed by 20% → 70° × 0.8 = 56°;
[0314] Initial gain 15dB → Increase 2dB → 17dB;
[0315] Adjusted parameters: beamwidth 56°, gain 17dB;
[0316] Link 4 (θ=30°, small range, since θ=30° is at the upper limit of the small range):
[0317] Initial beamwidth 65° → Compressed by 20% → 65° × 0.8 = 52°;
[0318] Initial gain 15dB → Increase 2dB → 17dB;
[0319] Adjusted parameters: beamwidth 52°, gain 17dB.
[0320] If a link has an angle of θ = 70° (large range), the initial beamwidth is 60° → expanded by 20% → 72°, and the initial gain is 15dB → reduced by 2dB → 13dB. After adjustment, the parameters meet the boundary and are deemed qualified.
[0321] Collaborative verification of adaptability involves checking the matching between the adjusted antenna parameters and the relay node locations, and correcting any conflicting parameters. The relay node and link information in the topology optimization strategy are associated with the corresponding antenna parameters by UAV ID to generate a collaborative adaptation scheme for antenna parameters and topology.
[0322] Collaborative verification is crucial to ensuring that "antenna parameters do not conflict with the topology". It is necessary to check whether the parameters are adapted to the relay node location (such as whether the beam is blocked by the rock wall in the canyon), correct the conflict and associate the ID to form an executable collaborative adaptation scheme, thus solving the problem of "parameters and topology being out of sync, resulting in execution failure".
[0323] I. Core dimensions of compatibility check:
[0324] The compatibility between the adjusted parameters and the relay node locations was checked from three dimensions: coverage area, terrain conflict, and energy consumption balance.
[0325] Coverage matching: Verify whether the signal strength of the adjusted beam parameters within the coverage area meets the standard (≥-85dBm). The formula is: Signal strength = Gain - 20log(d) - Attenuation coefficient × d, where d is the coverage distance, and the attenuation coefficient is obtained from the digital twin. For example, after adjustment, the gain of link 1 is 17dB, d=200m, and the attenuation coefficient is 0.185dB / m. The signal strength is: 17-20log200-0.185×200≈17-46.02-37≈-66.02dBm≥-85dBm, and the coverage is qualified.
[0326] Terrain conflict check: Simulate the beam coverage area in the digital twin to check for situations where the beam is completely blocked by the terrain. For example, if relay node R2 is located on the west side of a canyon, and after adjusting the beamwidth of link 2 to 80° (azimuth angle 150°, pointing to the east rock wall), the simulation shows that 30% of the beam is blocked by the rock wall, indicating a conflict.
[0327] Energy balance check: Ensure the total energy consumption of a single repeater node is less than or equal to the hardware limit (e.g., R2's antenna has a maximum energy consumption of 10W). The energy consumption calculation formula is: Energy Consumption = Gain 2 × Beamwidth / 1000, for example, the power consumption of Link 1 in R2 = 17 2 ×48 / 1000≈13.87W, Link 2 energy consumption = 14 2×80 / 1000≈15.68W, total energy consumption≈29.55W>10W, there is an energy consumption conflict.
[0328] II. Methods for correcting conflicting parameters:
[0329] In response to the conflicts identified during the inspection, corrections were made according to the principle of "prioritizing coverage resolution and rebalancing energy consumption."
[0330] Terrain conflict correction: Link 2 beam is blocked by 30% by the rock wall. The beam width needs to be expanded to diffract and cover the area. The beam width is expanded from 80° to 100° (still within the 120° boundary). The gain is kept at 14dB. Coverage is re-simulated, and the blocking ratio is reduced to 10% (acceptable).
[0331] Energy consumption conflict correction: R2's total energy consumption is too high. The gain of non-core links needs to be reduced. Link 2 (non-core, only covering UAV 3) gain is reduced from 14dB to 12dB, energy consumption = 12. 2 ×100 / 1000 = 14.4W, Link 1 power consumption is 13.87W, total power consumption ≈ 28.27W, still exceeding the limit. Further reduce the gain of Link 1 from 17dB to 16dB, power consumption = 16. 2 ×48 / 1000≈12.29W, total power consumption≈12.29+14.4=26.69W, still exceeding the limit. Finally, the beamwidth of link 2 was adjusted to 90°, gain 12dB, and power consumption = 12. 2 ×90 / 1000=12.96W, total energy consumption≈12.29+12.96=25.25W. Although it does not fully meet the standard, it is close and the coverage is qualified, so it is considered acceptable (subsequent iterations and optimizations).
[0332] The corrected parameters for Link 2 are: beamwidth 90°, gain 12dB; the parameters for Link 1 are: beamwidth 48°, gain 16dB.
[0333] III. Generation of Collaborative Adaptation Scheme:
[0334] By associating the drone ID with the relay node, link information, and corrected antenna parameters, a structured scheme is generated. Example:
[0335] "UAV swarm antenna parameters in complex terrain - topology cooperative adaptation scheme (canyon terrain)":
[0336] Relay node R2 (UAV No. 2, coordinates 110.08°, 25.07°, 520m):
[0337] Link 1 (R2→UAV No. 1): Azimuth angle 30°, coverage range 200m, antenna parameters (beamwidth 48°, gain 16dB).
[0338] Link 2 (UAV R2→3): Azimuth angle 150°, coverage range 250m, antenna parameters (beamwidth 90°, gain 12dB).
[0339] Node power consumption: ≈25.25W (close to the hardware limit of 10W, further optimization is needed);
[0340] Relay node R4 (UAV No. 4, coordinates 110.03°, 25.02°, 510m):
[0341] Link 3 (UAV R4→5): Azimuth angle 330°, coverage range 180m, antenna parameters (beamwidth 56°, gain 17dB).
[0342] Link 4 (R4→UAV No. 1): Azimuth angle 60°, coverage range 220m, antenna parameters (beamwidth 52°, gain 17dB).
[0343] Node energy consumption: ≈17 2 ×56 / 1000 + 17 2 ×52 / 1000≈16.46+15.14=31.6W (further optimization is needed);
[0344] The plan is valid for 30 seconds and needs to be updated at intervals to adapt to drone movement.
[0345] Coverage verification: All link signal strength ≥ -75dBm, coverage is qualified.
[0346] The solution needs to be marked with "items to be optimized" (such as energy consumption overrun) to provide direction for subsequent iterative optimization, while ensuring that the parameters can be directly called by the antenna adjustment module (such as the duty cycle of the hardware control signal corresponding to a beamwidth of 48°).
[0347] S204. Based on the aforementioned collaborative adaptation scheme, the cluster communication quality and navigation continuity indicators are collected in real time. The indicator deviations are fed back to the terrain digital twin for scene correction, driving the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjusting the reconfigurable antenna parameters, and generating a cluster topology dynamic optimization result that adapts to complex terrain changes.
[0348] Specifically, a collaborative adaptation scheme can be loaded, and the antenna parameters and topology strategies in the scheme can be extracted as a baseline configuration to clarify the expected communication coverage and navigation and positioning accuracy requirements of each UAV and form a baseline parameter table for the scheme.
[0349] Loading the collaborative adaptation scheme is the foundation for subsequent indicator collection and deviation analysis. It requires accurately extracting baseline parameters from the "antenna-topology" dual dimensions, while clearly defining the expected performance targets to avoid a lack of reference for subsequent data collection. This step needs to address the problem of "unclear baselines leading to a lack of standards for deviation judgment," ensuring that each parameter has a clear expected value and physical meaning.
[0350] I. Loading and core parameter extraction of the collaborative adaptation scheme:
[0351] By calling the collaborative adaptation solution through the system interface (storage format is JSON, local cache latency ≤50ms), two types of core baseline configurations are extracted:
[0352] Antenna parameter benchmark: Extract adjusted antenna parameters based on UAV ID (relay node and regular node), including "beamwidth (unit: °), gain (unit: dB), and beam direction angle (unit: °)," which must correspond one-to-one with the link. For example:
[0353] Relay node R2 (UAV 2): Link 1 (R2→1) beamwidth 48°, gain 16dB, azimuth angle 30°; Link 2 (R2→3) beamwidth 90°, gain 12dB, azimuth angle 150°;
[0354] Relay node R4 (UAV No. 4): Link 3 (R4→No. 5) beamwidth 56°, gain 17dB, azimuth angle 330°; Link 4 (R4→No. 1) beamwidth 52°, gain 17dB, azimuth angle 60°;
[0355] Ordinary nodes (No. 1, No. 3, No. 5): The antenna parameters follow the relay node configuration by default (e.g., UAV No. 1 receives the beam of R2, and the parameters are consistent with R2 link 1).
[0356] Topology strategy baseline: Extract the three-dimensional coordinates of relay nodes, link priority ranking, and link distance, for example:
[0357] Relay node coordinates: R2 (110.08°, 25.07°, 520m), R4 (110.03°, 25.02°, 510m);
[0358] Link priority: Link 3 (R4→5) > Link 1 (R2→1) > Link 4 (R4→1) > Link 2 (R2→3);
[0359] Link distances: R2→1 180m, R2→3 220m, R4→5 150m, R4→1 200m.
[0360] II. Definitions of expected communication coverage and navigation positioning accuracy:
[0361] Expected communication coverage: Calculated based on link distance and antenna gain, using the formula "Coverage radius = Link distance × 1.2" (1.2 is a redundancy factor to prevent coverage failure due to small drone movements), unit: meters. For example:
[0362] The distance of link R2→1 is 180m, and the expected coverage radius is 180×1.2=216m;
[0363] The distance between R4 and 5 is 150m, and the expected coverage radius is 150 × 1.2 = 180m.
[0364] At the same time, the "minimum signal strength threshold" (≥-85dBm, to ensure communication quality) within the coverage area is clearly defined. This threshold is set based on the receiving sensitivity (-90dBm) of the UAV communication module, with a 5dB redundancy.
[0365] Expected navigation and positioning accuracy: Referring to the "Technical Requirements for Unmanned Aerial Vehicle (UAV) Swarm Navigation" and considering the characteristics of complex terrain, two types of indicators are set:
[0366] Positioning error: ≤3 meters for relay nodes (high-precision positioning is required to ensure accurate beam pointing), ≤5 meters for ordinary nodes;
[0367] Navigation continuity: The duration of continuous positioning interruption is ≤1 second (to avoid topology chaos caused by positioning interruption).
[0368] III. Textual presentation of the scheme's baseline parameter table:
[0369] Organized according to the logic of "Drone ID - Parameter Type - Baseline Value - Expected Target", example:
[0370] "Scheme Baseline Parameter Table (Canyon Topographic Cluster, Data Collection Time: 2024-11-01 10:10:00):"
[0371] Drone 2 (Relay Node R2):
[0372] Antenna parameters: Link 1 (→ No. 1) beamwidth 48°, gain 16dB, azimuth 30°; Link 2 (→ No. 3) beamwidth 90°, gain 12dB, azimuth 150°;
[0373] Topology reference: coordinates (110.08°, 25.07°, 520m), link 2 priority 4;
[0374] Expected goals: Link 1 coverage radius 216m (signal strength ≥ -85dBm), positioning error ≤ 3m;
[0375] Drone 4 (Relay Node R4):
[0376] Antenna parameters: Link 3 (→ 5) beamwidth 56°, gain 17dB, azimuth 330°; Link 4 (→ 1) beamwidth 52°, gain 17dB, azimuth 60°;
[0377] Topology reference: coordinates (110.03°, 25.02°, 510m), link 3, priority 1;
[0378] Expected goals: Link 3 coverage radius 180m (signal strength ≥ -85dBm), positioning error ≤ 3m;
[0379] Drone 1 (Ordinary Node):
[0380] Antenna parameters: Receive R2 Link 1 (48° / 16dB), R4 Link 4 (52° / 17dB);
[0381] Expected goals: Positioning error ≤ 5m, navigation interruption duration ≤ 1s;
[0382] Drones 3 and 5 (ordinary nodes):
[0383] Expected targets: Positioning error ≤ 5m, continuous navigation interruption duration ≤ 1s, received signal strength ≥ -85dBm.
[0384] Based on the baseline parameter table of the scheme, the performance indicators are collected, and the communication quality indicators and navigation continuity indicators of the UAV cluster are collected at preset time intervals. The actual collected values are compared with the expected values in the baseline parameter table of the scheme, and the deviation rate of each indicator is calculated to form an indicator deviation set.
[0385] Performance metric collection is crucial for verifying the effectiveness of the solution. It needs to cover both the "communication and navigation" dimensions at fixed intervals, quantifying the gap between actual and expected results through deviation rates, and providing data support for subsequent adjustments. This step needs to address the issues of "incomplete metric collection and lack of standardized deviation calculations," ensuring that deviation results are reproducible.
[0386] I. Setting the preset time interval:
[0387] The interval needs to balance "real-time performance" and "data stability." Referring to the dynamic movement speed of the drone swarm (approximately 5 m / s in a canyon) and the frequency of terrain changes, it is set to 10 seconds (i.e., data is collected every 10 seconds). If a link signal fluctuates frequently (e.g., absolute deviation rate > 10%), it can be dynamically shortened to 5 seconds to avoid missing instantaneous anomalies; if the indicators are stable (absolute deviation rate < 5%), it can be extended to 20 seconds to reduce computational burden.
[0388] II. Methods for collecting communication quality and navigation continuity indicators:
[0389] Communication quality indicators:
[0390] Link signal strength (unit: dBm): Collected through the signal detection interface of the UAV communication module. 10 data points are collected for each link each time and the average value is taken. For example, the collected values for link R2→1 are -70dBm, -72dBm...-68dBm, with an average of -70dBm.
[0391] Link packet loss rate (unit: %): The number of lost packets is calculated through a ping test (sending 100 data packets). The formula is "packet loss rate = number of lost packets / total number of data packets × 100%". For example, if 5 out of 100 packets are lost, the packet loss rate is 5%.
[0392] Communication connection duration (unit: s): Records the time during which the link remains continuously connected (signal strength ≥ -85dBm, packet loss rate ≤ 10%), for example, the R4→5 link remains continuously connected for 120 seconds.
[0393] Navigation continuity metrics:
[0394] Positioning error (unit: m): Calculated by fusion data of RTK-GPS and IMU, comparing the Euclidean distance between the actual position of the UAV and the true value (provided by the ground base station), for example, the true coordinates (110.08°, 25.07°, 520m) and the actual coordinates (110.08°, 25.07°, 523m). =3m;
[0395] Navigation interruption duration (unit: s): Records the continuous time when the positioning error is >5m (ordinary node) or >3m (relay node). For example, if the positioning error of UAV No. 1 is >5m for 8 consecutive seconds, the interruption duration is 8 seconds.
[0396] III. Calculation of Deviation Rate and Formation of Index Deviation Set:
[0397] The deviation rate is used to quantify the degree of deviation between the actual value and the expected value. The formula is "Deviation rate = (Actual collected value - Expected value) / Expected value × 100%". The result is rounded to two decimal places; a positive number indicates that the actual value exceeds the expectation, and a negative number indicates that the actual value does not meet the expectation. An example of data collection and calculation is provided below, using a benchmark parameter as an example:
[0398] Signal strength of link 1 (R2→1): expected ≥ -85dBm (take -80dBm as the target value), actual -70dBm, deviation rate = (-70 - (-80)) / (-80)×100% = -12.50% (better than expected);
[0399] Packet loss rate of Link 2 (R2→3): Expected ≤5%, actual 8%, deviation rate = (8-5) / 5×100%=60.00% (exceeding expectations);
[0400] The positioning error of UAV No. 1 was expected to be ≤5m, but actually 6m. The deviation rate was (6-5) / 5×100%=20.00% (exceeding expectations).
[0401] R4 navigation interruption duration: expected ≤1s, actual 0s, deviation rate = (0-1) / 1×100% = -100.00% (better than expected).
[0402] The indicator deviation set is organized according to "Indicator Type - Drone ID - Actual Value - Expected Value - Deviation Rate", example:
[0403] "Indicator Deviation Set (Collection Time: 2024-11-01 10:10:10):"
[0404] Communication quality indicators:
[0405] Link 1 (R2→1): Actual signal strength -70dBm, expected -80dBm, deviation rate -12.50%;
[0406] Link 2 (R2→3): Actual packet loss rate 8%, expected 5%, deviation rate 60.00%;
[0407] Link 3 (R4→5): Actual connection time 100s, expected 120s, deviation rate -16.67%;
[0408] Navigation continuity metrics:
[0409] UAV No. 1: Positioning error: actual 6m, expected 5m, deviation rate 20.00%;
[0410] Drone No. 3: The actual downtime was 2 seconds, the expected downtime was 1 second, and the deviation rate was 100.00%.
[0411] R2 (No. 2): Positioning error: actual 2m, expected 3m, deviation rate -33.33%.
[0412] Feedback indicator deviation correction twin, filter out anomalies where the indicator deviation concentration exceeds the fault tolerance range of the scheme, associate the anomaly deviation with the corresponding grid cell of the terrain digital twin, update the shading area boundary and signal attenuation coefficient of the cell, and generate the corrected terrain digital twin.
[0413] Deviation feedback and twin correction are key to achieving "dynamic optimization." It requires locating the root cause of anomalies (often inaccurate terrain parameter labeling) and updating the static features of the twin to provide accurate scene input for subsequent strategy iterations. This step needs to address the problems of "lack of root cause location for abnormal deviations" and "disconnect between the twin and the actual terrain," ensuring that the corrected scene closely resembles the real environment.
[0414] I. Setting the Fault Tolerance Range of the Solution:
[0415] The tolerance range is the critical value that distinguishes between "normal fluctuations" and "abnormal deviations." It is set based on industry standards for UAV communication and navigation and the characteristics of terrain interference. Different indicators have different tolerance ranges due to varying sensitivities.
[0416] Communication metrics: signal strength deviation ±15%, packet loss rate deviation ±20%, connection time deviation ±25%;
[0417] Navigation metrics: Positioning error deviation rate ±30%, Interruption duration deviation rate ±50%.
[0418] Deviations exceeding this range are classified as "abnormal items" and require further analysis of their root causes.
[0419] II. Screening and root cause identification of anomalies:
[0420] Anomalies are filtered from the set of indicator deviations, and the corresponding terrain grids are located by associating them through "link path - grid cell". Example:
[0421] Anomaly 1: The packet loss rate deviation rate of link 2 (R2→3) is 60.00% (exceeding ±20%). The link path covers terrain grid 103 (between R2 and UAV 3). It is preliminarily judged that the occlusion area of this grid is insufficiently labeled or the attenuation coefficient is too low.
[0422] Anomaly 2: The downtime deviation rate of UAV No. 3 is 100.00% (exceeding ±50%). UAV No. 3 is located in grid 103. It is speculated that the navigation signal of this grid is severely blocked. The original marked blockage of 70% may be too low.
[0423] Normal item: Link 1 deviation rate -12.50% (within ±15%), no correction required.
[0424] III. Correction operations for terrain digital twins:
[0425] For grid cells associated with anomalies, static features are updated through "field survey data completion + electromagnetic wave propagation model recalculation":
[0426] Boundary update of the obscured area: For grid 103, a rescan using LiDAR on a drone revealed that the original obscured area only marked the western rock face, omitting the protruding rock on the northern side (530m high, higher than the flight altitude of drone 3 at 510m). The boundary of the obscured area was expanded from "10m west" to "10m west + 8m north", and the obscuration was updated from 70% to 85%.
[0427] Signal attenuation coefficient recalculation: Based on the updated shielding degree (85%) and environmental electromagnetic parameters (interference intensity -75dBm), the attenuation coefficient is recalculated using the free space propagation model. The formula is "attenuation coefficient = 0.185×(1 + shielding degree / 100)" (the attenuation coefficient increases by 10% for every 10% increase in shielding degree). The original coefficient was 0.185dB / m, and the updated coefficient is 0.185×(1+85 / 100)=0.185×1.85≈0.342dB / m.
[0428] Other grid verifications: Simultaneously verify the shading and attenuation parameters of the associated adjacent grids 102 (where R2 is located) and 104 (where R4 is located) to ensure no chain errors.
[0429] IV. Revision of the terrain digital twin generation:
[0430] Integrate all updated grid parameters to generate a new terrain feature map and digital twin. Example:
[0431] "The corrected digital twin of the terrain (canyon terrain, grid size 5m×5m):"
[0432] Grid 103 (related anomalies 1 and 2):
[0433] Shielding area: 10m on the west side + 8m on the north side (originally 10m on the west side), shielding degree 85% (originally 70%).
[0434] Signal attenuation coefficient: 0.342dB / m (originally 0.185dB / m);
[0435] Elevation values: max 530m, min 500m, avg 515m;
[0436] Grid 102 (where R2 is located):
[0437] Shielding level 30% (no change), attenuation coefficient 0.185dB / m (no change);
[0438] Grid 104 (where R4 is located):
[0439] Shielding level 40% (no change), attenuation coefficient 0.190dB / m (no change);
[0440] Other grids: Parameters remain unchanged, keep original configuration.
[0441] The corrected twin needs to pass the "signal strength measurement verification". For example, the measured signal strength of link 2 in grid 103 drops from -75dBm to -88dBm, which is consistent with the updated attenuation coefficient calculation result (error ≤2dB), and the correction is deemed qualified.
[0442] The iterative optimization strategy and parameters are input into the distributed reinforcement learning model, and the topology optimization strategy is adjusted in combination with parameter conflicts in the collaborative adaptation scheme. At the same time, the antenna parameters are corrected according to the new topology strategy, and the adjusted indicators are collected to verify whether the deviation rate meets the scheme requirements. After the target is met, the optimized topology strategy and antenna parameters are integrated to generate a cluster topology dynamic optimization result adapted to complex terrain changes.
[0443] Iterative optimization is the final step in achieving "adaptation to complex terrain changes." It requires updating the topology strategy through reinforcement learning and simultaneously adjusting antenna parameters until the performance targets are met, forming a closed-loop optimization. This step needs to address the issues of "lack of dynamic adjustment of strategies and parameters, and substandard optimization results," ensuring that the final solution adapts to the corrected terrain.
[0444] I. Adjustment of Topology Optimization Strategy:
[0445] The corrected terrain digital twin is input into a distributed reinforcement learning model (FedRL framework). Addressing parameter conflicts in the original scheme (such as excessive attenuation in the R2→3 link), the strategy is adjusted through asynchronous decision-making by multiple agents.
[0446] Relay node location adjustment: The attenuation between the original relay node R2 (grid 102) and UAV No. 3 (grid 103) is too large. R2 will be moved to the adjacent low-shading grid 101 (shading degree 20%, attenuation coefficient 0.170dB / m), with new coordinates (110.07°, 25.06°, 520m).
[0447] Link priority rearrangement: Because the attenuation of link 2 (R2→3) is reduced after R2 is moved, its priority is increased from 4 to 2; link 3 (R4→5) retains priority 1;
[0448] Link coverage adjustment: The distance between links R2 and R3 has been shortened from 220m to 200m, and the expected coverage radius has been adjusted from 250m to 240m.
[0449] Example of the adjusted new topology optimization strategy:
[0450] "Iterative topology optimization strategy (based on modified twin):"
[0451] Relay node R2 (UAV No. 2): New coordinates (110.07°, 25.06°, 520m), located in grid 101 (occlusion 20%).
[0452] Relay node R4 (UAV No. 4): Coordinates remain unchanged (110.03°, 25.02°, 510m);
[0453] Link priority: Link 3 (R4→5) > Link 2 (R2→3) > Link 1 (R2→1) > Link 4 (R4→1).
[0454] II. Synchronous correction of antenna parameters:
[0455] Adjust the antenna parameters according to the link direction and distance in the new topology strategy to ensure adaptation to the new propagation path:
[0456] Link 2 (R2→3): After R2 is moved, the azimuth angle is adjusted from 150° to 145°, the attenuation coefficient is reduced to 0.170dB / m, no need to expand the beamwidth, maintain 90°, and the gain is increased from 12dB to 14dB (to compensate for the energy consumption of the shortened distance).
[0457] Link 1 (R2→1): After R2 is moved, the distance increases from 180m to 190m, the beamwidth is slightly adjusted from 48° to 50° (to expand coverage), and the gain remains at 16dB;
[0458] Other links: No significant parameter adjustments, keep the original configuration.
[0459] Example of corrected antenna parameters:
[0460] "Iterative antenna parameter configuration: "
[0461] Link 2 (R2→3): Beamwidth 90°, gain 14dB, azimuth angle 145°;
[0462] Link 1 (R2→1): Beamwidth 50°, gain 16dB, azimuth angle 32°;
[0463] Other links: Parameters remain the same as the original solution.
[0464] III. Indicator Validation and Result Integration:
[0465] The adjusted metrics are collected at 10-second intervals to verify whether the deviation rate meets the tolerance range.
[0466] Link 2 packet loss rate: actual 5%, expected 5%, deviation rate 0.00% (meets standards);
[0467] Interruption duration of UAV No. 3: Actual 1 second, expected 1 second, deviation rate 0.00% (meets the standard).
[0468] Link 1 signal strength: actual -72dBm, expected -80dBm, deviation rate -10.00% (meets standards).
[0469] All performance indicators are within the tolerance range. The optimized topology strategy and antenna parameters are integrated to generate the final dynamic optimization result.
[0470] "Results of dynamic optimization of UAV swarm topology under complex terrain (canyon terrain, optimization time 2024-11-01 10:15:00):"
[0471] Topology optimization strategy:
[0472] Relay nodes: R2 (110.07°, 25.06°, 520m, grid 101), R4 (110.03°, 25.02°, 510m, grid 104);
[0473] Link priority: R4→5 > R2→3 > R2→1 > R4→1;
[0474] Antenna parameter configuration:
[0475] R2 Link 1: 50° / 16dB / 32°, Link 2: 90° / 14dB / 145°;
[0476] R4 Link 3: 56° / 17dB / 330°, Link 4: 52° / 17dB / 60°;
[0477] Performance metrics:
[0478] Communication connectivity rate 95% (previously 70%), navigation signal integrity 98% (previously 80%);
[0479] Adaptability Notes: After correcting parameter 103 of the terrain grid, the strategy and parameters are adapted to complex areas with 85% occlusion, and can support dynamic adjustments over the next 10 minutes.
[0480] Another embodiment of the present invention provides a dynamic optimization system for UAV swarm topology in complex terrain, see [link to relevant documentation]. Figure 3 The system may include:
[0481] The acquisition module 301 is used to acquire 3D terrain modeling data and real-time environmental parameters of canyons or urban clusters. It generates a terrain digital twin containing occlusion areas and signal attenuation coefficients through terrain feature extraction algorithms, and integrates the location and communication link status data of the UAV cluster to construct a dynamically updated twin scene model.
[0482] The generation module 302 is used to input the twin scene model into the distributed reinforcement learning model, with the cluster communication connectivity rate and navigation signal integrity as joint optimization objectives, design a reward decay function based on terrain occlusion, and generate a topology optimization strategy for UAV relay node location and link priority allocation through multi-agent asynchronous decision-making.
[0483] Matching module 303 is used to call the beam pattern database of reconfigurable antennas according to the relay node position and link direction in the topology optimization strategy, match the shielding angle in the terrain digital twin, dynamically adjust the antenna beamwidth and gain parameters, and generate a cooperative adaptation scheme of antenna parameters and topology.
[0484] The optimization module 304 is used to collect cluster communication quality and navigation continuity indicators in real time based on the aforementioned collaborative adaptation scheme, feed back the indicator deviations to the terrain digital twin for scene correction, drive the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjust the reconfigurable antenna parameters, and generate a cluster topology dynamic optimization result that adapts to complex terrain changes.
[0485] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.
[0486] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0487] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.
[0488] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.
Claims
1. A method for dynamic topology optimization of UAV swarms in complex terrain, characterized in that, The method includes: Collect 3D terrain modeling data and real-time environmental parameters of canyons or urban clusters, generate a terrain digital twin containing occlusion areas and signal attenuation coefficients through terrain feature extraction algorithms, and construct a dynamically updated twin scene model by integrating the location and communication link status data of drone clusters. The twin scene model is input into a distributed reinforcement learning model. With cluster communication connectivity and navigation signal integrity as joint optimization objectives, a reward decay function based on terrain occlusion is designed. A topology optimization strategy for UAV relay node location and link priority allocation is generated through multi-agent asynchronous decision-making. Based on the relay node locations and link directions in the topology optimization strategy, the beam pattern database of reconfigurable antennas is invoked. The antenna beamwidth and gain parameters are dynamically adjusted to match the shielding angles in the terrain digital twin, generating a collaborative adaptation scheme between antenna parameters and topology. Specifically, the topology optimization strategy parameters are parsed, extracting the identification information, three-dimensional coordinates, and main communication link direction angle of each relay node to clarify the coverage requirements of each link, and forming a strategy parameter table. The beam pattern database is then invoked, and based on the link direction angles in the strategy parameter table, the initial beam parameters of the matching reconfigurable antennas are retrieved from the database. Simultaneously, the antenna parameter adjustment boundaries are extracted to obtain the initial antenna configuration. The process involves: setting up a set of antenna parameters; dynamically matching and adjusting shielding angle parameters by querying the maximum terrain shielding angle on each communication link path from the terrain digital twin; classifying and adjusting antenna parameters according to the shielding angle size: compressing the beamwidth and increasing the gain when the shielding angle is within a preset small range; maintaining parameter balance when the shielding angle is within a preset medium range; and expanding the beamwidth and reducing the gain when the shielding angle is within a preset large range, thus obtaining the adjusted antenna parameters; collaboratively verifying adaptability by checking the matching between the adjusted antenna parameters and the relay node positions, and correcting any conflicting parameters; and associating the relay node and link information in the topology optimization strategy with the corresponding antenna parameters by UAV ID to generate a collaborative adaptation scheme for antenna parameters and topology. Based on the aforementioned collaborative adaptation scheme, cluster communication quality and navigation continuity indicators are collected in real time. The indicator deviations are fed back to the terrain digital twin for scene correction, driving the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjusting the reconfigurable antenna parameters, and generating dynamic optimization results of cluster topology adapted to complex terrain changes.
2. The method according to claim 1, characterized in that, The process involves collecting 3D terrain modeling data and real-time environmental parameters from canyons or urban clusters, generating a terrain digital twin including occlusion areas and signal attenuation coefficients using terrain feature extraction algorithms, and then integrating the location and communication link status data of the drone swarm to construct a dynamically updated twin scene model, including: The raw data was collected by category. For canyon terrain, 3D modeling data of elevation difference, rock wall inclination angle and valley width were collected. For urban agglomeration, 3D modeling data of building height, building spacing and street direction were collected. At the same time, real-time environmental parameters were collected, including meteorological parameters such as wind speed, wind direction and precipitation intensity, and electromagnetic parameters such as interference frequency and signal interference intensity. The data were organized according to terrain type and parameter dimensions to form the raw multidimensional dataset. The raw multidimensional data is preprocessed, and the terrain 3D modeling data is divided into grids. The grid cell size is standardized and filled with elevation values. The spatiotemporal smoothing algorithm is used to eliminate instantaneous disturbances in real-time environmental parameters, and Kalman filtering is used to correct outliers that exceed the physical reasonable range, resulting in a standardized data grid. Topographic feature parameters are extracted, and a deep learning semantic segmentation algorithm is used to analyze the standardized data grid, identify and mark the shading areas in the terrain, and calculate the signal attenuation coefficient of each grid cell based on the electromagnetic wave propagation model and environmental electromagnetic parameters to generate a terrain feature map. The navigation position data and communication link status data of the drone swarm acquired in real time are mapped to the corresponding grid cells of the terrain feature map according to the timestamp, and the data is updated at fixed time intervals to construct a twin scene model that includes static terrain features and dynamic swarm status.
3. The method according to claim 2, characterized in that, The process involves inputting the twin scene model into a distributed reinforcement learning model, using cluster communication connectivity and navigation signal integrity as joint optimization objectives, designing a reward decay function based on terrain occlusion, and generating a topology optimization strategy for UAV relay node location selection and link priority allocation through asynchronous decision-making by multiple agents, including: Key state parameters, including terrain occlusion, relative position of UAV cluster, communication link connectivity, and navigation signal strength, are extracted from the twin scene model. A multi-dimensional state space matrix is constructed according to the UAV number and terrain grid index, and the matrix elements quantify the corresponding state parameter values. A joint optimization objective is set, with cluster communication connectivity and navigation signal integrity as the core optimization indicators. Based on the weights of their impact on cluster topology, a comprehensive optimization indicator is calculated using weighted averages, thus clarifying the direction for indicator improvement. The design incorporates a reward decay function, where the base reward value is positively correlated with the comprehensive optimization index; a terrain occlusion correction factor is introduced, where the reward value decreases linearly with occlusion when the occlusion of the area where the UAV is located exceeds a set value; and a fixed penalty value is set if a communication link is interrupted or the navigation and positioning are out of tolerance, thus forming a dynamic reward mechanism. The system performs asynchronous decision-making by multi-agents, splitting the multidimensional state space matrix into local state segments for each drone. Each drone acts as an independent agent, calculating decision priorities based on its local segment. The global coordination layer of the distributed reinforcement learning model integrates the decisions of each agent, prioritizing low-occlusion areas as relay nodes, and allocating link priorities in reverse order of signal attenuation coefficients. This generates a topology optimization strategy that includes the three-dimensional coordinates of relay nodes and the link priority ranking.
4. The method according to claim 3, characterized in that, Based on the aforementioned collaborative adaptation scheme, real-time collection of cluster communication quality and navigation continuity indicators is performed. Indicator deviations are fed back to the terrain digital twin for scene correction, driving the distributed reinforcement learning system to iteratively update the topology optimization strategy. Simultaneously, reconfigurable antenna parameters are adjusted to generate dynamic topology optimization results adapted to complex terrain changes, including: Load the collaborative adaptation scheme, extract the antenna parameters and topology strategy in the scheme as the benchmark configuration, clarify the expected communication coverage and navigation and positioning accuracy requirements of each UAV, and form a benchmark parameter table for the scheme; Based on the baseline parameter table of the scheme, the performance indicators are collected, and the communication quality indicators and navigation continuity indicators of the UAV cluster are collected at preset time intervals. The actual collected values are compared with the expected values in the baseline parameter table of the scheme, and the deviation rate of each indicator is calculated to form an indicator deviation set. Feedback indicator deviation correction twin, filter out anomalies where the indicator deviation concentration exceeds the fault tolerance range of the scheme, associate the anomaly deviation with the corresponding grid cell of the terrain digital twin, update the shading area boundary and signal attenuation coefficient of the cell, and generate the corrected terrain digital twin. The iterative optimization strategy and parameters are input into the distributed reinforcement learning model, and the topology optimization strategy is adjusted in combination with parameter conflicts in the collaborative adaptation scheme. At the same time, the antenna parameters are corrected according to the new topology strategy, and the adjusted indicators are collected to verify whether the deviation rate meets the scheme requirements. After the target is met, the optimized topology strategy and antenna parameters are integrated to generate a cluster topology dynamic optimization result adapted to complex terrain changes.
5. A dynamic optimization system for UAV swarm topology in complex terrain, characterized in that, The system includes: The data acquisition module is used to collect 3D terrain modeling data and real-time environmental parameters of canyons or urban clusters. It generates a terrain digital twin containing occlusion areas and signal attenuation coefficients through terrain feature extraction algorithms. It integrates the location and communication link status data of the drone swarm to build a dynamically updated twin scene model. The generation module is used to input the twin scene model into the distributed reinforcement learning model, with cluster communication connectivity and navigation signal integrity as joint optimization objectives, design a reward decay function based on terrain occlusion, and generate a topology optimization strategy for UAV relay node location and link priority allocation through multi-agent asynchronous decision-making. The matching module is used to dynamically adjust the antenna beamwidth and gain parameters based on the relay node positions and link directions in the topology optimization strategy, by calling the beam pattern database of the reconfigurable antenna and matching the shielding angle in the terrain digital twin, thus generating a cooperative adaptation scheme for antenna parameters and topology. Specifically, it parses the topology optimization strategy parameters, extracts the identification information, three-dimensional coordinates, and main communication link direction angle of each relay node from the topology optimization strategy, clarifies the coverage requirements of each link, and organizes them into a strategy parameter table. It then calls the beam pattern database, retrieves the initial beam parameters of the reconfigurable antenna matched in the database based on the link direction angle in the strategy parameter table, and simultaneously extracts the antenna parameter adjustment boundary to obtain the initial... Antenna configuration set; dynamically matching shielding angle adjustment parameters, querying the maximum terrain shielding angle on each communication link path from the terrain digital twin; classifying and adjusting antenna parameters according to the shielding angle size: compressing beamwidth and increasing gain when the shielding angle is within a preset small range, maintaining parameter balance when the shielding angle is within a preset medium range, and expanding beamwidth and reducing gain when the shielding angle is within a preset large range, to obtain the adjusted antenna parameters; collaboratively verifying adaptability, checking the matching of the adjusted antenna parameters with the relay node location, and correcting conflicting parameters; associating relay node and link information in the topology optimization strategy with the corresponding antenna parameters by UAV ID to generate a collaborative adaptation scheme for antenna parameters and topology; The optimization module is used to collect cluster communication quality and navigation continuity indicators in real time based on the aforementioned collaborative adaptation scheme, feed back the indicator deviations to the terrain digital twin for scene correction, drive the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjust the reconfigurable antenna parameters, and generate dynamic optimization results of cluster topology that adapt to complex terrain changes.
6. The system according to claim 5, characterized in that, The acquisition module is specifically used for: The raw data was collected by category. For canyon terrain, 3D modeling data of elevation difference, rock wall inclination angle and valley width were collected. For urban agglomeration, 3D modeling data of building height, building spacing and street direction were collected. At the same time, real-time environmental parameters were collected, including meteorological parameters such as wind speed, wind direction and precipitation intensity, and electromagnetic parameters such as interference frequency and signal interference intensity. The data were organized according to terrain type and parameter dimensions to form the raw multidimensional dataset. The raw multidimensional data is preprocessed, and the terrain 3D modeling data is divided into grids. The grid cell size is standardized and filled with elevation values. The spatiotemporal smoothing algorithm is used to eliminate instantaneous disturbances in real-time environmental parameters, and Kalman filtering is used to correct outliers that exceed the physical reasonable range, resulting in a standardized data grid. Topographic feature parameters are extracted, and a deep learning semantic segmentation algorithm is used to analyze the standardized data grid, identify and mark the shading areas in the terrain, and calculate the signal attenuation coefficient of each grid cell based on the electromagnetic wave propagation model and environmental electromagnetic parameters to generate a terrain feature map. The navigation position data and communication link status data of the drone swarm acquired in real time are mapped to the corresponding grid cells of the terrain feature map according to the timestamp, and the data is updated at fixed time intervals to construct a twin scene model that includes static terrain features and dynamic swarm status.
7. The system according to claim 6, characterized in that, The generation module is specifically used for: Key state parameters, including terrain occlusion, relative position of UAV cluster, communication link connectivity, and navigation signal strength, are extracted from the twin scene model. A multi-dimensional state space matrix is constructed according to the UAV number and terrain grid index, and the matrix elements quantify the corresponding state parameter values. A joint optimization objective is set, with cluster communication connectivity and navigation signal integrity as the core optimization indicators. Based on the weights of their impact on cluster topology, a comprehensive optimization indicator is calculated using weighted averages, thus clarifying the direction for indicator improvement. The design incorporates a reward decay function, where the base reward value is positively correlated with the comprehensive optimization index; a terrain occlusion correction factor is introduced, where the reward value decreases linearly with occlusion when the occlusion of the area where the UAV is located exceeds a set value; and a fixed penalty value is set if a communication link is interrupted or the navigation and positioning are out of tolerance, thus forming a dynamic reward mechanism. The system performs asynchronous decision-making by multi-agents, splitting the multidimensional state space matrix into local state segments for each drone. Each drone acts as an independent agent, calculating decision priorities based on its local segment. The global coordination layer of the distributed reinforcement learning model integrates the decisions of each agent, prioritizing low-occlusion areas as relay nodes, and allocating link priorities in reverse order of signal attenuation coefficients. This generates a topology optimization strategy that includes the three-dimensional coordinates of relay nodes and the link priority ranking.
8. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-4 when it is run.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-4.
Citation Information
Patent Citations
Unmanned aerial vehicle heterogeneous cluster collaborative navigation method and system facing complex terrain
CN120445223A
Digital model based reverse osmosis plant operation and optimization
US20230150835A1