Unmanned aerial vehicle cluster topology dynamic optimization method and system under complex terrain
By constructing a dynamic twin scene model and optimizing the topology of the UAV swarm using distributed reinforcement learning, the problem of insufficient communication and navigation performance in complex terrain was solved, and stable mission execution of the UAV swarm in complex terrain was achieved.
Patent Information
- Application Number
- CN202511316860.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-16
AI Technical Summary
In complex terrain, the communication and navigation performance of UAV swarms is affected by terrain obstruction and signal attenuation. Traditional topologies are difficult to adapt to dynamic environments, leading to communication link interruptions and reduced navigation accuracy, which cannot meet mission reliability requirements.
By collecting 3D terrain modeling data and real-time environmental parameters, a dynamic twin scene model is constructed. Combined with distributed reinforcement learning and reconfigurable antennas, relay node location and link priority are optimized, and antenna beam parameters are dynamically adjusted to achieve collaborative adaptation of cluster topology.
It has achieved dynamic improvement in the communication and navigation performance of drone clusters in complex terrain, ensuring the stable execution of tasks.
Smart Images

Figure CN120805745A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of unmanned aerial vehicles, and particularly relates to a method and system for dynamically optimizing the topology of an unmanned aerial vehicle cluster in complex terrain. BACKGROUND
[0002] In complex terrain such as canyons and urban clusters, an unmanned aerial vehicle cluster is easily affected by terrain shielding and signal attenuation, and a traditional fixed topology structure is difficult to adapt to a dynamic environment. For example, shielding areas can cause communication links to be interrupted, signal attenuation can reduce navigation accuracy, and thus cause the cluster to fail to cooperate. Existing topology optimization methods are mostly based on idealized terrain assumptions, and do not construct a dynamic correlation model of terrain and cluster state. They only optimize through static relay node deployment or fixed link priority allocation, and thus cannot respond to terrain shielding changes and signal fluctuations in real time. At the same time, they lack a coordinated adaptation design of antenna parameters and topology structure, and the beam adjustment of a reconfigurable antenna is not combined with dynamic adjustment of terrain shielding angles, further exacerbating the communication and navigation stability problems, and thus it is difficult to meet the reliability requirements of unmanned aerial vehicle cluster reconnaissance, rescue and other tasks in complex terrain. SUMMARY
[0003] The purpose of the present application is to provide a method and system for dynamically optimizing the topology of an unmanned aerial vehicle cluster in complex terrain, so as to solve the problems in the prior art and dynamically improve the communication and navigation performance of the cluster in complex terrain, and ensure stable execution of the cluster task.
[0004] One embodiment of the present application provides a method for dynamically optimizing the topology of an unmanned aerial vehicle cluster in complex terrain, the method comprising: Collecting terrain three-dimensional modeling data and real-time environment parameters of a canyon or an urban cluster, generating a terrain digital twin containing shielding areas and signal attenuation coefficients through a terrain feature extraction algorithm, and fusing position and communication link state data of the unmanned aerial vehicle cluster to construct a dynamically updated twin scene model; Inputting the twin scene model into a distributed reinforcement learning model, taking the cluster communication connectivity rate and navigation signal integrity as joint optimization objectives, designing a reward decay function based on terrain shielding degree, and generating a topology optimization strategy for unmanned aerial vehicle relay node site selection and link priority allocation through multi-agent asynchronous decision-making; According to the relay node position and link direction in the topology optimization strategy, calling a beam pattern database of a reconfigurable antenna, dynamically adjusting antenna beam width and gain parameters according to the shielding angles in the terrain digital twin, and generating a coordinated adaptation scheme of antenna parameters and topology structure; Based on the cooperative adaptation scheme, the cluster communication quality and navigation continuity indicators are collected in real time, the indicator deviation is fed back to the terrain digital twin for scene correction, the distributed reinforcement learning system is driven to iteratively update the topology optimization strategy, the reconfigurable antenna parameters are synchronously adjusted, and the cluster topology dynamic optimization result adapted to complex terrain changes is generated.
[0005] Optionally, the terrain three-dimensional modeling data and real-time environmental parameters of the canyon or urban cluster are collected, a terrain digital twin containing shielding areas and signal attenuation coefficients is generated through terrain feature extraction algorithm, a dynamically updated twin scene model is constructed by fusing the position and communication link state data of the UAV cluster, including: The original data is classified and collected, the three-dimensional modeling data of altitude difference, rock wall inclination and valley bottom width is collected for canyon terrain, the three-dimensional modeling data of building height, building spacing and street direction is collected for urban cluster, and real-time environmental parameters are collected, including wind speed, wind direction and precipitation intensity, electromagnetic parameters including interference frequency and signal interference strength, and the original multi-dimensional data set is formed by organizing the terrain type and parameter dimension; The original multi-dimensional data is preprocessed, the terrain three-dimensional modeling data is grid divided, the grid unit size is unified and the elevation value is filled; the real-time environmental parameters are smoothed by space-time smoothing algorithm to eliminate transient disturbance, and the abnormal values exceeding the physical reasonable range are corrected by Kalman filter to obtain standardized data grid; The terrain feature parameters are extracted, the standardized data grid is analyzed by using deep learning semantic segmentation algorithm, the shielding areas in the terrain are identified and labeled, the signal attenuation coefficient of each grid unit is calculated based on the electromagnetic wave propagation model combined with the environmental electromagnetic parameters, and the terrain feature map is generated; The real-time acquired navigation position data and communication link state data of the UAV cluster are mapped to the corresponding grid unit of the terrain feature map according to the time stamp, and the data is updated at fixed time intervals to construct the twin scene model containing static terrain features and dynamic cluster state.
[0006] Optionally, the twin scene model is input into the distributed reinforcement learning model, the cluster communication connectivity rate and navigation signal integrity are taken as the joint optimization target, the reward decay function based on terrain shielding degree is designed, the topology optimization strategy of UAV relay node site selection and link priority allocation is generated through multi-agent asynchronous decision making, including: The key state parameters including terrain shielding degree, relative position of UAV cluster, communication link connectivity state and navigation signal strength are extracted from the twin scene model, a multi-dimensional state space matrix is constructed according to the UAV number and terrain grid index, and the matrix element quantifies the corresponding state parameter value; Set joint optimization goal, set cluster communication connectivity and navigation signal integrity as core optimization indicators, and calculate comprehensive optimization indicators by weighting according to the influence weight of the two on cluster topology, and determine the index improvement direction; Design reward decay function, the basic reward value is positively correlated with the comprehensive optimization indicator; introduce the terrain shielding degree correction factor, when the shielding degree of the area where the unmanned aerial vehicle is located exceeds the set value, the reward value is linearly attenuated according to the shielding degree; if the communication link is interrupted or the navigation positioning is out of tolerance, a fixed penalty value is set, forming a dynamic reward mechanism; Execute multi-agent asynchronous decision, split the multi-dimensional state space matrix into local state fragments according to the unmanned aerial vehicle, and calculate the decision priority based on the local fragment for each unmanned aerial vehicle as an independent agent; the global coordination layer of the distributed reinforcement learning model integrates the decisions of each agent, preferentially selects a low-shielding area as a relay node, and assigns link priority according to the reverse order of the signal attenuation coefficient to generate a topology optimization strategy containing the three-dimensional coordinates of the relay node and the link priority order.
[0007] Optionally, according to the position of the relay node and the link direction in the topology optimization strategy, the beam pattern database of the reconfigurable antenna is called to dynamically adjust the antenna beam width and gain parameters by matching the shielding angle in the terrain digital twin, and a collaborative adaptation scheme of antenna parameters-topology structure is generated, including: Analyze the topology optimization strategy parameters, extract the identification information, three-dimensional coordinates and main communication link direction angle of each relay node from the topology optimization strategy, clearly define the coverage range requirement of each link, and form a strategy parameter table; Call the beam pattern database, retrieve the matching initial beam parameters of the reconfigurable antenna in the database according to the link direction angle in the strategy parameter table, and extract the antenna parameter adjustment boundary to obtain an initial antenna configuration set; Dynamically match the shielding angle adjustment parameters, query the maximum terrain shielding angle on each communication link path from the terrain digital twin; adjust the antenna parameters according to the shielding angle size, compress the beam width and increase the gain when the shielding angle is in a preset smaller range, keep the parameters balanced when the shielding angle is in a preset medium range, and expand the beam width and reduce the gain when the shielding angle is in a preset larger range, to obtain the adjusted antenna parameters; Collaboratively verify the adaptability, check the matching of the adjusted antenna parameters and the relay node position, and modify the parameters that exist conflicts; associate the relay node, link information in the topology optimization strategy and the corresponding antenna parameters according to the unmanned aerial vehicle ID to generate a collaborative adaptation scheme of antenna parameters-topology structure.
[0008] Optionally, based on the cooperative adaptation scheme, real-time collection of cluster communication quality and navigation continuity indicators, feedback of indicator deviation to the terrain digital twin for scene correction, driving the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronous adjustment of the reconfigurable antenna parameters, generation of cluster topology dynamic optimization results adapted to complex terrain changes, including: Load the cooperative adaptation scheme, extract the antenna parameters and topology strategy in the scheme as the reference configuration, clearly define the expected communication coverage and navigation positioning accuracy requirements of each UAV, and form a scheme reference parameter table; Collect performance indicators based on the scheme reference parameter table, collect communication quality indicators and navigation continuity indicators of the UAV cluster at preset time intervals; compare the actual collection value with the expected value in the scheme reference parameter table, calculate the deviation rate of each indicator, and form an indicator deviation set; Feedback the index deviation correction twin, filter the abnormal items in the index deviation set that exceed the scheme fault tolerance range, associate the abnormal deviation to the corresponding grid unit of the terrain digital twin, update the shielding area boundary and signal attenuation coefficient of the unit, and generate a corrected terrain digital twin; Iterative optimization strategy and parameters, input the corrected terrain digital twin into the distributed reinforcement learning model, adjust the topology optimization strategy in combination with the parameter conflicts in the cooperative adaptation scheme; simultaneously correct the antenna parameters according to the new topology strategy, collect the adjusted indicators to verify whether the deviation rate meets the scheme requirements; after meeting the requirements, integrate the optimized topology strategy and antenna parameters to generate the cluster topology dynamic optimization results adapted to complex terrain changes.
[0009] Another embodiment of the present application provides a UAV cluster topology dynamic optimization system in complex terrain, the system comprising: The collection module is used to collect terrain three-dimensional modeling data and real-time environmental parameters of the canyon or urban group, generate a terrain digital twin containing shielding areas and signal attenuation coefficients through terrain feature extraction algorithms, and construct a dynamically updated twin scene model by fusing the position and communication link state data of the UAV cluster; The generation module is used to input the twin scene model into a distributed reinforcement learning model, take the cluster communication connectivity rate and navigation signal integrity as the joint optimization target, design a reward decay function based on the terrain shielding degree, and generate a topology optimization strategy of UAV relay node site selection and link priority allocation through multi-agent asynchronous decision-making; The matching module is used to match the shielding angle in the terrain digital twin according to the relay node position and link direction in the topology optimization strategy, dynamically adjust the antenna beam width and gain parameters by calling the beam pattern database of the reconfigurable antenna, and generate a cooperative adaptation scheme of antenna parameters-topology structure; An optimization module is configured to collect cluster communication quality and navigation continuity indicators in real time based on the collaborative adaptation scheme, feed back indicator deviations to the terrain digital twin for scene correction, drive the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjust the reconfigurable antenna parameters, and generate the cluster topology dynamic optimization result adapted to complex terrain changes.
[0010] A further embodiment of the present application provides a storage medium having a computer program stored therein, wherein the computer program is configured to execute the method described in any of the above embodiments when executed.
[0011] A further embodiment of the present application provides an electronic device comprising a memory having a computer program stored therein and a processor configured to execute the computer program to execute the method described in any of the above embodiments.
[0012] Compared with the prior art, the unmanned aerial vehicle cluster topology dynamic optimization method provided by the present application collects terrain three-dimensional modeling data and real-time environmental parameters of a canyon or a city cluster, fuses position and communication link state data of the unmanned aerial vehicle cluster to construct a dynamically updated twin scene model, inputs the twin scene model into a distributed reinforcement learning model, generates a topology optimization strategy of unmanned aerial vehicle relay node site selection and link priority allocation through multi-agent asynchronous decision-making, generates a collaborative adaptation scheme of antenna parameters-topology structure according to the relay node position and link direction in the topology optimization strategy, collects cluster communication quality and navigation continuity indicators in real time based on the collaborative adaptation scheme, and generates a cluster topology dynamic optimization result adapted to complex terrain changes, thereby achieving dynamic improvement of cluster communication and navigation performance in complex terrain and ensuring stable execution of cluster tasks. BRIEF DESCRIPTION OF DRAWINGS
[0013] Figure 1 A hardware structure block diagram of a computer terminal of the unmanned aerial vehicle cluster topology dynamic optimization method in complex terrain provided by the embodiment of the present application is shown in the figure. Figure 2 A flowchart of the unmanned aerial vehicle cluster topology dynamic optimization method in complex terrain provided by the embodiment of the present application is shown in the figure. Figure 3 A structure diagram of the unmanned aerial vehicle cluster topology dynamic optimization system in complex terrain provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0014] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be explained as a limitation of the present application.
[0015] The embodiment of the present application firstly provides a UAV cluster topology dynamic optimization method under complex terrain, which can be applied to an electronic device, such as a computer terminal, specifically, a general computer and the like.
[0016] The UAV cluster topology dynamic optimization method under complex terrain provided by the embodiment of the present application is described in detail below by taking a computer terminal as an example. Figure 1 A hardware structure block diagram of a computer terminal of the UAV cluster topology dynamic optimization method under complex terrain provided by the embodiment of the present application is shown in FIG. 1. Figure 1 As shown in the figure, the computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory can include a non-volatile storage medium and an internal memory.
[0017] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can make the processor execute any UAV cluster topology dynamic optimization method under complex terrain.
[0018] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0019] The internal memory provides an environment for the execution of the computer program in the non-volatile storage medium, which, when executed by the processor, can make the processor execute any UAV cluster topology dynamic optimization method under complex terrain.
[0020] The network interface is used for network communication, such as sending assigned tasks and the like. Those skilled in the art can understand that Figure 1 The structure shown in FIG. 1 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. Specifically, the computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0021] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0022] Referring toFigure 2 Embodiments of the present application provide a method for dynamically optimizing the topology of a UAV cluster in complex terrain, which can include the following steps: S201, collect terrain three-dimensional modeling data and real-time environmental parameters of a canyon or urban group, generate a terrain digital twin containing shadow areas and signal attenuation coefficients through terrain feature extraction algorithms, and fuse the position and communication link state data of the UAV cluster to construct a dynamically updated twin scene model; Specifically, the original data can be classified and collected. For canyon terrain, three-dimensional modeling data of elevation difference, rock wall inclination, and valley bottom width are collected. For urban groups, three-dimensional modeling data of building height, building spacing, and street orientation are collected. At the same time, real-time environmental parameters are collected, including wind speed, wind direction, and precipitation intensity for meteorological parameters, and interference frequency and signal interference strength for electromagnetic parameters. The original multi-dimensional data set is formed by organizing the data according to terrain type and parameter dimension. Original data collection is the basis for building a terrain digital twin. It needs to follow the principle of "differential collection according to terrain type + comprehensive coverage of environmental parameters" to ensure that the data accurately reflects the physical characteristics of complex terrain and the real-time environmental impact, solving the problem of "data distortion due to single data". This step needs to clarify "what to collect, how to collect, and how to organize" to provide structured input for subsequent preprocessing.
[0023] I. Classification of terrain three-dimensional modeling data: For the two typical complex terrains of canyons and urban groups, the core parameters that can represent the terrain shielding and signal propagation characteristics are collected. Each parameter has a clear collection device, precision range, and sampling density to ensure that the data is quantifiable and reusable: Canyon terrain three-dimensional modeling data: Elevation difference: reflects the terrain undulation in the vertical direction of the canyon (directly affects signal shielding), uses a UAV LiDAR (LiDAR, model Velodyne VLP-16, ranging accuracy ±2cm, sampling density 10 points / ㎡) to fly along the canyon axis, for example, a canyon with a valley bottom elevation of 1000m and two sides of rock wall with a maximum elevation of 1500m, the elevation difference is recorded as 500m (maximum value) and 300m (average value); Rock wall inclination: reflects the steepness of the rock wall (the larger the inclination, the stronger the signal reflection and the more serious the shielding), the rock wall plane is fitted through LiDAR point cloud data, and the angle between the plane and the horizontal plane is calculated (accuracy ±0.5°), for example, the west rock wall inclination is 75° and the east rock wall inclination is 60°; Valley width: reflects the horizontal spatial scale of the canyon (the smaller the width, the more obvious the signal multipath effect), RTK-GPS (accuracy ±1cm) is used to collect the coordinates of the boundaries of the valley floor on both sides, the horizontal distance is calculated, for example, the width of the valley floor of a certain canyon is between 50m-100m, recorded as 50m, 55m, 60m…100m at intervals of 5m.
[0024] Urban terrain 3D modeling data: Building height: reflects the height of the vertical shielding source in the city (tall buildings are easy to block low-altitude UAV signals), a tilt camera (model DJI P1, 50 million pixels, modeling accuracy ±5cm) is used to shoot the city area, and building height data is generated through 3D modeling software (such as ContextCapture), for example, the building height in a certain business district is 30m-100m, recorded as 30m (low-rise), 60m (mid-rise), and 100m (high-rise); Building spacing: reflects the density of buildings in the horizontal direction (the smaller the spacing, the more difficult it is to diffract signals), RTK-GPS is used to collect the coordinates of the outer walls of adjacent buildings, and the horizontal distance is calculated, for example, the building spacing in a certain residential area is 20m-50m, recorded as 20m (dense area) and 50m (open area); Street orientation: reflects the main channel direction of signal propagation in the city (signal attenuation is smaller along the street orientation), the azimuth of the street (accuracy ±1°) is recorded through GIS maps and field measurements, for example, the orientation of a certain main road is east-west (azimuth 90°), and the orientation of a secondary road is north-south (azimuth 0°).
[0025] II. Collection of real-time environmental parameters: Environmental parameters directly affect the quality of UAV communication and navigation signals, and two key parameters, "weather + electromagnetic", need to be collected to ensure that external interference factors affecting signal propagation are covered: Weather parameters: Wind speed: affects the stability of the UAV (indirectly affects the stability of the communication link), a small weather station (model Davis Vantage Pro2, measurement range 0-60m / s, accuracy ±0.1m / s) is installed at the top of the terrain, with a sampling frequency of 1Hz, for example, the collected value is 3m / s (light wind) and 8m / s (strong wind); Wind direction: affects the signal propagation path (signal attenuation is smaller along the wind direction), collected by the wind direction sensor of the weather station (measurement range 0-360°, accuracy ±5°), recorded as azimuth, for example, northeast by east 30° (wind direction angle 60°); Precipitation intensity: Influences electromagnetic wave attenuation (rainwater increases signal absorption), collected with a tipping bucket rain gauge (resolution 0.1mm, accuracy ±0.5mm), recorded as intensity classification 0.5mm / h (light rain), 10mm / h (heavy rain).
[0026] Electromagnetic parameters: Interference frequency: Identifies electromagnetic interference sources (such as WiFi, Bluetooth, industrial equipment) that affect UAV communication, collected with a spectrum analyzer (model Keysight N9918A, frequency range 9kHz-18GHz, resolution 1Hz) at key terrain points (such as valley bottoms, city squares), for example, 2.4GHz (WiFi interference), 5.8GHz (UAV image transmission interference) is detected; Signal interference intensity: Quantifies the degree of influence of interference sources on UAV signals, collects the power of interference signals with a signal strength meter (measurement range -120dBm to -30dBm, accuracy ±1dBm), for example, 2.4GHz interference intensity is -80dBm (weak interference), 5.8GHz interference intensity is -60dBm (medium interference).
[0027] Three, arrangement of original multidimensional data sets: Arrange data according to the hierarchical structure of "terrain type - parameter dimension" to avoid confusion, for example: "Original multidimensional data set (terrain type: valley): Three-dimensional modeling data: elevation difference {maximum value 500m, average value 300m}, rock wall inclination {west side 75°, east side 60°}, valley bottom width {50m, 55m,..., 100m}; Real-time environmental parameters: weather {wind speed 3m / s, wind direction 60°, precipitation intensity 0.5mm / h}, electromagnetic {interference frequency 2.4GHz / 5.8GHz, interference intensity -80dBm / -60dBm}; Original multidimensional data set (terrain type: urban group): Three-dimensional modeling data: building height {30m, 60m, 100m}, building spacing {20m, 50m}, street orientation {90° (east-west), 0° (north-south)}; Real-time environmental parameters: weather {wind speed 5m / s, wind direction 180° (south wind), precipitation intensity 0mm / h}, electromagnetic {interference frequency 2.4GHz, interference intensity -70dBm}. Each data set is labeled with collection time (accurate to the minute), collection device number, ensuring data traceability and providing a basis for subsequent preprocessing of abnormal values.
[0028] Preprocess the original multidimensional data, grid the terrain three-dimensional modeling data, unify the grid cell size and fill in the elevation value; adopt the space-time smoothing algorithm to eliminate transient disturbance for real-time environmental parameters, and correct the abnormal values beyond the physical reasonable range through Kalman filtering to obtain standardized data grid; The original data may have problems such as "discrete terrain data and environmental parameter fluctuation", which need to be converted into unified and stable standardized data through preprocessing to provide reliable input for terrain feature extraction and solve the problem of "low data quality leading to feature extraction deviation". This step needs to be executed in two steps: "terrain data gridding + environmental parameter smoothing correction" to ensure that the data format is unified and the value is reasonable.
[0029] I. Grid division of terrain three-dimensional modeling data: Grid is the core means of converting discrete terrain data into continuous grid, which needs to unify the grid size and coordinate system to ensure that data of different terrain types can be compared and integrated: Grid cell size setting: based on the signal propagation characteristics (the wavelength of unmanned aerial vehicle communication is usually 10cm-1m, the grid size needs to be less than the wavelength to capture details), the grid cell size is set to 5m×5m (considering precision and computational efficiency), the canyon terrain extends the grid along the axis direction, and the urban cluster terrain covers the grid according to the rectangular area; Coordinate system: adopt the three-dimensional coordinate system combined with WGS84 geodetic coordinate system (longitude and latitude) and altitude, for example, the left lower corner coordinate of a canyon grid is (longitude 110.0°, latitude 25.0°, altitude 1000m), and the right upper corner coordinate is (longitude 110.1°, latitude 25.1°, altitude 1500m), which divides 20×20=400 grid cells; Elevation value filling: for each grid cell, extract the maximum, minimum and average values of elevation from the original LiDAR or oblique photography data, and fill them into the grid attributes, for example, the elevation value of a canyon grid cell (110.02°, 25.03°) is {max:1200m, min:1050m, avg:1125m}, and the elevation value of an urban cluster grid cell (120.05°, 30.02°) is {max:80m (building top), min:0m (ground), avg:40m}.
[0030] After gridding, the grid integrity needs to be verified to ensure that there is no data missing (missing grid uses adjacent grid interpolation to supplement), for example, a grid at the edge of the canyon has no LiDAR data, and the average elevation of the east and south adjacent grids is 1100m.
[0031] II. Space-time smoothing and abnormal value correction of real-time environmental parameters: The environmental parameters are susceptible to instantaneous disturbances (such as gusts, sudden electromagnetic interference), and need to be processed by algorithms to ensure data stability: Temporal and spatial smoothing algorithm: Adopt sliding average algorithm to eliminate instantaneous disturbances in time dimension, window size is set to 5 sampling points (10 seconds per sampling point, window 50 seconds), formula is: smoothed value = (x1+x2+x3+x4+x5) / 5, for example, the original wind speed data is 3m / s, 8m / s (gust), 4m / s, 3m / s, 5m / s, smoothed value = (3+8+4+3+5) / 5=4.6m / s, eliminating the influence of gust; Spatial dimension adopts Kriging interpolation method to complete the space of sparse collected electromagnetic parameters (such as only 3 points collected), to ensure that each grid cell has environmental parameter value; Kalman filter correction of abnormal value: For data beyond the physical reasonable range (such as sudden 100mm / h of precipitation intensity, far beyond the local climate level), Kalman filter is used for correction. Filter parameters are set: state variable is environmental parameter value, process noise covariance Q=0.01 (reflecting the natural fluctuation of parameters), measurement noise covariance R=0.1 (reflecting the sensor error), for example, the original precipitation intensity data is 0.5mm / h, 100mm / h (abnormal), 0.6mm / h, the abnormal value is corrected to 0.55mm / h after filtering, which is consistent with the physical reasonable range (the maximum precipitation intensity in the local area is 50mm / h).
[0032] The standardized data grid obtained after preprocessing contains "terrain elevation, environmental parameters" two types of attributes, the format is uniform and the value is stable, which lays a foundation for subsequent terrain feature extraction.
[0033] Extracting terrain feature parameters, using deep learning semantic segmentation algorithm to analyze the standardized data grid, identifying and labeling the shielding area in the terrain, calculating the signal attenuation coefficient of each grid cell based on the electromagnetic wave propagation model combined with environmental electromagnetic parameters, generating terrain feature atlas; Terrain feature extraction is the core link of transforming "original data" into "available features of twin body", which needs to identify shielding area through semantic segmentation and calculate attenuation coefficient through propagation model, solving the problem of "unable to quantify the influence of terrain on signal". This step needs to clarify "how to identify shielding and how to calculate attenuation", generating feature atlas that can directly support topology optimization.
[0034] I. Shielding area identification based on deep learning semantic segmentation: U-Net semantic segmentation model (suitable for segmentation tasks of structured data such as terrain) is adopted to divide the terrain area in the standardized data grid into "shielding area" and "non-shielding area", the specific process is: Data preprocessing: convert the elevation values of the normalized data grid into a grayscale image (the higher the elevation, the larger the grayscale value), for example, an elevation of 1000m corresponds to a grayscale of 0, and an elevation of 1500m corresponds to a grayscale of 255, forming an input image of 256x256 pixels; Model training and inference: train the U-Net model using 10,000 complex terrain labeled images (manually labeled occlusion areas such as canyon rock walls, city high-rise shadows), training parameters: batch size 16, number of iterations 50, learning rate 0.001, loss function using cross-entropy. Input the preprocessed image into the trained model, output the binary segmentation result (white for occlusion area, black for non-occlusion area), for example, the canyon west rock wall grid and the city high-rise grid below are marked as occlusion areas; Occlusion area boundary optimization: use morphological closing operation (first dilation then erosion) to eliminate small holes in the segmentation result (such as small white spaces caused by high-rise windows), the structure element is set to 3x3 pixels, ensuring that the occlusion area boundary is continuous and accurate.
[0035] After recognition, the segmentation accuracy needs to be verified, through on-site comparison, the occlusion area recognition accuracy needs to be ≥95%, for example, the overlap between the model labeled occlusion area and the actual high-rise shadow in a certain city area is 97%, which is determined to be qualified.
[0036] II. Signal attenuation coefficient calculation based on electromagnetic wave propagation model: Using the free space propagation model combined with environmental interference correction, calculate the signal attenuation coefficient (unit: dB / m) of each grid cell, quantify the signal propagation loss in that grid, specific steps: Basic attenuation coefficient calculation (free space model): formula is L0=32.45+20log (d)+20log (f), where d is the signal propagation distance (km), f is the signal frequency (MHz). For example, the unmanned aerial vehicle communication frequency f=2400MHz, the signal propagation distance d=0.5km in a certain grid cell, L0=32.45+20log0.5+20log2400≈32.45-6.02+67.6≈94.03dB, the basic attenuation coefficient = 94.03dB / 500m≈0.188dB / m; Environmental interference correction: Combine the interference intensity (I, unit dBm) in the electromagnetic parameters and the precipitation intensity (P, unit mm / h) in the meteorological parameters, the correction formula is L = L0 + a x I + b x P, where a = 0.01 (interference intensity correction coefficient), b = 0.1 (precipitation intensity correction coefficient). For example, interference intensity I = -80 dBm, precipitation intensity P = 0.5 mm / h, after correction L = 94.03 + 0.01 x (-80) + 0.1 x 0.5 ≈ 94.03 - 0.8 + 0.05 ≈ 93.28 dB, the corrected attenuation coefficient = 93.28 dB / 500 m ≈ 0.187 dB / m; Terrain shielding correction: If the grid cell is a shielding area, an additional shielding attenuation ΔL = 10 x log (θ / 90°) (θ is the inclination of the shielding area, unit °) is added, for example, the rock wall inclination θ = 75°, ΔL = 10 x log (75 / 90) ≈ -0.79 dB, the final attenuation coefficient = 0.187 dB / m + (-0.79 dB) / 500 m ≈ 0.185 dB / m (shielding area attenuation is slightly lower, because the signal reflection is enhanced).
[0037] III. Generation of terrain feature map: Integrate the shielding area label and the signal attenuation coefficient to generate a visual terrain feature map, each grid cell in the map is labeled with "shielding state (shielding / non-shielding), attenuation coefficient (dB / m)", for example: "Terrain feature map (canyon terrain, grid size 5m x 5m): Grid (110.02°, 25.03°): shielding state = shielding (west rock wall), attenuation coefficient = 0.185 dB / m; Grid (110.05°, 25.05°): shielding state = non-shielding (open area at the bottom of the valley), attenuation coefficient = 0.188 dB / m; Grid (110.08°, 25.07°): shielding state = shielding (east rock wall), attenuation coefficient = 0.186 dB / m." The map needs to support dynamic updating (update when integrating subsequent unmanned aerial vehicle data), and provide core feature input for building a twin scene model.
[0038] Map the navigation position data and communication link state data of the real-time acquired unmanned aerial vehicle cluster to the corresponding grid cells of the terrain feature map according to the timestamp, update the data at fixed time intervals, and build a twin scene model containing static terrain features and dynamic cluster state.
[0039] The twin scene model needs to integrate "static terrain" and "dynamic cluster" data, ensuring spatiotemporal synchronization through timestamp mapping, solving the problem of "terrain and cluster state separation leading to non-dynamic model". This step needs to clarify "how to obtain cluster data, how to map, and how to update", achieving the dynamic and realistic nature of the twin.
[0040] I. Real-time acquisition of UAV cluster data: Acquire core data reflecting the state of the cluster, ensuring that the data matches the spatiotemporal dimensions of the terrain feature map: Navigation position data: Use the RTK-GPS (precision ±1cm) and IMU (Inertial Measurement Unit, update frequency 100Hz) built into the UAV to output the three-dimensional coordinates (longitude, latitude, height) of the UAV and the timestamp (accurate to milliseconds), for example, the position data of UAV 1 is "Timestamp: 2024-11-01 10:00:00.123, Coordinates: (110.02°, 25.03°, 500m)". Communication link state data: Collect link parameters through the Ad Hoc network between UAVs, including link signal-to-noise ratio (SNR, unit dB, measurement range 0-40dB), link connectivity state (1 = connected, 0 = interrupted), data transmission rate (unit Mbps), for example, the link data between UAV 1 and UAV 2 is "Timestamp: 2024-11-01 10:00:00.123, SNR=25dB, connectivity state = 1, rate = 10Mbps".
[0041] Data is transmitted in real time to the ground processing system through a wireless data link (transmission rate 50Mbps, delay ≤100ms), ensuring data timeliness.
[0042] II. Timestamp mapping and grid association: Map the cluster data to the grid cells of the terrain feature map according to the timestamp, ensuring that "the UAV is in which grid, and which grid terrain feature is associated": Timestamp synchronization: Use NTP protocol to calibrate (synchronization accuracy ≤1ms) the UAV clock and the ground system clock, ensuring that the timestamp of the cluster data is consistent with the collection timestamp of the terrain feature map; Grid positioning: Calculate the terrain grid cell where the UAV is located (through coordinate range matching) according to the three-dimensional coordinates of the UAV, for example, the UAV coordinates (110.02°, 25.03°, 500m) fall within the grid (110.02°-110.07°, 25.03°-25.08°), and associate the shielding state and attenuation coefficient of the grid; Data association: add the navigation position of the UAV and the link state data as "dynamic attributes" to the corresponding grid cell, for example, the attributes of the grid (110.02°, 25.03°) are updated to "occlusion state = occlusion, attenuation coefficient = 0.185 dB / m, UAV 1 position = (110.02°, 25.03°, 500 m), link 1 (1-2) = SNR 25 dB, connectivity 1".
[0043] III. Dynamic updating and twin scene model construction: Update the cluster data at fixed time intervals (set according to the terrain change frequency and cluster movement speed, for example, 30 seconds) to ensure that the model can reflect the real-time state: Update frequency setting: the update frequency of static features of the terrain (such as canyon rock wall, building height) is 1 hour (slow change), and the update frequency of dynamic data of the cluster is 30 seconds (fast movement of UAV); Model construction: use Unity 3D engine to build a visual twin scene, superimpose terrain feature map (static layer) and cluster data (dynamic layer), render terrain elevation and occlusion area (different colors) in the static layer, and render UAV position (icon) and communication link (line, color reflects SNR: green = SNR≥20dB, yellow = 10dB≤SNR<20dB, red = SNR<10dB) in the dynamic layer; Model verification: through field flight verification, the deviation between the UAV position in the twin scene and the actual position is ≤1m, and the deviation between the link state and the actual SNR is ≤2dB, which determines that the model is qualified.
[0044] The finally constructed twin scene model can not only intuitively show the static features of complex terrain, but also can reflect the dynamic state of the UAV cluster in real time, providing high-fidelity scene input for subsequent distributed reinforcement learning.
[0045] S202, input the twin scene model into the distributed reinforcement learning model, take the cluster communication connectivity and navigation signal integrity as the joint optimization target, design a reward decay function based on the terrain occlusion degree, and generate a topology optimization strategy of UAV relay node site selection and link priority allocation through multi-agent asynchronous decision-making; Specifically, key state parameters including terrain occlusion degree, relative position of UAV cluster, communication link connectivity state, and navigation signal strength can be extracted from the twin scene model, a multi-dimensional state space matrix can be constructed according to the UAV number and terrain grid index, and the matrix elements can be quantized to the corresponding state parameter values. The key state parameters are the input basis of the distributed reinforcement learning model, which need to be accurately extracted and quantified from the twin scene model to ensure that the parameters objectively reflect the dynamic relationship between the "terrain-cluster". The multi-dimensional state space matrix is the structured carrier of the parameters, which solves the problem of "parameter dispersion leading to inefficient learning of the model". This step needs to clarify "which parameters to extract, how to quantify, and how to build the matrix" to provide structured input for subsequent optimization.
[0046] I. Extraction and quantification of key state parameters: Each parameter needs to be combined with the static terrain features and dynamic cluster data of the twin scene model through quantifiable index definitions to avoid subjective judgments. The specific extraction and quantification methods are as follows: Terrain shielding degree: Reflects the signal shielding degree of the grid cell where the UAV is located, quantified based on the "sheltered area proportion" in the terrain feature map generated in step two, with a value range of 0-100% (0% for no shielding, 100% for complete shielding). The extraction method is: locate the terrain grid cell where the UAV is located in the twin scene model, calculate the area proportion of the sheltered area (such as canyon rock wall, city high-rise shadow) in the cell, for example, UAV No. 1 is located in grid No. 101, the total area of the grid is 25㎡ (5m×5m), the sheltered area is 17.5㎡, then the terrain shielding degree = (17.5 / 25) ×100% = 70%; UAV cluster relative position: Reflects the spatial relationship between a single UAV and the cluster center (affecting the link connection distance), taking the cluster geometric center as the origin, calculating the three-dimensional Euclidean distance (unit: meters) of the UAV, the formula is: relative distance =√[(x1-x0) 2 +(y1-y0) 2 +(z1-z0) 2 ], where (x0, y0, z0) is the cluster center coordinate, (x1, y1, z1) is the UAV coordinate. For example, the cluster center coordinate (110.05°, 25.05°, 500m), UAV No. 2 coordinate (110.08°, 25.07°, 520m), after coordinate conversion (1° longitude ≈111km, 1° latitude ≈111km) calculation plane distance ≈360 meters, height difference 20 meters, relative distance =√(360 2 +20 2 ) ≈360.56 meters, quantified as 361 meters (rounded to the integer); Communication link connectivity: reflects the effectiveness of the inter-UAV link, combined with link connectivity and signal-to-noise ratio (SNR) quantization, value range 0-1 (0 for interruption, 1 for optimal connectivity). The quantization formula is: connectivity value = connectivity identifier x (SNR-10) / 30, where connectivity identifier = 1 (connected) or 0 (interrupted), SNR threshold 10dB (lower than which the link is unstable), upper limit 40dB (optimal). For example, the link SNR between UAV No. 3 and UAV No. 4 is 25dB, and the connectivity identifier is 1, then the connectivity value is 1x(25-10) / 30=0.5; if SNR=5dB, connectivity identifier=0, state value=0; Navigation signal strength: reflects the reliability of UAV navigation positioning, based on GPS signal strength (unit: dBm) and positioning error (unit: meters) comprehensive quantization, value range 0-1 (0 for the worst, 1 for the best). The quantization formula is: navigation signal value = 0.5x[(-signal strength-80) / 40]+0.5x[1-(positioning error / 10)], where signal strength ranges from -120dBm (worst) to -80dBm (best), and positioning error ranges from 0 to 10 meters (10 meters is the threshold for excessive error). For example, the signal strength of UAV No. 5 is -90dBm, and the positioning error is 2 meters, then the navigation signal value = 0.5x[(90-80) / 40]+0.5x[1-(2 / 10)]=0.5x0.25+0.5x0.8=0.125+0.4=0.525, rounded to 0.53.
[0047] II. Construction of multi-dimensional state space matrix: The matrix is indexed by "UAV number - terrain grid index" as a two-dimensional index, and the elements are the quantized values of the above four parameters, ensuring that the state of each UAV is strongly associated with the terrain grid where it is located, and the matrix dimension is set according to the cluster size and terrain range (example: 5 UAVs, 5 core grids): Index definition: row index is UAV number (1-5), column index is terrain grid index (101-105, corresponding to the core grid of UAV activity in the twin scene); Element filling: each matrix element is a "terrain occlusion, relative position, communication link state, navigation signal strength" four-tuple, for example: Matrix element (1,101) = (70,350,0.5,0.53) → UAV No. 1 in grid No. 101, occlusion 70%, relative to cluster center 350 meters, link state 0.5, navigation signal 0.53; Matrix element (2, 102) = (30, 280, 0.8, 0.75) → Uav 2 in grid 102, occlusion 30%, relative distance 280m, link status 0.8, navigation signal 0.75; Matrix element (3, 103) = (90, 420, 0.2, 0.35) → Uav 3 in grid 103, occlusion 90%, relative distance 420m, link status 0.2, navigation signal 0.35; Matrix verification: Check if there are outliers in the elements (such as occlusion > 100%, relative distance negative), if so, correct based on adjacent elements, for example, element (4, 104) occlusion 110%, corrected to 100%, ensure the matrix data is reasonable.
[0048] Set joint optimization target, set cluster communication connectivity and navigation signal integrity as core optimization indicators, according to the influence weight of the two on cluster topology, weighted calculation to get comprehensive optimization index, clear index promotion direction; Joint optimization target is the "guiding sign" of reinforcement learning, need to balance the importance of communication and navigation, avoid single optimization leading to cluster function loss (such as only optimizing communication leading to navigation inaccuracy). This step needs to determine the index weight, calculation method and promotion direction, to ensure that the target is quantifiable and achievable.
[0049] I. Definition and calculation of core optimization indicators: Cluster communication connectivity: reflects the overall connectivity of the links between the drones in the cluster, the calculation method is "actual connected link number / theoretical maximum link number ×100%". Theoretical maximum link number = n×(n-1) / 2 (n is the number of drones), the actual connected link number is the number of links with link status value ≥0.3 (0.3 is the link effective threshold). For example, 5 drones, theoretical maximum link number = 5×4 / 2=10, actual connected link number = 7 (3 link status values <0.3), then communication connectivity = 7 / 10×100%=70%; Navigation signal integrity: reflects the overall reliability of the navigation positioning of the drones in the cluster, the calculation method is "the number of drones with navigation signal value ≥0.5 / total number of drones ×100%" (0.5 is the navigation reliability threshold). For example, among 5 drones, 4 have navigation signal value ≥0.5, 1 <0.5, then navigation signal integrity = 4 / 5×100%=80%.
[0050] II. Determination of index weight and calculation of comprehensive optimization index: The weight is determined based on "cluster task demand" and "expert experience". The communication connectivity rate is more critical for cluster data transmission, and the weight is higher than the navigation signal integrity. The specific settings are as follows: Weight allocation: cluster communication connectivity rate weight ω1=0.6, navigation signal integrity weight ω2=0.4 (weight sum is 1 to ensure fairness); Comprehensive optimization index formula: comprehensive index S=ω1×(communication connectivity rate / 100)+ω2×(navigation signal integrity / 100), value range 0-1 (1 is the best).
[0051] Example calculation: if the communication connectivity rate is 70% and the navigation signal integrity is 80%, then S=0.6×0.7+0.4×0.8=0.42+0.32=0.74; if the communication connectivity rate is improved to 90% and the navigation integrity is improved to 85%, then S=0.6×0.9+0.4×0.85=0.54+0.34=0.88, the index is significantly improved.
[0052] Three, the clear direction of index improvement: Based on the current value and target value of the comprehensive optimization index (usually set S≥0.9 as the standard), the improvement direction is clear: Low connectivity rate scenario (such as S=0.74, communication connectivity rate 70%): the improvement direction is "increase relay nodes, reduce sheltered area links, and increase the number of connected links", for example, add relay nodes in low sheltered grid (such as grid 102, sheltered degree 30%) to improve the interrupted link between drones 3 and 4; Low navigation integrity scenario (such as S=0.65, navigation integrity 60%): the improvement direction is "move drones with poor navigation signal (such as drone 3, signal value 0.35) to low sheltered areas to reduce the shelter of terrain on navigation signal"; Both indexes are low: prioritize improving the communication connectivity rate with higher weight, and then optimize the navigation integrity to avoid resource dispersion.
[0053] Design a reward decay function, the basic reward value is positively related to the comprehensive optimization index; introduce a terrain sheltering degree correction factor, when the sheltering degree of the area where the drone is located exceeds the set value, the reward value is linearly decayed according to the sheltering degree; if the communication link is interrupted or the navigation positioning is out of tolerance, set a fixed penalty value to form a dynamic reward mechanism; The reward decay function is the "feedback mechanism" of the reinforcement learning model, which needs to be optimized through "reward guidance" and "risk avoidance through punishment" to solve the problem of "the model cannot perceive the impact of terrain and the consequences of failure". The function design needs to consider "optimization target, terrain impact, and failure penalty" to ensure dynamic and reasonable.
[0054] I. Calculation of basic reward value: The basic reward value is positively correlated with the comprehensive optimization indicator S, encouraging the model to improve the indicator. The formula is: basic reward R0=10×S (10 is the reward coefficient, ensuring that the reward value is between 0 and 10, facilitating gradient update). Example: When S=0.74, R0=10×0.74=7.4; When S=0.88, R0=10×0.88=8.8; When S=0.95 (up to standard), R0=10×0.95=9.5 (close to full score reward).
[0055] II. Design of terrain shielding degree correction factor: Introduce correction factor γ to weaken the reward of high shielding degree area (avoid the model selecting shielding area as relay node), formula is: If terrain shielding degree α≤50% (low shielding), γ=1 (no attenuation, encourage selection); If 50%<α≤80% (medium shielding), γ=1 - (α-50) / 100 (linear attenuation, the higher the shielding degree, the smaller γ); If α>80% (high shielding), γ=0.2 (minimum attenuation coefficient, significantly weaken the reward, avoid selection).
[0056] Example calculation: UAV in α=30% (low shielding) area, γ=1, corrected reward R1=R0×γ=7.4×1=7.4; UAV in α=70% (medium shielding) area, γ=1 - (70-50) / 100=0.8, corrected reward R1=7.4×0.8=5.92; UAV in α=85% (high shielding) area, γ=0.2, corrected reward R1=7.4×0.2=1.48.
[0057] III. Setting of fixed penalty value: For the two types of faults "communication link interruption" and "navigation positioning deviation", set the fixed penalty value P to offset part of the reward and avoid the model ignoring the risk of failure: Communication link interruption penalty: if the link connectivity state value = 0 (interruption), P1=-5 (penalty 5 for each interrupted link, the more links, the heavier the penalty); Navigation positioning deviation penalty: if the navigation signal value of a UAV is less than 0.3 (positioning error >7 meters, deviation), P2=-3 (penalty 3 for each UAV with deviation).
[0058] IV. Complete calculation of dynamic reward mechanism: Final reward value R = R1 + ∑P1 + ∑P2, example: Scenario 1: S = 0.74, α = 70% (γ = 0.8), 1 link interruption (P1 = -5), 1 UAV navigation out-of-tolerance (P2 = -3), then R = 5.92 -5 -3 = -2.08 (negative reward, model needs to adjust strategy); Scenario 2: S = 0.88, α = 30% (γ = 1), 0 link interruptions, 0 out-of-tolerance, R = 8.8 +0 +0 = 8.8 (positive reward, model strategy is reasonable); Scenario 3: S = 0.95, α = 40% (γ = 1), 0 interruptions, 0 out-of-tolerance, R = 9.5 +0 +0 = 9.5 (high reward, model strategy is optimal).
[0059] The reward mechanism needs to be verified through simulation to ensure that the model obtains high rewards in low occlusion, high connectivity, and no out-of-tolerance scenarios, and vice versa, obtaining low rewards or even penalties, guiding the model to learn the optimal strategy.
[0060] Perform multi-agent asynchronous decision-making, split the multi-dimensional state space matrix into local state segments based on UAVs, and calculate decision priorities based on local segments for each UAV as an independent agent; the global coordination layer of the distributed reinforcement learning model integrates the decisions of each agent, preferentially selects low occlusion areas as relay nodes, assigns link priority based on signal attenuation coefficient reverse sorting, and generates a topology optimization strategy containing relay node three-dimensional coordinates and link priority sorting.
[0061] Multi-agent asynchronous decision-making is the core of distributed reinforcement learning, requiring each UAV to make autonomous decisions, and then avoiding conflicts through global coordination to solve the problem of "low efficiency of centralized decision-making and inability to adapt to cluster size". This step needs to clarify "how to make local decisions, how to integrate globally, and how to generate strategies" to ensure efficient decision-making and compliance with optimization goals.
[0062] I. Splitting of local state segments: Split the multi-dimensional state space matrix into independent local state segments based on UAV numbers, and each segment only contains the state parameters of the UAV and the link state of 2-3 adjacent UAVs (reducing data transmission volume and achieving asynchronous decision-making). For example: 1st UAV local segment: (70% self-occlusion, 350 meters relative position, 0.5 / 0.6 link state with 2nd and 5th UAVs, 0.53 self-navigation signal); 2nd UAV local segment: (self-occlusion 30%, relative position 280 meters, link state with 1st and 3rd UAVs 0.5 / 0.7, self-navigation signal 0.75); 3rd UAV local segment: (self-occlusion 90%, relative position 420 meters, link state with 2nd and 4th UAVs 0.7 / 0.2, self-navigation signal 0.35).
[0063] After splitting, each UAV makes independent decisions through a local computing unit without waiting for other UAVs, achieving asynchrony (decision interval ≤100 ms, adapting to UAV dynamic movement).
[0064] II. Decision priority calculation of independent agents: Each UAV, as an agent, calculates the "relay node candidate priority" based on the local segment. The higher the priority, the more suitable it is as a relay node (relay node requires low occlusion, high link state, and high navigation signal). The calculation formula is: Priority P = (1 - a / 100) × average link state value × navigation signal value: Where a is the self-occlusion, and the average link state value is the average of the link states with adjacent UAVs in the local segment.
[0065] Example calculation: 1st UAV: a = 70%, average link state = (0.5 + 0.6) / 2 = 0.55, navigation signal 0.53, P = (1 - 0.7) × 0.55 × 0.53 = 0.3 × 0.2915 ≈ 0.087; 2nd UAV: a = 30%, average link state = (0.5 + 0.7) / 2 = 0.6, navigation signal 0.75, P = (1 - 0.3) × 0.6 × 0.75 = 0.7 × 0.45 = 0.315; 3rd UAV: a = 90%, average link state = (0.7 + 0.2) / 2 = 0.45, navigation signal 0.35, P = (1 - 0.9) × 0.45 × 0.35 = 0.1 × 0.1575 ≈ 0.016; 4th UAV (a = 40%, average link 0.65, navigation 0.68): P = 0.6 × 0.65 × 0.68 ≈ 0.265; 5th UAV (a = 60%, average link 0.58, navigation 0.62): P = 0.4 × 0.58 × 0.62 ≈ 0.143.
[0066] Priority ranking: No. 2 (0.315) > No. 4 (0.265) > No. 5 (0.143) > No. 1 (0.087) > No. 3 (0.016), No. 2 and No. 4 are the optimal relay node candidates.
[0067] III. Decision integration of the global coordination layer: The distributed reinforcement learning model (using the FedRL federated reinforcement learning framework, and the global coordination layer is deployed in the ground control center) receives the priority results of each agent and integrates them according to the following rules: Relay node site selection: Select the top 2 priority drones as relay nodes (5 drones require 2 relay nodes to ensure coverage of all drones), i.e. No. 2 and No. 4 drones. Obtain their three-dimensional coordinates through the twin scene model: No. 2 (110.08°, 25.07°, 520m), No. 4 (110.03°, 25.02°, 510m), and verify whether the site selection is in a low-shielding area (No. 2 α=30%, No. 4 α=40%, meeting the requirements); Link priority allocation: Extract the signal attenuation coefficients (calculated in step 2, unit: dB / m) of all links from the digital twin of the terrain, and sort them in reverse order according to the rule "the smaller the attenuation coefficient, the higher the priority" (smaller attenuation means less signal loss, better link). For example, link list and attenuation coefficient: 2-1 (0.185), 2-3 (0.190), 4-5 (0.182), 4-1 (0.188), priority ranking: 4-5 (1) > 2-1 (2) > 4-1 (3) > 2-3 (4).
[0068] IV. Generation of topology optimization strategy: Integrate relay node and link priority to generate a structured strategy, including "relay node information, link priority list", example: "Unmanned aerial vehicle cluster topology optimization strategy under complex terrain (valley terrain): Relay node configuration: Relay node 1: No. 2, three-dimensional coordinates (110.08°, 25.07° N, 520m altitude), grid shielding degree 30%, responsible for covering No. 1 and No. 3 drones; Relay node 2: No. 4, three-dimensional coordinates (110.03°, 25.02° N, 510m altitude), grid shielding degree 40%, responsible for covering No. 5 and No. 1 drones; Communication link priority (sorted in reverse order according to signal attenuation coefficient): Priority 1: No. 4 - No. 5 link (attenuation coefficient 0.182 dB / m, SNR=28 dB); Priority 2: 2-1 link (attenuation coefficient 0.185dB / m, SNR=25dB); Priority 3: 4-1 link (attenuation coefficient 0.188dB / m, SNR=23dB); Priority 4: 2-3 link (attenuation coefficient 0.190dB / m, SNR=21dB); Policy valid period: 30 seconds (need to update at fixed intervals, adapt to UAV movement). The policy needs to be output to the next step to provide relay nodes and link direction basis for antenna parameter adjustment, while ensuring that the policy can be executed (relay nodes cover all UAVs, and there is no conflict in the link).
[0069] S203, according to the relay node position and link direction in the topology optimization strategy, call the beam pattern database of the reconfigurable antenna, match the shielding angle in the digital twin of the terrain to dynamically adjust the antenna beam width and gain parameters, and generate an antenna parameter-topology structure cooperative adaptation scheme; Specifically, the topology optimization strategy parameters can be parsed, and the identification information, three-dimensional coordinates and main communication link direction angle of each relay node can be extracted from the topology optimization strategy to clearly define the coverage range requirements of each link and form a strategy parameter table. Parsing the topology optimization strategy parameters is a bridge connecting "topology strategy" and "antenna parameter adjustment", which needs to accurately extract the core information of relay nodes and links to ensure that the subsequent antenna parameters can accurately match the topology structure and avoid adaptation deviation due to information loss. This step needs to clarify "what information to extract, how to define information, and how to organize" to provide clear input basis for calling the beam database.
[0070] I. Extraction and definition of core parameters: From the generated topology optimization strategy, four types of key parameters are extracted, each of which needs to be clear about the physical meaning and quantitative standard to avoid ambiguity: Relay node identification information: used to uniquely distinguish different relay nodes, including "node number" (consistent with the UAV number, such as UAV No. 2 as a relay node, the number is R2), "topography grid index" (associated with the digital twin of the terrain, such as R2 is located in grid No. 102), "node type" (primary relay / secondary relay, the primary relay is responsible for the core link, and the secondary relay assists in coverage, such as R2 is the primary relay and R4 is the secondary relay); Relay node three-dimensional coordinates: Reflects the spatial position of the relay node, directly affects the coverage direction of the antenna beam, adopts the WGS84 coordinate system, the format is "longitude (°), latitude (°), altitude (m)", for example, the three-dimensional coordinates of R2 are (110.08°, 25.07°, 520m), and the three-dimensional coordinates of R4 are (110.03°, 25.02°, 510m); Main communication link direction angle: Defines the main beam pointing direction of the relay node and other unmanned aerial vehicles, expressed in azimuth angle (precision ±1°) with "0° as the north direction and clockwise as positive", and 2-3 main link direction angles (covering the main communication objects) need to be extracted for each relay node. For example, the main links of R2 include "R2→1 unmanned aerial vehicle" direction angle 30° (azimuth angle from R2 to 1 unmanned aerial vehicle) and "R2→3 unmanned aerial vehicle" direction angle 150°; The main links of R4 include "R4→5 unmanned aerial vehicle" direction angle 330° and "R4→1 unmanned aerial vehicle" direction angle 60°. Link coverage range requirement: Based on the relative distance between unmanned aerial vehicles (calculated in the above steps), the communication radius (unit: m) centered on the relay node is determined to ensure covering the target unmanned aerial vehicle and not wasting energy. For example, the relative distance between R2 and 1 unmanned aerial vehicle is 180m, and the coverage range requirement is 200m (leaving 20m redundancy); The relative distance between R2 and 3 unmanned aerial vehicle is 220m, and the coverage range requirement is 250m.
[0071] II. Textual arrangement of strategy parameter table: Due to the prohibition of using tables, the parameters need to be arranged in a coherent text according to "relay node grouping" to ensure the information is structured and easy to read: "Topological optimization strategy parameter table (canyon terrain cluster, 5 unmanned aerial vehicles, 2 relay nodes): Relay node R2 (main relay, unmanned aerial vehicle 2, belongs to grid 102): Three-dimensional coordinates: Longitude 110.08°, latitude 25.07°, altitude 520m; Main communication link: Link 1: R2→1 unmanned aerial vehicle, direction angle 30°, coverage range requirement 200m; Link 2: R2→3 unmanned aerial vehicle, direction angle 150°, coverage range requirement 250m; Relay node R4 (secondary relay, unmanned aerial vehicle 4, belongs to grid 104): Three-dimensional coordinates: Longitude 110.03°, latitude 25.02°, altitude 510m; Main communication link: Link 3: R4→UAV5, direction angle 330°, coverage requirement 180m; Link 4: R4→UAV1, direction angle 60°, coverage requirement 220m. Parameter table needs to be marked with “policy generation time” (such as 2024-11-01 10:05:00) and “validity period” (30 seconds) to ensure that the latest policy parameters are used in subsequent steps.
[0072] Call the beam pattern database to retrieve the matching reconfigurable antenna initial beam parameters in the database according to the link direction angle in the policy parameter table, and extract the antenna parameter adjustment boundary to obtain the initial antenna configuration set; The beam pattern database is a core resource that stores the correspondence between “direction angle - beam parameter”, calling this database can quickly obtain the initial antenna parameters that adapt to the link direction, avoiding blind adjustment; extracting the adjustment boundary can ensure that subsequent parameter adjustment does not exceed the hardware capability, solving the problem of “antenna parameter out of range leading to hardware failure”.
[0073] I. Structure and content of the beam pattern database: The database is based on the measured data of reconfigurable antennas (such as phased array antennas, model AD9361, supporting beam width adjustment of 15°-120° and gain adjustment of 8dB-20dB), storing the correspondence between “link direction angle - beam width - gain - pattern”, where: Link direction angle: divided by 5° interval (0°, 5°, 10°…360°), covering all possible communication directions; Initial beam parameters: for each direction angle, store “default beam width” (the optimal width that adapts to the direction without obstruction) and “default gain” (the optimal value that balances coverage and energy consumption), for example, direction angle 0°-60° (open direction) default beam width 60°, gain 15dB; direction angle 60°-120° (canyon east rock wall direction) default beam width 80°, gain 14dB; Pattern data: store the beam pattern (energy distribution curve) corresponding to each parameter, used for subsequent verification of coverage, but not called during the initial configuration phase, only beam width and gain are extracted.
[0074] The database uses local cache + cloud backup storage method, local cache ensures that the call delay is ≤50ms, meeting the real-time requirements.
[0075] II. Retrieval process of initial beam parameters: According to the link direction angle in the strategy parameter table, the initial parameters are "matched nearby" (direction angle deviation ≤ 5°) in the database, as follows: Link 1 (R2→1st drone, direction angle 30°): The corresponding parameters of direction angle 30° are retrieved in the database, and matched to the default beam width 60° and the default gain 15dB; Link 2 (R2→3rd drone, direction angle 150°): The direction angle 150° is retrieved, and matched to the default beam width 80° and the default gain 14dB; Link 3 (R4→5th drone, direction angle 330°): The direction angle 330° (i.e. -30°) is retrieved, and matched to the default beam width 70° and the default gain 15dB; Link 4 (R4→1st drone, direction angle 60°): The direction angle 60° is retrieved, and matched to the default beam width 65° and the default gain 15dB.
[0076] If there is no complete matching item for the direction angle (such as direction angle 33°), the average value of the parameters of the closest direction angles 30° or 35° is taken, for example, 33° takes the average value of 30° (60° / 15dB) and 35° (62° / 15dB), the beam width is 61° and the gain is 15dB.
[0077] III. Extraction of antenna parameter adjustment boundaries: The adjustment boundary is the hardware physical limitation of the reconfigurable antenna, which needs to be extracted from the database to ensure that subsequent adjustments do not exceed the capability range. The parameters include: Beam width adjustment boundary: minimum 15° (narrow beam, high directivity), maximum 120° (wide beam, large coverage), adjustment step 5° (minimum adjustment unit supported by hardware); Gain adjustment boundary: minimum 8dB (low gain, low energy consumption), maximum 20dB (high gain, far coverage), adjustment step 1dB; Correlation constraint: there is a negative correlation between beam width and gain (the narrower the beam, the higher the gain), for example, the maximum gain is 20dB when the beam width is 15°, and the minimum gain is 8dB when the beam width is 120°. The database will mark this constraint to avoid subsequent adjustment conflicts.
[0078] IV. Formation of the initial antenna configuration set: Integrate the retrieved initial parameters and adjustment boundaries, and arrange them according to "link number" to form the initial configuration set, as follows: "Initial antenna configuration set (based on beam pattern database retrieval): Link 1 (R2→1st, direction angle 30°): Initial beamwidth: 60°, adjustment boundary 15°-120°; Initial gain: 15dB, adjustment boundary 8dB-20dB; Correlation constraint: beamwidth decreases by 5°, gain increases by 1dB; Link 2 (R2→3, direction angle 150°): Initial beamwidth: 80°, adjustment boundary 15°-120°; Initial gain: 14dB, adjustment boundary 8dB-20dB; Correlation constraint: beamwidth decreases by 5°, gain increases by 1dB; Link 3 (R4→5, direction angle 330°): Initial beamwidth: 70°, adjustment boundary 15°-120°; Initial gain: 15dB, adjustment boundary 8dB-20dB; Correlation constraint: beamwidth decreases by 5°, gain increases by 1dB; Link 4 (R4→1, direction angle 60°): Initial beamwidth: 65°, adjustment boundary 15°-120°; Initial gain: 15dB, adjustment boundary 8dB-20dB; Correlation constraint: beamwidth decreases by 5°, gain increases by 1dB. The initial configuration set needs to verify whether the parameters meet the coverage requirements, for example, the initial beamwidth of link 1 is 60°, the gain is 15dB, and the signal strength within the 200m coverage range is ≥-85dBm (communication threshold), which is determined to be qualified.
[0079] Dynamic matching of the shielding angle adjustment parameter, querying the maximum terrain shielding angle on each communication link path from the terrain digital twin, and adjusting the antenna parameters according to the shielding angle size classification, compressing the beamwidth and increasing the gain when the shielding angle is in the preset smaller range, keeping the parameters balanced when the shielding angle is in the preset medium range, and expanding the beamwidth and reducing the gain when the shielding angle is in the preset larger range, to obtain the adjusted antenna parameters; The terrain shielding angle is a key factor affecting beam coverage (the larger the shielding angle, the more easily the beam is blocked by the terrain), and needs to be adjusted dynamically to adapt to the shielding situation, solving the problem of "initial parameters not considering terrain shielding leading to coverage failure". This step needs to clarify "how to query the shielding angle, how to classify and adjust, and how to calculate the adjusted parameters", to ensure that the parameters adapt to the terrain.
[0080] I. Query method of maximum terrain shielding angle: Query the maximum shadowing angle on the link path from the terrain digital twin, the shadowing angle is defined as "the angle between the line connecting the link starting point (relay node) and the highest point on the path and the horizontal direction of the link" (unit: °), the query steps are: Link path modeling: In the digital twin, connect the three-dimensional coordinates of the relay node and the target UAV to generate a link path segment (such as the path of R2 (520m) → UAV No. 1 (510m)); Terrain highest point identification: Along the path segment, take a terrain sampling point every 10m, extract the elevation value of each sampling point, find the point corresponding to the maximum value (i.e. the highest point), for example, the sampling point elevation on link 1 path is 520m, 525m, 530m, 528m…510m, the highest point elevation is 530m, located 80m east of R2; Shadowing angle calculation: Use trigonometric function to calculate the shadowing angle θ, the formula is θ=arctan [(H_h - H_s) / d], where H_h is the highest point elevation, H_s is the link starting point elevation, and d is the horizontal distance from the starting point to the highest point. For example, in link 1, H_h=530m, H_s=520m, d=80m, θ=arctan [(530-520) / 80]=arctan (0.125)≈7.13°, rounded to 7° (maximum terrain shadowing angle).
[0081] Query the maximum shadowing angle of all links by this method, the result example is: link 1 θ=7°, link 2 θ=45°, link 3 θ=15°, link 4 θ=30°.
[0082] II. Shadowing angle preset range and adjustment rules: According to the influence of shadowing angle on beam coverage, the shadowing angle is preset to three ranges, each range corresponds to different parameter adjustment strategies, the rules are based on "small shadowing angle uses narrow beam high gain (reduces interference), large shadowing angle uses wide beam low gain (expands coverage)": Preset small range: 0°≤θ≤30° (small shadowing, beam not easy to be blocked), adjustment strategy is "compress beam width (reduce sidelobe interference), improve gain (enhance signal strength)", adjustment amplitude: beam width reduces by 20%, gain increases by 2dB; Preset medium range: 30°<θ≤60° (medium shadowing, beam partially blocked), adjustment strategy is "keep parameters balanced (neither compress nor expand, balance coverage and interference)", adjustment amplitude: beam width remains unchanged, gain remains unchanged; Pre-set larger range: 60° < θ ≤ 90° (large shielding, beam is easy to be blocked), adjustment strategy is "expand beam width (diffraction cover blocked area), reduce gain (avoid energy waste)", adjustment amplitude: beam width increases by 20%, gain decreases by 2dB.
[0083] The adjustment amplitude needs to meet the antenna adjustment boundary. If it exceeds the boundary after adjustment, the boundary value is taken. For example, the initial beam width is 15° (the minimum boundary), which cannot be compressed when the shielding angle is small, and only the gain is increased.
[0084] III. Calculation example of adjusted antenna parameters: Combined with the initial parameters and the shielding angle query results, the adjusted parameters are calculated link by link: Link 1 (θ = 7°, small range): Initial beam width 60° → compressed by 20% → 60° × (1-20%) = 48° (meets the 15°-120° boundary); Initial gain 15dB → increased by 2dB → 15dB + 2dB = 17dB (meets the 8dB-20dB boundary); Adjusted parameters: beam width 48°, gain 17dB; Link 2 (θ = 45°, medium range): Initial beam width 80° → keep unchanged → 80°; Initial gain 14dB → keep unchanged → 14dB; Adjusted parameters: beam width 80°, gain 14dB; Link 3 (θ = 15°, small range): Initial beam width 70° → compressed by 20% → 70° × 0.8 = 56°; Initial gain 15dB → increased by 2dB → 17dB; Adjusted parameters: beam width 56°, gain 17dB; Link 4 (θ = 30°, small range, because θ = 30° belongs to the upper limit of small range): Initial beam width 65° → compressed by 20% → 65° × 0.8 = 52°; Initial gain 15dB → increased by 2dB → 17dB; Adjusted parameters: beam width 52°, gain 17dB.
[0085] If a link θ = 70° (large range), the initial beam width is 60° → expanded by 20% → 72°, the initial gain is 15dB → decreased by 2dB → 13dB, and the adjusted parameters meet the boundary, which is determined to be qualified.
[0086] Synergistic verification of adaptability, check the matching of the adjusted antenna parameters and the relay node position, modify the parameters with conflicts; associate the relay node, link information and corresponding antenna parameters in the topology optimization strategy with the UAV ID to generate a synergistic adaptation scheme of antenna parameters-topology structure.
[0087] Synergistic verification is the key to ensure that "antenna parameters and topology structure have no conflicts", which needs to check whether the parameters adapt to the relay node position (such as whether the beam is blocked by the rock wall in the canyon), modify the conflicts, associate the ID after modification, form an executable synergistic adaptation scheme, and solve the problem of "parameters and topology disconnection leading to unexecutability".
[0088] I. Core dimensions of adaptability check: Check the matching of the adjusted parameters and the relay node position from three dimensions of "coverage range, terrain conflict, energy balance": Coverage range matching: verify whether the signal strength of the adjusted beam parameters in the coverage range meets the standard (≥-85dBm), the formula is signal strength = gain - 20log (d) - attenuation coefficient × d, where d is the coverage distance, and the attenuation coefficient is obtained from the digital twin. For example, the gain of link 1 is adjusted to 17dB, d=200m, attenuation coefficient 0.185dB / m, signal strength = 17-20log200-0.185×200≈17-46.02-37≈-66.02dBm≥-85dBm, coverage is qualified; Terrain conflict check: simulate the beam coverage range in the digital twin and check whether there is a case of "beam being completely blocked by terrain". For example, relay node R2 is located on the west side of the canyon, and the adjusted beam width of link 2 is 80° (direction angle 150°, pointing to the east rock wall), simulation finds that 30% of the beam is blocked by the rock wall, which exists conflict; Energy balance check: ensure that the total energy consumption of a single relay node ≤ the hardware upper limit (such as the maximum antenna energy consumption of R2 is 10W), the energy consumption calculation formula is energy consumption = gain 2 × beam width / 1000, for example, the energy consumption of link 1 of R2 is 17 2 ×48 / 1000≈13.87W, the energy consumption of link 2 is 14 2 ×80 / 1000≈15.68W, total energy consumption≈29.55W>10W, there is energy consumption conflict.
[0089] II. Modification method of conflict parameters: According to the principle of "first solve coverage, then balance energy consumption", modify the conflicts found in the check: Terrain conflict correction: Link 2 beam is blocked by 30% by rock wall, need to expand beam width to diffract coverage, expand beam width from 80° to 100° (still within 120° boundary), gain remains 14dB, re-simulate coverage, blockage ratio drops to 10% (acceptable); Energy consumption conflict correction: R2 total energy consumption is too high, need to reduce gain of non-core links, link 2 (non-core, only covers drone 3) gain from 14dB to 12dB, energy consumption = 12 2 ×100 / 1000=14.4W, link 1 energy consumption 13.87W, total energy consumption ≈28.27W, still exceeds the upper limit, further reduce link 1 gain from 17dB to 16dB, energy consumption = 16 2 ×48 / 1000≈12.29W, total energy consumption ≈12.29+14.4=26.69W, still exceeds, finally adjust link 2 beam width to 90°, gain 12dB, energy consumption = 12 2 ×90 / 1000=12.96W, total energy consumption ≈12.29+12.96=25.25W, although not fully meet the standards, but close, and the coverage is qualified, determine acceptable (subsequent iteration optimization).
[0090] Corrected link 2 parameters: beam width 90°, gain 12dB; link 1 parameters: beam width 48°, gain 16dB.
[0091] III. Generation of collaborative adaptation scheme: According to the association of relay nodes, link information and corrected antenna parameters, generate a structured scheme, example: "UAV cluster antenna parameter-topology structure collaborative adaptation scheme under complex terrain (valley terrain): Relay node R2 (drone 2, coordinates 110.08°, 25.07°, 520m): Associated link 1 (R2→drone 1): direction angle 30°, coverage range 200m, antenna parameters (beam width 48°, gain 16dB); Associated link 2 (R2→drone 3): direction angle 150°, coverage range 250m, antenna parameters (beam width 90°, gain 12dB); Node energy consumption: ≈25.25W (close to hardware upper limit 10W, need to be optimized later); Relay node R4 (drone 4, coordinates 110.03°, 25.02°, 510m): Link 3 (R4→UAV 5): 330° direction angle, 180m coverage, antenna parameters (56° beamwidth, 17dB gain); Link 4 (R4 → UAV 1): azimuth angle 60°, coverage range 220m, antenna parameters (beamwidth 52°, gain 17dB); Node energy consumption: ≈17 2 ×56 / 1000 + 17 2 ×52 / 1000≈16.46+15.14=31.6W (subsequent optimization is required); The plan is valid for 30 seconds and needs to be updated at intervals to adapt to the movement of the drone. Coverage verification: All link signal strengths are ≥-75dBm, and coverage is qualified. The solution must be marked with "items to be optimized" (such as energy consumption overruns) to provide direction for subsequent iterative optimization, while ensuring that the parameters can be directly called by the antenna adjustment module (such as a beam width of 48° corresponding to the duty cycle of the hardware control signal).
[0092] S204, based on the collaborative adaptation scheme, collect cluster communication quality and navigation continuity indicators in real time, feed back indicator deviations to the terrain digital twin for scene correction, drive the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjust the reconfigurable antenna parameters, and generate cluster topology dynamic optimization results that adapt to complex terrain changes.
[0093] Specifically, you can load the collaborative adaptation solution, extract the antenna parameters and topology strategy in the solution as the benchmark configuration, clarify the expected communication coverage and navigation positioning accuracy requirements of each UAV, and form a solution benchmark parameter table; Loading the collaborative adaptation solution is the foundation for subsequent metric collection and deviation analysis. It requires precise extraction of baseline parameters across the antenna and topology dimensions, while also clarifying expected performance targets to avoid a lack of reference for subsequent data collection. This step addresses the issue of "unclear baselines leading to no standard for deviation assessment" and ensures that each parameter has a clear expected value and physical meaning.
[0094] 1. Loading the collaborative adaptation solution and extracting core parameters: The collaborative adaptation solution is called through the system interface (the storage format is JSON, and the local cache delay is ≤ 50ms) to extract two types of core benchmark configurations: Antenna parameter benchmark: Extract the adjusted antenna parameters by drone ID (relay node and ordinary node), including "beam width (unit: °), gain (unit: dB), beam direction angle (unit: °)", which must correspond to the link one by one. For example: Relay node R2 (drone 2): link 1 (R2→1) beamwidth 48°, gain 16dB, direction angle 30°; link 2 (R2→3) beamwidth 90°, gain 12dB, direction angle 150°; Relay node R4 (drone 4): link 3 (R4→5) beamwidth 56°, gain 17dB, direction angle 330°; link 4 (R4→1) beamwidth 52°, gain 17dB, direction angle 60°; Common node (1, 3, 5): antenna parameters follow the relay node configuration by default (e.g., drone 1 receives the beam of R2, the parameters are consistent with link 1 of R2).
[0095] Topology strategy benchmark: extract the three-dimensional coordinates of relay nodes, link priority ranking, and link distance, for example: Relay node coordinates: R2 (110.08°, 25.07°, 520m), R4 (110.03°, 25.02°, 510m); Link priority: link 3 (R4→5) > link 1 (R2→1) > link 4 (R4→1) > link 2 (R2→3); Link distance: R2→1 180m, R2→3 220m, R4→5 150m, R4→1 200m.
[0096] II. Definition of expected communication coverage range and navigation positioning accuracy: Expected communication coverage range: based on link distance and antenna gain calculation, the formula is "coverage radius = link distance × 1.2" (1.2 is a redundancy coefficient to avoid coverage failure due to small movement of drones), unit: meters. For example: R2→1 link distance 180m, expected coverage radius = 180×1.2=216m; R4→5 link distance 150m, expected coverage radius = 150×1.2=180m.
[0097] At the same time, the "minimum signal strength threshold" (≥-85dBm to ensure communication quality) is also defined within the coverage range, which is based on the receiving sensitivity (-90dBm) of the drone communication module, leaving a 5dB redundancy.
[0098] Expected navigation positioning accuracy: reference "Unmanned Aerial Vehicle Cluster Navigation Technology Requirements", combined with complex terrain characteristics, set two types of indicators: Positioning error: ≤3 meters (high-precision positioning required to ensure accurate beam pointing), ≤5 meters for ordinary nodes; Navigation continuity: ≤1 second of continuous positioning interruption (to avoid topology confusion caused by positioning interruption).
[0099] III. Verbal presentation of the scheme reference parameter table: Organized by "drone ID - parameter type - reference value - expected target" logic, example: "Scheme reference parameter table (canyon terrain cluster, collection time 2024-11-01 10:10:00): Drone 2 (relay node R2): Antenna parameters: Link 1 (→1) beam width 48°, gain 16dB, direction angle 30°; Link 2 (→3) beam width 90°, gain 12dB, direction angle 150°; Topology reference: coordinates (110.08°, 25.07°, 520m), link 2 priority 4; Expected target: Link 1 coverage radius 216m (signal strength ≥-85dBm), positioning error ≤3m; Drone 4 (relay node R4): Antenna parameters: Link 3 (→5) beam width 56°, gain 17dB, direction angle 330°; Link 4 (→1) beam width 52°, gain 17dB, direction angle 60°; Topology reference: coordinates (110.03°, 25.02°, 510m), link 3 priority 1; Expected target: Link 3 coverage radius 180m (signal strength ≥-85dBm), positioning error ≤3m; Drone 1 (ordinary node): Antenna parameters: receive R2 link 1 (48° / 16dB), R4 link 4 (52° / 17dB); Expected target: positioning error ≤5m, navigation continuous interruption duration ≤1s; Drone 3, 5 (ordinary nodes): Expected target: positioning error ≤5m, navigation continuous interruption duration ≤1s, received signal strength ≥-85dBm. Based on the scheme reference parameter table, collect performance indicators at preset time intervals; compare the actual collected values with the expected values in the scheme reference parameter table, calculate the deviation rate of each indicator, and form an indicator deviation set; Performance index collection is the core of verifying the effectiveness of the scheme, which needs to cover the "communication-navigation" double dimensions at fixed intervals, and quantify the gap between actual and expected through deviation rate, providing data support for subsequent correction. This step needs to solve the problems of "incomplete index collection" and "no standard for deviation calculation", and ensure the reproducibility of deviation results.
[0100] I. Setting of preset time interval: The interval needs to balance "real-time" and "data stability", and is set to 10 seconds (i.e. collecting data every 10 seconds) according to the dynamic moving speed of UAV cluster (about 5m / s in the canyon) and the frequency of terrain changes. If the signal of a link fluctuates frequently (e.g. absolute value of deviation rate >10%), it can be dynamically shortened to 5 seconds to avoid missing transient anomalies; if the index is stable (absolute value of deviation rate <5%), it can be extended to 20 seconds to reduce the computational burden.
[0101] II. Collection method of communication quality and navigation continuity indexes: Communication quality index: Link signal strength (unit: dBm): Collect through the signal detection interface of the UAV communication module, take the mean of 10 data points collected each time for each link, for example, R2→1 link collects values of -70dBm, -72dBm…-68dBm, the mean is -70dBm; Link packet loss rate (unit%): Calculate the number of lost packets by ping test (send 100 data packets), the formula is "packet loss rate = number of lost packets / total data packets ×100%", for example, send 100 packets and lose 5, the packet loss rate is 5%; Communication connectivity duration (unit: s): Record the time of continuous connectivity (signal strength ≥-85dBm, packet loss rate ≤10%) of the link, for example, R4→5 link is continuously connected for 120 seconds.
[0102] Navigation continuity index: Positioning error (unit: m): Calculate through RTK-GPS and IMU fusion data, compare the Euclidean distance between the actual position of the UAV and the true value (provided by the ground base station), for example, true value coordinates (110.08°, 25.07°, 520m), actual coordinates (110.08°, 25.07°, 523m), positioning error =√[(0) 2 +(0) 2 +(3) 2 ]=3m; Navigation interruption duration (unit: s): record the continuous time when positioning error > 5m (ordinary node) or > 3m (relay node), for example, the positioning error of UAV No. 1 is > 5m for 8 seconds continuously, and the interruption duration is 8 seconds.
[0103] III. Calculation of deviation rate and formation of index deviation set: The deviation rate is used to quantify the deviation of the actual value from the expected value, and the formula is "deviation rate = (actual value - expected value) / expected value x 100%", and the result is rounded to two decimal places. A positive number indicates that it exceeds the expectation, and a negative number indicates that it does not meet the expectation. Combined with the benchmark parameter example, the collection and calculation example is as follows: Link 1 (R2→1): signal strength, expected ≥-85dBm (target value is -80dBm), actual -70dBm, deviation rate = (-70 - (-80)) / (-80) x 100% = -12.50% (better than expected); Link 2 (R2→3): packet loss rate, expected ≤5%, actual 8%, deviation rate = (8-5) / 5 x 100% = 60.00% (exceeds expectation); UAV No. 1 positioning error: expected ≤5m, actual 6m, deviation rate = (6-5) / 5 x 100% = 20.00% (exceeds expectation); R4 navigation interruption duration: expected ≤1s, actual 0s, deviation rate = (0-1) / 1 x 100% = -100.00% (better than expected).
[0104] The index deviation set is sorted by "index type - UAV ID - actual value - expected value - deviation rate", and the example is as follows: "Index deviation set (collection time 2024-11-01 10:10:10): Communication quality index: Link 1 (R2→1): signal strength, actual -70dBm, expected -80dBm, deviation rate -12.50%; Link 2 (R2→3): packet loss rate, actual 8%, expected 5%, deviation rate 60.00%; Link 3 (R4→5): connectivity duration, actual 100s, expected 120s, deviation rate -16.67%; Navigation continuity index: UAV No. 1: positioning error, actual 6m, expected 5m, deviation rate 20.00%; UAV No. 3: interruption duration, actual 2s, expected 1s, deviation rate 100.00%; R2 (No. 2): positioning error actual 2m, expected 3m, deviation rate -33.33%. The feedback index deviation correction twin body screens abnormal items that exceed the scheme fault tolerance range, associates the abnormal deviation to the corresponding grid unit of the terrain digital twin body, updates the shielding area boundary and signal attenuation coefficient of the unit, and generates a corrected terrain digital twin body. Deviation feedback and twin body correction are the key to realizing "dynamic optimization", which needs to locate the root cause of the abnormality (mostly inaccurate terrain parameter labeling), update the static features of the twin body, and provide accurate scene input for subsequent strategy iteration. This step needs to solve the problems of "abnormal deviation without root cause positioning, twin body and actual terrain disconnection", and ensure that the corrected scene is close to the real environment.
[0105] I. Setting of scheme fault tolerance range: The fault tolerance range is the critical value that distinguishes "normal fluctuation" from "abnormal deviation", which is set based on the industry standards of unmanned aerial vehicle communication and navigation and the characteristics of terrain interference. Different indicators have different fault tolerance ranges due to different sensitivity: Communication indicators: signal strength deviation rate ±15%, packet loss rate deviation rate ±20%, connected time deviation rate ±25%; Navigation indicators: positioning error deviation rate ±30%, interruption time deviation rate ±50%.
[0106] Deviation beyond this range is considered "abnormal item" and needs further analysis of the root cause.
[0107] II. Screening and root cause positioning of abnormal items: Abnormal items are screened from the index deviation set, and through "link path-grid unit" association, the corresponding terrain grid of the abnormality is located. Examples: Abnormal item 1: link 2 (R2→3 No.) packet loss rate deviation rate 60.00% (exceeds ±20%), link path covers terrain grid No. 103 (between R2 and No. 3 unmanned aerial vehicle), preliminary judgment that the grid shielding area labeling is insufficient or the attenuation coefficient is too low; Abnormal item 2: No. 3 unmanned aerial vehicle interruption time deviation rate 100.00% (exceeds ±50%), No. 3 unmanned aerial vehicle is located in grid No. 103, it is speculated that the grid navigation signal is severely shielded, and the original shielding degree of 70% may be too low; Normal item: link 1 deviation rate -12.50% (within ±15%), no correction needed.
[0108] III. Correction operation of terrain digital twin body: For the grid cell associated with the abnormal item, update the static characteristics through "field survey data completion + electromagnetic wave propagation model recalculation": Sheltered area boundary update: For grid 103, through the LiDAR re-scanning carried by the UAV, it is found that the original sheltered area only labeled the west rock wall, missing the north protruding rock block (height 530m, higher than the flight height of UAV 3 510m), the sheltered area boundary is extended from "west 10m range" to "west 10m + north 8m range", and the sheltering degree is updated from 70% to 85%; Signal attenuation coefficient recalculation: based on the updated sheltering degree (85%) and environmental electromagnetic parameters (interference intensity -75dBm), the attenuation coefficient is recalculated using the free space propagation model, the formula is "attenuation coefficient = 0.185×(1 + sheltering degree / 100)" (sheltering degree increases by 10%, attenuation coefficient increases by 10%), the original coefficient is 0.185dB / m, and after updating = 0.185×(1+85 / 100)=0.185×1.85≈0.342dB / m; Other grid verification: for the associated adjacent grids 102 (R2 is located) and 104 (R4 is located), the sheltering and attenuation parameters are verified synchronously to ensure there is no chain error.
[0109] Four, generation of the corrected terrain digital twin: Integrate all updated grid parameters to generate new terrain feature map and digital twin, example: "Corrected terrain digital twin (canyon terrain, grid size 5m×5m): Grid 103 (associated with abnormal items 1 and 2): Sheltered area: west 10m + north 8m (original west 10m), sheltering degree 85% (original 70%); Signal attenuation coefficient: 0.342dB / m (original 0.185dB / m); Elevation value: max530m, min500m, avg515m; Grid 102 (R2 is located): Sheltering degree 30% (no change), attenuation coefficient 0.185dB / m (no change); Grid 104 (R4 is located): Sheltering degree 40% (no change), attenuation coefficient 0.190dB / m (no change); Other grids: no update of parameters, keep original configuration." The modified twin body needs to be verified by "signal strength measurement", for example, the measured signal strength of link 2 in grid 103 is reduced from -75dBm to -88dBm, which is consistent with the updated attenuation coefficient calculation result (error ≤2dB), and the modification is qualified.
[0110] Iterative optimization strategy and parameters, input the modified terrain digital twin into the distributed reinforcement learning model, and adjust the topology optimization strategy in the collaborative adaptation scheme; according to the new topology strategy, modify the antenna parameters, and collect the adjusted indicators to verify whether the deviation rate meets the scheme requirements; after meeting the requirements, integrate the optimized topology strategy and antenna parameters to generate the cluster topology dynamic optimization result that adapts to complex terrain changes.
[0111] Iterative optimization is the final link to achieve "adaptation to complex terrain changes", which needs to update the topology strategy through reinforcement learning and adjust the antenna parameters simultaneously until the indicators meet the requirements, forming a closed loop optimization. This step needs to solve the problem of "no dynamic adjustment of strategy and parameters, and the optimization result does not meet the requirements", to ensure that the final scheme adapts to the modified terrain.
[0112] I. Adjustment of topology optimization strategy: Input the modified terrain digital twin into the distributed reinforcement learning model (FedRL framework), and combine the parameter conflicts in the original scheme (such as R2→3 link attenuation is too large), adjust the strategy through multi-agent asynchronous decision: Relay node site adjustment: the attenuation between the original relay node R2 (grid 102) and the 3rd unmanned aerial vehicle (grid 103) is too large, move R2 to the adjacent low-shielding grid 101 (shielding degree 20%, attenuation coefficient 0.170dB / m), new coordinates (110.07°, 25.06°, 520m); Link priority rearrangement: after R2 moves, the attenuation of link 2 (R2→3) is reduced, so its priority is raised from 4 to 2; link 3 (R4→5) maintains priority 1; Link coverage range adjustment: the distance of R2→3 link is shortened from 220m to 200m, and the expected coverage radius is adjusted from 250m to 240m.
[0113] Example of the adjusted new topology optimization strategy: "Iterative topology optimization strategy (based on modified twin): Relay node R2 (unmanned aerial vehicle 2): new coordinates (110.07°, 25.06°, 520m), located in grid 101 (shielding degree 20%); Relay Node R4 (Drone 4): coordinates remain unchanged (110.03°, 25.02°, 510m); Link priority: Link 3 (R4→5) > Link 2 (R2→3) > Link 1 (R2→1) > Link 4 (R4→1). II. Synchronization correction of antenna parameters: According to the link direction and distance in the new topology strategy, adjust the antenna parameters to ensure adaptation to the new propagation path: Link 2 (R2→3): After R2 moves, the direction angle is adjusted from 150° to 145°, the attenuation coefficient is reduced to 0.170dB / m, the beam width does not need to be expanded and remains at 90°, and the gain is increased from 12dB to 14dB (to compensate for the energy consumption of the shortened distance); Link 1 (R2→1): After R2 moves, the distance increases from 180m to 190m, the beam width is slightly adjusted from 48° to 50° (to expand the coverage), and the gain remains at 16dB; Other links: No significant adjustment to parameters, maintain original configuration.
[0114] Example of corrected antenna parameters: "Antenna parameter configuration after iteration: Link 2 (R2→3): beam width 90°, gain 14dB, direction angle 145°; Link 1 (R2→1): beam width 50°, gain 16dB, direction angle 32°; Other links: parameters consistent with the original scheme." III. Index verification and result integration: Collect the adjusted indicators at 10-second intervals to verify whether the deviation rate meets the tolerance range: Link 2 packet loss rate: actual 5%, expected 5%, deviation rate 0.00% (meets the standard); 3rd drone interruption duration: actual 1s, expected 1s, deviation rate 0.00% (meets the standard); Link 1 signal strength: actual -72dBm, expected -80dBm, deviation rate -10.00% (meets the standard).
[0115] All index deviation rates are within the tolerance range. Integrate the optimized topology strategy and antenna parameters to generate the final dynamic optimization result: "Dynamic optimization result of unmanned aerial vehicle cluster topology under complex terrain (canyon terrain, optimization time 2024-11-0110:15:00): Topology optimization strategy: Relay nodes: R2 (110.07°, 25.06°, 520m, Grid 101), R4 (110.03°, 25.02°, 510m, Grid 104); Link priority: R4→5#> R2→3#> R2→1#> R4→1#; Antenna parameter configuration: R2 Link 1: 50° / 16dB / 32°, Link 2: 90° / 14dB / 145°; R4 Link 3: 56° / 17dB / 330°, Link 4: 52° / 17dB / 60°; Performance indicators: Communication connectivity rate 95% (original 70%), navigation signal integrity 98% (original 80%); Adaptability description: after modifying the terrain grid 103 parameters, the strategy and parameter adaptation degree of 85% in the complex area can support subsequent 10-minute dynamic adjustment. Another embodiment of the application provides a UAV cluster topology dynamic optimization system under complex terrain, referring to Figure 3 , the system can include: The acquisition module 301 is used for acquiring terrain three-dimensional modeling data and real-time environment parameters of a canyon or a city cluster, generating a terrain digital twin containing a shielding area and a signal attenuation coefficient through a terrain feature extraction algorithm, and fusing position and communication link state data of a UAV cluster to construct a dynamically updated twin scene model; The generation module 302 is used for inputting the twin scene model into a distributed reinforcement learning model, taking cluster communication connectivity rate and navigation signal integrity as joint optimization objectives, designing a reward decay function based on terrain shielding degree, and generating a topology optimization strategy of UAV relay node site selection and link priority allocation through multi-agent asynchronous decision-making; The matching module 303 is used for calling a beam pattern database of a reconfigurable antenna according to relay node positions and link directions in the topology optimization strategy, matching shielding angles in the terrain digital twin to dynamically adjust antenna beam width and gain parameters, and generating a coordinated adaptation scheme of antenna parameters-topology structure; The optimization module 304 is used for collecting cluster communication quality and navigation continuity indicators in real time based on the coordinated adaptation scheme, feeding back indicator deviations to the terrain digital twin for scene correction, driving the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjusting reconfigurable antenna parameters, and generating a cluster topology dynamic optimization result adapted to complex terrain changes.
[0116] The embodiment of the present application further provides a storage medium, wherein the storage medium stores a computer program, and the computer program is configured to execute the steps in any one of the method embodiments when running.
[0117] The embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to execute the steps in any one of the method embodiments.
[0118] Specifically, the electronic device can further comprise a transmission device connected with the processor and an input / output device connected with the processor.
[0119] The above embodiment according to the drawings illustrates the structure, features and effects of the present application, and the above description is only the preferred embodiment of the present application, but the present application is not limited to the drawings, any change or modification within the spirit of the present application, or equivalent embodiments within the scope of the present application, should be within the protection scope of the present application.
Claims
1. A method for dynamic optimization of drone cluster topology in complex terrain, characterized by: The method comprises: Collect 3D terrain modeling data and real-time environmental parameters of canyons or urban clusters, generate a digital twin of the terrain including obscured areas and signal attenuation coefficients through terrain feature extraction algorithms, and integrate the location and communication link status data of the drone cluster to build a dynamically updated twin scene model; The twin scenario model is input into a distributed reinforcement learning model. With cluster communication connectivity and navigation signal integrity as joint optimization objectives, a reward decay function based on terrain obscuration is designed. A topology optimization strategy for drone relay node site selection and link priority allocation is generated through multi-agent asynchronous decision-making. Based on the relay node positions and link directions in the topology optimization strategy, the beam pattern database of the reconfigurable antenna is called upon to dynamically adjust the antenna beam width and gain parameters to match the shielding angle in the terrain digital twin, thus generating a collaborative adaptation scheme for the antenna parameters and topology structure. Based on the collaborative adaptation scheme, cluster communication quality and navigation continuity indicators are collected in real time, and the indicator deviations are fed back to the terrain digital twin for scene correction. This drives the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjust the reconfigurable antenna parameters, and generate dynamic optimization results of the cluster topology that adapts to complex terrain changes.
2. The method according to claim 1, characterized in that The method involves collecting three-dimensional terrain modeling data and real-time environmental parameters of canyons or urban clusters, generating a terrain digital twin including shielded areas and signal attenuation coefficients through a terrain feature extraction algorithm, and integrating the location and communication link status data of the drone cluster to construct a dynamically updated twin scene model, including: Collect raw data by category. For canyon terrain, collect 3D modeling data on elevation difference, rock wall inclination, and valley bottom width. For urban agglomerations, collect 3D modeling data on building height, building spacing, and street orientation. Also collect real-time environmental parameters. Meteorological parameters include wind speed, wind direction, and precipitation intensity, and electromagnetic parameters include interference frequency and signal interference intensity. These are organized by terrain type and parameter dimension to form a raw multidimensional dataset. Preprocess the original multidimensional data, grid the terrain 3D modeling data, unify the grid cell size and fill in the elevation value; use the spatiotemporal smoothing algorithm to eliminate instantaneous disturbances in real-time environmental parameters, and use the Kalman filter to correct outliers that exceed the physical reasonable range to obtain a standardized data grid; Extract terrain characteristic parameters and use a deep learning semantic segmentation algorithm to analyze the standardized data grid, identify and mark the shielded areas in the terrain, and calculate the signal attenuation coefficient of each grid cell based on the electromagnetic wave propagation model and the environmental electromagnetic parameters to generate a terrain characteristic map; The real-time navigation position data and communication link status data of the drone cluster are mapped to the corresponding grid cells of the terrain feature map according to the timestamp. The data is updated at fixed time intervals to construct a twin scene model that includes static terrain features and dynamic cluster status.
3. The method according to claim 2, characterized in that The twin scenario model is input into a distributed reinforcement learning model, cluster communication connectivity and navigation signal integrity are used as joint optimization objectives, a reward decay function based on terrain shielding is designed, and a topology optimization strategy for drone relay node site selection and link priority allocation is generated through multi-agent asynchronous decision-making, including: Extract key state parameters from the twin scene model, including terrain obscuration, relative position of the UAV cluster, communication link connectivity status, and navigation signal strength. Construct a multidimensional state space matrix based on the UAV number and terrain grid index, and quantify the corresponding state parameter values with matrix elements. Set a joint optimization goal, taking cluster communication connectivity and navigation signal integrity as core optimization indicators. Based on the weight of their impact on cluster topology, a weighted calculation is performed to obtain a comprehensive optimization indicator, clarifying the direction of indicator improvement. Design a reward decay function, where the base reward value is positively correlated with the comprehensive optimization index; introduce a terrain obscuration correction factor; when the obscuration of the area where the drone is located exceeds the set value, the reward value decays linearly according to the obscuration; if the communication link is interrupted or the navigation positioning is out of tolerance, set a fixed penalty value to form a dynamic reward mechanism; Multi-agent asynchronous decision-making is performed, and the multidimensional state space matrix is split into local state segments according to the drone. Each drone acts as an independent agent to calculate the decision priority based on the local segment. The global coordination layer of the distributed reinforcement learning model integrates the decisions of each agent, prioritizes low-shading areas as relay nodes, assigns link priorities in reverse order according to the signal attenuation coefficient, and generates a topology optimization strategy that includes the three-dimensional coordinates of the relay nodes and the link priority ranking.
4. The method according to claim 3, characterized in that The method calls the beam pattern database of the reconfigurable antenna based on the relay node position and link direction in the topology optimization strategy, dynamically adjusts the antenna beam width and gain parameters according to the shielding angle in the terrain digital twin, and generates a collaborative adaptation scheme for the antenna parameters and topology structure, including: Analyze the topology optimization strategy parameters, extract the identification information, three-dimensional coordinates, and main communication link direction angle of each relay node from the topology optimization strategy, clarify the coverage requirements of each link, and organize them into a strategy parameter table; Call the beam pattern database, retrieve the matching reconfigurable antenna initial beam parameters in the database according to the link direction angle in the strategy parameter table, and extract the antenna parameter adjustment boundary to obtain the initial antenna configuration set; Dynamically match shielding angle adjustment parameters and query the maximum terrain shielding angle on each communication link path from the terrain digital twin. Adjust antenna parameters based on shielding angle. When the shielding angle is within the preset small range, the beam width is compressed and the gain is increased. When the shielding angle is within the preset medium range, the parameters are balanced. When the shielding angle is within the preset large range, the beam width is expanded and the gain is reduced, thus obtaining the adjusted antenna parameters. Collaboratively verify adaptability, check the matching of adjusted antenna parameters and relay node positions, and correct conflicting parameters; associate the relay nodes and link information in the topology optimization strategy with the corresponding antenna parameters according to the drone ID to generate a collaborative adaptation plan for antenna parameters and topology structure.
5. The method according to claim 4, characterized in that Based on the collaborative adaptation solution, cluster communication quality and navigation continuity indicators are collected in real time, and indicator deviations are fed back to the terrain digital twin for scene correction. This drives the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjust the reconfigurable antenna parameters, and generate cluster topology dynamic optimization results that adapt to complex terrain changes, including: Load the collaborative adaptation solution, extract the antenna parameters and topology strategy in the solution as the benchmark configuration, clarify the expected communication coverage and navigation positioning accuracy requirements of each UAV, and form a solution benchmark parameter table; Based on the solution benchmark parameter table, performance indicators are collected, and the communication quality indicators and navigation continuity indicators of the drone cluster are collected at preset time intervals. The actual collected values are compared with the expected values in the solution benchmark parameter table, and the deviation rate of each indicator is calculated to form an indicator deviation set. Feedback indicator deviation correction twin, filter out abnormal items in the indicator deviation set that exceed the scheme fault tolerance range, associate the abnormal deviation with the corresponding grid cell of the terrain digital twin, update the shielding area boundary and signal attenuation coefficient of the cell, and generate a corrected terrain digital twin; Iterate the optimization strategy and parameters, input the corrected terrain digital twin into the distributed reinforcement learning model, and adjust the topology optimization strategy based on the parameter conflicts in the collaborative adaptation scheme; simultaneously correct the antenna parameters according to the new topology strategy, and collect the adjusted indicators to verify whether the deviation rate meets the scheme requirements; after meeting the standards, integrate the optimized topology strategy and antenna parameters to generate a cluster topology dynamic optimization result that adapts to complex terrain changes.
6. A UAV swarm topology dynamic optimization system under complex terrain, characterized by: The system comprises: The acquisition module is used to collect three-dimensional terrain modeling data and real-time environmental parameters of canyons or urban clusters. It uses terrain feature extraction algorithms to generate a terrain digital twin that includes obscured areas and signal attenuation coefficients. It also integrates the location and communication link status data of the drone cluster to build a dynamically updated twin scene model. A generation module is used to input the twin scenario model into a distributed reinforcement learning model, take cluster communication connectivity and navigation signal integrity as joint optimization objectives, design a reward decay function based on terrain obscuration, and generate a topology optimization strategy for drone relay node site selection and link priority allocation through multi-agent asynchronous decision-making; A matching module is used to call the beam pattern database of the reconfigurable antenna based on the relay node position and link direction in the topology optimization strategy, dynamically adjust the antenna beam width and gain parameters to match the shielding angle in the terrain digital twin, and generate a collaborative adaptation solution for the antenna parameters and topology structure; The optimization module is used to collect cluster communication quality and navigation continuity indicators in real time based on the collaborative adaptation scheme, feed back indicator deviations to the terrain digital twin for scene correction, drive the distributed reinforcement learning system to iteratively update the topology optimization strategy, synchronously adjust the reconfigurable antenna parameters, and generate dynamic optimization results of the cluster topology that adapts to complex terrain changes.
7. The system according to claim 6, characterized in that The acquisition module is specifically used to: Collect raw data by category. For canyon terrain, collect 3D modeling data on elevation difference, rock wall inclination, and valley bottom width. For urban agglomerations, collect 3D modeling data on building height, building spacing, and street orientation. Also collect real-time environmental parameters. Meteorological parameters include wind speed, wind direction, and precipitation intensity, and electromagnetic parameters include interference frequency and signal interference intensity. These are organized by terrain type and parameter dimension to form a raw multidimensional dataset. Preprocess the original multidimensional data, grid the terrain 3D modeling data, unify the grid cell size and fill in the elevation value; use the spatiotemporal smoothing algorithm to eliminate instantaneous disturbances in real-time environmental parameters, and use the Kalman filter to correct outliers that exceed the physical reasonable range to obtain a standardized data grid; Extract terrain characteristic parameters and use a deep learning semantic segmentation algorithm to analyze the standardized data grid, identify and mark the shielded areas in the terrain, and calculate the signal attenuation coefficient of each grid cell based on the electromagnetic wave propagation model and the environmental electromagnetic parameters to generate a terrain characteristic map; The real-time navigation position data and communication link status data of the drone cluster are mapped to the corresponding grid cells of the terrain feature map according to the timestamp. The data is updated at fixed time intervals to construct a twin scene model that includes static terrain features and dynamic cluster status.
8. The system according to claim 7, characterized in that The generation module is specifically used to: Extract key state parameters from the twin scene model, including terrain obscuration, relative position of the UAV cluster, communication link connectivity status, and navigation signal strength. Construct a multidimensional state space matrix based on the UAV number and terrain grid index, and quantify the corresponding state parameter values with matrix elements. Set a joint optimization goal, taking cluster communication connectivity and navigation signal integrity as core optimization indicators. Based on the weight of their impact on cluster topology, a weighted calculation is performed to obtain a comprehensive optimization indicator, clarifying the direction of indicator improvement. Design a reward decay function, where the base reward value is positively correlated with the comprehensive optimization index; introduce a terrain obscuration correction factor; when the obscuration of the area where the drone is located exceeds the set value, the reward value decays linearly according to the obscuration; if the communication link is interrupted or the navigation positioning is out of tolerance, set a fixed penalty value to form a dynamic reward mechanism; Multi-agent asynchronous decision-making is performed, and the multidimensional state space matrix is split into local state segments according to the drone. Each drone acts as an independent agent to calculate the decision priority based on the local segment. The global coordination layer of the distributed reinforcement learning model integrates the decisions of each agent, prioritizes low-shading areas as relay nodes, assigns link priorities in reverse order according to the signal attenuation coefficient, and generates a topology optimization strategy that includes the three-dimensional coordinates of the relay nodes and the link priority ranking.
9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Digital twin distribution network fault scene generation method and system based on reinforcement learning
CN116933619A
Intelligent factory production scheduling method and system based on digital twinning
CN120295250A
Unmanned aerial vehicle heterogeneous cluster collaborative navigation method and system facing complex terrain
CN120445223A
Intelligent vibration digital twin systems and methods for industrial environments
US20210157312A1
Digital model based reverse osmosis plant operation and optimization
US20230150835A1
Cited By
Geological surveying and mapping method and system based on cooperation of multiple unmanned aerial vehicles
CN121230677A
Low-altitude station site selection digital optimization method, equipment and medium
CN121723882A
Intelligent AI-combined unmanned aerial vehicle air-ground integrated communication method
CN121887749A
Multi-mode 5G Internet of Things card signal intelligent coverage method and system
CN121908284A
A multi-mode 5G internet of things card signal intelligent coverage method and system
CN121908284B