A distributed communication method and system for low-altitude networks

By using observation time-series vectors and time-series analysis models to predict interference distribution in low-altitude networks, and combining interference correction and reinforcement learning to optimize spectrum allocation, the problems of spectrum resource congestion and electromagnetic environment changes in low-altitude networks are solved, achieving efficient spectrum utilization and communication reliability.

CN121334879BActive Publication Date: 2026-04-03SHENZHEN COTELL TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Low-altitude networks suffer from congested spectrum resources and highly variable electromagnetic environments, making it difficult for existing technologies to effectively cope with rapidly changing interference, resulting in low spectrum utilization efficiency and poor communication reliability.

Method used

By utilizing the interference spectrum, target statistics, and the observed time-series vectors of its own state, combined with a time-series analysis model, interference distribution is predicted, sub-band power allocation is performed, and spectrum allocation is optimized through interference correction and reinforcement learning to achieve distributed spectrum coordination.

Benefits of technology

It significantly improves the spectrum utilization efficiency and communication reliability of low-altitude networks in scenarios with strong interference and strong time-varying characteristics, reduces spectrum collisions, avoids central scheduling bottlenecks, and improves the ability to guarantee critical services and the overall service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121334879B_ABST
    Figure CN121334879B_ABST
Patent Text Reader

Abstract

This invention relates to the field of communication technology, specifically to a distributed communication method and system for low-altitude networks. A distributed communication system for low-altitude networks includes: an observation vector acquisition module, an interference timing analysis module, a sub-band allocation module, an interference correction module, and a reinforcement learning module. This invention utilizes an observation timing vector containing interference spectrum vectors, target statistical vectors, and its own state vector, combined with a timing analysis model to model the evolution of the electromagnetic environment and target distribution within a finite time window. It outputs a predicted interference distribution vector covering all sub-bands, enabling sensing nodes to proactively avoid deteriorating interference sub-bands during sub-band power allocation. This statistically reduces spectrum collisions, improves spectrum utilization efficiency, and enhances communication reliability. Furthermore, the pre-allocation results are exchanged between nodes, and a neighbor table containing neighbor location information and sub-band power pre-allocation vectors is maintained. Interference correction is then performed on the sub-band allocation network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and specifically to a distributed communication method and system for low-altitude networks. Background Technology

[0002] Low-altitude sensing networks typically consist of numerous sensing nodes such as drones and aerial base stations, simultaneously carrying multiple services including communication, sensing, and control within low-altitude airspace such as cities and industrial parks. Because low-altitude airspace already contains signals from navigation, radar, remote sensing, and various dedicated systems, coupled with frequent access by drone swarms, limited spectrum resources become highly congested in time, frequency, and space. The electromagnetic environment is highly time-varying, nodes move rapidly, topology changes frequently, there are multiple operators, and it is difficult to rely on a unified center for spectrum coordination. Existing technologies employ two main approaches: one relies on pre-planned or static / semi-static frequency band allocation and power masks, which struggle to reflect electromagnetic environment changes in a timely manner under the complex and rapidly evolving interference environment of low-altitude regions, easily leading to long-term congestion of some subbands and underutilization of others; the other approach uses centralized or quasi-centralized spectrum management, collecting interference observations from each sensing node through a central node and then uniformly allocating subbands and power. While this can reduce collisions to some extent, it suffers from high signaling overhead, high latency, single point of failure at the central point, and difficulty adapting to cross-domain scenarios with multiple operators. Summary of the Invention

[0003] This invention utilizes an observation time-series vector containing interference spectrum vectors, target statistical vectors, and its own state vector. Combined with a time-series analysis model, it models the evolution of the electromagnetic environment and target distribution within a finite time window, outputting a predicted interference distribution vector covering all subbands. This allows sensing nodes to proactively avoid deteriorating interference subbands during subband power allocation, thereby statistically reducing spectrum collisions and improving spectrum utilization efficiency and communication reliability. Furthermore, by feeding the predicted interference distribution vector into the subband allocation network to obtain a subband power pre-allocation vector, this pre-allocation result is exchanged between nodes, and a neighbor table containing neighbor location information and the subband power pre-allocation vector is maintained. Based on this, the subband allocation network is further optimized. Interference correction enables the final subband power allocation vector to avoid strong external interference while spontaneously shifting away from highly occupied subbands of neighbors and appropriately yielding to heavily loaded neighbors. This effectively reduces mutual interference between multiple nodes and avoids central scheduling bottlenecks. Simultaneously, the subband allocation network is updated using reinforcement learning based on the communication quality values ​​collected between each monitoring time slice. This allows the network to achieve an adaptive balance between reducing interference to neighbors and improving communication commands. As a result, the nodes converge to a distributed spectrum collaborative state with less interference and higher resource sharing efficiency through multiple rounds of game iteration. This significantly improves the spectrum utilization efficiency, critical service assurance capabilities, and overall service quality of the low-altitude sensing network in scenarios with strong interference and strong time-varying characteristics.

[0004] This invention provides a distributed communication method for low-altitude networks, comprising:

[0005] The sensing node collects the observation vector corresponding to the time slice at the current monitoring time point. The observation vector includes the response interference spectrum vector, the interference target statistical vector, and its own state vector.

[0006] The observation time series vector is composed of the currently acquired observation vector and the observation vectors acquired at the previous N-1 monitoring time points, where N is the preset time series window length. The observation time series vector is then fed into the interference time series analysis model for processing, and the predicted interference distribution vector is output. The predicted interference vector includes the predicted interference intensity of all sub-bands in the main working band.

[0007] The predicted interference distribution vector is fed into the subband allocation network for processing, and the output subband power pre-allocation vector includes the allocated power of all subbands in the main working band.

[0008] Based on the neighbor table corresponding to the sensing node, interference correction is performed on the sub-band allocation network. Then, the predicted interference distribution vector is sent to the interference-corrected sub-band allocation network for processing, and the sub-band power allocation vector is output. The signal transmission of the sensing node in the next time slice is configured through the sub-band power allocation vector.

[0009] Before outputting the subband power pre-allocation vector at the next monitoring time point, the subband allocation network is reinforced and updated by using the collected call quality values.

[0010] Preferably, based on the neighbor table corresponding to the sensing node, interference correction is performed on the sub-band allocation network, specifically including the following steps:

[0011] The predicted interference distribution vector is fed into the target subband allocation network for processing, outputting a target subband power pre-allocation vector. The predicted interference distribution vector and the target subband power pre-allocation vector are then concatenated and fed into the target allocation strategy evaluation network for processing, outputting a target strategy evaluation value. This evaluation value represents the proportion of subbands with weak interference selected by the subband allocation network. Based on the neighbor table corresponding to the sensing node and the subband power pre-allocation vector, the neighbor non-interference degree is calculated. The sum of the target strategy evaluation value and the non-interference degree is used as the gradient value of the target allocation strategy evaluation. The parameters within the allocation strategy evaluation network are iteratively updated using the gradient ascent method, maximizing the gradient value, until convergence. Finally, the predicted interference distribution vector and the subband power pre-allocation vector are concatenated and fed into the updated allocation strategy evaluation network. The network is processed, and the gradient value of the allocation strategy evaluation network for the subband power pre-allocation vector is calculated and denoted as the subband allocation gradient value. The node type corresponding to the synesthetic node is obtained. The node type includes aggressive and conservative. The upper limit of the gradient norm corresponding to the node type is determined according to the node type. The upper limit of the gradient norm corresponding to the aggressive type is higher than the upper limit of the gradient norm corresponding to the conservative type. It is determined whether the gradient norm corresponding to the subband allocation gradient value is higher than the upper limit of the gradient norm. If so, the parameters of the subband allocation network are iteratively updated using the gradient ascent method based on the subband allocation gradient value until convergence is completed, and the interference correction is completed. Otherwise, the subband allocation gradient value is scaled proportionally to the upper limit of the gradient norm. Then, the parameters of the subband allocation network are iteratively updated using the gradient ascent method based on the scaled subband allocation gradient value until convergence is completed, and the interference correction is completed.

[0012] Preferably, the neighbor non-interference degree is calculated based on the neighbor table corresponding to the sensing node and the sub-band power pre-allocation vector, specifically including the following steps:

[0013] Traverse all subbands in the main working band. For each subband, determine whether the sum of all allocated power in the subband's neighbor table and the subband power pre-allocation vector is higher than the rated allocated power. If the sum of all allocated power in the subband's neighbor table and the subband power pre-allocation vector is higher than the rated allocated power, mark the subband as an interfering subband. If the sum of all allocated power in the subband's neighbor table and the subband power pre-allocation vector is not higher than the rated allocated power, mark the subband as a low-interference subband. Calculate the ratio of the number of all low-interference subbands to the number of all subbands in the main working band, and record it as the neighbor non-interference degree.

[0014] Preferably, reinforcement learning operations are performed on the subband allocation network using the collected call quality values, specifically including the following steps:

[0015] The predicted interference distribution vector is fed into the target subband allocation network for processing, outputting a target subband power pre-allocation vector. The predicted interference distribution vector and the target subband power pre-allocation vector are then concatenated and fed into the target allocation policy evaluation network for processing, outputting a target policy evaluation value. Call quality values ​​are then obtained, and the sum of the target policy evaluation value and the call quality value is used as the gradient value for reinforcing the target allocation policy evaluation. The parameters within the allocation policy evaluation network are iteratively updated using the gradient ascent method in the direction that maximizes the gradient value, until convergence. The predicted interference distribution vector and the subband power pre-allocation vector are then concatenated and fed into the updated allocation policy evaluation network for processing, and the gradient value of the allocation policy evaluation network with respect to the subband power pre-allocation vector is calculated and denoted as the subband allocation gradient value. Based on the subband allocation gradient value, the parameters of the subband allocation network are iteratively updated using the gradient ascent method until convergence, completing interference correction and reinforcement learning operations. The target subband allocation network and the target allocation policy evaluation network are updated periodically.

[0016] Preferably, the target subband allocation network and the target allocation policy evaluation network are updated periodically, specifically including the following steps: the product of the parameters and the learning rate in the subband allocation network that has completed the reinforcement learning operation is directly added to the corresponding parameters in the target subband allocation network to complete the update of the target subband allocation network; the product of the parameters and the learning rate in the allocation policy evaluation network that has completed the reinforcement learning operation is directly added to the corresponding parameters in the target allocation policy evaluation network to complete the update of the target allocation policy evaluation network.

[0017] Preferably, training the subband allocation network specifically includes the following steps: obtaining several subband allocation training samples, each including an interference distribution vector; labeling the subband allocation training samples using a subband power allocation vector; forming a subband allocation training set from all labeled subband allocation training samples; and training the subband allocation network using the subband allocation training set.

[0018] Preferably, training the allocation strategy evaluation network specifically includes the following steps: obtaining several allocation strategy evaluation training samples, which include interference distribution vectors in the subband allocation training set and corresponding labeled subband power allocation vectors; labeling the allocation strategy evaluation training samples with allocation strategy evaluation values; forming an allocation strategy evaluation training set from all labeled allocation strategy evaluation training samples; and training the allocation strategy evaluation network using the allocation strategy evaluation training set.

[0019] The present invention also provides a distributed communication system for low-altitude networks, comprising:

[0020] The observation vector acquisition module is used by the sensing node to acquire the observation vector corresponding to the time slice at the current monitoring time point. The observation vector includes the response interference spectrum vector, the interference target statistical vector, and its own state vector.

[0021] The interference time series analysis module is used to combine the currently acquired observation vector with the observation vectors acquired at the previous N-1 monitoring time points to form an observation time series vector, where N is the preset time series window length. The observation time series vector is then fed into the interference time series analysis model for processing, and the predicted interference distribution vector is output. The predicted interference vector includes the predicted interference intensity of all sub-bands in the main working band.

[0022] The subband allocation module is used to feed the predicted interference distribution vector into the subband allocation network for processing and output the subband power pre-allocation vector, which includes the allocated power of all subbands in the main working band.

[0023] The interference correction module is used to perform interference correction on the sub-band allocation network based on the neighbor table corresponding to the sensing node, and then send the predicted interference distribution vector into the interference-corrected sub-band allocation network for processing, outputting the sub-band power allocation vector, and configuring the signal transmission of the sensing node in the next time slice through the sub-band power allocation vector.

[0024] The reinforcement learning module is used to perform reinforcement learning operations on the subband allocation network using the collected call quality values ​​before outputting the subband power pre-allocation vector at the next monitoring time point, thereby achieving reinforcement updates to the subband allocation network.

[0025] The present invention has the following advantages:

[0026] This invention utilizes an observation time-series vector containing interference spectrum vectors, target statistical vectors, and its own state vector. Combined with a time-series analysis model, it models the evolution of the electromagnetic environment and target distribution within a finite time window, outputting a predicted interference distribution vector covering all subbands. This allows sensing nodes to proactively avoid deteriorating interference subbands during subband power allocation, thereby statistically reducing spectrum collisions and improving spectrum utilization efficiency and communication reliability. Furthermore, by feeding the predicted interference distribution vector into the subband allocation network to obtain a subband power pre-allocation vector, this pre-allocation result is exchanged between nodes, and a neighbor table containing neighbor location information and the subband power pre-allocation vector is maintained. Based on this, the subband allocation network is further optimized. Interference correction enables the final subband power allocation vector to avoid strong external interference while spontaneously shifting away from highly occupied subbands of neighbors and appropriately yielding to heavily loaded neighbors. This effectively reduces mutual interference between multiple nodes and avoids central scheduling bottlenecks. Simultaneously, the subband allocation network is updated using reinforcement learning based on the communication quality values ​​collected between each monitoring time slice. This allows the network to achieve an adaptive balance between reducing interference to neighbors and improving communication commands. As a result, the nodes converge to a distributed spectrum collaborative state with less interference and higher resource sharing efficiency through multiple rounds of game iteration. This significantly improves the spectrum utilization efficiency, critical service assurance capabilities, and overall service quality of the low-altitude sensing network in scenarios with strong interference and strong time-varying characteristics. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the structure of a distributed communication system for low-altitude networks used in an embodiment of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this invention.

[0029] Example 1: A distributed communication method for low-altitude networks, comprising:

[0030] For drones that make up the low-altitude network, drones are regarded as sensing nodes. It should be noted that the low-altitude network consists of ground cellular base stations, airborne platforms (referring to drones) and satellites, and can achieve distributed communication through a multi-link topology.

[0031] The sensing node collects the observation vector corresponding to the time slice at the current monitoring time point. The observation vector includes the response interference spectrum vector, the interference target statistical vector, and its own state vector. The response interference spectrum vector includes the power corresponding to the interference in each sub-band. The response interference spectrum vector characterizes the interference received by the sub-band at the frequency domain level. It should be noted that during the execution of communication functions, the UAV has a main operating band according to the configuration file, such as 3.4–3.6 GHz, which is the main frequency range for signal transmission during communication. For the convenience of scheduling and multiplexing more users, the main operating band will be further divided into smaller segments. Each segmented frequency range is called a sub-band or subcarrier. For example, it may be divided into 20 sub-bands, each sub-band... The range is 10MHz. The interference target statistical vector refers to the position, distance, and power of the target corresponding to the interference source response after shortwave transmission. It should be noted that the targets responding to the interference source here include radar stations, communication base stations, and other drones that transmit signals themselves, as well as large buildings that strongly emit signals. If there is a strong interference target in a certain direction and at a certain distance, the interference in each subband in that direction will be higher. Therefore, the interference target statistical vector can reflect the existence of interference in the physical space. The self-state vector includes the drone's own position, speed, transmission power, and signal waveform, etc. The self-state vector represents the state of the drone's collected observation vector. And the time between adjacent monitoring time points is a time slice.

[0032] The observation time-series vector is composed of the currently acquired observation vector and the observation vectors acquired at the previous N-1 monitoring time points, where N is the preset time-series window length, used to cover the evolution of the electromagnetic environment and target distribution within a finite time period. The observation time-series vector is then fed into the interference time-series analysis model for processing, outputting a predicted interference distribution vector. The interference time-series analysis model can be any of the following: Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), Gated Recurrent Unit (GRU), or Temporal Convolutional Network (TCN). It is used to model the coupling relationship between the temporal correlation and multi-source features in the observation time-series vector, and to predict the interference intensity of all sub-bands in the main working band. The predicted interference intensity is generally described by the average interference power. It should be noted that the interference time-series analysis model is trained using a training sample composed of N observation vectors arranged in time from actual acquisition. The predicted output target and the last observation vector in the training sample are compared using the loss value calculated by MSE, and the model is trained based on the loss value.

[0033] The predicted interference distribution vector is fed into the subband allocation network for processing, outputting a subband power pre-allocation vector. This vector includes the allocated power of all subbands in the main operating band. It's important to note that allocated power refers to the signal transmission strength of the current sensing node in its corresponding subband. Generally, higher allocated power indicates less interference in the current subband, allowing it to perform critical data transmission services, while lower allocated power enables the transmission of low-speed control signals. The subband power pre-allocation vector is then sent to all neighboring nodes corresponding to the current sensing node. Simultaneously, the current sensing node receives the subband power pre-allocation vectors from its neighbors and updates its neighbor table. The neighbor table includes a one-to-one correspondence between location information and the subband power pre-allocation vector. This location information can be stored using GNSS coordinates. The subband allocation network typically employs a feedforward neural network, specifically consisting of an input layer, several hidden layers, and an output layer.

[0034] It should be added that neighbor nodes generally refer to sensing nodes that can perform a single wireless transmission with the current sensing node on the control channel. The general method of determination is that each sensing node periodically broadcasts a very short control packet on a dedicated control channel. If other sensing nodes can receive this control packet, they will merge the corresponding location information into the control packet and send it back. When two sensing nodes establish contact through the control packet, they can be determined to be neighbors and added to the neighbor table of the sensing node. The neighbor table includes the location information of the neighbor node, and the neighbor table is also updated periodically.

[0035] Based on the neighbor table corresponding to the sensing node, interference correction is performed on the subband allocation network. Then, the predicted interference distribution vector is fed into the interference-corrected subband allocation network for processing, and the subband power allocation vector is output. The signal transmission of the sensing node in the next time slice is configured through the subband power allocation vector. It should be noted that interference correction refers to adjusting the subband allocation network by considering the interference of the currently selected subband power pre-allocation vector to the neighbor nodes. This allows the subband power allocation vector output by the interference-corrected subband allocation network to avoid strong external interference, spontaneously avoid high neighbor occupancy, and appropriately give way to heavily loaded neighbor nodes.

[0036] Before outputting the subband power pre-allocation vector at the next monitoring time point, reinforcement learning operations are performed on the subband allocation network using the collected call quality values ​​to achieve reinforcement updates to the subband allocation network. It should be noted that by performing reinforcement learning operations on the subband allocation network, the subband allocation network will not get overly caught up in avoiding interference from neighboring nodes during the subband power allocation process. Furthermore, the game-like iterative interaction between interference correction and reinforcement learning allows all nodes to gradually converge to a spectrum division state with less interference and higher resource sharing efficiency, thereby achieving distributed spectrum collaboration without central control.

[0037] It should be added that, in the initial state of use, all the sensor nodes corresponding to the drones have a consistent sub-band allocation network built in.

[0038] This application utilizes observation time-series vectors containing interference spectrum vectors, target statistical vectors, and their own state vectors. Combined with a time-series analysis model, it models the evolution of the electromagnetic environment and target distribution within a finite time window, outputting a predicted interference distribution vector covering all subbands. This enables sensing nodes to proactively avoid deteriorating interference subbands during subband power allocation, thereby statistically reducing spectrum collisions and improving spectrum utilization efficiency and communication reliability. Furthermore, by feeding the predicted interference distribution vector into the subband allocation network to obtain a subband power pre-allocation vector, this pre-allocation result is exchanged between nodes, and a neighbor table containing neighbor location information and the subband power pre-allocation vector is maintained. Based on this, the subband allocation network is further optimized. Interference correction enables the final subband power allocation vector to avoid strong external interference while spontaneously shifting away from highly occupied subbands of neighbors and appropriately yielding to heavily loaded neighbors. This effectively reduces mutual interference between multiple nodes and avoids central scheduling bottlenecks. Simultaneously, the subband allocation network is updated using reinforcement learning based on the communication quality values ​​collected between each monitoring time slice. This allows the network to achieve an adaptive balance between reducing interference to neighbors and improving communication commands. As a result, the nodes converge to a distributed spectrum collaborative state with less interference and higher resource sharing efficiency through multiple rounds of game iteration. This significantly improves the spectrum utilization efficiency, critical service assurance capabilities, and overall service quality of the low-altitude sensing network in scenarios with strong interference and strong time-varying characteristics.

[0039] Based on the neighbor table corresponding to the sensing node, interference correction is performed on the sub-band allocation network, specifically including the following steps:

[0040] The predicted interference distribution vector is fed into the target subband allocation network for processing, outputting a target subband power pre-allocation vector. The predicted interference distribution vector and the target subband power pre-allocation vector are then concatenated and fed into the target allocation policy evaluation network for processing, outputting a target policy evaluation value. This target policy evaluation value represents the proportion of subbands with weak interference selected by the subband allocation network. Concatenation is typically done by joining the first and last subbands. Then, based on the neighbor table corresponding to the sensing node and the subband power pre-allocation vector, the neighbor non-interference degree is calculated. The sum of the target policy evaluation value and the non-interference degree is used as the target allocation policy evaluation gradient value. The gradient is then calculated by maximizing the target allocation policy evaluation gradient value. The ascending method iteratively updates the parameters within the allocation strategy evaluation network until convergence. Then, the predicted interference distribution vector and the subband power pre-allocation vector are concatenated and fed into the updated allocation strategy evaluation network for processing. The gradient value of the allocation strategy evaluation network with respect to the subband power pre-allocation vector is calculated and denoted as the subband allocation gradient value. The node type corresponding to the synesthetic node is obtained. The node type includes aggressive and conservative types, which are set by the operator. Aggressive type indicates that the corresponding synesthetic node is more biased towards subbands with weaker interference during subband power allocation, exhibiting aggressive subband contention behavior. Conservative type indicates that the node is more cautious during subband allocation... The strategy changes during the process are more stable. The upper limit of the gradient norm is determined according to the node type, with the upper limit of the gradient norm for the aggressive type being higher than that for the conservative type. It is then determined whether the gradient norm corresponding to the subband allocation gradient value is higher than the upper limit. If so, the parameters of the subband allocation network are iteratively updated using the gradient ascent method based on the subband allocation gradient value until convergence, completing the interference correction. Otherwise, the subband allocation gradient value is scaled proportionally to the upper limit of the gradient norm, and the parameters of the subband allocation network are iteratively updated using the gradient ascent method based on the scaled subband allocation gradient value until convergence, completing the interference correction. This process involves adjusting the gradient norm. The size of the upper limit can achieve different policy update styles: when the upper limit is large, the network parameters are updated more rapidly in a single iteration, the policy adjusts its preference for high-return subbands more quickly, and it exhibits a more "aggressive" spectrum contention behavior; when the upper limit is small, the step size of each network update is strictly limited, the policy preference changes more smoothly, it is easier to find policies with lower interference to neighbors, which helps to suppress frequent subband switching and excessive preemption, and reflects a more "conservative" spectrum access characteristic; in the initial state, the target subband allocation network is consistent with the subband allocation network, and the target allocation policy evaluation network is consistent with the allocation policy evaluation network.

[0041] The neighbor non-interference degree is calculated based on the neighbor table corresponding to the sensing node and the sub-band power pre-allocation vector, specifically including the following steps:

[0042] Traverse all sub-bands in the main working band. For each sub-band, determine whether the sum of all allocated power in the sub-band's neighbor table and sub-band power pre-allocation vector is higher than the rated allocated power. The rated allocated power is set by the operator. If the sum of all allocated power in the sub-band's neighbor table and sub-band power pre-allocation vector is higher than the rated allocated power, it indicates that the sub-band is overloaded or has high interference, and the sub-band is marked as an interfering sub-band. If the sum of all allocated power in the sub-band's neighbor table and sub-band power pre-allocation vector is not higher than the rated allocated power, it indicates that the sub-band has low interference, and the sub-band is marked as a low-interference sub-band. Calculate the ratio of the number of all low-interference sub-bands to the number of all sub-bands in the main working band, and record it as the neighbor non-interference degree.

[0043] The subband allocation network is subjected to reinforcement learning operations based on the collected call quality values, specifically including the following steps:

[0044] The predicted interference distribution vector is fed into the target subband allocation network for processing, outputting a target subband power pre-allocation vector. The predicted interference distribution vector and the target subband power pre-allocation vector are then concatenated and fed into the target allocation policy evaluation network for processing, outputting a target policy evaluation value. A call quality value, typically SINR, is then obtained. The sum of the target policy evaluation value and the call quality value is used as the gradient value for reinforcing the target allocation policy evaluation. The parameters within the allocation policy evaluation network are iteratively updated using the gradient ascent method, maximizing the direction of the gradient value, until convergence. The predicted interference distribution vector and the subband power pre-allocation vector are then concatenated and fed into the updated allocation policy evaluation network for processing, and the gradient value of the allocation policy evaluation network with respect to the subband power pre-allocation vector is calculated, denoted as the subband allocation gradient value. Based on the subband allocation gradient value, the parameters of the subband allocation network are iteratively updated using the gradient ascent method until convergence, completing interference correction and reinforcement learning operations. The target subband allocation network and the target allocation policy evaluation network are updated periodically.

[0045] The target subband allocation network and the target allocation strategy evaluation network are updated periodically, specifically including the following steps:

[0046] The product of the parameters and learning rate in the subband allocation network that has completed the reinforcement learning operation is directly added to the corresponding parameters in the target subband allocation network to update the target subband allocation network. The learning rate is set by the operator and is generally 0.01. The product of the parameters and learning rate in the allocation policy evaluation network that has completed the reinforcement learning operation is directly added to the corresponding parameters in the target allocation policy evaluation network to update the target allocation policy evaluation network.

[0047] Training the subband allocation network involves the following steps:

[0048] Several subband allocation training samples are obtained. These samples include interference distribution vectors, stored in the same format as the predicted interference distribution vectors. The difference is that the interference distribution vectors are obtained from actual data. The subband allocation training samples are labeled using subband power allocation vectors, which are set by the operator based on the interference distribution vectors, aiming for the lowest interference. All labeled subband allocation training samples are combined into a subband allocation training set. The subband allocation network is trained using this training set. During training, the difference between the predicted output of the subband allocation network and the labeled subband power allocation vectors is used to construct a loss value using the Mean Squared Error (MSE). Based on this loss value, parameters are updated using gradient descent. Training is then completed, and the accuracy of the subband allocation network is assessed. If the accuracy meets expectations, the trained subband allocation network is output; otherwise, training continues using the subband allocation training set.

[0049] Training the allocation policy evaluation network involves the following steps:

[0050] Several allocation strategy evaluation training samples are obtained. These samples include the interference distribution vector in the subband allocation training set and the corresponding labeled subband power allocation vector. The allocation strategy evaluation training samples are labeled with the allocation strategy evaluation value. Since the value is set by the operator based on the interference distribution vector, it tends to select the option with the lowest interference. The allocation strategy evaluation value is generally set to 1. All labeled allocation strategy evaluation training samples are combined into an allocation strategy evaluation training set. The allocation strategy evaluation network is trained using this training set. During training, the difference between the predicted output of the allocation strategy evaluation network and the labeled allocation strategy evaluation value of the allocation strategy evaluation training samples is used to construct a loss value using the method of MSE. Based on the loss value, the parameters are updated using the gradient descent method to complete the training. The accuracy of the allocation strategy evaluation network is then judged to see if it meets the expectations. If the accuracy meets the expectations, the trained allocation strategy evaluation network is output; otherwise, the allocation strategy evaluation network is trained again using the allocation strategy evaluation training set.

[0051] Example 2: A distributed communication system for low-altitude networks, such as... Figure 1 As shown, it includes:

[0052] The observation vector acquisition module is used by the sensing node to acquire the observation vector corresponding to the time slice at the current monitoring time point. The observation vector includes the response interference spectrum vector, the interference target statistical vector, and the self-state vector. The response interference spectrum vector includes the power corresponding to the interference in each sub-band. The response interference spectrum vector represents the interference situation of the sub-band at the frequency domain level. It should be noted that during the execution of communication functions, the UAV has a main operating band according to the configuration file, such as 3.4–3.6 GHz, which is the main frequency range for signal transmission during the UAV's communication process. For the convenience of scheduling and multiplexing more users, the main operating band will be further divided into smaller segments. Each segmented frequency range is called a sub-band or subcarrier. For example, it may be divided into 20 sub-bands. Each sub-band corresponds to a range of 10MHz. The interference target statistical vector refers to the position, distance, and power of the target corresponding to the interference source response after shortwave transmission. It should be noted that the targets responding to the interference source here include radar stations, communication base stations, and other drones that transmit signals themselves, as well as large buildings that strongly emit signals. If there is a strong interference target in a certain direction and at a certain distance, the interference in each sub-band in that direction will be higher. Therefore, the interference target statistical vector can reflect the existence of interference in the physical space. The self-state vector includes the drone's own position, speed, transmission power, and signal waveform, etc. The self-state vector represents the state of the drone's collected observation vector. And the time between adjacent monitoring time points is a time slice.

[0053] The interference time series analysis module is used to combine the currently acquired observation vector with the observation vectors acquired at the previous N-1 monitoring time points to form an observation time series vector, where N is a preset time series window length to cover the evolution of the electromagnetic environment and target distribution within a finite time period. The observation time series vector is then fed into the interference time series analysis model for processing, outputting a predicted interference distribution vector. This interference time series analysis model can be any one of the following: Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), Gated Recurrent Unit (GRU), or Temporal Convolutional Network (TCN). It is used to model the temporal correlation and coupling relationship between multi-source features in the observation time series vector, predicting the interference intensity of all sub-bands in the main working band. The predicted interference intensity is generally described by the average interference power. It should be noted that the interference time series analysis model is trained using a training sample consisting of N observation vectors arranged chronologically. The predicted output target and the last observation vector in the training sample are compared using the mean squared error (MSE) to calculate the loss value, and training is performed based on this loss value.

[0054] The subband allocation module is used to send the predicted interference distribution vector into the subband allocation network for processing and output the subband power pre-allocation vector. The subband power pre-allocation vector includes the allocated power of all subbands in the main working band. It should be noted that the allocated power refers to the signal transmission strength of the current sensing node on the corresponding subband. Generally, a high allocated power indicates that the current subband is less affected by interference and can perform critical data transmission services, while a low allocated power is used for the transmission of low-speed control signals. The subband power pre-allocation vector is sent to all neighboring nodes corresponding to the current sensing node. At the same time, the current sensing node receives the subband power pre-allocation vector from the neighboring nodes and updates the neighbor table of the current sensing node. The neighbor table includes one-to-one location information and subband power pre-allocation vector. The location information here can be stored in GNSS coordinates.

[0055] The interference correction module is used to perform interference correction on the subband allocation network based on the neighbor table corresponding to the sensing node. Then, the predicted interference distribution vector is sent to the interference-corrected subband allocation network for processing, and the output subband power allocation vector is used to configure the signal transmission of the sensing node in the next time slice. It should be noted that interference correction refers to adjusting the subband allocation network by taking into account the interference of the currently selected subband power pre-allocation vector to neighbor nodes. This allows the subband power allocation vector output by the interference-corrected subband allocation network to avoid strong external interference, spontaneously avoid high neighbor occupancy, and appropriately give way to heavily loaded neighbor nodes.

[0056] The reinforcement learning module is used to perform reinforcement learning operations on the subband allocation network using the collected call quality values ​​before outputting the subband power pre-allocation vector at the next monitoring time point. This enables reinforcement updates to the subband allocation network. It should be noted that by performing reinforcement learning operations on the subband allocation network, the network will not become overly involved in avoiding interference from neighboring nodes during the subband power allocation process. Furthermore, the iterative game between interference correction and reinforcement learning allows all nodes to gradually converge to a spectrum division state with less interference and higher resource sharing efficiency, thereby achieving distributed spectrum coordination without central control.

[0057] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims. Parts not described in detail in this specification are prior art known to those skilled in the art.

Claims

1. A distributed communication method for low-altitude networks, characterized in that, include: The sensing node collects the observation vector corresponding to the time slice at the current monitoring time point. The observation vector includes the response interference spectrum vector, the interference target statistical vector, and its own state vector. The observation time series vector is composed of the currently acquired observation vector and the observation vectors acquired at the previous N-1 monitoring time points, where N is the preset time series window length. The observation time series vector is then fed into the interference time series analysis model for processing, and the predicted interference distribution vector is output. The predicted interference distribution vector includes the predicted interference intensity of all sub-bands in the main working band. The predicted interference distribution vector is fed into the subband allocation network for processing, and the output subband power pre-allocation vector includes the allocated power of all subbands in the main working band. Based on the neighbor table corresponding to the sensing node, interference correction is performed on the sub-band allocation network. Then, the predicted interference distribution vector is sent to the interference-corrected sub-band allocation network for processing, and the sub-band power allocation vector is output. The signal transmission of the sensing node in the next time slice is configured through the sub-band power allocation vector. Before outputting the subband power pre-allocation vector at the next monitoring time point, the subband allocation network is reinforced and updated by using the collected call quality values.

2. The distributed communication method for low-altitude networks according to claim 1, characterized in that, Based on the neighbor table corresponding to the sensing node, interference correction is performed on the sub-band allocation network, specifically including the following steps: The predicted interference distribution vector is fed into the target subband allocation network for processing, outputting a target subband power pre-allocation vector. The predicted interference distribution vector and the target subband power pre-allocation vector are then concatenated and fed into the target allocation strategy evaluation network for processing, outputting a target strategy evaluation value. This evaluation value represents the proportion of subbands with weak interference selected by the subband allocation network. Based on the neighbor table corresponding to the sensing node and the subband power pre-allocation vector, the neighbor non-interference degree is calculated. The sum of the target strategy evaluation value and the non-interference degree is used as the gradient value of the target allocation strategy evaluation. The parameters within the allocation strategy evaluation network are iteratively updated using the gradient ascent method, maximizing the gradient value, until convergence. Finally, the predicted interference distribution vector and the subband power pre-allocation vector are concatenated and fed into the updated allocation strategy evaluation network. The network is processed, and the gradient value of the allocation strategy evaluation network for the subband power pre-allocation vector is calculated and denoted as the subband allocation gradient value. The node type corresponding to the synesthetic node is obtained. The node type includes aggressive and conservative. The upper limit of the gradient norm corresponding to the node type is determined according to the node type. The upper limit of the gradient norm corresponding to the aggressive type is higher than the upper limit of the gradient norm corresponding to the conservative type. It is determined whether the gradient norm corresponding to the subband allocation gradient value is higher than the upper limit of the gradient norm. If so, the parameters of the subband allocation network are iteratively updated using the gradient ascent method based on the subband allocation gradient value until convergence is completed, and the interference correction is completed. Otherwise, the subband allocation gradient value is scaled proportionally to the upper limit of the gradient norm. Then, the parameters of the subband allocation network are iteratively updated using the gradient ascent method based on the scaled subband allocation gradient value until convergence is completed, and the interference correction is completed.

3. The distributed communication method for low-altitude networks according to claim 2, characterized in that, The neighbor non-interference degree is calculated based on the neighbor table corresponding to the sensing node and the sub-band power pre-allocation vector, specifically including the following steps: Traverse all subbands in the main working band. For each subband, determine whether the sum of all allocated power in the subband's neighbor table and the subband power pre-allocation vector is higher than the rated allocated power. If the sum of all allocated power in the subband's neighbor table and the subband power pre-allocation vector is higher than the rated allocated power, mark the subband as an interfering subband. If the sum of all allocated power in the subband's neighbor table and the subband power pre-allocation vector is not higher than the rated allocated power, mark the subband as a low-interference subband. Calculate the ratio of the number of all low-interference subbands to the number of all subbands in the main working band, and record it as the neighbor non-interference degree.

4. A distributed communication method for low-altitude networks according to claim 3, characterized in that, The subband allocation network is subjected to reinforcement learning operations based on the collected call quality values, specifically including the following steps: The predicted interference distribution vector is fed into the target subband allocation network for processing, outputting a target subband power pre-allocation vector. The predicted interference distribution vector and the target subband power pre-allocation vector are then concatenated and fed into the target allocation policy evaluation network for processing, outputting a target policy evaluation value. Call quality values ​​are then obtained, and the sum of the target policy evaluation value and the call quality value is used as the gradient value for reinforcing the target allocation policy evaluation. The parameters within the allocation policy evaluation network are iteratively updated using the gradient ascent method in the direction that maximizes the gradient value, until convergence. The predicted interference distribution vector and the subband power pre-allocation vector are then concatenated and fed into the updated allocation policy evaluation network for processing, and the gradient value of the allocation policy evaluation network with respect to the subband power pre-allocation vector is calculated and denoted as the subband allocation gradient value. Based on the subband allocation gradient value, the parameters of the subband allocation network are iteratively updated using the gradient ascent method until convergence, completing interference correction and reinforcement learning operations. The target subband allocation network and the target allocation policy evaluation network are updated periodically.

5. A distributed communication method for low-altitude networks according to claim 4, characterized in that, The target subband allocation network and the target allocation policy evaluation network are updated periodically, specifically including the following steps: the product of the parameters and the learning rate in the subband allocation network that has completed reinforcement learning operations is directly added to the corresponding parameters in the target subband allocation network to complete the update of the target subband allocation network; the product of the parameters and the learning rate in the allocation policy evaluation network that has completed reinforcement learning operations is directly added to the corresponding parameters in the target allocation policy evaluation network to complete the update of the target allocation policy evaluation network.

6. A distributed communication method for low-altitude networks according to claim 5, characterized in that, Training the subband allocation network involves the following steps: obtaining several subband allocation training samples, each including an interference distribution vector; labeling the subband allocation training samples using the subband power allocation vector; forming a subband allocation training set from all labeled subband allocation training samples; and training the subband allocation network using the subband allocation training set.

7. A distributed communication method for low-altitude networks according to claim 6, characterized in that, Training the allocation policy evaluation network specifically includes the following steps: obtaining several allocation policy evaluation training samples, which include the interference distribution vector in the subband allocation training set and the corresponding labeled subband power allocation vector; labeling the allocation policy evaluation training samples with allocation policy evaluation values; forming an allocation policy evaluation training set with all labeled allocation policy evaluation training samples; and training the allocation policy evaluation network with the allocation policy evaluation training set.

8. A distributed communication system for low-altitude networks, characterized in that, The system employs a distributed communication method for low-altitude networks as described in any one of claims 1-7, comprising: The observation vector acquisition module is used by the sensing node to acquire the observation vector corresponding to the time slice at the current monitoring time point. The observation vector includes the response interference spectrum vector, the interference target statistical vector, and its own state vector. The interference time series analysis module is used to combine the currently acquired observation vector and the observation vectors acquired at the previous N-1 monitoring time points to form an observation time series vector, where N is the preset time series window length. The observation time series vector is then fed into the interference time series analysis model for processing, and the predicted interference distribution vector is output. The predicted interference distribution vector includes the predicted interference intensity of all sub-bands in the main working band. The subband allocation module is used to feed the predicted interference distribution vector into the subband allocation network for processing and output the subband power pre-allocation vector, which includes the allocated power of all subbands in the main working band. The interference correction module is used to perform interference correction on the sub-band allocation network based on the neighbor table corresponding to the sensing node, and then send the predicted interference distribution vector into the interference-corrected sub-band allocation network for processing, outputting the sub-band power allocation vector, and configuring the signal transmission of the sensing node in the next time slice through the sub-band power allocation vector. The reinforcement learning module is used to perform reinforcement learning operations on the subband allocation network using the collected call quality values ​​before outputting the subband power pre-allocation vector at the next monitoring time point, thereby achieving reinforcement updates to the subband allocation network.

Citation Information

Patent Citations

  • Dynamic rate matching pattern for spectrum sharing

    CN117397196A

  • Unmanned aerial vehicle cluster spectrum resource optimization method and related device

    CN120456276A