A hierarchical decision-making method for air-ground collaborative communication based on environmental cognition
Through a hierarchical decision-making method for air-ground collaborative communication based on environmental cognition, the multi-node communication of the air-ground collaborative network is optimized using graph convolutional networks and Markov decision processes, which solves the problem of limited transmission rate in complex urban environments and achieves efficient multi-node transmission rate and decision-making efficiency.
Patent Information
- Application Number
- CN202411059865.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-04
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-08-04
AI Technical Summary
In complex urban environments, traditional air-ground collaborative communication methods have difficulty in effectively utilizing spectrum resources, have limited transmission rates, and traditional reinforcement learning methods cannot converge quickly and cannot achieve efficient decision-making in multidimensional decision-making problems.
A hierarchical decision-making method for air-ground collaborative communication based on environmental cognition is adopted. Graph convolutional networks are used for node clustering. Combined with the Markov decision process and the actor-critic framework, the multi-node communication decision of the air-ground collaborative network is optimized. The hierarchical decision problem consists of the ground node cluster and UAV association decision layer, the UAV trajectory planning layer, and the ground node communication decision layer.
It significantly improves the transmission rate of multiple nodes in the air-ground collaborative communication network, effectively combats dynamic interference, and improves decision-making efficiency and transmission performance.
Smart Images

Figure CN119071817B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless communications, and in particular to a hierarchical decision-making method for air-ground collaborative communications based on environmental cognition. Background Art
[0002] The deployment of 6G networks faces numerous challenges, particularly in spectrum resource allocation. The growing demand for high-throughput applications and large-scale device access makes it crucial to effectively address congestion and establish intelligent connections between user devices. However, in complex urban environments, terrestrial channels are subject to building obstruction and small-scale fading, limiting transmission rates. Drones can assist in offloading tasks from ground nodes, alleviating the challenges of insufficient computing power, while also enabling dynamic network deployment and reducing resource inequality. Furthermore, utilizing line-of-sight links between drones and ground nodes, the higher channel quality and faster transmission rates can effectively reduce data transmission latency and power consumption. In the context of air-ground collaborative communications, scattered buildings create a complex electromagnetic environment, which can impact communication signal transmission but also serve as a barrier against malicious interference. Currently, there is a lack of air-ground collaborative communication methods that leverage environmental awareness, comprehensively utilize spatial and frequency domain resources in complex urban environments, and effectively combat dynamic interference.
[0003] In a complex and rapidly changing electromagnetic environment, since each node continuously transmits data throughout a time slot, communication decisions made within that time slot cannot capture information about the packet transmission volume generated in that time slot. Furthermore, due to the significant overhead of decision-making and data preparation, offloading scheduling decisions for processing / offloading these task queues can only be executed in the next time slot. Therefore, the joint learning strategy for air-ground collaborative communication decisions and UAV trajectory planning is formulated as a sequential decision problem, modeled as a Markov decision process. Traditional reinforcement learning methods, due to their sensitivity to hyperparameters, reliance on a single strategy, and limitations in solving multidimensional decision-making problems, struggle to achieve the fast convergence and multidimensional decision-making required for UAVs in complex urban environments. Furthermore, in large-scale multi-agent reinforcement learning, agents do not need to observe the states or actions of all participants; they only need to observe the states or actions of a subset of relevant agents in their neighborhood. Therefore, clustering agents with the same tasks in complex scenarios can reduce storage capacity and computational overhead, improving decision-making efficiency. However, in the context of air-ground collaborative communication, traditional clustering algorithms have high computational complexity and cannot take into account the complex and changeable electromagnetic environment in the city and the time-varying nature of cluster network topology. In addition, using only single features such as distance and density to cluster nodes is obviously insufficient. Summary of the Invention
[0004] The purpose of the present invention is to provide a hierarchical decision-making method for air-ground collaborative communication based on environmental cognition, which can effectively extract the characteristics of transmission performance from complex urban environments and significantly improve the transmission rate of multiple nodes in the air-ground collaborative communication network.
[0005] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:
[0006] A hierarchical decision-making method for air-ground collaborative communication based on environmental cognition, the method comprising the following steps:
[0007] Step 1: Set a complex urban environment as the scenario. In order to prevent malicious or unintentional dynamic interference in the scenario, drones are used to assist communication between multiple ground nodes. The data transmission tasks generated by each ground node use any of the following communication methods: no communication, direct communication between ground nodes, or drone-assisted communication.
[0008] Step 2: Build a graph convolutional network model and use it to aggregate multiple environmental features for node clustering. Based on environmental cognition, it comprehensively considers various environmental features, including building occlusion, node location, node task queue, and node usage frequency in the air-ground collaborative communication network, and optimizes clustering of multiple nodes.
[0009] Step 3: Based on the complexity of the task and the granularity of the decision, the UAV-assisted air-ground collaborative communication is divided into three subtasks: the ground node cluster and UAV association decision layer, the UAV trajectory planning layer, and the ground node communication decision layer. The complex decision-making problem of air-ground collaborative communication in this scenario is constructed; the complex decision-making problem of air-ground collaborative communication is modeled as a Markov decision process, and a hierarchical decision-making algorithm based on the actor-critic framework is used to optimize the communication decisions of multiple nodes in the air-ground collaborative network; until the hierarchical decision-making algorithm for air-ground collaborative communication based on environmental cognition converges, at which point the action selection decision is the optimal decision.
[0010] Step 1 further includes the following steps:
[0011] Set a complex urban environment as the scene, and the scene includes M drones with single antenna connections U = {U1, U2, ..., U m},m=1,2,...,M, and N ground nodes u={u1,u2,...,u n M drones are used to assist N ground nodes in performing air-ground collaborative communication in a dispersed building complex to deal with interference signals. The ground nodes are divided into K ground node clusters. The drones always fly at the same altitude H, and the ground nodes are on the horizontal plane and fixed in position.
[0012] When a ground node has a data transmission task, in each time slot, the data transmission task generated by each ground node adopts any of the following communication modes: no communication, direct communication between ground nodes, and drone-assisted communication. In addition, the data transmission task generated by a single ground node in a single time slot can be processed by only one drone.
[0013] The communication channel between the UAV and the ground node is described by the logarithmic range loss model, which includes a constant attenuation factor in small-scale fading and obstacle occlusion; the information rate of the nth,n∈{1,2,...,N}th ground node at a constant position on the ground is given by: Among them, the transmission power P, noise power N and path loss γ of the nth ground node are n ; The information rate when the UAV forwards the task queue is where γ m represents the path loss when the UAV forwards the task queue; when the path loss exponent in vacuum is set to α, the path loss of node n is Here, small-scale fading is modeled as a Rayleigh distributed random variable X with a scaling factor σ = 1 Rayleigh The attenuation due to the obstruction is modeled by a discrete factor β∈{1,0.01}, where β = 0.01 if there is a building obstruction and β = 1 otherwise. The path loss of the signal transmission between ground nodes is calculated as follows: Among them, P r Indicates the signal receiving power, P t Indicates the signal transmission power, G r Indicates the gain of the receiving antenna, G t represents the gain of the transmitting antenna, L represents the system loss, d represents the distance between the receiving end and the transmitting end, λ represents the wavelength of the signal, and γ represents the path loss exponent;
[0014] Assume that all ground nodes use the frequency band F = {f1,f2,...,f i},i∈{1,2,...,I}, there are I frequency points to choose from; suppose there is a jammer in the scene that transmits interference signals at different frequencies in each time slot, interfering with the data transmission between ground nodes, and the interference signal frequency of the jammer is repeated at the time slot period F j ={f1,f2,...,f j}, j∈{1,2,...,J}; When a ground node selects a frequency point for communication in each time slot, if the selected frequency point is the same as the frequency point of the interference signal sent by the jammer in the current time slot, it is considered that the current ground node communication is interfered with, and if multiple ground nodes select the same frequency point in the same time slot, it is considered that there is communication interference between these ground nodes; when interference and mutual interference occur, the transmission rate is calculated by the following formula Among them, G j is the channel gain from the interference signal to the interfered node, P j is the interference signal power, G i is the channel gain of the mutual interference signal, A is the set of mutual interference numbers of ground nodes, γ n is the path loss of node n, P is the transmission power of the nth ground node, and N is the noise power;
[0015] At the beginning of each time slot, the amount of data transmission randomly generated by each ground node is Δt. Each ground node has a data accumulation queue Δn, and each UAV has a data accumulation queue Δm. It is assumed that half of each time slot τ is used for transmission and half for reception. In addition, each ground node can only transmit data to one ground node per time slot, and the ground node can only receive data transmitted by the corresponding ground node.
[0016] At the beginning of the time slot, since the amount of data randomly generated by the ground node is Δt, the cumulative amount of task data to be transmitted by the ground node is Δ n =Δ n +Δ t , when the ground node chooses to use drones to assist in communication, if the transmission can be completed, that is, Δ t ≤R n At (t)·τ / 2, the amount of data received by the UAV and forwarded, the current amount of data received by the uplink UAV queue is: Δ m ≤ m +Δ t , the ground node queue minus the current data volume: Δ n =Δ n -Δ t ; Current data volume forwarded in the downlink UAV queue: Δ m =Δ m -Δ t ; If the transmission cannot be completed, Δ t >R n At (t)·τ / 2, the maximum amount of data that the UAV can currently receive in the uplink is: Δ m =Δ m +R n (t)·τ / 2, the maximum amount of data that can be sent by the ground node queue minus: Δ n =Δ n -R n (t)·τ / 2; the maximum amount of data that can be forwarded by the UAV in the downlink is: Δ m =Δ m -R m (t)·τ / 2, where R n (t) and R m(t) represents the information rate of the ground node and the UAV respectively; the unsent data is sent in the next time slot; when the ground node chooses direct communication, if the transmission can be completed, that is, Δ t ≤R n (t)·τ / 2, the ground node queue minus the current data volume: Δ n =Δ n -Δ t , if the transmission cannot be completed, that is, Δ t >R n At (t)·τ / 2, the maximum amount of data that can be sent by the ground node queue is subtracted: Δ n =Δ n -R n (t)·τ / 2; the unsent data is sent in the next time slot.
[0017] Step 3 further includes:
[0018] In the air-ground collaborative communication scenario, the joint learning strategy of air-ground collaborative communication decision-making and UAV trajectory planning is reduced to a sequential decision-making problem. A semi-Markov decision process is used to model the hierarchical decision-making problem in air-ground collaborative communication, and problem planning and strategy learning are performed in different time slots. Specifically:
[0019] Different tasks are modeled as semi-Markov decision processes with different time scales. A hierarchical structure is constructed through task relationships. This transforms the multi-task and multi-granularity complex decision-making problem in the air-ground collaborative communication scenario into a top-down three-layer hierarchical decision-making process: the ground node cluster and UAV association decision layer, the UAV trajectory planning layer, and the ground node communication decision layer.
[0020] Define a time slot interval Δ and deploy agents for node clustering, trajectory planning, and data transmission on drones and ground nodes. At the beginning of each Δ time slot, the ground nodes are randomly clustered, and drones are selected as auxiliary communication nodes for the ground node cluster. Each drone plans its trajectory for the next Δ time slots based on its state, and the ground nodes perform data transmission within each time slot to minimize the amount of data packets accumulated by each node.
[0021] Each ground node re-formulates its clustering decision every Δ time slots. Each ground node performs data transmission in each time slot t based on the latest cluster association decision issued by the ground node cluster. Once the UAV's trajectory plan is determined, its flight direction remains unchanged in the following Δ time slots until a new round of trajectory planning is generated.
[0022] In the air-ground collaborative communication scenario, a hierarchical decision-making algorithm based on the actor-critic framework uses two fully connected neural networks to approximate the policy function and the state-value function. The actor, composed of the policy network, participates in the generation of the agent's actions, and the critic evaluates the actions taken by the actor by using a parameterized state-value function.
[0023] At the end of Δ time slots, the temporal difference error (TD error) is calculated based on the accumulated rewards and the output of the value network. The TD mean square error is minimized to guide the update of the policy network and its own value network parameters.
[0024] Repeat the above steps until the algorithm converges, and take the action selection decision at this time as the optimal decision.
[0025] Furthermore, the UAV trajectory planning layer as an intermediate layer inherits the top-level ground node cluster and UAV associated cluster information a k,t , the ground node communication decision layer receives information from the top and middle layers as the bottom layer k,t and a m,t , for the ground node agent n,a n,t Determine the communication decisions of ground nodes;
[0026] In the decision layer of ground node cluster and UAV association, state s k,t For s k,t =[p1,p2,...,p M ,a k,t-1 ], where p1, p2, ..., p M Represents the positions of M drones, a k,t-1 represents the action decision of the kth ground node cluster in the time slot; the reward r obtained by ground node cluster k k,t+Δ Defined as the accumulation of queues between all ground nodes and drones in the cluster
[0027] In the UAV trajectory planning layer, the state s m,t For s m,t =[p m ,q m ,a k,t-1 ], where p m and q m Represents the location of UAV m and the accumulated task queue information, a k,t-1 represents the action decision of the kth ground node cluster in the time slot; the reward r obtained by drone m m,t+Δ It is defined as the queue accumulation between all ground nodes in the cluster associated with the UAV
[0028] In the communication decision layer of the ground node, the state s k,n,t For s k,n,t =[p m ,p n ,a n,t-1 ,f n ,q n ], where a n,t-1 represents the action decision of the time slot on the ground node n, p n 、f n ,q n They represent the location, communication frequency, and accumulated task queue information of ground node n respectively; the reward obtained by ground node n is defined as the accumulated task queue amount r of ground node n k,n,t =q n .
[0029] Furthermore, the process of optimizing the communication decision of multiple nodes in the air-ground collaborative network using the hierarchical decision algorithm based on the actor-evaluator framework includes the following steps:
[0030] Define the cluster association strategy network parameters θ of ground node cluster k k and cluster association value network parameter w k , the association decision between the ground node cluster and the UAV cluster in time slot t is expressed as a k,t =π(·|s k,t θ k ), where in state s k,t According to the cluster association strategy network parameter θ k The output result is selected with a certain probability k,t ,satisfy The decision to associate the ground node cluster with the drone cluster ends after Δ time slots, based on the accumulated reward r k,t+Δ And the output of the cluster association value network Calculating TD Error Where η is the loss coefficient; the cluster association value network parameters are updated with the goal of minimizing the TD mean square error, and the loss function is defined as The cluster association value network along the loss function L(w k ) The update formula of the gradient direction update parameter is expressed as Where α is the weight factor, the cluster association strategy function network is guided by the TD error along the loss function J(θ k ) updates the cluster association network parameters in the gradient direction, where the loss function is defined as J(θ k )=logπ(s k,t θ k )·δ k,t , the update formula of the cluster association strategy network parameters in each Δ time slot is expressed as
[0031] Define the trajectory planning strategy network parameters θ of UAV m m and trajectory planning value network parameter w m , the trajectory planning decision in time slot t is represented as a m,t =π(·|s m,t θ m ); When the trajectory planning decision ends after Δ time slots, the cumulative reward r obtained by the drone m,t+Δ And the output V(s) of the trajectory planning value network m,t ;w m ), calculate the TD error δ of UAV m m,t =r m,t+Δ +ηV(s m,t+Δ ;w m )-V(s m,t ;w m ); define the communication decision strategy network parameters θ of the ground node n in each cluster k,n and communication decision value network parameter w k,n , the communication decision action of the ground node is represented as a k,n,t =π(·|s k,n,t θ k,n ); The communication decision action of the ground node is made in each time slot, and the reward r is obtained at the end of the data transmission action. k,n,t , TD error is expressed as δ k,n,t =r k,n,t +ηV(s k,n,t+1 ;w k,n )-V(s k,n,t ;w k,n ).
[0032] Furthermore, step 2 includes the following sub-steps:
[0033] Step 2.1: Abstract all nodes into a graph model representation G = (V, E), where V represents the node set and E represents the edge set;
[0034] Step 2.2, construct the adjacency matrix A and the feature matrix X; the dimension of the adjacency matrix A is N×N, which is used to reflect the communication relationship between nodes. If there is a communication relationship between node i and node j, i, j∈{1,2,…,N}, then A ij =1, otherwise A ij =0, the dimension of the feature matrix X is N×F, where F represents the number of classes of node features;
[0035] Step 2.3: Build a graph convolutional model. The number of neurons in the input layer is equal to the number of columns in the feature matrix X. The number of neurons in the output layer is determined by the number of clusters in which the nodes are clustered. The Glorot method is used to initialize the uniformly distributed weight matrix W.
[0036] Step 2.4: pre-process the adjacency matrix A and feature matrix X and input them into the graph convolutional network, and obtain the prediction matrix Y through forward propagation. in and Represent the preprocessed adjacency matrix A and feature matrix X, W respectively (1) and W (2) Respectively represent the weight matrices of the first and second layers in the graph convolutional network, and ReLU and softmax represent two selected activation functions;
[0037] Step 2.5, calculate the cross entropy loss function between the prediction matrix and the label matrix output by the graph convolutional network to update the weight matrix parameters. Where N is the number of nodes, K is the number of clusters in which the nodes are clustered, and Y ij and P ij Represent the prediction matrix and label matrix respectively;
[0038] Step 2.6: Repeat steps 2.4-2.5 until the algorithm converges. If the prediction accuracy of the graph convolutional network meets the requirements, apply the pre-trained weight matrix W to the hierarchical decision-making algorithm based on the actor-critic framework to extract multiple environmental features to understand the complex urban environment.
[0039] Furthermore, in the decision layer for associating ground node clusters with UAVs, the adjacency matrix A reflects the queue transmission relationship between the ground node clusters and UAVs, and the feature matrix X includes the positions of all ground nodes and UAVs as well as the cumulative queue volume. In the decision layer for communicating with ground nodes, the adjacency matrix A reflects the queue transmission relationship between ground nodes, and the feature matrix X includes the positions of ground nodes and whether frequency interference occurs between them.
[0040] Furthermore, in step 2.4, the calculation formula of the preprocessing adjacency matrix A is Where D represents the degree matrix of the adjacency matrix A, I represents the identity matrix; the calculation formula of the preprocessing feature matrix X is in and δ X Represent the mean and variance of the feature matrix X respectively;
[0041] In the graph convolutional network, the forward propagation from hidden layer l to hidden layer l+1 is expressed as Where σ represents the ReLU activation function, H l and W l denote the output matrix and weight matrix of hidden layer l respectively.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] The present invention proposes a hierarchical decision-making method for air-ground collaborative communication based on environmental cognition, which effectively extracts the characteristics of transmission performance from complex urban environments and significantly improves the transmission rate of multiple nodes in the air-ground collaborative communication network. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a flow chart of the hierarchical decision-making method for air-ground collaborative communication based on environmental cognition of the present invention.
[0045] Figure 2 This is a model diagram of the air-ground collaborative communication hierarchical decision-making system based on environmental cognition of the present invention.
[0046] Figure 3 This is a specific operational flow chart of the air-ground collaborative communication hierarchical decision-making method based on environmental cognition of the present invention.
[0047] Figure 4 This is a specific operation flow chart of the pre-trained graph convolutional network of the present invention. DETAILED DESCRIPTION
[0048] The embodiments of the present invention are described in further detail below with reference to the accompanying drawings.
[0049] A hierarchical decision-making method for air-ground collaborative communication based on environmental cognition, characterized in that the hierarchical decision-making method for air-ground collaborative communication comprises the following steps:
[0050] Step 1: Set a complex urban environment as the scenario. In order to prevent malicious or unintentional dynamic interference in the scenario, drones are used to assist communication between multiple ground nodes. The data transmission tasks generated by each ground node use any of the following communication methods: no communication, direct communication between ground nodes, or drone-assisted communication.
[0051] Step 2: Build a graph convolutional network model and use it to aggregate multiple environmental features for node clustering. Based on environmental cognition, it comprehensively considers various environmental features, including building occlusion, node location, node task queue, and node frequency in the air-ground collaborative communication network, and optimizes clustering of multiple nodes.
[0052] Step 3: Based on the complexity of the task and the granularity of the decision, the UAV-assisted air-ground collaborative communication is divided into three subtasks: the ground node cluster and UAV association decision layer, the UAV trajectory planning layer, and the ground node communication decision layer. The complex decision-making problem of air-ground collaborative communication in this scenario is constructed; the complex decision-making problem of air-ground collaborative communication is modeled as a Markov decision process, and a hierarchical decision-making algorithm based on the actor-critic framework is used to optimize the communication decisions of multiple nodes in the air-ground collaborative network; until the hierarchical decision-making algorithm for air-ground collaborative communication based on environmental cognition converges, at which point the action selection decision is the optimal decision.
[0053] This embodiment proposes a hierarchical decision-making method for air-ground collaborative communication based on environmental cognition. Its main steps, system model, and specific operations are as follows: Figure 1-4 As shown, specifically including the following:
[0054] Set a complex urban environment as the background, such as Figure 2 As shown, the scenario includes M drones with single antenna connections U = {U1, U2, ..., U m},m=1,2,...,M, and N ground nodes u={u1,u2,...,u n}, n = 1, 2, ..., N, equipped with the same single antenna, perform air-ground collaborative communication against interference signals in a dispersed building cluster. The ground nodes are divided into K ground node clusters. Assume that in this scenario, as the time slot changes, the UAV always flies at the same altitude H, and the ground nodes are on the horizontal plane and fixed in position.
[0055] When a ground node has a data transmission task (receiving or sending a data packet), each ground node can choose from three communication modes in each time slot: no communication, direct communication between ground nodes, or drone-assisted communication. Furthermore, a data transmission task generated by a single ground node in a single time slot can only be handled by one drone. Assuming that all ground nodes can establish a connection with only one drone per time slot, if ground node cluster k is associated with drone m, then any data transmission task generated by nodes within cluster k that chooses drone communication can only be transmitted via drone m. When a new cluster association decision is made, the ground node cluster sends this command to all ground nodes within the cluster. Further data transmission by these nodes will strictly adhere to the cluster association decision for that time slot, and ground nodes can choose to communicate with other ground nodes in each time slot.
[0056] The communication channel between the UAV and the ground node is described by a logarithmic range loss model that includes a constant attenuation factor for small-scale fading and obstacle occlusion. The information rate of the nth,n∈{1,2,...,N}th ground node located at a constant position on the ground is given by: Among them, the transmission power P, noise power N and path loss γ of the nth user n , Similarly, the information rate when the UAV forwards the task queue is where γ m It represents the path loss when the UAV forwards the task queue. When the path loss exponent in vacuum is set to α, the path loss of node n is as follows Here, small-scale fading is modeled as a Rayleigh distributed random variable X with a scaling factor σ = 1 RayleighThe attenuation due to the obstruction is modeled by a discrete factor β∈{1,0.01}, where β=0.01 if there is a building obstruction and β=1 otherwise. The path loss for signal transmission between ground nodes (ground-to-ground transmission) is calculated as Among them, P r Indicates the signal receiving power, P t Indicates the signal transmission power, G r Indicates the gain of the receiving antenna, G t represents the gain of the transmitting antenna, L represents the system loss, d represents the distance between the receiving end and the transmitting end, λ represents the wavelength of the signal, and γ represents the path loss exponent.
[0057] Assume that all ground nodes use the frequency band F = {f1,f2,...,f i},i∈{1,2,...,I}, there are I frequency points to choose from; suppose there is a jammer in the scene that transmits interference signals at different frequencies in each time slot, interfering with the data transmission between ground nodes, and the interference signal frequency of the jammer is repeated at the time slot period F j ={f1,f2,...,f j}, j∈{1,2,...,J}; When a ground node selects a frequency point for communication in each time slot, if the selected frequency point is the same as the frequency point of the interference signal sent by the jammer in the current time slot, it is considered that the current ground node communication is interfered with, and if multiple ground nodes select the same frequency point in the same time slot, it is considered that there is communication interference between these ground nodes; when interference and mutual interference occur, the transmission rate is calculated by the following formula Among them, G j is the channel gain from the interference signal to the interfered node, P j is the interference signal power, G i is the channel gain of the mutual interference signal, A is the set of mutual interference numbers of ground nodes, γ n is the path loss of node n, P is the transmission power of the nth ground node, and N is the noise power.
[0058] The amount of data transmission randomly generated by each ground node in each time slot is Δt. Each ground node has a data accumulation queue Δn, and each drone has a data accumulation queue Δm. It is assumed that half of each time slot, i.e. τ / 2, is used for transmission and half of the time is used for reception. Moreover, each ground node can only transmit data to one ground node in each time slot, and the ground node can only receive data transmitted by the corresponding ground node.
[0059] At the beginning of the time slot, since the amount of data randomly generated by the ground node is Δt, the cumulative amount of task data to be transmitted by the ground node is Δ n =Δ n +Δ t, when the ground node chooses to use drones to assist in communication, if the transmission can be completed, that is, Δ t ≤R n At (t)·τ / 2, the UAV receives the data volume and completes the forwarding, that is, the uplink UAV queue receives the current data volume, and the ground node queue subtracts the current data volume: Δ m =Δ m +Δ t , Δ n =Δ n -Δ t ; Current data volume forwarded in the downlink UAV queue: Δ m =Δ m -Δ t ; If the transmission cannot be completed, Δ t >R n At (t)·τ / 2, the maximum amount of data that the UAV can currently receive in the uplink is subtracted from the maximum amount of data that the ground node queue can currently send: Δ m =Δ m +R n (t)·τ / 2,Δ n =Δ n -R n (t)·τ / 2; the maximum amount of data that can be forwarded by the UAV in the downlink is: Δ m =Δ m -R m (t)·τ / 2, where R n (t) and R m (t) represents the information rate of the ground node and the UAV respectively; the unsent data is sent in the next time slot. When the ground node chooses direct communication, if the transmission can be completed, that is, Δ t ≤R n (t)·τ / 2, the ground node queue minus the current data volume: Δ n =Δ n -Δ t , if the transmission cannot be completed, that is, Δ t >R n At (t)·τ / 2, the maximum amount of data that can be sent by the ground node queue is subtracted: Δ n =Δ n -R n (t)·τ / 2; the unsent data is sent in the next time slot.
[0060] In air-ground collaborative communication scenarios, since communication decisions made within a time slot cannot obtain information about the amount of data packets transmitted in that time slot, and since data transmission overhead is non-negligible, the joint learning strategy for air-ground collaborative communication decisions and UAV trajectory planning is reduced to a sequential decision problem and modeled as a Markov decision process. Furthermore, ground node communication decisions and UAV trajectory planning decisions have different time scales. Simply modeling multiple decision problems as a single Markov decision process cannot address the complex decision-making that spans multiple time slots. Therefore, a semi-Markov decision process (SMDP) is used to model the hierarchical decision problem in air-ground collaborative communication, enabling problem planning and policy learning across different time slots. The decision-making scheme in the SMDP can persist across multiple time steps to simulate decision-making behavior with multiple granularities. This allows the complex decision-making problem in air-ground collaborative communication scenarios to be divided into a small-slot ground node communication decision problem and a large-slot cluster association decision problem, enabling problem planning and policy learning across different time slots. For multi-task and multi-granularity complex decision-making problems, different tasks can be modeled as SMDPs with different time scales. A hierarchical structure can be constructed through task relationships. The hierarchical structure allows complex decision-making problems to be further subdivided into multiple layers of sub-problems for separate decision-making. The multi-task and multi-granularity complex decision-making problems in the air-ground collaborative communication scenario are transformed into a top-down three-layer hierarchical decision-making, namely, the ground node cluster and UAV association decision layer, the UAV trajectory planning layer, and the ground node communication decision layer.
[0061] A time slot interval Δ is defined, and agents responsible for node clustering, trajectory planning, and data transmission are deployed on both drones and ground nodes. At the beginning of each Δ time slot, each ground node is randomly clustered, and a drone is selected as an auxiliary communication node for that ground node cluster. Each drone plans its trajectory for the next Δ time slots based on its state, and the ground nodes transmit data within each time slot to minimize the amount of data packets accumulated by each node. Each ground node re-formulates its clustering decision every Δ time slots, and within each time slot t, each ground node transmits data based on the latest cluster association decision issued by the ground node cluster. Once a drone's trajectory plan is finalized, its flight direction remains unchanged for the next Δ time slots until a new trajectory plan is generated.
[0062] The UAV trajectory planning layer, as the middle layer, inherits the top-level ground node cluster and UAV associated cluster information. k,t , the ground node communication decision layer receives information from the top and middle layers as the bottom layer k,t and a m,t , for the ground node agent n,a n,t Determines the communication decision of the ground node. In the decision layer of the ground node cluster and the UAV, the state s k,t For sk,t =[p1,p2,...,p M ,a k,t-1 ], where p1, p2, ..., p M Represents the positions of M drones, a k,t-1 represents the action decision of the kth ground node cluster in the time slot; the reward r obtained by ground node cluster k k,t+Δ It can be defined as the accumulation of queues between all ground nodes and UAVs in the cluster, that is, In the UAV trajectory planning layer, the state s m,t For s m,t =[p m ,q m ,a k,t-1 ], where p m and q m Represents the location of UAV m and the accumulated task queue information, a k,t-1 represents the action decision of the kth ground node cluster in the time slot; the reward r obtained by drone m m,t+Δ It can be defined as the queue accumulation between the UAV and all ground nodes in the cluster, i.e. In the communication decision layer of the ground node, the state s k,n,t For s k,n,t =[p m ,p n ,a n,t-1 ,f n ,q n ], where p m represents the position of the UAV m, a n,t-1 represents the action decision of the time slot on the ground node n, p n 、f n ,q n They represent the location, communication frequency, and accumulated task queue information of ground node n respectively; the reward obtained by ground node n can be defined as the accumulated task queue amount of ground node n, that is, r k,n,t =q n .
[0063] In the air-ground collaborative communication scenario, Figure 3 As shown in the figure, based on the Actor-Critic (AC) framework, two fully connected neural networks are used to approximate the policy function π(s; θ) and the state-value function V(s; w). The actor, composed of the policy network π(s; θ), participates in the generation of the agent's actions, while the critic evaluates the actions taken by the actor using the parameterized state-value function V(s; w). For the ground node communication decisions, drone trajectory planning, and the decision to associate ground node clusters with drones in the scene, the action space is discrete, so a softmax strategy is adopted.
[0064] Define the policy network parameters θ of ground node cluster k k and the value network parameter w k , the association decision between the ground node cluster and the UAV cluster in time slot t can be expressed as a k,t =π(·|s k,t θ k ), where in state s k,t According to the policy network parameters θ k The output result is selected with a certain probability k,t ,satisfy The scheme ends after Δ time slots, based on the cumulative reward r k,t+Δ And the output of the value network V(s k,t ;w k ), calculate the TD error Where η is the loss coefficient. The value network parameters are updated with the goal of minimizing the TD mean square error. The loss function is defined as The update formula of the value network parameters can be expressed as Where α is the weight factor. Under the guidance of TD error, the policy function network follows the loss function J(θ k ) updates the network parameters in the gradient direction, where the loss function is defined as J(θ k )=logπ(s k,t θ k )·δ k,t The update formula of the policy network parameters in each Δ time slot can be expressed as
[0065] Similarly, define the policy network parameters θ of drone m m and the value network parameter w m , the trajectory planning decision in time slot t can be expressed as a m,t =π(·|s m,t θ m ). When the scheme ends after Δ time slots, the cumulative reward r obtained by the drone is m,t+Δ And the output of the value network V(s m,t ;w m ), calculate the TD error δ of UAV m m,t =r m,t+Δ +ηV(s m,t+Δ ;w m )-V(s m,t ;w m ). Define the communication decision strategy network parameters θ of the ground node n in each cluster k,n and the value network parameter w k,n The communication decision action of the ground node can be expressed as a k,n,t =π(·|sk,n,t θ k,n ). Different from the above two schemes, the communication decision action of the ground node needs to make a decision in each time slot and obtain the reward r at the end of the executed data transmission action. k,n,t , the TD error can be expressed as δ k,n,t =r k,n,t +ηV(s k,n,t+1 ;w k,n )-V(s k,n,t ;w k,n ). The updating method of the value network and policy network executed by drone m and ground node n is the same as that of the cluster association decision layer.
[0066] Repeat the above steps until the algorithm converges, at which point the action selection decision is the optimal decision.
[0067] Environmental cognition is introduced in the decision layer for association between ground node clusters and drones and the decision layer for communication between ground nodes. Based on environmental cognition, various environmental features such as building occlusion, node location, node task queue, and node frequency in the air-ground collaborative communication network are comprehensively considered. The ability of graph convolutional networks to aggregate various environmental features for node clustering is used to achieve multi-node clustering optimization. The process of pre-training graph convolutional networks is as follows: Figure 4 As shown in Figure 2, the output weight matrix W is the pre-trained graph convolutional network model.
[0068] First, all nodes are abstracted into a graph model representation G = (V, E), where V represents the node set and E represents the edge set. Construct the adjacency matrix A and the feature matrix X, where the dimension of the adjacency matrix A is N × N, reflecting the connection relationship between nodes. If there is a connection relationship between node i and node j, i, j∈{1,2,…,N}, then A ij =1, otherwise A ij =0, the dimension of the feature matrix X is N×F, where F represents the number of node feature classes. In the decision layer for the association between ground node clusters and UAVs, the adjacency matrix A reflects the queue transmission relationship between the ground node clusters and UAVs, and the feature matrix X includes the positions of all ground nodes and UAVs and the cumulative queue size. In the decision layer for the communication between ground nodes, the adjacency matrix A reflects the queue transmission relationship between ground nodes, and the feature matrix X includes the positions of ground nodes and whether frequency interference occurs between them.
[0069] Next, a graph convolution model is constructed. The number of neurons in the input layer is equal to the number of columns of the feature matrix X. The number of neurons in the output layer is determined by the number of clusters of the nodes. The Glorot method is used to initialize the uniformly distributed weight matrix W. The adjacency matrix A and the feature matrix X are preprocessed and input into the graph convolution network. The calculation formula of the preprocessed adjacency matrix A is: Where D represents the degree matrix of the adjacency matrix A, I represents the identity matrix; the calculation formula of the preprocessing feature matrix X is in and δ X Represent the mean and variance of the feature matrix X respectively.
[0070] Then, the prediction matrix Y is obtained through forward propagation, that is, in and Represent the preprocessed adjacency matrix A and feature matrix X, W respectively (1) and W (2) Represent the weight matrices of the first and second layers in the graph convolutional network, respectively. ReLU and softmax represent two selected activation functions. In the graph convolutional network, the specific calculation formula for the forward propagation from hidden layer l to hidden layer l+1 can be expressed as Where σ represents the ReLU activation function, H l and W l denote the output matrix and weight matrix of hidden layer l respectively.
[0071] Finally, the cross entropy loss function between the prediction matrix and the label matrix output by the graph convolutional network is calculated to update the weight matrix parameters, that is, Where N is the number of nodes, K is the number of clusters in which the nodes are clustered, and Y ij and P ij Denote the prediction matrix and label matrix, respectively. Repeat steps 3.5-3.6 until the algorithm converges. If the prediction accuracy of the graph convolutional network meets the requirements, the pre-trained weight matrix W can be applied to the hierarchical decision-making algorithm based on the actor-critic framework to extract multiple environmental features and realize the cognition of complex urban environments.
[0072] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal translation scripting language JavaScript, etc.
[0073] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0074] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0075] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions for executing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0076] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0077] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A hierarchical decision-making method for air-ground collaborative communication based on environmental cognition, characterized in that: The air-ground collaborative communication hierarchical decision-making method comprises the following steps: Step 1: Set a complex urban environment as the scenario. To address malicious or unintentional dynamic interference in the scenario, use drones to assist communication between multiple ground nodes. The data transmission tasks generated by each ground node use any of the following communication modes: no communication, direct communication between ground nodes, or drone-assisted communication. Step 2: Build a graph convolutional network model and use it to aggregate multiple environmental features for node clustering. Based on environmental cognition, it comprehensively considers various environmental features, including building occlusion, node location, node task queue, and node usage frequency in the air-ground collaborative communication network, and optimizes clustering of multiple nodes. Step 3: Based on the complexity of the task and the granularity of the decision, the UAV-assisted air-ground collaborative communication is divided into three subtasks: the ground node cluster and UAV association decision layer, the UAV trajectory planning layer, and the ground node communication decision layer. This constructs the complex decision-making problem of air-ground collaborative communication in this scenario. The complex decision-making problem of air-ground collaborative communication is modeled as a Markov decision process. A hierarchical decision-making algorithm based on the actor-critic framework is used to optimize the communication decisions of multiple nodes in the air-ground collaborative network. This is done until the hierarchical decision-making algorithm for air-ground collaborative communication based on environmental cognition converges, at which point the action selection decision is the optimal decision. Step 3 further includes: In the air-ground collaborative communication scenario, the joint learning strategy of air-ground collaborative communication decision-making and UAV trajectory planning is reduced to a sequential decision-making problem. A semi-Markov decision process is used to model the hierarchical decision-making problem in air-ground collaborative communication, and problem planning and strategy learning are performed in different time slots. Specifically: Different tasks are modeled as semi-Markov decision processes with different time scales. A hierarchical structure is constructed through task relationships. This transforms the multi-task and multi-granularity complex decision-making problem in the air-ground collaborative communication scenario into a top-down three-layer hierarchical decision-making process: the ground node cluster and UAV association decision layer, the UAV trajectory planning layer, and the ground node communication decision layer. Define a time slot interval Δ and deploy agents for node clustering, trajectory planning, and data transmission on drones and ground nodes. At the beginning of each Δ time slot, the ground nodes are randomly clustered, and drones are selected as auxiliary communication nodes for the ground node cluster. Each drone plans its trajectory for the next Δ time slots based on its state, and the ground nodes perform data transmission within each time slot to minimize the amount of data packets accumulated by each node. Each ground node re-formulates its clustering decision every Δ time slots. Each ground node performs data transmission in each time slot t based on the latest cluster association decision issued by the ground node cluster. Once the UAV's trajectory plan is determined, its flight direction remains unchanged in the following Δ time slots until a new round of trajectory planning is generated. In the air-ground collaborative communication scenario, a hierarchical decision-making algorithm based on the actor-critic framework uses two fully connected neural networks to approximate the policy function and the state-value function. The actor, composed of the policy network, participates in the generation of the agent's actions, and the critic evaluates the actions taken by the actor by using a parameterized state-value function. At the end of Δ time slots, the temporal difference error TD is calculated based on the accumulated rewards and the output of the value network. The goal of minimizing the TD mean square error is to guide the update of the policy network and its own value network parameters. Repeat the above steps until the algorithm converges, and take the action selection decision at this time as the optimal decision.
2. The hierarchical decision-making method for air-ground collaborative communication based on environmental cognition according to claim 1 is characterized in that: Step 1 further includes the following steps: Set a complex urban environment as the scene, and the scene includes M drones with single antenna connections U = {U1, U2, ..., U m },m=1,2,...,M, and N ground nodes u={u1,u2,...,u n M drones are used to assist N ground nodes in performing air-ground collaborative communication in a dispersed building complex to deal with interference signals. The ground nodes are divided into K ground node clusters. The drones always fly at the same altitude H, and the ground nodes are on the horizontal plane and fixed in position. When a ground node has a data transmission task, in each time slot, the data transmission task generated by each ground node adopts any of the following communication modes: no communication, direct communication between ground nodes, and drone-assisted communication. In addition, the data transmission task generated by a single ground node in a single time slot can be processed by only one drone. The communication channel between the UAV and the ground node is described by the logarithmic range loss model, which includes a constant attenuation factor in small-scale fading and obstacle occlusion; the information rate of the nth,n∈{1,2,...,N}th ground node at a constant position on the ground is given by: Among them, the transmission power P, noise power N and path loss γ of the nth ground node are n ; The information rate when the UAV forwards the task queue is where γ m represents the path loss when the UAV forwards the task queue; when the path loss exponent in vacuum is set to α, the path loss of node n is Here, small-scale fading is modeled as a Rayleigh distributed random variable X with a scaling factor σ = 1 Rayleigh The attenuation due to the obstruction is modeled by a discrete factor β∈{1,0.01}, where β = 0.01 if there is a building obstruction and β = 1 otherwise. The path loss of the signal transmission between ground nodes is calculated as follows: Among them, P r Indicates the signal receiving power, P t Indicates the signal transmission power, G r Indicates the gain of the receiving antenna, G t represents the gain of the transmitting antenna, L represents the system loss, d represents the distance between the receiving end and the transmitting end, λ represents the wavelength of the signal, and γ represents the path loss exponent; Assume that all ground nodes use the frequency band F = {f1,f2,...,f i },i∈{1,2,...,I}, there are I frequency points to choose from; suppose there is a jammer in the scene that transmits interference signals at different frequencies in each time slot, interfering with the data transmission between ground nodes, and the interference signal frequency of the jammer is repeated at the time slot period F j ={f1,f2,...,f j }, j∈{1,2,...,J}; When a ground node selects a frequency point for communication in each time slot, if the selected frequency point is the same as the frequency point of the interference signal sent by the jammer in the current time slot, it is considered that the current ground node communication is interfered with, and if multiple ground nodes select the same frequency point in the same time slot, it is considered that there is communication interference between these ground nodes; when interference and mutual interference occur, the transmission rate is calculated by the following formula Among them, G j is the channel gain from the interference signal to the interfered node, P j is the interference signal power, G i is the channel gain of the mutual interference signal, A is the set of mutual interference numbers of ground nodes, γ n is the path loss of node n, P is the transmission power of the nth ground node, and N is the noise power; At the beginning of each time slot, the amount of data transmission randomly generated by each ground node is Δt. Each ground node has a data accumulation queue Δn, and each UAV has a data accumulation queue Δm. It is assumed that half of each time slot τ is used for transmission and half for reception. In addition, each ground node can only transmit data to one ground node per time slot, and the ground node can only receive data transmitted by the corresponding ground node. At the beginning of the time slot, since the amount of data randomly generated by the ground node is Δt, the cumulative amount of task data to be transmitted by the ground node is Δ n =Δ n +Δ t , when the ground node chooses to use drones to assist in communication, if the transmission can be completed, that is, Δ t ≤R n At (t)·τ / 2, the amount of data received by the UAV and forwarded, the current amount of data received by the uplink UAV queue is: Δ m =Δ m +Δ t , the ground node queue minus the current data volume: Δ n =Δ n -Δ t ; Current data volume forwarded in the downlink UAV queue: Δ m =Δ m -Δ t ; If the transmission cannot be completed, Δ t >R n At (t)·τ / 2, the maximum amount of data that the UAV can currently receive in the uplink is: Δ m =Δ m +R n (t)·τ / 2, the maximum amount of data that can be sent by the ground node queue minus: Δ n =Δ n -R n (t)·τ / 2; the maximum amount of data that can be forwarded by the UAV in the downlink is: Δ m =Δ m -R m (t)·τ / 2, where R n (t) and R m (t) represents the information rate of the ground node and the UAV respectively; the unsent data is sent in the next time slot; when the ground node chooses direct communication, if the transmission can be completed, that is, Δ t ≤R n (t)·τ / 2, the ground node queue minus the current data volume: Δ n =Δ n -Δ t , if the transmission cannot be completed, that is, Δ t >R n At (t)·τ / 2, the maximum amount of data that can be sent by the ground node queue is subtracted: Δ n =Δ n -R n (t)·τ / 2; the unsent data is sent in the next time slot.
3. The hierarchical decision-making method for air-ground collaborative communication based on environmental cognition according to claim 1 is characterized in that: The UAV trajectory planning layer is an intermediate layer that inherits the top-level ground node cluster and UAV associated cluster information a k,t , the ground node communication decision layer receives information from the top and middle layers as the bottom layer k,t and a m,t , for the ground node agent n,a n,t Determine the communication decisions of ground nodes; In the decision layer of ground node cluster and UAV association, state s k,t For s k,t =[p1,p2,…,p M ,a k,t-1 ], where p1, p2, …, p M Represents the positions of M drones, a k,t-1 represents the action decision of the kth ground node cluster in the time slot; the reward r obtained by ground node cluster k k,t+Δ Defined as the accumulation of queues between all ground nodes and drones in the cluster In the UAV trajectory planning layer, the state s m,t For s m,t =[p m ,q m ,a k,t-1 ], where p m and q m Represents the location of UAV m and the accumulated task queue information, a k,t-1 represents the action decision of the kth ground node cluster in the time slot; the reward r obtained by drone m m,t+Δ It is defined as the queue accumulation between all ground nodes in the cluster associated with the UAV In the communication decision layer of the ground node, the state s k,n,t For s k,n,t =[p m ,p n ,a n,t-1 ,f n ,q n ], where a n,t-1 represents the action decision of the time slot on the ground node n, p n 、f n ,q n Respectively represent the location of ground node n, communication frequency, and accumulated task queue information; The reward obtained by ground node n is defined as the accumulated task queue amount r of ground node n k,n,t =q n .
4. The hierarchical decision-making method for air-ground collaborative communication based on environmental cognition according to claim 3 is characterized in that: The process of optimizing the communication decision of multiple nodes in an air-ground collaborative network using a hierarchical decision algorithm based on an actor-critic framework includes the following steps: Define the cluster association strategy network parameters θ of ground node cluster k k and cluster association value network parameter w k , the association decision between the ground node cluster and the UAV cluster in time slot t is expressed as a k,t =π(·|s k,t θ k ), where in state s k,t According to the cluster association strategy network parameter θ k The output result is selected with a certain probability k,t ,satisfy The decision to associate the ground node cluster with the drone cluster ends after Δ time slots, based on the accumulated reward r k,t+Δ And the output of the cluster association value network V(s k,t ;w k ), calculate the TD error δ k,t =r k,t+Δ +ηV(s k,t+Δ ;w k )-V(s k,t ;w k ), where η is the loss coefficient; the cluster association value network parameters are updated with the goal of minimizing the TD mean square error, and the loss function is defined as The cluster association value network along the loss function L(w k ) The update formula of the gradient direction update parameter is expressed as Where α is the weight factor, the cluster association strategy function network is guided by the TD error along the loss function J(θ k ) updates the cluster association network parameters in the gradient direction, where the loss function is defined as J(θ k )=logπ(s k,t θ k )·δ k,t , the update formula of the cluster association strategy network parameters in each Δ time slot is expressed as Define the trajectory planning strategy network parameters θ of UAV m m and trajectory planning value network parameter w m , the trajectory planning decision in time slot t is represented as a m,t =π(·|s m,t θ m ); When the trajectory planning decision ends after Δ time slots, the cumulative reward r obtained by the drone m,t+Δ And the output V(s) of the trajectory planning value network m,t ;w m ), calculate the TD error δ of UAV m m,t =r m,t+Δ +ηV(s m,t+Δ ;w m )-V(s m,t ;w m ); define the communication decision strategy network parameters θ of the ground node n in each cluster k,n and communication decision value network parameter w k,n , the communication decision action of the ground node is represented as a k,n,t =π(·|s k,n,t θ k,n ); The communication decision action of the ground node is made in each time slot, and the reward r is obtained at the end of the data transmission action. k,n,t , TD error is expressed as δ k,n,t =r k,n,t +ηV(s k,n,t+1 ;w k,n )-V(s k,n,t ;w k,n ).
5. The hierarchical decision-making method for air-ground collaborative communication based on environmental cognition according to claim 1 is characterized in that: Step 2 includes the following sub-steps: Step 2.1: Abstract all nodes into a graph model representation G = (V, E), where V represents the node set and E represents the edge set; Step 2.2, construct the adjacency matrix A and the feature matrix X; The dimension of the adjacency matrix A is N×N, which is used to reflect the communication relationship between nodes. If there is a communication relationship between node i and node j, i, j∈{1,2,…,N}, then A ij =1, otherwise A ij =0, the dimension of the feature matrix X is N×F, where F represents the number of classes of node features; Step 2.3: Build a graph convolutional model. The number of neurons in the input layer is equal to the number of columns in the feature matrix X. The number of neurons in the output layer is determined by the number of clusters in which the nodes are clustered. The Glorot method is used to initialize the uniformly distributed weight matrix W. Step 2.4: pre-process the adjacency matrix A and feature matrix X and input them into the graph convolutional network, and obtain the prediction matrix Y through forward propagation. in and Represent the preprocessed adjacency matrix A and feature matrix X, W respectively (1) and W (2) Respectively represent the weight matrices of the first and second layers in the graph convolutional network, and ReLU and softmax represent two selected activation functions; Step 2.5, calculate the cross entropy loss function between the prediction matrix and the label matrix output by the graph convolutional network to update the weight matrix parameters. Where N is the number of nodes, K is the number of clusters in which the nodes are clustered, and Y ij and P ij Represent the prediction matrix and label matrix respectively; Step 2.6: Repeat steps 2.4-2.5 until the algorithm converges. If the prediction accuracy of the graph convolutional network meets the requirements, apply the pre-trained weight matrix W to the hierarchical decision-making algorithm based on the actor-critic framework to extract multiple environmental features to understand the complex urban environment.
6. The hierarchical decision-making method for air-ground collaborative communication based on environmental cognition according to claim 5 is characterized in that: In the decision layer for associating ground node clusters with UAVs, the adjacency matrix A reflects the queue transmission relationship between the ground node clusters and UAVs, and the feature matrix X includes the positions of all ground nodes and UAVs and the cumulative queue volume; In the ground node communication decision layer, the adjacency matrix A reflects the queue transmission relationship between ground nodes, and the feature matrix X includes the location of the ground nodes and whether frequency interference occurs between the ground nodes.
7. The hierarchical decision-making method for air-ground collaborative communication based on environmental cognition according to claim 6 is characterized in that: In step 2.4, the calculation formula for the preprocessed adjacency matrix A is Where D represents the degree matrix of the adjacency matrix A, and I represents the identity matrix; The calculation formula of the preprocessing feature matrix X is in and δ X Represent the mean and variance of the feature matrix X respectively; In the graph convolutional network, the forward propagation from hidden layer l to hidden layer l+1 is expressed as Where σ represents the ReLU activation function, H l and W l denote the output matrix and weight matrix of hidden layer l respectively.