A distributed intelligent relay dynamic deployment and user scheduling method and system

By employing a distributed intelligent relay dynamic deployment and user scheduling method, and leveraging the collaborative communication between UAVs and ground sensor nodes, combined with federated learning and reinforcement learning, the low latency and privacy protection issues of IoT systems are addressed. This achieves efficient computing resource scheduling and data transmission, adapts to complex environmental changes, and enhances the system's adaptability and stability.

CN121077548BActive Publication Date: 2026-02-06WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511620740.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-06
Estimated Expiration
2045-11-07

AI Technical Summary

Technical Problem

IoT systems face challenges such as low-latency communication requirements, limited computing power, and data privacy protection issues, especially in collaborative detection involving multiple listeners where there is a lack of effective resource management and trajectory design methods.

Method used

A distributed intelligent relay dynamic deployment and user scheduling method is adopted. Through collaborative communication between multiple UAVs and ground sensor nodes, combined with federated learning and reinforcement learning algorithms, the three-dimensional deployment location of UAVs and user scheduling strategies are optimized to minimize system latency and protect data privacy.

Benefits of technology

It achieves low-latency and high-reliability data transmission in dynamic environments, improves the system's computing efficiency and privacy protection capabilities, adapts to complex environmental changes, avoids flight conflicts, and enhances the system's adaptability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121077548B_ABST
    Figure CN121077548B_ABST
Patent Text Reader

Abstract

The application provides a distributed intelligent relay dynamic deployment and user scheduling method and system, which comprises the following steps: based on a multi-unmanned aerial vehicle and multi-ground sensor node communication scene, the ground sensor node senses environment data and generates a calculation task request; the unmanned aerial vehicle receives the calculation task request, takes minimizing the average time delay as a target, and jointly solves the optimized unmanned aerial vehicle three-dimensional deployment position and user scheduling strategy through a double-layer closed loop mechanism; based on the optimized unmanned aerial vehicle three-dimensional deployment position and user scheduling strategy, data collection, calculation decision and result feedback are performed; the above steps are repeated until a maximum time frame is reached, and the final unmanned aerial vehicle three-dimensional deployment position and scheduling strategy are output. According to the application, the ground sensing device senses the surrounding environment and initiates a calculation request, the unmanned aerial vehicle collects and processes related data according to the request. After the calculation is completed, the unmanned aerial vehicle feeds back the decision result to the ground device, so that efficient and low-latency data processing and transmission are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of unmanned aerial vehicle communication, and particularly relates to a distributed intelligent relay dynamic deployment and user scheduling method and system. BACKGROUND

[0002] In recent years, with the rapid development of Internet of Things technology, a large number of sensor nodes are widely deployed in key areas to monitor the environmental state in real time. However, in this process, the Internet of Things application faces many challenges. First, many applications have very high requirements for the timeliness of data transmission, especially in industrial monitoring, disaster warning and other scenarios. Higher communication delay may have serious consequences, so low-latency transmission is an urgent need. Second, due to energy and cost constraints, ground sensor nodes usually have limited computing power and are difficult to independently complete complex data processing and decision-making tasks, and therefore need to rely on external resources for efficient computing. Finally, the data collected by Internet of Things devices often involves user privacy or critical infrastructure information, and the data is extremely vulnerable to leakage during transmission and processing, so data privacy protection has become an important problem to be solved. In summary, how to realize low-latency communication, enhance computing power support and strengthen data privacy protection is a key technical bottleneck that needs to be broken through in current Internet of Things systems.

[0003] To address the above problems, introducing unmanned aerial vehicles (UAV) as an aerial auxiliary platform has become a promising solution. Compared with traditional fixed infrastructure, UAVs have the advantages of low cost, flexible deployment, easy scheduling and management, and can dynamically adjust the position and service range according to the distribution of ground nodes. In addition, because UAVs are usually in a high-altitude flight state, the communication link between them and the ground nodes has a high line-of-sight (LOS) probability, which helps to reduce channel fading and interference, thereby realizing low-latency, high-reliability data transmission. Further, UAVs can integrate a certain amount of computing and storage resources, serving as mobile edge servers to provide computing offloading and data processing services for resource-constrained ground nodes, significantly alleviating the computing pressure of ground nodes.

[0004] In terms of data privacy, federated learning (FL) as a distributed machine learning paradigm can achieve collaborative training by exchanging local model parameters without uploading raw data, effectively reducing the risk of privacy leakage. Combined with the mobility and computing power of UAVs, the federated learning framework based on UAVs can further improve the comprehensive performance of Internet of Things systems in data privacy protection and intelligent decision-making. In summary, the cooperative communication, computing and privacy protection mechanism based on UAVs provides new technical support and research direction for Internet of Things applications that are sensitive to delay, limited in computing and focus on privacy protection.

[0005] Based on this, the present application provides a kind of distributed intelligent relay dynamic deployment and user scheduling method and system.The method is by optimizing the deployment position of unmanned aerial vehicle and the ground equipment served by each unmanned aerial vehicle jointly, aims at minimizing total delay.Specifically, ground sensing device first perceives and generates the computing task needing to be handled, then by the device with computing demand sends computing request to air platform.Unmanned aerial vehicle collects these computing requests according to scheduling strategy, and carries out data processing and decision calculation locally, finally returns the processing result to corresponding ground equipment.Through the above process, the computing efficiency and response speed of system can be effectively improved, meet the needs of delay-sensitive application in Internet of Things environment. SUMMARY

[0006] In view of the current lack of trajectory and resource management design methods for multi-listener cooperative detection in the integrated network of concealed sensing and communication, the present application provides a kind of distributed intelligent relay dynamic deployment and user scheduling method and system to improve the performance of concealed sensing and communication.

[0007] To solve the above technical problems, the present application provides the following technical solutions:

[0008] A kind of distributed intelligent relay dynamic deployment and user scheduling method, comprising the following steps:

[0009] Step S1: based on multi-unmanned aerial vehicle, multi-ground sensor node communication scene, ground sensor node perceives environmental data and generates computing task request;

[0010] Step S2: unmanned aerial vehicle receives the computing task request, with the goal of minimizing average delay, and solves the optimized three-dimensional deployment position of unmanned aerial vehicle and user scheduling strategy by double-layer closed-loop mechanism jointly;

[0011] Step S3: based on the optimized three-dimensional deployment position of unmanned aerial vehicle and user scheduling strategy, data collection, computing decision and result return are executed;

[0012] Step S4: repeat S1-S3 until the maximum time frame is reached, output the final three-dimensional deployment position of the UAV and the scheduling strategy.

[0013] Further, the step S1 comprises: a first ground sensor node perceiving environmental data in a time frame and generating a computing task request containing data size and priority .

[0014] Further, the average time delay in step S2 is:

[0015]

[0016] wherein, wherein , represents the priority of the task, the larger the value, the lower the priority, wherein 1 represents the first priority, 2 represents the second priority, represents the highest priority number; is the average time delay, represents the data transmission time of the node with priority to the UAV in a time frame, represents the time for the data of the node with priority to be calculated by the UAV in a time frame, represents the time for the UAV to transmit decision data to the node with priority in a time frame, represents the priority weight coefficient, represents the total number of UAVs, represents the total number of time frames, represents the total number of ground sensor nodes, represents that N ground sensor nodes are distributed in an area of Km, represents the total number of frames for the UAV to complete the transmission-computation-decision task; the sensor nodes are subject to a statistical continuous distribution function , respectively represent the numerical value of the axis, the numerical value of the axis and the numerical value of the axis. Further,

[0017] Further,

[0018] ​​

[0019]

[0020]

[0021] wherein, denotes the i-th ground sensor node and the priority is denotes the i-th ground sensor node and the priority is denotes the data transmission rate of the i-th ground sensor node in the time frame data size; denotes the i-th UAV denotes the i-th UAV denotes the data transmission rate of the i-th UAV in the time frame data size; denotes the amount of computing resources required by the i-th UAV to process per unit of data, denotes the computing capability of the i-th UAV . denotes the size of the decision data in the time frame .

[0022] Further, the double-layer closed-loop mechanism in the step S2 comprises:

[0023] outer position optimization: adopting FL-D3QN algorithm to dynamically adjust the three-dimensional position of the UAV wherein, denotes the horizontal coordinate of the i-th UAV in the time frame , denotes the vertical coordinate of the i-th UAV in the time frame , denotes the height of the i-th UAV in the time frame .

[0024] inner user scheduling: based on the updated UAV position, solving the user-UAV association mapping closed-form solution by using optimal transmission theory to complete the communication resource allocation and scheduling of each UAV.

[0025] Further, the outer position optimization adopts Markov decision model, comprising:

[0026] state space ;

[0027] action space ;

[0028] reward function

[0029] wherein, denotes the i-th UAV ​, unmanned aerial vehicle ,... unmanned aerial vehicle k ,... unmanned aerial vehicle in a time frame deployable three-dimensional position, respectively represent the flight direction of each unmanned aerial vehicle during the three-dimensional deployment process; is a negative constant, representing the penalty when the unmanned aerial vehicle is out of the boundary or the unmanned aerial vehicle collides, and respectively represent the average task completion time when the unmanned aerial vehicle is in state and state .

[0030] Further, the inner-layer user scheduling user-unmanned aerial vehicle association mapping closed-form solution is:

[0031]

[0032] That is, the user selects the unmanned aerial vehicle that communicates with it to make the total delay time shortest.

[0033] Further, the execution flow of the FL-D3QN algorithm is:

[0034] S21 The aggregation end unmanned aerial vehicle sends its model parameters to each unmanned aerial vehicle, that is, wherein is the model parameter of the unmanned aerial vehicle ;

[0035] S22 The optimal user scheduling of each unmanned aerial vehicle at a fixed position is analyzed and obtained;

[0036] S23 The results of state transition of each unmanned aerial vehicle are calculated respectively, and the results of state transition are saved to its experience pool;

[0037] S24 Each unmanned aerial vehicle trains the Dueling DDQN network based on the local experience pool;

[0038] S25 Periodically aggregate local parameters: wherein represents the model parameter after aggregation processing, is the model parameter sent by each unmanned aerial vehicle to the aggregation end for aggregation processing, is the size of the model parameter of all unmanned aerial vehicles, represents the size of the model parameter of the unmanned aerial vehicle ;

[0039] S26 Determine whether the maximum set exploration step number is reached, if yes, execute S27, otherwise repeat steps S21-S25;

[0040] S27 judges whether the maximum set time frame is reached, if not, the next time frame of the unmanned aerial vehicle three-dimensional deployment position and user scheduling is continuously optimized;

[0041] S28 outputs the optimization result.

[0042] Further, the Dueling DDQN network loss function is:

[0043]

[0044] wherein, represents the unmanned aerial vehicle in the time slot the reward value in the first transition, is the main network parameter, is the discount factor, represents the unmanned aerial vehicle in the time slot the state in the first transition, is the target network parameter, represents the action selected from the action space in the process of state transition, which makes the value maximum. .

[0045] On the other hand, the application provides a distributed intelligent relay dynamic deployment and user scheduling system, comprising:

[0046] A task request generation module: which is used for generating a computing task request based on a multi-unmanned aerial vehicle and multi-ground sensor node communication scene, and the ground sensor node perceives environmental data;

[0047] A joint solution module: which is used for receiving the computing task request by the unmanned aerial vehicle, and jointly solving the optimized unmanned aerial vehicle three-dimensional deployment position and user scheduling strategy through a double-layer closed-loop mechanism with the goal of minimizing the average time delay;

[0048] A result feedback module: which is used for executing data collection, computing decision and result feedback based on the optimized unmanned aerial vehicle three-dimensional deployment position and user scheduling strategy;

[0049] A final result output module: which is used for repeating the above steps until the maximum time frame is reached, and outputting the final unmanned aerial vehicle three-dimensional deployment position and scheduling strategy.

[0050] Compared with the prior art, the application has the following beneficial effects:

[0051] ​(1) Dynamic environment intelligent relay three-dimensional cooperative deployment: In view of the problem of dynamic change of data between different frames of ground sensor nodes and unstable communication environment, the three-dimensional relay cooperative deployment and user scheduling of multiple unmanned aerial vehicles are intelligently optimized by real-time sensing of the environment state and combining a federal reinforcement learning method. Each unmanned aerial vehicle adaptively adjusts the three-dimensional deployment position and user scheduling according to the environmental change, so as to minimize the system average delay. The mechanism can adapt to complex dynamic environment, and improves the adaptability, stability and overall service performance of the unmanned aerial vehicle network.

[0052] (2) Double-layer intelligent relay deployment and user scheduling method: The application provides a double-layer intelligent relay deployment and user scheduling method, wherein the outer layer adopts a FL-D3QN algorithm combined based on federal learning and a Dueling DDQN algorithm to optimize the three-dimensional deployment position of multiple unmanned aerial vehicles; and the inner layer completes the communication resource allocation and scheduling of each unmanned aerial vehicle based on a user scheduling closed-form solution derived based on the optimal transmission theory. The method can accelerate the convergence speed of the algorithm and improve the overall performance of the system, while effectively protecting data privacy.

[0053] (3) Heterogeneous communication network and safety control mechanism: In the process of multiple unmanned aerial vehicles cooperatively performing data collection, calculation and decision-making tasks, the application fully considers the heterogeneous characteristics of the system, including the differences in the computing power of unmanned aerial vehicles, the differences in the data size of different frames, and the differences in the priority of ground transmission nodes. At the same time, in view of the collision risk and boundary crossing problem that may occur in the cooperative flight of multiple unmanned aerial vehicles, a corresponding safety control mechanism is designed to effectively avoid flight conflicts and boundary violations. The design makes the multiple unmanned aerial vehicles more suitable for actual application requirements in the process of cooperative operation, and significantly improves the feasibility and stability of the system. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0055] Figure 1 The method flowchart for implementing the application.

[0056] Figure 2 The unmanned aerial vehicle three-dimensional position deployment and user scheduling optimization flowchart for implementing the application.

[0057] Figure 3 The system architecture schematic diagram for implementing the application. DETAILED DESCRIPTION

[0058] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the protection scope of the present application.

[0059] Embodiment 1

[0060] The present application will be further described below in conjunction with the drawings.

[0061] The present application will be further described below in conjunction with the drawings. Figure 1 The method for dynamically deploying an intelligent relay and scheduling a user based on federated reinforcement learning is introduced, and comprises the following steps:

[0062] Step S1: Based on a multi-unmanned aerial vehicle (UAV) and multi-ground sensor node communication scenario, a ground sensor node perceives environmental data and generates a computing task request;

[0063] Step S2: An UAV receives the computing task request, aims to minimize the average time delay, and jointly solves an optimized three-dimensional deployment position of the UAV and a user scheduling strategy through a double-layer closed-loop mechanism;

[0064] Step S3: Based on the optimized three-dimensional deployment position of the UAV and the user scheduling strategy, data collection, computing decision and result return are performed;

[0065] Step S4: Steps S1-S3 are repeated until a maximum time frame is reached, and the final three-dimensional deployment position of the UAV and the scheduling strategy are output.

[0066] In the embodiment, the multi-UAV and multi-ground sensor node communication scenario comprises K UAVs, N ground sensor nodes, The N ground sensor nodes are distributed in an area of K m2, and it is assumed that the total number of frames for the UAV to complete a transmission-computing-decision task is T. In the t th time frame, it is assumed that the position of the k th UAV is , where represents the horizontal coordinate of the k th UAV in the t th time frame, and represents the vertical coordinate of the k th UAV in the t th time frame. ​​​​​​​​​The height. Assume the sensor nodes follow a statistical continuous distribution function. It can be any continuous distribution, such as a uniform distribution. Assuming different sensor nodes have different task priorities, the first... There are 1 ground sensor node with a priority of 1 In time frame Data size is The data size and priority of each ground sensor node change with the time frame. Their location is represented as... ,in Represents sensor nodes x-coordinate Represents sensor nodes The vertical axis represents the time of the entire operating system. The time of the entire operating system is divided into three parts: the time for the UAV to collect data after receiving a computing request from the ground sensor node, the computing time, and the time to send the computing decision results to the ground user.

[0067] The average task completion time can be expressed as

[0068]

[0069] in The comma (,) indicates the task priority. A higher number indicates a lower priority, where 1 represents the first priority and 2 represents the second priority. Represents the highest priority number; where Indicates time frame Priority is nodes With drones Data transmission time, Indicates time frame Priority is nodes Data by drones Calculation time, Indicates drone In time frame Transmit decision data to the priority level. nodes The time. This represents the priority weight coefficient. Indicates the total number of drones. N Indicates the total number of time frames. This indicates the total number of ground sensor nodes. express The ground sensing nodes are distributed in Within a Km area denotes the total number of frames that the UAVs complete the transmission-computation-decision task; the sensor nodes obey a statistical continuous distribution function , denote the numerical value of the axis, the numerical value of the axis, respectively.

[0070] wherein , wherein denotes the UAV and the node data transmission rate in the time frame , which is related to the UAV position and user position; , wherein denotes the UAV amount of computing resources required to process each unit of data, denotes the computing power of the UAV ; wherein , wherein denotes the size of the decision data in the time frame . In the process of communication transmission, the UAV adopts the communication strategy of FDMA (frequency division multiple access), that is, a single UAV can communicate with multiple nodes in a time frame.

[0071] A. In the sensor data collection and task request module, the ground sensor device senses and collects data near it, and decides whether to issue a transmission-computation-decision task request in the time frame according to the integrity of the collected data.

[0072] B. After the ground sensor device senses the data near it and issues a transmission-computation-decision task request, the UAV flies to a fixed deployment position and begins to perform the data collection, computation, and decision transmission tasks. In this process, the three-dimensional deployment positions of multiple UAVs and the user scheduling of each UAV are jointly optimized. In a single time frame , the three-dimensional deployment positions of multiple UAVs are optimized based on the FL-D3QN (Federated Learning dueling Double Deep Q-Network) algorithm, and then the optimal user scheduling of each UAV is obtained based on the optimal transmission theory. Compared with the traditional DDQN (Double Deep Q-Network) algorithm, the proposed method can protect data privacy while accelerating the convergence of the algorithm, thereby reducing the average task delay of the system.

[0073] Specifically, in a single time frame In the process of optimizing the three-dimensional deployment positions of multiple UAVs in the time frame based on the FL-D3QN algorithm, the established time minimization problem is first converted into a Markov decision problem, and then the deployment positions of each UAV are solved based on the proposed FL-D3QN algorithm. One complete Markov decision process can be composed of four parts, namely , wherein is a state space, is an action space, is a state transition probability, is a reward function. The state space when optimizing the deployment of UAVs is , wherein respectively represent the three-dimensional positions of UAV , UAV , …, UAV k , …, UAV that can be deployed in the time frame .

[0074] The action space is , wherein respectively represent the flight directions of each UAV in the process of three-dimensional deployment. In this embodiment, 6 flight directions are taken, wherein , , , , , respectively represent the forward flight, backward flight, left flight, right flight, upward flight and downward flight of the UAV. The state transition probability represents the probability that the UAV in state transitions to the next state according to the selected action . The reward function can be represented as

[0075]

[0076] is a negative constant, representing the punishment when the UAV is out of the boundary or the UAV collides, and respectively represent the average task completion time when the UAV is in state and state . It can be seen that when the UAV transitions from state to state .If the task completion time decreases, a positive reward is given; otherwise, a negative reward is given. Then, after establishing the Markov state, the 3D deployment positions of each UAV in that time frame are optimized based on the FL-D3QN algorithm. The specific optimization process is as follows:

[0077] like Figure 2 As shown, the optimization of multi-UAV 3D location deployment and user scheduling allocation based on FL-D3QN and the optimal transmission mechanism is as follows:

[0078] The S21 aggregation terminal drone uses its model parameters Distribute to each drone, that is ,in For drones Model parameters;

[0079] S22 then analyzes and calculates the optimal user scheduling for each UAV at a fixed location. Since the UAV locations are fixed, only the association between nodes and UAVs needs to be optimized. The optimal user association strategy is derived below based on task time minimization. This strategy can be understood as finding a mapping relationship between users and UAVs within the framework of optimal transmission theory, which minimizes the average task latency. The specific proof is as follows:

[0080] First, we prove that the optimal solution exists. As mentioned earlier, the average delay time of the system is as follows:

[0081]

[0082] because and User location A continuous function, and calculate the time delay. It depends only on the data size and the drone's computing power, therefore the total delay function Since it is a continuous function, and a continuous function necessarily satisfies the lower semi-continuity, the total delay function is... It is also a lower semi-continuous function. For continuously distributed users and discrete UAV service nodes, the lower semi-continuous function guarantees the existence of the optimal transport mapping (user association strategy). Therefore, the optimal user association exists.

[0083] The user association function is defined below as follows: ,when The time indicates drone With priority nodes Related, otherwise ,and Since the user association problem is essentially a mixed-integer nonlinear programming problem with two variables, it is nonconvex and cannot be directly differentiated. Therefore, a relaxation method is introduced to transform it into a probabilistic variable problem. At this time, the user allocation problem is converted into a convex problem with respect to and satisfies the Slater condition. Based on the Lagrange multiplier method, first introduce the Lagrange multiplier , the total delay function is expressed as:

[0084]

[0085] wherein By taking the first-order derivative of and setting it to 0, we get:

[0086]

[0087] Therefore, for the UAV , the optimal association needs to satisfy:

[0088]

[0089] Therefore, at the optimal association, the user selects the solution of the UAV as:

[0090]

[0091] wherein at the intersection of the optimal partition boundary and , the Lagrange multiplier satisfies:

[0092]

[0093] wherein represents the data transmission time of the node with priority and the UAV at the time frame , represents the time for the data of the node with priority to be calculated by the UAV at the time frame , represents the time for the UAV to transmit the decision data to the node with priority at the time frame .

[0094] Accordingly, each user can be allocated to each UAV.

[0095] S23 in the time frame , assuming that the total training step number is , then the UAV The number of frames and parts in this time frame is Its state at that time is Its action space Choose one action Then transition to the next state. In the middle, the communication-computation-decision task is completed at this position, and then the average latency between the two states is calculated and a reward is obtained according to the reward function calculation method. Then each drone will transmit the state transition results. Save it to its experience pool.

[0096] Each S24 drone is randomly selected from the experience pool. The system uses sample data and then trains its neural network using gradient descent. The main goal during training is to reduce the loss function to obtain a larger reward. The loss function of DDQN is... Defined as:

[0097]

[0098] in, Indicates drone In the time slot No. The reward value at the time of the next transfer Main network parameters, As a discount factor, Indicates drone In the time slot No. The state at the time of the next transition. For the target network parameters, This indicates that during the state transition process, performance is selected from the action space to make... The action with the highest value Then use the target network To evaluate this action. When the Dueling layer is added to the network, it does not change the basic architecture of the loss function, but only... The function's structure has been modified within the Dueling DDQN structure. The function is represented as follows:

[0099]

[0100] in This represents the state value function, used to evaluate the overall quality of the current state regardless of the action chosen. It does not focus on the choice of action, but only on the value of the state itself. represents an action advantage function, which is used to evaluate the goodness of the selected action compared to the general action in the current state. represents the absolute value of the estimated state value, which is a constant.

[0101] S25 After each unmanned aerial vehicle trains the data locally, each unmanned aerial vehicle uploads its model parameters every time interval to the aggregation end, and the aggregation end aggregates the model parameters and distributes the parameters to each unmanned aerial vehicle, wherein the aggregation is as follows:

[0102]

[0103] wherein represents the model parameters after the aggregation processing, is the model parameters sent by each unmanned aerial vehicle to the aggregation end for aggregation processing, is the size of the model parameters of all unmanned aerial vehicles, represents the size of the model parameters of the unmanned aerial vehicle .

[0104] S26 It is judged whether the maximum set exploration step number is reached, if yes, S27 is executed, otherwise steps S21-S25 are repeated.

[0105] S27 It is judged whether the maximum set time frame is reached, if not, the unmanned aerial vehicle three-dimensional deployment position and user scheduling of the next time frame are continued to be optimized;

[0106] S28 The optimization result is output.

[0107] (2) It is judged whether the maximum training episode is reached, if yes, the three-dimensional deployment position of each unmanned aerial vehicle and the user scheduling of each time frame are output, if not, step (1) is repeated.

[0108] The optimized three-dimensional position deployment result of each unmanned aerial vehicle and the user scheduling result are output.

[0109] Embodiment 2

[0110] The embodiment provides a distributed intelligent relay dynamic deployment and user scheduling system, as shown in Figure 3 , which comprises:

[0111] A task request generation module: which is used for generating a computing task request based on a multi-unmanned aerial vehicle and multi-ground sensor node communication scene, and the ground sensor node perceives environment data and generates a computing task request;

[0112] A joint solution module: which is used for the unmanned aerial vehicle to receive the computing task request, to minimize the average delay, and to jointly solve the optimized unmanned aerial vehicle three-dimensional deployment position and user scheduling strategy through a double-layer closed-loop mechanism;

[0113] A result returning module is configured to perform data collection, calculation decision and result returning based on the optimized three-dimensional deployment position of the UAV and the user scheduling strategy execution data;

[0114] A final result output module is configured to repeat the above steps until a maximum time frame is reached, and output the final three-dimensional deployment position of the UAV and the scheduling strategy.

[0115] In this embodiment, the data sensing, task generation, uplink communication and downlink communication are included, wherein the data sensing includes that the ground sensing device performs real-time data sensing and collection on the surrounding environment or the monitored object through its own sensor device (such as a temperature sensor, a humidity sensor, a camera, an air quality detector, etc.). The task generation includes that the ground sensing device generates a calculation or processing task after completing the data sensing, and initiates a request to the UAV or the edge server, requiring it to perform data transmission, calculation and transmission of the calculation result back to the sensing device. The uplink communication includes that the ground sensing device transmits the sensed data to the UAV, so as to facilitate the calculation and processing of the UAV. The downlink communication includes that the UAV issues the processed result to the ground sensing device, so as to realize information feedback and application.

[0116] It should be understood that parts not described in detail in the specification are all prior art.

[0117] It should be understood that the above description of the preferred embodiments is more detailed, and therefore should not be considered as limiting the scope of patent protection of the present application. It is not necessary and impossible to enumerate all the embodiments. Those skilled in the art can make substitutions or modifications without departing from the scope of the present application, which shall fall within the scope of protection of the present application. The scope of protection of the present application shall be subject to the appended claims.

Claims

1. A method for dynamic deployment and user scheduling of distributed intelligent relays, characterized in that, Includes the following steps: Step S1: Based on the communication scenario of multiple UAVs and multiple ground sensor nodes, the ground sensor nodes perceive environmental data and generate computing task requests; Step S2: The UAV receives the computation task request, aims to minimize the average latency, and jointly solves the optimized UAV 3D deployment location and user scheduling strategy through a two-layer closed-loop mechanism; the average latency is: Among them, The comma (,) indicates the task priority. A higher number indicates a lower priority, where 1 represents the first priority and 2 represents the second priority. Indicates the highest priority number; For average delay, Indicates time frame Priority is nodes With drones Data transmission time, Indicates time frame Priority is nodes Data by drones Calculation time, Indicates drone In time frame Transmit decision data to the priority level. The time of the node; This represents the priority weight coefficient. Indicates the total number of drones. N Indicates the total number of time frames. This indicates the total number of ground sensor nodes. express The ground sensing nodes are distributed in Within a Km area This represents the total number of frames the UAV completes for the transmission-computation-decision task; sensor nodes follow a statistical continuous distribution function. , They represent The value of the axis, The values ​​of the axes and The value of the axis; Step S3: Based on the optimized 3D deployment location of the UAV and the user scheduling strategy, perform data collection, calculation decision-making, and result feedback; Step S4: Repeat S1-S3 until the maximum time frame is reached, and output the final three-dimensional deployment location and scheduling strategy of the UAV.

2. The method for dynamic deployment and user scheduling of distributed intelligent relays according to claim 1, characterized in that, Step S1 includes: the first Each ground sensor node in time frame Sensing environmental data, generating data including size Priority The computational task request.

3. The method for dynamic deployment and user scheduling of distributed intelligent relays according to claim 1, characterized in that, in, Indicates the first There are 1 ground sensor node with a priority of 1 In time frame Data size; Indicates drone With nodes In time frame The data transmission rate; Indicates drone The amount of computing resources required to process each unit of data Indicates drone Computational power; Represents time frame The size of the decision data.

4. The method for dynamic deployment and user scheduling of distributed intelligent relays according to claim 3, characterized in that, The dual-layer closed-loop mechanism in step S2 includes: Outer layer position optimization: The FL-D3QN algorithm is used to dynamically adjust the three-dimensional position of the UAV. ,in: Indicates drone In time frame x-coordinate Indicates drone In time frame The ordinate, Indicates drone In time frame Height; Inner User Scheduling: Based on the updated UAV locations, solve the closed-form solution of the user-UAV association mapping using optimal transport theory. This completes the allocation and scheduling of communication resources for each UAV.

5. The method for dynamic deployment and user scheduling of distributed intelligent relays according to claim 4, characterized in that, The outer layer position optimization employs a Markov decision model, including: state space ; Action space ; reward function in, They represent drones drones ...drones k ...drones In time frame Deployable 3D location, These represent the flight directions of each drone during the three-dimensional deployment process; This is a negative constant, representing the penalty when a drone goes out of bounds or collides with another drone. and These indicate the drone's current status. and state The average task completion time.

6. The method for dynamic deployment and user scheduling of distributed intelligent relays according to claim 4, characterized in that, The inner layer user scheduling user-UAV association mapping closed solution for: i.e., user Choose the drone that minimizes the total latency to communicate with.

7. The method for dynamic deployment and user scheduling of distributed intelligent relays according to claim 4, characterized in that, The execution flow of the FL-D3QN algorithm is as follows: The S21 aggregation terminal drone uses its model parameters Distribute to each drone, that is ,in For drones Model parameters; S22 analyzes and calculates the optimal user scheduling for each UAV when it is in a fixed position; S23 calculates the state transition results for each UAV and saves the state transition results to its experience pool; Each of the S24 drones trained the Dueling DDQN network based on its local experience pool; S25 Periodic Aggregation Local Parameters: ,in This represents the model parameters after aggregation. The model parameters sent by each drone to the aggregation terminal for aggregation processing. The number and size of model parameters for all drones. Indicates drone The size of the number of model parameters; S26 Determine whether the maximum set number of exploration steps has been reached. If yes, proceed to S27; otherwise, repeat steps S21-S25. S27 determines whether the maximum set time frame has been reached. If not, it continues to optimize the three-dimensional deployment position of the drone and user scheduling for the next time frame. S28 outputs the optimized results.

8. The method for dynamic deployment and user scheduling of distributed intelligent relays according to claim 1, characterized in that, Dueling DDQN network loss function for: in, Indicates drone In the time slot No. The reward value at the time of the next transfer Main network parameters, As a discount factor, Indicates drone In the time slot No. The state at the time of the next transition. For the target network parameters, This indicates that during the state transition process, performance is selected from the action space to make... The action with the highest value .

9. A distributed intelligent relay dynamic deployment and user scheduling system, characterized in that, include: Task request generation module: It is used in scenarios involving communication between multiple UAVs and multiple ground sensor nodes, where ground sensor nodes perceive environmental data and generate computation task requests; Joint Solver Module: It is used by the UAV to receive the computing task request, with the goal of minimizing the average latency, and jointly solves the optimized UAV 3D deployment position and user scheduling strategy through a two-layer closed-loop mechanism; Result feedback module: It is used to perform data collection, calculation and decision-making and result feedback based on the optimized three-dimensional deployment position of the UAV and the user scheduling strategy; Final result output module: It is used to repeat the above steps until the maximum time frame is reached, and output the final three-dimensional deployment position and scheduling strategy of the UAV; The distributed intelligent relay dynamic deployment and user scheduling system is used to execute the steps in the distributed intelligent relay dynamic deployment and user scheduling method according to any one of claims 1-8.